AI can now drive your accounting software. Three things to fix
GPT-6 Astra, released on 3 September 2026, operates software through the interface instead of through an API, which puts every system in your firm in reach whether or not it has an integration. The capability is real, it fails more than a quarter of the time on the standard benchmark, and the questions it raises about logins, audit trails and supervision matter more right now than the accuracy figure.
Trent McLaren · 4 September 2026 · 9 min read
In this article
- What computer use actually means in a practice
- The number that was not in the press release
- Problem one: the agent needs a login, not an API key
- Problem two: you cannot review what you cannot see
- Problem three: your AI policy is written about tools, not actors
- What to do this month
- Frequently asked questions
- Is GPT-6 Astra available to my firm right now?
- Does computer use mean I no longer need integrations?
- What happens if the agent makes a mistake in a client file?
- How is this different from the robotic process automation firms tried years ago?
- Should we tell clients that an AI agent touched their file?
Part of our AI in accounting coverage. See the full AI for accounting firms guide →
OpenAI released GPT-6 Astra on 3 September 2026. Most of the coverage went to the AGI talk. The part that changes your week is quieter: the model operates software the way a person does. It opens the browser, clicks the fields, types into them and moves on. No integration, no API key, no app store listing. If a human can do the job in a browser, the model can attempt the job in a browser.
Every piece of AI advice your firm has read for two years assumed the opposite. It assumed the AI needed a connector, and that the practical question was whether your stack exposed one. That question just stopped being the gate. The gate is now whether you are willing to hand a machine a seat at your software.
What computer use actually means in a practice
Strip the demos back and computer use is one capability: the model reads a screen, decides what to click, and clicks it. OpenAI's own examples include filling in forms, updating CRM records, working across spreadsheets and browser tabs, and drafting a federal tax return in a browser from a W-2.
Note the word draft. That is OpenAI's own framing, and it is the right one: the output is a starting position a person still has to stand behind.
For a firm, the jobs this points at are the ones that were never automatable because they lived in a user interface nobody would build an integration for. Re-keying a client's figures from a portal that has no export. Chasing a status across three systems that do not talk. Reconciling a list in one window against a list in another. This is the work that got quietly pushed to the least experienced person in the team, because it needed hands rather than judgement.
The number that was not in the press release
On OSWorld V2-Offline, the standard benchmark for models operating real computers, Astra scores 72.6%. Its predecessor scored 65.7%. It also finishes tasks in roughly 40 minutes where the older model took about 75.
Faster and better. Also wrong more than a quarter of the time.
That figure is the single most useful thing published about this release, and it deserves to sit at the front of any conversation your firm has about agents. A tool that fails 27% of the time is not unusable. A junior on their first week fails at plenty of tasks. The difference is that you know to check a junior's work, you know which tasks they are ready for, and they tell you when they are stuck.
Carry one more caveat. The eye-catching 99.9% on the ARC-AGI-3 reasoning benchmark came from a custom adapter built by OpenAI, and the same model scored 62.7% on the standard harness. Treat every launch number as a reason to run your own test, not as a result.
Problem one: the agent needs a login, not an API key
Here is the part almost nobody has thought through. An integration authenticates as an application, with scopes, a consent screen and a revocable token. An agent driving a browser authenticates as a person, using that person's username and password.
Which person? In practice, whoever is logged in on the machine it runs on.
That is a problem your terms of use may already have an opinion about. Xero's Terms of Use put the responsibility for login credentials on the account holder, and section 27 asks users to keep their credentials secure by "not letting any other person use them". Xero's separate Developer Platform terms are more direct about automation: apps must not use bots or browser extensions to undermine security controls or simulate user actions, and those updated terms applied to developers registered before 4 December 2025 from 2 March 2026.
Read plainly, neither clause was written with an AI agent in mind. The developer terms govern developers, not a practitioner running a model on their own desktop, and an agent is not obviously "another person". But nothing in either document points towards this being clearly permitted, and it would be optimistic to assume your software vendors are relaxed about a machine typing into their product under a human's credentials at superhuman speed. Xero has said it uses automated monitoring and may suspend access where it suspects a breach.
Check the terms for every system you would point an agent at, and check them again in six months, because this is the clause every vendor is about to rewrite. None of this is legal advice, and if the answer matters commercially, get it from someone who carries insurance for being wrong.
Problem two: you cannot review what you cannot see
When an integration writes to your ledger, there is a record. The API call happened, it is logged, the app is named in the audit trail, and you can revoke it.
When an agent clicks through the interface as Sarah from the client services team, the audit trail says Sarah did it. Every change looks like human work, because as far as the system is concerned it was.
This is the practical reason to slow down, and it is separate from any question about how good the model is. Even at 100% accuracy you would have a governance hole: no way to answer "who changed this" six months later, when the client asks. Firms that have already worked through how connected AI tools reach their data have at least a token to point at. Screen-driving agents give you nothing unless you build the record yourself.
The workable answer for now is boring. Run agents in a named account that exists for exactly that purpose, so the trail says what it is. Record the session. Keep the output in a review state that a human has to clear before it counts as done, the same way you would treat any other agent workflow you have set up.
Problem three: your AI policy is written about tools, not actors
Pull up your firm's AI policy. If you have one, it almost certainly covers what staff may paste into a chat window, which tools are approved, and what happens to client data. All still necessary. All aimed at a person using a tool.
An agent operating your software is not a tool in that sense. It takes actions with consequences under someone's identity. The policy questions change shape: which systems may an agent touch, under whose account, with what standing limits, and who signs off before its output leaves the firm.
If you are starting from nothing, the one-page policy shape is still the right starting point, with a new section for agents rather than a rewrite. For anyone doing tax agent work, the supervision question is not optional: your Code obligations follow the work, not the tool, and an agent operating under a practitioner's login is about as clear a supervision question as the profession has faced.
What to do this month
- Do not roll it out yet. Astra is live for organisations in OpenAI's Trusted Access and Daybreak programs, with wider availability following. Enterprise workspace administrators have to switch it on deliberately, so the default is off. Leave it off while you do the rest of this list.
- Write down the five jobs you would point it at. Not the impressive ones. The re-keying, the cross-system chasing, the list-against-list checking. If you cannot name five, you do not have a use case yet and that is a fine answer.
- Read the terms for each system on that list. Practice management, ledger, portal, payroll. Note which ones say something about automated access and which are silent.
- Decide the identity question before the tooling question. A named agent account with its own licence and its own scope is more work than reusing a staff login, and it is the difference between a reviewable process and an unauditable one.
- Test on last period's work. Give it a job you have already completed and compare. You will learn more from one week of that than from a quarter of reading about benchmarks, including this article.
The firms that get value out of this will treat it as a staffing decision rather than a software purchase. You are adding something that acts, at speed, under a name, and the whole question is what supervision it sits under. That has always been a practice management problem, and a system like FYI or AccountKit will tell you more about whether you are ready than any model card will.
For the wider picture on where this fits, our AI for accounting firms hub tracks the rest of it.
Frequently asked questions
Is GPT-6 Astra available to my firm right now?
Not to everyone. At launch on 3 September 2026 it went live for organisations in OpenAI's Trusted Access and Daybreak programs, with API, AWS and the Plus, Pro, Business and Enterprise plans following over the days after. On Enterprise it is off until an administrator enables it for the workspace, so check with whoever holds your admin account rather than assuming staff already have it.
Does computer use mean I no longer need integrations?
No, and treating it that way would be a mistake. An integration is faster, cheaper per task, logged, revocable and does not break when a vendor moves a button. Computer use is for the gaps: the systems that never had an integration and never will. Think of it as reaching work that was previously out of scope, not as replacing the connections you already rely on.
What happens if the agent makes a mistake in a client file?
The same thing that happens when a staff member does, with one difference: the audit trail will name whoever's login it used. Responsibility does not move to the vendor because a model did the typing. This is why the identity and review steps matter more than the accuracy percentage.
How is this different from the robotic process automation firms tried years ago?
Older automation followed a recorded script and broke the moment a screen changed. These models read the screen and decide what to do, so they survive interface changes and can handle a task they were not specifically configured for. The trade is predictability: a script fails the same way every time, while a model can fail in a new way each run, which is harder to test for.
Should we tell clients that an AI agent touched their file?
If it did work that affects their numbers, yes, and your engagement terms may already require it. The question of what to disclose and when is worth settling before the first agent runs rather than after a client asks, and it is the same question firms already worked through for consent around AI use on client data.
ai agentscomputer usegpt-6practice managementgovernance