How to assess a new AI model without losing a week to it
A major model now ships every few weeks and assessing each one properly is not sustainable without a fixed process. Six checks, run in about two hours, produce a decision rather than an opinion, and GPT-6 Astra is the worked example.
Trent McLaren · 4 September 2026 · 9 min read
In this article
- The six checks
- 1. Can you actually get it?
- 2. Find the real number, not the headline number
- 3. Check the knowledge cutoff before anything else touches tax
- 4. Work out what your actual usage comes to
- 5. Understand what it refuses, and how it refuses
- 6. Decide what you tell the team, and say it
- The two-hour test that beats any review
- What not to do
- Frequently asked questions
- How often should a firm review which AI model it uses?
- Why does a model's knowledge cutoff matter so much in accounting?
- Are benchmark scores worth paying attention to at all?
- Should we standardise on one AI model across the firm?
- What is the minimum a small firm should do when a new model launches?
Part of our AI in accounting coverage. See the full AI for accounting firms guide →
A major model ships roughly every few weeks now. Each one arrives with benchmark charts, a claim about a new era, and a queue of people in your firm asking whether it changes anything. Answering that properly every time is not sustainable, and ignoring it is how firms end up two years behind without noticing.
What works is a fixed set of checks you run each time, in about two hours, that produces a decision rather than an opinion. Below is the one we use, with GPT-6 Astra, released on 3 September 2026, as the worked example. The checks matter more than the example. The example will be out of date by the time you need this.
The six checks
| Check | The question | Astra, as at September 2026 |
|---|---|---|
| Access | Can we actually get it, and is it on? | Limited programs first, wider access following. Off by default on Enterprise until an admin enables it. |
| Real number | What does it score on the task we care about? | 72.6% on computer use. Some headline figures used a custom harness. |
| Cutoff | What does it not know? | Knowledge cutoff of 30 April 2026. |
| Cost | What does our real usage come to? | Listed at 10 dollars per million input tokens and 50 dollars per million output tokens. |
| Refusals | What will it stop doing, and how? | Restricted on advanced security work. Some API safety checks end a task outright. |
| People | What do we tell the team? | One message, this week, before someone asks. |
1. Can you actually get it?
Start here, because a surprising number of firm-wide conversations happen about a model nobody in the building can use yet. Astra went live for organisations in OpenAI's Trusted Access and Daybreak programs first, with the API, AWS and the consumer and business plans following over the days after. On Enterprise workspaces it is off until an administrator switches it on.
Two minutes with whoever holds your admin account tells you whether the rest of the exercise is theoretical. Do that before you read a single benchmark.
2. Find the real number, not the headline number
Every launch has one figure that describes what the model does for you and several that describe what it does for the press release. Learning to tell them apart is the most valuable skill in this whole process.
For Astra, the figure worth having is 72.6% on OSWorld V2-Offline, the standard benchmark for models operating real software. That is up from 65.7% for the previous model, and it means roughly one task in four still goes wrong.
The figures worth discounting are the ones with an asterisk. The widely quoted 99.9% on the ARC-AGI-3 reasoning benchmark was achieved using a custom adapter built by OpenAI; on the standard harness the same model scored 62.7%. Independent aggregation tells a flatter story too: Artificial Analysis put Astra at 61 on its intelligence index against 66 for Claude Fable 5.1, which is a very different picture from the launch presentation.
None of this makes the model bad. It makes the launch material a marketing document, which it always was. Take the number attached to the task you would actually give it, ignore the rest, and expect the truth to be less dramatic in both directions.
3. Check the knowledge cutoff before anything else touches tax
This is the check almost nobody runs, and in this profession it is the one most likely to hurt.
Astra's knowledge cutoff is 30 April 2026. Anything legislated, ruled on, indexed or announced since then is not in the model. It will still answer your question about it, fluently, because that is what these systems do.
For a firm, that means every rate, threshold and deadline that moved in the current year is a hazard unless the model is reading a document you gave it or searching a source you trust. This is not a flaw to be fixed in the next version. Every model has a cutoff and every model sounds equally confident on either side of it, which is exactly why the limits are worth understanding as a category rather than per release.
4. Work out what your actual usage comes to
Published token costs are close to meaningless on their own, because nobody knows what a million tokens looks like in their own work until they measure it.
Astra is listed at 10 dollars per million input tokens and 50 dollars per million output tokens, with cached input an order of magnitude cheaper. What matters is what your firm's real jobs consume. Take three tasks you would genuinely hand it, run them, and look at the usage. Two things usually surface: the expensive work is rarely the work people assumed, and caching changes the arithmetic more than any published number suggests.
Do this before anyone builds a business case, and again after a month of real use, because the first estimate is always wrong.
5. Understand what it refuses, and how it refuses
Newer models come with more gating, and the gating has operational consequences that never appear in a review.
OpenAI classified Astra as its first model to reach the Critical threshold for cybersecurity capability, having developed exploits for hardened browsers and operating systems during testing. Standard access refuses advanced security work as a result. More practically for a firm: some API safety checks stop a task outright rather than pausing for a human decision.
A task that halts mid-run is a different operational problem from a task that asks permission. If you are building anything that runs unattended, that behaviour belongs in your design, and it is a good reason to keep a person on the end of anything that matters until you know how it behaves.
6. Decide what you tell the team, and say it
Within a day of any major launch, someone in your firm has read a headline saying the profession is finished. Silence from leadership is not neutral in that situation, it is an answer, and it is the wrong one.
The message does not need to be long. What it is, whether the firm has it, whether staff may use it and on what, and when you will look again. Four sentences. If your one-page AI policy already covers approved tools and client data, this is a note rather than a rewrite.
The two-hour test that beats any review
Once the six checks are done, stop reading about the model and use it on work you already know the answer to.
Pick three jobs from last month that are finished and checked. Run them. Compare. You are not measuring whether the output looks impressive, because it always does. You are measuring how long the review took, and what kind of mistakes appeared. A model that is wrong in obvious ways is safer in a firm than one that is wrong in plausible ways, and no benchmark will ever tell you which one you have.
Write down what you found, with the date, and keep it. Six months later, when the next launch arrives, that page is worth more than the entire launch presentation, because it is about your work.
What not to do
- Do not switch your firm standard on launch week. Retraining people has a real cost and the ranking will move again within a quarter.
- Do not put client data into something to test it. Use your own files or a redacted example, the same rules that apply to any tool touching client information.
- Do not run the exercise on the strength of a headline. Vendor and third-party numbers frequently disagree, and the disagreement is the useful part.
- Do not decide nothing, twice in a row. Two consecutive releases with no decision is a decision, and firms that keep waiting for the tools to settle tend to still be waiting.
The point of a fixed process is that it makes each launch cheap to assess. The first time it takes an afternoon, and after that it is two hours and a note in a file. That is a sustainable posture towards a technology that is going to keep doing this, and it is a great deal better than reading benchmark threads at midnight. The rest of our coverage sits on the AI for accounting firms hub.
Frequently asked questions
How often should a firm review which AI model it uses?
Run the six checks at every major release, which takes about two hours once you have done it once, and revisit your actual firm standard once or twice a year. Changing the tool your team uses has retraining costs that a small capability improvement will not repay, so the review and the switch are separate decisions with different thresholds.
Why does a model's knowledge cutoff matter so much in accounting?
Because rates, thresholds, deadlines and rulings change every year, and a model has no awareness of what happened after its cutoff while sounding exactly as confident about it. Astra's cutoff is 30 April 2026. Any answer touching current-year figures needs to come from a document you supplied or a source the model searched, not from memory.
Are benchmark scores worth paying attention to at all?
Yes, for the specific task you care about, and treated as an upper bound. Scores are produced under favourable conditions by the party with an interest in the result, and some are run on custom harnesses that materially change the figure. Use them to decide whether something is worth two hours of your testing, never as the decision itself.
Should we standardise on one AI model across the firm?
For most firms, yes, on one primary tool, because training, policy, support and review are much simpler with a single default and the differences between leading models are smaller than the difference between using one well and using three badly. Keep the standard under annual review rather than treating it as permanent.
What is the minimum a small firm should do when a new model launches?
Check whether you can access it, check the knowledge cutoff, and send four sentences to your team. That is twenty minutes and it covers the two things most likely to cause a problem, which are staff using something nobody has assessed and staff assuming the firm has no view. Everything else can wait for a quieter week.
aimodel evaluationgovernancepractice managementgpt-6