🔍 Read the full analysis: How Three AI Tools Support My September 2026 Work on ThorstenMeyerAI.com
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Thorsten Meyer’s September 29 assessment assigns Opus 5.5 to building, GPT-6.1 Sol to detailed review, and a decision model called Jev to high-volume routing. The comparisons cite Artificial Analysis Intelligence Index v4.3.x and estimated task costs; Meyer says teams should test models against their own work before switching.
Meyer bases most of his score comparisons on the Artificial Analysis Intelligence Index v4.3.x, which he describes as a measure of general capability rather than a verdict on any particular workload. In his table, Opus 5.5 scores 58 at its top setting and costs an estimated $5.98 per task. GPT-6.1 Sol scores 51 at xhigh and costs $0.39 per task. These figures are index results and task-cost estimates reported by Meyer, not guarantees for other users or tasks.
His stated workflow gives Opus 5.5 the main development role: high for features, APIs and refactors, and xhigh for harder work such as architecture and migrations. He uses GPT-6.1 Sol at high or xhigh to examine a specific file or change and to review Opus’s work. He says a review model from a different family can provide a useful second perspective, and that Sol’s estimated cost makes it practical to run on meaningful changes.
Meyer assigns narrower tasks to other systems: Sonnet 5.5 at high for scoped subtasks and documents, Luna for classification and extraction, and Astra or Fable as second opinions when Sol and Opus disagree. Jev, which he describes as unable to write sentences, handles yes-or-no and routing decisions at high volume. The source does not provide benchmark or cost figures for Jev.
Opus builds. Sol reviews. Jev decides.
One price tape, six models
Score against cost, at every effort setting
The effort dial moves the bill more than the model
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol: near-Astra scores at a fraction of the price
Three published settings
| Setting | Index | Cost per task | Output tokens | First token |
|---|---|---|---|---|
| medium | 48 | $0.21 | 15M | 5.3 s |
| high | 50 | $0.32 | 25M | 57 s |
| xhigh | 51 | $0.39 | 36M | 69 s |
Same score band, very different bill
My stack: who builds, who reviews
Cheaper tokens are not cheaper work
Read the numbers with four warnings
Part 2: Jev, the model that decides instead of writing
One call in, typed answers out
Three question types
Confidence is the superpower
Three uses running in my publishing operation
The fit test, then the shadow test
- Replay 300 to 500 past decisions
- Compare overall and per confidence band
- Read 20 disagreements, decide who was right
- High band at 95% or better?
- Own flag, off by default
- Canary on 5 to 10 units
- Roll out in the confident band only
24 use cases, sorted by how well they fit
Proven in production
- 1Relevance gate
- 2Language check
- 3Classifier fallback
Publishing and content
- 4Thin-source detector
- 5Same-event dedupe
- 6Product fits roundup
- 7Disclosure present
- 8Headline quality
- 9Comment moderation
Commerce and support
- 10Support-ticket routing
- 11Return-reason coding
- 12Review to feature complaints
- 13Catalogue taxonomy
- 14Order-fraud pre-triage
Software and AI systems
- 15LLM guardrail
- 16RAG passage filter
- 17Citation check
- 18Tool and intent routing
- 19Log-line triage
- 20PR risk triage
Business ops and home
- 21Inbox triage
- 22Expense categorisation
- 23Lead qualification
- 24Smart-home intent
Limits, cost and one hard rule
How Cost Shapes Model Roles
The account illustrates a practical purchasing question for teams using AI in software and knowledge work: how much capability does a task require, and what does that level cost? Meyer’s figures put Sol xhigh at $0.39 per task, compared with $3.26 for GPT-6 Astra and $7.63 for Claude Fable 5.1, while their listed index scores are relatively close. If those estimates hold for a team’s own tasks, using a lower-cost model for routine review could make more frequent checks affordable.
Cost also changes with the effort setting. Meyer reports that Opus 5.5 rises from $1.82 per task at high to $3.46 at xhigh and $5.98 at max. He says the maximum setting adds two index points over xhigh for 73% more estimated cost. These comparisons suggest that settings deserve attention alongside model selection, but benchmark scores alone do not establish whether added expense improves outcomes in a specific workflow.
Meyer cautions that cheaper model use does not automatically mean cheaper work: additional human review can outweigh token savings. His recommendations are based on his own workflow and the cited index, so readers would need to measure quality, time and costs on their own tasks before adopting the same split.
The September Model Comparisons
Meyer frames his account around a shift in the AI market: several models have scores within roughly 20 index points, while their estimated cost per task varies widely. The source lists Opus 5.5, released September 22, with a top-setting score of 58; Sonnet 5.5, released September 28, at 56; and GPT-6.1 Sol, released September 29, at 51 for xhigh. It lists Fable 5.1 at 53, GPT-6 Astra at 53, and GPT-6 Luna at 37.
The reported comparisons have limits. Artificial Analysis’s index is a general-capability benchmark, and the source does not show how its tasks map to Meyer’s actual development and review work. Meyer says one index point may fall within measurement noise and recommends shadow-testing before changing a workflow. His article also distinguishes benchmark results from its estimated costs per task, which may depend on how a task is defined and run.
“Which model clears my quality bar at the lowest cost per task?”
— Thorsten Meyer
Limits of the Cost Estimates
The source does not provide enough detail to independently assess how its cost-per-task estimates were calculated or how closely the index tasks match Meyer’s workload. It also does not report a controlled comparison of the proposed workflow against alternatives, or measured changes in software quality, review time or total cost.
Meyer notes that GPT-6.1 Sol’s low and max settings were not yet listed in the index and says a one-point score difference may be within noise. The source’s final cost example is cut off after saying that halving model price saves 12.5% of real cost and that an extra minute of human review erases that saving; it labels the example illustrative, not measured. The full assumptions behind that calculation are therefore unclear.
The article gives no benchmark scores or task-cost estimates for Jev, and does not explain how the model handles errors in routing decisions. It remains unclear whether the recommended assignments would produce similar results for other teams, prompts or software tasks.
Test Before Changing Workflows
Meyer recommends shadow-testing models on a team’s own tasks before switching systems. That would let teams compare quality and cost on the work they actually need done, including whether a lower-cost review pass catches issues that matter and whether extra human checks erase the savings.
Further comparisons may change as index coverage expands. The source says low and max results for GPT-6.1 Sol were not yet available at publication. It does not give a date for those results or announce a formal next test, so the timing of additional evidence remains unknown.
Key Questions
Which models does Meyer use for building and review?
Meyer says he uses Opus 5.5 for building and GPT-6.1 Sol at high or xhigh for detailed analysis and review.
How much does GPT-6.1 Sol cost per task in the cited comparison?
The source lists estimated costs of $0.21 at medium, $0.32 at high and $0.39 at xhigh. These are the article’s figures based on the cited index, not a universal price for every task.
What does Meyer use Jev for?
He describes Jev as a decision model for high-volume yes-or-no and routing judgments. The source does not give its benchmark score or cost per task.
Do the benchmark rankings establish which model is best for every team?
No. Meyer says the Artificial Analysis Intelligence Index measures general capability and advises readers to shadow-test models on their own workloads before switching.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
