🔍 Read the full analysis: How My AI Stack Works In September 2026 on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
With six frontier AI models clustered within about 20 index points while their per-task costs differ by roughly 100x, Thorsten Meyer’s September 2026 stack assigns Claude Opus 5.5 as main builder, newly released GPT-6.1 Sol as detail-and-review model, and Jev for routing. Effort settings, not model choice, are the biggest cost lever.
Technology writer Thorsten Meyer published his September 2026 AI model stack on 29 September, built around a market shift he says has taken hold over the past four weeks: six leading models now sit within about 20 index points of each other on the Artificial Analysis Intelligence Index, while their cost per task differs by roughly 100x. His answer to the new question — “which model clears my quality bar at the lowest cost per task?” — pairs Claude Opus 5.5 as the main builder with GPT-6.1 Sol, released the same day, as a low-cost reviewer, plus a decision-only model for routing.
Meyer’s stack assigns Opus 5.5 at high or xhigh effort as the primary model for development work, on the grounds that high delivers 54 index points at $1.82 per task and xhigh adds 2 more points for hard problems such as architecture and migrations. GPT-6.1 Sol at high or xhigh handles detail work and an independent review pass at $0.32 to $0.39 per task. Astra, Fable, Sonnet 5.5 and Luna serve as alternates for specific jobs, while Jev, a decision model Meyer notes “cannot write a sentence,” takes over high-volume yes/no and routing judgements.
Three findings anchor the piece, all based on Artificial Analysis Intelligence Index v4.3.x scores. First, Opus 5.5 outscores its more expensive sibling Fable 5.1 by 5 points while costing less per task — $5.98 against $7.63 at top settings. Second, Sonnet 5.5 at max effort costs more per task than Opus at max for 2 fewer points, which Meyer argues disqualifies it from that setting. Third, GPT-6.1 Sol costs roughly one-eighth of Astra and one-twentieth of Fable per task for a score only 1 to 2 points lower.
GPT-6.1 Sol launched on 29 September at $2 / $10 per million tokens, matching its week-old predecessor, and its medium setting already matches GPT-6 Sol’s score of 48 at one-fifth of that model’s $1.06 per-task cost. The catches, per Meyer: high and xhigh settings take 57 to 69 seconds to first token, making Sol non-interactive at those levels, and Opus 5.5 still leads it by 5 points at xhigh.
Opus builds. Sol reviews. Jev decides.
One price tape, six models
Score against cost, at every effort setting
The effort dial moves the bill more than the model
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol: near-Astra scores at a fraction of the price
Three published settings
| Setting | Index | Cost per task | Output tokens | First token |
|---|---|---|---|---|
| medium | 48 | $0.21 | 15M | 5.3 s |
| high | 50 | $0.32 | 25M | 57 s |
| xhigh | 51 | $0.39 | 36M | 69 s |
Same score band, very different bill
My stack: who builds, who reviews
Cheaper tokens are not cheaper work
Read the numbers with four warnings
Part 2: Jev, the model that decides instead of writing
One call in, typed answers out
Three question types
Confidence is the superpower
Three uses running in my publishing operation
The fit test, then the shadow test
- Replay 300 to 500 past decisions
- Compare overall and per confidence band
- Read 20 disagreements, decide who was right
- High band at 95% or better?
- Own flag, off by default
- Canary on 5 to 10 units
- Roll out in the confident band only
24 use cases, sorted by how well they fit
Proven in production
- 1Relevance gate
- 2Language check
- 3Classifier fallback
Publishing and content
- 4Thin-source detector
- 5Same-event dedupe
- 6Product fits roundup
- 7Disclosure present
- 8Headline quality
- 9Comment moderation
Commerce and support
- 10Support-ticket routing
- 11Return-reason coding
- 12Review to feature complaints
- 13Catalogue taxonomy
- 14Order-fraud pre-triage
Software and AI systems
- 15LLM guardrail
- 16RAG passage filter
- 17Citation check
- 18Tool and intent routing
- 19Log-line triage
- 20PR risk triage
Business ops and home
- 21Inbox triage
- 22Expense categorisation
- 23Lead qualification
- 24Smart-home intent
Limits, cost and one hard rule
Why Per-Task Cost Now Drives Model Choice
The piece documents a practical consequence of a maturing frontier market: when capability gaps shrink to single index points, the deciding variables become cost per task and effort settings. Meyer shows that on Opus 5.5, moving from xhigh to max adds 2 index points but 73% more cost per task, and from medium to max raises cost 4.46x for 7 points — meaning the effort dial moves the bill more than the choice between most models.
The stack’s core idea is the affordable review seat: a different model family checking Opus’s output, at $0.39 per task, is both a better check than self-review and cheap enough to run on every meaningful change. Meyer also cautions that halving model price saves only 12.5% of real cost in his illustrative example, and a single extra minute of human review erases the saving — a figure he flags as illustrative rather than measured.
A Month of Front-Model Releases
The stack lands at the end of a crowded release month: Claude Fable 5.1 on 1 September, GPT-6 Astra on 3 September, Opus 5.5 and GPT-6 Luna on 22 September, Sonnet 5.5 on 28 September, and GPT-6.1 Sol on 29 September. Per-task costs across the group range from Luna’s $0.07 (1,429 tasks per $100) to Fable’s $7.63 (13 tasks per $100).
All capability figures come from the Artificial Analysis Intelligence Index v4.3.x, which Meyer describes as a map of general capability rather than a verdict on any specific workload — hence his standing advice to shadow-test before switching models. Sonnet 5.5 at max effort produced about 193k output tokens per task, the most Artificial Analysis has measured, which Meyer cites as evidence of poor value at that setting.
“In four weeks, the AI frontier stopped being a leaderboard and became a price curve.”
— Thorsten Meyer, ThorstenMeyerAI.com
Limits of the Index and the Data
Meyer is explicit that all scores come from a single benchmark family, Artificial Analysis Intelligence Index v4.3.x, and that one index point sits inside the noise — a gap that covers several of the comparisons he draws. The index measures general capability, not performance on any individual workload, and he recommends shadow-testing before switching models.
For GPT-6.1 Sol, released the same day as publication, Artificial Analysis has not yet published low or max effort settings, and only three effort levels are listed so far. Meyer’s cost-of-human-review example is labeled illustrative, not measured, and the piece represents one practitioner’s configuration rather than a vendor recommendation or independently verified evaluation.
Watch Settings, Prices and New Benchmarks
Expect Artificial Analysis to publish Sol’s remaining effort settings, which could shift its value assessment. Meyer notes that future model swaps in his stack depend on where new releases land on the price curve, and his four operating rules — effort is not capability, a different model reading the same flawed spec is not an independent review, passing tests are not approval to ship, and failing work goes back to the builder with evidence — will govern how the stack evolves through future releases.
Key Questions
Which model does Meyer use as his main builder?
Claude Opus 5.5 at high or xhigh effort — 54 index points at $1.82 per task for everyday development, or 56 points at $3.46 for architecture, migrations and trust boundaries.
Why use GPT-6.1 Sol instead of Opus for review?
Sol’s xhigh setting sits only 1 to 2 points below Astra and Fable at $0.39 per task instead of $3.26 or $7.63, making a routine second-opinion pass from a different model family affordable on every meaningful change.
What is the biggest cost lever in this stack?
The effort setting. On Opus 5.5, going from xhigh to max adds 2 index points and 73% more cost per task; medium to max raises cost 4.46x for 7 points.
What are GPT-6.1 Sol’s drawbacks?
High and xhigh settings take 57 to 69 seconds to produce a first token, so it is not interactive at those levels, and Opus 5.5 still leads it by 5 index points at xhigh.
What is Jev used for?
Jev is a decision model that Meyer says cannot write a sentence; it handles high-volume yes/no judgements and routing, such as classification and extraction work alongside GPT-6 Luna.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
