🔍 Read the full analysis: From Building To Decisions: My September 2026 AI Stack on ThorstenMeyerAI.com
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
TL;DR
In a Sept. 29 article, Thorsten Meyer describes using Claude Opus 5.5 for development and GPT-6.1 Sol for detailed review, with cheaper models assigned to narrower tasks. His comparison, based mainly on Artificial Analysis Intelligence Index v4.3.x, argues that task cost and effort settings can matter as much as capability scores. The figures are the author’s reported results and may not predict performance on other workloads.
Thorsten Meyer said on September 29, 2026, that he uses Claude Opus 5.5 for most development work and newly released GPT-6.1 Sol for detailed review, grounding the division in a comparison of model scores and estimated task costs. His account matters to teams choosing AI tools because it argues that models with closely grouped benchmark scores can carry sharply different costs.
Meyer says the comparison draws on the Artificial Analysis Intelligence Index v4.3.x, which he describes as a general capability measure rather than a verdict on a particular workload. In his table, Opus 5.5 scores 58 at its top setting and costs $5.98 per task; GPT-6.1 Sol at xhigh scores 51 and costs $0.39. The figures and role assignments are reported by Meyer, not independently verified here.
The article places Opus 5.5 at high for features, APIs, refactors and multi-file work, and at xhigh for more demanding architecture and migration problems. Meyer assigns Sol at high or xhigh to file-level investigation and review, describing its cost as low enough to use routinely. He lists Astra or Fable as alternatives when Sol and Opus disagree, and Sonnet 5.5 or Luna for scoped subtasks and routine checks.
His table also compares the other models’ top-setting results: Sonnet 5.5 scores 56 at $7.60 per task, Fable 5.1 scores 53 at $7.63, Astra scores 53 at $3.26, and Luna scores 37 at $0.07. Meyer cautions that a one-point gap falls within the index’s noise. He says Sol’s high and xhigh settings take 57 to 69 seconds to produce a first token, a delay that limits their fit for interactive use.
Opus builds. Sol reviews. Jev decides.
One price tape, six models
Score against cost, at every effort setting
The effort dial moves the bill more than the model
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol: near-Astra scores at a fraction of the price
Three published settings
| Setting | Index | Cost per task | Output tokens | First token |
|---|---|---|---|---|
| medium | 48 | $0.21 | 15M | 5.3 s |
| high | 50 | $0.32 | 25M | 57 s |
| xhigh | 51 | $0.39 | 36M | 69 s |
Same score band, very different bill
My stack: who builds, who reviews
Cheaper tokens are not cheaper work
Read the numbers with four warnings
Part 2: Jev, the model that decides instead of writing
One call in, typed answers out
Three question types
Confidence is the superpower
Three uses running in my publishing operation
The fit test, then the shadow test
- Replay 300 to 500 past decisions
- Compare overall and per confidence band
- Read 20 disagreements, decide who was right
- High band at 95% or better?
- Own flag, off by default
- Canary on 5 to 10 units
- Roll out in the confident band only
24 use cases, sorted by how well they fit
Proven in production
- 1Relevance gate
- 2Language check
- 3Classifier fallback
Publishing and content
- 4Thin-source detector
- 5Same-event dedupe
- 6Product fits roundup
- 7Disclosure present
- 8Headline quality
- 9Comment moderation
Commerce and support
- 10Support-ticket routing
- 11Return-reason coding
- 12Review to feature complaints
- 13Catalogue taxonomy
- 14Order-fraud pre-triage
Software and AI systems
- 15LLM guardrail
- 16RAG passage filter
- 17Citation check
- 18Tool and intent routing
- 19Log-line triage
- 20PR risk triage
Business ops and home
- 21Inbox triage
- 22Expense categorisation
- 23Lead qualification
- 24Smart-home intent
Limits, cost and one hard rule
Why Cost Per Task Matters
Meyer’s comparison focuses attention on the cost of completing a task at a given quality level, rather than token prices or a leaderboard position alone. In his reported figures, Sol xhigh costs $0.39 per task, compared with $3.26 for Astra and $7.63 for Fable. If those estimates hold for a team’s own work, the difference could affect how often it can afford to run a second model review or process high volumes of routine work.
The account also shows why a benchmark result should not be treated as a purchasing decision by itself. Meyer describes the index as a map of general capability and recommends shadow-testing before switching. A cheaper model may need more human checking, take longer to respond or perform differently on a team’s code and instructions. The relevant comparison for readers is the one their own tasks and review process produce.
His proposed two-model workflow makes review cost part of software quality control. A separate model family can provide another perspective on code, but Meyer notes that reviewers can still share a flawed specification. He says passing tests are not approval to ship, and recommends giving a failed review and its evidence back to the builder model. That is a process recommendation from the author, not a measured finding in the index.
How Meyer Built His Stack
Meyer’s article is dated September 29, 2026, and says GPT-6.1 Sol was released that day. It places that release against a fast-moving set of models: Opus 5.5 is listed as released Sept. 22, Sonnet 5.5 on Sept. 28, Fable 5.1 on Sept. 1, Astra on Sept. 3 and Luna on Sept. 22. These dates and model details come from the source account.
The comparison separates model choice from effort setting. For Opus 5.5, Meyer reports that moving from medium to max raises the index score from 51 to 58 while increasing estimated cost per task from $1.34 to $5.98. At xhigh, he reports a score of 56 for $3.46. He therefore chooses high for ordinary development and reserves xhigh for selected harder problems.
He makes a similar distinction for Sonnet 5.5. Its reported high setting scores 47 for $1.08 per task, while max scores 56 for $7.60. Meyer says max generated about 193,000 output tokens per task in the index, which he identifies as the highest measured there. He argues that the extra effort can make the most expensive setting difficult to justify for his use.
““The practical reading: Sol is not the model I ask to build. It is the model I can afford to run on everything.””
— Thorsten Meyer
Limits of the Model Comparison
The article does not establish how these models perform across independent workloads or whether other users would see the same cost per task. The index is a general benchmark, and Meyer advises readers to shadow-test before changing their systems. The source does not provide enough detail to reproduce every task estimate or assess how its workload compares with a particular organization’s work.
Some model settings were not available in the cited index at publication: Meyer says low and max results for GPT-6.1 Sol had not yet been published. He also says one index point is within the noise, so small score differences should not be read as decisive. The article’s concluding cost example is explicitly illustrative rather than measured, and the supplied source ends mid-sentence before giving its full calculation.
It is also unclear whether the reported release timing, prices and benchmark entries will remain current as providers update models and rates. Meyer lists token prices separately from per-task estimates; those measures do not by themselves establish a team’s total cost, which can include human review and waiting time.
Test Before Changing Models
Meyer’s stated next step for readers is to shadow-test candidate models on their own work before switching defaults. That would let teams compare output quality, latency, task cost and the amount of human review needed using the same prompts and acceptance criteria. The article does not announce a formal test plan or a follow-up date.
For his own workflow, Meyer says Opus remains the main builder, with Sol used for details and a second review. He suggests bringing a failed review, along with its evidence, back to Opus for correction. Whether that division remains useful will depend on future index updates and results on each team’s workload; the source does not report later testing.
Key Questions
What model does Thorsten Meyer use for building?
He says Claude Opus 5.5 at high effort is his main model for features, APIs, refactors and multi-file work. He uses xhigh for selected harder problems.
Why does he use GPT-6.1 Sol for review?
Meyer says Sol at high or xhigh can investigate files and review changes at a reported $0.32 to $0.39 per task. He also reports a 57-to-69-second time to first token at those settings.
Which benchmark supports the comparison?
The article cites the Artificial Analysis Intelligence Index v4.3.x. Meyer says it measures general capability and should not be treated as a verdict on a reader’s specific workload.
Does Meyer recommend switching based on the scores alone?
No. He recommends shadow-testing models on a team’s own tasks before changing its setup, since benchmark scores and reported task costs may not predict local results.
Source: ThorstenMeyerAI.com
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
