AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: From Building To Decisions: My September 2026 AI Stack on ThorstenMeyerAI.com

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

TL;DR

In a Sept. 29 article, Thorsten Meyer describes using Claude Opus 5.5 for development and GPT-6.1 Sol for detailed review, with cheaper models assigned to narrower tasks. His comparison, based mainly on Artificial Analysis Intelligence Index v4.3.x, argues that task cost and effort settings can matter as much as capability scores. The figures are the author’s reported results and may not predict performance on other workloads.

Thorsten Meyer said on September 29, 2026, that he uses Claude Opus 5.5 for most development work and newly released GPT-6.1 Sol for detailed review, grounding the division in a comparison of model scores and estimated task costs. His account matters to teams choosing AI tools because it argues that models with closely grouped benchmark scores can carry sharply different costs.

Meyer says the comparison draws on the Artificial Analysis Intelligence Index v4.3.x, which he describes as a general capability measure rather than a verdict on a particular workload. In his table, Opus 5.5 scores 58 at its top setting and costs $5.98 per task; GPT-6.1 Sol at xhigh scores 51 and costs $0.39. The figures and role assignments are reported by Meyer, not independently verified here.

The article places Opus 5.5 at high for features, APIs, refactors and multi-file work, and at xhigh for more demanding architecture and migration problems. Meyer assigns Sol at high or xhigh to file-level investigation and review, describing its cost as low enough to use routinely. He lists Astra or Fable as alternatives when Sol and Opus disagree, and Sonnet 5.5 or Luna for scoped subtasks and routine checks.

His table also compares the other models’ top-setting results: Sonnet 5.5 scores 56 at $7.60 per task, Fable 5.1 scores 53 at $7.63, Astra scores 53 at $3.26, and Luna scores 37 at $0.07. Meyer cautions that a one-point gap falls within the index’s noise. He says Sol’s high and xhigh settings take 57 to 69 seconds to produce a first token, a delay that limits their fit for interactive use.

At a glance
reportWhen: Published September 29, 2026; GPT-6.1 S…
The developmentThorsten Meyer published a Sept. 29, 2026 account of how he assigns AI models to software development, review and routine decisions based on reported capability scores and cost per task.

Opus builds. Sol reviews. Jev decides.

The September 2026 AI stack in one page: six frontier models on one price curve, and a decision model for the high-volume judgements that do not need a sentence.
Scores: Artificial Analysis Intelligence Index v4.3.x. Data as of 29 September 2026.
BuildsClaude Opus 5.5 at high or xhigh effort
Digs and reviewsGPT-6.1 Sol at high or xhigh effort
DecidesJev on high-volume yes/no and routing calls

One price tape, six models

Put every model on the same cost-per-task ruler and capability looks compressed. The bill does not.
Price tape: cost per task of six models on a log scale, from GPT-6 Luna at $0.07 to Fable 5.1 at $7.63$0.05$0.10$0.50$1$5$10cost per task, log scale: each tick is a different order of magnitudeGPT-6 Lunaindex 37 · $0.07GPT-6.1 Solindex 51 · $0.39 (xhigh)GPT-6 Astraindex 53 · $3.26Opus 5.5index 58 · $5.98Sonnet 5.5 · index 56 · $7.60Fable 5.1 · index 53 · $7.63about 100× from the cheapest to the priciest, but only 21 index points between them

Score against cost, at every effort setting

Each dot is an effort level. Opus 5.5 at high already matches Astra and Fable at max on this index, for less money.
Intelligence Index score against cost per task for each effort setting of six models$0.01$0.10$1$102030405060cost per Intelligence Index task, log scaleindexOpus high / xhigh: my defaultOpus 5.5Sonnet 5.5Fable 5.1GPT-6 AstraGPT-6.1 Sol (new)GPT-6 Sol (Sep 22), dashedGPT-6 Lunaup and to the left is better
Astra and Fable are shown at their top published setting. Luna starts at $0.0045 per task. GPT-6.1 Sol has no low or max setting published yet.

The effort dial moves the bill more than the model

Going from medium to max on Opus costs 4.46× more for 7 points. That is why I run high or xhigh.

Claude Opus 5.5

$0.55
42
$1.34
51
$1.82
54
$3.46
56
$5.98
58
low
medium
high
xhigh
max
Solid bars are where I run it. Max adds 2 points over xhigh for 73% more cost.

Claude Sonnet 5.5

$0.41
36
$0.59
41
$1.08
47
$2.74
52
$7.60
56
low
medium
high
xhigh
max
Best value is high. At max it writes about 193k output tokens per task, the most measured.

GPT-6.1 Sol: near-Astra scores at a fraction of the price

Launched 29 September at $2 in and $10 out per 1M tokens. It sits 1 to 2 points under Astra and Fable, and Opus xhigh still leads it by 5.

Three published settings

SettingIndexCost per taskOutput tokensFirst token
medium48$0.2115M5.3 s
high50$0.3225M57 s
xhigh51$0.3936M69 s
Median for comparable models is 82M output tokens. High and xhigh are not interactive: plan for a wait before the first token.

Same score band, very different bill

GPT-6.1 Sol xhigh
$0.39index 51
Opus 5.5 high
$1.82index 54
GPT-6 Astra max
$3.26index 53
Opus 5.5 xhigh
$3.46index 56
Fable 5.1 max
$7.63index 53
Cost per Intelligence Index task. A one-point gap is inside the noise.

My stack: who builds, who reviews

Opus does the work. A second model family reviews it, because a different reviewer catches what the author cannot see.
Stack diagram: Opus 5.5 builds at high effort, escalates to xhigh, and sends every change to GPT-6.1 Sol for review; Astra or Fable give a second opinionOpus 5.5 · xhighhard problems: architecture,migrations, trust boundariesOpus 5.5 · highMAIN BUILDERfeatures, APIs, multi-filework, refactorsescalate when it gets hardGPT-6.1 Solhigh or xhighdigs into details andreviews every change$0.32–0.39 per taskdifffindingsAstra or Fablesecond opinion, 8 to 20×the cost per taskif they disagreeSonnet 5.5 · Lunaside work: scopedsubtasks, bulk checksand routingFailed review? Hand Opus the failing case and the evidence.Never just “try harder”: effort cannot supply a missing requirement.
Effort is not capability. Turning the dial up does not make a model smarter.
Effort cannot fill gaps. A missing requirement stays missing at any setting.
Different model, same spec. That is not independent review if both read the same flawed brief.
Green tests are not approval. Passing tests only prove what the tests cover.

Cheaper tokens are not cheaper work

Illustrative, not measured: $1 of model time plus 4 minutes of review at $45 an hour. Halving the model price saves 12.5% of the total. One extra minute of review erases it.
$4.00
review $3.00
model $1.00
Baseline
$3.50
review $3.00
model $0.50
Model price cut 50%
$4.25
review $3.75
model $0.50
Cheaper model plus 1 extra minute of review
Track cost per accepted result: model, tools, review and rework, divided by the results someone actually uses.

Read the numbers with four warnings

The index movesFable scored 66 on an earlier version and 53 on v4.3. Compare within one version only.
Fallback is includedFlagged cyber and biology tasks route to older Anthropic models, now on Sonnet 5.5 too.
Max is not productionReal deployments run medium or high, where gaps narrow and costs fall.
Your work decidesShadow-test on your own tasks. Budget cost per task, not per token.

Part 2: Jev, the model that decides instead of writing

Jev cannot write, summarise or extract. It answers narrow typed questions with a probability and an honest confidence, in under a second, for about $0.04 per million input tokens.

One call in, typed answers out

Your code, not Jev, decides what to do with each answer, usually by confidence band.
Jev flow: state and typed questions go into one Jev call; typed answers with confidence come out; code acts alone, escalates the gray zone, or logsStatea ticket, a story,a site profile,a log line …+ typed questions,many per callJevone call0.3 to 0.9 s$0.042 / M tokens inAnswersnoul: 0.03choice: billing p 0.91, conf 0.86score: 2.7 of 3 conf 0.64code branches on thisAct aloneconf ≥ 0.8Escalategray zone toLLM or humanLogmeasure first

Three question types

noul
A yes/no question. Returns the probability of yes, 0 to 1.
gates, flags, filters
choice
Pick one option. Returns the choice, a probability per option, and a confidence.
routing, classification, taxonomy
score
Rate on your ordered levels. Returns a position (it can fall between levels) plus a confidence.
quality, fit, severity, priority

Confidence is the superpower

In my own measurement on a 31-topic classification, Jev agreed with a frontier LLM almost every time it was sure, and rarely when it was not. So: decide the clear cases, route the gray zone.
confidence 0.8 or higher
97–99%
all answers
89%
confidence below 0.5
42%
Agreement with a frontier LLM, my production data, September 2026, rounded.

Three uses running in my publishing operation

About 90,000 decisions so far. Checks I could only afford on a sample now cover everything.
$2.01
Language check
78,889 articles scanned overnight. 1,576 in the wrong language found, 1,553 fixed in place.
22%
Relevance gate
About 10,000 story-to-site pairings judged in 3 days. Only 22% were clearly on-topic.
89%
Classifier fallback
Agreement with the primary LLM across 31 topics, used when that LLM errors.

The fit test, then the shadow test

Use Jev only when all four hold. Then prove it on past decisions before it acts on anything.
High volumeThousands of small calls, not a handful of big ones.
Narrow questionNo multi-step reasoning needed.
Cheap errorsOr unsure cases go to something smarter.
Heuristic failsVisibly, and measured, not assumed.
  1. Replay 300 to 500 past decisions
  2. Compare overall and per confidence band
  3. Read 20 disagreements, decide who was right
  4. High band at 95% or better?
  5. Own flag, off by default
  6. Canary on 5 to 10 units
  7. Roll out in the confident band only

24 use cases, sorted by how well they fit

Start from the strong fits. The amber ones need a measurement before you trust them, and the red ones fail one of the four conditions.
in productionstrong fitmeasure firstpoor fit

Proven in production

  • 1Relevance gate
  • 2Language check
  • 3Classifier fallback

Publishing and content

  • 4Thin-source detector
  • 5Same-event dedupe
  • 6Product fits roundup
  • 7Disclosure present
  • 8Headline quality
  • 9Comment moderation

Commerce and support

  • 10Support-ticket routing
  • 11Return-reason coding
  • 12Review to feature complaints
  • 13Catalogue taxonomy
  • 14Order-fraud pre-triage

Software and AI systems

  • 15LLM guardrail
  • 16RAG passage filter
  • 17Citation check
  • 18Tool and intent routing
  • 19Log-line triage
  • 20PR risk triage

Business ops and home

  • 21Inbox triage
  • 22Expense categorisation
  • 23Lead qualification
  • 24Smart-home intent

Limits, cost and one hard rule

No writing, summarising or extractionPair it with an LLM for the write step.
No world knowledgePut a snippet in the state; a bare name means nothing.
Reads your wording literallyA rewording moved my results about 2 points. Freeze it, re-measure after changes.
Weaker on non-English, maths, datesKeep those checks on an LLM. Early access, hosted API only.
100,000 decisions ≈ $2.50
About 60M input tokens at $0.042 per million, output free, roughly 600 tokens per three-question call. Latency 0.3 to 0.9 seconds.
Never the sole decision-maker for consequences about people. Hiring, credit, medical and legal outcomes stay with a human. Jev can sort and flag. A person decides.
Sources. Model scores, cost per task and speeds: Artificial Analysis, Intelligence Index v4.3.x, including the GPT-6.1 Sol medium, high and xhigh pages, checked 29 September 2026. Astra and Fable scores from the Artificial Analysis v4.3 announcement. Jev figures are my own production measurements, September 2026, rounded. The review-bill example is illustrative. Read the full article on thorstenmeyerai.com.

Why Cost Per Task Matters

Meyer’s comparison focuses attention on the cost of completing a task at a given quality level, rather than token prices or a leaderboard position alone. In his reported figures, Sol xhigh costs $0.39 per task, compared with $3.26 for Astra and $7.63 for Fable. If those estimates hold for a team’s own work, the difference could affect how often it can afford to run a second model review or process high volumes of routine work.

The account also shows why a benchmark result should not be treated as a purchasing decision by itself. Meyer describes the index as a map of general capability and recommends shadow-testing before switching. A cheaper model may need more human checking, take longer to respond or perform differently on a team’s code and instructions. The relevant comparison for readers is the one their own tasks and review process produce.

His proposed two-model workflow makes review cost part of software quality control. A separate model family can provide another perspective on code, but Meyer notes that reviewers can still share a flawed specification. He says passing tests are not approval to ship, and recommends giving a failed review and its evidence back to the builder model. That is a process recommendation from the author, not a measured finding in the index.

How Meyer Built His Stack

Meyer’s article is dated September 29, 2026, and says GPT-6.1 Sol was released that day. It places that release against a fast-moving set of models: Opus 5.5 is listed as released Sept. 22, Sonnet 5.5 on Sept. 28, Fable 5.1 on Sept. 1, Astra on Sept. 3 and Luna on Sept. 22. These dates and model details come from the source account.

The comparison separates model choice from effort setting. For Opus 5.5, Meyer reports that moving from medium to max raises the index score from 51 to 58 while increasing estimated cost per task from $1.34 to $5.98. At xhigh, he reports a score of 56 for $3.46. He therefore chooses high for ordinary development and reserves xhigh for selected harder problems.

He makes a similar distinction for Sonnet 5.5. Its reported high setting scores 47 for $1.08 per task, while max scores 56 for $7.60. Meyer says max generated about 193,000 output tokens per task in the index, which he identifies as the highest measured there. He argues that the extra effort can make the most expensive setting difficult to justify for his use.

““The practical reading: Sol is not the model I ask to build. It is the model I can afford to run on everything.””

— Thorsten Meyer

Limits of the Model Comparison

The article does not establish how these models perform across independent workloads or whether other users would see the same cost per task. The index is a general benchmark, and Meyer advises readers to shadow-test before changing their systems. The source does not provide enough detail to reproduce every task estimate or assess how its workload compares with a particular organization’s work.

Some model settings were not available in the cited index at publication: Meyer says low and max results for GPT-6.1 Sol had not yet been published. He also says one index point is within the noise, so small score differences should not be read as decisive. The article’s concluding cost example is explicitly illustrative rather than measured, and the supplied source ends mid-sentence before giving its full calculation.

It is also unclear whether the reported release timing, prices and benchmark entries will remain current as providers update models and rates. Meyer lists token prices separately from per-task estimates; those measures do not by themselves establish a team’s total cost, which can include human review and waiting time.

Test Before Changing Models

Meyer’s stated next step for readers is to shadow-test candidate models on their own work before switching defaults. That would let teams compare output quality, latency, task cost and the amount of human review needed using the same prompts and acceptance criteria. The article does not announce a formal test plan or a follow-up date.

For his own workflow, Meyer says Opus remains the main builder, with Sol used for details and a second review. He suggests bringing a failed review, along with its evidence, back to Opus for correction. Whether that division remains useful will depend on future index updates and results on each team’s workload; the source does not report later testing.

Key Questions

What model does Thorsten Meyer use for building?

He says Claude Opus 5.5 at high effort is his main model for features, APIs, refactors and multi-file work. He uses xhigh for selected harder problems.

Why does he use GPT-6.1 Sol for review?

Meyer says Sol at high or xhigh can investigate files and review changes at a reported $0.32 to $0.39 per task. He also reports a 57-to-69-second time to first token at those settings.

Which benchmark supports the comparison?

The article cites the Artificial Analysis Intelligence Index v4.3.x. Meyer says it measures general capability and should not be treated as a verdict on a reader’s specific workload.

Does Meyer recommend switching based on the scores alone?

No. He recommends shadow-testing models on a team’s own tasks before changing its setup, since benchmark scores and reported task costs may not predict local results.

Source: ThorstenMeyerAI.com

EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Cybersecurity 2026: New Threats and How to Stay Safe

Learn about the emerging cybersecurity threats of 2026 and discover essential strategies to protect yourself from evolving digital dangers.

Improving Heuristics For A* Pathfinding

Researchers develop improved heuristics for A* pathfinding, promising faster and more efficient navigation in AI and robotics applications.

How Small Businesses Can Benefit From These 12 AI Automation Tools In 2026

Discover 12 AI automation tools that small businesses can leverage in 2026 to save time, reduce costs, and scale operations efficiently.

RoundupForge: The Data Layer

RoundupForge, a data layer, automates product deduplication and ranking for large-scale product roundups, ensuring trustworthy recommendations.