🔍 Read the full analysis: Can Jev Improve Your AI Decision Process? 24 Ways To Use It on ThorstenMeyerAI.com
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
TL;DR
Thorsten Meyer’s September 29 article maps 24 possible uses for Jev, a tool that returns typed answers to narrow questions so software can route routine decisions. Meyer says three uses are live in his publishing operation, 12 meet his fit test, seven need measurement and two are poor fits; those classifications and performance figures are his reported results, not independent evaluations.
Thorsten Meyer published a 24-use assessment of Jev on September 29, reporting that three applications are already running in his publishing operation and that 12 others meet his proposed fit test. The guide matters to teams considering AI for routine decisions because it argues for measuring whether a task suits the tool before wiring it into a workflow.
Meyer describes Jev as a system that takes a state, such as text or JSON, alongside typed questions and returns answers that software can act on. Its answer types include a yes-or-no probability, a choice with probabilities and confidence, and a score on ordered levels. Meyer says a call takes about 0.3 to 0.9 seconds and costs about $0.04 per million input tokens. These are figures reported in his article; the source material does not provide an independent cost or speed audit.
The three live applications Meyer lists are a relevance gate for matching stories to sites, a language check, and a fallback topic classifier. He reports that a scan of 78,889 articles cost $2.01, found 1,576 non-English articles and fixed 1,553. For classification across 31 topics, he reports 89% agreement with a frontier large language model overall, rising to 97% to 99% when Jev’s confidence was at least 0.8. The article does not specify an independent evaluation or full test methodology for those figures.
Meyer’s four conditions for adopting Jev are high volume, a narrow question, low-cost errors or a route for uncertain cases, and evidence that the existing heuristic fails. He recommends replaying 300 to 500 past decisions, comparing performance across confidence bands, and reviewing disagreements before deployment. His examples include disclosure checks and comment moderation as strong fits, while headline deduplication is a poor fit in his assessment because a canary found no duplicates to address.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
Where Routine AI Decisions Fit
The guide’s practical argument is that automation should target repeated, narrow decisions where a wrong answer can be contained. A confidence score can support routing: software handles clear cases and sends uncertain ones to a person or a more capable system. In Meyer’s design, the code that uses Jev’s answer sets the action; Jev itself does not determine the downstream outcome.
That distinction matters for publishers and operations teams weighing the cost of checking every item against the risk of missed errors. Meyer’s examples suggest potential applications in content review, commerce and customer operations, but his categories remain an assessment by the tool’s author. The reported production results provide a starting point for evaluation, not proof that the same accuracy, savings or safety will hold in another organization.
From Publishing Checks to a Fit Test
Meyer says the need for a thin-source detector became apparent because 88% of news items he processes begin with a bare headline. He classifies that detector as “measure first”: a model could judge whether a source includes enough verifiable facts, but he says the current heuristic’s failure needs to be demonstrated. The proposed rule would fetch the original source or skip the item at very low confidence and write only when confidence is high.
The article also describes uses that need more evidence before adoption. A product-to-roundup check could catch an item that does not match a reader’s search, while a headline-quality score could prompt review before publication. Meyer labels both “measure first.” He calls disclosure detection and comment moderation strong fits, with human review for uncertain disclosure cases and confidence thresholds for approving or hiding comments.
“Jev is the right tool wherever a system needs thousands of small judgements and can hand the unclear ones to something smarter.”
— Thorsten Meyer, article author
Evidence Still Needed Across Use Cases
The article’s claims about speed, cost and agreement are attributed to Meyer, and the provided material does not include independent verification, detailed datasets or a full description of the evaluation method. It is also unclear how results would vary across different content, languages, industries or error costs. Meyer’s reported performance on a 31-topic classification should not be taken as a general accuracy guarantee.
Seven of the 24 ideas are marked “measure first” because Meyer says the existing heuristic has not been shown to fail. The source excerpt also ends while describing commerce and customer operations, so it does not provide the full list of 24 use cases or the rationale for every rating. The article’s summary says two uses are poor fits, but the supplied details identify only the deduplication example.
Measure Before Switching On
Meyer recommends replaying 300 to 500 real past decisions, checking results overall and by confidence band, and reviewing 20 disagreements to determine which decision was correct. He advises wiring Jev into a workflow only where the high-confidence band reaches 95%, then testing behind a separate flag on 5% to 10% of units before wider rollout. These are his proposed steps; the article does not give a schedule for publishing further measurement results.
Key Questions
What does Jev do?
According to Meyer, Jev takes text or JSON plus typed questions and returns structured answers, such as a probability, a category choice or a score. Software can then use those answers to route a decision.
How many of the 24 uses does Meyer say are ready?
Meyer says three are live in his publishing operation and 12 more are strong fits. He marks seven for measurement first and two as poor fits. These are his classifications in the article.
What does Meyer’s fit test require?
The task should involve high volume and a narrow question, errors should be inexpensive or uncertain cases should go elsewhere, and measurement should show that the current heuristic fails.
Are Jev’s reported accuracy figures independently verified?
The supplied article material attributes the figures to Meyer and describes his measurement on a 31-topic classification. It does not provide independent verification or enough methodological detail to establish how broadly the results apply.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
