🔍 Read the full analysis: 24 Approaches To Decision Modeling With Jev on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Thorsten Meyer says he has mapped 24 uses for Jev, a tool that returns typed answers to narrow questions so software can act on them. He reports three publishing uses already running, 12 strong fits, seven that need measurement and two poor fits; the figures and performance results come from his own account.
Thorsten Meyer published a list of 24 decision-modeling uses for Jev, a tool he says returns typed answers to narrow questions that software can act on. Meyer reports that three uses are running in his publishing operation, while 12 more meet his criteria for a strong fit; the proposal matters to teams deciding whether to automate frequent, low-cost judgments.
Meyer describes Jev as a system that receives text or JSON state along with typed questions, then returns answers such as a yes probability, a choice among options, or a score. He says it does not write, summarize or extract information. Instead, the caller’s code decides what to do with each answer. A single call can carry the state and several questions, and Meyer estimates a response time of about 0.3 to 0.9 seconds and an input cost of about $0.04 per million tokens.
His three reported live uses are a relevance gate for stories and sites, an English-language check, and a fallback topic classifier. For the language check, Meyer says he scanned 78,889 articles for $2.01, identified 1,576 as non-English and fixed 1,553. He also reports about 10,000 story-and-site pairings judged over three days for the relevance gate. The classifier, he says, agreed with a frontier large language model 89% of the time overall and 97% to 99% when Jev’s confidence was at least 0.8.
Meyer’s list assigns each use a fit label. Besides the three live cases, he calls 12 strong fits, says seven need measurement and marks two poor fits. In publishing, examples include checking disclosures and moderating comments as strong fits; detecting thin sources, matching products to roundups and reviewing headlines need more measurement. He labels same-event deduplication a poor fit because his canary found no duplicates.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
Where Automated Checks May Help
The list offers a practical decision rule for teams weighing small automated checks: use them when judgments are frequent, narrowly defined and inexpensive to get wrong, and when existing rules show a measurable weakness. Meyer says uncertain cases should go to a person or a more capable system, while clear cases can be handled automatically. That approach limits the consequences of a mistaken answer and gives teams a way to keep existing workflows for borderline decisions.
The reported publishing results also make the costs and scope concrete. A scan of tens of thousands of articles cost Meyer $2.01, according to his account, but that figure alone does not establish the quality of the classifications or whether the same costs would apply elsewhere. For publishers facing large queues of routine checks, his examples suggest where a pilot might be useful. They do not establish that every listed task will work well in another organization.
His poor-fit example is a useful constraint: low cost and a simple question do not by themselves justify automation. If a test finds no current problem to correct, the system may add complexity without addressing a demonstrated need. Meyer argues that a failing heuristic should be measured before teams build around a replacement.
Meyer’s Four-Part Fit Test
Meyer proposes four conditions for using Jev: high volume, a narrow question, errors that are cheap or can be routed for review, and a heuristic that has been shown to fail. He advises replaying 300 to 500 past decisions, comparing results overall and across confidence bands, then reviewing 20 disagreements to judge which answer was right. He says teams should wire a use in only if the high-confidence band reaches 95% in that evaluation.
For rollout, Meyer recommends giving each use a separate flag that is off by default, testing it on 5% to 10% of units, then expanding. His account of a 31-topic classification test says agreement with a frontier model was 97% to 99% at confidence of 0.8 or higher, and 42% below 0.5. These are results from his measurement; the article excerpt does not provide the test data, evaluation method or an independent replication.
The examples shown in the source focus on publishing. Meyer says 88% of the news items he processes start from a bare headline, which motivated a proposed thin-source detector. The detector would check whether an item contains enough verifiable facts to support a report. His proposal is to fetch the original source or skip an item at confidence below 0.2, and write at 0.8 or higher. The middle range is not assigned an automatic action in the excerpt.
“Jev is the right tool wherever a system needs thousands of small judgements and can hand the unclear ones to something smarter.”
— Thorsten Meyer
Evidence Still Needs Review
The figures in the source are Meyer’s own reported measurements. The material provided does not include the underlying data, details of the evaluation process, or independent testing. It is therefore unclear how well the results would carry over to other publishers, different content or other systems.
Seven proposed uses still need measurement, according to Meyer, because he has not established that the existing heuristic fails. His excerpt identifies three live uses and gives publishing examples, but does not detail the other listed applications across commerce, software, business operations and the home. The two poor fits are not identified in the supplied portion. The full breakdown and results for those cases remain unclear from the available material.
Measure Before Wider Rollout
Meyer’s proposed next step for each unproven use is to replay real past decisions and compare Jev’s answers against the existing process, including by confidence band. Teams would review disagreements, then test a qualifying use behind a separate flag on a small share of units. The stated rollout path is to expand after that canary; the source does not give dates for further releases or report a broader rollout beyond the three publishing uses.
Key Questions
What does Jev do?
According to Meyer, Jev takes text or JSON state and typed questions, then returns answers such as probabilities, category choices or scores for software to use.
How many uses does Meyer say are ready or running?
He reports three live uses and 12 strong fits. He separately classifies seven as needing measurement and two as poor fits.
What results does Meyer report from the article scan?
He says a scan of 78,889 articles cost $2.01, found 1,576 non-English articles and fixed 1,553. The source does not provide independent verification of those results.
How does Meyer recommend testing a proposed use?
He recommends replaying 300 to 500 past decisions, checking performance across confidence bands, and reviewing 20 disagreements before a limited canary rollout.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
