🔍 Read the full analysis: Could Jev Help With AI Decisions? 24 Ways To Use It on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
A September 29 guide from Thorsten Meyer maps 24 possible uses for Jev, a tool that returns typed answers to narrow questions so software can route or filter cases. The author says three uses are running in his publishing operation, 12 meet his proposed fit test, seven need measurement and two are poor fits. Those figures and performance results are the author’s reported findings; independent validation is not provided in the supplied material.
Thorsten Meyer published a 24-use guide to Jev on September 29, 2026, reporting that three applications are already running in his publishing operation and that a scan of nearly 79,000 articles cost $2.01. The guide argues that Jev is suited to high-volume, narrow decisions where uncertain cases can be sent elsewhere, while labeling another 12 uses strong fits, seven as needing measurement and two as poor fits.
Meyer describes Jev as a system that takes text or JSON state plus typed questions and returns answers software can act on, rather than writing or summarizing prose. The answer formats include a yes-or-no probability, a choice among options with probabilities and confidence, or a score on ordered levels. The guide says one call carries the state and questions, takes about 0.3 to 0.9 seconds, and costs about $0.04 per million input tokens; those figures are presented by the author, with no independent measurement details in the supplied material.
The three reported live uses are a relevance check matching stories to sites, an English-language check, and a fallback topic classifier. Meyer says the language check scanned 78,889 articles, found 1,576 non-English items and fixed 1,553 for $2.01. For the classifier, he reports 89% agreement with a frontier large language model overall, rising to 97%–99% when Jev’s confidence was at least 0.8. He says his measurement covered 31 topics; the guide does not provide the underlying dataset or an outside evaluation.
The proposed fit test requires high volume, a narrow question, low-cost errors or a route for uncertain answers, and evidence that an existing heuristic fails. Meyer recommends replaying 300 to 500 past decisions, comparing results by confidence band and reviewing 20 disagreements before deployment. He says to wire in the tool only where the high-confidence band reaches 95%, then use a separate flag, begin with 5%–10% of units, and roll out gradually.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
Where Automated Checks Could Save Time
The guide’s central point is that a cheap decision call can make it practical to check many items that might otherwise pass through an imperfect rule. In publishing, the examples include language, relevance, disclosure and comment checks. Meyer says the value depends on acting on clear answers while routing ambiguous cases to a person or a more capable system; the tool’s output alone does not determine what happens.
That distinction matters for readers and operators because the proposed uses include decisions with different consequences. A relevance mismatch may affect what gets published, while a missed affiliate disclosure can create a compliance concern. The guide advises sending disclosure misses to human review and treating headline scoring as a prompt rather than the only publication gate. Its own poor-fit example, event deduplication, reportedly found no duplicates in a canary, illustrating that low cost does not establish a problem worth automating.
From Three Live Checks to 24 Uses
The article organizes its 24 proposed applications across publishing, commerce, software, business operations and the home. The supplied excerpt details the three live publishing applications and begins a set of six publishing and content examples. It identifies a thin-source detector, product matching in roundups and headline quality as cases needing measurement, while disclosure detection and comment moderation are labeled strong fits. Same-event deduplication is labeled a poor fit after the author’s canary reportedly found no duplicates.
Meyer says a relevance test judged about 10,000 story-and-site pairings in three days, with 22% clearly on-topic. For thin-source detection, he says 88% of the news items he processes start from a bare headline. The article proposes testing such systems against real past decisions and retaining existing behavior for uncertain cases. The supplied material does not include the remaining use cases or the details of the full measurement process.
“Jev is the right tool wherever a system needs thousands of small judgements and can hand the unclear ones to something smarter.”
— Thorsten Meyer, in the September 29 guide
Evidence Behind the Reported Results
The performance and cost figures in the guide are author-reported. The supplied material does not include a dataset, evaluation protocol, confidence calibration method, or independent replication for the reported agreement rates. It also does not explain how the 95% threshold was selected or whether results hold across different topics and organizations.
The excerpt does not provide the remaining 15 use cases in detail, nor does it establish whether all 12 strong-fit applications have been deployed. It says three are live, 12 are strong fits, seven need measurement and two are poor fits. The boundary between a wrong answer that is cheap and one requiring escalation will depend on each operator’s use case.
Measure Before Expanding Deployment
The guide’s proposed next step for each unproven application is a shadow evaluation using 300 to 500 past decisions, followed by review of disagreements and performance across confidence bands. Meyer recommends enabling a separate feature flag only where the high-confidence band reaches 95%, then starting with 5%–10% of units and expanding gradually. The article does not announce a product release or a schedule for adding the other applications.
Key Questions
What does Jev do?
According to Meyer, Jev takes text or JSON state and typed questions, then returns structured answers such as probabilities, classifications or scores that software can use to branch. It is described as a decision tool rather than a prose-writing or summarization system.
How many uses does the guide say are ready or running?
Meyer says three uses are live in his publishing operation and 12 meet his strong-fit test. Seven need measurement and two are poor fits. The guide does not say that all 12 strong fits are deployed.
What results does Meyer report for the article scan?
He says Jev scanned 78,889 articles for $2.01, identified 1,576 non-English items and fixed 1,553. These are figures reported by the author; the supplied material does not include an independent audit.
When does the guide recommend routing a decision to a person?
The proposed approach is to act on clear, high-confidence answers and route uncertain cases to a person or a more capable system. The specific threshold and action should be checked against past decisions and the consequences of errors for that use case.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
