🔍 Read the full analysis: 24 Ways Jev Can Help Structure AI Decisions on ThorstenMeyerAI.com
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Thorsten Meyer published a 24-use-case map for Jev, a tool he says returns typed, confidence-scored answers that software can use to make narrow decisions. His account says three publishing checks are already running, 12 uses meet his fit test, seven need measurement, and two are poor fits; the available material details only the publishing examples.
Thorsten Meyer published a 24-use-case map for Jev on September 29, describing how the tool could handle small, repeated AI decisions across publishing, commerce, software, business operations and the home. Meyer says three publishing checks are already running in his operation, while 12 proposed uses meet his four-part fit test, seven need measurement and two are poor fits.
Meyer describes Jev as a system that receives text or JSON plus typed questions and returns answers that software can act on. The response types include a yes-or-no probability, a choice among options with probabilities and confidence, or a score on ordered levels. The tool does not write or summarize material, according to his description; the calling code sets the rule for what happens next.
Meyer says a call takes about 0.3 to 0.9 seconds and costs about $0.04 per million input tokens. Those are figures reported in his post, not independently verified here. He also reports that, in his measurement on a 31-topic classification, Jev agreed with a frontier LLM 97% to 99% of the time when confidence was at least 0.8, compared with 42% agreement below 0.5.
For a language check, Meyer says a scan of 78,889 articles cost $2.01 and found 1,576 non-English items, of which 1,553 were fixed. He reports that three uses are live in his publishing operation and that the fleet has handled about 90,000 decisions so far. The source material does not specify how that total is divided among the three uses.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
Where Confidence Changes the Workflow
The proposed approach is designed for high-volume, narrow decisions where a mistaken answer has limited cost or uncertain cases can be sent to a person or a more capable system. Meyer’s recommended pattern is to let software act on clear answers and preserve the existing path for uncertain ones. That can make checks practical across many items, while keeping the final action under the application’s control.
The examples also show why low cost alone is not a reason to automate. Meyer labels duplicate-event detection a poor fit after a canary found no duplicates. In his account, that result means the team has not shown a problem for the check to solve. For publishers and other operators, measuring existing errors before adding a new classifier may help avoid wiring in an unneeded system.
Some proposed uses also carry greater consequences than routine classification. Meyer identifies disclosure checks for affiliate links and provided products as a strong fit, but recommends routing missed disclosures to human review and never auto-publishing on Jev’s answer alone. The distinction matters: a confidence threshold can guide workflow, but the organization still defines acceptable risk and safeguards.
The Four Tests Behind Meyer’s Map
Meyer says a use should meet four conditions before teams build around Jev: it involves thousands of small calls, asks a narrow question without multi-step reasoning, has cheap errors or a route for uncertain cases, and addresses a heuristic that is visibly failing. He says the last condition must be demonstrated with evidence rather than assumed. If a keyword rule works, his advice is to keep it.
For a candidate use, Meyer recommends replaying 300 to 500 real past decisions, comparing results overall and by confidence band, and reviewing 20 disagreements to judge which answer was correct. He says to wire the tool in only where the high-confidence band reaches 95%. His rollout proposal is a separate feature flag that is off by default, followed by a canary on 5% to 10% of units.
The three live examples use that cautious pattern in different ways. A relevance gate checks whether a story suits a site and publishes the gray-area cases through the previous process. A language check flags items for rewriting in place when its probability falls below 0.1. A topic classifier acts as a fallback when the primary LLM fails, with a keyword rule as the last resort. Meyer reports about 10,000 relevance pairings judged in three days, with 22% clearly on-topic, and 89% agreement with a frontier LLM for the classifier overall.
The post also lists six publishing and content applications. Besides the three running checks, it describes a thin-source detector, product-to-roundup matching, disclosure review, headline-quality scoring and comment moderation. In the supplied material, thin-source detection, product matching and headline scoring are marked for measurement first; disclosure review and comment moderation are described as strong fits. The duplicate-event detector is marked a poor fit.
““Use Jev only when all four conditions hold.””
— Thorsten Meyer
Evidence Still Needed for Most Uses
The post’s supplied material names a total of 24 potential uses across five areas, but the excerpt provides details only for the production examples and publishing section. It does not identify the individual commerce, software, business-operations or home use cases, so their specific questions, rules and evidence cannot be assessed from this material.
The reported performance figures are Meyer’s own measurements. The material does not describe the full evaluation method, the composition of the comparison set, or how accuracy was judged across different confidence bands. It also does not provide independent verification of the scan’s cost or outcomes. Seven proposed uses still need a measurement of whether the existing heuristic fails, according to Meyer, and the supplied excerpt does not show which evidence would change each assessment.
It is also unclear whether the reported production results have been independently audited or how the system performs on topics outside the described classification test. Meyer’s stated thresholds and rollout steps are recommendations from his post; they are not universal performance guarantees.
Measure Before Expanding Jev
Meyer’s proposed next step for teams considering a use is to test it against historical decisions, review disagreements and confirm that high-confidence answers meet the stated 95% threshold. A separate flag and limited canary would then let a team observe behavior before expanding use. For cases that do not pass the fit test, the existing process remains the relevant comparison.
The next developments for the 24-use map are not stated in the supplied material. Meyer does not give launch dates for the uses marked strong fit or measure first, nor does he set out a schedule for publishing results from the remaining categories. Further detail on those use cases, and evidence from their measurement or deployment, would clarify how far the framework extends beyond his publishing operation.
Key Questions
What is Jev?
Meyer describes Jev as a tool that takes text or JSON and typed questions, then returns probabilities, choices or scores with confidence information. Software determines what action follows.
How many Jev uses does Meyer say are ready or running?
Meyer says three uses are live in his publishing operation and 12 meet all four conditions in his fit test. He classifies seven as needing measurement and two as poor fits. The supplied material does not detail every use in the 24-item map.
What does the fit test require?
The proposed use should involve high volume, ask a narrow question, have low-cost errors or a path for uncertain cases, and address a heuristic that has been shown to fail.
What did the reported article scan find?
Meyer says a scan of 78,889 articles cost $2.01, identified 1,576 non-English articles and led to 1,553 being fixed. These are figures reported in his post.
Does Meyer recommend automatically acting on every Jev answer?
No. His proposed pattern is to act on clear cases and route uncertain ones to a person or another system. For disclosure checks, he specifically recommends human review of misses and says they should never be auto-published.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
