Some of Your AI Calls Were Never Conversations: What Typed-Decision Models Like Jev Do, and Don't, Change in the Automation Stack
Introduction
On 15 September 2026, the startup TypeSafe AI launched Jev, a typed-decision model that picks answers from a list its developer writes and attaches a probability to each.
It arrived with an announced $40 million seed round led by the venture firm DCVC and a homepage promising "193.6x Faster, 444.6x Cheaper", footnoted as "based on workflows for System One tasks", and "Zero Hallucinations". Its chief executive, Diogo Almeida, is a primary author of OpenAI's InstructGPT paper, the 2022 study behind ChatGPT's training methods. A day later the trade site AI News reported speeds "up to 193.6 times faster", measured "against consensus baselines from GPT-6 Astra and Fable 5.1", models from the AI labs OpenAI and Anthropic. Within 24 hours Jev had reached nearly 13% of paid teams on the gateway of Vercel, a cloud platform that resells model access, during a free promotion running to 25 September, which left retention as Vercel's "next test".
For an operator running classification, routing or screening through general-purpose models, Jev is a cheaper, faster engine for one class of call. The stakes sit in three questions: how many calls fit that class, what the honest comparison is once vendor benchmarks are set aside, and when a probability rather than a person should decide that an answer is safe to act on.
What Jev Returns: A Pick From Your List, With Probabilities Attached
Choice, Score and Noul: Three Ways to Ask One Question of One Piece of Text
TypeSafe calls Jev a "System One model", after Daniel Kahneman's System 1, and the developer fixes the answer space in advance through three primitives.
**Choice:** the model picks one of up to 255 options and returns a probability for each, plus a summary `confidence` statistic.
**Score:** the model places the input on an ordered rubric, such as a severity scale, with a probability for each level.
**Noul:** short for Bernoulli, the model returns one probability that a stated proposition is true.
Many questions can run against one shared `state` per request, with input limited to text of up to 64,000 tokens. Jev returns no strings at all, a restriction that, Almeida says, lets every output be computed in parallel with no output-token cost, and the same weights serve every customer without fine-tuning.
When a commenter on Hacker News, the technology forum, called Jev "basically a zero-shot classifier", Almeida replied "exactly right". A zero-shot classifier sorts text into categories it was never specifically trained on, and Jev works as a hosted one.
The mechanism underneath is undisclosed. Almeida calls the architecture "close to the chest for now", the training method TypeSafe calls RLCD has no published paper according to his interview with the AI engineering podcast Latent Space, and TypeSafe withholds public-benchmark results. The technology site TechCrunch reports that "outside observers suspect" an open-weight language model underneath, a reading TypeSafe rejects and nothing on the record settles, so an operator's own evaluation is the only assurance available.
The Nine Things TypeSafe Says Jev Gets Wrong
TypeSafe's own jaggedness page, reviewed on 17 September 2026, lists nine failure modes, including literal reading, arithmetic and counting, date comparison, adversarial content and generation. For extraction it advises finding candidates with a regex or a generative model and letting Jev choose, so an invoice number still needs an upstream step and date or arithmetic logic belongs in code.
Schema-Valid Is Not Correct, and LLMs Have Been Schema-Valid Since August 2024
Answering its own launch promise that Jev "can't hallucinate", TypeSafe's FAQ says Jev "guarantees the shape of its answers, not that every decision is correct". The community audit scienthoon recorded no invalid responses in 4,621 calls but also a benchmark question on which Jev put probability 0.00 on the right answer and 1.00 on a wrong one, so "can't hallucinate" means schema-valid, not correct.
Language models have had the same schema guarantee since 6 August 2024, when OpenAI introduced structured outputs, a form of constrained decoding that blocks any token breaking the developer's JSON schema. OpenAI reported 100% adherence for gpt-4o-2024-08-06 against under 40% for the older gpt-4-0613, and Anthropic, the developer of Claude, offers the same constrained-decoding guarantee from Haiku 4.5 to Fable 5.1, with documented exceptions such as refusals and the capitalization of enum values.
What constrained decoding does not supply is a probability per option, calibration, or lower cost and latency, since the model still generates token by token. Log-probabilities can supply class probabilities, as OpenAI's cookbook documents, which narrows Jev's real differences to native per-option probabilities, parallel questions over one input, input-only billing and a sub-second floor.
Whether the constraint costs accuracy is disputed, with vendors on both sides. Almeida says constrained decoding makes models "dumber" and a 2024 preprint by Tam and colleagues found reasoning declined under format restrictions, while .txt, maintainer of the Outlines structured-generation library, found "an improvement across the board" on a re-run. The narrow reading is that constraining reasoning may hurt and constraining a final label has not been shown to, a question each operator settles on labeled data.
That retires one cost the launch charged to language-model decision calls, the tax of parsing and policing free text, for operators on structured outputs. Price and latency remain, weighted by how many calls are decisions.
How Many of Your Calls Are Decisions? No Published Dataset Counts Them
Using a large model as a one-label classifier happens inside a model developer's own research, since Anthropic's Economic Index, its published series on Claude usage, prompts Claude for "exactly one of the answer options". Armin Ronacher, chief technology officer of Earendil, told TechCrunch: "We should have seen this earlier in many ways, but presumably because the LLMs are so cheap and subsidized, you often don't have to be creative yet." Neither example measures how common the pattern is.
The best usage data measures adjacent things: on Claude's first-party API, 77% of transcripts showed automation patterns in August 2025, which is not the same as bounded labels. The Index's June 2026 output classifier has no label or decision category, so its 16% "no clear output" share is not a decision share. A token study by the model marketplace OpenRouter and the venture firm a16z classifies prompt topics and weights by tokens, which under-counts short calls.
No public figure exists, and the number lives in each operator's logs as the share of calls, by count, whose output is one of a fixed set of values. TypeSafe predicts a Jevons effect in which falling cost opens up "orders of magnitude more use cases", but Anthropic's preliminary estimate for Claude's API is that a 10% cost cut for a task goes with only about 3% more usage. Price and latency then decide whether moving that share pays, which is what the launch figures claim to settle.
193.6x Faster Than What?
How TypeSafe Built the Headline Multipliers
The multipliers come from four workflows written by TypeSafe's model capabilities team. They were scored against a reference answer averaged from GPT-6 Astra and Fable 5.1 at high reasoning settings, so the reported "baselines" defined what counted as right and were not the models being timed. The timed rivals ran at default settings inside a TypeSafe wrapper that the company concedes "tends to be slower and more expensive".
TypeSafe names no comparator for either multiplier. On its own figures, by this article's arithmetic, 193.6x fits Jev's time per case against Claude Sonnet 5, which matched Jev's 67.8% agreement with the consensus, and 444.6x fits its cost against Claude Opus 5, which agreed 5.3 points more often; other pairings in the rounded data also fit, so the attribution is an inference, not TypeSafe's. On invoice processing, the most finance-like workflow, Jev's agreement trailed OpenAI's sol configuration by 17.3 points, and TypeSafe itself expects its multipliers to be "on the higher end of real world gains".
TypeSafe argues consensus scoring ["likely underestimate[s]"](https://typesafe.ai/blog/introducing-system-one-models-and-jev) Jev, while the automation agency ayautomate, which ran its own labeled test, found agreement running 6 to 8 points above accuracy, and with no human re-scoring of the workflows the net bias is unknown.
Measured Independently: About 2–5x Faster and 5–50x Cheaper per Decision
Measuring between 18 and 22 September 2026, four independent testers, ayautomate, jujumilk3, AnthusAI and jevbench, recorded median latencies of 0.24 to 0.39 seconds, which confirms the Jev half of the speed claim. The rival "3 to 329 seconds" came from a mirror of the Artificial Analysis leaderboard, a third-party benchmark of time to first token that includes maximum-reasoning modes. LLMs configured for a decision left Jev only 2.0 to 3.6 times faster.
Jev lists at $0.042 per million input tokens, with output free, as of 23 September 2026, but because tokenizers differ, billed cost per decision is the figure that matters, and on that measure ayautomate found Jev roughly 5 to 50 times cheaper than the models it tested.

Both sets of multipliers can be true, because the headline compares Jev with slow, costly configurations an operator would rarely deploy for a decision layer. Whether even the smaller gap lasts is unclear, since TypeSafe has said both "We can't prove it isn't subsidized" and "We can serve Jev profitably at our current prices".
As Accurate as a Small LLM, Not as a Classifier Trained on Your Labels
Cheaper answers matter only at acceptable accuracy, and Jev's rivals include much older classifiers. Zero-shot classifiers built on natural-language inference reuse a model trained to judge whether one sentence implies another. A fine-tuned small encoder such as DistilBERT learns a fixed taxonomy from labeled examples, in about ten minutes in one test, then runs locally at almost no cost per call.
In ayautomate's test, Jev beat OpenAI's GPT-5.4 nano on prompt injection but trailed it on confusable eight-way intents, and finished about five points behind the larger GPT-5.6 Terra on 77-way routing. In jevbench, a community benchmark, fine-tuned DistilBERT beat Jev on topic and intent tasks, most sharply on Banking77's 77 banking intents, at 88.0% against 76.4%. All of these tests ran on jev-1.13.0 between 19 and 22 September, on small samples and without peer review.
The literature agrees: Bucher and Martini found fine-tuned small models "consistently and significantly outperform" larger zero-shot ones, and Pecher and colleagues put the break-even at about 100 labeled examples on average, varying by task.

Where Jev Earns Its Place: No Labels Yet, or Categories That Keep Changing
With no labels at all, a community comparison by zhuyansen found that Jev beat clean zero-shot BERT on all seven sets and was worth about 230 labels of a trained model on AG News and Banking77, and more than 2,048 on three other tasks. That supports a sequence: start zero-shot, with Jev or a small model under structured outputs, while a taxonomy is new or shifting, and train a classifier once labels justify it.
Vercel, a reseller of both Jev and OpenAI's models, sets the limit: "If you already classify tickets with Astra, a typed Jev answer alone isn't a reason to migrate."
A Probability You Can Rank On, Not One You Can Use Raw
Once accuracy proves par, what remains distinctive is the probability on each answer. A model is calibrated when outcomes given probability *p* occur about *p* of the time, a group property that, as TypeSafe notes, guarantees nothing about one answer and is usually summarized as expected calibration error (ECE).
GPT-4's technical report documents the problem in language models, with post-training raising ECE on an MMLU knowledge-benchmark subset from 0.007 to 0.074, and verbalized confidence runs overconfident, though one study found it better calibrated than the token probabilities of models tuned with human feedback. Calibration is also routinely repaired after the fact, as temperature scaling refits probabilities on a small labeled set, so Jev's contribution is cheap native probabilities to refit.
TypeSafe claims more, calling Jev "Calibrated", while its cookbook calls a threshold band "neither a calibrated guarantee nor an optimized threshold" and Almeida, asked on Latent Space whether he was claiming Jev's calibration is perfect, replied "I didn't say that".
What Three Independent Audits Found in Jev's Probabilities
Three community audits on GitHub, using different data, agree on direction but not on a precise number. AnthusAI tested 8,801 labeled sentiment examples and, on a 1,000-example subset scored for both models, found strong ranking, with an area-under-curve score (AUROC), the chance a right answer outscores a wrong one, of 0.83 against 0.72 for Llama 3.1 8B. Its raw ECE of 0.064 to 0.160 fell to 0.006 to 0.018 after isotonic recalibration, with a few hundred labels capturing most of the benefit.
Removing the abstain option pushed ECE from 0.023 to 0.793, and where the label hinged on an unstated rule, Jev was right 44.7% of the time at a mean stated probability of 0.74. Probabilities were not coherent either, with TypeSafe's own "refund" and "not refund" Nouls summing to 1.19 and audited complementary pairs summing to 0.71 to 1.42, and 50 identical requests returned 15 distinct answers.
Why the `confidence` Field Is Not the Chance of Being Right
The `confidence` field on Choice answers invites the wrong use. It measures how concentrated the distribution is, and AnthusAI found that it equals twice the top probability minus one on a two-option Choice, and that Choice answers were right only 50% to 57% of the time whenever the chosen option's probability fell between 50% and 95%.
Taken together, the audits advise recalibrating each question on a few hundred of the operator's labels, including an "other" or "unknown" option, and never thresholding on the `confidence` field. By this article's own inference, not any source's, those few hundred labels are roughly what a fine-tuned classifier needs on a stable taxonomy, so calibration work doubles as a training set.
Gate and Escalate: Frontier Accuracy at a Quarter of the Cost, in One Test
A score that ranks well enables selective classification, which Geifman and El-Yaniv extended to deep networks in 2017, answering only above a threshold set for a target error rate. In a cascade the rest go to a larger model, a pattern TypeSafe's documentation describes as act, confirm and escalate bands.
In ayautomate's test on 19 September 2026, Jev answered routing items when its `confidence` value reached 0.80, and the rest went to GPT-5.6 Terra. The cascade matched Terra's accuracy at roughly a quarter of its cost, but the gate was Jev's raw `confidence` field, and thresholds were scored on the same items, so ayautomate says to "treat the exact figures as an estimate". This article's recommendation, a threshold chosen on a separate validation set using recalibrated probabilities, is stricter than what the test did.

Confident errors still got through: 5 of 112 answers accepted at 0.90 or above were wrong on eight-way routing, all one confusion between two overlapping labels, and 12 of 153 were wrong on 77-way. As ayautomate puts it, "The confidence score cannot tell you that your label set overlaps." Governance therefore sits in label design and, as Vercel advises, in auditing accepted answers as well as rejected ones. Ronacher warns the design "delegates the hallucination problem a little bit to the user", and Almeida, sketching "calibration plus a cascade of models" on Latent Space, added: "I don't really know how that's gonna go".
The Pattern Is Portable; the Vendor Is One Week Old
The cascade's logic does not depend on Jev, and TypeSafe's vendor-reported evaluation found "every model is more accurate, cheaper and faster" when a policy is split into typed questions rather than written as one prompt. Within a week, Vercel's AI SDK offered the same typed questions across TypeSafe, OpenAI, Anthropic and Google models through an evaluate function, OpenRouter opened an alpha Decisions API, and six open clones appeared within two days. The large labs had shipped no decision-only model by 23 September, offering cheap generative tiers such as gpt-6-luna instead.
That portability matters because TypeSafe is "not promising long-term support", its rate limits "can change without notice", it briefly could not serve users under launch demand, and its model lost 6.5 points on Korean in one audit.
Audit, Test, Gate, Pin: Sizing Your Own Decision Layer
The evidence reduces to four steps.
**Audit:** count the decision calls in the operator's own logs, with billed cost and latency per call.
**Test:** on a labeled sample of real traffic, compare Jev, a small model with structured outputs and, where labels exist for a stable taxonomy, a fine-tuned encoder, with an "other" option in every answer set. Cost is measured the same way, per billed decision, and whether gpt-6-luna's $0.01 per million cached input tokens undercuts Jev on long static prompts remains untested.
**Gate:** anything that acts without a person should clear a threshold on a recalibrated probability, with escalation below it and sampled review above it.
**Pin:** fix the version ID, keep a fallback route and re-validate thresholds on every release.

Conclusion
The evidence supports a modest claim: some of an operator's calls are decisions, and only their own logs can say how many. For those calls, jev-1.13.0, one week after launch, is schema-valid, sub-second, roughly 5 to 50 times cheaper per billed decision than the models tested against it, and about as accurate as a small language model. Its scores rank answers well enough that, in one test, a cascade gated on its raw confidence value matched a frontier model at about a quarter of the cost.
The evidence will not carry 193.6x as an expectation, the parsing tax as a saving unique to Jev, calibration out of the box, an edge over a classifier trained on an operator's labels, or "can't hallucinate" beyond schema validity. What looks durable is architectural, namely narrow typed questions, logic kept in code and a recalibrated probability deciding what escalates, and as of 23 September 2026 Jev is the cheapest hosted first rung measured so far, not the pattern itself.