Skip to Content

Jev’s Paradox: The hidden cost of cheap AI decisions

This article was first published on LinkedIn.

Nash Borges

Consider a cybersecurity alert for a potentially malicious PowerShell script that downloads and runs a file. An AI agent in a security operations center (SOC) can handle it by picking from a fixed set of actions: close the alert as benign, collect more evidence, or send it to an analyst. Before letting the agent act on such alerts, the SOC needs to know how often it picks correctly and whether its confidence helps identify its mistakes.

TypeSafe.ai built Jev for this kind of decision-making work. A developer supplies structured context, a question, and the permitted answers, and Jev returns probabilities over those answers without generating a long-form text response. TypeSafe pitches that design as much faster and cheaper per decision than a general-purpose LLM. In a Latent Space interview, TypeSafe cofounder and CEO Diogo Almeida described a class of models where “the goal is for code to be the consumer.” He wants these judgments to become reliable enough that developers can call them as routinely as a database query. TypeSafe named Jev after Jevons paradox, in which greater efficiency stimulates enough additional demand to increase total resource consumption. That points to Jev’s paradox. If Jev makes automated decisions cheap enough to run on everything, teams will automate far more decisions, and any error rate above the current baseline, human or automated, will add up to more total mistakes. 

Always returning a valid answer makes Jev easy to integrate. How often it chooses the correct answer determines whether it is useful. If it also knows when a proposed answer is unreliable, that can influence how much work can be safely automated. Those are three separate properties, and the attention around Jev’s launch has blurred them a bit.

Classifiers that accept their labels at runtime are not new. Since at least 2019, natural-language inference models have handled zero-shot classification by checking whether an input supports a statement such as “this alert is malicious” without needing to retrain the model. A March 2026 benchmark compared these models with embedding models, rerankers, and instruction-tuned LLMs. TypeSafe has not published Jev’s architecture, so it is unclear where Jev fits among them, though its stated training goal of calibrated decisions resembles published research discussed below. Regardless, novelty should matter less to a SOC than whether Jev improves its workflow, and early independent tests published after the launch are worth a closer look.

Jev’s advantages depend on what it replaces

One independent evaluation compared Jev to OpenAI’s models on two intent-classification benchmarks. The evaluation tested 200 requests from CLINC150, a benchmark for identifying what a user is asking a chatbot to do. Jev reached 87% accuracy with no training examples, compared with 80% for the small gpt-5.4-nano model and 92% for GPT-5.6 Terra. Terra sits below GPT-5.6 Sol and GPT-6 Astra in OpenAI’s lineup, so Jev is a decent zero-shot classifier, but not frontier model quality.

The second comparison tested 208 requests from Banking77, a benchmark for classifying banking customer-support requests, and added a trained classifier to the mix using the bge-small embedding model with logistic regression. Without any training examples, Jev reached 83%, while the supervised encoder reached 93%. The trained classifier learned the dataset’s idiosyncrasies and labeling conventions from thousands of examples, while Jev did not perform as well from its pre-training alone. So an ML team with training and evaluation data should still start with a supervised classifier before trusting zero-shot models with important business decisions. On the other hand, if your questions change often, or you have no labeled data, then your options are more limited. Both Jev and frontier models are quite accurate on many zero-shot tasks, but accuracy isn’t everything.

Speed and cost matter too. On Banking77, the same evaluation recorded median response times of 0.44 seconds for Jev, more than 3x faster than GPT-5.6 Terra at 1.51 seconds in the tested API configurations. A separate phishing benchmark found Jev about 12 times cheaper per email than Claude Haiku 4.5 at list prices. Lower costs per request and faster responses can make several use-cases more affordable. That is the demand TypeSafe is betting on, although more model responses can compound mistakes depending on how the answers are combined. A supervised classifier running locally can also be cheaper, faster, and more accurate. The catch is that it answers only one question and needs labeled data and expertise to build and maintain.

Confidence trained against outcomes

When a weather app says there is a 30% chance of rain, the forecast is not wrong if it rains. It is supposed to rain on roughly three of every ten days that receive that forecast. That property is called calibration, and weather forecasters have measured it for decades by comparing stated probabilities with observed outcomes and correcting their models when the two drift apart. Hamill et al. showed how those corrections substantially improved the reliability of precipitation guidance for the U.S. National Blend of Models, which adopted the method operationally in 2017. LLMs rarely get checked against observed outcomes in the same way.

Jev answers in three fixed formats. A Choice picks one option from a list, such as the triage decision from the opening. A Score places something on a scale of levels you describe, such as an incident’s severity from informational to critical. A Noul, short for Bernoulli, gives the probability that something is true, such as whether an alert is malicious. For Choice and Score answers, Jev returns probabilities over the permitted options or levels and a separate confidence statistic summarizing how concentrated those probabilities are. A Noul’s single probability plays the same role. A value of 1.0 is the top of the confidence statistic’s scale, but it only represents the model’s certainty. Its relationship to correctness has to be measured. In the evaluation described above, Jev’s confidence was exactly 1.0 on 102 of 200 CLINC150 items, and six of those answers were wrong.

TypeSafe has not published how it trains Jev’s probabilities. It calls its approach reinforcement learning for calibrated decisions (RLCD), and Almeida described RLCD in the interview as a new training goal, without specifying an algorithm. Academic research shows one way to pursue a similar goal. 

Fourteen months before Jev’s launch, in July 2025, Damani et al. at MIT introduced reinforcement learning with calibration rewards (RLCR), building on earlier work such as SaySelf. Under RLCR’s correctness-and-confidence reward, each answer earns a correctness reward minus a penalty for the gap between its confidence and the outcome. The penalty grows with the square of that gap, so overconfident mistakes cost much more than cautious ones. This correctness-and-confidence reward ranges from −1 to 1, with higher values being better. The experiments also used a separate output formatting reward not represented in Table 1.

Calibration rewards add feedback about the confidence attached to an answer. Two incorrect answers can receive different rewards if one expresses much greater confidence than the other. Training can then discourage overconfident mistakes while continuing to reward correct answers. 

As another point of comparison, reinforcement learning with verifiable rewards (RLVR) scores an answer against a known outcome, typically giving one point for a correct answer and zero for an incorrect one. That reward encourages accuracy while leaving any reported confidence unscored.

Take the PowerShell alert from the opening, and imagine a model classifying the activity as malicious or benign. Suppose prior investigations have verified labels for training. Table 1 shows illustrative responses, confidence estimates, and different reward signals. Correctness-only rewards like RLVR treat both mistakes equally and both correct answers equally, while RLCR gives the cautious mistake a higher reward than the confident mistake. It also gives the confident correct answer a higher reward than the hesitant correct answer.

 

1790648063687.png
Table 1: Illustrative rewards for four answers to the PowerShell alert under correctness-only (RLVR) and calibration (RLCR) training.

The grader scores each answer and its confidence against the verified outcome. Training uses these rewards to update the model’s parameters. Across many training examples, those differences can help the model learn which evidence supports stronger confidence and which leaves room for doubt.

The confidence penalty in the example comes from the Brier score, introduced for weather forecasting in 1950. Brier was concerned that forecasters would game the system, reporting whatever earned the best score even when it differed from what they believed. His scoring rule encouraged realistic probability estimates by penalizing the squared gap between predicted probabilities and observed outcomes.

Damani et al. trained RLCR on HotpotQA, a benchmark that requires combining information from multiple Wikipedia pages, with some supporting paragraphs removed from the training examples. They compared it with correctness-only reinforcement learning. Tested on HotpotQA questions with both supporting paragraphs present, expected calibration error fell from 37% with correctness-only training to 3% with RLCR, while answer accuracy stayed around 62-63%. Confidence became far more informative while the models answered about the same share of questions correctly.

The improvement only partly transferred to other domains like factual recall, mathematics, commonsense, and science related questions. Across those datasets, its average calibration error rose to 21%, while it was still lower than the baselines on average. 

Independent tests of Jev show a similar pattern. An out-of-distribution calibration audit found that Jev’s probabilities needed little correction on three public benchmarks it may have seen during training. Another test used synthetic customer-support tickets that it could not have trained on. Jev was asked which department should handle each ticket, how urgent it was, and whether the customer sounded angry. For 300 urgency questions, the correct answer depended on an organizational rule Jev was not given. It answered only 45% correctly, yet assigned its chosen answers an average probability of 74%.

Confidence also varied by task. Jev generally overstated its certainty about routing and urgency, but understated it when identifying angry customers. Like any pre-trained system, probabilities should be validated for any task with serious implications.

Confidence as a decision threshold

Calibrated confidence becomes operationally interesting when it connects to actions with different costs. An agentic SOC has to decide whether an automated investigation can be safely closed or should go to a human. Closing a case of a real intrusion risks missing an attack. Escalating everything creates alert fatigue and wastes analyst attention. Making that choice requires three separate capabilities.

Calibration: does the number mean what it says? For illustration, suppose an agent investigates an alert and concludes, “This is benign. I’m 99.9% confident.” If that confidence is calibrated, the benign conclusion should be correct about 99.9% of the time across comparable cases receiving that score. That frequency provides a risk estimate for a comparable alert, although it cannot guarantee the outcome of any individual case. Whether the remaining risk is acceptable for automatic closure is a separate decision.

Discrimination: can it distinguish its stronger conclusions from its weaker ones? Imagine a model that gets 90 of 100 decisions right and reports 90% confidence every time. It is well calibrated overall, but its confidence provides no help identifying the 10 mistakes that could benefit from further review. To support selective automation, confidence should tend to be higher for conclusions that turn out correct and lower for conclusions that turn out wrong. The separation will not be perfect, but it should be useful.

Decision policy: what should the system do with that information? Even a well-calibrated model with useful discrimination does not determine what level of risk an organization is willing to accept. Strong evidence of intrusion may warrant escalation or an automated response. Strong evidence of benign activity may make an investigation eligible for automatic closure. Insufficient evidence may call for further investigation or human review. The closure threshold depends on the consequences of missing an intrusion, the cost of review, how critical the affected system is, and whether other controls would still catch an attack. The policy may allow the agent to recommend closure for an analyst to review while requiring stronger evidence before it can close an investigation automatically. In Sophos MDR, AI agents close cases end-to-end only inside boundaries that analysts set, and higher-stakes decisions go to a human-in-the-loop before a risky action is taken.

The CLINC150 evaluation tested exactly this kind of policy. Each request went first to a cheaper model, either Jev or gpt-5.4-nano, and any answer below a confidence cutoff was escalated to the more expensive GPT-5.6 Terra. The question was how much traffic still had to reach Terra. When the goal was to come within one percentage point of Terra’s accuracy, Jev-first routing escalated 22% of requests and nano-first routing escalated 49%. When the goal was to match Terra exactly, the result flipped. Jev-first routing had to send every request to Terra, while nano-first routing sent 73%.

The flip came from Jev’s maximum-confidence answers. Jev gave 102 of the 200 answers a confidence of exactly 1.0, and six of those were wrong. A cutoff can only separate answers with different scores, so recovering the one of those six mistakes that Terra got right meant escalating all 102, and by then the cheaper first stage was doing no useful work. The results were also fragile. Cutoffs chosen on half the data fell short of the one-point target on the other half, and neither model’s confidence proved consistently better at flagging its own mistakes across the two benchmarks. One promising result is too thin a basis for a production cutoff. The threshold has to be chosen on held-out cases from the environment where it will run.

Smaller questions and trained decision rules

A public phishing benchmark provides another comparison. Its author asked Jev and Claude Haiku 4.5 whether an email agent should click the link in each of 2,000 emails, whose synthetic bodies were built around real phishing and legitimate URLs. Asked directly, Jev reached 63% accuracy and caught 43% of phishing emails, while Haiku reached 81% and caught 76%. 

In a separate experiment within the same study, the author used Jev to generate features for another classifier. Jev supplied probability estimates for five narrower questions, including whether a link pointed to a URL shortener or free hosting platform. A separate logistic regression classifier learned from those features and the labels of 1,000 emails, then reached 95% accuracy on the remaining 1,000. 

The second experiment combined dataset-informed questions with supervised training: the author designed the questions after studying the dataset’s catalog of URL evasion techniques, and the separate classifier learned from labeled examples. Its accuracy was measured on the held-out half, whereas the direct verdicts were scored on all 2,000 emails. The experiment supports Almeida’s advice to break decisions into small, testable questions, although it does not show that decomposition alone produced the gain. It suggests testing model answers as features beside conventional rules and classifiers, then letting measured performance decide which components belong in the workflow.

Validation on your own alerts

Weather forecasts are useful because forecasters spent decades building a discipline of evaluation. They plot predicted probabilities against observed frequencies, identify biases, and adjust. Their trust comes from measuring each model against reality, however sophisticated the model is, and they keep measuring.

For a security team, that means evaluating Jev on their own alerts compared to rules, trained classifiers, and generative models. Choose confidence thresholds on a separate calibration set for each question, then measure missed intrusions, analyst workload, and total cost on an untouched test set, including retries, extra model requests, and human review. Statistical methods such as selective prediction can turn that calibration set into a bound on the error rate among automated decisions. Repeat the checks across customers, time periods, and new attack techniques. Test whether reordering the answer choices or adding an irrelevant option changes the probabilities enough to alter an action. Language models are known to favor certain answer positions and labels, and some have observed both kinds of sensitivity in Jev’s outputs. Include adversarial text in alert fields, because attackers can often write part of the evidence the model reads. Use Jev wherever those measurements show it improves the workflow. The cheaper each judgment gets, the more of the workflow a team can afford to automate. But always remember, cheaper decisions are only a bargain if the mistakes don’t cost more than the automation saves.

Research and blogging assistance provided by Hermes, an AI Assistant. It gives this quip a 30% chance of landing and asks to be judged across a larger sample.