← back to the log

Jev: What It Is and Where It Applies

Written by
Jonathan Major · Founder, Agentic Studio Labs

Senior AI engineer focused on production agent workflows, secure tool integrations, and decision systems with practical human-in-the-loop controls. More background on the about page.

What Jev is

Jev is a decision model, not a text model: you send it a state plus typed questions, and it returns typed answers with probabilities that your code can branch on directly. TypeSafe AI calls this class a System One model, after Kahneman's fast, intuitive mode of thinking (see the next section), in contrast to the slow, deliberate generation of an LLM.

It launched on September 15, 2026 from TypeSafe AI, founded by Diogo Almeida, a former OpenAI researcher and co-inventor of RLHF and InstructGPT, with a $40M round led by DCVC. Early access was waitlisted for the first five days; on September 20, 2026 TypeSafe opened it to everyone with no waitlist. The name references William Stanley Jevons and the Jevons paradox: cheaper tokens will mean more tokens consumed, not fewer.

The simplest mental model

An LLM is a writer, Jev is a judge. It answers the kind of gut-check question a knowledgeable person could settle in a few seconds given the right context, and nothing more.

System 1, System 2, and where the models sit

The names come from Daniel Kahneman's Thinking, Fast and Slow (2011): System 1 is the fast, automatic, intuitive mode (recognizing a face, finishing "tips and...", reading a stop sign), System 2 is the slow, effortful, deliberate mode (multiplying 343 by 57, weighing a contract clause). There is no System 3 in Kahneman's framework; the numbers stop at two, and "System One model" is TypeSafe's coinage for a model class, not a step in a series.

Kahneman mode Model class Trained with Returns Examples
System 1: fast, intuitive, calibrated gut-check System One model Reinforcement learning for calibrated decisions (RLCD) Typed answers plus probabilities and confidence, in one pass Jev
Between the two: fluent, general, conversational Chat LLM Next-token prediction plus RLHF Free-form text, token by token GPT-5.6 Terra, Claude Sonnet 5
System 2: slow, deliberate, multi-step Reasoning LLM RLHF plus reinforcement learning on verifiable rewards (RLVR) Text after an extended thinking trace GPT-6 Astra, Claude Fable 5.1

The useful reading is that an agent needs both modes, the way a person does. Most steps in a loop are System 1 calls (is this relevant, which tool, how urgent, is this safe), and a few are System 2 (write the fix, plan the migration, explain the tradeoff). Chat LLMs have been doing both jobs; Jev exists to take the first kind back at System 1 speed and cost, leaving the reasoning model for the second.

One caveat on the analogy: Kahneman's System 1 is prone to bias precisely because it is fast, and Jev's failure modes (literal reading, no arithmetic, distractible by irrelevant state) are the machine version of the same tradeoff. Calibrated confidence is what makes it safe to build on: the model can say how sure it is, so code knows when to slow down and escalate.

How it works

One request carries a state and any number of typed questions; Jev ingests the state once and evaluates every question against it in parallel, in isolation, in a single forward pass. Adding questions barely changes response time, and because each question is evaluated independently there is no context rot between them.

state + questions
↓ one request
Jev: evaluate each question in parallel
↓ one response
typed answers + probabilities + confidence
↓
your code: branch, sort, route

Code stays in control; Jev supplies narrow judgments inside the loop.

The three primitives

Question type Asks Returns
Choice Pick one option from a defined set (up to 255) choice, a probability per option, confidence
Score Rate the state against ordered, descriptive levels (2 to 10) score, a probability per level, confidence
Noul Is this statement true? noul, a 0 to 1 probability

All three can be mixed in one call. Confidence is a second axis distinct from probability: the answer says what, confidence says whether to act on it.

Training

Jev is trained with Reinforcement Learning for Calibrated Decisions (RLCD) rather than next-token prediction, so the probabilities it returns are meant to be calibrated, not just ranked. It is not fine-tuned per customer; you shape it through the state (your records and reference material) and each question's instructions and criteria, which accept JSON structure. Same weights serve every account.

Limits (Jev 1.13, jev-1.13.0)

Item Value
Context 64k tokens per request; 32k for state plus the longest question
Input Text only (string, JSON object, or array); pre-process images, audio, binaries to text
Rate limits 250,000 tokens/sec, 1,200 requests/min, adjusting dynamically
Language English primary; other languages handled but less accurately
Aliases jev-latest, jev-preview (both point to 1.13.0 today); pin the versioned ID if you tune confidence thresholds

Source: TypeSafe docs, Models and Introduction.

Speed and cost

Jev costs $0.042 per million input tokens with output free, and answers in roughly 70 to 500 ms regardless of how many questions ride in the request. That is the whole pitch: decisions that used to cost an LLM round trip now cost about as much as a database query.

Comparison Figure Who reports it
Input price $0.042 / Mtok (a Btok is $42), output $0 TypeSafe Models page
Latency, demo on typesafe.ai 0.114 s vs 8.566 s for GPT-5.6 Terra The Register
Cost vs Claude Fable 5.1 238x cheaper The Register
Workflow evals vs LLMs up to 193.6x faster, 444.6x cheaper Vercel changelog, citing TypeSafe
Batching 13 questions in one call vs 13 calls 12.2x cheaper, 10.0x faster, same answers TypeSafe cookbook

Read the multipliers with care: they are TypeSafe's own numbers on classification and routing tasks, where an LLM is overkill by construction. On anything that needs generation or multi-step reasoning Jev does not compete, it is simply the wrong tool. The honest framing is that it collapses the cost of the 80 percent of agent steps that are yes/no or pick-one, and leaves the LLM for the rest.

One more cost lever: rate limits (250k tokens/sec, 1,200 req/min) are being adjusted dynamically during launch demand, so a production integration should assume 429s and retry with backoff.

Design patterns and failure modes

The skill with Jev is decomposition: break a judgment into atomic questions, ask them all in one call, and reassemble the answer in code you control. TypeSafe documents four patterns that cover most systems.

Pattern What it does Why
Speculative fan-out Pack many questions, including ones you may not need, into one call; code decides what is relevant Adding questions is nearly free in latency and cost
Confidence-gated routing Act on high-confidence answers automatically, send low-confidence ones to review or a bigger model Safety without giving up speed on clear cases
Composite scoring Score each dimension separately (market size, feasibility, differentiation), weight them in code Change a coefficient instead of rewriting a prompt
Intent routing Classify an incoming request and route it to deterministic logic, a specialist LLM, or a human Most traffic never touches an expensive model

The cascade is the pattern people keep arriving at: Jev classifies and routes cheaply, code handles what it can, a frontier model takes the hard minority. Jev is not a replacement for the LLM; it removes the LLM from the steps that never needed one.

Known failure modes (Jev 1.13, reviewed 2026-09-17)

TypeSafe publishes a jaggedness page for each release. The current list:

Failure mode Do this instead
Literal reading: answers the words written, not the intent Write the exact condition and boundary cases in instructions and criteria
Math and counting: recognizes the shape of an answer rather than tallying Keep arithmetic in code; one Noul per item, sum in code
Date and time comparison: reads dates as text Extract components as Choices, compare in code
Indirection: double negatives, property-of-a-property, multi-hop Reduce hops; name the relevant part of state
Large state full of irrelevant detail: accuracy falls with distractors Filter in code first; send only what the question needs
Adversarial content: state is not treated as hostile Precise criteria; test injection cases before deploying
Contradictory instructions vs criteria Align them; avoid a Noul where true means no
Structural invariants: P(yes) from a Noul and from a Choice are not comparable, and P(noul) + P(not noul) need not sum to 1 Word each question directly; never carry a threshold across question types
Generation: it cannot write text Use an LLM, or turn extraction into a Choice over candidates

The adversarial-content line matters for security use: a process name or command line crafted to argue for its own benign classification can move the answer, so criteria need to be explicit and the model should not be the only gate.

Applications

Every good Jev use case has the same shape: a bounded decision, made often, where speed or volume made an LLM impractical. TypeSafe's use-case map lists ten decision shapes, condensed to eight rows below (search, retrieval, and ranking share one); the community builds from the first week map onto the same list.

Decision shapes (from TypeSafe's use-case map)

Shape Reach for it when Examples
Classification One known category should win Intent, topic, department, risk type
Detection You need the probability one property is present Spam, fraud, urgency, jailbreaks, sensitive data
Scoring The answer sits on an ordered rubric Severity, relevance, quality, frustration
Routing A category selects the next code path Tool choice, escalation, model routing, support queues
Search and ranking Items need ordering by semantic relevance Reranking, semantic find, candidate prioritization
Verification An artifact must be checked for specific failure modes Citation support, policy violations, tool-call errors
Feature extraction A classical ML model needs semantic signals Purchase intent, churn signals, risk indicators
Structured extraction Known fields must be recovered from text Order fields, dates, entity attributes

Source: Example use cases.

Harness engineering (the dominant early use)

Agents run in a loop: an LLM decides, a tool executes, something evaluates the result, repeat. Most of those evaluations are classifications, and Jev takes them out of the LLM's hands. Examples shipped in the first week:

Build Who What Jev decides
Building a Harness with Jev LangChain (Sydney Runkle) Which model or subagent handles the next step; whether to continue, retry, ask, or stop
Jev Model Router for Claude Code Daniel San Per-request subagent and main model selection via TypeSafe API or Vercel AI Gateway
Relevance-based compaction LiteLLM Whether each completed tool exchange is still relevant; blanks the rest before the LLM sees it
Skill suggestion TypeSafe cookbook Picks at most one of 182 Hermes skills per turn, or none
Guardrails for LLMs TypeSafe cookbook Jailbreak probability and harm severity on every input and output

Business workflows

Build Who What Jev decides
Bank transaction categorization Charlie Barmore, CPA Account coding on 250 synthetic transactions, scored against a key, compared to LLMs on cost and timing
Classifying SEC reports TypeSafe cookbook One of 75 industry groups per filing; low confidence falls back to the broader division
Re-ranking legal passages TypeSafe cookbook Query-candidate relevance; top-1 accuracy 5% to 18%, top-10 38% to 62% over BM25
Entity alignment TypeSafe cookbook Merge, leave, or hand to a curator, for 450 product pairs, no threshold to fit

Project note: bookkeeping reconciliation (QBO and Mercury)

Field note from an Agentic Studio Labs bookkeeping pipeline (QuickBooks Online plus Mercury): Jev fits one step, matching, and at the pipeline's current volume it would not change the result. It is worth running only as a low-stakes comparison test alongside the security-monitoring evaluation.

Where it fits: matching questions that can run in parallel with only the low-confidence answers escalated.

Question Type Example
Deposit to invoice Noul Does this $6,400 deposit on this date pay invoice X?
Account mapping Choice Which Books ledger account does this QBO account map to?

Why it is marginal today:

  • At small-business volume, most matches resolve exactly on amount, date, and customer name; the few leftovers can be judged by the agent or by a person in seconds.
  • Using Jev means sending client names and amounts to another vendor, a data-sharing cost with no offsetting gain at this volume.
  • Categorization, owner-transfer handling, and close readiness already work through the approval-gated agent and the JSONL audit trail; nothing there is bottlenecked on decision speed or cost.

If tested: run deposit-to-invoice matching through Jev next to the exact-match code on the same month and compare agreement, confidence on the disagreements, and whether any low-confidence answer was actually wrong. Keep the bank connector read-only, keep every ledger write approval-gated, and keep amounts and dates matched in code; Jev reads dates as text.

Sources: TypeSafe docs: Introduction, Introducing System One Models and Jev.

Real-time and machine-native

Build Who What Jev decides
Trading bot on Monad Jarrod Watts Buy or sell from a price feed, every 300 ms block
macOS clipboard watcher Marcel Pociot Whether clipboard content is a terminal command, and which quick action to offer
Doom, Minecraft bot, drone obstacle course, self-driving sim TypeSafe launch demos Next action from game or sensor state, faster than human perception

The thread running through all three tables: code owns the loop, Jev owns the judgment call inside it, and confidence decides whether the judgment is trusted or escalated.

Project note: macOS process-monitoring agent

Field note from an Agentic Studio Labs endpoint agent under evaluation. The agent design already matches Jev's shape: an ES NOTIFY path (detect, not block), raw events and per-user baselines kept locally in SQLite, and compressed metadata or delta packets sent out for judgment. Each packet is the state; each of the three use cases becomes a small set of typed questions asked in one call.

Use case Question type Example question Code owns
User-behavior baselining Noul Is this process, parent, and argument pattern consistent with the attached baseline summary for this user? Baseline statistics, frequency counts, time-of-day math
Antivirus gray zone Choice Benign tooling / dual-use admin tool / likely unwanted / needs analyst, with criteria per option Hash lookups, signature checks, allowlists
Malware behavior Score (2 to 10 levels) Severity of this event sequence against a MITRE-style rubric (persistence, injection, exfil) Correlation across events, thresholds, alert routing
ES NOTIFY event
↓
Swift agent: enrich + baseline lookup in SQLite
↓
Cheap local rules pass?
yes →
Log only
↓ no
Delta packet as state + fan-out questions to Jev
↓
Confidence high?
yes →
Act on typed answer
↓ no
Escalate: frontier LLM or analyst

Jev sits between deterministic rules and the expensive tier; confidence gating decides which way an event goes.

What the docs say to watch for in this workload:

  • Filter before sending. Accuracy drops with irrelevant state, so the delta packet should carry only the fields each question needs, not the raw event stream.
  • Do the arithmetic locally. Rate spikes, time windows, and counts belong in the agent; Jev reads them as semantic facts once computed ("3x the user's hourly baseline").
  • Prefer semantic over numeric representations. Send "unsigned binary in /tmp launched by a browser" rather than raw paths and hashes alone.
  • Treat state as untrusted. A process argument written to look benign can move a classification. Keep the model as one signal, keep allow/deny logic in code, and test injection cases.
  • Pin the model version once thresholds are tuned; jev-latest can move.
  • Latency budget: 70 to 500 ms per call fits a detection path, not an inline blocking path, which matches the NOTIFY decision. Batching several events into one state with per-event questions is cheaper than one call per event.

Open question: whether per-user baselines can be summarized small enough to fit the 32k state budget alongside the event, or whether the agent should send only the deviation summary.

Getting started

Access is open to everyone as of September 20, 2026, no waitlist (announcement); keys come from console.typesafe.ai. Everything hits one endpoint, POST https://api.typesafe.ai/v1/systemone, and the SDKs wrap it.

Route Install or call Notes
Python SDK typesafe_sdk, TypeSafeClient().system_one(state, questions) Sync and async clients, retries with backoff built in
JavaScript SDK @typesafe-ai/sdk, choice(), score(), noul() helpers Typed question and response interfaces
Claude Code plugin claude plugin marketplace add typesafe-ai/skills then claude plugin install typesafe@typesafe-ai Invoke with /typesafe:typesafe-ai; teaches the agent the primitives and patterns
Other agents npx skills add typesafe-ai/skills --skill typesafe-ai Project-local by default, -g for global
LangChain langchain-typesafe, TypeSafeClassifier Drops into a harness as the routing step
Vercel AI Gateway AI SDK 7.0.105+, experimental_evaluate with model typesafe-ai/jev ZDR option via provider settings
LiteLLM Pass-through at LITELLM_PROXY_BASE_URL/typesafe, plus a compaction guardrail Proxy holds the key; cost tracking and logging
OpenRouter playground Browser Mini demos, no code
Braintrust Tracing integration Inspect questions, answers, and probability distributions per decision

Swift has no official SDK; the HTTP API is a single JSON POST, so a thin URLSession wrapper covers the macOS agent. Keep the questions and threshold constants in one file so they are the thing a reviewer reads.

A sensible first experiment: export a few hundred real events from the SQLite store, write three or four questions, and measure calibration against your own labels before wiring anything into the agent.

Have an agent loop paying LLM prices for yes/no decisions? We design harnesses that put the right model on each step and keep code in control.

./book-discovery