What Jev is
Jev is a decision model, not a text model: you send it a state plus typed questions, and it returns typed answers with probabilities that your code can branch on directly. TypeSafe AI calls this class a System One model, after Kahneman's fast, intuitive mode of thinking (see the next section), in contrast to the slow, deliberate generation of an LLM.
It launched on September 15, 2026 from TypeSafe AI, founded by Diogo Almeida, a former OpenAI researcher and co-inventor of RLHF and InstructGPT, with a $40M round led by DCVC. Early access was waitlisted for the first five days; on September 20, 2026 TypeSafe opened it to everyone with no waitlist. The name references William Stanley Jevons and the Jevons paradox: cheaper tokens will mean more tokens consumed, not fewer.
The simplest mental model
An LLM is a writer, Jev is a judge. It answers the kind of gut-check question a knowledgeable person could settle in a few seconds given the right context, and nothing more.
System 1, System 2, and where the models sit
The names come from Daniel Kahneman's Thinking, Fast and Slow (2011): System 1 is the fast, automatic, intuitive mode (recognizing a face, finishing "tips and...", reading a stop sign), System 2 is the slow, effortful, deliberate mode (multiplying 343 by 57, weighing a contract clause). There is no System 3 in Kahneman's framework; the numbers stop at two, and "System One model" is TypeSafe's coinage for a model class, not a step in a series.
| Kahneman mode | Model class | Trained with | Returns | Examples |
|---|---|---|---|---|
| System 1: fast, intuitive, calibrated gut-check | System One model | Reinforcement learning for calibrated decisions (RLCD) | Typed answers plus probabilities and confidence, in one pass | Jev |
| Between the two: fluent, general, conversational | Chat LLM | Next-token prediction plus RLHF | Free-form text, token by token | GPT-5.6 Terra, Claude Sonnet 5 |
| System 2: slow, deliberate, multi-step | Reasoning LLM | RLHF plus reinforcement learning on verifiable rewards (RLVR) | Text after an extended thinking trace | GPT-6 Astra, Claude Fable 5.1 |
The useful reading is that an agent needs both modes, the way a person does. Most steps in a loop are System 1 calls (is this relevant, which tool, how urgent, is this safe), and a few are System 2 (write the fix, plan the migration, explain the tradeoff). Chat LLMs have been doing both jobs; Jev exists to take the first kind back at System 1 speed and cost, leaving the reasoning model for the second.
One caveat on the analogy: Kahneman's System 1 is prone to bias precisely because it is fast, and Jev's failure modes (literal reading, no arithmetic, distractible by irrelevant state) are the machine version of the same tradeoff. Calibrated confidence is what makes it safe to build on: the model can say how sure it is, so code knows when to slow down and escalate.
How it works
One request carries a state and any number of typed questions; Jev ingests the state once and evaluates every question against it in parallel, in isolation, in a single forward pass. Adding questions barely changes response time, and because each question is evaluated independently there is no context rot between them.
Code stays in control; Jev supplies narrow judgments inside the loop.
The three primitives
| Question type | Asks | Returns |
|---|---|---|
| Choice | Pick one option from a defined set (up to 255) | choice, a probability per option, confidence |
| Score | Rate the state against ordered, descriptive levels (2 to 10) | score, a probability per level, confidence |
| Noul | Is this statement true? | noul, a 0 to 1 probability |
All three can be mixed in one call. Confidence is a second axis distinct from probability: the answer says what, confidence says whether to act on it.
Training
Jev is trained with Reinforcement Learning for Calibrated Decisions (RLCD) rather than next-token prediction, so the probabilities it returns are meant to be calibrated, not just ranked. It is not fine-tuned per customer; you shape it through the state (your records and reference material) and each question's instructions and criteria, which accept JSON structure. Same weights serve every account.
Limits (Jev 1.13, jev-1.13.0)
| Item | Value |
|---|---|
| Context | 64k tokens per request; 32k for state plus the longest question |
| Input | Text only (string, JSON object, or array); pre-process images, audio, binaries to text |
| Rate limits | 250,000 tokens/sec, 1,200 requests/min, adjusting dynamically |
| Language | English primary; other languages handled but less accurately |
| Aliases | jev-latest, jev-preview (both point to 1.13.0 today); pin the versioned ID if you tune confidence thresholds |
Source: TypeSafe docs, Models and Introduction.
Speed and cost
Jev costs $0.042 per million input tokens with output free, and answers in roughly 70 to 500 ms regardless of how many questions ride in the request. That is the whole pitch: decisions that used to cost an LLM round trip now cost about as much as a database query.
| Comparison | Figure | Who reports it |
|---|---|---|
| Input price | $0.042 / Mtok (a Btok is $42), output $0 | TypeSafe Models page |
| Latency, demo on typesafe.ai | 0.114 s vs 8.566 s for GPT-5.6 Terra | The Register |
| Cost vs Claude Fable 5.1 | 238x cheaper | The Register |
| Workflow evals vs LLMs | up to 193.6x faster, 444.6x cheaper | Vercel changelog, citing TypeSafe |
| Batching 13 questions in one call vs 13 calls | 12.2x cheaper, 10.0x faster, same answers | TypeSafe cookbook |
Read the multipliers with care: they are TypeSafe's own numbers on classification and routing tasks, where an LLM is overkill by construction. On anything that needs generation or multi-step reasoning Jev does not compete, it is simply the wrong tool. The honest framing is that it collapses the cost of the 80 percent of agent steps that are yes/no or pick-one, and leaves the LLM for the rest.
One more cost lever: rate limits (250k tokens/sec, 1,200 req/min) are being adjusted dynamically during launch demand, so a production integration should assume 429s and retry with backoff.
Design patterns and failure modes
The skill with Jev is decomposition: break a judgment into atomic questions, ask them all in one call, and reassemble the answer in code you control. TypeSafe documents four patterns that cover most systems.
| Pattern | What it does | Why |
|---|---|---|
| Speculative fan-out | Pack many questions, including ones you may not need, into one call; code decides what is relevant | Adding questions is nearly free in latency and cost |
| Confidence-gated routing | Act on high-confidence answers automatically, send low-confidence ones to review or a bigger model | Safety without giving up speed on clear cases |
| Composite scoring | Score each dimension separately (market size, feasibility, differentiation), weight them in code | Change a coefficient instead of rewriting a prompt |
| Intent routing | Classify an incoming request and route it to deterministic logic, a specialist LLM, or a human | Most traffic never touches an expensive model |
The cascade is the pattern people keep arriving at: Jev classifies and routes cheaply, code handles what it can, a frontier model takes the hard minority. Jev is not a replacement for the LLM; it removes the LLM from the steps that never needed one.
Known failure modes (Jev 1.13, reviewed 2026-09-17)
TypeSafe publishes a jaggedness page for each release. The current list:
| Failure mode | Do this instead |
|---|---|
| Literal reading: answers the words written, not the intent | Write the exact condition and boundary cases in instructions and criteria |
| Math and counting: recognizes the shape of an answer rather than tallying | Keep arithmetic in code; one Noul per item, sum in code |
| Date and time comparison: reads dates as text | Extract components as Choices, compare in code |
| Indirection: double negatives, property-of-a-property, multi-hop | Reduce hops; name the relevant part of state |
| Large state full of irrelevant detail: accuracy falls with distractors | Filter in code first; send only what the question needs |
| Adversarial content: state is not treated as hostile | Precise criteria; test injection cases before deploying |
| Contradictory instructions vs criteria | Align them; avoid a Noul where true means no |
| Structural invariants: P(yes) from a Noul and from a Choice are not comparable, and P(noul) + P(not noul) need not sum to 1 | Word each question directly; never carry a threshold across question types |
| Generation: it cannot write text | Use an LLM, or turn extraction into a Choice over candidates |
The adversarial-content line matters for security use: a process name or command line crafted to argue for its own benign classification can move the answer, so criteria need to be explicit and the model should not be the only gate.
Applications
Every good Jev use case has the same shape: a bounded decision, made often, where speed or volume made an LLM impractical. TypeSafe's use-case map lists ten decision shapes, condensed to eight rows below (search, retrieval, and ranking share one); the community builds from the first week map onto the same list.
Decision shapes (from TypeSafe's use-case map)
| Shape | Reach for it when | Examples |
|---|---|---|
| Classification | One known category should win | Intent, topic, department, risk type |
| Detection | You need the probability one property is present | Spam, fraud, urgency, jailbreaks, sensitive data |
| Scoring | The answer sits on an ordered rubric | Severity, relevance, quality, frustration |
| Routing | A category selects the next code path | Tool choice, escalation, model routing, support queues |
| Search and ranking | Items need ordering by semantic relevance | Reranking, semantic find, candidate prioritization |
| Verification | An artifact must be checked for specific failure modes | Citation support, policy violations, tool-call errors |
| Feature extraction | A classical ML model needs semantic signals | Purchase intent, churn signals, risk indicators |
| Structured extraction | Known fields must be recovered from text | Order fields, dates, entity attributes |
Source: Example use cases.
Harness engineering (the dominant early use)
Agents run in a loop: an LLM decides, a tool executes, something evaluates the result, repeat. Most of those evaluations are classifications, and Jev takes them out of the LLM's hands. Examples shipped in the first week:
| Build | Who | What Jev decides |
|---|---|---|
| Building a Harness with Jev | LangChain (Sydney Runkle) | Which model or subagent handles the next step; whether to continue, retry, ask, or stop |
| Jev Model Router for Claude Code | Daniel San | Per-request subagent and main model selection via TypeSafe API or Vercel AI Gateway |
| Relevance-based compaction | LiteLLM | Whether each completed tool exchange is still relevant; blanks the rest before the LLM sees it |
| Skill suggestion | TypeSafe cookbook | Picks at most one of 182 Hermes skills per turn, or none |
| Guardrails for LLMs | TypeSafe cookbook | Jailbreak probability and harm severity on every input and output |
Business workflows
| Build | Who | What Jev decides |
|---|---|---|
| Bank transaction categorization | Charlie Barmore, CPA | Account coding on 250 synthetic transactions, scored against a key, compared to LLMs on cost and timing |
| Classifying SEC reports | TypeSafe cookbook | One of 75 industry groups per filing; low confidence falls back to the broader division |
| Re-ranking legal passages | TypeSafe cookbook | Query-candidate relevance; top-1 accuracy 5% to 18%, top-10 38% to 62% over BM25 |
| Entity alignment | TypeSafe cookbook | Merge, leave, or hand to a curator, for 450 product pairs, no threshold to fit |
Project note: bookkeeping reconciliation (QBO and Mercury)
Field note from an Agentic Studio Labs bookkeeping pipeline (QuickBooks Online plus Mercury): Jev fits one step, matching, and at the pipeline's current volume it would not change the result. It is worth running only as a low-stakes comparison test alongside the security-monitoring evaluation.
Where it fits: matching questions that can run in parallel with only the low-confidence answers escalated.
| Question | Type | Example |
|---|---|---|
| Deposit to invoice | Noul | Does this $6,400 deposit on this date pay invoice X? |
| Account mapping | Choice | Which Books ledger account does this QBO account map to? |
Why it is marginal today:
- At small-business volume, most matches resolve exactly on amount, date, and customer name; the few leftovers can be judged by the agent or by a person in seconds.
- Using Jev means sending client names and amounts to another vendor, a data-sharing cost with no offsetting gain at this volume.
- Categorization, owner-transfer handling, and close readiness already work through the approval-gated agent and the JSONL audit trail; nothing there is bottlenecked on decision speed or cost.
If tested: run deposit-to-invoice matching through Jev next to the exact-match code on the same month and compare agreement, confidence on the disagreements, and whether any low-confidence answer was actually wrong. Keep the bank connector read-only, keep every ledger write approval-gated, and keep amounts and dates matched in code; Jev reads dates as text.
Sources: TypeSafe docs: Introduction, Introducing System One Models and Jev.
Real-time and machine-native
| Build | Who | What Jev decides |
|---|---|---|
| Trading bot on Monad | Jarrod Watts | Buy or sell from a price feed, every 300 ms block |
| macOS clipboard watcher | Marcel Pociot | Whether clipboard content is a terminal command, and which quick action to offer |
| Doom, Minecraft bot, drone obstacle course, self-driving sim | TypeSafe launch demos | Next action from game or sensor state, faster than human perception |
The thread running through all three tables: code owns the loop, Jev owns the judgment call inside it, and confidence decides whether the judgment is trusted or escalated.
Project note: macOS process-monitoring agent
Field note from an Agentic Studio Labs endpoint agent under evaluation. The agent design already matches Jev's shape: an ES NOTIFY path (detect, not block), raw events and per-user baselines kept locally in SQLite, and compressed metadata or delta packets sent out for judgment. Each packet is the state; each of the three use cases becomes a small set of typed questions asked in one call.
| Use case | Question type | Example question | Code owns |
|---|---|---|---|
| User-behavior baselining | Noul | Is this process, parent, and argument pattern consistent with the attached baseline summary for this user? | Baseline statistics, frequency counts, time-of-day math |
| Antivirus gray zone | Choice | Benign tooling / dual-use admin tool / likely unwanted / needs analyst, with criteria per option | Hash lookups, signature checks, allowlists |
| Malware behavior | Score (2 to 10 levels) | Severity of this event sequence against a MITRE-style rubric (persistence, injection, exfil) | Correlation across events, thresholds, alert routing |
Jev sits between deterministic rules and the expensive tier; confidence gating decides which way an event goes.
What the docs say to watch for in this workload:
- Filter before sending. Accuracy drops with irrelevant state, so the delta packet should carry only the fields each question needs, not the raw event stream.
- Do the arithmetic locally. Rate spikes, time windows, and counts belong in the agent; Jev reads them as semantic facts once computed ("3x the user's hourly baseline").
- Prefer semantic over numeric representations. Send "unsigned binary in /tmp launched by a browser" rather than raw paths and hashes alone.
- Treat state as untrusted. A process argument written to look benign can move a classification. Keep the model as one signal, keep allow/deny logic in code, and test injection cases.
- Pin the model version once thresholds are tuned;
jev-latestcan move. - Latency budget: 70 to 500 ms per call fits a detection path, not an inline blocking path, which matches the NOTIFY decision. Batching several events into one state with per-event questions is cheaper than one call per event.
Open question: whether per-user baselines can be summarized small enough to fit the 32k state budget alongside the event, or whether the agent should send only the deviation summary.
Getting started
Access is open to everyone as of September 20, 2026, no waitlist (announcement); keys come from console.typesafe.ai. Everything hits one endpoint, POST https://api.typesafe.ai/v1/systemone, and the SDKs wrap it.
| Route | Install or call | Notes |
|---|---|---|
| Python SDK | typesafe_sdk, TypeSafeClient().system_one(state, questions) |
Sync and async clients, retries with backoff built in |
| JavaScript SDK | @typesafe-ai/sdk, choice(), score(), noul() helpers |
Typed question and response interfaces |
| Claude Code plugin | claude plugin marketplace add typesafe-ai/skills then claude plugin install typesafe@typesafe-ai |
Invoke with /typesafe:typesafe-ai; teaches the agent the primitives and patterns |
| Other agents | npx skills add typesafe-ai/skills --skill typesafe-ai |
Project-local by default, -g for global |
| LangChain | langchain-typesafe, TypeSafeClassifier |
Drops into a harness as the routing step |
| Vercel AI Gateway | AI SDK 7.0.105+, experimental_evaluate with model typesafe-ai/jev |
ZDR option via provider settings |
| LiteLLM | Pass-through at LITELLM_PROXY_BASE_URL/typesafe, plus a compaction guardrail |
Proxy holds the key; cost tracking and logging |
| OpenRouter playground | Browser | Mini demos, no code |
| Braintrust | Tracing integration | Inspect questions, answers, and probability distributions per decision |
Swift has no official SDK; the HTTP API is a single JSON POST, so a thin URLSession wrapper covers the macOS agent. Keep the questions and threshold constants in one file so they are the thing a reviewer reads.
A sensible first experiment: export a few hundred real events from the SQLite store, write three or four questions, and measure calibration against your own labels before wiring anything into the agent.
Have an agent loop paying LLM prices for yes/no decisions? We design harnesses that put the right model on each step and keep code in control.
./book-discovery