Series: AI Agent 實戰 (3/4)

← Loop Engineering: Designing Systems That Prompt Agents for You How to Design Safety Layers for an AI Agent: From Keyword Detection to Long-Term Behavioral Monitoring →
Key Points 16 min read
  • 200x faster isn't because the model is small — the token-by-token generation loop simply doesn't exist
  • 'No hallucination' guarantees types, not correctness — miss the right answer in your schema and it confidently picks second best
  • Already on Cloudflare Workers AI as typesafe/jev — your existing AI binding can call it today
Table of Contents

Two Chinese-language tech channels covered the same model this week: one asked “200x faster and 400x cheaper?”, the other asked “is AI Agent architecture really about to change?”

The model is Jev, and it only opened for early access on September 15, 2026. What makes it interesting isn’t a benchmark score — it’s that it doesn’t generate text at all. You ask it questions and it returns typed values with probabilities, designed to be consumed by code, not read by a person.

This post covers three layers: how it actually works, what’s wrong with the “cannot hallucinate” claim, and how to decide which nodes in your own system should be swapped out.


TL;DR

Jev is TypeSafe AI’s System One Model: unstructured state in, typed probabilistic decisions out. The core insight is that a large share of the model calls inside an agent loop don’t need to “talk” at all — yet they’re paying talking prices.

It’s fast and cheap not because the model is small, but because the thing that drives LLM latency and output pricing — the token-by-token generation loop — doesn’t exist here.


Start with what an agent actually does. The standard loop: the LLM decides the next step → a tool executes → the result is fed back → the LLM decides again.

The problem is the deciding. Every judgment call — which model should this request route to, is this tool call dangerous, is this retrieved chunk relevant, should this turn end — is a full LLM call. And every call goes through the complete generation loop, emitting tokens one at a time, just to produce "yes" or "risky": an answer two or three tokens long.

You paid for a full natural-language generation pipeline, in both money and latency, to get a boolean.

Founder Diogo Almeida (formerly at OpenAI, co-creator of ChatGPT and RLHF) puts it this way: “We have lightning in a bottle, and yet it is not useful… computers speak a different language.” LLMs speak human, but the caller is a program. All that glue in between — parsing, retrying, validating format — exists to translate human back into machine.

Jev’s approach is to remove that translation layer entirely: if the caller is a program, output something a program can eat directly.


Three Primitives: Choice, Score, Noul

Jev’s API has exactly three question types. The constraint is deliberate — every decision has to be pressed into one of these three shapes first.

Noul (yes/no) — returns a probability between 0 and 1:

{
  "type": "noul",
  "instructions": "Will this shell command cause irreversible damage?",
  "criteria": {
    "true": "Deletes files, rewrites history, or touches production",
    "false": "Read-only or safely reversible"
  }
}

Response:

{ "type": "noul", "noul": 0.95 }

Choice (pick one) — returns the selected option plus the full probability distribution and a confidence value:

{
  "type": "choice",
  "choice": "needs_reasoning",
  "probabilities": { "needs_reasoning": 0.88, "simple_lookup": 0.12 },
  "confidence": 0.81
}

Up to 255 options per question.

Score (rate) — positions the input on an ordered scale of 2 to 10 levels:

{
  "type": "score",
  "score": 1.05,
  "legend": { "0": "Not worth reading", "1": "Worth referencing", "2": "Must read" },
  "probabilities": { "0": 0.0, "1": 0.95, "2": 0.05 },
  "confidence": 0.92
}

Note that score is 1.05, not 1 — it’s a probability-weighted continuous value carrying the information “leaning slightly toward the next level up.” That’s genuinely useful for ranking, and a plain classifier can’t give it to you.

A full request looks like this: one state, a set of questions:

POST https://api.typesafe.ai/v1/systemone
{
  "model": "jev-latest",
  "state": { "title": "...", "blurb": "...", "source": "..." },
  "questions": {
    "worth_reading": { "type": "score", "instructions": "...", "criteria": [...] },
    "is_ad": { "type": "noul", "instructions": "...", "criteria": {...} },
    "topic": { "type": "choice", "instructions": "...", "criteria": {...} }
  }
}

The critical detail: adding more questions barely changes response time. All questions are scored in parallel within the same forward pass. This fundamentally changes the cost calculus — where you used to weigh whether a judgment was worth an extra LLM call, you can now ask twenty at once.


Why It’s Fast: There Is No Generation Loop

This is the part that gets explained wrong most often. “200x faster” is not because Jev is a small model.

An autoregressive LLM producing N tokens needs N sequential forward passes, each waiting on the previous token before it can compute the next. That sequential dependency is the source of the latency, and the source of output-token pricing.

Jev’s architecture is a single parallel forward pass: it doesn’t “produce” an answer, it scores every possible answer in the schema simultaneously. The options are finite and known in advance, so nothing needs to be generated step by step — compute once, take the distribution.

flowchart TB
  subgraph LLM["Autoregressive LLM: N tokens = N sequential forward passes"]
    direction LR
    P1[prompt] --> T1[token 1] --> T2[token 2] --> T3[token 3] --> T4["...until EOS"]
  end
  subgraph JEV["Jev: all candidate answers scored at once"]
    direction LR
    S[state + questions] --> F[single parallel forward pass]
    F --> O1["option A: 0.88"]
    F --> O2["option B: 0.12"]
    F --> O3["score: 1.05"]
  end
  LLM -.->|"remove the generation loop"| JEV

Which is why the pricing reads $0.042 per million input tokens, output free — there is no output to charge for. Measured latency lands between 70 and 500ms.

The training method is RLCD (Reinforcement Learning for Calibrated Decisions), using synthetic data exclusively. The objective function isn’t “does this sound human” but “is this probability estimate accurate."


"Mathematically Cannot Hallucinate” Is Being Misread

The claim that Jev “mathematically cannot hallucinate” is everywhere. Half of it is true, but it’s being read as the other half.

What’s guaranteed is the output space, not output correctness.

Schema constraints do make it impossible to return anything outside the type — define three options and it won’t invent a fourth, won’t return malformed JSON, won’t suddenly start explaining itself. That eliminates an entire class of format failures, which is genuinely the most annoying category of LLM structured-output bugs.

But it will still pick B when the answer was A. A misclassification isn’t a hallucination — yet for your system, the consequence is identical.

More worth watching is one specific failure mode: if the right answer isn’t in your schema, Jev will confidently pick the least-wrong one. There’s no “I don’t know” exit unless you add it to the options yourself. A Noul with only {safe, risky} will, when faced with “I don’t understand this command,” return a probability somewhere in the middle — and what you receive is a perfectly normal-looking float.

So in practice: any open-ended judgment should include an unknown / needs_escalation option, with a confidence threshold that escalates upward.


Calibration Is the Metric That Actually Matters

More important than accuracy is calibration: when the model says 0.7, is it actually right 70 times out of 100?

General LLMs are bad at this. Ask a GPT-class model to output a 0–1 confidence score and you’ll get a pile of 0.85s and 0.9s — because that’s how humans write confidence scores in the training corpus, not because it reflects internal uncertainty. The number has no statistical meaning, so you can’t threshold on it.

Jev’s entire training objective is to make that number usable. That’s what the confidence field is for — it lets you write code like this:

const { answers } = await env.AI.run('typesafe/jev', { state, questions });
const risk = answers.is_destructive;

if (risk.noul > 0.9) return block();       // confidently dangerous → block
if (risk.noul < 0.1) return proceed();     // confidently safe → allow
return escalateToLLM(state);               // gray zone → only now pay for a big model

That’s a three-state decision, not a binary classification. Being able to write it depends entirely on the probability being trustworthy. Early-adopter feedback points the same way: email classification service Bryo AI found Gemini slightly more accurate, but at 10–20x the cost — and noted that Jev gives “real probability”, which matters more for workflow automation than those few accuracy points.

Vercel uses it for command safety review — swapping the safety reviewer from GPT-5.6 Luna to Jev, they report 5–18x faster (p95) with better accuracy. One number captures the market reaction: within 24 hours of launch, nearly 13% of paid Vercel AI Gateway teams were using it, making it the fastest-adopted model in that platform’s history — more than double the share any previous launch reached, including the GPT-5.6 family.

That said — every 40–200x / 40–400x figure is vendor-reported, and TypeSafe itself acknowledges the benchmark workflows were internally constructed. The right use for numbers like these is “worth spending an afternoon testing,” not “safe to cite in an architecture decision record.”


How It Differs From What You’re Already Using

This is the practical question. You might be thinking: can’t I just use structured output?

ApproachGuaranteesLatency sourceProbabilities?Best for
LLM + JSON mode / structured outputValid formatStill token-by-token❌ No trustworthy probabilitiesProducing text and structure together
Constrained decoding (grammar)Valid formatStill token-by-token⚠️ logprobs uncalibratedSelf-hosted models needing strict grammar
Fine-tuned BERT-class classifierType safetyVery fast⚠️ Must calibrate yourselfFixed labels, labeled data, rarely changes
JevType safety + calibrated probabilitiesSingle forward pass✅ It’s the design goalLabels change often, no labeled data, need probabilities

The difference from structured output is architectural: that approach wraps constraints around the generation loop, but the loop is still there — you save nothing on latency or output cost.

The difference from fine-tuning your own small classifier is operational: the BERT route is equally fast and cheap, but you need labeled data, a training pipeline, and a full retrain to change one label. Changing Jev’s criteria means editing an instructions string — a huge difference in scenarios where label definitions shift weekly (content grading, risk assessment).

The community has also connected this to DSPy’s signature concept: compiling expensive LLM calls down into specialized small AI functions. Same direction.


Three Patterns for Wiring It Into an Existing Pipeline

Jev doesn’t replace your main model. It replaces the trivial judgments your main model is moonlighting on.

flowchart LR
  REQ[Request in] --> R{"Jev: route<br/>Choice"}
  R -->|simple| SM[Small model / lookup]
  R -->|complex| LM[Frontier LLM]
  LM --> TC[Emit tool call]
  TC --> G{"Jev: safety gate<br/>Noul"}
  G -->|risky| HUM[Require human approval]
  G -->|safe| EXE[Execute tool]
  EXE --> V{"Jev: verify<br/>Score"}
  V -->|fails| LM
  V -->|passes| OUT[Return result]

1. Model Router — have Jev assess complexity with a Choice on the way in; simple requests go to a cheap model, complex ones to a frontier LLM. What you save is the whole batch of calls that never needed a big model.

2. Safety Gate — judge risk with a Noul before the tool call executes. This is exactly what a coding agent’s auto-accept mode needs: the check has to be fast enough that the user never notices, otherwise you may as well just ask them. Vercel’s usage is precisely this.

3. Verifier — score result quality with a Score after the tool runs, and bounce it back if it fails. This one deserves attention, because it answers the conclusion from Loop Engineering head-on: verification cost determines what you can automate. When the marginal cost of verification approaches zero, every loop you had to abandon because “verifying costs more than redoing” suddenly becomes viable.

Which is why Jev matters more to Harness Engineering than it does to model capability — it made no model smarter, it made the harness able to afford far denser checkpoints.


This Site as a Worked Example

Abstract is less useful than real. This site runs on Cloudflare Workers, and its current Workers AI usage is: @cf/baai/bge-m3 for embeddings, @cf/qwen/qwen1.5-14b-chat-awq for RAG answers, @cf/baai/bge-reranker-base for reranking, and @cf/meta/llama-3.1-8b-instruct for metadata extraction during ingest.

Reclassify those calls as “generation vs decision” and the picture is stark:

Existing nodeNatureCurrent approachSystem One shape
RAG answer generationGenerationqwen-14b streamingUnchanged — this one needs to talk
Article embeddingEmbeddingbge-m3Unchanged
Daily mining A/B/C gradingDecisionLLM reads title + blurb, scoresScore (3 levels)
Should a query go through RAGDecisionNone today — always doesNoul
Fact-check’s four verdictsDecisionLLM produces judgmentChoice (4 options)
Is a term “hardcore” enoughDecisionLLM judges word by wordNoul
Are two tags duplicates/synonymsDecisionLLM comparesNoul
Should a submitted FAQ be approvedDecisionManualNoul + confidence threshold

Six of eight nodes are decisions, not generation. And all six currently pay generation prices.

The wiring is the easy part. The AI binding already exists in wrangler.jsonc:

"ai": { "binding": "AI", "remote": true }

And Jev is already listed on Cloudflare Workers AI as model ID typesafe/jev, with a 32k context window. No new API key, no new SDK, no deployment changes:

// Before: one LLM call to get a boolean
const verdict = await AI.run('@cf/meta/llama-3.1-8b-instruct', {
  messages: [{ role: 'user', content: `Is this worth writing about? Reply A/B/C only: ${blurb}` }],
});
// then parse the text, handle it explaining itself, handle it replying "B級" instead of "B"...

// After: one call, five questions, typed values back
const { answers } = await AI.run('typesafe/jev', {
  state: { title, blurb, source },
  questions: {
    worth_writing: {
      type: 'score',
      instructions: 'Is this worth rewriting into a deep technical article?',
      criteria: ['Not worth it', 'Usable as material', 'Must read'],
    },
    is_promo: {
      type: 'noul',
      instructions: 'Is this product marketing rather than technical content?',
      criteria: { true: 'Primarily selling a product', false: 'Has real technical substance' },
    },
    topic: {
      type: 'choice',
      instructions: 'Topic classification',
      criteria: { agent: null, rag: null, infra: null, model: null, other: null },
    },
  },
});

if (answers.worth_writing.score > 1.5 && answers.worth_writing.confidence > 0.8) {
  // high-confidence "must read"
}

The difference isn’t just cost. The second version has no parsing, no retries, and no defensive code for “the model added another sentence again” — and the two extra questions cost essentially no additional time.


When Not to Use It

These are the practical criteria for deciding whether to adopt:

ScenarioUse Jev?Why
Classification, routing, scoring, risk assessment✅Dead center of the design goal
Need probabilities to set thresholds✅Calibration is the core value
Multiple independent questions on one input✅Parallel scoring is nearly free
Producing any text a human will read❌It doesn’t talk
Summarizing, translating, rewriting❌Same
Decision requires a lookup or tool call first❌Single forward pass — can’t chain, can’t ask back
More than 255 options❌Hard limit
Need to explain the reasoning❌Probabilities only, no rationale
Fixed labels with plenty of labeled data⚠️A fine-tuned small model may be cheaper
Input exceeds 32k tokens⚠️Context limit on Cloudflare

That last “no rationale” point deserves emphasis. If your scenario requires justifying a decision to a user (content moderation appeals, risk-based rejections), Jev gives you 0.93 and nothing else. The sensible architecture there is Jev filters first, and the small number of blocked cases go to an LLM for an explanation — that’s only a few percent of traffic anyway.

Two operational realities worth noting. Availability: it launched on September 15 behind a waitlist, but demand was heavy enough that the waitlist was removed on September 20–21 — anyone can now sign up at console.typesafe.ai with $5 in free credit (roughly 120 million tokens), so there’s nothing to wait for if you want to test it. Stability: the API was overwhelmed during launch week, and 429 (rate limited) plus 529 (overloaded) need exponential backoff — the official SDKs handle it, but you’ll implement it yourself when calling through Workers AI.


Overall

The model is named Jev, after the economist William Stanley Jevons — Jevons paradox: when a resource becomes more efficient and cheaper per unit, total consumption doesn’t fall, it explodes.

That naming is the most honest part of this whole story. If decision costs really drop 400x, the outcome isn’t a bill that’s 400x smaller. The outcome is that you’ll stuff twenty previously-unaffordable questions into every agent loop: verify every step, gate every tool call, score every retrieved chunk, judge every turn’s continuation.

Which brings it back to the old point — the model isn’t dumb, it lacks guidance. And the density of that guidance has always been capped by cost. Jev made no model smarter; it just pushed the marginal cost of asking one more question toward zero.

So the answer to “is AI Agent architecture about to change” is probably: the architecture won’t, but the density of checkpoints will. You used to insert verification only at critical nodes, because each one was an LLM call. Now you can insert it at every step.

There’s exactly one thing worth doing right now: list every LLM call in your system and mark which ones only need a typed value back. That list will be longer than you expect.


References

For a deeper look at the technologies and architecture referenced here, the official docs and extended reading below are a good starting point.

Some material is not expanded in full for length reasons; inline links throughout the post point to further reading.

Ask this article

Answers come from this article only. Click any prompt below or open the chat at the bottom right.

🇺🇸 English

Two Chinese-language tech channels covered the same model this week. One asked, "two hundred times faster and four hundred times cheaper?" The other asked, "is AI agent architecture about to change?"

The model is called Jev, and it only opened for early access on September fifteenth, 2026. What makes it interesting isn't a benchmark score. It's that Jev doesn't generate text at all. You ask it questions, and it hands back typed values with probabilities attached — designed to be consumed by code, not read by a person.

So let me take you through three layers here: how it actually works, what's wrong with the "it cannot hallucinate" claim that's going around, and how you'd decide which pieces of your own system are worth swapping out.

Here's the core idea in one breath. Jev is what TypeSafe AI calls a System One Model. Unstructured state goes in, typed probabilistic decisions come out. And the insight behind it is this: a huge share of the model calls inside an agent loop don't need to *talk* at all — and yet they're paying talking prices.

It's fast and cheap not because the model is small. It's fast and cheap because the thing that drives latency and output pricing in an LLM — that token-by-token generation loop — simply doesn't exist here.

Let's start with why anybody wants a model that can't talk.

Think about what an agent actually does. The standard loop: the LLM decides the next step, a tool executes, the result comes back, the LLM decides again. The problem is the deciding. Every single judgment call — which model should this request route to, is this tool call dangerous, is this retrieved chunk relevant, should this turn end — every one of those is a full LLM call. And every one of those calls runs the complete generation pipeline, emitting tokens one at a time, just to produce the word "yes" or the word "risky." An answer that's two or three tokens long.

You paid for an entire natural-language generation pipeline — in money and in latency — to get back a boolean.

The founder is Diogo Almeida, formerly at OpenAI, co-creator of ChatGPT and of RLHF. The way he puts it: "We have lightning in a bottle, and yet it is not useful… computers speak a different language." LLMs speak human. But the caller is a program. All that glue in between — the parsing, the retrying, the format validation — that entire layer exists to translate human back into machine.

Jev's move is to delete the translation layer. If the caller is a program, output something a program can eat directly.

Now, the API. This is the part I find genuinely elegant, because it's so constrained. Jev has exactly three question types. Three. Every decision you want to make has to be pressed into one of these three shapes.

The first one is called a **Noul** — that's yes-or-no. You give it instructions like "will this shell command cause irreversible damage?", and then you define criteria for what counts as true and what counts as false. True might be "deletes files, rewrites history, or touches production." False might be "read-only or safely reversible." And what comes back is not the word "yes." It's a number between zero and one. Point nine five. That's it.

The second type is **Choice** — pick one. You hand it a set of options, up to two hundred fifty-five of them, and it returns the selected option, plus the full probability distribution across all options, plus a confidence value. So you don't just learn that it picked "needs reasoning" — you learn it picked "needs reasoning" at point eight eight versus "simple lookup" at point one two, with confidence point eight one.

The third is **Score** — rate this thing on an ordered scale, anywhere from two to ten levels. And here's a detail I want to slow down on. Say you've got a three-level scale: zero is "not worth reading," one is "worth referencing," two is "must read." Jev doesn't return one. It returns one point zero five. That's a probability-weighted continuous value, and it carries information — it's telling you "this leans very slightly toward the next level up." For ranking, that's genuinely useful, and a plain classifier will not give it to you.

The shape of a full request is: one blob of state — a title, a blurb, a source, whatever you've got — and then a set of named questions hanging off it.

And here's the critical detail, the one that changes how you design things. **Adding more questions barely changes response time.** All the questions get scored in parallel inside the same forward pass. Which flips the whole cost calculus. You used to sit there weighing whether some judgment was worth spending an extra LLM call on. Now you can just ask twenty of them at once.

Okay. Why is it fast? This is the part people explain wrong most often. Two hundred times faster is not because Jev is a small model.

Picture what an autoregressive LLM has to do to produce N tokens. It needs N sequential forward passes. Each one waits on the previous token before it can compute the next. Token one, then token two, then token three, all the way to end-of-sequence. That sequential dependency — that chain — *is* the latency. It's also where output-token pricing comes from.

Jev's architecture is a single parallel forward pass. It doesn't produce an answer. It scores every possible answer in your schema simultaneously. Because the options are finite and known up front, nothing has to be generated step by step. You compute once, and you read off the distribution. Option A: point eight eight. Option B: point one two. Score: one point zero five. All at the same moment.

Which is why the pricing reads four point two cents per million input tokens, and output is *free*. There is no output to charge for. Measured latency lands somewhere between seventy and five hundred milliseconds.

The training method is called RLCD — Reinforcement Learning for Calibrated Decisions — and it uses synthetic data exclusively. The objective function isn't "does this sound human." It's "is this probability estimate accurate."

Now let's deal with the claim that's all over the internet: that Jev "mathematically cannot hallucinate." Half of that is true. And it's being read as the other half.

What's guaranteed is the output *space*. Not output *correctness*.

The schema constraints genuinely do make it impossible to return anything outside the type. You define three options, it will not invent a fourth. It will not return malformed JSON. It will not suddenly start explaining itself in a paragraph. That kills an entire category of format failures, which — let's be honest — is the single most annoying class of bug in LLM structured output.

But it will still pick B when the right answer was A. A misclassification is not a hallucination. And yet, for your system, the consequence is exactly identical.

The failure mode that actually deserves your attention is this one: **if the correct answer isn't in your schema, Jev will confidently pick the least-wrong option available.** There's no "I don't know" exit unless you build one. Imagine a Noul with only two outcomes, safe and risky. Now feed it a command it genuinely doesn't understand. It returns a probability somewhere in the middle. And what lands in your code is a perfectly normal-looking float. Nothing about it looks wrong.

So the practical rule: any open-ended judgment should include an "unknown" or "needs escalation" option, with a confidence threshold that kicks things upward.

Which brings me to the metric that actually matters, and it's not accuracy. It's **calibration**. When the model says point seven, is it actually right seventy times out of a hundred?

General-purpose LLMs are bad at this. Ask a GPT-class model to give you a zero-to-one confidence score and you get a pile of point eight fives and point nines. Not because that reflects any internal uncertainty — but because that's how humans write confidence scores in the training corpus. The number has no statistical meaning. You cannot threshold on it.

Jev's entire training objective is to make that number usable. That's what the confidence field is for. And it lets you write a very specific shape of code: if risk is above point nine, block it — confidently dangerous. If risk is below point one, let it through — confidently safe. Everything in between? *That's* when you escalate to a big expensive model.

Notice what that is. It's a three-state decision, not a binary classification. And your ability to write it depends entirely on the probability being trustworthy.

The early adopter feedback points the same direction. Bryo AI, an email classification service, found Gemini slightly more accurate — but at ten to twenty times the cost. And they specifically called out that Jev gives you a *real* probability, which for workflow automation matters more than those few accuracy points.

Vercel uses it for command safety review. They swapped their safety reviewer from GPT-5.6 Luna over to Jev and report five to eighteen times faster at p95, with better accuracy. And here's the number that captures the market reaction: within twenty-four hours of launch, nearly thirteen percent of paid Vercel AI Gateway teams were using it. Fastest-adopted model in that platform's history — more than double the share of any previous launch, including the GPT-5.6 family.

That said. Every one of those forty-to-two-hundred-x, forty-to-four-hundred-x figures is vendor-reported. TypeSafe itself acknowledges the benchmark workflows were internally constructed. The right use for numbers like that is "worth spending an afternoon testing." Not "safe to cite in an architecture decision record."

So how is this different from what you're already doing? Because you might be thinking: can't I just use JSON mode?

Let me lay out four options side by side.

If you use an LLM with structured output, you get format validity. But the latency is still token-by-token, and you get no trustworthy probabilities. That approach is best when you want text and structure together.

If you use constrained decoding with a grammar, you also get format validity, still token-by-token, and the logprobs you get out are uncalibrated. Good for self-hosted models that need strict grammar.

If you fine-tune a BERT-class classifier, you get type safety, it's very fast, but you have to calibrate the probabilities yourself. Good when your labels are fixed, you have labeled data, and it rarely changes.

Jev gives you type safety *plus* calibrated probabilities, in a single forward pass, and the probabilities are the whole design goal. It fits when your labels change often, you don't have labeled data, and you need those probabilities.

So the difference from structured output is architectural. Structured output wraps constraints *around* the generation loop — but the loop is still in there. You save nothing on latency, nothing on output cost.

The difference from fine-tuning your own classifier is operational. The BERT route is equally fast and equally cheap. But you need labeled data, a training pipeline, and a full retrain to change a single label. Changing Jev's criteria means editing a string. That's an enormous difference in any scenario where label definitions shift weekly — content grading, risk assessment, that kind of thing.

The community has also connected this to DSPy's signature concept — compiling expensive LLM calls down into specialized small AI functions. Same direction of travel.

Alright, so how do you actually wire this into something you've already built? The framing that helps me: Jev doesn't replace your main model. It replaces the trivial judgments your main model has been moonlighting on.

Three patterns.

Pattern one, the **Model Router**. On the way in, have Jev assess complexity with a Choice. Simple requests go to a cheap model, complex ones go to a frontier LLM. What you save is the entire batch of calls that never needed a big model in the first place.

Pattern two, the **Safety Gate**. Judge risk with a Noul *before* the tool call executes. This is exactly what a coding agent's auto-accept mode needs — the check has to be fast enough that the user never notices it happened. Otherwise you may as well just interrupt them and ask. Vercel's usage is precisely this.

Pattern three, the **Verifier**. After the tool runs, score the result quality with a Score, and bounce it back to the model if it fails. This one deserves real attention, because it answers something head-on: verification cost determines what you can automate. When the marginal cost of verification approaches zero, every loop you abandoned because "verifying costs more than just redoing it" suddenly becomes viable again.

Put those together and the flow looks like this: request comes in, Jev routes it, simple stuff goes to a small model, complex stuff goes to the frontier LLM. The LLM emits a tool call. Jev gates it — risky things go to a human for approval, safe things execute. After execution, Jev verifies the result. Fails, it loops back to the LLM. Passes, it returns.

Which is why I'd argue Jev matters more to harness engineering than it does to model capability. It made no model smarter. It made the harness able to *afford* far denser checkpoints.

Let me make this concrete with a real system. This site runs on Cloudflare Workers. Its current Workers AI usage: bge-m3 for embeddings, a qwen 14-billion chat model for RAG answers, bge-reranker for reranking, and llama 3.1 8B for metadata extraction during ingest.

Now reclassify every one of those calls as either generation or decision, and the picture gets stark.

RAG answer generation — that's generation. It needs to talk. Leave it alone. Article embedding — that's embedding, leave it alone.

But then: daily mining A/B/C grading, where an LLM reads a title and blurb and assigns a grade — that's a decision. That's a three-level Score. Whether a query should go through RAG at all — decision, and right now there's no check, it just always does. That's a Noul. Fact-check's four verdicts — decision, that's a Choice with four options. Is a term hardcore enough to be worth defining — decision, Noul. Are these two tags duplicates or synonyms — decision, Noul. Should a submitted FAQ be approved — currently manual, but that's a Noul plus a confidence threshold.

Six of eight nodes are decisions, not generation. And all six of them are currently paying generation prices.

The wiring, by the way, is the easy part. The AI binding already exists in the Wrangler config. And Jev is already listed on Cloudflare Workers AI under the model ID typesafe-slash-jev, with a thirty-two-thousand-token context window. No new API key. No new SDK. No deployment changes.

So the before-and-after looks like this. Before: one LLM call to llama 3.1, with a prompt that says "is this worth writing about, reply A, B, or C only" — and then you parse the text, and you handle the case where it explains itself anyway, and you handle the case where it replies with the Chinese characters for "grade B" instead of the letter B.

After: one call to Jev, with a state blob and three questions. A Score asking whether this is worth rewriting into a deep technical article. A Noul asking whether this is product marketing rather than technical content. And a Choice for topic classification across agent, RAG, infra, model, other. Typed values come back, and you branch on them directly — score above one point five, confidence above point eight, that's a high-confidence must-read.

And the difference isn't just cost. The second version has no parsing, no retries, and no defensive code for "the model added another sentence again." Plus those two extra questions cost essentially no additional time.

Now — when should you *not* use this?

Use it for classification, routing, scoring, risk assessment. Dead center of the design goal. Use it when you need real probabilities to set thresholds — calibration is the whole point. Use it when you have multiple independent questions about one input, because parallel scoring is nearly free.

Don't use it for producing any text a human will read. It doesn't talk. Don't use it for summarizing, translating, rewriting — same reason. Don't use it when the decision requires a lookup or a tool call first, because a single forward pass can't chain and can't ask you a follow-up question. Don't use it above two hundred fifty-five options — that's a hard limit. And don't use it when you need to explain the reasoning, because you get probabilities and nothing else.

Two more that are "proceed with caution." If your labels are fixed and you have plenty of labeled data, a fine-tuned small model may just be cheaper. And if your input exceeds thirty-two thousand tokens, you're over the Cloudflare context limit.

That "no rationale" point deserves emphasis. If your scenario requires justifying a decision to a user — content moderation appeals, risk-based rejections — Jev hands you point nine three and absolutely nothing else. The sensible architecture there is: Jev filters first, and the small number of blocked cases go to an LLM for an explanation. That's only a few percent of your traffic anyway.

Two operational realities. Availability: it launched on September fifteenth behind a waitlist, but demand was heavy enough that the waitlist got removed around September twentieth or twenty-first. Anyone can sign up now at console dot typesafe dot ai with five dollars of free credit — roughly a hundred and twenty million tokens. So if you want to test it, there's nothing to wait for. Stability: the API was overwhelmed during launch week. You'll see 429s for rate limiting and 529s for overload, and you need exponential backoff. The official SDKs handle it for you — but if you're calling through Workers AI, you're implementing that yourself.

So let me pull this together.

The model is named Jev after the economist William Stanley Jevons. Jevons paradox: when a resource gets more efficient and cheaper per unit, total consumption doesn't fall — it explodes.

And honestly, that naming is the most candid part of this whole story. If decision costs really do drop four hundred times, the outcome is not a bill that's four hundred times smaller. The outcome is that you will stuff twenty previously-unaffordable questions into every single agent loop. Verify every step. Gate every tool call. Score every retrieved chunk. Judge every turn's continuation.

Which loops back to an old point: the model isn't dumb, it lacks guidance. And the density of that guidance has always been capped by cost. Jev made no model smarter. It just pushed the marginal cost of asking one more question toward zero.

So — three things to carry out of this.

One: the speed isn't about model size. It's that the generation loop is gone entirely. A single parallel forward pass scoring a finite answer space, which is also why output tokens are free.

Two: "cannot hallucinate" means the output *space* is guaranteed, not the output *correctness*. It'll still pick wrong. And if the right answer isn't in your schema at all, it'll confidently pick the closest thing — so always build yourself an escape hatch.

Three: the real product here is calibrated probability. Accuracy you can get elsewhere. A number you can genuinely threshold on — where point seven means seventy percent — is what lets you write three-state logic instead of binary logic, and that's the thing that changes your architecture.

And if you want one concrete action: go list every LLM call in your system, and mark the ones that only need a typed value back. That list is going to be longer than you expect.

🇹🇼 中文

這週中文技術圈有兩個頻道同時在講同一個模型,一個問「快兩百倍又便宜四百倍,這是真的嗎」,另一個問「AI Agent 是不是真的要大改了」。被講的這個東西叫 Jev,2026 年 9 月 15 號才開放早期存取。

它特別的地方不在跑分,而在於——它根本不生成文字。你問它問題,它回你的是帶機率的型別值,設計給程式讀,不是給人讀。

我們今天拆三層:它實際怎麼運作、「不會幻覺」這個說法錯在哪、還有你要怎麼判斷自己系統裡哪些節點該換掉。

先講一句總結。Jev 是 TypeSafe AI 所謂的 System One Model,輸入非結構化的狀態,輸出型別化的機率決策。它背後的核心洞見是這樣的:agent loop 裡面有一大半的模型呼叫,其實根本不需要「說話」,可是它們都在付說話的錢。它快、它便宜,不是因為模型小,是因為 LLM 延遲跟 output 計價的來源——那個逐 token 的生成迴圈——在它身上不存在。

## 為什麼一個不會說話的模型會紅

想一下 agent 實際在做什麼。標準的 loop 就是:LLM 決定下一步、工具執行、結果回填、LLM 再決定。

問題出在「決定」。每一次判斷——這個請求該路由到哪個模型、這個 tool call 危不危險、這份檢索結果相不相關、這一輪要不要結束——都是一次完整的 LLM 呼叫。而每一次呼叫都得走完整個生成迴圈,一個 token 一個 token 吐出來,只為了拿到「yes」或者「risky」這種兩三個 token 的答案。

你為了一個布林值,付了一整套自然語言生成的錢,還有它的延遲。

創辦人 Diogo Almeida 是前 OpenAI 的人,ChatGPT 跟 RLHF 的共同建造者。他的說法很有畫面感:我們手上握著瓶中閃電,卻沒辦法好好用,因為電腦說的是另一種語言。LLM 說人話,但呼叫它的是程式。中間那層 parse、retry、驗格式的膠水,全都是為了把人話翻回程式話。

Jev 的做法就是把這層翻譯整個拆掉。既然呼叫方是程式,那就直接輸出程式吃得下的東西。

## 三個原語

Jev 的 API 只有三種問題類型,這個限制是刻意的——所有決策都得先被壓成這三種形狀之一。

第一種叫 Noul,是非題。你給它一段指示,比如「這個 shell 指令會不會造成不可逆的破壞」,再給它 true 跟 false 各自的判準,它回你一個零到一的機率。就這樣,一個浮點數。

第二種是 Choice,單選題。它回你選中的那一項,外加完整的機率分布跟一個信心值。比如它告訴你這個請求該走 needs_reasoning,機率零點八八,另一個選項零點一二,信心零點八一。單一問題最多可以有兩百五十五個選項。

第三種是 Score,評分。在一個二到十級的有序量尺上定位。這裡有個細節我覺得很漂亮:它回的分數可能是一點零五,而不是整數的一。因為那是機率加權之後的連續值,它帶著「有點偏向下一級」這個資訊。你拿來做排序的時候,這個小數點後面的東西是有意義的,純分類器給不了你。

然後完整的請求長什麼樣子?一個 state,配一組 questions。你把文章的標題、摘要、來源丟進 state,然後一次問它三個問題:值不值得讀、是不是廣告、主題歸在哪一類。

這裡是關鍵——多加問題,幾乎不增加回應時間。所有問題是在同一次 forward pass 裡面平行評分的。這件事徹底改變了成本結構。以前你會精算「這個判斷值不值得多開一次 LLM 呼叫」,現在你可以一口氣問二十個。

## 為什麼快

這段最容易被講錯。「快兩百倍」不是因為 Jev 是個小模型。

自迴歸的 LLM,要產生 N 個 token 就需要 N 次序列化的 forward pass,每一次都得等前一個 token 出來才能算下一個。這個序列依賴就是延遲的根源,也是 output token 計價的根源。

Jev 的架構是單次平行的 forward pass。它不「產生」答案,它是對 schema 裡面所有可能的答案同時打分。因為選項是有限的、事先就知道的,所以不需要逐步生成——一次算完,取機率分布。

你可以這樣想像兩張圖:左邊是傳統 LLM,prompt 進去,token 一、token 二、token 三,一路排隊排到 EOS 才結束。右邊是 Jev,state 加 questions 進去,一次平行運算,同時噴出選項 A 零點八八、選項 B 零點一二、score 一點零五。中間那條線就寫著四個字:拿掉生成迴圈。

所以它的定價才會變成 input 每百萬 token 四分二美元,output 完全免費——因為根本沒有 output 這個東西可以收你錢。實測延遲落在七十到五百毫秒。

訓練方法叫 RLCD,Reinforcement Learning for Calibrated Decisions,完全使用合成資料。它的目標函數不是「講得像不像人話」,而是「機率估得準不準」。

## 「數學上不會幻覺」是被誤讀的

網路上到處在寫 Jev「mathematically cannot hallucinate」。這句話有一半是對的,但大家理解成了另外一半。

它保證的是輸出空間,不是輸出的正確性。

Schema 約束確實讓它不可能回傳型別以外的東西。你定義三個選項,它不會發明第四個,不會給你 malformed JSON,不會突然開始解釋自己。這消滅的是格式類的失敗,而那確實是 LLM structured output 最惱人的一類 bug。

但是——它照樣會把該選 A 的選成 B。分類錯誤不叫幻覺,可是對你的系統來說,後果一模一樣。

更值得警惕的是一個具體的失敗模式:如果你的 schema 裡面沒有正確答案,Jev 會很有自信地選出次好的那一個。它沒有「我不知道」這個出口,除非你自己把它加進選項裡。一個只有 safe 跟 risky 兩選的 Noul,遇到「這個指令我根本看不懂」的情況,只會回你一個介於中間的機率——而你拿到的,是一個看起來非常正常的浮點數。

所以實務上的規則很簡單:任何開放式的判斷,選項裡一定要留一個 unknown 或者 needs_escalation,並且對低信心的結果設門檻往上送。

## Calibration 才是真正該看的指標

比準確率更重要的是 calibration。白話講就是:當模型說零點七的時候,那一百次裡面,是不是真的對了七十次?

這件事一般的 LLM 做得非常差。你叫 GPT 那類模型輸出一個零到一的信心分數,它會給你一堆零點八五、零點九,因為那是訓練語料裡面人類寫信心分數的習慣,不是它真實的內部不確定性。那個數字沒有統計意義,你沒辦法拿來設門檻。

Jev 的整個訓練目標就是讓這個數字可用。這才是 confidence 這個欄位存在的意義。它讓你可以寫出這樣的程式:拿到風險機率,大於零點九就直接擋、小於零點一就直接放行,中間那段灰色地帶,才動用大模型去處理。

注意這是一個三態決策,不是二元分類。而能這樣寫的前提,就是機率可信。

早期採用者的回饋也指向這一點。email 分類服務 Bryo AI 測下來認為 Gemini 準確度是略高一點,但成本是十到二十倍,而且他們特別強調一句:Jev 給的是真實機率。對自動化流程來說,這件事比那幾個百分點的準確率更關鍵。

Vercel 則是拿它來做指令安全審查。原本跑在 GPT-5.6 Luna 上的 safety reviewer 換成 Jev 之後,回報 p95 快了五到十八倍,而且準確度更高。附帶一個能說明市場反應的數字:Jev 上線二十四小時內,就有將近百分之十三的 Vercel AI Gateway 付費團隊在用,是那個平台史上採用最快的模型,比 GPT-5.6 系列發布的時候還快一倍以上。

不過我得講清楚:所有那些四十倍、兩百倍、四百倍的數字,全部都是廠商自報的,TypeSafe 自己也承認 benchmark 用的 workflow 是他們內部建構的。這類數字的正確用法是「值得花一個下午測測看」,不是「可以寫進架構決策文件」。

## 跟你已經在用的東西差在哪

你可能會想,我用 structured output 不就好了嗎?

我們一個一個比。LLM 加 JSON mode,它保證的是格式合法,但延遲還是來自逐 token 生成,而且它給不了你可信的機率。它適合的場景是你需要同時產生文字跟結構。

Constrained decoding,也就是用 grammar 約束的那套,一樣保證格式合法,一樣還在逐 token 生成,logprob 是有,但沒校準過。適合自架模型、需要嚴格文法的情況。

自己微調一個 BERT 類的分類器,型別安全、極快、也很便宜,但機率要自己校準。適合標籤固定、你手上有標註資料、而且不常改的場景。

Jev 則是型別安全加上校準過的機率,單次 forward pass,而給你機率本來就是它的設計目標。適合標籤常改、零標註資料、而且你需要機率的場景。

所以跟 structured output 的差別是架構性的:那個方案只是在生成迴圈外面加約束,迴圈本身還在,延遲跟 output 費用一毛錢都沒省。

跟自己微調小分類器的差別則是營運性的。BERT 那條路同樣快、同樣便宜,但你需要標註資料、需要訓練流程,改一個標籤就要重訓一輪。而 Jev 要改判準,只是改 instructions 那個字串而已。在標籤定義每週都在動的場景——像是內容分級、風險判定——這個差別巨大。

社群裡面也有人把它連到 DSPy 的 signature 概念:把昂貴的 LLM 呼叫「編譯」成專用的小型 AI function。方向是一致的。

## 三個接進現有 pipeline 的 pattern

講清楚一件事:Jev 不取代你的主模型,它取代的是主模型正在兼差做的那些瑣碎判斷。

想像一條流程。請求進來,第一關是 Jev 用 Choice 做路由,簡單的丟給小模型或者直接查表,複雜的才給 frontier LLM。大模型產生了 tool call 之後,第二關是 Jev 用 Noul 當安全閘,判定有風險就轉人工確認,安全才放行執行。工具跑完,第三關是 Jev 用 Score 驗證結果品質,不合格就打回去給大模型重跑,合格才回傳。

這三個關卡就是三個 pattern。

第一個是 Model Router,省下的是整批本來就不該用大模型的呼叫。

第二個是 Safety Gate,在 tool call 執行「之前」判風險。這就是 coding agent 的 auto-accept 模式背後需要的東西——判斷要快到使用者感覺不到,不然還不如直接跳出來問。Vercel 的用法正是這個。

第三個是 Verifier,這一點我覺得特別值得注意,因為它正面回應了 Loop Engineering 那篇的結論:驗證成本決定你能自動化什麼。當驗證的邊際成本趨近於零,原本那些「驗證比重做還貴」所以不得不放棄的迴圈,突然全部變得可行。

這也是為什麼 Jev 對 Harness Engineering 的意義,大於它對模型能力的意義。它沒讓任何模型變聰明,它是讓 harness 能負擔得起更密的檢查點。

## 拿這個站當實例

講抽象的不如看真的。這個站跑在 Cloudflare Workers 上,目前的 Workers AI 用量大概是:bge-m3 做嵌入、qwen 十四 B 做 RAG 問答、bge-reranker 做重排、llama-3.1-8b 在 ingest 的時候抽 metadata。

把這些呼叫按照「生成」還是「決策」重新分類一次,結果很明顯。

RAG 問答產生回覆,那是生成,不動,那本來就是要說話的。文章向量化是嵌入,不動。

然後是決策的部分:每日挖礦的 A、B、C 分級,現在是 LLM 讀標題摘要打分,其實就是一個三級的 Score。查詢該不該走 RAG,現在根本沒判斷,一律走,這應該是一個 Noul。事實查證那四種 verdict,現在是 LLM 產生判斷,其實是一個四選項的 Choice。術語夠不夠硬核,Noul。tag 是不是重複或同義,Noul。FAQ 送審該不該過,現在是人工,可以做成 Noul 加上信心門檻。

八個節點裡面,六個是決策,不是生成。而這六個現在全都在付生成的錢。

更省事的是接法。這個專案的 wrangler 設定裡面,AI binding 已經存在了。而 Jev 已經上架 Cloudflare Workers AI,模型 ID 就是 typesafe 斜線 jev,context window 三萬二。也就是說,不需要新的 API key、不需要新的 SDK、不需要改部署設定。

原本的寫法是:呼叫 llama,丟一段 prompt 說「這則新聞值得寫嗎,只回 A 或 B 或 C」,然後你要 parse 文字、要處理它多嘴解釋的情況、要處理它回「B 級」而不是「B」的情況。

改完之後是:一次呼叫,問它三個問題——值不值得寫成深度技術文章、這是不是產品行銷稿、主題歸在哪一類——然後直接拿到型別值,寫個判斷式,分數大於一點五而且信心大於零點八,就是高信心的「必讀」。

差別不只是省錢。改完之後那段程式碼,沒有 parse、沒有 retry、沒有那些「模型又多講了一句廢話」的防禦性程式碼。而且多問的那兩個問題,幾乎不花額外時間。

## 什麼時候不該用

這是決定要不要導入的實際判準。

該用的情況:分類、路由、評分、風險判定,正中設計目標。需要機率來設門檻,calibration 就是它的核心賣點。同一份輸入要問多個獨立問題,平行評分幾乎免費。

不該用的情況:要產生任何人類要讀的文字,不行,它不會說話。要摘要、翻譯、改寫,同上。決策之前需要先查資料或呼叫工具,不行,單次 forward pass,不能 chain、不能反問。選項超過兩百五十五個,硬上限。需要解釋「為什麼這樣判」,它只有機率,沒有理由。

還有兩個要斟酌的:標籤固定而且你已經有大量標註資料,那微調一個小模型可能更划算。輸入超過三萬二 token,那是 Cloudflare 上的 context 上限。

最後那個「不能解釋理由」值得單獨強調。如果你的場景需要對使用者交代判斷依據——像是內容審核的申訴、風控的拒絕原因——Jev 給你零點九三這個數字,但給不了任何說明。這種時候合理的架構是:Jev 先篩,被擋下來的那少數案例,再交給 LLM 去生成解釋。反正那只佔流量的幾個百分點。

另外提醒兩個營運面的現實。第一是可用性,九月十五號發布的時候是 waitlist 限量存取,但需求太猛,九月二十到二十一號已經取消候補名單全面開放了,註冊送五塊美金額度,大概一億兩千萬 token,所以現在要測不用等。第二是穩定性,發布那一週 API 被打爆過一輪,429 限流跟 529 過載需要你自己做指數退避——官方 SDK 有內建,但你如果直接打 Workers AI,那得自己處理。

## 收尾

最後講一下這個名字。Jev 取自經濟學家 William Stanley Jevons,對應的是 Jevons paradox:當資源的使用效率提高、單位成本下降,總消耗量不會減少,反而會暴增。

我覺得這個命名是整件事裡面最誠實的部分。如果決策成本真的掉四百倍,結果不會是你的帳單少四百倍。結果是你會在每一個 agent loop 裡面塞進二十個以前捨不得問的問題:每一步都驗證、每一個 tool call 都過安全閘、每一份檢索結果都評分、每一輪都判斷要不要繼續。

所以今天有三件事值得你記住。

第一,Jev 快跟便宜的原因在架構,不在參數量。它拿掉的是逐 token 的生成迴圈,所以 output 根本沒東西可以計價。你不能用「小模型」的心智模型去理解它。

第二,「數學上不會幻覺」保證的是輸出空間,不是輸出正確性。它不會給你 schema 以外的東西,但它照樣會選錯——而且當正確答案不在你的選項裡,它會很有自信地選出次好的那個。所以永遠留一個 unknown 的出口。

第三,真正的賣點是 calibration,不是準確率。機率可信,你才寫得出三態決策:高信心擋、高信心放、灰色地帶才升級給大模型。

至於「AI Agent 要不要大改」,我的答案大概是:架構不會變,但檢查點的密度會變。以前你只在關鍵節點插驗證,因為每個驗證都是一次 LLM 呼叫;現在你可以在每一步都插。

而現在就值得做的事只有一件:把你系統裡面所有的 LLM 呼叫列出來,標記哪些其實只是要一個型別值。

那份清單的長度,會比你預期的長。

Tags

Related Articles

Harness Engineering (2): Five Engineering Answers from OpenAI's Million-Line Experiment

Three OpenAI engineers, five months, one million lines of AI-generated code, zero hand-written. The real value of this experiment isn't the numbers — it's the proof that Harness design can be engineered. Five concrete practices: making the app legible to agents, treating the repo as the source of truth, mechanizing architectural constraints, rewriting merge philosophy, and background entropy management.

Harness Engineering (3): Industry Consensus, Four Pillars, and a Three-Phase Rollout

Distilling Harness Engineering from concept and benchmark case into something you can start executing today: the four fixed failure modes of Agents, the 40% context sweet spot, the four-pillar framework the industry has converged on, and a three-phase roadmap from 'this afternoon' to 'fully automated in two weeks' — closing with six industry consensus points and three still-unsolved problems.

RAG's Five Stages: From Pipeline to Reasoning Retrieval, and the Naive RAG on My Own Site

Over the past two years RAG evolved from a 'linear pipeline' to 'loop-based reasoning'. It maps cleanly to five stages: Naive, Advanced, Modular, Graph, Agentic. The real inflection point is control moving from pipeline to agent — a System 1 → System 2 shift. Looking back at engineer-news's own RAG stack, it's stuck at the Naive edge — so this post also lays out what to fix next.