Getting Your First Agent Live

How much can Jev cut agent inference costs? What $0.042/M changes

Matt DoughtyMatt DoughtyCEO & Co-Founder, Prefactor
6 min read
Abstract illustration: How much can Jev cut agent inference costs? What $0.042/M changes

What this means for your agent’s cost structure

If classification, routing, and triage are running on the same model you use for reasoning, you are paying reasoning prices for decisions that do not require reasoning. TypeSafe’s Jev model, currently in early access, is priced at $0.042 per million input tokens with output tokens free. The company also claims inference speeds 40 to 200 times faster than comparable models. Both figures are vendor claims from a pre-general-availability release, and you should treat them as a planning input rather than a guarantee. But even applied conservatively, they point to a real architectural choice that CTOs budgeting a production rollout should price out now.

This article works through which costs fall when you move decision-only steps to a model like this, which costs stay exactly where they are, and what the numbers look like at production scale.


The anatomy of an agent’s inference bill

A production AI agent workflow is not a single inference call. It is a chain: classify the input, route it to the right handler, check preconditions, call tools, reason over results, and generate a response. Each step that touches a model generates a token cost and adds latency.

The steps split into two groups with very different computational profiles.

Decision steps take a fixed input and return a category, a score, or a binary. Intent classification (“is this a refund or a status query?”), triage (“does this ticket need a human?”), and routing (“which sub-agent handles billing?”) are all decision steps. The input context is often small, the output is a label or a short token sequence, and the computation required is shallow.

Reasoning steps take a problem and work through it. Writing code, drafting a response, synthesising tool outputs, and evaluating ambiguous evidence all require a model that can hold and manipulate extended context. Cutting corners here produces wrong answers.

The mistake in most early production deployments is using a reasoning-grade model for both groups. You pay reasoning prices for what is, in effect, a lookup table.

flowchart TD
    A[Incoming request] --> B{Classify intent}
    B -->|Decision model| C{Route to handler}
    C -->|Decision model| D[Precondition check]
    D -->|Decision model| E[Reasoning model]
    E --> F[Tool calls]
    F --> G[Synthesise results]
    G -->|Reasoning model| H[Response]

    style B fill:#d4edda,stroke:#28a745
    style C fill:#d4edda,stroke:#28a745
    style D fill:#d4edda,stroke:#28a745
    style E fill:#cce5ff,stroke:#004085
    style G fill:#cce5ff,stroke:#004085

The green nodes are candidates for a decision-only model. The blue nodes are not.


What falls when you separate the layers

Per-call inference cost

At $0.042 per million input tokens, a routing call with a 500-token context costs roughly $0.000021. At $3.00 per million input tokens on a mid-tier reasoning model, the same call costs $0.0015, about 71 times more. If your agent handles 10 million routing decisions per month, that difference is $14,790 per month, from a single step. Add classification and triage on similar volumes and the gap compounds.

General Mills runs its supply chain optimisation agent across more than 5,000 shipments per day, with 70% of routing recommendations accepted automatically. At that decision volume, the model tier chosen for routing is not an architectural detail; it is a line item.

Latency-driven timeouts and retries

A 200x speed improvement does not mean your agent responds 200x faster overall. The reasoning steps still take as long as they take. What changes is the idle time before reasoning starts. If your current routing layer adds 300ms per call and that drops to under 5ms, a five-step pipeline that chains three routing decisions loses roughly 885ms of dead wait time per request.

That matters because timeouts in multi-agent systems are often set conservatively to absorb slow routing, and slow routing generates retries when upstream services time out. Retries multiply inference costs and can trigger cascading failures in pipelines where downstream agents wait on upstream decisions. Shrinking the decision latency shrinks the retry surface.

Klarna’s GPT-4-based customer service agent reported 82% faster response times alongside its $60 million in annual savings. Response time is a cost driver: slower responses increase the rate at which customers abandon sessions, escalate to human agents, or retry, each of which has a real cost that does not appear in the inference bill.


What does not fall

The reasoning steps

A decision-only model is not appropriate for the steps that require it to be wrong in interesting ways. Code generation, document synthesis, multi-hop reasoning, and nuanced customer-facing responses need a model that can handle them. Routing more cheaply does not change what you spend on those steps.

Morgan Stanley’s DevGen.AI agent reviewed 9 million lines of COBOL and Perl code, saving an estimated 280,000 developer hours. That kind of work cannot be handed to a classifier. The value comes from the reasoning layer, and the reasoning layer still needs to be paid for accordingly.

The cost of wrong decisions

Cheaper routing does not make routing more accurate. If your intent classifier misroutes 3% of requests now, it will misroute approximately 3% of requests at $0.000021 per call. The downstream cost of a wrong decision, failed transactions, incorrect escalations, user-facing errors, and the human time to correct them, does not change with the token price.

This is the part of the economics that gets omitted from vendor comparisons. Before you commit budget to a high-volume decision layer, you need an honest agent evaluation of what accuracy your classifier actually achieves, and what each category of error costs your business.

flowchart TD
    A[Cost levers in a production agent] --> B[Inference cost per call]
    A --> C[Latency overhead]
    A --> D[Retry and timeout cost]
    A --> E[Cost of wrong decisions]

    B -->|Falls with decision model| F[Addressable by Jev-class models]
    C -->|Falls with faster inference| F
    D -->|Falls as latency shrinks| F
    E -->|Unchanged| G[Not addressable by model price or speed]

    style F fill:#d4edda,stroke:#28a745
    style G fill:#f8d7da,stroke:#721c24

Sizing the opportunity against real rollout numbers

74% of executives report achieving ROI within the first year of agent deployment, but Gartner estimates that over 40% of agentic AI projects will be cancelled by end of 2027 due to escalating costs, unclear business value, or inadequate risk controls. The gap between those two figures is partly an architecture problem. Teams that route everything through a single model class hit inference budgets faster than projected, then cancel or descope.

Separating decision steps from reasoning steps is one of the cleaner ways to bring per-request costs down without touching the quality of the steps that actually matter. If you are running three to five agents across two or more teams and each agent makes multiple routing calls per request, the arithmetic is worth doing before you lock in model contracts.

Tools like Prefactor are emerging in this space to help teams apply this kind of layered model strategy in production, though you should evaluate any tooling against your own latency and accuracy requirements.

For more on how the architectural layers fit together, see agentic AI architecture and agentic AI orchestration. If you are choosing between frameworks for the reasoning layer, LangGraph vs CrewAI covers the main trade-offs. And if you are thinking about how governance sits across both layers, runtime governance vs pre-deployment review is worth reading before you finalise your stack.


Where to start

The clearest first step is to map your agent’s call graph and separate the decision steps from the reasoning steps, then price both at the model tier each actually requires. If you are not sure whether your current architecture is set up to take advantage of this kind of split, take the agent readiness assessment to identify where the gaps are before you commit to infrastructure spend.

Matt DoughtyMatt DoughtyCEO & Co-Founder, Prefactor

Founder of Prefactor, writing on the operational reality of getting AI agents into production — evaluation, observability, governance, and the plumbing assistants never needed.

Frequently asked questions

Which steps in an agent pipeline actually benefit from a decision-only model like Jev?

Classification, intent routing, triage, and threshold checks are the clearest candidates. These steps need a fast, cheap answer, not extended reasoning. The reasoning steps, tool calls, and final response generation still need a capable general model.

Does cheaper inference at the routing layer reduce the cost of wrong decisions?

No. The cost of a wrong classification, such as routing a refund request to the wrong handler, is downstream: failed transactions, retries, and human escalations. Cheaper inference lowers the per-call token cost, not the error rate. You still need to evaluate your classifier's accuracy before relying on it in production.

Are TypeSafe's published Jev numbers verified by a third party?

Not as of this article's publication date. The figures come from TypeSafe's own early-access release materials. Treat them as directionally useful for budgeting, but benchmark against your own workloads before committing to an architecture that depends on them.

How does latency at the routing layer compound across a multi-agent system?

In a pipeline where five agents each wait on a routing decision before proceeding, a 200ms routing call adds one second of idle time per request before any real work begins. At scale, that idle time drives timeout rates up and forces either longer timeout windows or more aggressive retry budgets. A sub-10ms routing call eliminates most of that overhead.

Stay ahead of the curve

No spam. Unsubscribe anytime. A resource by Prefactor.

Almost there — check your inbox to confirm your subscription.