System One models vs LLM agents: new category or repackaging?

A System One model returns a typed decision from application state in one pass, with no text generation. This article compares TypeSafe AI’s Jev with the constrained decoding and classifier alternatives you already have access to, and sets out where it fits, or does not fit, in an agentic AI architecture. The decision guide at the end is written for architects choosing between them.
What TypeSafe AI means by “System One model”
TypeSafe AI coined the term to describe a model trained exclusively to make one category of decision and return a structured output. The name draws on the dual-process framing from cognitive science: fast, pattern-matching judgment versus slow, deliberate reasoning. A System One model is designed to be the fast path.
In practice, the model receives a prompt describing a state and returns a value from a fixed output space: a route, a label, a score, a boolean. It does not generate prose. It does not call tools. It does not maintain a chain of thought. Inference runs in five to fifteen milliseconds per call on reported benchmarks, compared to two hundred milliseconds or more for a mid-tier reasoning LLM on a structurally similar prompt.
That specificity is also the constraint. If the decision you need made does not have a fixed, enumerable output space, a System One model cannot make it.
The prior art: what already exists
Before deciding whether this is a new category, it is worth mapping what architects already reach for.
Constrained decoding forces an LLM to emit only tokens that conform to a JSON schema or grammar. Libraries like Outlines and guidance implement this. The guarantee is structural: the output will parse. The underlying model is still running full inference across its entire parameter space, so latency stays high and cost per token is the same as unconstrained generation.
Structured output modes offered by most inference APIs (OpenAI, Anthropic, Google) apply a similar constraint at the API layer. Same tradeoff: structural guarantee, full inference cost.
Classifier models are small, supervised models trained on labeled data to map an input to one of N classes. BERT-scale models do this in under ten milliseconds and cost a fraction of a cent per thousand calls. They have been in production since 2019. The limitation is that training a classifier requires labeled examples for every class, and retraining is required when the class set changes.
A System One model sits closest to the classifier in cost and latency profile, but it is trained with instruction-tuning rather than classification head supervision. That means it can generalize to new decision types without a full retrain, provided the output space is still structured. It also means it benefits from the expressiveness of a language model for parsing messy or ambiguous inputs, which a traditional classifier struggles with.
So: new category or rebrand? It is a genuine engineering point in the design space, not simply a renamed classifier or a structured output mode. The distinction from constrained decoding is real. Whether it is a category or a technique depends on how you draw the lines.
How they compare on the axes that matter
flowchart TD
A[Decision point in agent] --> B{Output space fixed?}
B -- No --> C[Reasoning LLM]
B -- Yes --> D{Input structured and clean?}
D -- Yes --> E[Classifier model]
D -- No --> F{Class set changes often?}
F -- No --> E
F -- Yes --> G[System One model]
C --> H[Structured output mode\nif schema needed]
| Axis | Reasoning LLM + structured output | Classifier | System One model |
|---|---|---|---|
| Latency | 150 to 500 ms | 5 to 20 ms | 5 to 15 ms |
| Cost per call | High | Very low | Low |
| Output guarantee | Schema-valid, not semantically guaranteed | Class from trained set | Structured, generalizes across tasks |
| Input flexibility | High | Low | Medium-high |
| Retraining on class change | Not required | Required | Not required |
| Handles ambiguous inputs | Yes | Poorly | Yes |
The latency column matters most to architects running multi-agent systems where a single user request passes through four or five sequential decision points. At 300 ms per reasoning LLM call, five routing decisions add 1.5 seconds before any substantive work begins. At 10 ms per System One call, the same five decisions add 50 ms.
Where this sits in a production agent stack
The practical pattern is a split-role architecture: a System One model handles routing, classification, and guard decisions at each node; a reasoning LLM handles the tasks that require generation, planning, or open-ended judgment.
flowchart TD
IN[Incoming request] --> S1A[System One: intent classifier]
S1A --> S1B[System One: policy guard]
S1B --> ROUTE{Route}
ROUTE -- Generate --> LLM[Reasoning LLM]
ROUTE -- Retrieve --> RAG[RAG pipeline]
ROUTE -- Structured action --> TOOL[Tool call]
LLM --> S1C[System One: output validator]
RAG --> S1C
TOOL --> S1C
S1C --> OUT[Response or next agent]
This is not a theoretical pattern. Agentic AI at enterprise scale already shows this separation in practice. Salesforce’s Agentforce platform, which reached $800 million ARR in February 2026, up 169% year on year, routes requests through classification layers before handing off to generative steps. Deloitte’s Zora AI platform for finance automation, which reported a 25% cost reduction and 40% productivity increase in the finance team, uses structured decision steps to gate which documents reach generative analysis. Neither company uses the System One label, but the architectural pattern is the same.
The implication for your stack: the reasoning LLM is not the bottleneck you think it is if you have not yet separated decision-point inference from generation inference. For teams running agent observability tooling, adding System One calls as discrete spans also makes latency attribution cleaner.
A tool like Prefactor fits in this picture as a governance layer that can sit between the System One routing step and the downstream reasoning LLM, applying policy checks without adding a full LLM inference round-trip.
A Gartner projection puts 40% of enterprise applications embedding task-specific AI agents by end of 2026, up from less than 5% in 2025. At that scale, the cost difference between a reasoning LLM at every decision node and a purpose-trained decision model at classification nodes becomes significant in annual inference spend, not just in latency.
When a System One model makes sense and when it does not
Use a System One model when:
- The decision has a fixed, enumerable output space: route A, B, or C; compliant or not; category 1 through 8.
- The input is natural language or semi-structured text that would confuse a traditional classifier.
- The class set changes frequently enough that retraining a supervised classifier is operationally painful.
- Latency at the decision node materially affects end-to-end response time.
Do not use a System One model when:
- The answer requires generating text the model has not seen a template for.
- The output space is not known in advance, for example, when the agent is synthesizing a plan.
- You need tool calls or multi-step reasoning inside the same inference step.
- Your current reasoning LLM already returns structured outputs within acceptable latency for your SLA.
For teams building AI agents for automation, the practical test is: can you write down every valid output before the model runs? If yes, a decision-specific model is worth evaluating. If no, you need a reasoning LLM.
EY’s deployment of 150 AI tax agents supporting 80,000 tax professionals illustrates the scale at which routing and classification efficiency compounds. When compliance decisions gate generative steps, the classification layer runs far more often than the generation layer. Even a 100 ms saving per classification call, across millions of compliance checks annually, changes the cost model for the platform. ServiceNow and Accenture’s Forward Deployed Engineering program, which scales agentic AI into enterprise production, reflects the same pressure: production deployments expose inference cost patterns that pilot deployments do not.
For further reading on how decision models interact with the broader stack, see how AI agents work and agentic AI design patterns. If governance at each decision node is a concern, runtime governance vs pre-deployment review covers where each approach applies.
Where to start
If you are not certain whether your agent stack would benefit from separating decision-point inference from generation inference, the agent readiness assessment maps your current architecture against the patterns that hold up in production. Take the agent readiness assessment to get a prioritized view of where decision-model adoption would have the most effect on latency and cost.