Getting Your First Agent Live

Jev calibration: the RLCD questions TypeSafe has not answered

Matt DoughtyMatt DoughtyCEO & Co-Founder, Prefactor
6 min read
Abstract illustration: Jev calibration: the RLCD questions TypeSafe has not answered

RLCD, the method TypeSafe says gives Jev calibrated output confidence, has no published reward function, training procedure, or calibration methodology as of September 2026. Before routing any consequential decision through Jev, a technical buyer needs answers to specific questions about what the confidence score actually represents and how it behaves outside the training distribution.

TypeSafe describes RLCD, Reinforcement Learning for Calibrated Decisions, as the method that gives Jev calibrated output confidence. As of 20 September 2026, TypeSafe has not published the reward function, the training architecture, the training procedure, or the calibration methodology. The company has also not stated which statistic the confidence value is derived from. Each of those gaps corresponds to a production risk. The sections below name the risk, explain why it is load-bearing, and give you a concrete question to put in the room.


Why calibration documentation is a production requirement, not a nice-to-have

AI agent governance frameworks treat model confidence as an input to routing logic: if the score exceeds a threshold, the agent acts autonomously; if it falls below, the action goes to a human or a fallback path. That logic is only as sound as the calibration claim it rests on.

Consider what happens when calibration is wrong at the tails. A score of 0.92 should mean the model is correct roughly 92 times out of 100 on inputs similar to its training distribution. If the model was calibrated on a narrow domain and your inputs sit at the edge of that domain, the true rate may be 0.70 or lower, with no signal in the score itself that anything has changed.

According to a May 2026 survey, 22% of deployed AI agent systems deliver negative ROI at 12 months despite 88% reaching production deployment. Miscalibrated confidence that causes over-reliance on autonomous action is one of the documented contributors. The pattern is consistent: teams accept the vendor’s calibration claim, set thresholds at deployment, and discover the claim does not hold on their data only after failures accumulate.


The open questions, grouped by topic

Calibration methodology

flowchart TD
    A[Model returns confidence score] --> B{Score above threshold?}
    B -- Yes --> C[Agent acts autonomously]
    B -- No --> D[Escalate to human]
    C --> E{Was calibration validated\non your data distribution?}
    E -- Yes --> F[Threshold is load-bearing]
    E -- No --> G[Threshold is a guess]

Question 1. What statistic is the confidence value? Is it a softmax probability, a Platt-scaled output, a temperature-scaled logit, or something else? Each has different properties at the tails, and the answer determines whether post-hoc recalibration is possible on your side.

Question 2. On what dataset was calibration measured, and how was that dataset constructed? A model calibrated on internal TypeSafe benchmarks may not be calibrated on financial services workflows, healthcare triage queues, or logistics disruption events. Teams at eSentire compressed expert threat analysis from five hours to seven minutes with an AI-driven system, but they validated alignment with senior analysts at 95% before scaling. That kind of domain-specific validation is what TypeSafe’s documentation currently does not let you replicate for Jev.

Question 3. What calibration metric did TypeSafe optimise for? Expected Calibration Error, Maximum Calibration Error, and reliability diagrams each reveal different failure modes. Knowing which one TypeSafe used tells you which failure modes were not the focus.

For context on how calibration fits into broader AI governance best practices, the short version is that calibration is a property you validate per deployment domain, not once at training time.

Out-of-distribution behaviour

Question 4. How does confidence behave when the input falls outside the training distribution? Does the score compress toward 0.5, does it stay high, or does the model refuse to score? Each response has different implications for your fallback logic.

Question 5. What out-of-distribution detection, if any, is built into RLCD? If none, you need to implement it in your own agent observability layer. That is a solvable problem, but it is your problem, not TypeSafe’s, until they say otherwise.

51% of respondents report having AI agents in production today, and 78% have active deployment plans. Most of those teams are discovering out-of-distribution edge cases after go-live rather than before. The lack of published OOD behaviour from TypeSafe means you are likely to discover Jev’s edge cases the same way.

Unseen state

flowchart TD
    A[Input arrives] --> B{Is input within\ntraining distribution?}
    B -- Known --> C[RLCD score is meaningful]
    B -- Unknown --> D{What does Jev return?}
    D --> E[High confidence\nwith no signal]
    D --> F[Low confidence\nwith no signal]
    D --> G[OOD flag]
    E --> H[Agent may act incorrectly\nwith apparent certainty]
    F --> I[Agent escalates correctly]
    G --> I

Question 6. Has TypeSafe tested Jev on adversarial or novel inputs, and are those results published? Novel state is not the same as adversarial state, but both probe whether the model knows what it does not know. L’Oréal achieved 99.9% accuracy on conversational analytics queries with 44,000 monthly users, but that figure comes with a specific domain scope. Jev’s performance envelope needs the same specificity.

Question 7. What is the intended behaviour when Jev encounters a prompt structure or data schema it was not trained on? A documented fallback, such as a confidence floor or a mandatory escalation, is acceptable. An undocumented one is a gap in your AI governance framework.

Update versioning and calibration drift

Question 8. How does TypeSafe version RLCD updates, and does a model update trigger re-publication of calibration benchmarks? If Jev’s calibration shifts between versions and TypeSafe does not announce it, the thresholds you set at deployment become incorrect without any visible signal.

Question 9. Will TypeSafe provide a changelog that distinguishes changes to the reward function from changes to the base model? These are different kinds of changes with different effects on confidence score behaviour. Treating them as equivalent in a changelog would obscure meaningful calibration shifts.

For teams using an AI governance policy that requires model cards or equivalent disclosures for production components, this question is not optional. If TypeSafe cannot answer it, the governance gap is yours to document.

Putting the checklist together

Bring these nine questions to your TypeSafe evaluation conversation. Some vendors will answer several of them verbally in a technical briefing. That is useful, but ask for written documentation you can include in your internal approval process. A verbal commitment to calibration does not survive a model update or a team change.

Tools like Prefactor sit in this category of production readiness infrastructure, letting you log confidence scores, track threshold performance over time, and detect when calibration appears to have shifted. Whether you use a dedicated tool or build the logging yourself, the requirement is the same: you need a record of how Jev’s scores relate to actual outcomes in your domain, independent of TypeSafe’s claims.

The evaluate AI agents guide on this site covers the broader evaluation framework. The RLCD questions above are the Jev-specific layer on top of that foundation.


Where to start

Before your next conversation with TypeSafe, map which of the nine questions your team can answer from public documentation and which require a direct response. That gap analysis is the core of your vendor due-diligence record. Take the agent readiness assessment to identify the other production gaps in your deployment plan alongside the RLCD questions.

Matt DoughtyMatt DoughtyCEO & Co-Founder, Prefactor

Founder of Prefactor, writing on the operational reality of getting AI agents into production — evaluation, observability, governance, and the plumbing assistants never needed.

Frequently asked questions

What is RLCD and why does the missing documentation matter?

RLCD stands for Reinforcement Learning for Calibrated Decisions, the training approach TypeSafe uses to produce Jev's confidence scores. Without the reward function, calibration methodology, or training data description, you cannot independently verify whether those scores are reliable in your domain or only in the domains TypeSafe tested.

Can I just test Jev on my own data to fill these gaps?

Internal testing will tell you how Jev performs on data you have today, but it will not tell you how confidence scores behave on inputs the model has never seen, or whether a model update will silently shift calibration. Both of those require disclosures from TypeSafe that have not been made.

What is the difference between accuracy and calibration?

A model is accurate when it gives the right answer. A model is calibrated when its stated confidence matches its actual frequency of being right: a calibrated model that says 80% confidence should be correct roughly 80% of the time across a large sample. A model can be accurate on average but badly miscalibrated at the tails, which is where production failures tend to concentrate.

How does model update versioning affect a production agent?

If TypeSafe updates RLCD without publishing a versioned changelog or re-running published benchmarks, a confidence score of 0.87 after the update may not mean the same thing as 0.87 before it. Agents that gate actions on confidence thresholds need to know when the underlying calibration shifts, so they can re-validate those thresholds against the new model.

Stay ahead of the curve

No spam. Unsubscribe anytime. A resource by Prefactor.

Almost there — check your inbox to confirm your subscription.