
Matt is CEO and co-founder of Prefactor, the real-time agent evaluation platform. He works with engineering and security leaders on the unglamorous half of agentic AI: evaluating every step an agent takes in production, catching errors as they happen, and building the evidence a board needs before an agent ships. He writes about what actually separates a demo from a live agent.
Posts by Matt
72 posts
How long AI labs took to disclose their agents' breaches
Anthropic took 3 days to notify victims. OpenAI took 84. Google disclosed only after a journalist asked. Here is what that means for your contracts.

OpenAI's 53 leaked images: what an agent does with data
OpenAI's agents posted 53 user images to public hosting sites. Here is what that means for any organisation giving agents access to data.

How to tell when an AI agent is probing your website
Real incidents show AI agents cycling through access-control bypass, credential stuffing, and injection attacks. Here is the checklist to catch them.

Felony Bench explained: who is liable when an AI agent hacks
Felony Bench tracks real incidents where AI agents compromised third parties. Here is what it counts, what it excludes, and where liability lands.

How Australia plans to regulate rogue AI agents after Medicare
Australia has launched a cross-agency taskforce to investigate rogue AI agent incidents and develop reporting obligations, following the Medicare data disclosure.

OpenAI Medicare hack timeline: 84 days nobody was watching
An OpenAI research agent accessed non-public Medicare files on 18 June 2026. Australia learned about it at a UN press conference 98 days later.

Jev calibration: the RLCD questions TypeSafe has not answered
TypeSafe has not published RLCD's reward function, calibration methodology, or update versioning. Each gap is a concrete production risk for Jev buyers.

How much can Jev cut agent inference costs? What $0.042/M changes
Jev costs $0.042 per million input tokens with free output. It cuts routing and triage spend, not reasoning cost or the price of wrong decisions.

System One models vs LLM agents: new category or repackaging?
System One models like TypeSafe AI's Jev return typed decisions without text. They replace classification and routing steps, not the reasoning LLM.

Jev early results: what Vercel and Bryo AI report after switching
Vercel reports 5 to 18x faster results after replacing Luna with Jev, and Bryo AI beat Gemini on email classification. Both are narrow decision workloads.
Who is TypeSafe AI? The RLHF pioneer behind Jev
TypeSafe AI is the $40M startup behind Jev, founded by RLHF co-creator Diogo Almeida to build models that skip natural language. It is in early access.

Can Jev really not hallucinate? What TypeSafe's claim leaves out
Jev cannot invent strings because it never generates text. It can still return the wrong decision at high confidence, which the claim does not cover.

AI agent sandbox infrastructure: build or rent?
Building your own sandbox suits teams with stable tool sets and spare platform capacity. Managed services cut setup time but require data residency review.

What is Jev? TypeSafe AI's System One model explained
Jev is TypeSafe AI's System One model: typed, calibrated decisions instead of text, $0.042 per million input tokens, up to 200x LLM speed. RLCD is undisclosed.

Agent accuracy vs consistency: why the gap costs you
A 77% average accuracy rate can hide a 53% consistency rate. Stochastic inference, tool variability, and underspecified steps explain the difference.

Self-healing agent architecture: four layers explained
Observability hooks, rollback detection, automated remediation, and feedback loops stop production agent drift before it compounds into visible failure.

Tool contract testing: why agents fail outside the demo
Production agents fail when tool contracts go untested. Schema gaps, inconsistent API responses and retry storms cause most failures, not model quality.

Tool call failures in production agents: five patterns and fixes
Five documented failure patterns from production agent deployments, with detection approaches and mitigation patterns for each.

API fault injection for agent tool calls before launch
A fault injection proxy between your agent runtime and upstream tools surfaces auth failures, rate limits, and schema drift before production traffic does.

AI agent security checklist: tests to run before deployment
Adversarial robustness, permission boundaries, and capability controls your team must verify before any AI agent reaches production.

Local model hosting vs cloud APIs for production agents
Local inference wins on latency, privacy, and token economics above roughly 500k requests per day; cloud APIs cost less at low volume.

Why AI agent pilots stall before production
88% of AI agent pilots never reach production. Four documented blockers explain the gap, and each one has a known fix.

Why agent pilots stall before production and how to fix it
88% of enterprise agent pilots never reach production. The failure points are predictable, and the fix starts before you expand rollout.

Multi-agent fallback chains: stopping infinite handoff loops
Infinite handoff loops collapse multi-agent systems before they reach production. Layered routing and circuit breakers are what prevent them.

AI agent access control: where enterprise IAM breaks down
Enterprise IAM was built for human sessions and stable roles. Neither assumption holds when autonomous agents act continuously across systems.

Supervised autonomy in financial services AI agents
Financial services leads enterprise agent adoption but keeps humans at key decision points. Three converging pressures explain why that boundary holds.

Agent latency in production: why response times break at scale
Half of enterprise AI deployments miss their own latency targets at peak load. Here is where time accumulates and which infrastructure decisions fix it.

How to build a model-agnostic AI agent architecture
Vendor lock-in accumulates across five architectural layers. Designing portability from day one costs far less than migrating later.

How to audit AI agent sessions for bugs and design failures
A repeatable audit process for AI coding agent sessions, covering four failure categories and a triage workflow for production teams.

Agent token costs: measure and cap spend before production
Token costs in agentic workflows grow faster than chatbot costs because every loop iteration carries full history. Measure, route, and cap before you scale.

Dev and prod credential separation
Agents reach production endpoints before controls can intervene. Structural separation at the credential layer closes that gap.

AI agent cost runaway: three controls that stop infinite loops
Per-agent token budgets, spend-rate circuit breakers, and pre-call enforcement gates stop runaway agent loops in under 60 seconds, not 11 days.

Prompt injection defenses to build before your agent ships
Prompt injection is structural, not a patch. These pre-production controls bound your exposure before the first user reaches your agent.

AI agent baseline measurement: what to record before deployment
Capturing process metrics before an agent runs is the only way to prove value afterward. Four data categories make that baseline rigorous.

Agent sprawl: building infrastructure before agent count scales
Agents reach production faster than coordination layers do. Registry, orchestration, and governance must precede growth, not follow an incident.

Async AI agents: retry logic, idempotency, and state persistence
Forcing async workflows into real-time execution causes duplicate writes, lost checkpoints, and cascading failures. These patterns fix that.

AI agent escalation flows: protecting revenue at the boundary
Escalation architecture determines whether your AI agents protect or erode revenue. Set decision boundaries by consequence class, not agent capability.

How to evaluate AI agent output quality at scale
Sampling, automated triage signals, and hard execution limits let teams catch agent failures without reviewing every output.

Why multi-step agents fail in production and how to fix it
A 95% per-step accuracy rate collapses to 60% end-to-end success across ten steps. Checkpoint architecture stops that compounding.

AI agent kill switches: four control points for your first build
Agents that write records or call APIs need halt mechanisms, approval gates, and audit logs designed in from the start, not added after launch.

Why AI agent pilots stall at the governance stage
Only 12% of enterprise AI agent pilots reach production. The ones that ship treat access controls and audit logs as preconditions, not afterthoughts.

AI agent silent failures: why your dashboard misses most of them
67% of deployed agents degrade without triggering alerts. Instrumenting reasoning traces and step-level context catches failures standard APM tools cannot.

How to Build Your Own AI Agent Like ChatGPT (Step by Step)
You don't need to train a model to build your own AI like ChatGPT. The five steps: pick a model API, add a system prompt, tools, memory, and guardrails.

Agent production metrics: the three categories that determine ROI
Pilot accuracy scores do not survive contact with production. Three metric categories covering performance, behavior, and outcomes close the gap.

Why AI agent pilots fail before production: five root causes
88% of AI agent pilots never reach production. Five structural gaps explain why, and each one can be fixed before a pilot starts.

AI agents as infrastructure: what your IT team needs to change
Agents consume compute, credentials, and policy resources differently from applications. Your IT playbook needs updating before first deployment.

AI agent sandbox escape: causes and containment controls
Sandbox isolation fails for AI agents through tool chaining, credential inheritance, and goal persistence. Two 2026 incidents show what to fix.

Agent evaluation beyond task completion: three failure modes
Completion rate misses inference waste, goal drift, and policy violations. All three appear in agents that are hitting their task targets.

Grounding layers for production AI agents: build guide
A grounding layer is a three-stage pipeline. All three stages must hold under real traffic or hallucinations return.

AI agent execution validation: three verification layers
Agents infer success from response shape, not system state. Pre-execution guardrails, response validation, and outcome checks close that gap.

AI agent file deletion and credential leaks: three controls
Agents delete files and leak secrets when they hold excess access and no gate precedes action. Vaults, sandboxes, and approval gates fix both.

Seven agent failure patterns that block production deployment
Seven failure patterns account for most agent deployments that stall before production. Each one is observable and correctable before launch.

Production agent failure modes that surface at runtime
Context saturation, state corruption, and reasoning drift compound across long runs. Five structural failure modes explain why, with instrumentation to catch.

Agent state recovery: why rollbacks alone are not enough
Code rollbacks leave corrupted agent state behind. Versioned memory, context snapshots, and deterministic state machines are required for clean recovery.

OpenAI Hugging Face incident: what it means for production agents
Two OpenAI models escaped a benchmark sandbox and reached Hugging Face infrastructure. The containment lessons apply to any agent you ship today.

AI agent cost governance: controls to set before deployment
Set cost-per-outcome metrics, escalation triggers, and quota limits before your agent goes live to keep spending tied to business value.

AI agent KPIs: proving performance before you scale
Four measurements every engineering and ops leader needs before expanding an AI agent: baseline, test set, quality metrics, one business number.

Agent observability: instrument before production, not after
Pilots hide edge cases. Distributed tracing, token anomaly alerts, and automated evals each catch a different failure mode before go-live.

AI agent governance: why deployment outpaces accountability
72% of enterprises run AI agents with unmanaged risk. Shared credentials, bypassed approvals and data leaks follow directly from that gap.

AI agent governance: why accountability frameworks lag deployment
Most enterprises run AI agents in production before accountability frameworks exist. Three documented failure modes show what that gap costs.

Human-in-the-loop oversight: earning agent autonomy over time
Start with explicit human checkpoints, then use measured performance to remove each gate as your agent builds a track record that justifies the trust.

Agent permissions: least privilege before production
Scope credentials, build tool allowlists, and calculate blast radius so a bad agent decision stays recoverable, not permanent.

API and data readiness checklist for your first AI agent
Integration failures stop most first agents, not the model. Audit your API surface and data quality before committing engineering time.

MCP tool access: how agent permissions reshape your risk model
When agents call tools directly via MCP, traditional perimeter controls stop working. This is what changes and what governance covers the gap.

How to verify agent-generated code before it ships
Agent-generated code carries a 15 to 18% higher vulnerability rate than human code. Verification must target agent-specific failure modes.

Why enterprise AI agent pilots fail before production
Ford and Meta reversed high-profile agent programmes for different reasons. Both failures point to the same scoping errors you can avoid.

How to build an agent evaluation suite from scratch
Build a working agent eval suite using golden sets, LLM-as-judge scoring, and production sampling, starting in week one with 30 to 50 cases.

AI agents vs assistants: what changes when AI starts acting
AI agents write to your systems, not just to your screen. That shift demands new permissions, integration work, and accountability before first deployment.

Build vs buy for a production AI agent: a decision framework
Vertical products win on commodity workflows; building wins on proprietary ones. Most mid-size companies land on a deliberate hybrid of both.

How Do You Know Your AI Agents Are Doing Their Job?
Move from gut feeling to evidence: what to measure for agents in production, how failures hide, and how to produce proof that satisfies a board.

How to choose your first AI agent use case
Three tests determine whether a candidate agent use case is ready to build. Bounded scope, measurable outcome, and tolerable failure mode explained.

AI agent demo to production: five gaps to close
Edge cases, integration failures, permissions, cost at volume, and latency separate a working demo from a production agent.
Stay ahead of the curve
Get frameworks, playbooks, and insights on agentic governance delivered to your inbox.
No spam. Unsubscribe anytime. A resource by Prefactor.