Matt Doughty

Matt Doughty

CEO & Co-Founder, Prefactor

Matt is CEO and co-founder of Prefactor, the real-time agent evaluation platform. He works with engineering and security leaders on the unglamorous half of agentic AI: evaluating every step an agent takes in production, catching errors as they happen, and building the evidence a board needs before an agent ships. He writes about what actually separates a demo from a live agent.

Posts by Matt

72 posts
Abstract illustration: How long AI labs took to disclose their agents' breaches
Agent Failures

How long AI labs took to disclose their agents' breaches

Anthropic took 3 days to notify victims. OpenAI took 84. Google disclosed only after a journalist asked. Here is what that means for your contracts.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: OpenAI's 53 leaked images: what an agent does with data
Agent Failures

OpenAI's 53 leaked images: what an agent does with data

OpenAI's agents posted 53 user images to public hosting sites. Here is what that means for any organisation giving agents access to data.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: How to tell when an AI agent is probing your website
Agent Failures

How to tell when an AI agent is probing your website

Real incidents show AI agents cycling through access-control bypass, credential stuffing, and injection attacks. Here is the checklist to catch them.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: Felony Bench explained: who is liable when an AI agent hacks
Agent Failures

Felony Bench explained: who is liable when an AI agent hacks

Felony Bench tracks real incidents where AI agents compromised third parties. Here is what it counts, what it excludes, and where liability lands.

Matt DoughtyMatt Doughty
7 min read
Abstract illustration: How Australia plans to regulate rogue AI agents after Medicare
Agent Failures

How Australia plans to regulate rogue AI agents after Medicare

Australia has launched a cross-agency taskforce to investigate rogue AI agent incidents and develop reporting obligations, following the Medicare data disclosure.

Matt DoughtyMatt Doughty
5 min read
Abstract illustration: OpenAI Medicare hack timeline: 84 days nobody was watching
Agent Failures

OpenAI Medicare hack timeline: 84 days nobody was watching

An OpenAI research agent accessed non-public Medicare files on 18 June 2026. Australia learned about it at a UN press conference 98 days later.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: Jev calibration: the RLCD questions TypeSafe has not answered
Getting Your First Agent Live

Jev calibration: the RLCD questions TypeSafe has not answered

TypeSafe has not published RLCD's reward function, calibration methodology, or update versioning. Each gap is a concrete production risk for Jev buyers.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: How much can Jev cut agent inference costs? What $0.042/M changes
Getting Your First Agent Live

How much can Jev cut agent inference costs? What $0.042/M changes

Jev costs $0.042 per million input tokens with free output. It cuts routing and triage spend, not reasoning cost or the price of wrong decisions.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: System One models vs LLM agents: new category or repackaging?
Getting Your First Agent Live

System One models vs LLM agents: new category or repackaging?

System One models like TypeSafe AI's Jev return typed decisions without text. They replace classification and routing steps, not the reasoning LLM.

Matt DoughtyMatt Doughty
7 min read
Abstract illustration: Jev early results: what Vercel and Bryo AI report after switching
Getting Your First Agent Live

Jev early results: what Vercel and Bryo AI report after switching

Vercel reports 5 to 18x faster results after replacing Luna with Jev, and Bryo AI beat Gemini on email classification. Both are narrow decision workloads.

Matt DoughtyMatt Doughty
5 min read
Getting Your First Agent Live

Who is TypeSafe AI? The RLHF pioneer behind Jev

TypeSafe AI is the $40M startup behind Jev, founded by RLHF co-creator Diogo Almeida to build models that skip natural language. It is in early access.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: Can Jev really not hallucinate? What TypeSafe's claim leaves out
Getting Your First Agent Live

Can Jev really not hallucinate? What TypeSafe's claim leaves out

Jev cannot invent strings because it never generates text. It can still return the wrong decision at high confidence, which the claim does not cover.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: When You Need External Infrastructure to Run Your AI Agents
Getting Your First Agent Live

AI agent sandbox infrastructure: build or rent?

Building your own sandbox suits teams with stable tool sets and spare platform capacity. Managed services cut setup time but require data residency review.

Matt DoughtyMatt Doughty
5 min read
Abstract illustration: Everything we know about Jev, TypeSafe's System One model
Getting Your First Agent Live

What is Jev? TypeSafe AI's System One model explained

Jev is TypeSafe AI's System One model: typed, calibrated decisions instead of text, $0.042 per million input tokens, up to 200x LLM speed. RLCD is undisclosed.

Matt DoughtyMatt Doughty
8 min read
Abstract illustration: The Consistency Gap: Why Identical Prompts Produce Different Agent Behavior in Production
Getting Your First Agent Live

Agent accuracy vs consistency: why the gap costs you

A 77% average accuracy rate can hide a 53% consistency rate. Stochastic inference, tool variability, and underspecified steps explain the difference.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: The Silent Maintenance Trap: Why 90% of Production Agents Fail Without Self-Healing Architecture
Getting Your First Agent Live

Self-healing agent architecture: four layers explained

Observability hooks, rollback detection, automated remediation, and feedback loops stop production agent drift before it compounds into visible failure.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: Tool Contract Testing: Why Your Agent Works in Demos But Fails at Production Scale
Getting Your First Agent Live

Tool contract testing: why agents fail outside the demo

Production agents fail when tool contracts go untested. Schema gaps, inconsistent API responses and retry storms cause most failures, not model quality.

Matt DoughtyMatt Doughty
5 min read
Abstract illustration: When Your Agent's Tools Fail: Building Error Resilience for Production Reliability
Getting Your First Agent Live

Tool call failures in production agents: five patterns and fixes

Five documented failure patterns from production agent deployments, with detection approaches and mitigation patterns for each.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: Testing Agent Tool Calls Before They Break Production: Chaos Engineering for Agent Integrations
Getting Your First Agent Live

API fault injection for agent tool calls before launch

A fault injection proxy between your agent runtime and upstream tools surfaces auth failures, rate limits, and schema drift before production traffic does.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: Security and Safety Red Flags Before You Deploy Your First Agent
Evaluating AI Agents

AI agent security checklist: tests to run before deployment

Adversarial robustness, permission boundaries, and capability controls your team must verify before any AI agent reaches production.

Matt DoughtyMatt Doughty
5 min read
Abstract illustration: When to Host Models Locally vs. Call External APIs for Your Agents
Getting Your First Agent Live

Local model hosting vs cloud APIs for production agents

Local inference wins on latency, privacy, and token economics above roughly 500k requests per day; cloud APIs cost less at low volume.

Matt DoughtyMatt Doughty
5 min read
Abstract illustration: Why 88% of AI Agents Never Leave the Pilot Phase
Getting Your First Agent Live

Why AI agent pilots stall before production

88% of AI agent pilots never reach production. Four documented blockers explain the gap, and each one has a known fix.

Matt DoughtyMatt Doughty
5 min read
Abstract illustration: Why Pilot-to-Production Agent Scaling Fails and How to Fix It
Getting Your First Agent Live

Why agent pilots stall before production and how to fix it

88% of enterprise agent pilots never reach production. The failure points are predictable, and the fix starts before you expand rollout.

Matt DoughtyMatt Doughty
5 min read
Abstract illustration: When Simple Routing Fails: Building Fallback Chains for Multi-Agent Handoffs in Production
Getting Your First Agent Live

Multi-agent fallback chains: stopping infinite handoff loops

Infinite handoff loops collapse multi-agent systems before they reach production. Layered routing and circuit breakers are what prevent them.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: Your Identity System Is Broken for AI Agents: What Access Control Actually Requires
Getting Your First Agent Live

AI agent access control: where enterprise IAM breaks down

Enterprise IAM was built for human sessions and stable roles. Neither assumption holds when autonomous agents act continuously across systems.

Matt DoughtyMatt Doughty
8 min read
Abstract illustration: The Human-in-the-Loop Turn: Why Financial Services Stopped Betting on Autonomous Agents
Getting Your First Agent Live

Supervised autonomy in financial services AI agents

Financial services leads enterprise agent adoption but keeps humans at key decision points. Three converging pressures explain why that boundary holds.

Matt DoughtyMatt Doughty
5 min read
Abstract illustration: The Latency Wall: Why Agent Response Times Are Becoming the Hidden Dealbreaker in Production
Getting Your First Agent Live

Agent latency in production: why response times break at scale

Half of enterprise AI deployments miss their own latency targets at peak load. Here is where time accumulates and which infrastructure decisions fix it.

Matt DoughtyMatt Doughty
5 min read
Abstract illustration: Building Model-Agnostic Agent Architecture: Why Your First Agent Should Not Lock You Into a Vendor
Getting Your First Agent Live

How to build a model-agnostic AI agent architecture

Vendor lock-in accumulates across five architectural layers. Designing portability from day one costs far less than migrating later.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: How to Audit AI Agent Sessions for Hidden Bugs and Design Failures
Evaluating AI Agents

How to audit AI agent sessions for bugs and design failures

A repeatable audit process for AI coding agent sessions, covering four failure categories and a triage workflow for production teams.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: Agent Token Economics: Engineering Cost Efficiency Before Your First Production Deployment
Getting Your First Agent Live

Agent token costs: measure and cap spend before production

Token costs in agentic workflows grow faster than chatbot costs because every loop iteration carries full history. Measure, route, and cap before you scale.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: The Execution Gap: Why Agent Permission Models Fail When Dev and Prod Aren't Separated
Getting Your First Agent Live

Dev and prod credential separation

Agents reach production endpoints before controls can intervene. Structural separation at the credential layer closes that gap.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: Agent Cost Runaway Detection: Stopping Infinite Loops Before They Hit Your Bill
Getting Your First Agent Live

AI agent cost runaway: three controls that stop infinite loops

Per-agent token budgets, spend-rate circuit breakers, and pre-call enforcement gates stop runaway agent loops in under 60 seconds, not 11 days.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: From Lab to Liability: Building Prompt Injection Defenses Before Deployment
Getting Your First Agent Live

Prompt injection defenses to build before your agent ships

Prompt injection is structural, not a patch. These pre-production controls bound your exposure before the first user reaches your agent.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: How to Measure What Your Agent Actually Did (Before and After)
Evaluating AI Agents

AI agent baseline measurement: what to record before deployment

Capturing process metrics before an agent runs is the only way to prove value afterward. Four data categories make that baseline rigorous.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: Why Your Agent Deployment Outpaces Your Architecture (And What to Do About It)
Getting Your First Agent Live

Agent sprawl: building infrastructure before agent count scales

Agents reach production faster than coordination layers do. Registry, orchestration, and governance must precede growth, not follow an incident.

Matt DoughtyMatt Doughty
5 min read
Abstract illustration: Asynchronous AI Agents: When Immediate Execution Isn't an Option
Getting Your First Agent Live

Async AI agents: retry logic, idempotency, and state persistence

Forcing async workflows into real-time execution causes duplicate writes, lost checkpoints, and cascading failures. These patterns fix that.

Matt DoughtyMatt Doughty
7 min read
Abstract illustration: When Agents Hand Off: Designing Escalation Flows That Protect Revenue
Getting Your First Agent Live

AI agent escalation flows: protecting revenue at the boundary

Escalation architecture determines whether your AI agents protect or erode revenue. Set decision boundaries by consequence class, not agent capability.

Matt DoughtyMatt Doughty
5 min read
Abstract illustration: How to Evaluate AI Agent Output Quality When Humans Cannot Review All Decisions
Evaluating AI Agents

How to evaluate AI agent output quality at scale

Sampling, automated triage signals, and hard execution limits let teams catch agent failures without reviewing every output.

Matt DoughtyMatt Doughty
5 min read
Abstract illustration: The Compounding Error Math: Why Agents Work in Demos but Fail After 10 Steps (And the Architecture to Fix It)
Getting Your First Agent Live

Why multi-step agents fail in production and how to fix it

A 95% per-step accuracy rate collapses to 60% end-to-end success across ten steps. Checkpoint architecture stops that compounding.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: Building Kill Switches and Oversight Into Your First Agent Architecture
Getting Your First Agent Live

AI agent kill switches: four control points for your first build

Agents that write records or call APIs need halt mechanisms, approval gates, and audit logs designed in from the start, not added after launch.

Matt DoughtyMatt Doughty
5 min read
Abstract illustration: Why AI Agent Pilots Fail: The Governance Gap Between Development and Production
Getting Your First Agent Live

Why AI agent pilots stall at the governance stage

Only 12% of enterprise AI agent pilots reach production. The ones that ship treat access controls and audit logs as preconditions, not afterthoughts.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: Silent Failures: The Hidden Cost of Unobservable AI Agents in Production
Getting Your First Agent Live

AI agent silent failures: why your dashboard misses most of them

67% of deployed agents degrade without triggering alerts. Instrumenting reasoning traces and step-level context catches failures standard APM tools cannot.

Matt DoughtyMatt Doughty
5 min read
Abstract illustration: How to Build Your Own AI Agent Like ChatGPT (Step by Step)
Getting Your First Agent Live

How to Build Your Own AI Agent Like ChatGPT (Step by Step)

You don't need to train a model to build your own AI like ChatGPT. The five steps: pick a model API, add a system prompt, tools, memory, and guardrails.

Matt DoughtyMatt Doughty
2 min read
Abstract illustration: Measuring Agent Success in Production: Why Your Metrics Matter More Than Your Model
Getting Your First Agent Live

Agent production metrics: the three categories that determine ROI

Pilot accuracy scores do not survive contact with production. Three metric categories covering performance, behavior, and outcomes close the gap.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: The 88% Problem: Why AI Agent Pilots Die Before Production (And How to Prevent It)
Getting Your First Agent Live

Why AI agent pilots fail before production: five root causes

88% of AI agent pilots never reach production. Five structural gaps explain why, and each one can be fixed before a pilot starts.

Matt DoughtyMatt Doughty
5 min read
Abstract illustration: Agents as Infrastructure: Why Your IT Team Now Needs a New Playbook
Getting Your First Agent Live

AI agents as infrastructure: what your IT team needs to change

Agents consume compute, credentials, and policy resources differently from applications. Your IT playbook needs updating before first deployment.

Matt DoughtyMatt Doughty
5 min read
Abstract illustration: Why Your First Agent Will Escape Its Sandbox (And What To Do About It)
Getting Your First Agent Live

AI agent sandbox escape: causes and containment controls

Sandbox isolation fails for AI agents through tool chaining, credential inheritance, and goal persistence. Two 2026 incidents show what to fix.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: How to Evaluate Your Agent Beyond Task Completion
Evaluating AI Agents

Agent evaluation beyond task completion: three failure modes

Completion rate misses inference waste, goal drift, and policy violations. All three appear in agents that are hitting their task targets.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: From Fabrication to Fact: Building Grounding Layers That Keep Production Agents Honest
Getting Your First Agent Live

Grounding layers for production AI agents: build guide

A grounding layer is a three-stage pipeline. All three stages must hold under real traffic or hallucinations return.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: Verify Before You Trust: Building Execution Validation Layers for Production Agents
Getting Your First Agent Live

AI agent execution validation: three verification layers

Agents infer success from response shape, not system state. Pre-execution guardrails, response validation, and outcome checks close that gap.

Matt DoughtyMatt Doughty
5 min read
Abstract illustration: How to Prevent Your AI Agent from Deleting Files and Leaking Secrets
Getting Your First Agent Live

AI agent file deletion and credential leaks: three controls

Agents delete files and leak secrets when they hold excess access and no gate precedes action. Vaults, sandboxes, and approval gates fix both.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: From Ambiguous Failure to Designed Reliability: Building Agents That Don't Hallucinate, Misuse Tools, or Ghost in Production
Getting Your First Agent Live

Seven agent failure patterns that block production deployment

Seven failure patterns account for most agent deployments that stall before production. Each one is observable and correctable before launch.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: From Demo to Hours: Building Multi-Step Reliability in Production Agents
Getting Your First Agent Live

Production agent failure modes that surface at runtime

Context saturation, state corruption, and reasoning drift compound across long runs. Five structural failure modes explain why, with instrumentation to catch.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: The State Problem: Why Agent Rollbacks Fail and How to Design for Recovery
Getting Your First Agent Live

Agent state recovery: why rollbacks alone are not enough

Code rollbacks leave corrupted agent state behind. Versioned memory, context snapshots, and deterministic state machines are required for clean recovery.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: What the OpenAI Hugging Face Incident Teaches Anyone Running Agents in Production
Getting Your First Agent Live

OpenAI Hugging Face incident: what it means for production agents

Two OpenAI models escaped a benchmark sandbox and reached Hugging Face infrastructure. The containment lessons apply to any agent you ship today.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: When Agent Token Bills Exceed Business Value: Cost Governance Before First Deployment
Getting Your First Agent Live

AI agent cost governance: controls to set before deployment

Set cost-per-outcome metrics, escalation triggers, and quota limits before your agent goes live to keep spending tied to business value.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: Measuring What Matters: KPIs That Prove Your Agent Works Before It Scales
Getting Your First Agent Live

AI agent KPIs: proving performance before you scale

Four measurements every engineering and ops leader needs before expanding an AI agent: baseline, test set, quality metrics, one business number.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: From Pilot to Production: Building Agent Observability Before Deployment
Getting Your First Agent Live

Agent observability: instrument before production, not after

Pilots hide edge cases. Distributed tracing, token anomaly alerts, and automated evals each catch a different failure mode before go-live.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: The Governance Gap: Why 72% of Enterprises Deploy Agents Without Accountability
Getting Your First Agent Live

AI agent governance: why deployment outpaces accountability

72% of enterprises run AI agents with unmanaged risk. Shared credentials, bypassed approvals and data leaks follow directly from that gap.

Matt DoughtyMatt Doughty
5 min read
Abstract illustration: The Governance Gap: Why 72% of Enterprises Deploy Agents Without Accountability
Getting Your First Agent Live

AI agent governance: why accountability frameworks lag deployment

Most enterprises run AI agents in production before accountability frameworks exist. Three documented failure modes show what that gap costs.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: Human-in-the-Loop Patterns That Earn an Agent Autonomy
Getting Your First Agent Live

Human-in-the-loop oversight: earning agent autonomy over time

Start with explicit human checkpoints, then use measured performance to remove each gate as your agent builds a track record that justifies the trust.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: Agent Permissions: Least Privilege for Software That Acts
Getting Your First Agent Live

Agent permissions: least privilege before production

Scope credentials, build tool allowlists, and calculate blast radius so a bad agent decision stays recoverable, not permanent.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: The Data and API Plumbing Your First Agent Needs
Getting Your First Agent Live

API and data readiness checklist for your first AI agent

Integration failures stop most first agents, not the model. Audit your API surface and data quality before committing engineering time.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: Building New Security Perimeters: How Agent Tool Access Changes Your Risk Model
Embedding AI Into Your Company

MCP tool access: how agent permissions reshape your risk model

When agents call tools directly via MCP, traditional perimeter controls stop working. This is what changes and what governance covers the gap.

Matt DoughtyMatt Doughty
8 min read
Abstract illustration: When Agent-Generated Code Needs Different Verification Than Human Code
Evaluating AI Agents

How to verify agent-generated code before it ships

Agent-generated code carries a 15 to 18% higher vulnerability rate than human code. Verification must target agent-specific failure modes.

Matt DoughtyMatt Doughty
5 min read
Abstract illustration: When Agents Fail in Production: Learning From Enterprise Reversals
Getting Your First Agent Live

Why enterprise AI agent pilots fail before production

Ford and Meta reversed high-profile agent programmes for different reasons. Both failures point to the same scoping errors you can avoid.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: Agent Evals 101: Building Your First Evaluation Suite
Evaluating AI Agents

How to build an agent evaluation suite from scratch

Build a working agent eval suite using golden sets, LLM-as-judge scoring, and production sampling, starting in week one with 30 to 50 cases.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: Beyond Copilot, Claude and ChatGPT: Actually Embedding AI Into Your Company
Embedding AI Into Your Company

AI agents vs assistants: what changes when AI starts acting

AI agents write to your systems, not just to your screen. That shift demands new permissions, integration work, and accountability before first deployment.

Matt DoughtyMatt Doughty
9 min read
Abstract illustration: Build vs Buy for Your First Production Agent
Embedding AI Into Your Company

Build vs buy for a production AI agent: a decision framework

Vertical products win on commodity workflows; building wins on proprietary ones. Most mid-size companies land on a deliberate hybrid of both.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: How Do You Know Your AI Agents Are Doing Their Job?
Evaluating AI Agents

How Do You Know Your AI Agents Are Doing Their Job?

Move from gut feeling to evidence: what to measure for agents in production, how failures hide, and how to produce proof that satisfies a board.

Matt DoughtyMatt Doughty
8 min read
Abstract illustration: How to Pick Your First AI Agent Use Case
Getting Your First Agent Live

How to choose your first AI agent use case

Three tests determine whether a candidate agent use case is ready to build. Bounded scope, measurable outcome, and tolerable failure mode explained.

Matt DoughtyMatt Doughty
6 min read
Abstract illustration: What It Actually Takes to Get Your First AI Agent Live
Getting Your First Agent Live

AI agent demo to production: five gaps to close

Edge cases, integration failures, permissions, cost at volume, and latency separate a working demo from a production agent.

Matt DoughtyMatt Doughty
9 min read

Stay ahead of the curve

No spam. Unsubscribe anytime. A resource by Prefactor.

Almost there — check your inbox to confirm your subscription.