AI engineer

Agent readiness for AI engineers

The AI engineer ships the agent and carries the pager when it misbehaves. Readiness from this seat is concrete: instrumentation that makes failures debuggable, evals that catch regressions before users do, and tool boundaries that hold when a prompt injection lands. Governance done well is scaffolding you build once; done badly it is a sprint tax on every release.

Key concerns

Instrumenting on day one, not after the incident

Traces, cost tracking, eval hooks, and audit events at the tool-call boundary are cheap to wire into a new agent and expensive to retrofit under incident pressure. The first version of the agent should already answer: what did it do, what did it cost, and where is the trace. Anything less is borrowing debugging time from your future self at a bad rate.

Tool boundaries that assume steering

Any agent reading untrusted content — tickets, emails, web pages — will eventually receive input crafted to redirect it. Prompt hygiene does not solve this; the tool boundary does. Scope each tool to the minimum action, validate parameters server-side, and keep irreversible operations behind gates the model cannot talk its way through.

Evals as the regression suite

The model provider will update the model whether or not you are ready, and the agent's behaviour will move. An evaluation set that runs on every model, prompt, or tool change is the only mechanism that catches the drift before production does. Every incident and every surprising trace is a new case; an eval set that is not growing is decaying.

Readiness checklist

  • New agents start from a scaffold with tracing, cost tracking, and audit events already wired
  • Every tool call is logged with inputs and outputs, with redaction at write time
  • Irreversible actions sit behind server-side gates, not prompt instructions
  • An eval set exists, runs on every change, and gates promotion
  • The agent has its own identity and short-lived credentials — nothing shared, nothing in config
  • You can pull the full trace for any production task from the last 30 days in minutes

Frequently asked questions

How is this different from the head of AI's readiness?

Altitude. The head of AI owns the portfolio: which agents exist, what evidence wins governance reviews, how quality is argued. The AI engineer owns the artifact: the instrumentation, the eval cases, the tool scopes. The head of AI's evidence is made of things the AI engineer built.

Find out where your organisation stands on agent readiness.

Take the assessment →