AI agents
How to Evaluate an AI Agent for Operations: 12 Questions
Most AI agent evaluations test the demo rather than the deployment. Twelve questions that separate systems you can run against real vendors from systems that only work in a controlled scenario.
Evaluate an operations agent on restraint, boundaries and evidence rather than on capability. The most revealing questions are how often it declines to surface something, what it can do without a human, and what happens when it is wrong.
- Ask how often it declines to surface something. Restraint is the feature.
- A boundary that needs qualifiers to explain is not a boundary.
- Evaluate against your own history, not the vendor's scenario.
Access and scope
1. What does it read, and is access read-only? Read-only is a meaningful control an assessor can verify. Write access needs a specific justification.
2. Can exclusions be enforced before ingestion? There is a real difference between content that is filtered after arriving and content that never enters the platform. Only the second is defensible when a customer asks what you shared.
3. Is customer content used for training? Should be a flat no, in the contract rather than the marketing page.
Boundaries
4. State in one sentence what it can do without a human. If the answer requires conditions or a confidence threshold, the boundary is not testable. See where the approval boundary belongs.
5. Can it widen its own permissions? Must be no. An agent that can expand its own scope has no boundary whatever the configuration says.
6. Who approves, and is the decision recorded? A named person and a retained record, which is what an auditor will ask for.
Restraint
7. How often does it decline to surface something? The most revealing question on the list, and the one vendors are least prepared for. A system that surfaces everything has relocated alert fatigue, not solved it. Expect a number.
8. What does it do on the third occurrence of the same fault? If the answer is the same as the first occurrence, there is no recurrence memory, and recurrence is where most of the recoverable time sits.
"Show me something it decided not to raise, and explain why." A system with genuine judgement has these examples and they are interesting. A system without judgement has nothing to show, because it raised everything.
Evidence
9. Can it show its working, linked to original records? Every conclusion should trace back to the signals that produced it, in the systems they came from. Without that you cannot verify it and cannot defend it.
10. What does an escalation it drafts actually contain? Reproduction, the provider's own request identifier, UTC timestamps, scope, and references to prior occurrences. Ask to see a real one. Thin drafts mean someone still does the work.
Failure
11. What happens when it is wrong, and is that recoverable? It will be wrong. The question is whether being wrong produces a bad draft a human catches, or a sent message you cannot retract.
12. What happens when it breaks or you leave? Does detection stop silently? Is the history exportable? A record you cannot take with you is a record you do not own.
How to run the evaluation
Not on the vendor's scenario, which is selected to work.
- Pick an integration that has already failed more than once. You know the answer, which makes the test meaningful.
- Run it in observe-only mode first. Compare what it would have raised against what actually happened. Cheapest and most informative stage.
- Count both error types. Things it raised that did not matter, and things that mattered that it missed. Vendors report the second; the first is what determines whether anyone keeps reading.
- Check recurrence explicitly. Feed it a fault you know recurred, and see whether it says so.
That sequence is what a phased rollout describes in more detail, and most of the value turns out to be visible before any outbound action is enabled at all.
Frequently asked questions
What should you ask an AI agent vendor first?
What the agent reads, whether access is read-only, and what it can do without a human stated in one sentence. If the boundary needs conditions or a confidence threshold to explain, it is not a control anyone can approve.
How do you test whether an agent exercises judgement?
Ask it to show something it decided not to raise, and why. Systems with genuine judgement have interesting examples. Systems without have none, because they raised everything.
How should an evaluation be structured?
Against an integration that has already failed more than once, in observe-only mode, comparing what it would have raised against what actually happened. Count false positives as well as misses, since false positives determine whether anyone keeps reading.
Stop rediscovering the same integration failure
Traxivo correlates the signals your tools already produce into one incident timeline, recognises a recurrence as a recurrence, and drafts the follow-up with the evidence attached. Nothing is sent without a named approver.
See how Traxivo works Browse use cases

