How to Ground AI Agent Findings in Evidence: A Cloudflare Pattern
How to ground AI agent findings in evidence: validate every citation in code. Cloudflare's 3-state pattern plus a TypeScript checker you can copy.

TL;DR — Here is how to ground AI agent findings in evidence: collect the evidence with deterministic code first, make the model cite items from that package, and verify every citation in code before anything is reported. Cloudflare’s Managed Defense harness works this way, and its key idea is a three-state record (not checked, checked and empty, checked and absent) so a timeout never reads as “all clear”. Below is the pattern, the failure it prevents, and a TypeScript validator you can copy.
Grounding an agent in evidence is the practice of letting the agent assert only what it can point to: each claim carries a reference to a pre-collected evidence item, and non-model code confirms that the item exists, belongs to this investigation, and supports the claim.
Why Do Single-Agent Investigations Fail?
Cloudflare’s agentic security operations write-up starts with a failed prototype: one general-purpose agent was handed the whole alert investigation and produced unsupported claims. The authors name three causes, and each one generalizes beyond security.
| Failure | What happened | Fix in the final design |
|---|---|---|
| Context became authority | The alert text was treated as fact. “A detection is a hypothesis, not proof that an exploit succeeded.” | Detections enter as hypotheses; only admitted evidence can support a claim |
| Scope drift | The agent could query the wrong account, time range, or source | Scope is fixed in application code before any model sees results |
| Lost failure states | A timeout looked identical to “checked, found nothing” | Three explicit evidence states in every report |
The quote that matters most for builders: “You can’t rely on a language model prompt to be a boundary.” If your scope lives in a system prompt, it is a suggestion. I made the same argument about tool access in scoping an AI agent’s Cloudflare Workers access: the boundary has to be enforced where the model cannot talk its way past it.
What Does It Mean to Ground an AI Agent in Evidence?
Cloudflare’s pipeline has six stages, and the order is the point. Evidence collection happens before any inference, and validation happens before any report.
- Deterministic recon. Fixed workflows make versioned API calls and store each item with its source, version, and timestamp.
- Noise filtering. A lightweight triage model (Clef, on Workers AI) scores the alert against recon data and parks likely false positives.
- Specialist investigation. A coordinator runs four specialists in parallel: traffic, customer context, global telemetry, and threat intelligence.
- Synthesis. One agent merges typed findings into an advisory. It cannot fetch new evidence or invent classifications outside an approved vocabulary.
- Decision scoring. Clef checks whether the evidence is sufficient and whether anything contradicts it.
- Advisory report. An LLM writes the analyst-facing summary, and a human still makes the final call.
The design choice underneath: the recon snapshot is replayable, so, in the authors’ words, “differences between specialist AI agents’ findings come from interpretation rather than retrieval.” That is a debugging superpower. When two runs disagree, you know the model disagreed, not the data. It pairs naturally with the parallel tool-call harness pattern, where independent calls run concurrently and results merge in deterministic code.
How to Ground AI Agent Findings in Evidence: A 5-Step Build
Cloudflare did not publish code, so what follows is my own minimal implementation of the same idea. It is illustrative, not their source.
- Freeze scope in code. Resolve account, time window, and sources before the first model call, and pass only those handles to tools.
- Build a versioned evidence package. Every item gets a stable ID, a source, a timestamp, and a collection status.
- Require citations in a typed schema. A finding is
{ claim, evidenceIds[] }, never free prose. - Validate citations in code. Reject any ID that is missing, from another investigation, or that does not support the claim.
- Report limits explicitly. Failed findings become stated limitations, and thin evidence yields no classification.
type Status = "not_checked" | "checked_empty" | "checked_absent";
interface Evidence {
id: string;
investigationId: string;
source: string;
collectedAt: string; // ISO timestamp
status: Status;
facts: Record<string, string | number | boolean>;
}
interface Finding {
claim: string;
evidenceIds: string[];
assertsFact: { key: string; equals: string | number | boolean };
}
export function validate(findings: Finding[], pkg: Map<string, Evidence>, invId: string) {
const accepted: Finding[] = [];
const limitations: string[] = [];
for (const f of findings) {
const items = f.evidenceIds.map((id) => pkg.get(id));
const ok =
items.length > 0 &&
items.every((e) => e && e.investigationId === invId && e.status !== "not_checked") &&
items.some((e) => e!.facts[f.assertsFact.key] === f.assertsFact.equals);
if (ok) accepted.push(f);
else limitations.push(`Unsupported or unverifiable: ${f.claim}`);
}
return { accepted, limitations };
}Two details carry the weight. The not_checked guard means a source that timed out can never back a claim. And assertsFact forces the model to say exactly which fact it relies on, so “supports the claim” is a mechanical comparison instead of a judgment call. If you orchestrate this on a durable runner, Cloudflare Workflows gives you the checkpointing that Cloudflare uses so a failed stage reuses validated results instead of restarting.
How Should an Agent Report Missing Evidence?
Cloudflare’s advisory distinguishes three states, and this is the part I would copy first.
| State | Meaning | Example wording |
|---|---|---|
| Not checked | The source was never successfully queried (timeout, error, out of scope) | “Global telemetry was unavailable; widespread activity cannot be assessed” |
| Checked, no matching result | The query ran and returned nothing | “No prior detections for this path in the window” |
| Checked, evidence of absence | The data positively shows the thing did not happen | “Enforcement logs show the request was blocked” |
When evidence is insufficient, the harness makes no classification or disposition at all. That restraint is rarer than it sounds: most agent demos always produce an answer. Silence with a stated reason is more useful to an on-call analyst than a confident guess, and it is the same discipline as treating benchmark claims skeptically, which I covered in verifying AI agent benchmark claims.
When Is This Worth Building?
Use the full pattern when a wrong answer is expensive and a human acts on the output: security triage, incident review, compliance checks, financial reconciliation. Skip the parallel specialists for low-stakes summarization; keep the citation validator regardless, because it is about 30 lines. One honest caveat: Cloudflare reports no accuracy or false-positive numbers, so the benefit here is a design argument, not a measured result. Build a replayable snapshot first and measure your own error rate against it. For the wider security model around agents, see the agentic AI enterprise security model, and for the harness mindset see agent harness design and the LLM engineering hub.
FAQ
What does it mean to ground an AI agent in evidence? Every claim points at a pre-collected evidence item, and application code verifies the item exists and supports the claim. The model interprets evidence but never invents it.
Why can’t a system prompt enforce an agent’s scope? A prompt is an instruction, not a boundary. Fix account, time range, and sources in code before the model sees any results.
What are the three evidence states? Not checked, checked with no matching result, and checked with evidence supporting absence. A timeout is always the first, never the second.
Do I need multiple agents? No. The grounding comes from deterministic evidence collection plus citation validation. Multiple specialists help when data sources differ.
Does Cloudflare publish accuracy numbers? No. The post reports architecture and rationale only, with no accuracy, latency, or alert-volume figures.
Sources
- Building an evidence-grounded agentic security operations harness on Cloudflare: Cloudflare Blog. Primary source for the pipeline, the three failure modes, the quotes, and the three evidence states. The code in this post is my own illustration, not Cloudflare’s.
- OWASP Top 10 for LLM Applications: background on excessive agency and overreliance, the risks that citation validation and code-enforced scope address.
- Cloudflare Workflows documentation: durable multi-step execution used for stage checkpointing.
Frequently asked questions
Google Search · Preferred sources
Prefer this site on Google
If you already read this writing, add umesh-malik.com as a Preferred Source. Google can then highlight it with a preferred badge in Top Stories, AI Overviews, and AI Mode — for you, not as a site-wide ranking boost.
Related Articles

AI Engineering
Agent context compaction: keep what the 150K cutoff drops
Agent context compaction drops every block before the summary at 150K tokens. What survives, what instructions silently replaces, and the usage field that lies.

AI Engineering
Agent Harness Design: Why an ARC-AGI-3 Score Tripled
Agent harness design decided a benchmark: OpenAI's ARC-AGI-3 score went 13.3% → 38.3% with zero model changes. What that means for your agent loop.

AI Engineering
Cloudflare Wallets and x402: How AI Agents Pay for APIs
Cloudflare Wallets and x402 explained: how AI agents get a spending identity, how HTTP 402 payments work, and what breaks when your agent holds a budget.
Keep reading
Get new posts on AI, Claude Code & LLMs
New deep-dives on AI engineering, Claude Code, and developer tooling — follow along however you prefer.
About the Author
Software engineer writing about AI, Claude Code, LLMs, OpenAI, Anthropic, and developer tooling. 5+ years building production systems at Expedia Group, Tekion, and BYJU'S.