Evidence Assurance is a framework developed by John Koblinsky, founder of Marsh Island Group, in 2026. It describes the degree of evidentiary assurance a person can reasonably expect from an AI-generated answer: Directional, Grounded, or Auditable.
| Level | User should expect | Minimum evidence behavior |
|---|---|---|
| Directional | Useful synthesis or guidance, with an incomplete evidence chain | The system may use prior synthesized knowledge or incomplete document access |
| Grounded | An answer based on evidence retrieved at query time | Relevant sources are retrieved, but coverage and interpretation may still be incomplete |
| Auditable | An answer that can be inspected and defended | Important claims trace to a specific study, source location, and relevant audience; coverage boundaries are explicit |
Which Level Do You Need?
The right level depends on what the answer will be used for and what happens if it is wrong. Evidence Assurance is a proportional standard: require more traceability as the consequences and scrutiny increase.
| Typical situation | Reasonable level | What to require |
|---|---|---|
| Exploration, brainstorming, or a reversible internal task | Directional | Clear limits on what the answer can support |
| Synthesis expected to come from a defined research library | Grounded | Retrieved sources and enough context to inspect what informed the answer |
| A client, executive, legal, or high-consequence decision | Auditable | Claim-level source locations, audience fit, and explicit coverage boundaries |
A Quick Evidence Assurance Screen
Before relying on an AI-generated research answer, ask three questions:
-
01 · Access
Can you tell what evidence the system could actually see?
An answer cannot be stronger than the evidence made available to it.
-
02 · Traceability
Can each important claim be traced to a source location and relevant audience?
A citation is not enough if it does not support the specific claim being made.
-
03 · Coverage
Does the answer disclose what was not searched, retrieved, or represented?
Auditable evidence makes its boundaries visible instead of implying completeness.
If the evidence chain is unavailable, treat the answer as Directional. If sources were retrieved but coverage remains uncertain, treat it as Grounded. Auditable requires all three conditions and still does not guarantee that the underlying evidence is correct.
One Question at Three Levels
Consider a product marketer asking: “Will this AI-automation headline resonate with German CFOs?” A system at any level can produce a confident answer. What changes is the evidence beneath it and, therefore, what the answer can reasonably be used for.
Directional
Prior synthesis
“CFOs tend to reject vague AI claims. Lead with ROI and operational proof.”
What sits underneath: broad patterns summarized earlier. The system did not retrieve German CFO evidence for this question. Useful for an early draft; insufficient for a launch decision.
Grounded
Evidence retrieved now
“Relevant German CFO reactions show resistance to generic AI language and repeated requests for business proof.”
What sits underneath: relevant records retrieved from a defined research library at query time. The user can inspect what informed the answer, but coverage or source-verification limits may remain unclear.
Auditable
Traceable evidence chain
“The evidence in scope supports revising the headline. Each conclusion links to the relevant study passage, audience, and verification status.”
What sits underneath: claim-level source locations, audience metadata, source-verification state, and an explicit account of which studies and markets were—and were not—searched. Suitable for a consequential or externally scrutinized decision.
Illustrative scenario adapted from earlier work building AI-assisted research tools. The wording is generalized; no internal research findings or participant quotations are reproduced.
What It Is Not
Evidence Assurance is not a model benchmark. It does not rate a model's general capability or compare AI systems against each other.
It is not a maturity model. A more technically complex AI system is not automatically more trustworthy — the levels describe the outcome a user receives, not the architecture underneath it.
It is not a guarantee of correctness. An Auditable answer can be traced to its sources and still be wrong if those sources are wrong; the framework grades traceability, not truth.
Related Approaches
Evidence Assurance was developed independently, alongside John Koblinsky's work building AI research systems at SAP. A later review of market-research and AI literature found adjacent approaches addressing related problems from different angles:
Merciv's confidence tiers and auditability guidance address source citation and auditability within enterprise AI research tools directly. Fuel Cycle's Grounded AI focuses on reducing hallucination in AI-driven insights platforms. RAGAS is a technical evaluation framework for scoring retrieval-augmented generation quality.
Evidence Assurance differs by separating the assurance a user should expect from an answer from the technical architecture used to produce it — a user-facing outcome scale rather than a system benchmark or an implementation guide.
Citation
When referring to this framework, cite:
Koblinsky, John. "Evidence Assurance." Marsh Island Group, 2026.
For the full origin narrative, evidence, and application context, see How Much Can You Trust an AI Research Answer?
Follow the Work
Evidence Assurance is part of a larger body of work on AI-assisted research.
Follow the framework as it develops, including practical examples and companion work on building research tools that make their evidence behavior visible.
Related Reading
Why Human Oversight of AI Fails: The Judgment Gap
The companion problem: whether the humans reviewing AI output are still equipped to catch what Evidence Assurance would flag as unsupported.
How AI Sycophancy Undermines Executive Judgment
Confidence in an AI's tone is not evidence of its reliability — the distinction Evidence Assurance formalizes for research specifically.