Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents

2026-10-02 · Hugging Face

Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents

Background and Challenges

Tool-using LLM agents have evolved beyond reading from a single retrieved passage. Through the Model Context Protocol (MCP), an agent can call a search tool, inspect structured records, query databases, and pull metadata, weaving all of it into a single answer. This makes factuality checking more subtle. Most existing verification systems, from RAGAS faithfulness to fine-grained checkers like MiniCheck, AlignScore, and SummaC, ask whether a claim is supported by available evidence only after that evidence has been pooled together. They fail to indicate which MCP tool output supports each claim or whether that source matches the one named in the answer.

The Cross-Source Conflation Problem

Multiverse Computing's latest paper, *ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents*, targets this gap. The core failure mode they identify is cross-source conflation: a claim that is true somewhere in the evidence but attributed to the wrong source.

For instance, a customer support agent might state, "According to the account record, this plan includes a 30-day refund window." The refund window may be real, but stated in a policy document, not the account record. Pooling them makes the claim look supported, but the attribution is wrong. In data-sensitive settings, wrong attribution can be as damaging as a wrong fact. The same applies to clinical agents where patient-specific details might be misleadingly presented as findings from medical literature.

How ProvenanceGuard Works

ProvenanceGuard is a post-generation verification layer that sits on top of a black-box MCP agent. It runs after an answer is produced and never collapses evidence into an anonymous context. Instead, it preserves the source identity throughout the pipeline. It reads captured MCP traces, including tool outputs and source IDs, without retraining the agent. The process involves five sequential steps:

1. Decomposition: Breaks the answer into specific claims.

2. Routing: Uses a MiniLM model to find the source most relevant to each claim.

3. Support Scoring: A DeBERTa NLI verifier model checks whether that source actually supports the claim.

4. Attribution Checking: Compares the actual supporting source with the one the answer explicitly names or implies.

5. Decision: Emits a per-claim source verdict and a global, answer-level allow or block decision.

The verifier also checks literal values closely: a number, date, or identifier absent from the source cannot pass just because the sentence sounds plausible. If an answer is blocked, a RARR-style repair step can attempt a source-grounded revision or safe fallback, which is then re-verified.

Experimental Results

The team tested ProvenanceGuard on a medical agent that used patient records, research articles, and other tools, yielding 281 real traces. Medicine is an ideal test case because facts from a patient's record and general research cannot be treated as the same source. Human experts checked 361 claims from 40 answers.

Key findings include:

  • Experts identified 139 claims that should not pass, and ProvenanceGuard caught 138 of them, letting only one through.
  • It also held 67 claims that experts considered supported, sending them for review or repair. This reflects the system's cautious setting, favoring a second look at supported claims over letting unsupported ones through.
  • For claims with an identifiable source, it picked the correct source about 86% of the time.

This conservative decision policy is well-suited for data-sensitive reviews where getting the source right matters more than producing the fastest possible answer.

Source