Multiverse Computing has released ProvenanceGuard, a post-generation verification layer designed to ensure that claims made by LLM agents using the Model Context Protocol (MCP) are attributed to the correct source. The tool addresses a specific failure mode known as cross-source conflation, where an answer is factually supported by the pooled evidence but incorrectly attributed to a specific tool or document.
What Happened
Traditional factuality checkers like RAGAS, MiniCheck, AlignScore, and SummaC typically evaluate whether a claim is supported by any available evidence. However, in MCP-based agents that pull from multiple tools—such as search engines, databases, and structured records—this approach can mask attribution errors. ProvenanceGuard operates by preserving source identity throughout the verification pipeline rather than collapsing evidence into a single context. It decomposes the agent's answer into specific claims, routes each claim to the most relevant source, verifies support, and checks if the source matches the one named or implied by the answer.
The system uses a sequence of steps including claim decomposition, source routing, support scoring, attribution checking, and optional repair. In the paper's experiments, the team utilized local models, including MiniLM for routing, a DeBERTa NLI verifier for support checks, and a local language model for claim decomposition. The verifier strictly checks literal values, ensuring that numbers, dates, or identifiers absent from the source cannot pass. If an answer is blocked, the system can attempt a RARR-style repair or provide a safe fallback.
Why It Matters
For developers and enterprises deploying agents in data-sensitive fields like healthcare and finance, accurate source attribution is critical. The paper reports that in a medical agent test involving 281 real traces, human experts identified 139 claims that should not pass. ProvenanceGuard caught 138 of these, missing only one, while holding 67 supported claims for further review. This reflects a conservative policy prioritizing safety over speed.
In comparative benchmarks against source-blind checkers, ProvenanceGuard achieved a reject/block F1 score of 0.802, outperforming MiniCheck (0.783), RAGAS Faithfulness (0.758), AlignScore (0.662), and SummaC-ZS (0.436). Crucially, ProvenanceGuard is the only system among these that emits per-claim source verdicts. In a controlled test of wrong attribution, the system correctly caught all 50 instances where the named source was swapped. The company notes that while the reported results used local models, the architecture can be adapted to hosted models, though such setups require separate testing and calibration.
The Bottom Line
ProvenanceGuard offers a mechanism to audit MCP agents by verifying not just the truth of a statement, but the provenance of the evidence supporting it. The approach has already been adapted for NVIDIA NVFlow’s finance agent to verify answers against SEC excerpts. While the system adds overhead—roughly half a second per answer in local configurations—it provides a layer of accountability for agents operating in complex, multi-source environments.