Aivora
Hugging FaceAI AgentIntermediate

Stopping Cross-Source Conflation: ProvenanceGuard Brings Source-Aware Verification to MCP Agents

拒絕張冠李戴:ProvenanceGuard 為 MCP 代理帶來「來源感知」事實查核

2 min read
Stopping Cross-Source Conflation: ProvenanceGuard Brings Source-Aware Verification to MCP Agents
The 30-second version

When LLM agents retrieve data from multiple tools, they often suffer from "cross-source conflation"—stating a true fact but attributing it to the wrong source. Traditional verifiers ignore sources if the fact exists in the pooled context. ProvenanceGuard solves this by intercepting the agent's MCP trace post-generation without retraining. It preserves source IDs, decomposes answers into claims, routes them to specific sources, verifies support using NLI, and checks attribution. In medical-domain tests, it caught 138 of 139 invalid claims, achieving a 0.802 block F1 score and outperforming source-blind baselines.

Key points

01

Prevents Cross-Source Conflation

Traditional verifiers are "source-blind", letting claims pass if they exist in the context pool, even if attributed to the wrong source.

02

No Retraining Required

Works as a post-generation layer on black-box agents by reading MCP tool traces and source IDs directly, requiring no model retraining.

03

Five-Step Pipeline & Repair

Sequentially performs claim decomposition, routing, support scoring, attribution checking, and decision emitting, with optional RARR-style repair.

04

State-of-the-Art Performance

Outperforms baselines on medical traces with a 0.802 block F1 score, catching 138 out of 139 unsupported claims under a conservative policy.

How it works

ProvenanceGuard Verification and Repair Flow
Read MCP traceSingle claimsFind sourceNLI evaluationCompare source IDsAgent AnswerClaim DecompSource RoutingSupport CheckAttr CheckVerdict & Repair

Why it matters

In data-sensitive industries like medicine and finance, attributing a fact to the wrong document is as hazardous as hallucination. ProvenanceGuard makes agent actions auditable by mapping claims directly to tool IDs. Its adoption in NVIDIA's NVFlow finance agent proves its readiness for high-stakes enterprise applications where accountability is paramount.

Who it affects

  • AI Developer
  • AI Researcher
  • Enterprise Leader
  • Product Manager

How to use it

  1. 1Medical AI Agents: Separating patient-specific clinical history from general medical literature to avoid misattribution.
  2. 2Financial Compliance Agents: Verifying SEC filing citations and financial summaries against exact regulatory documents.

Limitations & caveats

  • Identifying the exact source among highly similar documents remains challenging, dropping to 50.3% accuracy in stress tests.
  • The evaluated setup relies heavily on local models (MiniLM, DeBERTa) and requires recalibration if transitioned to hosted cloud APIs.

Related

AutoSynthData: Generating Targeted Training Data from Enterprise Agent Failures
Hugging FaceAI Agent

AutoSynthData: Generating Targeted Training Data from Enterprise Agent Failures

AutoSynthData:以企業 Agent 的失敗為師,自動生成高規格微調訓練資料

AutoSynthData is a framework by ServiceNow that analyzes an enterprise agent's failures against a stronger teacher to automatically generate and validate high-quality synthetic training data, bridging crucial performance gaps.

2 min read
KaliBench: Evaluating and Boosting LLM Command Generation on Kali Linux
arXivAI Agent

KaliBench: Evaluating and Boosting LLM Command Generation on Kali Linux

KaliBench:首個 Kali Linux 資安工具指令生成基準測試,助 8B 模型直逼 685B 巨獸

KaliBench is a fine-grained benchmark for evaluating LLM CLI command generation on Kali Linux. While open-weight models score below 42% accuracy, training an 8B model using KaliBench's runtime-free verifiable rewards allows it to rival a 685B MoE model.

2 min read
VISTA: Empowering Multimodal Agents with a Long-Horizon Visual Harness
arXivAI Agent

VISTA: Empowering Multimodal Agents with a Long-Horizon Visual Harness

VISTA:為多模態 AI 打造的「長視域」互動式視覺輔助框架

VISTA is a visual harness that equips multimodal models with long-horizon vision and lossless memory, dramatically improving reasoning and efficiency in interactive environments.

2 min read