Building Reliable Data Analytics Agents: Lessons from NVIDIA's KDD Cup Harness Design
NVIDIA 於 KDD Cup 的實務經驗:打造可靠資料分析 Agent 的 Harness 設計原則

In the KDD Cup Data Agents competition, teams used a small, fixed LLM to solve analytics tasks spanning databases, documents, and videos. NVIDIA's KGMON team unified CSV and JSON data into a single SQLite interface and added a schema scouting preflight step before the main agent loop. By constraining the agent to a tight toolset within a persistent Python environment and inspecting execution traces, they reduced reasoning turns and overall error rates.
Key points
Unified Data Access & Schema Scouting
Converts heterogeneous CSV and JSON files into a single SQLite interface and runs a preflight schema check to save discovery turns and prevent incorrect joins.
Constrained Tools & Persistent State
Limits the agent to a minimal toolset while retaining variables in a persistent Python environment, backed by middleware that fixes malformed tool calls.
Context-Conscious Doc & Video Processing
Blocks direct full-file reading and leverages specialized doc/video helper tools to extract targeted evidence without overloading the main context window.
Trajectory Inspection & Selective Ensembling
Logs full execution traces for root-cause inspection while selectively grouping answers across multiple runs to boost output accuracy on contested tasks.
How it works
Why it matters
Agent design often over-relies on larger foundation models, but enterprise settings face strict compute constraints. NVIDIA's approach shows that building a clear, constrained, and inspectable harness around smaller models delivers superior reliability, offering a practical implementation model for small and open LLMs.
Who it affects
- AI Developer
- AI Researcher
- Product Manager
- Enterprise Leader
How to use it
- 1Enterprise automated data analytics across heterogeneous databases and unstructured documents
- 2Lightweight domain-specific AI Agent development powered by small open LLMs
- 3Automated trajectory diagnosis and error taxonomy for AI agent systems
Limitations & caveats
- Repeated runs and answer ensembling increase compute costs and latency, requiring selective usage based on task criticalness.
- Video keyframe processing and table extractions add upfront implementation complexity to the harness.
Related
BrickBench: Evaluating Agentic Text-Conditioned LEGO Design
BrickBench:評測 AI Agent 進行文字導向積木設計的全新基準
BrickBench is a benchmark for evaluating agentic text-to-LEGO design, equipped with the BrickAgent environment to assess agents' physical validity, semantic alignment, and design quality.
RECAST: Active Evidence Construction via Adaptive Routing and Computation
RECAST:透過自適應證據路由,為 LLM 計算出正確的上下文
RECAST is an agentic framework that treats evidence construction as a sequential decision process, utilizing a Router-Compiler loop to actively compute and derive the optimal context for RAG rather than relying on passive retrieval.

Microsoft Releases Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Real-Harness Training
微軟開源 Agent Lightning v1.0:僅 3,500 行程式碼,直接用生產環境 Harness 訓練 Agent 的強化學習框架
Microsoft Research Asia has introduced Agent Lightning v1.0, a lightweight open-source framework that trains AI agents using their actual deployment harnesses via a transparent LLM proxy.