Aivora
NVIDIA DeveloperAI AgentIntermediate

Building Reliable Data Analytics Agents: Lessons from NVIDIA's KDD Cup Harness Design

NVIDIA 於 KDD Cup 的實務經驗:打造可靠資料分析 Agent 的 Harness 設計原則

2 min read
Building Reliable Data Analytics Agents: Lessons from NVIDIA's KDD Cup Harness Design
The 30-second version

In the KDD Cup Data Agents competition, teams used a small, fixed LLM to solve analytics tasks spanning databases, documents, and videos. NVIDIA's KGMON team unified CSV and JSON data into a single SQLite interface and added a schema scouting preflight step before the main agent loop. By constraining the agent to a tight toolset within a persistent Python environment and inspecting execution traces, they reduced reasoning turns and overall error rates.

Key points

01

Unified Data Access & Schema Scouting

Converts heterogeneous CSV and JSON files into a single SQLite interface and runs a preflight schema check to save discovery turns and prevent incorrect joins.

02

Constrained Tools & Persistent State

Limits the agent to a minimal toolset while retaining variables in a persistent Python environment, backed by middleware that fixes malformed tool calls.

03

Context-Conscious Doc & Video Processing

Blocks direct full-file reading and leverages specialized doc/video helper tools to extract targeted evidence without overloading the main context window.

04

Trajectory Inspection & Selective Ensembling

Logs full execution traces for root-cause inspection while selectively grouping answers across multiple runs to boost output accuracy on contested tasks.

How it works

KGMON Reliable Data Analytics Agent Architecture
NormalizeSchema ContextTargeted QueriesExtracted RulesLog TracesAnswer EvaluationHeterogeneous DataEnsembling & OutputSQLite & SchemaScoutingPersistent Env & ToolsDoc & Video HelpersTrace & Inspection

Why it matters

Agent design often over-relies on larger foundation models, but enterprise settings face strict compute constraints. NVIDIA's approach shows that building a clear, constrained, and inspectable harness around smaller models delivers superior reliability, offering a practical implementation model for small and open LLMs.

Who it affects

  • AI Developer
  • AI Researcher
  • Product Manager
  • Enterprise Leader

How to use it

  1. 1Enterprise automated data analytics across heterogeneous databases and unstructured documents
  2. 2Lightweight domain-specific AI Agent development powered by small open LLMs
  3. 3Automated trajectory diagnosis and error taxonomy for AI agent systems

Limitations & caveats

  • Repeated runs and answer ensembling increase compute costs and latency, requiring selective usage based on task criticalness.
  • Video keyframe processing and table extractions add upfront implementation complexity to the harness.

Related

BrickBench: Evaluating Agentic Text-Conditioned LEGO Design
arXivAI Agent

BrickBench: Evaluating Agentic Text-Conditioned LEGO Design

BrickBench:評測 AI Agent 進行文字導向積木設計的全新基準

BrickBench is a benchmark for evaluating agentic text-to-LEGO design, equipped with the BrickAgent environment to assess agents' physical validity, semantic alignment, and design quality.

2 min read
RECAST: Active Evidence Construction via Adaptive Routing and Computation
arXivAI Agent

RECAST: Active Evidence Construction via Adaptive Routing and Computation

RECAST:透過自適應證據路由,為 LLM 計算出正確的上下文

RECAST is an agentic framework that treats evidence construction as a sequential decision process, utilizing a Router-Compiler loop to actively compute and derive the optimal context for RAG rather than relying on passive retrieval.

2 min read