Aivora
NVIDIA DeveloperAI AgentIntermediate

Tracing AI Agent Trajectories and Performance with NVIDIA NeMo Relay

使用 NVIDIA NeMo Relay 追蹤 AI Agent 的執行軌跡與效能

2 min read
Tracing AI Agent Trajectories and Performance with NVIDIA NeMo Relay
The 30-second version

Standard evaluations only check if an AI agent succeeded, ignoring inefficient hidden steps like redundant tool calls and high token usage. NVIDIA NeMo Relay addresses this by introducing structured observability for agent harnesses like Hermes Agent. It outputs event-level logs (ATOF) and step-by-step trajectories (ATIF), and exports OpenTelemetry traces to platforms like Arize Phoenix. This allows developers to analyze exact LLM calls, tool errors, and duration, transforming agent evaluation from simple success checks into deterministic performance optimization.

Key points

01

Expose Hidden Inefficiencies

Correct final answers can mask redundant search queries, tool retries, and high latency that waste processing tokens.

02

Dual ATOF & ATIF Formats

ATOF logs raw lifecycle events and timing, while ATIF compiles these events into a sequential, step-by-step interaction history.

03

Standard OpenTelemetry Export

Seamlessly export traces via the OpenInference exporter to visualization backends like Arize Phoenix or LangSmith.

04

Systematic Harness Evaluation

Compare baseline and modified agent runs systematically to verify if changes truly optimize behavior or just alter retry thresholds.

How it works

Comparison of Three Observability Formats in NeMo Relay
ATOFATIFOpenTelemetry
Format StructureJSONL 串流日誌JSON 步驟紀錄Parent-Child Spans 結構
Core Content生命週期事件與 UUID 關聯彙整後的工具調用與觀察Token、延遲、錯誤屬性指標
Best Use Case除錯個別事件與時間排序逐步分析評估 Agent 解決路徑在 Phoenix 等工具中視覺化分析

Why it matters

In enterprise environments, relying solely on task completion rates masks high token costs and latency bottlenecks. NeMo Relay provides a structured evidence layer for safety governance and performance auditing. This telemetry allows organizations to optimize agent prompts, tools, and configurations deterministically, ensuring that bug fixes do not inadvertently lead to bloated LLM calls or runaway execution paths.

Who it affects

  • AI Developer
  • AI Researcher
  • Enterprise Leader

How to use it

  1. 1Trace and debug infinite retry loops where an agent repeatedly runs failing terminal commands.
  2. 2Benchmark and compare different models on complex tool-use tasks to measure total token consumption and execution speed.

Limitations & caveats

  • Exported trace payloads can contain highly sensitive information like prompts, tool outputs, and internal directory paths.
  • Rigorous harness evaluations require frozen external environments (such as cached web searches) to prevent dynamic network fluctuations from corrupting benchmark data.

Related

AutoSynthData: Generating Targeted Training Data from Enterprise Agent Failures
Hugging FaceAI Agent

AutoSynthData: Generating Targeted Training Data from Enterprise Agent Failures

AutoSynthData:以企業 Agent 的失敗為師,自動生成高規格微調訓練資料

AutoSynthData is a framework by ServiceNow that analyzes an enterprise agent's failures against a stronger teacher to automatically generate and validate high-quality synthetic training data, bridging crucial performance gaps.

2 min read
KaliBench: Evaluating and Boosting LLM Command Generation on Kali Linux
arXivAI Agent

KaliBench: Evaluating and Boosting LLM Command Generation on Kali Linux

KaliBench:首個 Kali Linux 資安工具指令生成基準測試,助 8B 模型直逼 685B 巨獸

KaliBench is a fine-grained benchmark for evaluating LLM CLI command generation on Kali Linux. While open-weight models score below 42% accuracy, training an 8B model using KaliBench's runtime-free verifiable rewards allows it to rival a 685B MoE model.

2 min read
VISTA: Empowering Multimodal Agents with a Long-Horizon Visual Harness
arXivAI Agent

VISTA: Empowering Multimodal Agents with a Long-Horizon Visual Harness

VISTA:為多模態 AI 打造的「長視域」互動式視覺輔助框架

VISTA is a visual harness that equips multimodal models with long-horizon vision and lossless memory, dramatically improving reasoning and efficiency in interactive environments.

2 min read