Tracing AI Agent Trajectories and Performance with NVIDIA NeMo Relay
使用 NVIDIA NeMo Relay 追蹤 AI Agent 的執行軌跡與效能

Standard evaluations only check if an AI agent succeeded, ignoring inefficient hidden steps like redundant tool calls and high token usage. NVIDIA NeMo Relay addresses this by introducing structured observability for agent harnesses like Hermes Agent. It outputs event-level logs (ATOF) and step-by-step trajectories (ATIF), and exports OpenTelemetry traces to platforms like Arize Phoenix. This allows developers to analyze exact LLM calls, tool errors, and duration, transforming agent evaluation from simple success checks into deterministic performance optimization.
Key points
Expose Hidden Inefficiencies
Correct final answers can mask redundant search queries, tool retries, and high latency that waste processing tokens.
Dual ATOF & ATIF Formats
ATOF logs raw lifecycle events and timing, while ATIF compiles these events into a sequential, step-by-step interaction history.
Standard OpenTelemetry Export
Seamlessly export traces via the OpenInference exporter to visualization backends like Arize Phoenix or LangSmith.
Systematic Harness Evaluation
Compare baseline and modified agent runs systematically to verify if changes truly optimize behavior or just alter retry thresholds.
How it works
| ATOF | ATIF | OpenTelemetry | |
|---|---|---|---|
| Format Structure | JSONL 串流日誌 | JSON 步驟紀錄 | Parent-Child Spans 結構 |
| Core Content | 生命週期事件與 UUID 關聯 | 彙整後的工具調用與觀察 | Token、延遲、錯誤屬性指標 |
| Best Use Case | 除錯個別事件與時間排序 | 逐步分析評估 Agent 解決路徑 | 在 Phoenix 等工具中視覺化分析 |
Why it matters
In enterprise environments, relying solely on task completion rates masks high token costs and latency bottlenecks. NeMo Relay provides a structured evidence layer for safety governance and performance auditing. This telemetry allows organizations to optimize agent prompts, tools, and configurations deterministically, ensuring that bug fixes do not inadvertently lead to bloated LLM calls or runaway execution paths.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
How to use it
- 1Trace and debug infinite retry loops where an agent repeatedly runs failing terminal commands.
- 2Benchmark and compare different models on complex tool-use tasks to measure total token consumption and execution speed.
Limitations & caveats
- Exported trace payloads can contain highly sensitive information like prompts, tool outputs, and internal directory paths.
- Rigorous harness evaluations require frozen external environments (such as cached web searches) to prevent dynamic network fluctuations from corrupting benchmark data.
Related

AutoSynthData: Generating Targeted Training Data from Enterprise Agent Failures
AutoSynthData:以企業 Agent 的失敗為師,自動生成高規格微調訓練資料
AutoSynthData is a framework by ServiceNow that analyzes an enterprise agent's failures against a stronger teacher to automatically generate and validate high-quality synthetic training data, bridging crucial performance gaps.
KaliBench: Evaluating and Boosting LLM Command Generation on Kali Linux
KaliBench:首個 Kali Linux 資安工具指令生成基準測試,助 8B 模型直逼 685B 巨獸
KaliBench is a fine-grained benchmark for evaluating LLM CLI command generation on Kali Linux. While open-weight models score below 42% accuracy, training an 8B model using KaliBench's runtime-free verifiable rewards allows it to rival a 685B MoE model.
VISTA: Empowering Multimodal Agents with a Long-Horizon Visual Harness
VISTA:為多模態 AI 打造的「長視域」互動式視覺輔助框架
VISTA is a visual harness that equips multimodal models with long-horizon vision and lossless memory, dramatically improving reasoning and efficiency in interactive environments.