Aivora
NVIDIA DeveloperAI AgentIntermediate

From Tool Calls to Task Completion: How to Systematically Evaluate AI Agents

從工具調用到任務完成:如何系統化評估 AI Agent

2 min read
From Tool Calls to Task Completion: How to Systematically Evaluate AI Agents
The 30-second version

Traditional benchmarks only evaluate single tool calls, but real-world AI agents must execute multi-turn sequences, manage errors, and maintain state. Modern evaluation requires a live execution environment to measure both End-to-End (E2E) outcomes (did the task finish?) and step-level process metrics (was each tool call valid?), enabling developers to optimize reliability, cost, and latency.

Key points

01

Shift to Execution Environments

Correct individual tool calls do not guarantee task success. Evaluation must occur in an environment that updates state, verifying whether the final state meets requirements.

02

Dual-Layer Scoring

End-to-End (E2E) scoring evaluates the final state (e.g., did the refund post?) for release gating, while step-level scoring traces path validity to guide debugging and fine-tuning.

03

Three Axes of Metrics

Beyond accuracy (success rate and consistency over 3-5 trials), teams must track verbosity (steps per success) and cost (spend per successful task) to measure efficiency.

04

Dimensions of Quality Benchmarks

Robust benchmarks feature high task complexity, statefulness, and rely on executable verification (e.g., running unit tests) rather than just LLM-as-a-Judge.

How it works

Step-Level (Process) vs. End-to-End (Outcome) Evaluation
步驟級(過程)評估 / Process Scoring端到端(結果)評估 / Outcome Scoring
Evaluation Focus評估單次工具呼叫是否有效、相關、且對當前狀態有用忽略路徑,僅確認最終的系統或環境狀態是否達成目標
Primary Use Case開發、除錯與模型微調(找出是哪一步驟導致工具調用鏈斷裂)生產環境發佈的檢驗閘口(衡量用戶經歷的真實成效)
Handling of Redundant Steps會扣減「工具調用精準度」與步驟分數(例如做了無效的檔案檢視)不影響分數,只要最終結果正確,即判定為成功(成功率為 1)
Evaluation Target追蹤紀錄(Trace)中的每一行對話與個別工具參數任務結束後的資料庫狀態、通過的程式測試或關閉的工單

Why it matters

As enterprises deploy AI agents, evaluating whether a model 'sounds right' is no longer sufficient. Agents that cannot interact with live environments or recover from step-level errors cause direct business failures. A systematic evaluation framework combining E2E validation and step-level tracing allows teams to quantify reliability, predict operational costs, and pinpoint exactly which tool call broke the chain before shipping.

Who it affects

  • AI Developer
  • AI Researcher
  • Product Manager
  • Enterprise Leader

How to use it

  1. 1Using SWE-bench Verified to evaluate an AI coding agent's ability to navigate codebases, write patches, and pass unit tests.
  2. 2Building private domain evaluations based on internal APIs, databases, and actual tickets to gate agent releases.
  3. 3Tracing agent trajectories to eliminate redundant steps (such as unnecessary file views) to reduce token costs and latency.

Limitations & caveats

  • Evaluation is highly susceptible to data contamination, such as web-searching agents retrieving live answer keys or public datasets being scraped into pretraining corpora.
  • When executable checks are absent, LLM-as-a-Judge can fill the gap, but its scores remain provisional and must be continuously validated against human expert ratings.

Related

AutoSynthData: Generating Targeted Training Data from Enterprise Agent Failures
Hugging FaceAI Agent

AutoSynthData: Generating Targeted Training Data from Enterprise Agent Failures

AutoSynthData:以企業 Agent 的失敗為師,自動生成高規格微調訓練資料

AutoSynthData is a framework by ServiceNow that analyzes an enterprise agent's failures against a stronger teacher to automatically generate and validate high-quality synthetic training data, bridging crucial performance gaps.

2 min read
KaliBench: Evaluating and Boosting LLM Command Generation on Kali Linux
arXivAI Agent

KaliBench: Evaluating and Boosting LLM Command Generation on Kali Linux

KaliBench:首個 Kali Linux 資安工具指令生成基準測試,助 8B 模型直逼 685B 巨獸

KaliBench is a fine-grained benchmark for evaluating LLM CLI command generation on Kali Linux. While open-weight models score below 42% accuracy, training an 8B model using KaliBench's runtime-free verifiable rewards allows it to rival a 685B MoE model.

2 min read