Aivora
Hugging FaceAI AgentIntermediate

AutoSynthData: Generating Targeted Training Data from Enterprise Agent Failures

AutoSynthData:以企業 Agent 的失敗為師,自動生成高規格微調訓練資料

2 min read
AutoSynthData: Generating Targeted Training Data from Enterprise Agent Failures
The 30-second version

General LLMs often struggle with specific enterprise workflows and tools. ServiceNow's AutoSynthData addresses this by comparing a target model's failures with a stronger teacher's successes to produce "capability specification cards." It then generates brand-new, realistic training tasks. Each task undergoes rigorous positive and negative verification gates, with failed tasks repaired via a critic loop. Tested on EnterpriseOps Gym, this targeted synthetic data approach boosted the target model's Hybrid Pass@1 rate by 7.2 percentage points, representing a 35% relative improvement.

Key points

01

Target Capability Gaps

Identifies specific weaknesses by comparing the target model's failures with a stronger teacher's successes, distilling them into structured capability cards.

02

Double-Gate Verification & Repair

Uses positive gates (running the reference solution) and negative gates (mutating outcomes to catch weak verifiers) combined with a critic-led repair loop.

03

Target & Multiply Strategy

Scales datasets by first generating vetted core samples, then expanding them into diverse, safe variants without suffering from generational drift.

04

Proven Enterprise Performance

Successfully raised Gemma's Pass@1 rates in EnterpriseOps Gym Hybrid (by 35% relatively) and ITSM environments, closing up to 59% of the teacher gap.

How it works

AutoSynthData Generation & Verification Workflow
Analyze failuresGuide generationSubmit candidatePassFailPassFailRepaired retryEvaluate GapCreate Spec CardGenerate TaskNegative GateCritic & RepairSFT DatasetPositive Gate

Why it matters

The biggest hurdle for deploying enterprise agents is the lack of high-quality, environment-specific training data. AutoSynthData turns failures into valuable training signals systematically. Beyond SFT, this automated, closed-loop paradigm can dynamically feed reinforcement learning (RL) pipelines with difficulty-calibrated tasks, drastically lowering the cost of adapting LLMs to complex, stateful corporate environments.

Who it affects

  • AI Developer
  • AI Researcher
  • Enterprise Leader
  • Product Manager

How to use it

  1. 1Automatically generating SFT data for ITSM agents to strictly adhere to customized corporate workflows and tool chains.
  2. 2Creating verified tool-use tasks and evaluation benchmarks for agents when onboarding new proprietary APIs or databases.

Limitations & caveats

  • Highly dependent on a stronger teacher model; if the teacher itself cannot solve the task in highly specialized environments, data generation fails.
  • Designing robust and unbiased positive/negative verifiers for complex, stateful environments remains challenging.

Related

KaliBench: Evaluating and Boosting LLM Command Generation on Kali Linux
arXivAI Agent

KaliBench: Evaluating and Boosting LLM Command Generation on Kali Linux

KaliBench:首個 Kali Linux 資安工具指令生成基準測試,助 8B 模型直逼 685B 巨獸

KaliBench is a fine-grained benchmark for evaluating LLM CLI command generation on Kali Linux. While open-weight models score below 42% accuracy, training an 8B model using KaliBench's runtime-free verifiable rewards allows it to rival a 685B MoE model.

2 min read
VISTA: Empowering Multimodal Agents with a Long-Horizon Visual Harness
arXivAI Agent

VISTA: Empowering Multimodal Agents with a Long-Horizon Visual Harness

VISTA:為多模態 AI 打造的「長視域」互動式視覺輔助框架

VISTA is a visual harness that equips multimodal models with long-horizon vision and lossless memory, dramatically improving reasoning and efficiency in interactive environments.

2 min read