AutoSynthData: Generating Targeted Training Data from Enterprise Agent Failures
AutoSynthData:以企業 Agent 的失敗為師,自動生成高規格微調訓練資料

General LLMs often struggle with specific enterprise workflows and tools. ServiceNow's AutoSynthData addresses this by comparing a target model's failures with a stronger teacher's successes to produce "capability specification cards." It then generates brand-new, realistic training tasks. Each task undergoes rigorous positive and negative verification gates, with failed tasks repaired via a critic loop. Tested on EnterpriseOps Gym, this targeted synthetic data approach boosted the target model's Hybrid Pass@1 rate by 7.2 percentage points, representing a 35% relative improvement.
Key points
Target Capability Gaps
Identifies specific weaknesses by comparing the target model's failures with a stronger teacher's successes, distilling them into structured capability cards.
Double-Gate Verification & Repair
Uses positive gates (running the reference solution) and negative gates (mutating outcomes to catch weak verifiers) combined with a critic-led repair loop.
Target & Multiply Strategy
Scales datasets by first generating vetted core samples, then expanding them into diverse, safe variants without suffering from generational drift.
Proven Enterprise Performance
Successfully raised Gemma's Pass@1 rates in EnterpriseOps Gym Hybrid (by 35% relatively) and ITSM environments, closing up to 59% of the teacher gap.
How it works
Why it matters
The biggest hurdle for deploying enterprise agents is the lack of high-quality, environment-specific training data. AutoSynthData turns failures into valuable training signals systematically. Beyond SFT, this automated, closed-loop paradigm can dynamically feed reinforcement learning (RL) pipelines with difficulty-calibrated tasks, drastically lowering the cost of adapting LLMs to complex, stateful corporate environments.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
- Product Manager
How to use it
- 1Automatically generating SFT data for ITSM agents to strictly adhere to customized corporate workflows and tool chains.
- 2Creating verified tool-use tasks and evaluation benchmarks for agents when onboarding new proprietary APIs or databases.
Limitations & caveats
- Highly dependent on a stronger teacher model; if the teacher itself cannot solve the task in highly specialized environments, data generation fails.
- Designing robust and unbiased positive/negative verifiers for complex, stateful environments remains challenging.
Related
KaliBench: Evaluating and Boosting LLM Command Generation on Kali Linux
KaliBench:首個 Kali Linux 資安工具指令生成基準測試,助 8B 模型直逼 685B 巨獸
KaliBench is a fine-grained benchmark for evaluating LLM CLI command generation on Kali Linux. While open-weight models score below 42% accuracy, training an 8B model using KaliBench's runtime-free verifiable rewards allows it to rival a 685B MoE model.
VISTA: Empowering Multimodal Agents with a Long-Horizon Visual Harness
VISTA:為多模態 AI 打造的「長視域」互動式視覺輔助框架
VISTA is a visual harness that equips multimodal models with long-horizon vision and lossless memory, dramatically improving reasoning and efficiency in interactive environments.

Google Announces Gemini 4 Argon: Frontier Model with 1M Output Tokens and Long-Horizon Reasoning
Google 發表新一代前沿模型 Gemini 4 Argon:具備 100 萬 Token 輸出與強大自主 Agent 推理能力
Google has unveiled Gemini 4 Argon, featuring an industry-leading 1-million output token limit designed to sustain deep reasoning across complex, long-horizon workflows.