ServiceNow CoreAI has introduced AutoSynthData, a new pipeline designed to generate high-quality training data for enterprise AI agents by targeting specific capability gaps. The system identifies weaknesses in a target model’s performance and uses a stronger teacher model to create, validate, and scale synthetic tasks that address those specific deficiencies.
What Happened
AutoSynthData operates by evaluating a target model in a specific environment to identify patterns in tasks it struggles to complete. It distills these findings into capability specification cards, which guide the generation of new tasks that are feasible, realistic, and difficult enough to provide training signal. The pipeline employs a two-phase process: a 'target' phase that creates core training samples and a 'multiply' phase that expands the dataset by generating novel variants of accepted samples.
To ensure data quality, each candidate task undergoes rigorous verification, including positive checks to confirm the reference solution works and negative checks to ensure incorrect outcomes fail. A critic and repair loop handles failed candidates by diagnosing issues such as inconsistent state or bad reference trajectories. The system also performs batch-level reviews to prevent overrepresentation of easy tasks and to direct generation efforts toward missing capabilities.
In experiments using EnterpriseOps Gym, ServiceNow tested the pipeline on the Hybrid domain using Gemma-4-26B-A4B-it as the target model and Qwen3.8-27B as the teacher. AutoSynthData generated 2,000 synthetic training samples in approximately 18 hours. Fine-tuning the Gemma model on this dataset resulted in a best checkpoint at epoch 5.
Why It Matters
The introduction of AutoSynthData addresses a critical bottleneck in enterprise AI: the scarcity of high-quality, domain-specific training data. While broadly capable models exist, they often fail in specific enterprise environments due to unique workflows, tools, and constraints. By automating the creation of tasks that target these specific failures, organizations can more efficiently fine-tune models to handle their proprietary systems.
The results from the EnterpriseOps Gym experiments demonstrate significant potential. In the Hybrid domain, the synthetic supervised fine-tuning (SFT) checkpoint improved mean Pass@1 by 7.2 percentage points, representing a 35% relative improvement. Verifier success rates rose from 63.01% to 68.55%, closing 59% of the original performance gap between the target model and the reference model. The approach was also validated in the ITSM domain, where it raised mean Pass@1 from 18.77% to 27.18% using a different teacher model, DeepSeek-V4.1-Flash.
The Bottom Line
AutoSynthData provides a structured method for turning agent failures into actionable training data, validated through rigorous execution and verification loops. By focusing on the boundary of a model's capabilities, it enables more efficient post-training for enterprise agents, with ServiceNow reporting substantial performance gains in controlled experimental settings.