AutoSynthData turns agent failures into training data
ServiceNow's CoreAI team has built a pipeline that reads a model's failures, has a stronger teacher show the fix, and generates validated tasks at training scale, with measured gains in two enterprise environments.

A model can be broadly capable and still fall over on the particular systems a company actually runs. ServiceNow's CoreAI team built AutoSynthData to find those weak spots and turn them into training data, and it has published the pipeline along with measured results in two enterprise environments.
The method pairs a target model with a stronger teacher. The target's failures and the teacher's successes decide what the model should learn next, the pipeline generates fresh tasks that exercise that capability, and every candidate is validated before it reaches the training set. As the model improves, the curriculum follows it into whatever it still cannot do.
What makes a task usable
ServiceNow defines a task as three parts: a system specification, a user prompt and a verifier. A generated task has to be feasible in the environment, realistic enough that somebody might genuinely ask for it, and difficult enough to expose a weakness. Tasks the model already solves reliably teach it nothing, so the useful region is the narrow band of work that is possible and plausible but not yet dependable.
Candidates clear a two-sided gate. The reference solution is executed to confirm it really solves the task, and the expected outcome is then mutated to confirm that wrong answers fail verification. Anything rejected goes to a critic, which diagnoses the fault and drives a targeted repair instead of forcing generation to start from scratch. Batches are reviewed too, so a dataset does not collapse into a handful of easy task families.
The measured gains
In the Hybrid domain of EnterpriseOps Gym, with Gemma-4-26B-A4B-it as the target and Qwen3.8-27B as the teacher, AutoSynthData generated 2,000 training samples in about 18 hours. The best checkpoint, at epoch five, lifted mean Pass@1 by 7.2 percentage points - a 35% relative improvement - and raised verifier success from 63.01% to 68.55%, closing 59% of the original gap between the target model and the reference.
On the ITSM domain, this time with DeepSeek-V4.1-Flash as the teacher, 1,994 samples took 66 hours and moved mean Pass@1 from 18.77% to 27.18%. The training tasks were built from capability specifications; the generator never saw the original evaluation prompts.
Our opinion
Synthetic data has a reputation problem, and it is earned: most pipelines can produce plausible-looking tasks but cannot prove those tasks are solvable, and a lazy verifier will happily reward the wrong answer. The interesting part of AutoSynthData is that the negative gate does the heavy lifting, mutating the expected outcome and checking that it fails. The honest framing of the training frontier as something that moves as the model improves is the other detail worth copying; a dataset is a snapshot, and treating it as a fixed asset is how post-training quietly stalls.