🔍 Read the full analysis: AutoSynthData Offers A New Approach To Enterprise Agent Training Data on ThorstenMeyerAI.com
Get monitors, keyboards and dev gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
ServiceNow CoreAI describes AutoSynthData, a system that uses an enterprise agent’s failures and a stronger teacher model’s successful runs to generate new training tasks. The company points to EnterpriseOps Gym as an example, but the supplied material reports no measured performance gains, task counts or comparisons with other methods.
ServiceNow CoreAI says it has built AutoSynthData, a system that turns an enterprise agent’s observed failures into new training tasks and checks those tasks in the target environment, as described in the original analysis. The company illustrates the method with the released EnterpriseOps Gym dataset, but the supplied account reports no measured improvement in model performance or comparison with other training approaches.
According to ServiceNow CoreAI’s description, AutoSynthData first evaluates a target model on diagnostic tasks in an enterprise environment. A stronger teacher model attempts the same tasks, giving the system examples of both where the target struggles and how a more capable model can succeed. The resulting information is used to identify the capability being tested, relevant tools and workflows, the target’s failure points and the conditions for a valid result.
The system distills those findings into what the company calls sanitized capability specification cards. AutoSynthData’s task generators use the cards, rather than the original prompts, entities, execution traces or verifier details, to create new tasks. The tasks vary wording, starting conditions, entities, tools, workflow combinations and difficulty. Each includes an environment specification, a user prompt and a verifier that checks whether the request was completed within the environment’s constraints.
ServiceNow CoreAI says generated tasks are checked in the environment and accepted samples can be used for post-training. The updated model can then be evaluated again, with remaining weaknesses informing another generation round. The supplied material does not say how many tasks were generated or accepted, which models were used, or how the target’s results changed after training.
Training Agents for Local Workflows
Enterprise agents must do more than produce plausible text: they may need to update records, follow access rules, use particular tools in a permitted order and leave a system in a required state. A model that performs well on broad tests may still fail at organization-specific workflows. AutoSynthData is intended to create practice tasks around those local gaps rather than relying only on generic examples or manually written training data.
The verifier is a central part of that proposal. If it accepts an incorrect outcome, training could reward the wrong behavior; if it rejects a valid solution because it expects one specific sequence of actions, it could penalize sound behavior. ServiceNow’s description says verifiers should reflect the request and environment, reject failures and policy violations, and accept valid solutions without requiring an exact trajectory.
That matters because enterprise agents can change operational data, not just generate answers. If the approach works as described, it could give teams a way to generate varied tasks that test whether an agent can act safely and correctly in a particular system. Whether it improves reliability, lowers data costs or transfers to other environments remains unreported.
enterprise AI training data generation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
From Evaluation Failures to Tasks
In the framework described by ServiceNow CoreAI, an agentic environment sets out what an agent can observe and change, which tools and APIs it can access, and how its actions affect the system. A task combines that environment’s instructions and policies with a user-facing request and a verifier. It may also include setup such as a seeded database or knowledge articles.
The distinction between an executable task and a realistic one is part of the problem AutoSynthData aims to address. A request may be possible to carry out but implausible as enterprise work; another may sound realistic but be impossible because a tool or required information is unavailable, a state change cannot be made, or policy forbids it. The company says its generated tasks are meant to be feasible, realistic and challenging for the current model.
For its example, ServiceNow CoreAI cites the released EnterpriseOps Gym dataset and references Malay et al. (2026). The supplied account does not describe the dataset’s size, the specific workflows assessed or model scores. It also gives no publication date for the AutoSynthData description.
““A model may be broadly capable and still struggle with a particular environment.””
— ServiceNow CoreAI
AI model testing and verification software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Performance Evidence Is Missing
The supplied material describes a method, not an evaluation demonstrating its effectiveness. It gives no before-and-after performance results, measurement method or comparison against a baseline, such as manually written tasks or another data-generation approach. Task generation and verifier acceptance counts are also absent, so readers cannot judge the volume or quality of the training data described.
The account does not identify the target or teacher models, the amount of post-training, or the time and cost required to run the pipeline. It is also unclear whether results in EnterpriseOps Gym would carry over to other enterprise systems and workflows. Although the generator is said to receive capability cards rather than original evaluation details, the material provides no analysis of possible overlap between generated tasks and evaluation material.
Without those details, the strength of the EnterpriseOps Gym example and the practical value of the approach cannot be assessed from the supplied information alone. Claims about improved reliability, lower costs or broader generalization would be premature.
enterprise agent workflow automation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Results Needed to Test the Method
The next useful evidence would be a reported evaluation of models trained with AutoSynthData. That should include task-generation totals, verifier acceptance rates, the target model’s performance before and after training, and a clear comparison with an appropriate baseline. Results broken out across workflows would help show whether gains reflect better task-specific performance or broader improvement.
Evaluation across more than one environment would also help establish whether the method applies beyond EnterpriseOps Gym. Details about the teacher and target models, training volume, costs and safeguards against overlap with evaluation tasks would make the results easier to interpret. No such results or timeline are provided in the material supplied, so further evidence remains pending.
AI model performance evaluation platforms
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is AutoSynthData?
AutoSynthData is a system described by ServiceNow CoreAI that uses an enterprise agent’s observed failures, alongside a stronger teacher model’s successful attempts, to generate and check new training tasks.
How does the system create training tasks?
ServiceNow says the system distills observed capability gaps into specification cards. A generator uses those cards to create varied tasks, each with an environment specification, a user prompt and a verifier; the tasks are then checked in the target environment.
Has AutoSynthData been shown to improve agent performance?
The supplied material reports no measured improvement. It provides no before-and-after results, benchmark scores or comparison with other training methods.
What dataset does ServiceNow cite?
ServiceNow CoreAI points to the released EnterpriseOps Gym dataset as an example. The supplied account does not give its size, the specific workflows tested or model scores.
What information is still needed to assess the approach?
Readers would need task and verifier acceptance counts, model identities, training volume, before-and-after results, baseline comparisons, costs and evidence across multiple enterprise environments. The material does not provide those details.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
