AutoSynthData: Generating Training Data for Enterprise Agents

2026-10-03 · Hugging Face

AutoSynthData: Generating Training Data for Enterprise Agents

Enterprises require agents that perform effectively within their specific environments. Even a broadly capable model may struggle in a particular setting due to poorly handled workflows, misused tool combinations, or unrespected constraints. Transforming these weaknesses into training data is crucial. To address this, ServiceNow CoreAI developed AutoSynthData, a pipeline designed to convert capability gaps into actionable training data.

What Makes a Useful Agentic Task?

An agentic environment defines the world an agent operates in, including observable states, modifiable conditions, invocable tools, and state transitions. A task instantiated within this environment is abstracted as: `task = (system specification, user prompt, verifier)`.

System Specification

This defines the constraints under which the agent operates, including system instructions, environment policies, and task-specific initializations like seeded database states. The specification must be compatible with the environment's tools and actions, with clear instructions that avoid arbitrary difficulty.

Agent-Facing Task

The user prompt specifies the user's goal and constraints. A generated task must satisfy three properties:

  • Feasibility: There must exist at least one trajectory satisfying the prompt while respecting the system specification, ruling out tasks dependent on unavailable tools or policies.
  • Realism: The prompt should resemble a plausible user request within the target environment.
  • Difficulty: The task should expose a weakness of the current agent, as already solved tasks provide little training signal.

Verifier

The verifier determines if a trajectory successfully completes the task, requiring three properties:

  • Consistency: It must agree with the user prompt, system specification, and environment state.
  • Soundness: It must reject trajectories that fail the task or violate constraints.
  • Completeness: It must accept valid solutions rather than encoding a single reference trajectory.

Overview of AutoSynthData

Given an environment and a target model, AutoSynthData generates training tasks comprising a system specification, user prompt, and verifier. These tasks are grounded in the environment to provide useful training signals.

AutoSynthData first evaluates the target model using diagnostic tasks to identify patterns of struggle. A stronger teacher model helps characterize solvable tasks and successful behaviors. AutoSynthData then translates these capability gaps into new executable tasks, validates them in the environment, and uses accepted samples for post-training. Evaluating the updated model reveals remaining gaps, guiding the next generation cycle.

From Model Failures to a Curriculum

AutoSynthData uses evaluation runs to identify what the model needs to learn next. In the EnterpriseOps Gym experiment, both the target and teacher models are run on evaluation tasks to identify:

  • The capability being tested
  • The tools and workflow structure involved
  • Where the target model fails and how the teacher succeeds
  • Properties a correct final state must satisfy
  • Dimensions that can vary while preserving the tested capability

These findings are distilled into sanitized capability specification cards. The generator receives these cards—without original prompts, entities, or trajectories—to create new tasks with different prompts, states, and solution paths.

Generating and Scaling Tasks

Identifying a capability gap tells us what to teach, but training requires many varied tasks to exercise it. AutoSynthData uses the specification card to generate these tasks at scale. As the model improves, the curriculum shifts toward remaining difficulties, creating a continuous optimization loop.

Source