Enterprises need agents that work well in their own environments. A model may be broadly capable but struggle with a particular environment. This includes workflows it handles poorly or tools it misuses.
AutoSynthData turns these capability gaps into training data. It uses the target model's failures and a stronger teacher's successes to decide what the model should learn next. The system generates and validates new tasks that exercise those capabilities.
An agentic environment defines the world in which an agent operates. A task consists of a system specification, a user prompt, and a verifier. The system specification defines constraints like instructions and policies.
The user prompt specifies what the user wants accomplished. Generated tasks must be feasible, realistic, and difficult enough to expose weaknesses. Feasibility ensures at least one trajectory satisfies the prompt while respecting specifications.
Realism requires the prompt to resemble something a user would plausibly ask. Difficulty ensures the task is not yet consistently solved by the current agent. The verifier determines whether the trajectory successfully completes the task.
AutoSynthData evaluates the target model using diagnostic tasks. It identifies patterns in tasks the model struggles to complete. A stronger teacher helps characterize which tasks are solvable and what successful behavior looks like.
The system distills findings into sanitized capability specification cards. The generator receives these cards but not original prompts or trajectories. It creates new tasks with different prompts, states, and solution paths.
Why it matters
Organizations can improve agent quality by generating relevant training data from real failures.
Source: https://huggingface.co/blog/ServiceNow-AI/autosynthdata