Google researchers have introduced ToolGrad, a novel framework that reimagines how artificial intelligence systems learn to use tools efficiently. Rather than the traditional approach of generating user queries first and then searching for corresponding tool-use solutions—a process prone to failures and inefficiency—ToolGrad inverts the workflow by generating successful tool-use chains before creating the associated user prompts. This answer-first methodology significantly reduces the computational cost and time required to generate training datasets at scale.
The framework leverages "textual gradients," feedback provided by language models in plain text, to iteratively construct complex API workflows from large tool libraries. ToolGrad employs four sequential modules: an API Proposer that identifies promising candidates, API Executors that test selections in parallel, an API Selector that chooses the best performing call, and an LLM Updater that revises the synthetic user query. In experiments using real-world APIs from ToolBench's 16,000+ API database, ToolGrad demonstrated higher pass rates and lower generation costs compared to traditional approaches.
When fine-tuned models based on ToolGrad-generated datasets were evaluated on the Berkeley Function Calling Leaderboard—a benchmark featuring different tools than those used in training—Gemma-3 models of all sizes showed consistent improvements over base models and competitive performance against proprietary systems from OpenAI, Google, and Anthropic. The results suggest that more efficient dataset generation can translate to more capable and cost-effective AI assistants.
Key Points
ToolGrad inverts the typical dataset generation workflow, creating tool-use solutions before user queries, reducing cost and complexity
Uses textual gradient feedback to iteratively refine API workflows, adapting prompt optimization techniques to synthetic data generation
Fine-tuned Gemma-3 models on ToolGrad data matched or exceeded proprietary LLMs on out-of-distribution tool-use benchmarks