Anthropic Adds Claude API Tools for Automated Evaluations and Iterative App Improvement
Anthropic adds build-eval and hillclimb commands to Claude Code’s API skill, enabling automated evaluation creation and iterative app optimization with held-out tests; on a 44-ticket support benchmark, the workflow reaches 98.9% accuracy at about 1 cent per ticket.