Anthropic Adds Claude API Tools for Automated Evaluations and Iterative App Improvement
Summary
Anthropic adds build-eval and hillclimb commands to Claude Code’s API skill, enabling automated evaluation creation and iterative app optimization with held-out tests; on a 44-ticket support benchmark, the workflow reaches 98.9% accuracy at about 1 cent per ticket.
Key Points
- Anthropic adds /claude-api build-eval and /claude-api hillclimb commands to its Claude API skill, letting Claude Code create evaluations and iteratively improve applications while using held-out tests to detect overfitting.
- The build-eval workflow samples production transcripts, bug reports, hand-written cases and codebase-derived synthetic inputs, then validates graders by checking repeat verdicts, infrastructure errors and whether baseline performance exceeds roughly 95%.
- On a 44-ticket support benchmark, hillclimbing moves from Opus 4.8 at 74.4% accuracy and 4.6 cents per ticket to Sonnet 5 at 98.9% and about 1 cent per ticket; held-out accuracy rises from 78.6% to 90.5%.