Skip to content

Anthropic Adds Claude API Tools for Automated Evaluations and Iterative App Improvement

Sep 29, 2026
claude.dev Blog
Article image for Anthropic Adds Claude API Tools for Automated Evaluations and Iterative App Improvement

Summary

Anthropic adds build-eval and hillclimb commands to Claude Code’s API skill, enabling automated evaluation creation and iterative app optimization with held-out tests; on a 44-ticket support benchmark, the workflow reaches 98.9% accuracy at about 1 cent per ticket.

Key Points

  • Anthropic adds /claude-api build-eval and /claude-api hillclimb commands to its Claude API skill, letting Claude Code create evaluations and iteratively improve applications while using held-out tests to detect overfitting.
  • The build-eval workflow samples production transcripts, bug reports, hand-written cases and codebase-derived synthetic inputs, then validates graders by checking repeat verdicts, infrastructure errors and whether baseline performance exceeds roughly 95%.
  • On a 44-ticket support benchmark, hillclimbing moves from Opus 4.8 at 74.4% accuracy and 4.6 cents per ticket to Sonnet 5 at 98.9% and about 1 cent per ticket; held-out accuracy rises from 78.6% to 90.5%.

Tags

Read Original Article