Taste-Bench Launches 502-Question Test of AI Agents’ Long-Horizon Decision-Making
Taste-Bench launches a 502-question benchmark for long-horizon AI-agent decisions, with GPT-5.6 Sol leading at 59.7% under a paired-order test that requires correct choices in both option orders.