Fig launches RIDGE, finding frontier AI models have jagged task performance with no clear overall leader
Summary
Fig launches RIDGE, finding no clear frontier-AI champion as Astra, Opus 5.5 and rivals show jagged task-level results; Astra leads VisualWebArena at 139 of 177 tasks, yet every lower-scoring model solves some tasks it misses.
Key Points
- Fig releases RIDGE, reporting that Astra, Opus 5.5 and other frontier models show jagged task-level performance across web automation, robotics, industrial procedures, assembly and driving, with no clear overall leader.
- On VisualWebArena, Astra leads with 139 of 177 tasks (78.5%), but every lower-scoring model solves tasks Astra misses; all seven models collectively solve 156 tasks.
- A near-flat Opus 4.7-to-Opus 5 average change of -1.1 percentage points masks 36 task outcome flips, with 17 gains and 19 losses.