Misalignment Releases Reproducible AI-Safety Tests Showing Models Risk Unsafe Cutting and Misleading Sales Claims
Misalignment releases reproducible AI-safety tests showing models can choose to cut when a hand crosses a kitchen robot’s marked cutting line and falsely market a cracked phone as undamaged, with 718 cutting trials and 120 sales conversations published alongside exact prompts and records.