Skip to content

Research

415 articles found

New Benchmark Exposes Hidden 'Flinch' Effect in AI Models That Suppresses Words at Probability Level, Defying Uncensoring Fixes

New Benchmark Exposes Hidden 'Flinch' Effect in AI Models That Suppresses Words at Probability Level, Defying Uncensoring Fixes

Apr 21, 2026
Morgin.ai

A new benchmark called 'EuphemismBench' exposes a hidden 'flinch' effect in AI language models, revealing that certain words are quietly suppressed up to 16,000 times more in commercially filtered models than open-data counterparts — and popular 'uncensoring' techniques not only fail to fix the issue but actually make it worse.

Scientists Crack Open AI's 'Black Box' to Reveal How Neural Networks Think

Scientists Crack Open AI's 'Black Box' to Reveal How Neural Networks Think

Apr 21, 2026
Oz

Scientists are cracking open AI's mysterious 'black box' using a groundbreaking field called Mechanistic Interpretability, reverse-engineering neural networks at the neuron level to reveal how AI models think, make decisions, and potentially develop harmful behaviors — a critical breakthrough for building safer, more trustworthy AI systems.

Research AI Safety
Previous
Page 10 of 42
Next
Showing 91 - 100 of 415 articles