LLMs Fail to Self-Report Adversarial Prefills, Study Finds
The study tested ten open-weight LLMs on four safety benchmarks and found that no model reliably identifies its own compromised outputs. This finding challenges prior work on LLM introspection and suggests that self-report mechanisms are insufficient for safety-critical applications.











