When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space
Weimeng Wang, Ziqiang Wang, Zihang Zhan, Chuanpu Fu, Qi Li, Ke Xu
Read on arXiv →Key claim
PRISM effectively distinguishes content and physical dangers in LLMs.
In plain English
Large language models can misinterpret benign instructions as dangerous when applied in the physical world. Current models struggle to differentiate between content danger and physical danger, leading to potential safety issues. This paper introduces PRISM, a method that effectively distinguishes these dangers and achieves high accuracy on a new benchmark for assessing physical risks. Builders should be interested in this advancement as it improves the safety of LLMs in practical scenarios.
Introduces a new method for distinguishing between content and physical dangers in LLMs.
Demonstrates strong performance on multiple benchmarks with rigorous testing.
Deep reliability assessment
The methodology supports the claim that physical jailbreaks form a distinct signal in LLM representations, but the generalizability across different models and tasks may be overclaimed.
Reproducibility
No open source code or dataset is mentioned in the paper.
Key figure
Figure 1 illustrates the distinction between textual jailbreak and physical jailbreak, emphasizing the need for physical-safety questions in addition to text-safety moderation.
