← Back to feed
2026-07-16agentsalignment

When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space

Weimeng Wang, Ziqiang Wang, Zihang Zhan, Chuanpu Fu, Qi Li, Ke Xu

PDF preview for When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space
Read on arXiv →

Key claim

PRISM effectively distinguishes content and physical dangers in LLMs.

In plain English

Large language models can misinterpret benign instructions as dangerous when applied in the physical world. Current models struggle to differentiate between content danger and physical danger, leading to potential safety issues. This paper introduces PRISM, a method that effectively distinguishes these dangers and achieves high accuracy on a new benchmark for assessing physical risks. Builders should be interested in this advancement as it improves the safety of LLMs in practical scenarios.

Novelty
8.0/10

Introduces a new method for distinguishing between content and physical dangers in LLMs.

Reliability
8.0/10

Demonstrates strong performance on multiple benchmarks with rigorous testing.

Deep reliability assessment

The methodology supports the claim that physical jailbreaks form a distinct signal in LLM representations, but the generalizability across different models and tasks may be overclaimed.

Reproducibility

No open source code or dataset is mentioned in the paper.

Key figure

Figure 1 illustrates the distinction between textual jailbreak and physical jailbreak, emphasizing the need for physical-safety questions in addition to text-safety moderation.

Benchmark results

SafeAgentBenchaccuracy: 87.7vs LLM judgesN/ASOTA