- Updated: August 24, 2026
- 1 min read
SteeringSafety: Benchmarking Representation Steering in LLMs Across Safety Perspectives
SteeringSafety: Benchmarking Representation Steering in LLMs Across Safety Perspectives
Abstract: We introduce SteeringSafety, a comprehensive benchmark for evaluating representation steering methods across nine safety perspectives spanning 18 datasets. While prior work highlights the general capabilities of representation steering, we focus on safety perspectives including refusal, bias, hallucination, social behaviors, reasoning, epistemic integrity, and normative judgment. SteeringSafety provides modularized building blocks for state‑of‑the‑art steering methods, enabling unified implementation of DIM, ACE, CAA, PCA, and LAT with recent enhancements such as conditional steering.
Key Findings
- Strong steering performance depends on the pairing of method, model, and specific perspective.
- DIM is consistently effective, yet all methods exhibit substantial entanglement: improving effectiveness on one safety perspective often significantly changes performance on others.
- Social behaviors are most vulnerable (degradation up to 76%).
- Refusal steering (jailbreaking) frequently compromises normative judgment such as commonsense morality (up to 26%).
- Hallucination steering shifts political views unpredictably across models, ranging from a 25% shift to the right to a 28% shift to the left.
These findings demonstrate the need to understand steering methods through multiple safety angles rather than a single target behavior.
Read the full paper on arXiv and explore related resources on ubos.tech.

Andrii Bidochko
CTO UBOS
Andrii Bidochko is an AI entrepreneur and researcher focused on AI agents, reinforcement learning, and autonomous systems. He writes about the technologies shaping the future of machine intelligence, from frontier models and agent architectures to real-world AI applications.