✨ From vibe coding to vibe deployment. UBOS MCP turns ideas into infra with one message.

Learn more
Andrii Bidochko
  • Updated: August 24, 2026
  • 1 min read

SteeringSafety: Benchmarking Representation Steering in LLMs Across Safety Perspectives

SteeringSafety: Benchmarking Representation Steering in LLMs Across Safety Perspectives

Abstract: We introduce SteeringSafety, a comprehensive benchmark for evaluating representation steering methods across nine safety perspectives spanning 18 datasets. While prior work highlights the general capabilities of representation steering, we focus on safety perspectives including refusal, bias, hallucination, social behaviors, reasoning, epistemic integrity, and normative judgment. SteeringSafety provides modularized building blocks for state‑of‑the‑art steering methods, enabling unified implementation of DIM, ACE, CAA, PCA, and LAT with recent enhancements such as conditional steering.

Key Findings

  • Strong steering performance depends on the pairing of method, model, and specific perspective.
  • DIM is consistently effective, yet all methods exhibit substantial entanglement: improving effectiveness on one safety perspective often significantly changes performance on others.
  • Social behaviors are most vulnerable (degradation up to 76%).
  • Refusal steering (jailbreaking) frequently compromises normative judgment such as commonsense morality (up to 26%).
  • Hallucination steering shifts political views unpredictably across models, ranging from a 25% shift to the right to a 28% shift to the left.

These findings demonstrate the need to understand steering methods through multiple safety angles rather than a single target behavior.

Read the full paper on arXiv and explore related resources on ubos.tech.

SteeringSafety benchmark illustration


Andrii Bidochko

CTO UBOS

Andrii Bidochko is an AI entrepreneur and researcher focused on AI agents, reinforcement learning, and autonomous systems. He writes about the technologies shaping the future of machine intelligence, from frontier models and agent architectures to real-world AI applications.

Sign up for our newsletter

Stay up to date with the roadmap progress, announcements and exclusive discounts feel free to sign up with your email.

Sign In

Register

Reset Password

Please enter your username or email address, you will receive a link to create a new password via email.