✨ From vibe coding to vibe deployment. UBOS MCP turns ideas into infra with one message.

Learn more
Andrii Bidochko
  • Updated: July 7, 2026
  • 2 min read

Oyster‑II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models

Oyster‑II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models

Large language models (LLMs) have demonstrated remarkable capabilities across diverse applications, yet ensuring their simultaneous safety, helpfulness, and trustworthiness remains a persistent challenge. Conventional refusal‑oriented alignment strategies mitigate harmful content generation but systematically fail to serve legitimate user needs, often withholding information that could safely and constructively address the underlying intent of sensitive queries.

Building upon the constructive safety paradigm pioneered by Oyster‑I, which moves beyond blanket refusal toward thoughtful, response‑oriented safety alignment, we identify two critical limitations of its Supervised Fine‑Tuning (SFT)‑based scheme: insufficient safety generalization to out‑of‑distribution scenarios and a phenomenon we term safety chain‑of‑thought (CoT) over‑generalization, wherein safety‑oriented reasoning patterns are excessively applied to benign queries, degrading helpfulness and user experience.

To address these limitations, we propose Oyster‑II, a reinforcement learning (RL)‑based constructive safety alignment framework that adopts a Zero‑RL paradigm combined with a multi‑stage reinforcement learning strategy. Evaluated across extensive benchmarks, Oyster‑II comprehensively surpasses both Qwen3‑14B and its predecessor Oyster‑I on safety dimensions, achieving cross‑scale performance comparable to Qwen3‑Max and Qwen3.5‑397B.

Read the full paper on arXiv. For more insights, explore related content on ubos.tech/ai-safety and ubos.tech/research.

Constructive Safety Alignment Illustration


Andrii Bidochko

CTO UBOS

Andrii Bidochko is an AI entrepreneur and researcher focused on AI agents, reinforcement learning, and autonomous systems. He writes about the technologies shaping the future of machine intelligence, from frontier models and agent architectures to real-world AI applications.

Sign up for our newsletter

Stay up to date with the roadmap progress, announcements and exclusive discounts feel free to sign up with your email.

Sign In

Register

Reset Password

Please enter your username or email address, you will receive a link to create a new password via email.