arXiv cs.AI
· Papers
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models
arXiv:2607.02914v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable capabilities across diverse applications, yet ensuring their simultaneous safety, helpfulness, and trustworthiness remains a persistent challenge. Conventional refusal-oriented alignment strategies mitigate harmful