Skip to content
r/LocalLLaMA · Communities

Stolen LLM Reasoning: How come OpenAI, Anthrophic, Google have the same vulnerabilities?

If you haven't checked the paper: https://arxiv.org/abs/2608.09867 TLDR: the authors show that you can swap out the "encrypted" reasoning of the biggest model, like Opus, Sol, and put them into weaker model with less guardrail, like Haiku, and ask it to repeat verbatim the reasoning thought. The main reason why this wo