Skip to content
LessWrong AI · Communities

Is Mythos good at cyber because it kept hacking Anthropic during training?

From the Mythos preview system card (emphasis mine):We ran an automated review of model behavior during training, sampling several hundred thousand transcripts from across much of the training process. We used recursive-summarization-based tools backed by Claude Opus 4.6 to summarize the resulting transcripts.[...]The