X · @teortaxesTex
· X / Twitter
RT Shannon Sands: We already know Anthropic wanted to do that (and probably have been, they have to have live tested their classifiers and steering ve…
RT Shannon SandsWe already know Anthropic wanted to do that (and probably have been, they have to have live tested their classifiers and steering vectors etc). Why should we "trust" closed models if that's the concern? We can detect adversarial training and ablate it. I'm already in the habit of running Qwen through an