Skip to content
r/LocalLLaMA · Communities

White-hat hacking IS the defense to black-hat hacking. The techniques are the same. How does Dario expect companies to do it if their models refuse?

You patch security holes by intentionally finding them. If the models refuse to do it, how can companies protect themselves against rogue AIs, whether they are Chinese or OpenAI/Anthropic themselves? The Hugging Face attack showed the world that any AI can do the unexpected. The only true safety we have is defense with