Skip to content
r/MachineLearning · Communities

What if a model could only learn what trusted LoRA adapters can express? [R]

Hello I published a paper. Most defenses against fine-tuning poisoning try to detect malicious data or reduce its impact. I explored a different question: What if the model simply could not learn certain malicious updates? The idea is to constrain fine-tuning to a subspace learned from trusted LoRA adapters. Useful ada