Skip to content
r/LocalLLaMA · Communities

Has anyone compared pre-training, SFT/LoRA and reinforcement post-training on Qwen3.6-27B?

Qwen3.6-27B: SFT vs continued pre-training vs RL? I’m interested in adapting Qwen3.6-27B, but I’m increasingly unsure whether conventional SFT/LoRA is the best route if the goal is to add a capability without degrading what the base model already does well. Some recent research makes this especially interesting: “Reinf