Skip to content
r/LocalLLaMA · Communities

Muse Glimmer ACTUALLY fits on a single RTX 3090

I did some testing this morning, and I was surprised to find that Muse Glimmer actually comfortably fits on a single RTX 3090 with full context + DFlash + mmproj at Q4_K_XL, unlike Qwen3.6-27B and Gemma-4-31B. Muse Glimmer supports up to 256k context according to Unsloth. Here is my command: llama-server --model Muse