Skip to content
r/LocalLLaMA · Communities

Muse Glimmer 30B running locally in-browser with custom WebGPU kernels at ~25 tok/s on an M4 Max (same speed as llama.cpp)

submitted by /u/xenovatech [link] [comments]