Skip to content
r/LocalLLaMA · Communities

LFM 2.5 230M running at 1440 tok/s in-browser through a custom backend

Everything runs through WebGPU, in-browser or in electron/tauri apps. It's fully portable and supports either Nvidia and Apple Silicon (Metal). The actual kernels are optimized for the specific hardware of the device. The Nvidia kernels are aggressively fused into a multi-pass architecture, while the Apple Silicon kern