Skip to content
r/LocalLLaMA · Communities

Ornith-397B running at Q4 on a single RTX PRO 6000 Blackwell 96GB – 2,354 tok/s prefill, ~20–24 tok/s decode

I've been building Krasis, an MoE-focused runtime for streaming big models through limited VRAM on NVIDIA consumer/workstation GPUs, and I think this is the most interesting result so far: Ornith-1.0-397B running interactively on one GPU. Hardware: 1× RTX PRO 6000 Blackwell 96GB + AMD EPYC 7742 (64c although the CPU is