Skip to content
r/LocalLLaMA · Communities

543 tok/s single-request Qwen3.6-35B-A3B on one RTX 5090 over a 65K-token decode

An example TL;DR I have open-sourced NInfer, a from-scratch C++/CUDA inference engine currently specialized for two exact Qwen3.6 checkpoints on a single RTX 5090. Both the engine and the converted model artifacts are publicly available: Github: https://github.com/Neroued/ninfer The main result: Qwen3.6-35B-A3B sustain