r/LocalLLaMA
· Communities
I built a Triton backend for Falcon3-10B-1.58bit: 97.5 tok/s decode on an RTX 5070
Hi r/LocalLLaMA — I’m sharing an experimental GPU-only inference backend and looking for independent reproductions, not just stars. Model: tiiuae/Falcon3-10B-Instruct-1.58bit GPU: NVIDIA RTX 5070 Batch: 1 Measured after warmup: • Hybrid packed decode: 97.51 tok/s • Stock Transformers BitLinear decode: 9.89 tok/s • Obse