r/LocalLLaMA
· Communities
Building a zero-dependency C inference engine for BitNet (1.58-bit) – lessons from hitting 36 tok/s on a Xeon CPU
Over the past few months I have been building a CPU-first inference engine from scratch in pure C99 (no Python, no CUDA, no BLAS, just GCC and make). The focus has been running 1.58-bit ternary models natively without heavy runtime overhead. Currently it hits 36.25 tok/s on BitNet b1.58-2B-4T on an Intel Xeon using 4 t