Skip to content
r/LocalLLaMA · Communities

Tested Nemotron 3.5 Lightning locally on coding, Hermes Agent and agentic work

Ran the model with quants (Q5) and MTP by bartowski with llama.cpp server. It takes ~24GB ram running on M5 Pro with 48GB at about 65t/s. On some tasks it was quite the overthinker. Overall, the quality of the code output was way below what you can expect for the size (but this is somewhat disclosed by the authors and