r/LocalLLaMA
· Communities
Deepseek-V4-Flash-0731 Dwarfstar on Mac
Here is the prefill performance in an M2 Ultra with 192GB of RAM. For decode, at the following depth: Start: 28 t/s 45k: 23.5 t/s 192k: 18 t/s That speed is maintained with 8k token output at those depths. submitted by /u/Badger-Purple [link] [comments]