r/LocalLLaMA
· Communities
DeepSeek V4 Flash, up to 32 tok/s on AMD Ryzen AI MAX+ 395
Hey fellow llamas. we have something new for Strix Halo owners we thought would be useful to share. i'll keep it short: We were able to fit DeepSeek V4 Flash plus its speculative draft on a single Ryzen AI MAX+ 395 with 128 GB of unified memory, and got it to a usable decode rate. Blog post with all details here: https