DeepSeek V4 Flash 0731 at 10–17 t/s (nothink) on MacBook M5 Pro **64GB***, partly via SSD streaming
Inspired by a post from u/giveen I motivated claude (no patinence on my side to work through everything myself) to help me get DS running on…
Inspired by a post from u/giveen I motivated claude (no patinence on my side to work through everything myself) to help me get DS running on…
It feels kinda redundant almost given that the MacBook is already mobile, but using it as a server would allow me to leave it always on…
Anyone recommend a source to get started with video/image gen? I pretty much only have experience with llms and right now I just tell the llm…
Most agent memory setups run a model call on the way in. Something reads the turn, decides whether it's worth keeping, rewrites it into a "memory",…
In tests: ~80 tok/s decoding 2,500–3,500 tok/s long-input prefilling Smooth use by 3–4 concurrent users Private, on-device inference for coding, agents, and offline batch jobs submitted…
submitted by /u/MoneyPowerNexis [link] [comments]
As you all know the model is 2.69B parameters with a 128K context window and purpose-built for multi-step agent workflows. What you are seeing is the…
submitted by /u/Afraid-Yoghurt6731 [link] [comments]
*I felt the need to write this post because it seems like very few people on this sub are aware of Chinese laws and how they're…
A new 460M vision model called VisionPsy-Nano-460M-Flash is taking a slightly different approach to on-device VLM speed. Instead of making the language model much smaller, it…