Dear poor people of this subreddit
I see people with multi-gpu setups but I'm sure there's a potato LLM runner out there somewhere. I have an old macbook pro (i5 8th gen,…
I see people with multi-gpu setups but I'm sure there's a potato LLM runner out there somewhere. I have an old macbook pro (i5 8th gen,…
Hey everyone, We just released our first release candidate from Spectral Labs: a Qwen3.5 0.8B Q4_K_M built using a new calibration-aware quantization approach we're calling SpectralQuant.…
"Hi all, we are finalized with our testing and are preparing the release pipeline. We will be releasing support for the Qwen3.5, Qwen3.6, and Gemma4 very…
We're visiting Shenzhen right now, and visited the Huaqiangbei electronics market. I've seen reports of 96 gig 5090s popping up on AliExpress, but never saw confirmation…
Hello guys, it seems like DeepSeek added a new vision mode to their application. Does this mean, that they will release a new vision model? Edit:…
I’m an owner of single dgx spark with 128 gb unified memory. and I’m hosting through all my local network my ppm over lmstudio. I’m mainly…
It's pretty popular to finetune qwen models but I never hear anyone say anything positive about them. submitted by /u/MrMrsPotts [link] [comments]
https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-DSpark https://github.com/deepseek-ai/DeepSpec/blob/main/DSpark_paper.pdf submitted by /u/External_Mood4719 [link] [comments]
fine-tuned LiquidAI’s LFM2.5-230M on Fable-5 traces and shipped it as GGUF tiny 230M coding-agent model. trained at 4096 ctx. exported Q4_K_M / Q8_0 / F16. runs…
sched : reintroduce less synchronizations during split compute (#20793) CUDA: Improve performance via less synchronizations between token (#17795) Adds CPU-to-CUDA copy capability to ggml_backend_cuda_cpy_tensor_async() Adds function…