Running Qwen3.5-122B on Mac Studio 96GB: Fixed 3 bugs that made long-context inference usable
Hey everyone, I recently switched from DS4 Flash to Qwen3.5-122B on my M3 Ultra Mac Studio for long-context agentic coding. While the model fit better, I…
Hey everyone, I recently switched from DS4 Flash to Qwen3.5-122B on my M3 Ultra Mac Studio for long-context agentic coding. While the model fit better, I…
Am I missing something? It seems like some people think distillation is magic and will raise the quality of output above what the base model is…
Hey friends, I started a project i think we can all benefit from. I rented a few h200's and am fine tuning a hui Qwen 35b…
Hey guys. Just a curious and maybe been feeling a little burn out on this whole experiencing local AI and stuff. I own a rig with…
b9978 Claude in one sentence what does this fix llama.cpp b9978 fixes a checkpoint bug that hit agentic workloads hardest: every agent turn created a new…
submitted by /u/fallingdowndizzyvr [link] [comments]
I Recently did a bunch of tests and wrote them all up on here, but the short version is that Vulkan basically doesn't work (or when…
Moondream 3.1 is a vision language model with a mixture-of-experts architecture (9B total parameters, 2B active). It delivers state-of-the-art visual reasoning and detection while staying fast…
For whom is this tread: Everyone with a 24GB GPU (rtx 3090, 7900xtx, rtx 4090) What this Thread is for: Sharing proven/well working llama-server start configs.…
I was surprised to see that Codex is actually open source, so to my understanding, if you use a local model, it works fully locally. How…