Hy3 1Bit 89-93 GB
AngelSlim/Hy3-GGUF at hugging face (I am not affiliated just testing) I dont really make posts but I wanted to make sure everyone is aware that there…
AngelSlim/Hy3-GGUF at hugging face (I am not affiliated just testing) I dont really make posts but I wanted to make sure everyone is aware that there…
https://preview.redd.it/nrxkobib5hdh1.png?width=1147&format=png&auto=webp&s=a36a282635bcbb23c0c2bdffa855689eab5e76f9 I'm running Unsloth's Qwen3.6:35B UD Q4_K_M with a 100k context on a single NVIDIA P40 (24GB) using TheTom's TurboQuant fork of llama.cpp. I know this…
I've got 64GB of regular RAM, and I ran Qwen3 Next 80B at UD-Q4_K_XL, enjoying both nice performance and a lot better internal knowledge than Qwen3.5…
Two DGX Spark and a Connect-X7 cable give you about 250GB of usable memory for $7000 8000 USD. This allows using some interesting models at 4-bit.…
I've been using DS4 flash as a local "big brain" for when my qwen 3.6 35b a3b gets stuck on something, and for planning. I'm blown…
submitted by /u/FreemanDave [link] [comments]
Inkling by Thinking Machines Lab is a huge step forward for US open weight models to catchup w/ China. Inkling solidly beats all US open models…
submitted by /u/AloneCoffee4538 [link] [comments]
Nvm ignore the image links here is the source: https://x.com/googlegemma/status/2077449152062247219 https://huggingface.co/spaces/google/gemma4_vision_token_budget submitted by /u/Iwaku_Real [link] [comments]
https://preview.redd.it/qd7drw8mqfdh1.png?width=2223&format=png&auto=webp&s=4df7e6f4002cf6c6508463f40ebdb6a39c6d1805 https://preview.redd.it/86s7i9dpqfdh1.png?width=2411&format=png&auto=webp&s=3f69ccc368fe0212fb4392f14b0d60e38261e186 So My idea is to instead of spending bigger models compute on easier tasks, why we jus