Tensor Split Fix for intel GPU’s llama.cpp release b9788
sycl : support --split-mode tensor #24152 I'd like to see some numbers if anyone has 2xintel gpus and tries this out submitted by /u/Bulky-Priority6824 [link] [comments]
sycl : support --split-mode tensor #24152 I'd like to see some numbers if anyone has 2xintel gpus and tries this out submitted by /u/Bulky-Priority6824 [link] [comments]
Including 9B Dense, 31B Dense, 35B MoE, and 397B MoE and reporting sota on different benchmark (let's see if this holds). https://huggingface.co/collections/deepreinforce-ai/ornith-10 submitted by /u/paf1138 [link]…
While working on an self-educational exercise tinkering with local models and trying my hand at setting up agents, I went down a rabbit hole: to see…
https://preview.redd.it/ilg8oj3uvf9h1.png?width=4096&format=png&auto=webp&s=a8acfcea75b12f165c9b5fcf12606156735c7946 https://openai.com/index/openai-broadcom-jalapeno-inference-chip/ submitted by /u/thecowmilk_ [link] [comments]
https://preview.redd.it/00o5xtaznf9h1.png?width=696&format=png&auto=webp&s=60a3306ea86a9b0d1f58c435b7dbb0a42761a415 Apple raised the prices across the product line this morning: https://www.reuters.com/world/asia-pacific/apple-raises-prices-macbooks-ipads-memory-costs-skyrocket-2026-06-25/ Beyond the base price, the co
Hi everyone, I just released Gemma-4-12B-Uncensored-Opus4.7-CoT. To remove the safety filters without destroying the model's reasoning, I combined a precise ablation method with a CoT (Chain-of-Thought)…
tested my model on kebab bench and it performs very well: https://huggingface.co/spaces/AlexWortega/hermes-agent-zerogpu submitted by /u/Mysterious_Hearing14 [link] [comments]
Are all modern LLMs tuned for chat? Are there any that do bare text completion? I honestly couldn't find any on hugging face. submitted by /u/MackThax…
I read it with a little bit of effort The tiny model result is insane, theoretically this could make make a 0.5b on-par with a 2/3/4b…
NVIDIA has released Nemotron-TwoTower-30B-A3B-Base-BF16, an unusual diffusion-based language model built from the Nemotron 3 Nano 30B-A3B backbone. Instead of generating strictly one token at a time,…