Would extremely high decode tok/s even be useful?
If you were able to get an inference machine that could do decode at 1k toks/s or even 10k tok/s, would that even be helpful? Would…
If you were able to get an inference machine that could do decode at 1k toks/s or even 10k tok/s, would that even be helpful? Would…
276B total parameters, 12B active, 1M context window. Blog post: https://thinkingmachines.ai/news/inkling-small/ NVFP4: https://huggingface.co/thinkingmachines/Inkling-Small-NVFP4 GGUF's by Unsloth: https://huggingface.co/unsloth/Inkling-Small-GGUF submitted by /u/rerri [link] [comments]
Hey fellow llamas, sorry for posting again this week but i thought this was interesting to showcase to share with y'all. Lucebox partnered up with AMD…
I'm guessing 3 years, what do you think? In other words: many of us will be able to afford a general purpose robot in 3 years…
submitted by /u/niacolhealth [link] [comments]
It was developed under Phase 2 of Korea's Sovereign AI Foundation Model Project. Size: 750B parameters (3x larger than their 236B v1 model). - License: Apache…
I've tested Nanbeige-4.2-3B. On paper, the benchmarks promise it blows away Qwen3.5-9B and Gemma4-12B. My goal was to have something very light and fast to replace…
I'm doing something a bit silly and trying to get hermes agent fully local on minimal resources. I've got a lenovop520 64gb quad channel ram and…
Would you use something like this? One pattern I've seen while building agentic apps is that agents sometimes get stuck in tool loops. Example: search →…
I ran into a doubt when using the following command. It seems that the System RAM usage keeps increasing even though there is >10GB of space…