Why Speculative Decoding went mature in 2026?
Spec-dec has been a thing for a while, in fact, it's wasn't an idea that was born for LLM inference. E.g. Uber's https://github.com/uber/submitqueue applied it to…
Spec-dec has been a thing for a while, in fact, it's wasn't an idea that was born for LLM inference. E.g. Uber's https://github.com/uber/submitqueue applied it to…
MiniMax H3 is an open-weight, general-purpose multimodal video generation model that works across text, images, video, and audio. In ComfyUI, you can use H3 for text-to-video,…
I also have 64 gb ddr4 ryzen 5600 Using llama.cpp Ubuntu distro Settings are as follows --n-gpu-layers 999 --n-cpu-moe 37 --no-mmap -ctk q8_0 -ctv q8_0 -fa…
VLX-Seek-1.5-10B VLX-Seek-1.5-10B is the open-source 10B model in the VLX-Seek 1.5 family, designed for fine-grained perception and visual grounding in embodied scenarios. It targets practical settings…
Aiming for RTX 6000 like performance at 25% of the cost. https://preview.redd.it/mi6fqpdd5hih1.png?width=631&format=png&auto=webp&s=9a639bcc79c7eb0a3834a88317220c289b7a52b6 2x3090s in tensor parallelism, then put those two in a pipeline feeding into another…
Sub is drowning in slop posts. Shortly after the new rule it was better. But it's gotten unbearable in the past month or so. submitted by…
I'm a retired platform engineer, coding mainly in Rust, and involved with a ML study group. I developed a ML programming language (alternative to Python, Colab)…
I'm a retired platform engineer, coding mainly in Rust, and involved with a ML study group. I developed a ML programming language (alternative to Python, Colab)…
TL;DR: I spent another 8 days following my last post making major improvements to the WinterMix method for MLX models. At 20k+ context this 59 GiB…
submitted by /u/etherd0t [link] [comments]