DFlash makes Qwen3.6 27B 2.2x faster with no quality loss
We ran the same Qwen3.6-27B locally three ways on one RTX 6000: baseline, MTP, DFlash. The tasks were: quicksort, write a Steam library in JSON, solve…
We ran the same Qwen3.6-27B locally three ways on one RTX 6000: baseline, MTP, DFlash. The tasks were: quicksort, write a Steam library in JSON, solve…
Everyone's probably seen the remotion thing that went viral a couple months back with CC. Its basically that with Kimi K3 as the model provider. Prev.…
submitted by /u/Charuru [link] [comments]
submitted by /u/WhyLifeIs4 [link] [comments]
Luciole-23B-Instruct-1.1 is a fine-tuned and aligned version of Luciole-23B-Base, an open-source, multilingual causal language model created by OpenLLM-France. Luciole-23B-Instruct-1.1 was developed by LINAGORA and the OpenLLM-Franceconsortium…
I was hopping that Kimi would be 2t, but nope is huge @ 2.8t!! (tears falling) That will make it more difficult to run decently. I…
Hey r/LocalLLaMA! Apparently, if you draw enough arrows between proxy rankings like KLD, perplexity, and BPW, and real deployment measurements, quantization evaluation starts to look like…
https://platform.kimi.ai/docs/guide/kimi-k3-quickstart submitted by /u/WhyLifeIs4 [link] [comments]
Hey, it's my first time posting here and I thought I'd share my progress on getting Antirez's imatrix Q2 DeepSeek V4 Flash GGUF (86.7 GB) running…
Someone had shared similar tools. Here is one in C++. I have been thinking of this "ocean" graph for a long time. Yes, it is useless,…