Longcat 2 model weights have been published
https://huggingface.co/meituan-longcat/LongCat-2.0-INT8 https://huggingface.co/meituan-longcat/LongCat-2.0-FP8 submitted by /u/RhubarbSimilar1683 [link] [comments]
https://huggingface.co/meituan-longcat/LongCat-2.0-INT8 https://huggingface.co/meituan-longcat/LongCat-2.0-FP8 submitted by /u/RhubarbSimilar1683 [link] [comments]
found this here https://huggingface.co/ideogram-ai/ideogram-4-fp8/discussions/3#6a2070c21bfe5300af2887b2 submitted by /u/po_stulate [link] [comments]
I plan to download a few of them in full 16-bit safetensors (so I'd be able to turn those into whatever quant size/quant-style I want later…
Seeing a lot of "what hardware for local RAG" threads lately, and the framing that keeps getting missed is: decode tok/s is not the bottleneck for…
https://huggingface.co/bartowski/DeepSeek-V4-Flash-GGUF This model is in MXFP4 and as such has only been provided in MXFP4 format! Original model: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash But on original's page none of tensors…
I have an intel n100 mini pc that's on 24/7 running proxmox. I want to use llama.cpp server with gemma 4 E2B for small tasks. Should…
Just a datapoint I wanted to share.Qwen 27b, at q6kxl, with multi-token prediction, on a 4090+3090 system, using lcpp, puts out 50-90 tokens/s decode and 1500-2200…
MongoDB has vector search and hybrid capabilities in the community (free, not cloud) build. Has anybody tried it in the real world and seen how it…
The fact that this feels less like a joke and more like a roadmap is the funniest part. Edit: Sorry, it seems the comment function is…
I didnt see any mention here. Source: https://portugal.gov.pt/en/gc25/communication/news/llm-amalia-shows-portugals-potential HF link SFT: https://huggingface.co/amalia-llm/AMALIA-9B-0626-SFT HF link DFO (Direct Preference Optimization): amalia-llm/AMALIA-9B-0626-DPO · Hugging Face Paper: https://arxiv.org/pdf/2603.265