AntLing-3.0-flash is now live on OpenRouter, and free to use through August 3, 2026
https://openrouter.ai/inclusionai/ling-3.0-flash submitted by /u/derspenti [link] [comments]
https://openrouter.ai/inclusionai/ling-3.0-flash submitted by /u/derspenti [link] [comments]
https://preview.redd.it/u8alj38wr0fh1.png?width=1860&format=png&auto=webp&s=3578776c60d9548a135a018702f44d0fddedd4b0 Little side project I'm doing so I can easily transfer any model I want fast to my AI Rig from my NAS. submitted by /u/TyedalWaves…
I wanted to know how cheap you can go and still run local models, so I ran Ollama CPU-only on a Youyeetoo X1S. It's a single-board…
In all my fiddling around with code and local models nothing has matched the speed and quality of DeepSeek V4 Flash DSpark on dual DGX Spark…
At the moment MLX (and Llama.cpp for Macs) run 16bit activations everywhere. Despite this, the M5 generation silicon actually does support INT8 activations - it actually…
Link to their official GGUF repo: https://huggingface.co/poolside/Laguna-S-2.1-GGUF/tree/main All the GGUFs received this fix 5ish hours ago - correct yarn_attn_factor to 1.0 (llama.cpp derives mscale) And the…
Hi everyone, About a month ago I publish my very first research paper on my neural network architecture called Silia. You can look at the model…
I work with financial documents for my job, and I wanted a small model I could run locally to enrich them before they hit a RAG…
I ran Laguna S 2.1 (118B MoE, 71 GB Q4) with DFlash speculative decoding on 2× RTX 5090. It doesn't fit, experts spill to CPU RAM.…
I could be reading it wrong but it looks like they REAPed their own 250B model down to a 32B model. They claim their 250B model…