Qwen Developers’ responses from their recent Twitter/X AMA
Questions & Responses(in BOLD) below. Favorite question(s) moved to end of the thread with combined responses(removed duplicates). Be optimistic folks. I'm sure we're getting other models…
Questions & Responses(in BOLD) below. Favorite question(s) moved to end of the thread with combined responses(removed duplicates). Be optimistic folks. I'm sure we're getting other models…
It's the purple cluster on the top left (the good corner...) I'm running the MXFP4 version from Bartoswski with Dspark at 1K t/s prefill and 90…
There are a lot of LLM benchmarks but few, if any, harness benchmarks. I am thinking this would be a really good community project to build…
I know many labs trying to shrink deployment costs and increase efficiency. While I do think that that is fine and dandy, I do sometimes question…
Hey everyone, I’ve been building Speechfony - a desktop app for reading PDFs (and EPUBs) with offline text-to-speech. Open a document, listen sentence-by-sentence with highlighting, or…
People may remember the Qwen3-TTS llama.cpp demo from a few months ago. That PR said it probably wouldn’t be merged because llama.cpp was missing some of…
A Qwen3.5-35B derived model with an interesting architectural difference that results in larger throughput and less token consumption (allegedly): https://huggingface.co/internlm/Intern-S2-Mobius submitted by /u/Miserable-Dare5090 [link] [comments]
submitted by /u/Megneous [link] [comments]
submitted by /u/fallingdowndizzyvr [link] [comments]
I speed up the generation part of the demo in case you get bored 😄 I also tested another long-form generation, and the VRAM usage looks…