RTX 5090 96GB spotted on Alibaba?
submitted by /u/panchovix [link] [comments]
submitted by /u/panchovix [link] [comments]
Intel's Optane technology as we know it has been discontinued. But while that's the case, I see it being highly promising still due to its role…
Pasted the same HTML/JS code (330 lines) into Qwen 35B A3B and Gemma 26B A4B. Qwen: tokenized the input to 1609 tokens Gemma: tokenized the input…
Firstly a big thanks to the poster hellohazine, he basically only removed the multi-lingual fat of the model and just kept the english language intact. It…
Hey everyone, I could use some advice on setting up speculative decoding correctly with llama-server. My Hardware: GPUs: RTX 4090 + RTX 6000 Pro (120GB total…
Phi was one of my favorite models with a bit of a mixed reputation with some claiming it's benchmaxxed and others seeing its potential and usecases.…
Disclaimer - no LLM was used to write this post/note As larger post about my setup will come later, want to give heads-up to folks who…
Curious what people think are the ideal 4-bit quantization types on MLX These quants seem to be the most popular, at least for Gemma4 and Qwen3.6:…
I have various quants of this model and am curious how they perform. can anyone recommend which benchmark would be a good test case for quantization…
I’m planning a dedicated home AI server, mainly for local LLM inference, agents/tool use, Docker services, and eventually larger MoE models with CPU offload. My plan…