DeepSeek V4 Flash with DSpark via SGLang
Hello guys. Sharing my experience with deploying DS-4-Falsh with DSpark on HGX-H200 For context of my setups etc (including how the hell i have access to…
Hello guys. Sharing my experience with deploying DS-4-Falsh with DSpark on HGX-H200 For context of my setups etc (including how the hell i have access to…
submitted by /u/paf1138 [link] [comments]
https://www.theinformation.com/briefings/exclusive-chinas-minimax-plans-launch-2-7-trillion-parameter-model According to The Information, MiniMax plans to launch a new-generation large language model with 2.7 trillion parameters. Sources revealed that the internal codename for this…
I noticed while using both, mimo was often better, after benchmarking mimo v2.5 via open code endpoint in diff harness like codex, oh my pi, hermes.…
Hello all, it has been 3 months since I made the initial post, where I wanted ideas to try out. The subreddit has been amazing with…
Just finished reading the paper: LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load I am starting to benchmark LLMs…
Me: you have that tool, use it. Opencode: no I don't see that tool in the context, I'll do it the retarded way. Me: *opening an…
Our new open-source AI model for Ancient Egyptian hieroglyph translation. Available in two versions: Horus Hiero 9B Horus Hiero Mini 4B (optimized for CPUs and mobile…
Hey guys, Despite browsing this sub, and the myriad of other LLM reddits, I feel no closer to understanding what the best memory system is these…
Warning: I am an accountant and not an ML engineer of any kind, and I'm potentially missing some important points. I wrote all this by hand,…