llama.cpp b9966 for sm-tensor
B9966 If you run -sm tensor in production you might want to grab this fix which removes 29 regex recompilations per tensor per token on the…
B9966 If you run -sm tensor in production you might want to grab this fix which removes 29 regex recompilations per tensor per token on the…
Saw this on a random social media post. I believe what we are seeing here is two sxm2 gpus, right? V100 or a100? Assuming two v100…
I occasionally comment and people ask how do I run something, how fast the model x on hardware y. There you go. Added a separate website…
Here is the upper limit of what can be done with $100 bucks worth of video cards. You can have 3 concurrent users with plenty of…
Hello guys, hoping you're doing fine! I'm continuing after this post some time ago, comparing stock MaxQ performance and such on Anima here. This time, I…
WEIRD DISCLAIMER: none of this was written by an LLM until you get to the Github repo/site, which was obviously assembled by your friend and mine,…
Which settings would suffice to work with it ? submitted by /u/soyalemujica [link] [comments]
hello real people and less-real bots, i'd appreciate if any of you people who have fine-tuned (either full or peft) more than half a model could…
I know about DEEPSWE but it lacks many models :( submitted by /u/9r4n4y [link] [comments]
Built around a recent sd.cpp release, aims to expose most of what the backend can do (generate, edit, video paths, models, hardware options), Windows + Linux…