r/LocalLLaMA
· Communities
Would there be a use case for running a 405B on a single 8xA100 node with up to 30 fine tuned specialists loaded hot at sub 200ms switching?
I know people consider llama 405b and others to be old now, lol, but I'm wondering if there would be a use case for it. I had a use case for a project I was building and I wanted to share what I got and get some feedback which would be much appreciated. base model: llama 3.1 405b (awq-int4, 202gb) hardware: single 8xa1