Skip to content
r/LocalLLaMA · Communities

Am I just hallucinating

Or is there any reason why I feel like model output quality seems to be better when I use higher micro-batch values (ub) in llama-cpp? I don't really have any hard numbers or anything (just running the same prompts), it's all just vibes. Some context, I'm running the latest build of llama-cpp, vulkan, 6900xt 16gb, 64gi