r/LocalLLaMA
· Communities
Some testing on RTX Pro 4500 (With Oculink) on PrismaQuant, INT4 Autoround and NVFP4 W4A4 quantized model
This little beast has been around for a while after asking about whether it's possible to set up in this subreddit. Beelink SER 8 8745 HS, AooStar eg01, and RTX Pro 4500 32GB The Sakamakismile model I've been running was throwing tool call errors and getting stuck in thinking loops in both Opencode and Cline across vLL