r/LocalLLaMA
· Communities
Making small local models actually useful for coding
Hey! Like a lot of people here with consumer GPUs (RTX 4060 8GB in my case), I wanted to see if I could use local models for daily coding tasks instead of burning cloud credits on simple boilerplate. The issue with running coding agents (Hermes for example) on 8GB VRAM is: Context bloat: as soon as you feed a decent ch