r/LocalLLaMA
· Communities
I’ve added Maple-Preview to Mference, got 40 tps generation with 500MB of used RAM on Air M4
I like the idea of running local models, but I don’t like the idea of having them eat up all of my memory. I’ve always thought that the best way to build an edge model would be to make something smart enough to reason over data, but without requiring much knowledge of its own. Why should a model carry all that knowledg