Using the Bonsai 27b 1b quant locally – regularly.
I've been using the 1bit quant of prismml's bonsai 27b for local conversation, casual chat/ literature review for fun (i throw random stuff from my notes…
I've been using the 1bit quant of prismml's bonsai 27b for local conversation, casual chat/ literature review for fun (i throw random stuff from my notes…
I'm distilling from Gemini Pro 3.1 as the teacher, the task has a mixture of data extraction and analysis. I need to process about 90 million…
mii-llm, an open source AI lab, released Zagreus-0.4B-por, a compact bilingual Portuguese–English language model pretrained entirely from scratch. The model has approximately 400 million parameters and…
The Open Letter was initiated by Microsoft and published today: “Open Weights and American AI Leadership”. It argues against broad or premature restrictions on open-weight models…
Is it a thing? i use kernel-anvil added to llama.cpp, but wondering if theres others out there ? Heres a link i found for vLLM -…
I'm downloading it again now. So far, the model hasn't performed well with reasoning tasks, but I really appreciate the work being done to fix this.…
"If you're walking, just take your car keys with you." Just wanted to try the ternary Bonsai model as it had its hype for some time…
From my understanding, most current LLMs are trained on trillions and trillions of tokens of mostly AI-generated data. Are there any recent models that are trained…
From Anton Lozhkov on 𝕏: https://x.com/anton_lozhkov/status/2080254608639701222 Two ways in: stack-v3-train - near-deduplicated, quality-filtered, PII-redacted, contents inline. Point load_dataset at it and go. https://huggingface.co/datasets/HuggingFaceCode/stack-v3-train stack-v3-full - the…
I think this is an areas which could utilize crowdfunding/ crowdcomputing. Imagine Qwen team comIng here and telling: ”Guys we need 10M$ to train qwen4-27b, if…