What am I missing?
Training my own micro-llama-model on a dataset I have published with my own program, I somehow fail to get it working in Unsloth Studio and other…
Training my own micro-llama-model on a dataset I have published with my own program, I somehow fail to get it working in Unsloth Studio and other…
I've been benchmarking a two-card box for a few weeks and I still can't quite get over some of these numbers, so I'm dumping them here.…
I've tried grok and Claude's deep research mode and I was amazed with the speed considering the amount of sources analysed. Is there anything as fast…
submitted by /u/Hannibalj2ca [link] [comments]
Hello! Does anyone tried text embedding models on Jetson Nano 2/4Gb? I need it for the RAG. I want to know the speed. For example microsoft/harrier-oss-v1-0.6b…
I bought an RTX 5090 last year just to run 27B models natively. I even fine-tuned it with my own data using LoRA, building RAGs and…
Has anyone yet tried to extract experts from kimi (or GLM 5.2) per chance? There is REAP that removes experts based on routing, but I could…
Hello Can you please share the tempature you get on your rtx 3090 under active llm load? Im trying to findout if my rtx 3090's tempatures…
we're experimenting with our own dynamic GGUF quants of kimi k3, made from the original weights with our llama.cpp fork. Q3_K_S is done and works 1114.76…
I think the claims about having Mythos level-model in our laptops in 1-2 years might not be so crazy of a theory submitted by /u/SilverRegion9394 [link]…