Turns out open AI is a coalition, not a company.
submitted by /u/InternationalGap3698 [link] [comments]
submitted by /u/InternationalGap3698 [link] [comments]
I started tracking my local model usage about four weeks ago and was wondering if anyone else here keeps track of how much they use them.…
What are some medium sized MoE models (up to 60B parameters in float8/110B in mxfp4) that are currently worth using? As far as I am aware…
submitted by /u/jah242 [link] [comments]
submitted by /u/MysteryWra [link] [comments]
Hello, I've been learning how to use local LLMs for a year or so on my workstation, using a RTX3090. Current setup : - i5 12400…
RTX 4070, 32gb system ram, Linux. NVIDIA-SMI 610.43.03, KMD Version: 610.43.03, CUDA UMD Version: 13.3 Systemd service modifications: [Service] Environment="OLLAMA_HOST=0.0.0.0:11434" Environment="OLLAMA_FLASH_ATTENTION=true" Environment="OLLAMA_KV_CACHE_TYPE=q8_0" Environment="OLLAMA_NUM_PARALLEL=1" Environment="OLLAM
Hallo Everyone, i need to generate big amount of high quality data, for that i need some cheap API providers. who is the cheapest,most reliable provider…
Wanted to consolidate where things actually stand now that the dust has settled on speculative decoding in llama.cpp, since the discourse a few months back was…
If you run local agentic coding harnesses (Aider, Claude Code, etc.), prompt evaluation usually eats up most of your execution time. Every turn re-evaluates thousands of…