Are there any qwen finetunes that were genuinely stronger than the base?
It's pretty popular to finetune qwen models but I never hear anyone say anything positive about them. submitted by /u/MrMrsPotts [link] [comments]
It's pretty popular to finetune qwen models but I never hear anyone say anything positive about them. submitted by /u/MrMrsPotts [link] [comments]
It is like pytest but for statistical tests: it ensures no regression of your metrics at a statistical level. It manages tedious things such that seeds,…
Article URL: https://www.economist.com/culture/2026/06/25/is-america-becoming-a-gerontocracy Comments URL: https://news.ycombinator.com/item?id=48695694 Points: 29 # Comments: 34
https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-DSpark https://github.com/deepseek-ai/DeepSpec/blob/main/DSpark_paper.pdf submitted by /u/External_Mood4719 [link] [comments]
fine-tuned LiquidAI’s LFM2.5-230M on Fable-5 traces and shipped it as GGUF tiny 230M coding-agent model. trained at 4096 ctx. exported Q4_K_M / Q8_0 / F16. runs…
Article URL: https://github.com/schlae/IBM_MCGA Comments URL: https://news.ycombinator.com/item?id=48695363 Points: 5 # Comments: 1
I submitted one of my NeurIPS review ~6 hrs later than the official deadline. Will this still affect my own submission? Asking because I’m a first…
sched : reintroduce less synchronizations during split compute (#20793) CUDA: Improve performance via less synchronizations between token (#17795) Adds CPU-to-CUDA copy capability to ggml_backend_cuda_cpy_tensor_async() Adds function…
Article URL: https://www.openttd.org/news/2026/06/25/openttd-16-0-beta1 Comments URL: https://news.ycombinator.com/item?id=48695149 Points: 60 # Comments: 5
I am relatively new here, I have little experience in how long support development takes. I know there are forks. But not merged status means AFAIK…
Article URL: https://www.sfwriter.com/wordstar.htm Comments URL: https://news.ycombinator.com/item?id=48694853 Points: 4 # Comments: 0
For those who are using llms on macbook, Want to understand how macbook is different than dedicated GPU in running those models? and how to know…
Article URL: https://grack.com/blog/2026/06/25/dissecting-a-failed-nation-state-attack/ Comments URL: https://news.ycombinator.com/item?id=48694631 Points: 4 # Comments: 0
I quantized deepreinforce-ai/Ornith-1.0-35B down to Q3_K_M so it fits comfortably on a single GPU. Produced locally with llama-quantize from the upstream BF16 GGUF — the quantizer…
Are there any local LLM based speech to text that can challenge Dragon NaturallySpeaking or Dragon Professional? Notably with regards to being able to change/delete words…
I’ve been working on long term memory for LLMs for a while and kept hitting the same issue. Everything works for demos, then falls apart. So…
I dumped the last week deep diving this and I’m I’ve been using Linux for 14 years and am a cloud systems engineer with a focus…
Although the page itself is more just fun to have made and look at (I like the flip sound), the fun part is how I made…
I've been using cloud provided models for agentic theorem proving a lot, and cost is becoming an issue for me. I have funding for hardware cost…
Article URL: https://news.mccombs.utexas.edu/research/foreign-funds-help-make-housing-unaffordable/ Comments URL: https://news.ycombinator.com/item?id=48693420 Points: 13 # Comments: 1
Article URL: https://daringfireball.net/2026/06/om Comments URL: https://news.ycombinator.com/item?id=48693391 Points: 28 # Comments: 2
TL;DREvaluation awareness — an AI recognizing it's being evaluated — is a widely discussed concept in AI safety. But there is a closely related concept that…
Article URL: https://twitter.com/Techmeme/status/2070638481265905837 Comments URL: https://news.ycombinator.com/item?id=48692995 Points: 18 # Comments: 5
Article URL: https://physics.stackexchange.com/questions/535/why-does-kinetic-energy-increase-quadratically-not-linearly-with-speed Comments URL: https://news.ycombinator.com/item?id=48692946 Points: 8 # Comments: 0