Can Gemma and Qwen models catch hallucinations by looking at their own logprobs?
Hi! I'm really obsessed with LLM hallucinations for the last 6 days 😭 I started by designing system prompts to attack hallucinations but failed, obviously. Now…
Every story across every category, newest first. Each card links to the original publisher; daily-brief posts open as editorial pages.
Hi! I'm really obsessed with LLM hallucinations for the last 6 days 😭 I started by designing system prompts to attack hallucinations but failed, obviously. Now…
RT khipu.aiBehind a lot of great AI projects, there is a good @JeffDean story. Here is mine 😉Back in 2019, even before Khipu AI officially existed,…
On-policy (Self-)Distillation (OPD / OPSD) has shown strong potential for post-training large language models (LLMs). However, existing methods still rely heavily on external supervision, including ground-truth…
What's Changed Make it easier to debug nested tensors. by @comfyanonymous in #15383 Update workflow templates to v0.11.37 by @comfyui-wiki in #15415 Minimum officially supported pytorch…
Article URL: https://tencent-hunyuan.github.io/Hunyuan3D-WorldClaw/ Comments URL: https://news.ycombinator.com/item?id=49265051 Points: 19 # Comments: 5
The least verifiable part of AI R&D (and the thing which may bottleneck the automated researchers) is making calls on large experiments and big training runs.Interestingly,…
Article URL: https://clickhouse.com/blog/pg_clickhouse-whats-new-july-2026 Comments URL: https://news.ycombinator.com/item?id=49265031 Points: 5 # Comments: 0
The U.S. VC firm still has more than 55% of its previous $650 million India fund available for deployment.
Daybreak Red and Daybreak Blue from OpenAI, specialized cyber defense models from OpenAI, are now available on Amazon Bedrock to eligible customers. Both models run with…
We converted the model from the original safetensors and found two issues. The first one made our quantization fail several times, the second one does not…
Article URL: https://www.suzanne3d.com/ Comments URL: https://news.ycombinator.com/item?id=49264755 Points: 4 # Comments: 0
RT Ryan GreenblattMy median for full automation of AI R&D is around late 2030/early 2031. But my "modal"/best guess prediction for this milestone would be significantly…
Article URL: https://arxiv.org/abs/2601.01828 Comments URL: https://news.ycombinator.com/item?id=49264583 Points: 4 # Comments: 1
Link: https://arxiv.org/pdf/2604.27883 Hi, Most of use are familiar with the headache of training a neural network using gradient descent where the training error may go to…
Pharma is suddenly paying for Bio × AI tools, and Chai is leading the pack with four deals closed this summer. Cofounder Matt McPartlon and Product…
Article URL: https://www.effort.news/p4ne Comments URL: https://news.ycombinator.com/item?id=49264352 Points: 3 # Comments: 1
Ryan believes we could get 3-6 years of AI progress at the current pace, but in a single year, once we automated AI R&D.So a jump…
Release: datasette-upload-dbs 0.5a0 This plugin has been around for a while - it lets users upload a brand new SQLite database to a hosted Datasette instance,…
These are single stream numbers Following on from my previous post about v100s (here) and inspired by this comment (here) I decided to work on kernels…
Cross-posted from my Substack. Basically, I created a text-based adventure game benchmark in April, and this morning my agent harness using Claude Opus 5 solved it…
This update brings persistent memory, local model access, and more enterprise controls to GitHub Copilot for JetBrains. It also improves everyday chat workflows and resolves reliability…
chat : fix muse-glimmer swallowing a trailing tool call into content Muse Glimmer routinely answers the user and calls a tool in a single generation. The…
Needs a lot of requests compared to Qwen (almost twice) and Gemma (almost x3). Final score is fine, even though it is "not a coding model"…
Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to…