Toward A Public Science of Model Behavior
This is a linkpost for our essay "Toward A Public Science of Model Behavior" on transluce.org. The full text is reproduced below.Today’s AI systems frequently behave…
This is a linkpost for our essay "Toward A Public Science of Model Behavior" on transluce.org. The full text is reproduced below.Today’s AI systems frequently behave…
GPT-5.6 Terra and Sol are now available in Perplexity and Perplexity Computer.
RT Tesla Owners Silicon ValleyGROK 4.5 LEADS ON REAL PROFESSIONAL WORK BENCHMARKNew data from Snorkel shows Grok 4.5 outperforming other frontier models on real-world professional tasks.On…
Article URL: https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2026.1766665/full Comments URL: https://news.ycombinator.com/item?id=48854247 Points: 23 # Comments: 1
I built this because I see that grocery savings are achievable in NYC. People usually just go to the store they're used to going to, and…
OpenAI's new family of models will continue to power Microsoft's suite of workplace and productivity apps.
I completed my law degree at a working-class London university. In my first year, I was 18 years old, and I was often the youngest person…
Grok Build improves almost every dayJason Ginsberg: some special features in Grok Build if you're new/dashboard: shows you every agent running in your TUI. no need…
Grok gamingtetsuo: Grok 4.5 in Grok Build created an FPS game in under an hour.The prompt was simple. I told it to write a game design…
Microsoft may once again be struggling to keep up with its own climate goals, according to its 2026 sustainability report. As reported by GeekWire, the report…
Grok doesn’t give upComposio: Grok 4.5 is the most persistent agent model we've tested. Here's one example: In one of our evals, we asked 3 models…
At the end of May, I attended the SciFM26 conference hosted at UChicago which had many speakers discussing agentic science and the future of AI+Science. This…
Learn how to use ChatGPT, start your first conversation, and discover simple ways to write, brainstorm, and solve problems with AI.
This paper was accepted at the AI4TCI (Workshop on AI for Secure and Trustworthy Critical Infrastructure Systems) Workshop at the International Conference on Availability, Reliability and…
Working at the frontier: How Cognition trusts Claude Fable 5 to work through the night
I like choices... but now I have:2x modes (Codex vs. Work mode)3x GPT-5.6 models (Sol, Terra, Luna)5x effort levels (Light, Medium, High, Extra High, Ultra)That's 2…
Grok 4.5Kevin Bass: Grok 4.5 scores significantly better than other frontier models Opus 4.8 and GPT 5.5 across ALL professional work benchmarks.It performs exceptionally well at…
I've been running some systematic tests on a few models comparing FP16 vs various GGUF quant levels, and instead of looking at one aggregate benchmark score,…
RT steve hsuResearch produced by LLMs tends to be combinatorial - a synthesis of ideas that already exist in the literature, rather than something truly original.…
OpenAI's No. 2 executive, Fidji Simo, is stepping down from her full-time role after her medical leave proved longer than expected — a leadership vacuum that…
easy to overlook the implications for how Sol can accelerate your engineering workflowsAndrew Curran: Quote from OpenAI on the livestream.'Already, Sol has been transforming our research…
RT Jeff BezosWally Funk waited 60 years to get to space, and no one ever earned it more. She trained with the Mercury 13 in 1961,…
OpenAI shipped GPT-5.6 across three price tiers and launched ChatGPT Work, while Meta released Muse Spark 1.1 and its first commercial API on the same day.
Hint for all AI Labs as they branch out from work for programming to general knowledge work: non-coders are not just dumber codersTaking away a bunch…