Before We Defer Research to AI: Measuring Apparent-Success-Seeking
Recently, I was improving a small LLM-powered classifier and noticed a few continuously failing test cases. As many would, I asked my AI code assistant to…
Recently, I was improving a small LLM-powered classifier and noticed a few continuously failing test cases. As many would, I asked my AI code assistant to…
In 2015 I made a little webapp that would use the NextBus API to show predictions for the MBTA: A few years later they moved from…
This is a pilot experiment, done on one model, with around $14 worth of compute, and a single seed per condition. The full writeup with all…
hate how DS has not increased prices yet and is just handwringing, “sorry, we’re ashamed, we’ll return the money… run, run away little ones! To better…
Article URL: https://hypercritical.co/hyperspace/ Comments URL: https://news.ycombinator.com/item?id=49252493 Points: 11 # Comments: 4
Article URL: https://tokenstead.ai/ Comments URL: https://news.ycombinator.com/item?id=49252486 Points: 6 # Comments: 1
Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visual context. Existing perception-aware post-training methods encourage image use…
Article URL: https://www.floppydisk.com/recycle Comments URL: https://news.ycombinator.com/item?id=49252462 Points: 8 # Comments: 0
Concerning Satanyahu’s reach is vast indeedZaldrīzes buzdari iksos daor: @KenKirtland17 I don’t know a single person who thinks China has a benign space program. Every aspect…
Yesterday I sat down with GPT 5.6 Sol High to do some brainstorming. The topic was one of the less appreciated Millenium Problems (the Birch and…
Article URL: https://github.com/antirez/h3.c Comments URL: https://news.ycombinator.com/item?id=49252179 Points: 45 # Comments: 4
Testing with 3-5 agents. Decode performance is superb, however if one performs a web search and needs to process a few thousand tokens, ALL other agents…
Introduction and Related workThe first person perspective of various experiences are subjective experiences. For Large language models, the study of subjective experiences was recently studied by…
Article URL: https://www.nytimes.com/2026/08/10/us/flock-cameras-can-track-every-car-in-america-police-love-them-citizens-dont.html Comments URL: https://news.ycombinator.com/item?id=49251978 Points: 4 # Comments: 1
Grasping at strawsWell at least it’ll accelerate the brain drainSelect Committee on China: Are you an American student, postdoctoral researcher, or academic who believes you were…
These numbers were captured during a real feature implementation task in Next.js and Nest.js (adding a theme switching system across components). The structural predictability of UI/state…
This really happenedNew York Post: Fauci privately warned of miscarriage risk linked to COVID vaccine while publicly claiming no issues, newly released texts show https://trib.al/bDcp5GG
Just downloaded the model, UD-Q5_K_XL quant, asked it to generate a long story to test out reasoning and speed with dflash (super fast btw, ~ 90…
I don’t think China is desperate for this to end, BillThey can take a few more months of less oilYou, on the other hand…Bill Mitchell: Here's…
Article URL: https://code.call-cc.org/releases/6.0.0/NEWS Comments URL: https://news.ycombinator.com/item?id=49251702 Points: 4 # Comments: 0
In retrospect it’s amusing that Leninism, as a dictatorship of the “vanguard party”, didn’t really depend on an economic model and could be Fine Actually if…
The most Spiritually Chinese of the AnglosElon Musk: @ARKInvest China is awesome. I strongly encourage people to visit.
Article URL: https://www.404media.co/ice-to-pay-lexisnexis-millions-for-data-to-feed-to-palantir/ Comments URL: https://news.ycombinator.com/item?id=49251643 Points: 12 # Comments: 1
San Francisco's housing market is in trouble again.