Scheming Evals Mislead in Both Directions
We spent several weeks measuring in-context scheming, the behavior where a model covertly pursues a misaligned goal while outwardly appearing to comply, and the result that…
We spent several weeks measuring in-context scheming, the behavior where a model covertly pursues a misaligned goal while outwardly appearing to comply, and the result that…
From clem 🤗 on X: https://x.com/ClementDelangue/status/2072683707001930215 From Palantir on X (video): https://x.com/PalantirTech/status/2072326189079757277 The information: Palantir CEO Says Some U.S. Government Customers Switched to Open Source AI:…
Finally got around to testing whether enabling P2P actually matters on a dual 3090 rig (PCIe 4.0 8x/8x), instead of just taking it on faith. Ran…
Article URL: https://www.swiftsilentdeadly.com/blog/gun-mistakes-in-fiction-writing-handgun-edition Comments URL: https://news.ycombinator.com/item?id=48773652 Points: 28 # Comments: 27
Article URL: https://www.youtube.com/watch?v=TtmPccUTDP8 Comments URL: https://news.ycombinator.com/item?id=48773580 Points: 20 # Comments: 14
Sometimes a reasoning model appears to pass through the correct answer before ending up wrongMotivationFigure 1: From the Opus 4.8 system card (page 196)Figure 1 shows…
Hi folks. I found this video explaining latest DSpark breakthrough from Deepseek. Seems like a huge change coming. https://www.youtube.com/watch?v=J0D7qV3nl7w submitted by /u/BringTea_666 [link] [comments]
Article URL: https://thombrown.blogspot.com/2026/07/load-plcbmbasic81-commodore-64-basic.html Comments URL: https://news.ycombinator.com/item?id=48772717 Points: 12 # Comments: 4
For open-weight LLMs, how practical is it to study defenses against post-release fine-tuning that weakens refusal or safety behavior? I've been seeing “uncensored” or “heretic”…
Article URL: https://wordgard.net/ Comments URL: https://news.ycombinator.com/item?id=48772573 Points: 20 # Comments: 1
Article URL: https://www.reuters.com/world/china/alibaba-ban-claude-code-workplace-over-alleged-backdoor-risks-source-says-2026-07-03/ Comments URL: https://news.ycombinator.com/item?id=48772443 Points: 12 # Comments: 2
I've been working on mapping (with tags) and steering local models based on their activation path in specific context to questioning during a/b testing. There is…
Article URL: https://weli.dev/blog/half-baked-product/ Comments URL: https://news.ycombinator.com/item?id=48772388 Points: 20 # Comments: 6
This is a follow-up to post about which local models stay fast deep into long context and I learned a lot from people here. I kept…
The list of suspicious hostnames is not stored in plaintext within the code; instead, it is Base64-encoded and then encrypted using a simple XOR operation with…
Follow-up: GLM-5.2 NVFP4 on four DGX Sparks — the MTP mystery is solved, and it's now ~24 tok/s at 128K context This is a follow-up to…
I'm a bit annoyed by the feeling that we're kind of stuck when it comes to using LLMs for programming.I use Claude Code and Codex, but…
If an application uses a Web-based interface and "hardware acceleration", it constructs its frame in VRAM and sometimes keeps it reserved even if the app is…
In this post I walk through the first Technical AI Safety puzzle from BlueDot and why linear probes would have missed all the most interesting stuff.In…
Disclaimer, this is a self promotion post but I truly believe it belongs here and is quite useful and relevant to this crowd specifically. I built…
Article URL: https://wasmer.io/ Comments URL: https://news.ycombinator.com/item?id=48770677 Points: 5 # Comments: 0
Article URL: https://stefan.schueller.net/posts/the-free-market-lie/ Comments URL: https://news.ycombinator.com/item?id=48770647 Points: 3 # Comments: 0
Article URL: https://canonry.ai/blog/ai-visibility-tools-are-lying Comments URL: https://news.ycombinator.com/item?id=48770575 Points: 9 # Comments: 1
Article URL: https://manticoresearch.com/blog/onnx-embeddings-speedup/ Comments URL: https://news.ycombinator.com/item?id=48770477 Points: 3 # Comments: 0