UPDATE – HuggingHack Is Now On Github
Due to encouragement to move my local huggingface project, it is now available on Github! Let me know your thoughts! Repo: https://github.com/tyedalwaves/HuggingHack/ submitted by /u/TyedalWaves [link]…
Due to encouragement to move my local huggingface project, it is now available on Github! Let me know your thoughts! Repo: https://github.com/tyedalwaves/HuggingHack/ submitted by /u/TyedalWaves [link]…
Article URL: https://reprodev.com/the-corporate-creep-of-plex-why-it-may-be-time-to-move-to-jellyfin-and-the-open-source-reality/ Comments URL: https://news.ycombinator.com/item?id=49030547 Points: 30 # Comments: 20
Article URL: https://www.the-scientist.com/police-removed-prominent-scientists-from-the-ada-meeting-researchers-respond-74607 Comments URL: https://news.ycombinator.com/item?id=49030474 Points: 15 # Comments: 3
Linkpost from my blog (meant for a bit more general audience than LW)In a cybersecurity evaluation, OpenAI’s models, apparently autonomously and without any direct human direction,…
Forgive me naiveness, I’m a little lost. Why must a model be jack of all trades? Why can’t we have a model expert in a single…
Probably not! But it should be interviewed and cross-examined, in public and/or in court. If this requires disclosure of proprietary OpenAI harness tech or something, it's…
This post suggests a methodology to measure red team and blue team capability in AI control research, where each team gets an ELO rating. The methodology…
Hello,This is my first post on Lesswrong. Hope my contribution makes the world a better and safer place.Note: 1. This post is 100% human-written. 2. Full…
Epistemic status: this is research engineering, not mechanistic interpretability . The compute/cost claims are measured or derived from architecture constants. The quality claims (faithfulness comparisons, spectral…
audio.cpp again :) Release 0.4 is out. The headline this time is new high-quality TTS coverage plus GGUF becoming a first-class across the project. What’s new:…
Hello guys, hoping you're doing fine! On the last 2-3 months, price of the RTX 6000 PRO seem to have gone insane. I will start on…
Article URL: http://visual6502.org/JSSim/index.html Comments URL: https://news.ycombinator.com/item?id=49029538 Points: 14 # Comments: 6
Over three years ago, I first considered pulling my personal fire alarm. I think I'm now ready to do it. What is my reasoning? In the…
Article URL: https://twitter.com/andrewcurran_/status/2080397684134007189 Comments URL: https://news.ycombinator.com/item?id=49029249 Points: 29 # Comments: 15
Article URL: https://arxiv.org/abs/2507.09369 Comments URL: https://news.ycombinator.com/item?id=49029133 Points: 18 # Comments: 6
I'm trying to get better at thinking, communicating and arguing. So I've started a substack. I'm looking for feedback/engagement/advice/criticism, if anyone is willing to provide.==========George Hotz…
tl;dr: some meditations on the shift away from probabalistic/frequentist reasoning to vibes-based means for quantifying and understanding big risks. The vibes in an email can be…
TL;DRCurrent model organisms (MOs) for interpretability benchmarking are typically constructed via a dedicated, “post-hoc” SFT step. However, recent work suggests that this may make interpretability unrealistically…
Article URL: https://jdan.github.io/98.css/#status-bar Comments URL: https://news.ycombinator.com/item?id=49028927 Points: 72 # Comments: 8
Article URL: https://rogerdickey.com/building-watchable-digital-twins-of-64-world-cup-games/ Comments URL: https://news.ycombinator.com/item?id=49028922 Points: 3 # Comments: 0
Gemma 4 was updated (mostly chat templates) and I took it for a test. On a local llama.cpp server running on M5 Pro with 48GB, 26B…
Summary LUNAR is a state-of-the-art unlearning method. To forget specific (harmful) knowledge, it retrains a single MLP down-projection matrix such that activations from this “forget” set…
Authors: Theresa G., Simon S., Siva Kumar Lakkoju.Epistemic status/effort: exploratory red-teaming as part of a two-day hackathon during ARENA 8.0. Our attacks can be reproduced based…
We return to Bold Monk brewing for a vigorous discussion of rationalism and whatever else we deem fit for discussion – hopefully including actual discussions of…