Community Polls on Alignment Controversies II
Please spend
Please spend
OverviewThese are my non-expert notes on the compute verification section of AIFP’s Plan A. I cover interconnect limits, memory wipes, network taps + replay, and ZKPs.…
Thanks to Fabien Roger (Anthropic), who pointed out this system card mistake to me. This mistake will likely be fixed in the relevant system cards after…
tl;dr:Lindsey 2025 found models can modulate their internal states: when instructed to “think about” a concept while writing an unrelated sentence, the representation of the concept…
As AI capabilities rapidly advance, we face critical information gaps in effective AI risk management: What are the risks from AI, which are most important, and…
Humans make carbon dioxide. Carbon dioxide is bad for cognition. But plants turn carbon dioxide back into oxygen. And plants are the one true home decoration…
Consider the following situations:when you are a small, growing startup in a big market, standard advice is not to worry too much about your competitors or…
The text as follows:see the below—makes Claude think that the prompt is unfinished, and fill in its own prompt.It will subsequently claim that it recieved what…
I spent the past week designing a test that I hoped would serve as a benchmark. But LLMs are improving faster than I expected, and my…
TL;DR. LessWrong's decision-theory debates (Newcomb, FDT vs CDT, counterfactual muggings) are almost entirely about what we suppose when we consider a candidate action or policy. There…