Skip to content
arXiv cs.AI · Papers

When Search Teaches Style: Causal Internalization of Tactical Priors in AlphaZero

arXiv:2504.14636v3 Announce Type: replace-cross Abstract: AlphaZero is normally evaluated as one agent: a policy-value network fused with Monte Carlo tree search. That fusion hides a causal question. When self-play search is given a useful prior, does the network absorb the induced behavior, or does the behavior stay r