Skip to content
arXiv cs.LG · Papers

CausalGate: Causal Importance Distillation for Transformer Module Pruning

arXiv:2607.22720v1 Announce Type: new Abstract: Existing adaptive inference methods for Large Language Models rely on observational heuristics, such as hidden-state similarity or activation magnitudes, to drop redundant modules. However, these correlation-based metrics often fail to capture subtle, non-linear structura