Skip to content
arXiv cs.AI · Papers

AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally

arXiv:2607.19363v1 Announce Type: new Abstract: Rotary Position Embedding (RoPE) is widely adopted in Transformers to encode positional information, yet standard implementations enforce a uniform frequency schedule and scaling across all attention heads. Using simplified retrieval tasks and length generalization scenar