arXiv cs.AI
· Papers
AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally
arXiv:2607.19363v1 Announce Type: new Abstract: Rotary Position Embedding (RoPE) is widely adopted in Transformers to encode positional information, yet standard implementations enforce a uniform frequency schedule and scaling across all attention heads. Using simplified retrieval tasks and length generalization scenar