r/LocalLLaMA
· Communities
arXiv publication: "Skip a Layer or Loop It? Learning Program-of-Layers in LLMs"
In this paper, Li, Li, and Zhou review formal theory describing how inference-time compute can be traded off for higher or lower inference competence, and apply that theory to a handful of familiar open-weight LLMs (Llama-3.2, Qwen1.5, Qwen2.5, and Qwen3): https://arxiv.org/abs/2606.06574v1 > Large language models (LLM