TensorX
返回文献探索

Paper · arXiv 2606.06574

Skip a Layer or Loop It? Learning Program-of-Layers in LLMs

Ziyue Li, Yang Li, Tianyi Zhou

26 upvotesJune 4, 2026arXiv 预印本
AI 摘要

Pretrained language models can execute layers dynamically through flexible program-of-layers strategies that improve accuracy while reducing computational overhead compared to standard fixed-depth inference.

large language modelsprogram-of-layersdynamic program-of-layersPoLarpretrained layerslayer skippinglayer loopingexecution programsdynamic depthlatent reasoning capacity

Abstract

Large language models (LLMs) perform inference by following a fixed depth and order, non-recurrent execution of all layers. We reveal the wide existence of training-free, flexible, dynamic program-of-layers (PoLar), where pretrained layers can be packed as modules and then skipped or looped to form a customized program for each input. For most inputs, substantially shorter program executions can achieve the same or better accuracy, while incorrect predictions of the original LLM can be corrected by alternative programs with fewer layers. These observations indicate that inference admits multiple valid latent computations beyond the standard forward pass. To efficiently achieve PoLar in practice, we propose a lightweight PoLar prediction network, which learns to generate execution programs that dynamically skip or repeat pretrained layers for each input. Experiments on mathematical reasoning benchmarks demonstrate that PoLar consistently improves accuracy over standard inference and prior dynamic-depth methods, often while executing fewer layers, and that these gains persist under out-of-distribution evaluation. Our results suggest that fixed-depth execution captures only a narrow subset of an LLM's latent reasoning capacity.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Skip a Layer or Loop It? Learning Program-of-Layers in LLMs | TensorX