TY - JOUR
T1 - You Only Scan Once
T2 - Efficient Multi-dimension Sequential Modeling with LightNet
AU - Qin, Zhen
AU - Mao, Yuxin
AU - Shen, Xuyang
AU - Li, Dong
AU - Zhang, Jing
AU - Dai, Yuchao
AU - Zhong, Yiran
N1 - Publisher Copyright:
© 2026, Transactions on Machine Learning Research. All rights reserved.
PY - 2026
Y1 - 2026
N2 - Linear attention mechanisms have gained prominence in causal language models due to their linear computational complexity and enhanced speed. However, the inherent decay mechanism in linear attention presents challenges when applied to multi-dimensional sequence modeling tasks, such as image processing and multi-modal learning. In these scenarios, the utilization of sequential scanning to establish a global receptive field necessitates multiple scans for multi-dimensional data, thereby leading to inefficiencies. This paper identifies the inefficiency caused by a “multiplicative decay” linear recurrence and proposes an efficient alternative “additive decay” linear recurrence to avoid the issue, as it can handle multidimensional data within a single scan. We further develop an efficient multi-dimensional sequential modeling framework called LightNet based on the new recurrence. Moreover, we present two new multi-dimensional linear relative positional encoding methods, MD-TPE and MD-LRPE to enhance the model’s ability to discern positional information in multidimensional scenarios. Our empirical evaluations across various tasks, including image classification, image generation, bidirectional language modeling, and autoregressive language modeling, demonstrate the efficacy of LightNet, showcasing its potential as a versatile and efficient solution for multi-dimensional sequential modeling.
AB - Linear attention mechanisms have gained prominence in causal language models due to their linear computational complexity and enhanced speed. However, the inherent decay mechanism in linear attention presents challenges when applied to multi-dimensional sequence modeling tasks, such as image processing and multi-modal learning. In these scenarios, the utilization of sequential scanning to establish a global receptive field necessitates multiple scans for multi-dimensional data, thereby leading to inefficiencies. This paper identifies the inefficiency caused by a “multiplicative decay” linear recurrence and proposes an efficient alternative “additive decay” linear recurrence to avoid the issue, as it can handle multidimensional data within a single scan. We further develop an efficient multi-dimensional sequential modeling framework called LightNet based on the new recurrence. Moreover, we present two new multi-dimensional linear relative positional encoding methods, MD-TPE and MD-LRPE to enhance the model’s ability to discern positional information in multidimensional scenarios. Our empirical evaluations across various tasks, including image classification, image generation, bidirectional language modeling, and autoregressive language modeling, demonstrate the efficacy of LightNet, showcasing its potential as a versatile and efficient solution for multi-dimensional sequential modeling.
UR - https://www.scopus.com/pages/publications/105041078041
M3 - 文章
AN - SCOPUS:105041078041
SN - 2835-8856
VL - 2026-May
JO - Transactions on Machine Learning Research
JF - Transactions on Machine Learning Research
ER -