TY - JOUR
T1 - SCVM3D
T2 - A Hybrid Backbone with Sparse Convolutions and Mamba-2 for Voxel-based 3D Object Detection
AU - Zhang, Lei
AU - Wang, Zhaozhong
AU - Zhao, Yongqiang
AU - Fan, Bin
AU - Liu, Nian
AU - Wang, Binglu
N1 - Publisher Copyright:
© 1991-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Accurate and efficient 3D object detection is crucial for autonomous driving, where LiDAR-based methods are essential for capturing high-precision spatial information in large-scale outdoor environments. Among existing approaches, voxel-based methods, which partition point clouds into structured grids, offer a computationally efficient paradigm. However, these approaches struggle to model long-range dependencies and capture global contextual information. Recent advances in state-space models (SSMs), particularly the Mamba architecture, offer a promising solution by enabling efficient global feature modeling with linear complexity. Nevertheless, the inherently sequential nature of SSMs is misaligned with the spatially unordered structure of voxel data. In this paper, we propose SCVM3D, a novel backbone architecture that combines sparse convolution and the SSM-based Mamba-2 model to address the limitations of voxel-based 3D object detection. SCVM3D integrates three core components: 1) a hierarchical sparse convolution module that efficiently aggregates local geometric cues, 2) an importance-aware adaptive voxel reordering scheme that optimizes the token sequence, and 3) a Mamba-2 module that captures global contextual relationships across the scene. By dynamically prioritizing voxel features and using the linear-complexity modeling capability of Mamba-2 for long-range dependencies, SCVM3D effectively captures both local and global information from voxelized LiDAR data. Extensive experiments on both the KITTI and nuScenes datasets demonstrate that SCVM3D consistently improves the performance of strong baseline detectors across diverse environments. These results underscore the potential of SCVM3D as an effective and efficient backbone for large-scale 3D object detection.
AB - Accurate and efficient 3D object detection is crucial for autonomous driving, where LiDAR-based methods are essential for capturing high-precision spatial information in large-scale outdoor environments. Among existing approaches, voxel-based methods, which partition point clouds into structured grids, offer a computationally efficient paradigm. However, these approaches struggle to model long-range dependencies and capture global contextual information. Recent advances in state-space models (SSMs), particularly the Mamba architecture, offer a promising solution by enabling efficient global feature modeling with linear complexity. Nevertheless, the inherently sequential nature of SSMs is misaligned with the spatially unordered structure of voxel data. In this paper, we propose SCVM3D, a novel backbone architecture that combines sparse convolution and the SSM-based Mamba-2 model to address the limitations of voxel-based 3D object detection. SCVM3D integrates three core components: 1) a hierarchical sparse convolution module that efficiently aggregates local geometric cues, 2) an importance-aware adaptive voxel reordering scheme that optimizes the token sequence, and 3) a Mamba-2 module that captures global contextual relationships across the scene. By dynamically prioritizing voxel features and using the linear-complexity modeling capability of Mamba-2 for long-range dependencies, SCVM3D effectively captures both local and global information from voxelized LiDAR data. Extensive experiments on both the KITTI and nuScenes datasets demonstrate that SCVM3D consistently improves the performance of strong baseline detectors across diverse environments. These results underscore the potential of SCVM3D as an effective and efficient backbone for large-scale 3D object detection.
KW - 3D object detection
KW - Autonomous driving
KW - mamba-2
KW - state space model
KW - voxel
UR - https://www.scopus.com/pages/publications/105045209881
U2 - 10.1109/TCSVT.2026.3712403
DO - 10.1109/TCSVT.2026.3712403
M3 - 文章
AN - SCOPUS:105045209881
SN - 1051-8215
JO - IEEE Transactions on Circuits and Systems for Video Technology
JF - IEEE Transactions on Circuits and Systems for Video Technology
ER -