Skip to main navigation Skip to search Skip to main content

SCVM3D: A Hybrid Backbone with Sparse Convolutions and Mamba-2 for Voxel-based 3D Object Detection

  • Northwestern Polytechnical University Xian
  • Mohamed Bin Zayed University of Artificial Intelligence

Research output: Contribution to journalArticlepeer-review

Abstract

Accurate and efficient 3D object detection is crucial for autonomous driving, where LiDAR-based methods are essential for capturing high-precision spatial information in large-scale outdoor environments. Among existing approaches, voxel-based methods, which partition point clouds into structured grids, offer a computationally efficient paradigm. However, these approaches struggle to model long-range dependencies and capture global contextual information. Recent advances in state-space models (SSMs), particularly the Mamba architecture, offer a promising solution by enabling efficient global feature modeling with linear complexity. Nevertheless, the inherently sequential nature of SSMs is misaligned with the spatially unordered structure of voxel data. In this paper, we propose SCVM3D, a novel backbone architecture that combines sparse convolution and the SSM-based Mamba-2 model to address the limitations of voxel-based 3D object detection. SCVM3D integrates three core components: 1) a hierarchical sparse convolution module that efficiently aggregates local geometric cues, 2) an importance-aware adaptive voxel reordering scheme that optimizes the token sequence, and 3) a Mamba-2 module that captures global contextual relationships across the scene. By dynamically prioritizing voxel features and using the linear-complexity modeling capability of Mamba-2 for long-range dependencies, SCVM3D effectively captures both local and global information from voxelized LiDAR data. Extensive experiments on both the KITTI and nuScenes datasets demonstrate that SCVM3D consistently improves the performance of strong baseline detectors across diverse environments. These results underscore the potential of SCVM3D as an effective and efficient backbone for large-scale 3D object detection.

Keywords

  • 3D object detection
  • Autonomous driving
  • mamba-2
  • state space model
  • voxel

Fingerprint

Dive into the research topics of 'SCVM3D: A Hybrid Backbone with Sparse Convolutions and Mamba-2 for Voxel-based 3D Object Detection'. Together they form a unique fingerprint.

Cite this