TY - JOUR
T1 - COOPMamba
T2 - Efficient Vehicle-to-Vehicle Cooperative Perception Based on 3-D Point Clouds
AU - Zhang, Peng
AU - Chen, Xinju
AU - Liang, Yunji
AU - Yan, Xiaokai
AU - Yu, Zhiwen
N1 - Publisher Copyright:
© 2001-2012 IEEE.
PY - 2026/5/1
Y1 - 2026/5/1
N2 - Vehicle-to-vehicle (V2V) cooperative perception promises enhanced environmental awareness but still suffers from two critical gaps: 1) existing intermediate-fusion models typically rely on single-branch feature aggregation, which struggles to simultaneously capture localized discriminative cues and long-range contextual structures and 2) attention-based fusion modules, like Transformers, incur high computational and communication overhead, limiting real-time deployment. To address these limitations, we propose COOPMamba in this article, a lightweight yet expressive collaborative-perception framework that introduces a dual-branch decomposition of salient and global information and a CMamba cross-branch state-space interaction block. The proposed CMamba block performs direction-aware, long-range feature propagation, enabling more robust multivehicle feature alignment and mitigating missed detections caused by incomplete observations. Extensive experiments on V2XSet and OPV2V demonstrate the effectiveness of our design: COOPMamba achieves 89.1% AP@0.5 and 77.1% AP@0.7 on V2XSet, outperforming state-of-the-art (SOTA) approaches by 2.5%-6.4% while maintaining the lowest MACs and inference latency among existing fusion methods. The results confirm that our cross-branch state-space modeling substantially improves collaborative 3-D object detection under both ideal and noisy real-world conditions. The source code is available at https://github.com/npunancy/coopmamba
AB - Vehicle-to-vehicle (V2V) cooperative perception promises enhanced environmental awareness but still suffers from two critical gaps: 1) existing intermediate-fusion models typically rely on single-branch feature aggregation, which struggles to simultaneously capture localized discriminative cues and long-range contextual structures and 2) attention-based fusion modules, like Transformers, incur high computational and communication overhead, limiting real-time deployment. To address these limitations, we propose COOPMamba in this article, a lightweight yet expressive collaborative-perception framework that introduces a dual-branch decomposition of salient and global information and a CMamba cross-branch state-space interaction block. The proposed CMamba block performs direction-aware, long-range feature propagation, enabling more robust multivehicle feature alignment and mitigating missed detections caused by incomplete observations. Extensive experiments on V2XSet and OPV2V demonstrate the effectiveness of our design: COOPMamba achieves 89.1% AP@0.5 and 77.1% AP@0.7 on V2XSet, outperforming state-of-the-art (SOTA) approaches by 2.5%-6.4% while maintaining the lowest MACs and inference latency among existing fusion methods. The results confirm that our cross-branch state-space modeling substantially improves collaborative 3-D object detection under both ideal and noisy real-world conditions. The source code is available at https://github.com/npunancy/coopmamba
KW - 3-D object detection
KW - autonomous driving
KW - state-space model (SSM)
KW - vehicle-to-vehicle (V2V) cooperative perception
UR - https://www.scopus.com/pages/publications/105036331719
U2 - 10.1109/JSEN.2026.3682367
DO - 10.1109/JSEN.2026.3682367
M3 - 文章
AN - SCOPUS:105036331719
SN - 1530-437X
VL - 26
SP - 16479
EP - 16489
JO - IEEE Sensors Journal
JF - IEEE Sensors Journal
IS - 10
ER -