跳到主要导航 跳到搜索 跳到主要内容

Sliding space-disparity transformer for stereo matching

  • Northwestern Polytechnical University Xian
  • Baidu Inc

科研成果: 期刊稿件文章同行评审

12 引用 (Scopus)

摘要

Transformers have achieved impressive performance in natural language processing and computer vision, including text translation, semantic segmentation, etc. However, due to excessive self-attention computation and memory occupation, the stereo matching task does not share its success. To promote this technology in stereo matching, especially with limited hardware resources, we propose a sliding space-disparity transformer named SSD-former. According to matching modeling, we simplify transformer for achieving faster speed, memory-friendly, and competitive performance. First, we employ the sliding window scheme to limit the self-attention operations in the cost volume for adapting to different resolutions, bringing efficiency and flexibility. Second, our space-disparity transformer remarkably reduces memory occupation and computation, only computing the current patch’s self-attention with two parts: (1) all patches of current disparity level at the whole spatial location and (2) the patches of different disparity levels at the exact spatial location. The experiments demonstrate that: (1) different from the standard transformer, SSD-former is faster and memory-friendly; (2) compared with 3D convolution methods, SSD-former has a larger receptive field and provides an impressive speed, showing great potential in stereo matching; and (3) our model obtains state-of-the-art performance and a faster speed on the multiple popular datasets, achieving the best speed–accuracy trade-off.

源语言英语
页(从-至)21863-21876
页数14
期刊Neural Computing and Applications
34
24
DOI
出版状态已出版 - 12月 2022

学术指纹

探究 'Sliding space-disparity transformer for stereo matching' 的科研主题。它们共同构成独一无二的学术指纹。

引用此