TY - JOUR
T1 - BEMN
T2 - Balanced Bias Enhanced Multi-Branch Network for Cross-View Geo-Localization
AU - Bi, Cheng
AU - Sun, Bo
AU - Wang, Jingfeng
AU - Yuan, Yuan
AU - Liu, Ganchao
N1 - Publisher Copyright:
© 1991-2012 IEEE.
PY - 2026/5/1
Y1 - 2026/5/1
N2 - Cross-view geo-localization (CVGL) offers a promising alternative for positioning in GNSS-constrained environments through visual matching techniques. Extreme viewpoint variations and the complexity of real-world scenes present significant challenges to this task. However, current methods primarily focus on learning single-scale features, which may be inadequate for practical applications. Although some approaches attempt to incorporate multi-scale representations, they may suffer from unimodal bias arising from structural discrepancies among model branches, limiting effective multi-scale feature extraction. To address these issues, we propose a fully multi-branch network architecture, named BEMN, which is designed to learn multi-scale robust feature representations. Specifically, we construct a multi-branch backbone network based on pretrained visual models and design a two-stage training strategy. In the first stage, a separate training scheme is employed to thoroughly optimize each branch of the network, and a joint feature alignment (JFA) module is introduced to align cross-view features. The entire network is fine-tuned in the second stage, where a frequency domain adjustment (FDA) module is designed to improve performance. To further assess the generalization ability of CVGL methods, we establish Xian-37, a highly challenging CVGL test dataset featuring complex real scenes captured from diverse platforms and viewpoints. Experimental results across multiple public benchmarks validate the superiority of our approach, achieving state-of-the-art performance and demonstrating outstanding generalization capabilities. Our code and model are available at https://github.com/VERYBC/BEMN
AB - Cross-view geo-localization (CVGL) offers a promising alternative for positioning in GNSS-constrained environments through visual matching techniques. Extreme viewpoint variations and the complexity of real-world scenes present significant challenges to this task. However, current methods primarily focus on learning single-scale features, which may be inadequate for practical applications. Although some approaches attempt to incorporate multi-scale representations, they may suffer from unimodal bias arising from structural discrepancies among model branches, limiting effective multi-scale feature extraction. To address these issues, we propose a fully multi-branch network architecture, named BEMN, which is designed to learn multi-scale robust feature representations. Specifically, we construct a multi-branch backbone network based on pretrained visual models and design a two-stage training strategy. In the first stage, a separate training scheme is employed to thoroughly optimize each branch of the network, and a joint feature alignment (JFA) module is introduced to align cross-view features. The entire network is fine-tuned in the second stage, where a frequency domain adjustment (FDA) module is designed to improve performance. To further assess the generalization ability of CVGL methods, we establish Xian-37, a highly challenging CVGL test dataset featuring complex real scenes captured from diverse platforms and viewpoints. Experimental results across multiple public benchmarks validate the superiority of our approach, achieving state-of-the-art performance and demonstrating outstanding generalization capabilities. Our code and model are available at https://github.com/VERYBC/BEMN
KW - Cross-view geo-localization (CVGL)
KW - dataset
KW - feature representation
KW - image retrieval
KW - multi-branch network
UR - https://www.scopus.com/pages/publications/105024787482
U2 - 10.1109/TCSVT.2025.3642751
DO - 10.1109/TCSVT.2025.3642751
M3 - 文章
AN - SCOPUS:105024787482
SN - 1051-8215
VL - 36
SP - 5968
EP - 5982
JO - IEEE Transactions on Circuits and Systems for Video Technology
JF - IEEE Transactions on Circuits and Systems for Video Technology
IS - 5
ER -