TY - GEN
T1 - Understanding Global Structure Relation via Reversible Visual State Space Model for Robust Cross-View Geo-Localization
AU - Ma, Peiyuan
AU - Fu, Yimin
AU - Lyu, Jialin
AU - Liu, Zhunga
N1 - Publisher Copyright:
© 2025 owner/author(s)
PY - 2025/11/3
Y1 - 2025/11/3
N2 - Cross-view geo-localization aims to match images captured from different views over the same geographic region. Existing methods typically determine spatial correlations between cross-view images according to the similarity of representations extracted from salient areas. However, such local appearance representations fail to capture the underlying structural relationships among the corresponding regions, which severely undermines the reliability of localization results in complex scenes. To address this problem, we propose a reversible visual state-space model to enhance the understanding of global structural relations inherent in images captured from different views. Specifically, we design a progressive spatial analysis approach, which incrementally integrates geometric dependencies exploited at different levels to improve the understanding of the global structure. Moreover, we introduce a reversible rotational scanning mechanism based on the 2D-selective-scan (SS2D) module to facilitate the exploitation of geometric dependencies between cross-view images. Finally, we adopt the cross-dimension interaction strategy to enrich the informativeness of representations in the common space, thereby reinforcing the discriminability of cross-view representations between different regions. Extensive experiments on the University-1652 and University160k-WX datasets demonstrate that the proposed method achieves state-of-the-art performance while maintaining robustness under complex environmental conditions.
AB - Cross-view geo-localization aims to match images captured from different views over the same geographic region. Existing methods typically determine spatial correlations between cross-view images according to the similarity of representations extracted from salient areas. However, such local appearance representations fail to capture the underlying structural relationships among the corresponding regions, which severely undermines the reliability of localization results in complex scenes. To address this problem, we propose a reversible visual state-space model to enhance the understanding of global structural relations inherent in images captured from different views. Specifically, we design a progressive spatial analysis approach, which incrementally integrates geometric dependencies exploited at different levels to improve the understanding of the global structure. Moreover, we introduce a reversible rotational scanning mechanism based on the 2D-selective-scan (SS2D) module to facilitate the exploitation of geometric dependencies between cross-view images. Finally, we adopt the cross-dimension interaction strategy to enrich the informativeness of representations in the common space, thereby reinforcing the discriminability of cross-view representations between different regions. Extensive experiments on the University-1652 and University160k-WX datasets demonstrate that the proposed method achieves state-of-the-art performance while maintaining robustness under complex environmental conditions.
KW - Cross-view Geo-localization
KW - State Space Model
KW - Structure Relation
UR - https://www.scopus.com/pages/publications/105023071894
U2 - 10.1145/3728482.3757390
DO - 10.1145/3728482.3757390
M3 - 会议稿件
AN - SCOPUS:105023071894
T3 - UAVM 2025 - Proceedings of the 3rd International Workshop on UAVs in Multimedia: Capturing the World from a New Perspective, Co-located with MM 2025
SP - 42
EP - 46
BT - UAVM 2025 - Proceedings of the 3rd International Workshop on UAVs in Multimedia
PB - Association for Computing Machinery, Inc
T2 - 3rd International Workshop on UAVs in Multimedia: Capturing the World a from New Perspective, UAVM 2025
Y2 - 27 October 2025 through 31 October 2025
ER -