TY - JOUR
T1 - Causal-Aware Feature Refinement with Dual Encoders for Remote Sensing Visual Question Answering
AU - Yang, Zhigang
AU - Yao, Huiguang
AU - Tian, Linmao
AU - Liu, Shanji
AU - Ni, Weiping
AU - Wu, Junzheng
AU - Li, Qiang
AU - Wang, Qi
N1 - Publisher Copyright:
© 1999-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Remote sensing visual question answering has emerged as a critical task bridging computer vision and natural language processing in geospatial understanding, yet it faces significant limitations in supporting complex reasoning and multilingual interaction. Existing datasets are predominantly English-only, focus on low-level tasks, and lack high-resolution imagery, restricting real-world applicability. To address these gaps, we introduce NWPU-RSVQA, the first large-scale bilingual dataset for high-order relational reasoning in RSVQA. It comprises 10,718 high-resolution (1,024-2,048px) global images with 107,180 human-verified QA pairs, enabling reasoning across attribute-relation-scene levels. Despite these advances, existing RSVQA methods struggle to perform high-order relational reasoning due to three fundamental limitations: (i) susceptibility to spurious correlations between visual patterns and answers, (ii) inefficient processing of high-resolution visual tokens, and (iii) underutilization of semantic priors. To address these issues, a causality-aware framework with three key innovations: (i) A dual-path vision encoder integrating ViT for global semantics and SAM for segmentation priors, generating comprehensive visual features. (ii) A Causal Embedding Module that explicitly disentangles answer-relevant causal features from non-causal noise via a gating mechanism, mitigating spurious correlations. (iii) A Dynamic Cross-Attention Token Sampler that adaptively compresses visual tokens by importance, reducing computational burden while preserving critical information. Extensive experiments on NWPU-RSVQA demonstrate that CausalVisQA outperforms state-of-the-art models across semantic consistency and factual accuracy metrics, particularly in complex relational reasoning tasks.
AB - Remote sensing visual question answering has emerged as a critical task bridging computer vision and natural language processing in geospatial understanding, yet it faces significant limitations in supporting complex reasoning and multilingual interaction. Existing datasets are predominantly English-only, focus on low-level tasks, and lack high-resolution imagery, restricting real-world applicability. To address these gaps, we introduce NWPU-RSVQA, the first large-scale bilingual dataset for high-order relational reasoning in RSVQA. It comprises 10,718 high-resolution (1,024-2,048px) global images with 107,180 human-verified QA pairs, enabling reasoning across attribute-relation-scene levels. Despite these advances, existing RSVQA methods struggle to perform high-order relational reasoning due to three fundamental limitations: (i) susceptibility to spurious correlations between visual patterns and answers, (ii) inefficient processing of high-resolution visual tokens, and (iii) underutilization of semantic priors. To address these issues, a causality-aware framework with three key innovations: (i) A dual-path vision encoder integrating ViT for global semantics and SAM for segmentation priors, generating comprehensive visual features. (ii) A Causal Embedding Module that explicitly disentangles answer-relevant causal features from non-causal noise via a gating mechanism, mitigating spurious correlations. (iii) A Dynamic Cross-Attention Token Sampler that adaptively compresses visual tokens by importance, reducing computational burden while preserving critical information. Extensive experiments on NWPU-RSVQA demonstrate that CausalVisQA outperforms state-of-the-art models across semantic consistency and factual accuracy metrics, particularly in complex relational reasoning tasks.
KW - benchmark
KW - causal embedding
KW - Remote sensing
KW - visual question answering
UR - https://www.scopus.com/pages/publications/105045721235
U2 - 10.1109/TMM.2026.3714932
DO - 10.1109/TMM.2026.3714932
M3 - 文章
AN - SCOPUS:105045721235
SN - 1520-9210
JO - IEEE Transactions on Multimedia
JF - IEEE Transactions on Multimedia
ER -