Skip to main navigation Skip to search Skip to main content

Causal-Aware Feature Refinement with Dual Encoders for Remote Sensing Visual Question Answering

  • Zhigang Yang
  • , Huiguang Yao
  • , Linmao Tian
  • , Shanji Liu
  • , Weiping Ni
  • , Junzheng Wu
  • , Qiang Li
  • , Qi Wang
  • Northwestern Polytechnical University Xian
  • Northwest Institute of Nuclear Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Remote sensing visual question answering has emerged as a critical task bridging computer vision and natural language processing in geospatial understanding, yet it faces significant limitations in supporting complex reasoning and multilingual interaction. Existing datasets are predominantly English-only, focus on low-level tasks, and lack high-resolution imagery, restricting real-world applicability. To address these gaps, we introduce NWPU-RSVQA, the first large-scale bilingual dataset for high-order relational reasoning in RSVQA. It comprises 10,718 high-resolution (1,024-2,048px) global images with 107,180 human-verified QA pairs, enabling reasoning across attribute-relation-scene levels. Despite these advances, existing RSVQA methods struggle to perform high-order relational reasoning due to three fundamental limitations: (i) susceptibility to spurious correlations between visual patterns and answers, (ii) inefficient processing of high-resolution visual tokens, and (iii) underutilization of semantic priors. To address these issues, a causality-aware framework with three key innovations: (i) A dual-path vision encoder integrating ViT for global semantics and SAM for segmentation priors, generating comprehensive visual features. (ii) A Causal Embedding Module that explicitly disentangles answer-relevant causal features from non-causal noise via a gating mechanism, mitigating spurious correlations. (iii) A Dynamic Cross-Attention Token Sampler that adaptively compresses visual tokens by importance, reducing computational burden while preserving critical information. Extensive experiments on NWPU-RSVQA demonstrate that CausalVisQA outperforms state-of-the-art models across semantic consistency and factual accuracy metrics, particularly in complex relational reasoning tasks.

Original languageEnglish
JournalIEEE Transactions on Multimedia
DOIs
StateAccepted/In press - 2026

Keywords

  • Remote sensing
  • benchmark
  • causal embedding
  • visual question answering

Fingerprint

Dive into the research topics of 'Causal-Aware Feature Refinement with Dual Encoders for Remote Sensing Visual Question Answering'. Together they form a unique fingerprint.

Cite this