Abstract
Remote sensing visual question answering has emerged as a critical task bridging computer vision and natural language processing in geospatial understanding, yet it faces significant limitations in supporting complex reasoning and multilingual interaction. Existing datasets are predominantly English-only, focus on low-level tasks, and lack high-resolution imagery, restricting real-world applicability. To address these gaps, we introduce NWPU-RSVQA, the first large-scale bilingual dataset for high-order relational reasoning in RSVQA. It comprises 10,718 high-resolution (1,024-2,048px) global images with 107,180 human-verified QA pairs, enabling reasoning across attribute-relation-scene levels. Despite these advances, existing RSVQA methods struggle to perform high-order relational reasoning due to three fundamental limitations: (i) susceptibility to spurious correlations between visual patterns and answers, (ii) inefficient processing of high-resolution visual tokens, and (iii) underutilization of semantic priors. To address these issues, a causality-aware framework with three key innovations: (i) A dual-path vision encoder integrating ViT for global semantics and SAM for segmentation priors, generating comprehensive visual features. (ii) A Causal Embedding Module that explicitly disentangles answer-relevant causal features from non-causal noise via a gating mechanism, mitigating spurious correlations. (iii) A Dynamic Cross-Attention Token Sampler that adaptively compresses visual tokens by importance, reducing computational burden while preserving critical information. Extensive experiments on NWPU-RSVQA demonstrate that CausalVisQA outperforms state-of-the-art models across semantic consistency and factual accuracy metrics, particularly in complex relational reasoning tasks.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Multimedia |
| DOIs | |
| State | Accepted/In press - 2026 |
Keywords
- Remote sensing
- benchmark
- causal embedding
- visual question answering
Fingerprint
Dive into the research topics of 'Causal-Aware Feature Refinement with Dual Encoders for Remote Sensing Visual Question Answering'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver