Skip to main navigation Skip to search Skip to main content

Vision–Language Models Empowered Nighttime Object Detection With Consistency Sampler and Hallucination Feature Generator

  • Lihuo He
  • , Junjie Ke
  • , Zhenghao Wang
  • , Jie Li
  • , Kai Zhou
  • , Qi Wang
  • , Xinbo Gao
  • School of Electronic Engineering, Xidian University
  • Tsinghua University
  • The Chinese University of Hong Kong, Shenzhen

Research output: Contribution to journalArticlepeer-review

Abstract

Current object detectors often suffer performance degradation when applied to cross-domain scenarios, particularly under challenging visual conditions such as nighttime scenes. This is primarily due to the I3 problems: Inadequate sampling of instance-level features, Indistinguishable feature representation across domains and Inaccurate generation for identical category participation. To address these challenges, we propose a domain-adaptive detection framework that enables robust generalization across different visual domains without introducing any additional inference overhead. The framework comprises three key components. Specifically, the centerness–category consistency sampler alleviates inadequate sampling by selecting representative instance-level features, while the paired centerness consistency loss enforces alignment between classification and localization. Second, VLM-based orthogonality enhancement leverages frozen vision–language encoders with an orthogonal projection loss to improve cross-domain feature distinguishability. Third, hallucination feature generator synthesizes robust instance-level features for missing categories, ensuring balanced category participation across domains. Extensive experiments on multiple datasets covering various domain adaptation and generalization settings demonstrate that our method consistently outperforms state-of-the-art detectors, achieving up to 5.5 mAP improvement, with particularly strong gains in nighttime adaptation.

Original languageEnglish
Pages (from-to)8345-8360
Number of pages16
JournalIEEE Transactions on Image Processing
Volume34
DOIs
StatePublished - 15 Dec 2025

Keywords

  • Domain adaptive object detection
  • hallucination feature generator
  • vision-language models

Fingerprint

Dive into the research topics of 'Vision–Language Models Empowered Nighttime Object Detection With Consistency Sampler and Hallucination Feature Generator'. Together they form a unique fingerprint.

Cite this