Abstract
Visible-infrared person re-identification (VI-ReID) aims to match identities across heterogeneous spectra. Current methods often bridge the modality gap via intermediate modalities or specific losses but overlook semantic alignment, particularly for local features. To address this, we propose the CLIP-driven Dual-level Semantic Alignment network (CDSA). CDSA leverages textual information to guide semantic representation learning at both global and local levels. Specifically, we design a Global Semantic Learning Module (GSLM) to capture identity-level semantics, and a Text prototype-based Semantic Refinement Module (TSRM) to extract discriminative local details using learnable text tokens. By integrating dual-level text guidance, our method effectively enhances modality-invariant visual representations. Extensive experiments on benchmark datasets demonstrate the effectiveness of our approach and its competitive performance against state-of-the-art methods.
| Original language | English |
|---|---|
| Pages (from-to) | 1-11 |
| Number of pages | 11 |
| Journal | IEEE MultiMedia |
| Early online date | 6 May 2026 |
| DOIs | |
| Publication status | Published online - 6 May 2026 |
Bibliographical note
Publisher Copyright:© 1994-2012 IEEE.
Keywords
- Contacts
- Circuits and systems
- Protocols
- Communication systems
- Wide area networks
- Pixel
- TV
- Videos
- Computer Networks
- Video equipment
Fingerprint
Dive into the research topics of 'CLIP-Driven Dual-level Semantic Alignment Networks for Visible-Infrared Person Re-Identification'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver