activity
20242026
collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

ALADIN:Attribute-Language Distillation Network for Person Re-Identification

Wang Zhou, Boran Duan, Haojun Ai +2

Recent vision-language models such as CLIP provide strong cross-modal alignment, but current CLIP-guided ReID pipelines rely on global features and fixed prompts. This limits their…

cs.CV2025

Structure-Aware Feature Rectification with Region Adjacency Graphs for Training-Free Open-Vocabulary Semantic Segmentation

Qiming Huang, Hao Ai, Jianbo Jiao

Benefiting from the inductive biases learned from large-scale datasets, open-vocabulary semantic segmentation (OVSS) leverages the power of vision-language models, such as CLIP, to…

cs.CV2025

Dynamic Pyramid Network for Efficient Multimodal Large Language Model

Hao Ai, Kunyi Wang, Zezhou Wang +7

Multimodal large language models (MLLMs) have demonstrated impressive performance in various vision-language (VL) tasks, but their expensive computations still limit the real-world…

cs.CV2024

InstantIR: Blind Image Restoration with Instant Generative Reference

Jen-Yuan Huang, Haofan Wang, Qixun Wang +4

Handling test-time unknown degradation is the major challenge in Blind Image Restoration (BIR), necessitating high model generalization. An effective strategy is to incorporate pri…

cs.CV2024

CSGO: Content-Style Composition in Text-to-Image Generation

Peng Xing, Haofan Wang, Yanpeng Sun +5

The diffusion model has shown exceptional capabilities in controlled image generation, which has further fueled interest in image style transfer. Existing works mainly focus on tra…