2 papers
cs.CV2025
GEA: Generation-Enhanced Alignment for Text-to-Image Person Retrieval
Hao Zou, Runqing Zhang, Xue Zhou +1
Text-to-Image Person Retrieval (TIPR) aims to retrieve person images based on natural language descriptions. Although many TIPR methods have achieved promising results, sometimes t…
cs.CV2025
AirCache: Activating Inter-modal Relevancy KV Cache Compression for Efficient Large Vision-Language Model Inference
Kai Huang, Hao Zou, Bochen Wang +3
Recent advancements in Large Visual Language Models (LVLMs) have gained significant attention due to their remarkable reasoning capabilities and proficiency in generalization. Howe…