Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
EgoSound: Benchmarking Sound Understanding in Egocentric Videos
Bingwen Zhu, Yuqian Fu, Qiaole Dong +6
Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in vision-language understanding. Yet, human perception is inherently multisensory, integrating…
cs.CV2025
HF-VTON: High-Fidelity Virtual Try-On via Consistent Geometric and Semantic Alignment
Ming Meng, Qi Dong, Jiajie Li +5
Virtual try-on technology has become increasingly important in the fashion and retail industries, enabling the generation of high-fidelity garment images that adapt seamlessly to t…
cs.CV2025
Non-autoregressive Sequence-to-Sequence Vision-Language Models
Kunyu Shi, Qi Dong, Luis Goncalves +2
Sequence-to-sequence vision-language models are showing promise, but their applicability is limited by their inference latency due to their autoregressive way of generating predict…