2 papers
cs.CV2025
Action Dubber: Timing Audible Actions via Inflectional Flow
Wenlong Wan, Weiying Zheng, Tianyi Xiang +2
We introduce the task of Audible Action Temporal Localization, which aims to identify the spatio-temporal coordinates of audible movements. Unlike conventional tasks such as action…
cs.LG2025
FedVLMBench: Benchmarking Federated Fine-Tuning of Vision-Language Models
Weiying Zheng, Ziyue Lin, Pengxin Guo +3
Vision-Language Models (VLMs) have demonstrated remarkable capabilities in cross-modal understanding and generation by integrating visual and textual information. While instruction…