4 papers
Show-Harness: Just a VLM Agent Can Play Robots
Yanzhe Chen, Zechen Bai, Zhijun Cao +7
Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harne…
Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning
Wenzheng Zeng, Siyi Jiao, Chen Gao +2
Dense video captioning aims to generate temporally grounded descriptions of video events, benefiting both event-level video understanding and generation. In this domain, autoregres…
Factorized Learning for Temporally Grounded Video-Language Models
Wenzheng Zeng, Difei Gao, Mike Zheng Shou +1
Recent video-language models have shown great potential for video understanding, but still struggle with accurate temporal grounding for event-level perception. We observe that two…
SlideTailor: Personalized Presentation Slide Generation for Scientific Papers
Wenzheng Zeng, Mingyu Ouyang, Langyuan Cui +1
Automatic presentation slide generation can greatly streamline content creation. However, since preferences of each user may vary, existing under-specified formulations often lead…