5 papers
Before Thinking, Learn to Decide: Proactive Routing for Efficient Visual Reasoning
Yinan Zhou, Haokun Lin, Yichen Wu +7
Large multimodal models have achieved strong reasoning on complex visual tasks, but their inference efficiency is often restricted by long chains of thought. A promising solution i…
LLMVA-GEBC: Large Language Model with Video Adapter for Generic Event Boundary Captioning
Yolo Yunlong Tang, Jinrui Zhang, Xiangchen Wang +2
Our winning entry for the CVPR 2023 Generic Event Boundary Captioning (GEBC) competition is detailed in this paper. Unlike conventional video captioning tasks, GEBC demands that th…
Multi-modal Segment Assemblage Network for Ad Video Editing with Importance-Coherence Reward
Yolo Yunlong Tang, Siting Xu, Teng Wang +3
Advertisement video editing aims to automatically edit advertising videos into shorter videos while retaining coherent content and crucial information conveyed by advertisers. It m…
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
Haokun Lin, Teng Wang, Yixiao Ge +6
Pioneering token-based works such as Chameleon and Emu3 have established a foundation for multimodal unification but face challenges of high training computational overhead and lim…
Auto-ABSA: Cross-Domain Aspect Detection and Sentiment Analysis Using Auxiliary Sentences
Teng Wang, Bolun Sun, Yijie Tong
After transformer is proposed, lots of pre-trained language models have been come up with and sentiment analysis (SA) task has been improved. In this paper, we proposed a method th…