5 papers
Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models
Yolo Y. Tang, Jing Bi, Pinxin Liu +24
Video understanding represents the most challenging frontier in computer vision, requiring models to reason about complex spatiotemporal relationships, long-term dependencies, and…
The Photographer Eye: Teaching Multimodal Large Language Models to Understand Image Aesthetics like Photographers
Daiqing Qi, Handong Zhao, Jing Shi +5
While editing directly from life, photographers have found it too difficult to see simultaneously both the blue and the sky. Photographer and curator, Szarkowski insightfully revea…
ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL
Yu Zhang, Yunqi Li, Yifan Yang +7
Although chain-of-thought reasoning and reinforcement learning (RL) have driven breakthroughs in NLP, their integration into generative vision models remains underexplored. We intr…
VRMDiff: Text-Guided Video Referring Matting Generation of Diffusion
Lehan Yang, Jincen Song, Tianlong Wang +4
We propose a new task, video referring matting, which obtains the alpha matte of a specified instance by inputting a referring caption. We treat the dense prediction task of mattin…
Less is More: Data-Efficient Complex Question Answering over Knowledge Bases
Yuncheng Hua, Yuan-Fang Li, Guilin Qi +3
Question answering is an effective method for obtaining information from knowledge bases (KB). In this paper, we propose the Neural-Symbolic Complex Question Answering (NS-CQA) mod…