2 citations · 2 across the 8 of their papers we have counts for
6 papers · 1 filter
HFS: Holistic Query-Aware Frame Selection for Efficient Video Understanding
Yiqing Yang, Yun Li, Daiqing Qi +5
Key frame selection is essentially a set-level optimization problem: the quality of the selected subset depends on the interactions among frames, rather than the score of any singl…
Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models
Yolo Y. Tang, Jing Bi, Pinxin Liu +24
Video understanding represents the most challenging frontier in computer vision, requiring models to reason about complex spatiotemporal relationships, long-term dependencies, and…
The Photographer Eye: Teaching Multimodal Large Language Models to Understand Image Aesthetics like Photographers
Daiqing Qi, Handong Zhao, Jing Shi +5
While editing directly from life, photographers have found it too difficult to see simultaneously both the blue and the sky. Photographer and curator, Szarkowski insightfully revea…
ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL
Yu Zhang, Yunqi Li, Yifan Yang +7
Although chain-of-thought reasoning and reinforcement learning (RL) have driven breakthroughs in NLP, their integration into generative vision models remains underexplored. We intr…
VRMDiff: Text-Guided Video Referring Matting Generation of Diffusion
Lehan Yang, Jincen Song, Tianlong Wang +4
We propose a new task, video referring matting, which obtains the alpha matte of a specified instance by inputting a referring caption. We treat the dense prediction task of mattin…
Reminding Multimodal Large Language Models of Object-aware Knowledge with Retrieved Tags
Daiqing Qi, Handong Zhao, Zijun Wei +1
Despite recent advances in the general visual instruction-following ability of Multimodal Large Language Models (MLLMs), they still struggle with critical problems when required to…