From the 1 of 7 linked papers with an AI index.
7 papers
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
Yuhan Zhu, Changlian Ma, Xiangyu Zeng +12
Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when the supporting evidence occurs. We study generalist video temporal gro…
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
Xinhao Li, Yuhan Zhu, Xiangyu Zeng +24
VideoChat3 is a fully open, 4B-parameter video-centric multimodal large language model that combines an efficient Inflated 3D Vision Transformer and adaptive frame resolution with…
FlowErase-RL: Rethinking Concept Erasure as Reward Optimization in Flow Matching Models
Yi Sun, Zhiqi Zhang, Xinhao Zhong +5
Recent advances in flow matching models have significantly improved text-to-image generation quality, but also introduce growing safety risks due to the generation of harmful or un…
SimInsert: Seamless Video Object Insertion via Regional Sparse Attention Fusion
Xinyu Chen, Yuyi Qian, Jiang Lin +9
Video object insertion requires ensuring spatio-temporal coherence and interactive realism, extending far beyond simple content placement. However, current approaches are often hin…
Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning
Xiangyu Zeng, Zhiqiu Zhang, Yuhan Zhu +12
Existing multimodal large language models for long-video understanding predominantly rely on uniform sampling and single-turn inference, limiting their ability to identify sparse y…
FreeControl: Efficient, Training-Free Structural Control via One-Step Attention Extraction
Jiang Lin, Xinyu Chen, Song Wu +7
Controlling the spatial and semantic structure of diffusion-generated images remains a challenge. Existing methods like ControlNet rely on handcrafted condition maps and retraining…