5 papers
Let ViT Speak: Generative Language-Image Pre-training
Yan Fang, Mengcheng Lan, Zilong Huang +7
In this paper, we present \textbf{Gen}erative \textbf{L}anguage-\textbf{I}mage \textbf{P}re-training (GenLIP), a minimalist generative pretraining framework for Vision Transformers…
EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models
Shiqi Huang, Ziyue Wang, Zhongrong Zuo +3
Recent Video Large Language Models (Video-LLMs) have demonstrated strong capabilities in video reasoning through reinforcement learning (RL). However, existing RL pipelines rely he…
TextSculptor: Training and Benchmarking Scene Text Editing
Yiheng Lin, Siyu Jiao, Xiaohan Lan +12
Recent advances in Multimodal Large Language Models (MLLMs) and diffusion-based generative models have substantially improved prompt-driven image editing. However, scene text editi…
CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoning
Qi Song, Honglin Li, Yingchen Yu +6
Recent releases such as o3 highlight human-like "thinking with images" reasoning that combines tool use with stepwise verification, yet most open-source approaches still rely on te…
ThinkGen: Generalized Thinking for Visual Generation
Siyu Jiao, Yiheng Lin, Yujie Zhong +9
Recent progress in Multimodal Large Language Models (MLLMs) demonstrates that Chain-of-Thought (CoT) reasoning enables systematic solutions to complex understanding tasks. However,…