4 papers
ReUnit: Multi-Granularity Visual Unitization for Long Video Understanding
Biao Tang, Xu Chen, Shuxiang Gou +4
Long-video understanding is constrained by the limited visual input capacity of video multimodal large language models (Video-MLLMs). Existing methods mainly optimize which content…
FineViT: Progressively Unlocking Fine-Grained Perception with Dense Recaptions
Peisen Zhao, Xiaopeng Zhang, Mingxing Xu +10
While Multimodal Large Language Models (MLLMs) have experienced rapid advancements, their visual encoders frequently remain a performance bottleneck. Conventional CLIP-based encode…
Fleximo: Towards Flexible Text-to-Human Motion Video Generation
Yuhang Zhang, Yuan Zhou, Zeyu Liu +4
Current methods for generating human motion videos rely on extracting pose sequences from reference videos, which restricts flexibility and control. Additionally, due to the limita…
DreamStory: Open-Domain Story Visualization by LLM-Guided Multi-Subject Consistent Diffusion
Huiguo He, Huan Yang, Zixi Tuo +7
Story visualization aims to create visually compelling images or videos corresponding to textual narratives. Despite recent advances in diffusion models yielding promising results,…