4 papers
Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding
Biao Tang, Xu Chen, Shuxiang Gou +3
Long-video understanding remains challenging for multimodal large language models, because temporally extended videos often contain thousands of frames and are therefore expensive…
FineViT: Progressively Unlocking Fine-Grained Perception with Dense Recaptions
Peisen Zhao, Xiaopeng Zhang, Mingxing Xu +10
While Multimodal Large Language Models (MLLMs) have experienced rapid advancements, their visual encoders frequently remain a performance bottleneck. Conventional CLIP-based encode…
DreamStory: Open-Domain Story Visualization by LLM-Guided Multi-Subject Consistent Diffusion
Huiguo He, Huan Yang, Zixi Tuo +7
Story visualization aims to create visually compelling images or videos corresponding to textual narratives. Despite recent advances in diffusion models yielding promising results,…
Fleximo: Towards Flexible Text-to-Human Motion Video Generation
Yuhang Zhang, Yuan Zhou, Zeyu Liu +4
Current methods for generating human motion videos rely on extracting pose sequences from reference videos, which restricts flexibility and control. Additionally, due to the limita…