5 papers
VideoWeaver: Evaluating and Evolving Skills for Agentic Long Video Generation
Jianhui Wei, Jie Tan, Hengchuan Zhu +6
Recent agent frameworks such as Claude Code, Codex, and OpenClaw are strong at tool use and orchestration, but whether they can handle long video generation, a long-horizon multimo…
How Far Are Video Models from True Multimodal Reasoning?
Xiaotian Zhang, Jianhui Wei, Yuan Wang +9
Despite remarkable progress toward general-purpose video models, a critical question remains unanswered: how far are these models from achieving true multimodal reasoning? Existing…
Action Draft and Verify: A Self-Verifying Framework for Vision-Language-Action Model
Chen Zhao, Zhuoran Wang, Haoyang Li +6
Vision-Language-Action (VLA) models have recently demonstrated strong performance across embodied tasks. Modern VLAs commonly employ diffusion action experts to efficiently generat…
Boosting Multi-modal Keyphrase Prediction with Dynamic Chain-of-Thought in Vision-Language Models
Qihang Ma, Shengyu Li, Jie Tang +5
Multi-modal keyphrase prediction (MMKP) aims to advance beyond text-only methods by incorporating multiple modalities of input information to produce a set of conclusive phrases. T…
Conf-Profile: A Confidence-Driven Reasoning Paradigm for Label-Free User Profiling
Yingxin Li, Jianbo Zhao, Xueyu Ren +8
User profiling, as a core technique for user understanding, aims to infer structural attributes from user information. Large Language Models (LLMs) provide a promising avenue for u…