7 papers · 1 filter
IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning
Chenghao Li, Fusheng Hao, Xikai Zhang +5
Multimodal large language models via reinforcement learning (RL) have demonstrated remarkable capabilities in complex visual reasoning tasks, yet they remain limited in long-horizo…
AutoCut: End-to-end advertisement video editing based on multimodal discretization and controllable generation
Milton Zhou, Sizhong Qin, Yongzhi Li +2
Short-form videos have become a primary medium for digital advertising, requiring scalable and efficient content creation. However, current workflows and AI tools remain disjoint a…
LongCat-Next: Lexicalizing Modalities as Discrete Tokens
Meituan LongCat Team, Bin Xiao, Chao Wang +86
The prevailing Next-Token Prediction (NTP) paradigm has driven the success of large language models through discrete autoregressive modeling. However, contemporary multimodal syste…
Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal Conditioning
Zhengjian Yao, Yongzhi Li, Xinyuan Gao +3
We present "Narrative Weaver", a novel framework that addresses a fundamental challenge in generative AI: achieving multi-modal controllable, long-range, and consistent visual cont…
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy
Te Yang, Jian Jia, Xiangyu Zhu +9
Large Language Models (LLMs) have strong instruction-following capability to interpret and execute tasks as directed by human commands. Multimodal Large Language Models (MLLMs) hav…
Spatiotemporal Graph Guided Multi-modal Network for Livestreaming Product Retrieval
Xiaowan Hu, Yiyi Chen, Yan Li +5
With the rapid expansion of e-commerce, more consumers have become accustomed to making purchases via livestreaming. Accurately identifying the products being sold by salespeople,…