activity
20242026
collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning

Chenghao Li, Fusheng Hao, Xikai Zhang +5

Multimodal large language models via reinforcement learning (RL) have demonstrated remarkable capabilities in complex visual reasoning tasks, yet they remain limited in long-horizo…

cs.CV2026

AutoCut: End-to-end advertisement video editing based on multimodal discretization and controllable generation

Milton Zhou, Sizhong Qin, Yongzhi Li +2

Short-form videos have become a primary medium for digital advertising, requiring scalable and efficient content creation. However, current workflows and AI tools remain disjoint a…

cs.CV2026

LongCat-Next: Lexicalizing Modalities as Discrete Tokens

Meituan LongCat Team, Bin Xiao, Chao Wang +86

The prevailing Next-Token Prediction (NTP) paradigm has driven the success of large language models through discrete autoregressive modeling. However, contemporary multimodal syste…

cs.CV2026

Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal Conditioning

Zhengjian Yao, Yongzhi Li, Xinyuan Gao +3

We present "Narrative Weaver", a novel framework that addresses a fundamental challenge in generative AI: achieving multi-modal controllable, long-range, and consistent visual cont…

cs.CV2024

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy

Te Yang, Jian Jia, Xiangyu Zhu +9

Large Language Models (LLMs) have strong instruction-following capability to interpret and execute tasks as directed by human commands. Multimodal Large Language Models (MLLMs) hav…

cs.CV2024

Spatiotemporal Graph Guided Multi-modal Network for Livestreaming Product Retrieval

Xiaowan Hu, Yiyi Chen, Yan Li +5

With the rapid expansion of e-commerce, more consumers have become accustomed to making purchases via livestreaming. Accurately identifying the products being sold by salespeople,…