From the 1 of 9 linked papers with an AI index.
9 papers
VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System
Haodong Li, Tianfei Ren, Xiaoxiao Ma +25
The paper presents VideoCoCo, a system that generates physically consistent videos by having a coding agent produce executable Blender code that defines the scene and its dynamics,…
OmniOPSD: Rationale-Privileged On-Policy Self-Distillation for Affective Computing
Zebang Cheng, Shuimu Chen, Boxue Yang +7
Reinforcement learning for multimodal large language models (MLLMs) is often hindered by severe reward sparsity in complex reasoning tasks. This challenge is particularly pronounce…
4DVLT: Dynamic Scene Understanding with Worldline-Centered Vision-Language Tracking
Chaoyue Li, Boxue Yang, Shengyao Zhou +3
4D dynamic scene understanding requires grounding language to a persistent worldline that binds identity, metric 3D motion, and synchronized multi-view 2D projections. Existing par…
EvoStreaming: Your Offline Video Model Is a Natively Streaming Assistant
Zichen Wen, Boxue Yang, Junlong Ke +5
Streaming video understanding demands more than watching longer videos: assistants must decide when to speak in real time, balancing responsiveness against verbosity. Yet most vide…
ASPECT: Node-Level Adaptive Spectral Fusion for Graph Contrastive Learning
Zhuolong Li, Boxue Yang, Haopeng Chen
Spectral graph contrastive learning often constructs low- and high-frequency views to capture complementary graph signals, but these views are commonly combined by graph-level or n…
Flash-Unified: A Training-Free and Task-Aware Acceleration Framework for Native Unified Models
Junlong Ke, Zichen Wen, Boxue Yang +6
Native unified multimodal models, which integrate both generative and understanding capabilities, face substantial computational overhead that hinders their real-world deployment.…