29 papers
MoWorld: A Flash World Model
Team Moxin, Deyi Ji, Tianrun Chen +37
The future of World Models depends not only on scaling model capability, but also on scaling practicality and inference efficiency. High-frame-rate inference enables responsive per…
Decoding Multimodal Cues: Unveiling the Implicit Meaning Behind Hateful Videos
Junyu Lu, Deyi Ji, Liqun Liu +9
Hateful videos have become prevalent on online platforms, highlighting an urgent need for effective detection. However, existing studies primarily focus on binary classification an…
CamGeo: Sparse Camera-Conditioned Image-to-Video Generation with 3D Geometry Priors
Xuanyi Liu, Deyi Ji, Liqun Liu +8
Sparse camera-conditioned image-to-video generation presents a pivotal challenge: synthesizing geometrically consistent 3D motion from minimal pose cues. Existing methods, which la…
Self-signals Driven Multi-LLM Debate for Efficient and Accurate Reasoning
Xuhang Chen, Zhifan Song, Deyi Ji +2
Large Language Models (LLMs) have exhibited impressive capabilities across diverse application domains. Recent work has explored Multi-LLM Agent Debate (MAD) as a way to enhance pe…
Claw AI Lab: An Autonomous Multi-Agent Research Team
Fan Wu, Cheng Chen, Zhenshan Tan +12
We present Claw AI Lab, a lab-native autonomous research platform that advances automated research from a hidden prompt-to-paper pipeline into an interactive AI laboratory. Rather…
Video-Zero: Self-Evolution Video Understanding
Ruixu Zhang, Deyi Ji, Lanyun Zhu +4
Self-evolution offers a promising path for improving reasoning models without relying on intensive human annotation. However, extending this paradigm to video understanding remains…