collaborators

6 papers

cs.CV2025

UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist

Zhengyang Liang, Daoan Zhang, Huichi Zhou +8

While specialized AI models excel at isolated video tasks like generation or understanding, real-world applications demand complex, iterative workflows that combine these capabilit…

cs.CV2025

Disc3D: Automatic Curation of High-Quality 3D Dialog Data via Discriminative Object Referring

Siyuan Wei, Chunjie Wang, Xiao Liu +3

3D Multi-modal Large Language Models (MLLMs) still lag behind their 2D peers, largely because large-scale, high-quality 3D scene-dialogue datasets remain scarce. Prior efforts hing…

cs.MA2025

Adaptive Coopetition: Leveraging Coarse Verifier Signals for Resilient Multi-Agent LLM Reasoning

Rui Jerry Huang, Wendy Liu, Anastasia Miin +1

Inference-time computation is a critical yet challenging paradigm for enhancing the reasoning performance of large language models (LLMs). While existing strategies improve reasoni…

cs.AI2025

Video-STR: Reinforcing MLLMs in Video Spatio-Temporal Reasoning with Relation Graph

Wentao Wang, Heqing Zou, Tianze Luo +8

Recent progress in Multimodal Large Language Models (MLLMs) has demonstrated strong semantic understanding capabilities, but struggles to perform precise spatio-temporal understand…

cs.IR2025

OneRec-V2 Technical Report

Guorui Zhou, Hengrui Hu, Hongtao Cheng +72

Recent breakthroughs in generative AI have transformed recommender systems through end-to-end generation. OneRec reformulates recommendation as an autoregressive generation task, a…

cs.IR2025

OneRec Technical Report

Guorui Zhou, Jiaxin Deng, Jinghao Zhang +62

Recommender systems have been widely used in various large-scale user-oriented platforms for many years. However, compared to the rapid developments in the AI community, recommenda…