collaborators

9 papers

cs.CV2026

DeepEyesV2: Toward Agentic Multimodal Model

Jack Hong, Chenxiao Zhao, ChengLin Zhu +3

Agentic multimodal models should not only comprehend text and images, but also actively invoke external tools, such as code execution environments and web search, and integrate the…

cs.SD2025

Semantic Codebooks as Effective Priors for Neural Speech Compression

Liuyang Bai, Weiyi Lu, Li Guo

Speech codecs are traditionally optimized for waveform fidelity, allocating bits to preserve acoustic detail even when much of it can be inferred from linguistic structure. This le…

cs.AI2025

PEAR: Phase Entropy Aware Reward for Efficient Reasoning

Chen Huang, Wei Lu, Wenxuan Zhang

Large Reasoning Models (LRMs) have achieved impressive performance on complex reasoning tasks by generating detailed chain-of-thought (CoT) explanations. However, these responses a…

cs.CL2025

Through the Valley: Path to Effective Long CoT Training for Small Language Models

Renjie Luo, Jiaxi Li, Chen Huang +1

Long chain-of-thought (CoT) supervision has become a common strategy to enhance reasoning in language models. While effective for large models, we identify a phenomenon we call Lon…

cs.CV2025

MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources

Sicong Leng, Jing Wang, Jiaxi Li +12

Large multimodal reasoning models have achieved rapid progress, but their advancement is constrained by two major limitations: the absence of open, large-scale, high-quality long c…

cs.CV2025

Vidi: Large Multimodal Models for Video Understanding and Editing

Vidi Team, Celong Liu, Chia-Wen Kuo +20

Humans naturally share information with those they are connected to, and video has become one of the dominant mediums for communication and expression on the Internet. To support t…