collaborators

5 papers

cs.CV2026

OmniReasoner: Thinking with Long Audio-Video via Native Tool Use

Yu Chen, Caorui Li, Ziyu Xiong +8

Long audio-video reasoning is difficult for omnimodal LLMs because the decisive evidence is often sparse, cross-modal, and too expensive to preserve with uniformly high-fidelity in…

cs.CV2026

Claim-Level Rubric Rewards for Video Caption Reinforcement Learning

Mingqi Gao, Hongyuan Dong, Yifei Chen +6

In this paper, we introduce Claim-Level Rubric Rewards (CuRe), a structured reward framework designed to address the reward-design bottleneck in reinforcement learning for dense vi…

cs.AI2026

OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs

Caorui Li, Yu Chen, Yiyan Ji +40

Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively eva…

cs.CV2025

PixCLIP: Achieving Fine-grained Visual Language Understanding via Any-granularity Pixel-Text Alignment Learning

Yicheng Xiao, Yu Chen, Haoxuan Ma +5

While the Contrastive Language-Image Pretraining(CLIP) model has achieved remarkable success in a variety of downstream vison language understanding tasks, enhancing its capability…

cs.LG2025

A Multi-scale Representation Learning Framework for Long-Term Time Series Forecasting

Boshi Gao, Qingjian Ni, Fanbo Ju +2

Long-term time series forecasting (LTSF) offers broad utility in practical settings like energy consumption and weather prediction. Accurately predicting long-term changes, however…