activity
20242026
most citedMiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

1 citations · 1 across the 2 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding

Fuwen Luo, Shengfeng Lou, Chi Chen +9

Video temporal understanding is crucial for multimodal large language models (MLLMs) to reason over events in videos. Despite recent advances in general video understanding, curren…

cs.CV2026

KARL: Knowledge-Aware Reasoning and Reinforcement Learning for Knowledge-Intensive Visual Grounding

Xinyu Ma, Ziyang Ding, Zhicong Luo +6

Knowledge-Intensive Visual Grounding (KVG) requires models to localize objects using fine-grained, domain-specific entity names rather than generic referring expressions. Although…

cs.CV2025

EscapeCraft: A 3D Room Escape Environment for Benchmarking Complex Multimodal Reasoning Ability

Ziyue Wang, Yurui Dong, Fuwen Luo +5

The rapid advancing of Multimodal Large Language Models (MLLMs) has spurred interest in complex multimodal reasoning tasks in the real-world and virtual environment, which require…

cs.CV2025

ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models

Ziyue Wang, Chi Chen, Fuwen Luo +6

Active perception, a crucial human capability, involves setting a goal based on the current understanding of the environment and performing actions to achieve that goal. Despite si…

cs.CV2024

StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Junming Lin, Zheng Fang, Chi Chen +5

The rapid development of Multimodal Large Language Models (MLLMs) has expanded their capabilities from image comprehension to video understanding. However, most of these MLLMs focu…

cs.CV2024

Model Composition for Multimodal Large Language Models

Chi Chen, Yiyang Du, Zheng Fang +8

Recent developments in Multimodal Large Language Models (MLLMs) have shown rapid progress, moving towards the goal of creating versatile MLLMs that understand inputs from various m…