works on

From the 1 of 13 linked papers with an AI index.

collaborators

13 papers

cs.CV2026

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

Zichao Lin, Yifeng Xie, Bowen Qu +30

We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmar…

cs.RO2026

Towards Predictive, Aligned, and Scalable Robot Learning

Peijun Tang, Shangjin Xie, Baifu Huang +6

The paper introduces Lumo-2, a latent world-action model that reasons about future physical dynamics in a shared latent space to generate robot actions, using a multi‑stage alignme…

cs.CL2026

Kimi K2.5: Visual Agentic Intelligence

Kimi Team, Tongtong Bai, Yifan Bai +339

We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. K2.5 emphasizes the joint optimization of text and vision so that…

cs.CV2026

WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models

Runjie Zhou, Youbo Shao, Haoyu Lu +16

We introduce WorldVQA, a benchmark designed to evaluate the atomic visual world knowledge of Multimodal Large Language Models (MLLMs). Unlike current evaluations, which often confl…

cs.CV2026

Towards Pixel-Level VLM Perception via Simple Points Prediction

Tianhui Song, Haoyu Lu, Hao Yang +8

We present SimpleSeg, a strikingly simple yet highly effective approach to endow Multimodal Large Language Models (MLLMs) with native pixel-level perception. Our method reframes se…

cs.CL2025

Operation Veja: Fixing Fundamental Concepts Missing from Modern Roleplaying Training Paradigms

Yueze Liu, Ajay Nagi Reddy Kumdam, Ronit Kanjilal +2

Modern roleplaying models are increasingly sophisticated, yet they consistently struggle to capture the essence of believable, engaging characters. We argue this failure stems from…