collaborators

10 papers

cs.CV2026

World Reasoning Arena

PAN Team, Qiyue Gao, Kun Zhou +15

World models (WMs) are intended to serve as internal simulators of the real world that enable agents to understand, anticipate, and act upon complex environments. Existing WM bench…

cs.AI2026

Learning Modal-Mixed Chain-of-Thought Reasoning with Latent Embeddings

Yifei Shao, Kun Zhou, Ziming Xu +5

We study how to extend chain-of-thought (CoT) beyond language to better handle multimodal reasoning. While CoT helps LLMs and VLMs articulate intermediate steps, its text-only form…

cs.CV2025

PAN: A World Model for General, Interactable, and Long-Horizon World Simulation

PAN Team, Jiannan Xiang, Yi Gu +31

A world model enables an intelligent agent to imagine, predict, and reason about how the world evolves in response to its actions, and accordingly to plan and strategize. While rec…

cs.AI2025

Auto-scaling Continuous Memory for GUI Agent

Wenyi Wu, Kun Zhou, Ruoxin Yuan +4

We study how to endow GUI agents with scalable memory that help generalize across unfamiliar interfaces and long-horizon tasks. Prior GUI agents compress past trajectories into tex…

cs.CV2025

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation

Yuheng Zha, Kun Zhou, Yujia Wu +7

Despite their success, current training pipelines for reasoning VLMs focus on a limited range of tasks, such as mathematical and logical reasoning. As a result, these models face d…

cs.LG2025

Towards General Continuous Memory for Vision-Language Models

Wenyi Wu, Zixuan Song, Kun Zhou +3

Language models (LMs) and their extension, vision-language models (VLMs), have achieved remarkable performance across various tasks. However, they still struggle with complex reaso…