collaborators

28 papers

cs.CV2026

LogiShot: Logically Coherent Cross-Shot Video Generation

Shuai Guo, Yuhang Yang, Zeyu Zhang +4

Generating cross-shot videos that are logically connected is essential for content creation. Currently, most cross-shot video-generation workflows, such as short-drama production,…

cs.CV2026

EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning

Chengjun Yu, Xuhan Zhu, Chaoqun Du +4

Multimodal large language models (MLLMs) are increasingly considered as a foundation for embodied agents, yet it remains unclear whether they can reliably reason about the long-ter…

cs.CL2026

Back on Track: Aligning Rewards and States for Reasoning in Diffusion Large Language Models

Yawen Shao, Jie Xiao, Kai Zhu +6

Reinforcement learning (RL) holds immense promise for enhancing the reasoning capabilities of diffusion large language models (dLLMs). However, progress is fundamentally constraine…

cs.CV2026

Dual-Pathway Geometry-Aware MLLM for Spatial Intelligence

Yufei Zheng, Xuhan Zhu, Zide Liu +9

Spatial understanding of the physical world from 2D visual inputs hinges on two complementary forms of geometric knowledge: holistic 3D structural perception and fine-grained metri…

cs.CV2026

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model

Chenfeng Wang, Wei He, Xuhan Zhu +10

In language reasoning, longer chains of thought consistently yield better performance, which naturally suggests that visual latent reasoning may likewise benefit from longer latent…

cs.CV2026

VMAD: Visual-enhanced Multimodal Large Language Model for Zero-Shot Anomaly Detection

Huilin Deng, Hongchen Luo, Wei Zhai +2

Zero-shot anomaly detection (ZSAD) recognizes and localizes anomalies in previously unseen objects by establishing feature mapping between textual prompts and inspection images, de…