collaborators

16 papers

cs.CV2026

Clinical Cognition Alignment for Gastrointestinal Diagnosis with Multimodal LLMs

Huan Zheng, Yucheng Zhou, Tianyi Yan +6

Multimodal Large Language Models (MLLMs) have demonstrated remarkable potential in medical image analysis. However, their application in gastrointestinal endoscopy is currently hin…

cs.CV2026

World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoning

Yucheng Zhou, Wei Tao, Yiwen Guo +1

World models and multimodal large language models (MLLMs) provide complementary capabilities for predicting future outcomes from static visual observations. World models can genera…

cs.CV2026

CogOmniControl: Reasoning-Driven Controllable Video Generation via Creative Intent Cognition

Hongji Yang, Songlian Li, Yucheng Zhou +4

Recent diffusion models achieve strong photorealism and fluency in video generation, yet remain fragile under abstract, sparse or complex conditions, leading to poor performance in…

cs.CL2026

Compatibility-Aware Dynamic Fine-Tuning for Large Language Models

Yucheng Zhou, Junwei Sheng, Qianning Wang +1

Supervised Fine-Tuning (SFT) is the predominant paradigm for aligning large language models (LLMs), yet it suffers from optimization instability and limited generalization. Recent…

cs.LG2026

Multimodal Large Language Models for Multi-Subject In-Context Image Generation

Yucheng Zhou, Dubing Chen, Huan Zheng +1

Recent advances in text-to-image (T2I) generation have enabled visually coherent image synthesis from descriptions, but generating images containing multiple given subjects remains…

cs.LG2026

Accelerating Training of Autoregressive Video Generation Models via Local Optimization with Representation Continuity

Yucheng Zhou, Jianbing Shen

Autoregressive models have shown superior performance and efficiency in image generation, but remain constrained by high computational costs and prolonged training times in video g…