collaborators

7 papers

astro-ph.HE2026

State-dependent broadband X-ray Timing Reconfiguration in the Changing-look AGN NGC 1566

Yu Tao, Jie Tang, Xuan Wei +1

NGC 1566 has shown dramatic X-ray spectral changes during its recent changing-look outburst, but the evolution of its broadband X-ray timing properties remains poorly constrained.…

cs.CL2026

GLM-OCR Technical Report

Shuaiqi Duan, Yadong Xue, Weihan Wang +20

GLM-OCR is an efficient 0.9B-parameter compact multimodal model designed for real-world document understanding. It combines a 0.4B-parameter CogViT visual encoder with a 0.5B-param…

cs.CV2026

VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation

Jiazheng Xu, Yu Huang, Jiale Cheng +19

Visual generative models have achieved remarkable progress in synthesizing photorealistic images and videos, yet aligning their outputs with human preferences across critical dimen…

cs.CV2025

LVBench: An Extreme Long Video Understanding Benchmark

Weihan Wang, Zehai He, Wenyi Hong +9

Recent progress in multimodal large language models has markedly enhanced the understanding of short videos (typically under one minute), and several evaluation datasets have emerg…

cs.CL2025

AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models

Yuhang Wu, Wenmeng Yu, Yean Cheng +5

Evaluating the alignment capabilities of large Vision-Language Models (VLMs) is essential for determining their effectiveness as helpful assistants. However, existing benchmarks pr…

cs.CV2025

CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Zhuoyi Yang, Jiayan Teng, Wendi Zheng +15

We present CogVideoX, a large-scale text-to-video generation model based on diffusion transformer, which can generate 10-second continuous videos aligned with text prompt, with a f…