activity
20242026
most citedVideo-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

1 citations · 1 across the 5 of their papers we have counts for

collaborators
Showing cs.CVShow all

16 papers · 1 filter

cs.CV2026

SlerpFlow: Spherical Trajectory Correction for Rectified Flow Inversion

Wenbin Duan, Yan Shu, Zhuoyuan Fu +4

Rectified-flow-based diffusion transformers, particularly FLUX, have demonstrated outstanding performance in high-quality image generation. However, achieving fast and accurate inv…

cs.CV2026

Towards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language Models

Zhuoyuan Fu, Zeshang Li, Yiqiong Zhang +5

While Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in 2D medical image understanding, their extension to 3D volumetric imaging remains hindered by…

cs.CV2026

Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation

Zekai Zhang, Jiahao Li, Jie Zhang +18

While text-to-image (T2I) models have achieved remarkable progress, they struggle with real-world requests that are often underspecified, implicit, or dependent on up-to-date knowl…

cs.CV2026

Qwen-Image-2.0-RL Technical Report

Yixian Xu, Kaiyuan Gao, Yuxiang Chen +25

We present Qwen-Image-2.0-RL, a post-training pipeline that applies reinforcement learning from human feedback (RLHF) and on-policy distillation (OPD) to improve both the visual qu…

cs.CV2026

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation

Jie Zhang, Xiaoyue Chen, Anzhe Chen +36

We introduce Qwen-RobotWorld, a language-conditioned video world model for embodied intelligence. With natural language as a unified action interface, it predicts physically ground…

cs.CV2026

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA

Mengzhuo Chen, Yan Shu, Chi Liu +4

We study whether grounded reasoning supervision from abundant 2D medical images can improve 3D medical VQA when both input types are aligned through a common reasoning interface. W…