1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CV2025
Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment
Zhuoxuan Cai, Jian Zhang, Xinbin Yuan +7
Recent studies demonstrate that multimodal large language models (MLLMs) can proficiently evaluate visual quality through interpretable assessments. However, existing approaches ty…
cs.AI2025★ 1 cited
Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning
Xinbin Yuan, Jian Zhang, Kaixin Li +8
Graphical User Interface (GUI) agents have made substantial strides in understanding and executing user instructions across diverse platforms. Yet, grounding these instructions to…
cs.CV2024
CPA: Camera-pose-awareness Diffusion Transformer for Video Generation
Yuelei Wang, Jian Zhang, Pengtao Jiang +3
Despite the significant advancements made by Diffusion Transformer (DiT)-based methods in video generation, there remains a notable gap with controllable camera pose perspectives.…