1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CV2025
Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment
Zhuoxuan Cai, Jian Zhang, Xinbin Yuan +7
Recent studies demonstrate that multimodal large language models (MLLMs) can proficiently evaluate visual quality through interpretable assessments. However, existing approaches ty…
cs.AI2025★ 1 cited
Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning
Xinbin Yuan, Jian Zhang, Kaixin Li +8
Graphical User Interface (GUI) agents have made substantial strides in understanding and executing user instructions across diverse platforms. Yet, grounding these instructions to…
cs.CV2023
ChatAnything: Facetime Chat with LLM-Enhanced Personas
Yilin Zhao, Xinbin Yuan, Shanghua Gao +4
In this technical report, we target generating anthropomorphized personas for LLM-based characters in an online manner, including visual appearance, personality and tones, with onl…