collaborators

7 papers

cs.CV2026

Enhancing Gaze Reasoning in Vision Foundation Models for Gaze Following

Shijing Wang, Yaping Huang, Chaoqun Cui +4

Gaze following requires both scene understanding and gaze reasoning to localize the gaze target of an in-scene person. Recently, vision foundation models (VFMs) have demonstrated s…

cs.CV2026

CineSRD: Leveraging Visual, Acoustic, and Linguistic Cues for Open-World Visual Media Speaker Diarization

Liangbin Huang, Xiaohua Liao, Chaoqun Cui +4

Traditional speaker diarization systems have primarily focused on constrained scenarios such as meetings and interviews, where the number of speakers is limited and acoustic condit…

cs.CL2026

From Utterance to Vividity: Training Expressive Subtitle Translation LLM via Adaptive Local Preference Optimization

Chaoqun Cui, Shijing Wang, Liangbin Huang +4

The rapid development of Large Language Models (LLMs) has significantly enhanced the general capabilities of machine translation. However, as application scenarios become more comp…

cs.CL2026

Hermes the Polyglot: A Unified Framework to Enhance Expressiveness for Multimodal Interlingual Subtitling

Chaoqun Cui, Shijing Wang, Liangbin Huang +4

Interlingual subtitling, which translates subtitles of visual media into a target language, is essential for entertainment localization but has not yet been explored in machine tra…

cs.RO2026

Agentic Reward Modeling: Verifying GUI Agent via Progressive Trajectory-Grounded Interaction

Chaoqun Cui, Jing Huang, Shijing Wang +3

Reinforcement learning with verifiable rewards (RLVR) provides a promising pathway for continuously advancing GUI agents, yet existing reward modeling paradigms face complementary…

cs.CV2025

VL4Gaze: Unleashing Vision-Language Models for Gaze Following

Shijing Wang, Chaoqun Cui, Yaping Huang +2

Human gaze provides essential cues for interpreting attention, intention, and social interaction in visual scenes, yet gaze understanding remains largely unexplored in current visi…