collaborators

9 papers

cs.LG2026

CertBind from Multimodal Connectivity to Certifiable Retrieval Decisions

Shuheng Cao, Zhenhao Zhang, Ruiqi Chen +7

Lightweight connectors make frozen multimodal encoders composable at the representation level. Deployment exposes a second problem at the level of task decisions. A connected route…

cs.AI2026

SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents

Yang Wan, Zhenhao Zhang, Jierui Wang +1

Deciding whether a trajectory actually fulfills its instruction governs how we measure computer-use agents on long-horizon graphical-user-interface tasks and how we train them with…

cs.CV2026

Diffusion Models are Open-World Affordance Learners: Leveraging Generative Priors for 3D Affordance Learning

Hanqing Wang, Zhenhao Zhang, Kaiyang Ji +12

3D affordance grounding aims to understand how diverse objects can be manipulated, making it a cornerstone of embodied interaction. However, prior works struggle to generalize to o…

cs.CV2026

Event-VLA: Action-Conditioned Event Fusion for Robust Vision-Language-Action Model

Jiaxin Liu, Xun Xu, Zhenhao Zhang +5

Vision-Language-Action (VLA) models have become an important paradigm of embodied AI. However, existing VLA models typically assume well-lit and stable indoor settings, while real-…

cs.AI2026

Mitigating Conversational Inertia in Multi-Turn Agents

Yang Wan, Zheng Cao, Zhenhao Zhang +4

Large language models excel as few-shot learners when provided with appropriate demonstrations, yet this strength becomes problematic in multiturn agent scenarios, where LLMs erron…

cs.CL2026

Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention Rewards

Ming Li, Pei Chen, Zhenhao Zhang +10

Large Language Models demonstrate strong capabilities in single-turn instruction following but suffer from Lost-in-Conversation (LiC), a degradation in performance as information i…