collaborators

10 papers

cs.AI2026

SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents

Yang Wan, Zhenhao Zhang, Jierui Wang +1

Deciding whether a trajectory actually fulfills its instruction governs how we measure computer-use agents on long-horizon graphical-user-interface tasks and how we train them with…

cs.AI2026

VISTA: View-Consistent Self-Verified Training for GUI Grounding

Xinyu Qiu, Yunzhu Zhang, Heng Jia +3

When applying Group Relative Policy Optimization (GRPO) for GUI Grounding, rollouts are sampled from a single screenshot view; groups often become either all failures on difficult…

physics.chem-ph2026

A Fixed-Point Neural Operator for Size- and Functional-Transferable Hamiltonian Prediction

Yunhong Lou, Xihang Yue, Xinran Wei +2

Predicting the Kohn-Sham Hamiltonian with machine learning can accelerate density functional theory while retaining access to molecular orbitals, energy levels, and electronic-stru…

cs.AI2026

Mitigating Conversational Inertia in Multi-Turn Agents

Yang Wan, Zheng Cao, Zhenhao Zhang +4

Large language models excel as few-shot learners when provided with appropriate demonstrations, yet this strength becomes problematic in multiturn agent scenarios, where LLMs erron…

cs.CV2026

GPD: Guided Progressive Distillation for Fast and High-Quality Video Generation

Xiao Liang, Yunzhu Zhang, Linchao Zhu

Diffusion models have achieved remarkable success in video generation; however, the high computational cost of the denoising process remains a major bottleneck. Existing approaches…

cs.CV2026

MTC-VAE: Multi-Level Temporal Compression with Content Awareness

Yubo Dong, Linchao Zhu

Latent Video Diffusion Models (LVDMs) rely on Variational Autoencoders (VAEs) to compress videos into compact latent representations. For continuous Variational Autoencoders (VAEs)…