3 papers
cs.CV2026
Benchmarking Living-Screen-Native GUI Agents on Short-Video Platforms
Jiashu Yao, Heyan Huang, Daiqing Wu +5
GUI agents today assume a static screen, where the world is frozen between two actions. However, real interfaces such as short-video applications violate this assumption, as their…
cs.CV2025
HomeSafeBench: Benchmarking Embodied Vision-Language Models in Free-Exploration Home Safety Inspection
Siyuan Gao, Jiashu Yao, Haoyu Wen +3
Safety hazards in the home are a leading cause of preventable domestic injuries, motivating an automated inspector that actively explores a home and reports hazards before they cau…
cs.CL2024
ReFF: Reinforcing Format Faithfulness in Language Models across Varied Tasks
Jiashu Yao, Heyan Huang, Zeming Liu +4
Following formatting instructions to generate well-structured content is a fundamental yet often unmet capability for large language models (LLMs). To study this capability, which…