6 papers
PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails
Mingyang Song, Luxin Xu, Haoyu Sun +3
Image guardrails are typically trained and evaluated under a fixed safety policy, implicitly treating safety as an intrinsic property of an image. Real deployments are different: t…
SubtleMemory: A Benchmark for Fine-Grained Relational Memory Discrimination in Long-Horizon AI Agents
Wenxuan Wang, Haoyu Sun, Fukuan Hou +4
Persistent AI assistants, such as OpenClaw, accumulate large collections of related memories over long-term interactions. As these memories grow, they may reinforce one another, di…
Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents
Haoyu Sun, Wenxuan Wang, Mingyang Song +5
Planning is central to LLM agents: before acting, an agent must decompose goals, select tools, reason over constraints, and decide when a task is infeasible. Yet existing agent eva…
AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning
Mingyang Song, Haoyu Sun, Jiawei Gu +4
When humans face problems beyond their immediate capabilities, they rely on tools, providing a promising paradigm for improving visual reasoning in multimodal large language models…
FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow
Haoyu Sun, Huichen Will Wang, Jiawei Gu +2
Front-end engineering involves a complex workflow where engineers conceptualize designs, translate them into code, and iteratively refine the implementation. While recent benchmark…
VarGes: Improving Variation in Co-Speech 3D Gesture Generation via StyleCLIPS
Ming Meng, Ke Mu, Yonggui Zhu +4
Generating expressive and diverse human gestures from audio is crucial in fields like human-computer interaction, virtual reality, and animation. Though existing methods have achie…