17 papers · 1 filter
VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing
Ziyun Zeng, Zixuan Wang, Yongsheng Yu +2
Evaluating generated videos remains challenging because existing benchmarks rely on fixed evaluation content, cover only a subset of generation and editing settings, and provide li…
MemoBench: Benchmarking World Modeling in Dynamically Changing Environments
Haoyu Chen, Kaichen Zhou, Hang Hua +11
Video generation models aspire to simulate dynamic environments, and several benchmarks now evaluate memory consistency across frames. However, most assess consistency only while t…
Agent Skills Should Go Beyond Text: The Case for Visual Skills
Binxiao Xu, Ruichuan An, Bocheng Zou +1
Reusable skills are a key mechanism for extending agent capabilities, allowing agents to accumulate experience and solve increasingly complex tasks. Yet most existing skill-learnin…
Aurora: Unified Video Editing with a Tool-Using Agent
Yongsheng Yu, Ziyun Zeng, Zhiyuan Xiao +4
Recent video editing models have converged on a unified conditioning design: a single diffusion transformer jointly consumes text, source video, and reference images, and one set o…
MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
Ziyun Zeng, Hang Hua, Bocheng Zou +3
Recent GUI agents have made substantial progress in visual grounding and action prediction, yet they remain brittle in long-horizon tasks that require maintaining task state across…
Visual Aesthetic Benchmark: Can Frontier Models Judge Beauty?
Yichen Feng, Yuetai Li, Chunjiang Liu +14
Multimodal large language models (MLLMs) are now routinely deployed for visual understanding, generation, and curation. A substantial fraction of these applications require an expl…