4 papers
ManimAgent: Self-Evolving Multimodal Agents for Visual Education
Wenjia Jiang, Zongyuan Cai, Yuanhang Shao +7
Multi-round reflection lets agents built on large language models recover from failures within a single task, but each task remains an isolated episode: lessons learned across many…
Hard to Read, Easy to Jailbreak: How Visual Degradation Bypasses MLLM Safety Alignment
Zhixue Song, Boyan Han, Yiwei Wang +1
Recent advancements in visual context compression enable MLLMs to process ultra-long contexts efficiently by rendering text into images. However, we identify a critical vulnerabili…
AppAgent-Claw: CLI Is All You Need for GUI Automation
Zhixue Song, Zhiheng Zhang, Yi Song +1
The OpenClaw platform provides a practical foundation for automation through its skill-oriented architecture, organizing external capabilities into lightweight, reusable components…
What can LLM tell us about cities?
Zhuoheng Li, Yaochen Wang, Zhixue Song +4
This study explores the capabilities of large language models (LLMs) in providing knowledge about cities and regions on a global scale. We employ two methods: directly querying the…