4 papers
SoftSkill: Behavioral Compression for Contextual Adaptation
Xijia Tao, Yihua Teng, Xinyu Fu +6
Agent skills are commonly deployed as natural-language Markdown files that encode answer policies, evidence-use habits, and task procedures. These files are readable and portable,…
MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents
Xijia Tao, Yihua Teng, Xinxing Su +7
Existing multimodal browsing benchmarks often fail to require genuine multimodal reasoning, as many tasks can be solved with text-only heuristics without vision-in-the-loop verific…
DocPuzzle: A Process-Aware Benchmark for Evaluating Realistic Long-Context Reasoning Capabilities
Tianyi Zhuang, Chuqiao Kuang, Xiaoguang Li +4
We present DocPuzzle, a rigorously constructed benchmark for evaluating long-context reasoning capabilities in large language models (LLMs). This benchmark comprises 100 expert-lev…
Android in the Zoo: Chain-of-Action-Thought for GUI Agents
Jiwen Zhang, Jihao Wu, Yihua Teng +5
Large language model (LLM) leads to a surge of autonomous GUI agents for smartphone, which completes a task triggered by natural language through predicting a sequence of actions o…