11 papers
SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration
Yinhao Tang, Youqing Fang, Yanan Sun +6
Scientific research involves complex information-seeking and reasoning workflows across heterogeneous sources. However, existing benchmarks primarily emphasize general-domain retri…
World Models as Group Actions
Zijie Wang, Wei Zhang, Weiming Zhang +4
Video world models have achieved strong visual realism, but this does not ensure that their dynamics are truly governed by actions. In this work, we argue that action faithfulness…
MindCopilot: Towards Formalizing and Evaluating Granular Human-LLM Co-Writing
Youqing Fang, Yinhao Tang, Yanan Sun +8
Recent writing assistants are increasingly shifting from passive, prompt-driven interaction to proactive, suggestion-based completion, which integrates localized continuations into…
ColorBrowserAgent: Complex Long-Horizon Browser Agent with Adaptive Knowledge Evolution
Jihong Wang, Jiamu Zhou, Weiming Zhang +7
With the advancement of vision-language models, web automation has made significant progress. However, deploying autonomous agents in real-world settings remains challenging, prima…
Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering
Chenyu Zhou, Huacan Chai, Wenteng Chen +18
Large language model (LLM) agents are increasingly built less by changing model weights than by reorganizing the runtime around them. Capabilities that earlier systems expected the…
Can LLMs Deobfuscate Binary Code? A Systematic Analysis of Large Language Models into Pseudocode Deobfuscation
Li Hu, Xiuwei Shang, Jieke Shi +6
Deobfuscating binary code remains a fundamental challenge in reverse engineering, as obfuscation is widely used to hinder analysis and conceal program logic. Although large languag…