14 papers
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill
Zhuoyang Qian, Biao Wu, Yiran Wang +6
Turning a research idea into a complete paper requires more than text generation: the system must retrieve literature, design and execute experiments, revise claims according to ev…
Agent Safety Should Be a Runtime Contract
Albus W. Ng, Yi Han, Jusheng Zhang +1
The dominant paradigm treats AI safety as a property to be instilled during model training via RLHF, DPO, or Constitutional AI. We argue this is structurally insufficient for auton…
What Memory Do GUI Agents Really Need? From Passive Records to Active Task-Driving States
Chen Liu, Ling Chen, Hanzhang Zhou +7
Mobile GUI agents increasingly face long-horizon tasks that require reading, updating, and reusing task-relevant data across pages and applications. Existing methods treat memory l…
Towards Recursive Self-Evolving Agentic Literature Retrieval
Yuwen Du, Tian Jin, Jing Kang +8
Scientific literature retrieval must understand complex search intents while preserving source authenticity. Traditional keyword and embedding-based systems return authentic source…
PaperJury: Due-Process Review for Bounded LaTeX Revision
Yiran Wang, Ruixuan An, Biao Wu +1
Pre-submission hardening of human-authored LaTeX computer science papers differs from drafting assistance because it requires adversarial whole-paper review, explicit no-fix outcom…
MCP-Persona: Benchmarking LLM Agents on Real-World Personal Applications via Environment Simulation
Wenhao Wang, Peizhi Niu, Gongyi Zou +9
The Model Context Protocol (MCP) has emerged as a transformative standard for connecting large language models (LLMs) with external data sources and tools, and has been rapidly ado…