6 papers · 1 filter
NaturalGAIA: A Verifiable Benchmark and Hierarchical Framework for Long-Horizon GUI Tasks
Zihan Zheng, Tianle Cui, Taoran Wang +4
Despite significant advances in LLM-driven GUI agents, the field remains constrained by the challenge of reconciling high-fidelity realism with verifiable evaluation accuracy. To a…
Difficulty-Aware Agentic Orchestration for Query-Specific Multi-Agent Workflows
Jinwei Su, Qizhen Lan, Yinghui Xia +6
Large Language Model (LLM)-based agentic systems have shown strong capabilities across various tasks. However, existing multi-agent frameworks often rely on static or task-level wo…
MAPGD: Multi-Agent Prompt Gradient Descent for Collaborative Prompt Optimization
Yichen Han, Yuhang Han, Siteng Huang +7
Prompt engineering is crucial for fully leveraging large language models (LLMs), yet most existing optimization methods follow a single trajectory, resulting in limited adaptabilit…
ComfySearch: Autonomous Exploration and Reasoning for ComfyUI Workflows
Jinwei Su, Qizhen Lan, Zeyu Wang +7
AI-generated content has progressed from monolithic models to modular workflows, especially on platforms like ComfyUI, allowing users to customize complex creative pipelines. Howev…
DebFlow: Automating Agent Creation via Agent Debate
Jinwei Su, Yinghui Xia, Yiqun Duan +4
Large language models (LLMs) have demonstrated strong potential and impressive performance in automating the generation and optimization of workflows. However, existing approaches…
eSapiens: A Platform for Secure and Auditable Retrieval-Augmented Generation
Isaac Shi, Zeyuan Li, Fan Liu +4
We present eSapiens, an AI-as-a-Service (AIaaS) platform engineered around a business-oriented trifecta: proprietary data, operational workflows, and any major agnostic Large Langu…