6 papers
Agentic Auto-Research is Fuzz Testing
Yifeng He, Jicheng Wang, Yinzhe Zhao +2
Autonomous research agents can generate experiments faster than researchers can validate them. Researchers have responded by scaling the proposer and ranking more samples with a le…
Is Progressive Disclosure All You Need for Long-Context Agents?
Yifeng He, Yinzhe Zhao, Jicheng Wang +1
Long-document question answering usually forces a choice between loading the whole document into the context window and bolting on a separate retriever. Agentic AI suggests a broad…
Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters
Ailin Huang, Ang Li, Aobo Kong +213
We introduce Step 3.5 Flash, a sparse Mixture-of-Experts (MoE) model that bridges frontier-level agentic intelligence and computational efficiency. We focus on what matters most wh…
Adaptive Milestone Reward for GUI Agents
Congmin Zheng, Xiaoyun Mo, Xinbei Ma +10
Reinforcement Learning (RL) has emerged as a mainstream paradigm for training Mobile GUI Agents, yet it struggles with the temporal credit assignment problem inherent in long-horiz…
ColorAgent: Building A Robust, Personalized, and Interactive OS Agent
Ning Li, Qiqiang Lin, Zheng Wu +19
With the advancements in hardware, software, and large language model technologies, the interaction between humans and operating systems has evolved from the command-line interface…
ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks
Yuanyi Song, Heyuan Huang, Qiqiang Lin +9
The rapid advancement of multimodal large language models has enabled agents to operate mobile devices by directly interacting with graphical user interfaces, opening new possibili…