9 papers
OPERA: Operator-residual feedback for reliable autonomous optical experiments with language-model agents
Ning Xu, Xiang Zheng, Fuqiang Zhong +4
Autonomous agents choose actions using scores that may not reflect experimental success. We developed OPERA, an operator-residual framework for optical experiments. It represents e…
DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models
Cheng Yin, Yankai Lin, Wang Xu +4
Does Chain-of-Thought (CoT) reasoning genuinely improve Vision Language Action (VLA) models, or does it merely add overhead? Existing CoT-VLA systems report limited and inconsisten…
AgentCPM-Report: Interleaving Drafting and Deepening for Open-Ended Deep Research
Yishan Li, Wentong Chen, Yukun Yan +12
Generating deep research reports requires large-scale information acquisition and the synthesis of insight-driven analysis, posing a significant challenge for current language mode…
AgentCPM-Explore: Realizing Long-Horizon Deep Exploration for Edge-Scale Agents
Haotian Chen, Xin Cong, Shengda Fan +16
While Large Language Model (LLM)-based agents have shown remarkable potential for solving complex tasks, existing systems remain heavily reliant on large-scale models, leaving the…
Learning to Focus: Causal Attention Distillation via Gradient-Guided Token Pruning
Yiju Guo, Wenkai Yang, Zexu Sun +3
Large language models (LLMs) have demonstrated significant improvements in contextual understanding. However, their ability to attend to truly critical information during long-cont…
AppCopilot: Toward General, Accurate, Long-Horizon, and Efficient Mobile Agent
Jingru Fan, Yufan Dang, Jingyao Wu +5
With the raid evolution of large language models and multimodal models, the mobile-agent landscape has proliferated without converging on the fundamental challenges. This paper ide…