7 papers
SeClaw: Spec-Driven Security Task Synthesis for Evaluating Autonomous Agents
Hao Cheng, Changtao Miao, Tianle Song +21
Autonomous LLM agents increasingly operate in stateful environments where they access tools, files, memory, and external services. While such capabilities enable complex real-world…
IFPV: An Integrated Multi-Agent Framework for Generative Operational Planning and High-Fidelity Plan Verification
Zhigao Huang, Zhengqing Hu, Dong Chen +5
Operational plan generation and verification are critical for modern complex and rapidly changing battlefield environments, yet traditional generation and verification methods stil…
Grounding LLMs in Scientific Discovery via Embodied Actions
Bo Zhang, Jinfeng Zhou, Yuxuan Chen +3
Large Language Models (LLMs) have shown significant potential in scientific discovery but struggle to bridge the gap between theoretical reasoning and verifiable physical simulatio…
Resolving State Ambiguity in Robot Manipulation via Adaptive Working Memory Recoding
Qingda Hu, Ziheng Qiu, Zijun Xu +7
State ambiguity is common in robotic manipulation. Identical observations may correspond to multiple valid behavior trajectories. The visuomotor policy must correctly extract the a…
SCP: Accelerating Discovery with a Global Web of Autonomous Scientific Agents
Yankai Jiang, Wenjie Lou, Lilong Wang +17
We introduce SCP: the Science Context Protocol, an open-source standard designed to accelerate discovery by enabling a global network of autonomous scientific agents. SCP is built…
Intern-S1: A Scientific Multimodal Foundation Model
Lei Bai, Zhongrui Cai, Yuhang Cao +173
In recent years, a plethora of open-source foundation models have emerged, achieving remarkable progress in some widely attended fields, with performance being quite close to that…