9 papers
ASPIRE: Agentic /Skills Discovery for Robotics
Runyu Lu, Yubo Wu, Ethan Kou +11
Traditional robot programming is challenging: it requires orchestrating multimodal perception, managing physical contact dynamics, and handling diverse configurations and execution…
EpiAgent: An Agent-Centric System for Ancient Inscription Restoration
Shipeng Zhu, Ang Chen, Na Nie +3
Ancient inscriptions, as repositories of cultural memory, have suffered from centuries of environmental and human-induced degradation. Restoring their intertwined visual and textua…
Efficient Distributed MLLM Training with Cornstarch
Insu Jang, Runyu Lu, Nikhil Bansal +2
Multimodal large language models (MLLMs) extend the capabilities of large language models (LLMs) by combining heterogeneous model architectures to handle diverse modalities like im…
The Last Human-Written Paper: Agent-Native Research Artifacts
Jiachen Liu, Jiaxin Pei, Jintao Huang +34
Scientific publication compresses a branching, iterative research process into a linear narrative, discarding the majority of what was discovered along the way. This compilation im…
Experiment-as-Code Labs: A Declarative Stack for AI-Driven Scientific Discovery
Zhenning Yang, Yuhan Chen, Patrick Tser Jern Kon +5
To unleash the full potential of AI for Science, we must untether the agents from a purely digital environment. The agent's ability to control and explore in real-world labs is ess…
Ambig-IaC: Multi-level Disambiguation for Interactive Cloud Infrastructure-as-Code Synthesis
Zhenning Yang, Kaden Gruizenga, Tongyuan Miao +3
The scale and complexity of modern cloud infrastructure have made Infrastructure-as-Code (IaC) essential for managing deployments. While large Language models (LLMs) are increasing…