10 papers
OpenForgeRL: Train Harness-native Agents in Any Environment
Xiao Yu, Baolin Peng, Ruize Xu +7
Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While power…
RIME: Enabling Large-Scale Agentic Music Post-Production
Noah Schaffer, Nikhil Singh
Almost every piece of recorded music you have ever heard was modified before it reached you; commercial releases rarely spring fully-formed from the mind of a musician. Despite the…
DMV-Bench: Diagnosing Long-Horizon Multimodal Agents' Visual Memory with Incidental Cue Injection
Yujin Tang, Chenming Shang, Ruize Xu +1
Research on agent memory has matured rapidly, but almost entirely on the text side: few existing benchmarks ask, in an interactive environment, when an agent genuinely needs to rem…
Visual Persuasion: What Influences Decisions of Vision-Language Models?
Manuel Cherep, Pranav M R, Pattie Maes +1
The web is littered with images, once created for human consumption and now increasingly interpreted by agents using vision-language models (VLMs). These agents make visual decisio…
Discovering and Steering Interpretable Concepts in Large Generative Music Models
Nikhil Singh, Manuel Cherep, Pattie Maes
The fidelity with which neural networks can now generate content such as music presents a scientific opportunity: these systems appear to have learned implicit theories of such con…
A Framework for Studying AI Agent Behavior: Evidence from Consumer Choice Experiments
Manuel Cherep, Chengtian Ma, Abigail Xu +3
Environments built for people are increasingly operated by a new class of economic actors: LLM-powered software agents making decisions on our behalf. These decisions range from ou…