2 papers
cs.AI2026
Online Agent-as-a-Judge: Situation-Generating Evaluation for Interactive Agents
Hyogon Ryu, Jeonghwan Kim, Yewon Lim +3
Evaluating LLM-powered interactive social agents is challenging because socially relevant behaviors depend not only on isolated outputs, but also on prior interactions, social role…
cs.AI2026
Efficient Skill Grounding via Code Refactoring with Small Language Models
Sera Choi, Wonje Choi, Saehun Chun +4
Effective skill grounding is essential for deploying reusable skills in embodied agents, as even minor embodiment or environmental differences can render an entire skill incompatib…