15 papers
An AI Co-Data-Scientist for Prioritizing Candidate Biomarkers from Wearable Sensor Data
Yubin Kim, Salman Rahman, Samuel Schmidgall +33
Wearable devices generate continuous physiological and behavioral data, but converting these signals into clinically reviewable biomarker hypotheses remains labor-intensive. We int…
EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments
Jundong Xu, Qingchuan Li, Jiaying Wu +11
Large language model (LLM) agents have achieved strong performance on a wide range of benchmarks, yet most evaluations assume static environments. In contrast, real-world deploymen…
Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models' Social Reasoning
Dong Won Lee, Yubin Kim, Denison Guvenoz +5
Our work focuses on the social reasoning capabilities of foundation models for real-world human-robot interactions. We introduce the Social Human Robot Embodied Conversation (SHREC…
TeamBench: Evaluating Agent Coordination under Enforced Role Separation
Yubin Kim, Chanwoo Park, Taehan Kim +9
Agent systems often decompose a task across multiple roles, but these roles are typically specified by prompts rather than enforced by access controls. Without enforcement, a team…
InvThink: Premortem Reasoning for Safer Language Models
Yubin Kim, Taehan Kim, Eugene Park +4
We present InvThink, a training and prompting framework that requires the model to enumerate, analyze, and constrain potential failures before generating its final response. Unlike…
Towards a Science of Scaling Agent Systems
Yubin Kim, Ken Gu, Chanwoo Park +17
Agents, language model-based systems capable of reasoning, planning, and acting are widely adopted in real-world tasks, yet how their performance changes as these systems scale acr…