6 papers
Making Embodied AI Reliable: A Community Agenda from Testing to Formal Verification
Xi Zheng, Dulanga Weerakoon, Yintong Huo +8
Embodied AI systems are increasingly deployed in open-world environments, yet ensuring their reliability remains a fundamental challenge. Drawing on discussions from the AAAI'26 Br…
An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models
Mingzhong Sun, Teresa Yeo, Armando Solar-Lezama +1
Studies of human reasoning have shown that people are typically stronger at evaluating reasoning than producing it from scratch. In contrast, large reasoning models (LRMs) are trai…
Adaptive Problem Generation via Symbolic Representations
Teresa Yeo, Myeongho Jeon, Dulaj Weerakoon +4
We present a method for generating training data for reinforcement learning with verifiable rewards to improve small open-weights language models on mathematical tasks. Existing da…
Towards Adaptive Environment Generation for Training Embodied Agents
Teresa Yeo, Dulaj Weerakoon, Dulanga Weerakoon +1
Embodied agents struggle to generalize to new environments, even when those environments share similar underlying structures to their training settings. Most current approaches to…
Controlled Training Data Generation with Diffusion Models
Teresa Yeo, Andrei Atanov, Harold Benoit +4
We present a method to control a text-to-image generative model to produce training data useful for supervised learning. Unlike previous works that employ an open-loop approach and…
An Analysis of Model Robustness across Concurrent Distribution Shifts
Myeongho Jeon, Suhwan Choi, Hyoje Lee +1
Machine learning models, meticulously optimized for source data, often fail to predict target data when faced with distribution shifts (DSs). Previous benchmarking studies, though…