34 papers
Consistent but Miscalibrated: Evaluating LLM Limitations for Risk Communication in Natural Language
Diego Cerda-Mardini, Sarath Chandar, Sreenath Madathil
LLMs are increasingly deployed as post-hoc explainers of AI-generated outputs, yet it remains unclear whether they can reliably communicate probabilistic information in natural lan…
Failed Reasoning Traces Tell You What Is Fixable (But Not by Reading Them)
Nizar Islah, Istabrak Abbes, Irina Rish +2
When post-trained language models fail on reasoning problems, the common test-time-scaling response is to spend more compute on additional attempts, and the failed traces play no f…
Dialectics of Alignment: Harnessing Unsafe Knowledge for Dynamic Safety Routing
Maryam Hashemzadeh, Jerry Huang, Minseon Kim +2
The prevailing paradigm in large language model (LLM) alignment operates via erasure, filtering unsafe data or training models to strictly refuse harmful prompts. While effective a…
Probabilistic Calibration Is a Trainable Capability in Language Models
Davide Baldelli, Sruthi Kuriakose, Maryam Hashemzadeh +2
Language models are increasingly used in settings where outputs must satisfy user-specified randomness constraints, yet their generation probabilities are often poorly calibrated t…
Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models
Nilaksh, Saurav Jha, Artem Zholus +1
World model-based policy evaluation is a practical proxy for testing real-world robot control by rolling out candidate actions in action-conditioned video diffusion models. As thes…
TAPNext++: What's Next for Tracking Any Point (TAP)?
Sebastian Jung, Artem Zholus, Martin Sundermeyer +6
Tracking-Any-Point (TAP) models aim to track any point through a video which is a crucial task in AR/XR and robotics applications. The recently introduced TAPNext approach proposes…