Showing cs.AIShow all
3 papers · 1 filter
cs.AI2025
ReasoningWeekly: A General Knowledge and Verbal Reasoning Challenge for Large Language Models
Zixuan Wu, Francesca Lucchetti, Aleksander Boruch-Gruszecki +5
Existing benchmarks for frontier models often test specialized, "PhD-level" knowledge that is difficult for non-experts to grasp. In contrast, we present a benchmark with 613 probl…
cs.AI2025
Joint Verification and Refinement of Language Models for Safety-Constrained Planning
Yunhao Yang, Neel P. Bhatt, William Ward +3
Large language models possess impressive capabilities in generating programs (e.g., Python) from natural language descriptions to execute robotic tasks. However, these generated pr…
cs.AI2024
Dynamic Model Predictive Shielding for Provably Safe Reinforcement Learning
Arko Banerjee, Kia Rahmani, Joydeep Biswas +1
Among approaches for provably safe reinforcement learning, Model Predictive Shielding (MPS) has proven effective at complex tasks in continuous, high-dimensional state spaces, by l…