6 papers
An Empirical Investigation of Robustness in Large Language Models under Tabular Distortions
Avik Dutta, Harshit Nigam, Hosein Hasanbeig +2
We investigate how large language models (LLMs) fail when tabular data in an otherwise canonical representation is subjected to semantic and structural distortions. Our findings re…
Distributions In, Distributions Out: The Case for Soft-Label Training
Agamdeep Singh, Ashish Tiwari, Hosein Hasanbeig +1
Supervised classifiers output a distribution over classes but are typically trained against a single label obtained by collapsing multiple annotators into a majority vote. On tasks…
ConDABench: Interactive Evaluation of Language Models for Data Analysis
Avik Dutta, Priyanshu Gupta, Hosein Hasanbeig +6
Real-world data analysis tasks often come with under-specified goals and unclean data. User interaction is necessary to understand and disambiguate a user's intent, and hence, esse…
Mission-driven Exploration for Accelerated Deep Reinforcement Learning with Temporal Logic Task Specifications
Jun Wang, Hosein Hasanbeig, Kaiyuan Tan +2
This paper addresses the problem of designing control policies for agents with unknown stochastic dynamics and control objectives specified using Linear Temporal Logic (LTL). Recen…
Are LLMs Good Cryptic Crossword Solvers?
Abdelrahman Sadallah, Daria Kotova, Ekaterina Kochmar
Cryptic crosswords are puzzles that rely not only on general knowledge but also on the solver's ability to manipulate language on different levels and deal with various types of wo…
Progressive Safeguards for Safe and Model-Agnostic Reinforcement Learning
Nabil Omi, Hosein Hasanbeig, Hiteshi Sharma +2
In this paper we propose a formal, model-agnostic meta-learning framework for safe reinforcement learning. Our framework is inspired by how parents safeguard their children across…