11 papers
Used Car Salesbots? Honesty and Credulity of LLMs as Bargaining Agents under Partial Information
Antonio Valerio Miceli-Barone, Vaishak Belle, Shay B. Cohen
In this work we study agents in simulated bargaining scenarios, where a buyer and a seller communicate through a text channel and attempt to negotiate mutually beneficial trades, u…
Can LLMs Compress (and Decompress)? Evaluating Code Understanding and Execution via Invertibility
Nickil Maveli, Antonio Vergari, Shay B. Cohen
LLMs demonstrate strong performance on code benchmarks, yet consistent reasoning across forward and backward execution remains elusive. We present RoundTripCodeEval (RTCE), a bench…
Theorem Prover as a Judge for Synthetic Data Generation
Joshua Ong Jun Leang, Giwon Hong, Wenda Li +1
The demand for synthetic data in mathematical reasoning has increased due to its potential to enhance the mathematical capabilities of large language models (LLMs). However, ensuri…
CoMAT: Chain of Mathematically Annotated Thought Improves Mathematical Reasoning
Joshua Ong Jun Leang, Aryo Pradipta Gema, Shay B. Cohen
Mathematical reasoning remains a significant challenge for large language models (LLMs), despite progress in prompting techniques such as Chain-of-Thought (CoT). We present **Chain…
Integrating Counterfactual Simulations with Language Models for Explaining Multi-Agent Behaviour
Bálint Gyevnár, Christopher G. Lucas, Stefano V. Albrecht +1
Autonomous multi-agent systems (MAS) are useful for automating complex tasks but raise trust concerns due to risks such as miscoordination or goal misalignment. Explainability is v…
One More Question is Enough, Expert Question Decomposition (EQD) Model for Domain Quantitative Reasoning
Mengyu Wang, Sotirios Sabanis, Miguel de Carvalho +2
Domain-specific quantitative reasoning remains a major challenge for large language models (LLMs), especially in fields requiring expert knowledge and complex question answering (Q…