3 papers
cs.AI2026
Prompt-Robust Language Models: Which Training Strategies Work?
Frederic Sadrieh, Michal Štefánik
Despite their strong performance, large language models remain highly sensitive to prompt formulation. Prior work addresses this through refined data construction or through dedica…
cs.LG2026
CooperBench: Why Coding Agents Cannot be Your Teammates Yet
Arpandeep Khatua, Hao Zhu, Peter Tran +8
Resolving team conflicts requires not only task-specific competence, but also social intelligence to find common ground and build consensus. As AI agents increasingly collaborate o…
cs.CL2024
Learning to Predict Usage Options of Product Reviews with LLM-Generated Labels
Leo Kohlenberg, Leonard Horns, Frederic Sadrieh +7
Annotating large datasets can be challenging. However, crowd-sourcing is often expensive and can lack quality, especially for non-trivial tasks. We propose a method of using LLMs a…