Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
MANTRA: Synthesizing SMT-Validated Compliance Benchmarks for Tool-Using LLM Agents
Ashwani Anand, Ivi Chatzi, Ritam Raha +1
Tool-using large language model (LLM) agents are increasingly deployed in settings where their reliable behavior is governed by strict procedural manuals. Ensuring that such agents…
cs.CL2026
Evaluation of Large Language Models via Coupled Token Generation
Nina Corvelo Benz, Stratis Tsirtsis, Eleni Straitouri +4
State of the art large language models rely on randomization to respond to a prompt. As an immediate consequence, a model may respond differently to the same prompt if asked multip…
cs.CL2026
Tokenization Multiplicity Leads to Arbitrary Price Variation in LLM-as-a-service
Ivi Chatzi, Nina Corvelo Benz, Stratis Tsirtsis +1
Providers of LLM-as-a-service have predominantly adopted a simple pricing model: users pay a fixed price per token. Consequently, one may think that the price two different users w…