4 papers
Preference Optimization Drives Monoculture in LLM Prediction Markets
James Begin, Brendan Gho, Suman Muppavarapu +6
Prediction markets rest on the independence of participant errors. As LLM agents become active traders on platforms like Kalshi and Polymarket, we ask: does this independence hold…
Interpreting Latent CoT Reasoning as Dynamical Systems
Sabari Iyyappan Duraipandian, Shreya Sanjay Boyane, Manju Nagesh +3
Recent latent reasoning methods, such as CODI and COCONUT, face a fundamental interpretability problem: they maintain multiple superimposed candidate traces in the hidden space at…
From Competition to Coordination: Market Making as a Scalable Framework for Safe and Aligned Multi-Agent LLM Systems
Brendan Gho, Suman Muppavarapu, Afnan Shaik +6
As foundation models are increasingly deployed as interacting agents in multi-agent systems, their collective behavior raises new challenges for trustworthiness, transparency, and…
HumanMCP: A Human-Like Query Dataset for Evaluating MCP Tool Retrieval Performance
Shubh Laddha, Lucas Changbencharoen, Win Kuptivej +3
Model Context Protocol (MCP) servers contain a collection of thousands of open-source standardized tools, linking LLMs to external systems; however, existing datasets and benchmark…