5 papers
The Value of Variance: Mitigating Debate Collapse in Multi-Agent Systems via Uncertainty-Driven Policy Optimization
Luoxi Tang, Yuqiao Meng, Joseph Costa +3
Multi-agent debate (MAD) systems improve LLM reasoning through iterative deliberation, but remain vulnerable to debate collapse, a failure type where final agent decisions are comp…
Smart Privacy Policy Assistant: An LLM-Powered System for Transparent and Actionable Privacy Notices
Sriharshini Kalvakuntla, Luoxi Tang, Yuqiao Meng +1
Most users agree to online privacy policies without reading or understanding them, even though these documents govern how personal data is collected, shared, and monetized. Privacy…
POLAR: Automating Cyber Threat Prioritization through LLM-Powered Assessment
Luoxi Tang, Yuqiao Meng, Ankita Patra +3
Large Language Models (LLMs) are intensively used to assist security analysts in counteracting the rapid exploitation of cyber threats, wherein LLMs offer cyber threat intelligence…
From Test-taking to Cognitive Scaffolding: A Pedagogical Diagnostic Benchmark for LLMs on English Standardized Tests
Luoxi Tang, Tharunya Sundar, Yuqiao Meng +7
As large language models (LLMs) are increasingly integrated into educational tools, current evaluations on standardized tests predominantly focus on binary outcome accuracy. Instea…
On the Eligibility of LLMs for Counterfactual Reasoning: A Decompositional Study
Shuai Yang, Qi Yang, Luoxi Tang +4
Counterfactual reasoning has emerged as a crucial technique for generalizing the reasoning capabilities of large language models (LLMs). By generating and analyzing counterfactual…