activity
20242026
most citedAdaptive Linguistic Prompting (ALP) Enhances Phishing Webpage Detection in Multimodal Large Language Models

3 citations · 9 across the 17 of their papers we have counts for

collaborators

27 papers

cs.CL2026

MUDDLE: Measuring Understanding of Documents under Distractor and Length Effects

Jason Luo, Saibilila Abudukelimu, Judy Song +4

Document question-answering systems increasingly answer questions over collections of retrieved documents rather than one clean source, so robustness to distracting context matters…

cs.AI2026

CBMAS: Cognitive Behavioral Modeling via Activation Steering

Ahmed H. Ismail, Anthony Kuang, Ayo Akinkugbe +2

Large language models (LLMs) often encode cognitive behaviors unpredictably across prompts, layers, and contexts, making them difficult to diagnose and control. We present CBMAS, a…

cs.AI2025

Reasoning Relay: Evaluating Stability and Interchangeability of Large Language Models in Mathematical Reasoning

Leo Lu, Jonathan Zhang, Sean Chua +4

Chain-of-Thought (CoT) prompting has significantly advanced the reasoning capabilities of large language models (LLMs). While prior work focuses on improving model performance thro…

cs.CL2025

Direct Confidence Alignment: Aligning Verbalized Confidence with Internal Confidence In Large Language Models

Glenn Zhang, Treasure Mayowa, Jason Fan +4

Producing trustworthy and reliable Large Language Models (LLMs) has become increasingly important as their usage becomes more widespread. Calibration seeks to achieve this by impro…

cs.AI2025

SMAGDi: Socratic Multi Agent Interaction Graph Distillation for Efficient High Accuracy Reasoning

Aayush Aluru, Myra Malik, Samarth Patankar +4

Multi-agent systems (MAS) often achieve higher reasoning accuracy than single models, but their reliance on repeated debates across agents makes them computationally expensive. We…

cs.AI2025

SwiftSolve: A Self-Iterative, Complexity-Aware Multi-Agent Framework for Competitive Programming

Adhyayan Veer Singh, Aaron Shen, Brian Law +4

Correctness alone is insufficient: LLM-generated programs frequently satisfy unit tests while violating contest time or memory budgets. We present SwiftSolve, a complexity-aware mu…