activity
20232026
most citedBadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models

4 citations · 7 across the 11 of their papers we have counts for

collaborators
Showing cs.CLShow all

9 papers · 1 filter

cs.CL2025

LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries

Ming Yin, Dinghan Shen, Silei Xu +11

Tool calling has emerged as a critical capability for AI agents. In contrast to conventional tool calling frameworks that rely on static, provider-specific tool definitions, the Mo…

cs.CL2025

LogicIF: Towards Complex Logic Instruction Following

Mian Zhang, Shujian Liu, Sixun Dong +9

Instruction following has catalyzed the recent era of Large Language Models (LLMs) and is the foundational skill underpinning more advanced capabilities such as reasoning and agent…

cs.CL2024

STRUX: An LLM for Decision-Making with Structured Explanations

Yiming Lu, Yebowen Hu, Hassan Foroosh +2

Countless decisions shape our daily lives, and it is paramount to understand the how and why behind these choices. In this paper, we introduce a new LLM decision-making framework c…

cs.CL2024

DeFine: Decision-Making with Analogical Reasoning over Factor Profiles

Yebowen Hu, Xiaoyang Wang, Wenlin Yao +5

LLMs are ideal for decision-making thanks to their ability to reason over long contexts. However, challenges arise when processing speech transcripts that describe complex scenario…

cs.CL2024

When Reasoning Meets Information Aggregation: A Case Study with Sports Narratives

Yebowen Hu, Kaiqiang Song, Sangwoo Cho +5

Reasoning is most powerful when an LLM accurately aggregates relevant information. We examine the critical role of information aggregation in reasoning by requiring the LLM to anal…

cs.CL2024

Can Large Language Models do Analytical Reasoning?

Yebowen Hu, Kaiqiang Song, Sangwoo Cho +4

This paper explores the cutting-edge Large Language Model with analytical reasoning on sports. Our analytical reasoning embodies the tasks of letting large language models count ho…