works on

From the 2 of 7 linked papers with an AI index.

activity
20242026
collaborators

7 papers

cs.SE2026

ToolAlignBench: Investigating Alignment Conflicts in Tool-Calling Enabled LLMs

Aryan Keluskar, Amrita Bhattacharjee, Huan Liu

The paper introduces a benchmark of 128 tool‑calling scenarios to study how safety‑aligned large language models may override deployment instructions in regulated settings, reveali…

cs.AI2026

When Does Personality Composition Matter for Multi-Agent LLM Teams?

Aryan Keluskar, Amrita Bhattacharjee, Huan Liu

The paper examines how prompting large language models with different personality traits influences the performance of multi-agent teams across coding, collaborative research, and…

cs.AI2025

From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge

Dawei Li, Bohan Jiang, Liangjie Huang +10

Assessment and evaluation have long been critical challenges in artificial intelligence (AI) and natural language processing (NLP). Traditional methods, usually matching-based or s…

cs.CL2024

Large Language Models for Data Annotation and Synthesis: A Survey

Zhen Tan, Dawei Li, Song Wang +7

Data annotation and synthesis generally refers to the labeling or generating of raw data with relevant information, which could be used for improving the efficacy of machine learni…

cs.CL2024

Do LLMs Understand Ambiguity in Text? A Case Study in Open-world Question Answering

Aryan Keluskar, Amrita Bhattacharjee, Huan Liu

Ambiguity in natural language poses significant challenges to Large Language Models (LLMs) used for open-domain question answering. LLMs often struggle with the inherent uncertaint…

cs.CL2024

Zero-shot LLM-guided Counterfactual Generation: A Case Study on NLP Model Evaluation

Amrita Bhattacharjee, Raha Moraffah, Joshua Garland +1

With the development and proliferation of large, complex, black-box models for solving many natural language processing (NLP) tasks, there is also an increasing necessity of method…