From the 2 of 7 linked papers with an AI index.
7 papers
ToolAlignBench: Investigating Alignment Conflicts in Tool-Calling Enabled LLMs
Aryan Keluskar, Amrita Bhattacharjee, Huan Liu
The paper introduces a benchmark of 128 tool‑calling scenarios to study how safety‑aligned large language models may override deployment instructions in regulated settings, reveali…
When Does Personality Composition Matter for Multi-Agent LLM Teams?
Aryan Keluskar, Amrita Bhattacharjee, Huan Liu
The paper examines how prompting large language models with different personality traits influences the performance of multi-agent teams across coding, collaborative research, and…
From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
Dawei Li, Bohan Jiang, Liangjie Huang +10
Assessment and evaluation have long been critical challenges in artificial intelligence (AI) and natural language processing (NLP). Traditional methods, usually matching-based or s…
Large Language Models for Data Annotation and Synthesis: A Survey
Zhen Tan, Dawei Li, Song Wang +7
Data annotation and synthesis generally refers to the labeling or generating of raw data with relevant information, which could be used for improving the efficacy of machine learni…
Do LLMs Understand Ambiguity in Text? A Case Study in Open-world Question Answering
Aryan Keluskar, Amrita Bhattacharjee, Huan Liu
Ambiguity in natural language poses significant challenges to Large Language Models (LLMs) used for open-domain question answering. LLMs often struggle with the inherent uncertaint…
Zero-shot LLM-guided Counterfactual Generation: A Case Study on NLP Model Evaluation
Amrita Bhattacharjee, Raha Moraffah, Joshua Garland +1
With the development and proliferation of large, complex, black-box models for solving many natural language processing (NLP) tasks, there is also an increasing necessity of method…