From the 2 of 8 linked papers with an AI index.
8 papers
ToolAlignBench: Investigating Alignment Conflicts in Tool-Calling Enabled LLMs
Aryan Keluskar, Amrita Bhattacharjee, Huan Liu
The paper introduces a benchmark of 128 tool‑calling scenarios to study how safety‑aligned large language models may override deployment instructions in regulated settings, reveali…
When Does Personality Composition Matter for Multi-Agent LLM Teams?
Aryan Keluskar, Amrita Bhattacharjee, Huan Liu
The paper examines how prompting large language models with different personality traits influences the performance of multi-agent teams across coding, collaborative research, and…
AgentOS: From Application Silos to a Natural Language-Driven Data Ecosystem
Rui Liu, Tao Zhe, Dongjie Wang +5
The rapid emergence of open-source, locally hosted intelligent agents marks a critical inflection point in human-computer interaction. Systems such as OpenClaw demonstrate that Lar…
Generalization in Online Reinforcement Learning for Mobile Agents
Li Gu, Zihuan Jiang, Zhixiang Chi +5
Graphical user interface (GUI)-based mobile agents automate digital tasks on mobile devices by interpreting natural-language instructions and interacting with the screen. While rec…
Do LLMs Understand Ambiguity in Text? A Case Study in Open-world Question Answering
Aryan Keluskar, Amrita Bhattacharjee, Huan Liu
Ambiguity in natural language poses significant challenges to Large Language Models (LLMs) used for open-domain question answering. LLMs often struggle with the inherent uncertaint…
Zero-shot LLM-guided Counterfactual Generation: A Case Study on NLP Model Evaluation
Amrita Bhattacharjee, Raha Moraffah, Joshua Garland +1
With the development and proliferation of large, complex, black-box models for solving many natural language processing (NLP) tasks, there is also an increasing necessity of method…