From the 1 of 8 linked papers with an AI index.
8 papers
Align AI to Dynamic Human-AI Workflows
Valerie Chen, Cleotilde Gonzalez, Anita Williams Woolley +4
The paper proposes moving from static, preference‑emulating AI alignment toward interactive, complementary alignment where human and AI behaviors co‑evolve over time, drawing on so…
When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games
Jerick Shi, Terry Jingcheng Zhang, Bernhard Schölkopf +2
As large language models are deployed as autonomous agents that communicate intentions before acting, a critical safety question is whether agents that publicly commit to actions w…
Why Search When You Can Transfer? Amortized Agentic Workflow Design from Structural Priors
Shiyi Du, Jiayuan Liu, Weihua Du +6
Automated agentic workflow design currently relies on per-task iterative search, which is computationally prohibitive and fails to reuse structural knowledge across tasks. We obser…
The Consensus Trap: Rescuing Multi-Agent LLMs from Adversarial Majorities via Token-Level Collaboration
Jiayuan Liu, Shiyi Du, Weihua Du +2
Multi-agent large language model (LLM) architectures increasingly rely on response-level aggregation, such as Majority Voting (MAJ), to raise reasoning ceilings. However, in open e…
From Sycophancy to Deception: A Unified Taxonomy for LLM Spontaneous Misalignment
Jerick Shi, Terry Jingcheng Zhang, Zhijing Jin +1
Large language models (LLMs) could produce systematically misaligned output, from hallucinated citations to strategic deception of evaluators, yet these phenomena are studied by se…
Cheap Talk, Empty Promise: Frontier LLMs easily break public promises for self-interest
Jerick Shi, Terry Jingcheng Zhang, Zhijing Jin +1
Large language models are increasingly deployed as autonomous agents in multi-agent settings where they communicate intentions and take consequential actions with limited human ove…