Publications (67)
Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation
Amruta Parulekar, Jinu Lee, Dilek Hakkani-Tür +1
Can a Single Model Master Both Multi-turn Conversations and Tool Use? CoALM: A Unified Conversational Agentic Language Model
Emre Can Acikgoz, Jeremiah Greer, Akul Datta +6
CESAR: Automatic Induction of Compositional Instructions for Multi-turn Dialogs
Taha Aksu, Devamanyu Hazarika, Shikib Mehri +4
User Preference Modeling for Conversational LLM Agents: Weak Rewards from Retrieval-Augmented Interaction
Yuren Hao, Shuhaib Mehri, ChengXiang Zhai +1
DocCHA: Towards LLM-Augmented Interactive Online diagnosis System
Xinyi Liu, Dachun Sun, Yi R. Fung +2
MultiSessionCollab: Learning User Preferences with Memory to Improve Long-Term Collaboration
Shuhaib Mehri, Priyanka Kargupta, Tal August +1
MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts?
Muntasir Wahed, Xiaona Zhou, Kiet A. Nguyen +5
Language Specific Knowledge: Do Models Know Better in X than in English?
Ishika Agarwal, Nimet Beyza Bozdag, Nisval Patel +1
Zero-Shot Controlled Generation with Encoder-Decoder Transformers
Devamanyu Hazarika, Mahdi Namazifar, Dilek Hakkani-Tür
Tool-R0: Self-Evolving LLM Agents for Tool-Learning from Zero Data
Emre Can Acikgoz, Cheng Qian, Jonas Hübotter +3
From Documents to Segments: A Contextual Reformulation for Topic Assignment
Hoonsang Yoon, Takyoung Kim, Wonkee Lee +3
Measuring and Mitigating the Distributional Gap Between Real and Simulated User Behaviors
Shuhaib Mehri, Philippe Laban, Sumuk Shashidhar +4
Towards Universal Dialogue Act Tagging for Task-Oriented Dialogues
Shachi Paul, Rahul Goel, Dilek Hakkani-Tür
A Rising Tide Lifts All Boats: MTQE Rewards for Idioms Improve General Translation Quality
Ishika Agarwal, Zhenlin He, Dhruva Patil +1
Plan Verification for LLM-Based Embodied Task Completion Agents
Ananth Hariharan, Vardhan Dongre, Dilek Hakkani-Tür +1
Drift No More? Context Equilibria in Multi-Turn LLM Interactions
Vardhan Dongre, Ryan A. Rossi, Viet Dac Lai +3
Learning Question-Guided Video Representation for Multi-Turn Video Question Answering
Guan-Lin Chao, Abhinav Rastogi, Semih Yavuz +3
A Desideratum for Conversational Agents: Capabilities, Challenges, and Future Directions
Emre Can Acikgoz, Cheng Qian, Hongru Wang +5
The Alignment Veto: How Safety Training Suppresses Cultural Knowledge in LLMs
Pardis Sadat Zahraei, Gokhan Tur, Dilek Hakkani-Tür +1
Do We Need Adam? Surprisingly Strong and Sparse Reinforcement Learning with SGD in LLMs
Sagnik Mukherjee, Lifan Yuan, Pavan Jayasinha +2
Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue
Vardhan Dongre, Dilek Hakkani-Tür
KILM: Knowledge Injection into Encoder-Decoder Language Models
Yan Xu, Mahdi Namazifar, Devamanyu Hazarika +3
Overview of the Ninth Dialog System Technology Challenge: DSTC9
Chulaka Gunasekara, Seokhwan Kim, Luis Fernando D'Haro +36
ToolRL: Reward is All Tool Learning Needs
Cheng Qian, Emre Can Acikgoz, Qi He +5
PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems
Jiayu Liu, Qihan Lin, Cheng Qian +8
Better Slow than Sorry: Introducing Positive Friction for Reliable Dialogue Systems
Mert İnan, Anthony Sicilia, Suvodip Dey +6
Premise-Augmented Reasoning Chains Improve Error Identification in Math reasoning with LLMs
Sagnik Mukherjee, Abhinav Chinta, Takyoung Kim +2
Neural Networks for Learnable and Scalable Influence Estimation of Instruction Fine-Tuning Data
Ishika Agarwal, Dilek Hakkani-Tür
DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking Dialogues
Takyoung Kim, Kang-wook Kim, Sang Hoon Woo +3
The paper presents DuplexGen, a framework that uses a small set of human preference annotations to adaptively generate turn‑taking behaviors in AI‑human dialogues across different…
Trajectory-Level Redirection Attacks on Vision-Language-Action Models
Gokul Puthumanaillam, Vardhan Dongre, Pranay Thangeda +3
Revisiting the Boundary between ASR and NLU in the Age of Conversational Dialog Systems
Manaal Faruqui, Dilek Hakkani-Tür
Building a Conversational Agent Overnight with Dialogue Self-Play
Pararth Shah, Dilek Hakkani-Tür, Gokhan Tür +4
LLMs are Vulnerable to Malicious Prompts Disguised as Scientific Language
Yubin Ge, Neeraja Kirtane, Hao Peng +1
Using In-Context Learning to Improve Dialogue Safety
Nicholas Meade, Spandana Gella, Devamanyu Hazarika +5
SpeakRL: Synergizing Reasoning, Speaking, and Acting in Language Models with Reinforcement Learning
Emre Can Acikgoz, Jinoh Oh, Jie Hao +7
Self-Improving LLM Agents at Test-Time
Emre Can Acikgoz, Cheng Qian, Heng Ji +2
PSI-Bench: Towards Clinically Grounded and Interpretable Evaluation of Depression Patient Simulators
Nguyen Khoi Hoang, Shuhaib Mehri, Tse-An Hsu +4
Persuade Me if You Can: A Framework for Evaluating Persuasion Effectiveness and Susceptibility Among Large Language Models
Nimet Beyza Bozdag, Shuhaib Mehri, Gokhan Tur +1
MAC: A Multi-Agent Framework for Interactive User Clarification in Multi-turn Conversations
Emre Can Acikgoz, Jinoh Oh, Joo Hyuk Jeon +7
From Context to Action: Analysis of the Impact of State Representation and Context on the Generalization of Multi-Turn Web Navigation Agents
Nalin Tiwary, Vardhan Dongre, Sanil Arun Chawla +2
When Attention Closes: How LLMs Lose the Thread in Multi-Turn Interaction
Vardhan Dongre, Joseph Hsieh, Viet Dac Lai +3
EiCAP: Beyond Fluency, Probing and Improving Emotional Intelligence in LLMs via Psychologically Grounded Multi-Turn Dialogue
Nizi Nazar, Pardis Sadat Zahraei, Dilek Hakkani-Tür +2
MIRAGE: A Benchmark for Multimodal Information-Seeking and Reasoning in Agricultural Expert-Guided Conversations
Vardhan Dongre, Chi Gui, Shubham Garg +4
Goal Alignment in LLM-Based User Simulators for Conversational AI
Shuhaib Mehri, Xiaocheng Yang, Takyoung Kim +3
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise
Abdulrahman AlRabah, Xiaocheng Yang, Dilek Hakkani-Tür +1
Simulating User Agents for Embodied Conversational-AI
Daniel Philipov, Vardhan Dongre, Gokhan Tur +1
VISITRON: Visual Semantics-Aligned Interactively Trained Object-Navigator
Ayush Shrivastava, Karthik Gopalakrishnan, Yang Liu +4
YourBench: Easy Custom Evaluation Sets for Everyone
Sumuk Shashidhar, Clémentine Fourrier, Alina Lozovskia +3
Few-Shot Synthetic Accented Speech for ASR Fine-Tuning: What Helps and When?
Yurii Halychanskyi, Nimet Beyza Bozdag, Mark Hasegawa-Johnson +2
HyST: A Hybrid Approach for Flexible and Accurate Dialogue State Tracking
Rahul Goel, Shachi Paul, Dilek Hakkani-Tür
ReIn: Conversational Error Recovery with Reasoning Inception
Takyoung Kim, Jinseok Nam, Chandrayee Basu +5
Beyond Sample-Level Feedback: Using Reference-Level Feedback to Guide Data Synthesis
Shuhaib Mehri, Xiusi Chen, Heng Ji +1
TD-EVAL: Revisiting Task-Oriented Dialogue Evaluation by Combining Turn-Level Precision with Dialogue-Level Comparisons
Emre Can Acikgoz, Carl Guo, Suvodip Dey +4
SMART: Self-Aware Agent for Tool Overuse Mitigation
Cheng Qian, Emre Can Acikgoz, Hongru Wang +5
From Fact to Judgment: Investigating the Impact of Task Framing on LLM Conviction in Dialogue Systems
Parisa Rabbani, Nimet Beyza Bozdag, Dilek Hakkani-Tür
ReSpAct: Harmonizing Reasoning, Speaking, and Acting Towards Building Large Language Model-Based Conversational AI Agents
Vardhan Dongre, Xiaocheng Yang, Emre Can Acikgoz +3
DialDefer: A Framework for Detecting and Mitigating LLM Dialogic Deference
Parisa Rabbani, Priyam Sahoo, Ruben Mathew +4
AcquisitionSynthesis: Targeted Data Generation using Acquisition Functions
Ishika Agarwal, Sofia Stoica, Emre Can Acikgoz +4
ATOD: An Evaluation Framework and Benchmark for Agentic Task-Oriented Dialogue Systems
Yifei Zhang, Hooshang Nayyeri, Rinat Khaziev +4
Must Read: A Comprehensive Survey of Computational Persuasion
Nimet Beyza Bozdag, Shuhaib Mehri, Xiaocheng Yang +7
Language Model is All You Need: Natural Language Understanding as Question Answering
Mahdi Namazifar, Alexandros Papangelis, Gokhan Tur +1
Measuring, Localizing, and Ablating Alignment Signatures in LLMs
Aniket Anand, Janvijay Singh, Zhewei Sun +2
Unsupervised Human Preference Learning
Sumuk Shashidhar, Abhinav Chinta, Vaibhav Sahai +1
Current Agents Fail to Leverage World Model as Tool for Foresight
Cheng Qian, Emre Can Acikgoz, Bingxuan Li +8
GBC: Gradient-Based Connections for Optimizing Multi-Agent Systems
Xiaocheng Yang, Abdulrahman Alrabah, Dilek Hakkani-Tür +1
Do LLMs Encode Functional Importance of Reasoning Tokens?
Janvijay Singh, Dilek Hakkani-Tür
AURA: A Diagnostic Framework for Tracking User Satisfaction of Interactive Planning Agents
Takyoung Kim, Janvijay Singh, Shuhaib Mehri +6