papers

Publications (21)

cs.CL2023

Investigating Reinforcement Learning for Communication Strategies in a Task-Initiative Setting

Baber Khalid, Matthew Stone

Many conversational domains require the system to present nuanced information to users. Such systems must follow up what they say to address clarification questions and repair misu…

cs.CL2021

COSMic: A Coherence-Aware Generation Metric for Image Descriptions

Mert İnan, Piyush Sharma, Baber Khalid +3

Developers of text generation models rely on automated evaluation metrics as a stand-in for slow and expensive manual evaluations. However, image captioning metrics have struggled…

eess.AS2025

Towards Robust Speech Recognition for Jamaican Patois Music Transcription

Jordan Madden, Matthew Stone, Dimitri Johnson +1

Although Jamaican Patois is a widely spoken language, current speech recognition systems perform poorly on Patois music, producing inaccurate captions that limit accessibility and…

cs.CL2022

Zero-shot Cross-Linguistic Learning of Event Semantics

Malihe Alikhani, Thomas Kober, Bashar Alhafni +6

Typologically diverse languages offer systems of lexical and grammatical aspect that allow speakers to focus on facets of event structure in ways that comport with the specific com…

math.CV2021

Asymptotics for Bergman projections with smooth weights: a direct approach

Michael Hitrik, Matthew Stone

We adapt the direct approach to the semiclassical Bergman kernel asymptotics, developed recently by A. Deleporte, J. Sjöstrand, and the first-named author for real analytic expone…

cs.RO2019

That and There: Judging the Intent of Pointing Actions with Robotic Arms

Malihe Alikhani, Baber Khalid, Rahul Shome +3

Collaborative robotics requires effective communication between a robot and a human partner. This work proposes a set of interpretive principles for how a robotic arm can use point…

cs.CL2019

CITE: A Corpus of Image-Text Discourse Relations

Malihe Alikhani, Sreyasi Nag Chowdhury, Gerard de Melo +1

This paper presents a novel crowd-sourced resource for multimodal discourse: our resource characterizes inferences in image-text contexts in the domain of cooking recipes in the fo…

cs.CL2002

Anaphora and Discourse Structure

Bonnie Webber, Matthew Stone, Aravind Joshi +1

We argue in this paper that many common adverbial phrases generally taken to signal a discourse relation between syntactically connected units within discourse structure, instead w…

cs.CL2020

Discourse Coherence, Reference Grounding and Goal Oriented Dialogue

Baber Khalid, Malihe Alikhani, Michael Fellner +2

Prior approaches to realizing mixed-initiative human--computer referential communication have adopted information-state or collaborative problem-solving approaches. In this paper,…

cs.LO2002

Disjunction and modular goal-directed proof search

Matthew Stone

This paper explores goal-directed proof search in first-order multi-modal logic. The key issue is to design a proof system that respects the modularity and locality of assumptions…

cmp-lg1998

Textual Economy through Close Coupling of Syntax and Semantics

Matthew Stone, Bonnie Webber

We focus on the production of efficient descriptions of objects, actions and events. We define a type of efficiency, textual economy, that exploits the hearer's recognition of infe…

cs.CL2024

Dialogue with Robots: Proposals for Broadening Participation and Research in the SLIVAR Community

Casey Kennington, Malihe Alikhani, Heather Pon-Barry +20

The ability to interact with machines using natural human language is becoming not just commonplace, but expected. The next step is not just text interfaces, but speech interfaces…

cmp-lg1995

CLiFF Notes: Research in the Language, Information and Computation Laboratory of the University of Pennsylvania

Editors, :, Matthew Stone +1

Short abstracts by computational linguistics researchers at the University of Pennsylvania describing ongoing individual and joint projects.

cs.CL2026

Chatbots Output Meaningful (but Problematic) Language

Matthew Stone, Una Stojnić

Are utterances by AI chatbots meaningful? Concretely, if a user asks, say, Anthropic's agent Claude, "What is the capital of Spain?" and Claude answers, "Madrid is the capital of S…

cs.CL2020

Clue: Cross-modal Coherence Modeling for Caption Generation

Malihe Alikhani, Piyush Sharma, Shengjie Li +2

We use coherence relations inspired by computational models of discourse to study the information needs and goals of image captioning. Using an annotation protocol specifically dev…

math.CO2017

Query Complexity of Mastermind Variants

Aaron Berger, Christopher Chute, Matthew Stone

We study variants of Mastermind, a popular board game in which the objective is sequence reconstruction. In this two-player game, the so-called \textit{codemaker} constructs a hidd…

cs.RO2023

Socially Cognizant Robotics for a Technology Enhanced Society

Kristin J. Dana, Clinton Andrews, Kostas Bekris +6

Emerging applications of robotics, and concerns about their impact, require the research community to put human-centric objectives front-and-center. To meet this challenge, we advo…

cs.CL2001

Microplanning with Communicative Intentions: The SPUD System

Matthew Stone, Christine Doran, Bonnie Webber +2

The process of microplanning encompasses a range of problems in Natural Language Generation (NLG), such as referring expression generation, lexical choice, and aggregation, problem…

cs.CV2022

Cross-Modal Coherence for Text-to-Image Retrieval

Malihe Alikhani, Fangda Han, Hareesh Ravi +3

Common image-text joint understanding techniques presume that images and the associated text can universally be characterized by a single implicit model. However, co-occurring imag…

cs.CL2020

AI2D-RST: A multimodal corpus of 1000 primary school science diagrams

Tuomo Hiippala, Malihe Alikhani, Jonas Haverinen +6

This article introduces AI2D-RST, a multimodal corpus of 1000 English-language diagrams that represent topics in primary school natural sciences, such as food webs, life cycles, mo…

cs.CL2020

Aspectuality Across Genre: A Distributional Semantics Approach

Thomas Kober, Malihe Alikhani, Matthew Stone +1

The interpretation of the lexical aspect of verbs in English plays a crucial role for recognizing textual entailment and learning discourse-level inferences. We show that two eleme…