papers

Publications (266)

cs.AI2026

FormGym: Doing Paperwork with Agents

Matthew Toles, Rattandeep Singh, Isaac Song +1

Completing paperwork is a challenging and time-consuming problem. Form filling is especially challenging in the pure-image domain without access to OCR, typeset PDF text, or a DOM.…

cs.CL2022

Affective Idiosyncratic Responses to Music

Sky CH-Wang, Evan Li, Oliver Li +2

Affective responses to music are highly personal. Despite consensus that idiosyncratic factors play a key role in regulating how listeners emotionally respond to music, precisely m…

cs.CV2026

Multiple Consistent 2D-3D Mappings for Robust Zero-Shot 3D Visual Grounding

Yufei Yin, Jie Zheng, Qianke Meng +7

Zero-shot 3D Visual Grounding (3DVG) is a critical capability for open-world embodied AI. However, existing methods are fundamentally bottlenecked by the poor quality of open-vocab…

cs.CL2021

Revealing Persona Biases in Dialogue Systems

Emily Sheng, Josh Arnold, Zhou Yu +2

Dialogue systems in the form of chatbots and personal assistants are being increasingly integrated into people's lives. Modern dialogue systems may consider adopting anthropomorphi…

cs.CL2020

Data Annealing for Informal Language Understanding Tasks

Jing Gu, Zhou Yu

There is a huge performance gap between formal and informal language understanding tasks. The recent pre-trained models that improved the performance of formal language understandi…

cs.CL2025

AutoSpec: An Agentic Framework for Automatically Drafting Patent Specification

Ryan Shea, Zhou Yu

Patents play a critical role in driving technological innovation by granting inventors exclusive rights to their inventions. However the process of drafting a patent application is…

cs.CL2022

Just Fine-tune Twice: Selective Differential Privacy for Large Language Models

Weiyan Shi, Ryan Shea, Si Chen +3

Protecting large language models from privacy leakage is becoming increasingly crucial with their wide adoption in real-world products. Yet applying differential privacy (DP), a ca…

cs.CL2020

Perception Score, A Learned Metric for Open-ended Text Generation Evaluation

Jing Gu, Qingyang Wu, Zhou Yu

Automatic evaluation for open-ended natural language generation tasks remains a challenge. Existing metrics such as BLEU show a low correlation with human judgment. We propose a no…

stat.ML2022

Optimal Model Averaging of Support Vector Machines in Diverging Model Spaces

Chaoxia Yuan, Chao Ying, Zhou Yu +1

Support vector machine (SVM) is a powerful classification method that has achieved great success in many fields. Since its performance can be seriously impaired by redundant covari…

stat.ME2025

Graph-based Square-Root Estimation for Sparse Linear Regression

Peili Li, Zhuomei Li, Yunhai Xiao +2

Sparse linear regression is one of the classic problems in the field of statistics, which has deep connections and high intersections with optimization, computation, and machine le…

cs.AR2025

UB-Mesh: a Hierarchically Localized nD-FullMesh Datacenter Network Architecture

Heng Liao, Bingyang Liu, Xianping Chen +31

As the Large-scale Language Models (LLMs) continue to scale, the requisite computational power and bandwidth escalate. To address this, we introduce UB-Mesh, a novel AI datacenter…

cs.CL2024

VarBench: Robust Language Model Benchmarking Through Dynamic Variable Perturbation

Kun Qian, Shunji Wan, Claudia Tang +4

As large language models achieve impressive scores on traditional benchmarks, an increasing number of researchers are becoming concerned about benchmark data leakage during pre-tra…

cs.CL2019

End-to-End Trainable Non-Collaborative Dialog System

Yu Li, Kun Qian, Weiyan Shi +1

End-to-end task-oriented dialog models have achieved promising performance on collaborative tasks where users willingly coordinate with the system to complete a given task. While i…

cs.CV2024

Research on Intelligent Aided Diagnosis System of Medical Image Based on Computer Deep Learning

Jiajie Yuan, Linxiao Wu, Yulu Gong +3

This paper combines Struts and Hibernate two architectures together, using DAO (Data Access Object) to store and access data. Then a set of dual-mode humidity medical image library…

cs.CL2021

Evaluation of In-Person Counseling Strategies To Develop Physical Activity Chatbot for Women

Kai-Hui Liang, Patrick Lange, Yoo Jung Oh +3

Artificial intelligence chatbots are the vanguard in technology-based intervention to change people's behavior. To develop intervention chatbots, the first step is to understand na…

cs.LG2026

Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training

Chenlu Ye, Zhou Yu, Ziji Zhang +5

Reinforcement Learning with Verifiable Rewards (RLVR) improves final-answer accuracy on reasoning tasks, but it does not reliably improve reasoning quality. Because outcome rewards…

cs.CL2018

Multimodal Hierarchical Reinforcement Learning Policy for Task-Oriented Visual Dialog

Jiaping Zhang, Tiancheng Zhao, Zhou Yu

Creating an intelligent conversational system that understands vision and language is one of the ultimate goals in Artificial Intelligence (AI)~\cite{winograd1972understanding}. Ex…

cs.CL2020

Paraphrase Augmented Task-Oriented Dialog Generation

Silin Gao, Yichi Zhang, Zhijian Ou +1

Neural generative models have achieved promising performance on dialog generation tasks if given a huge data set. However, the lack of high-quality dialog data and the expensive da…

cs.LG2022

Deep Dimension Reduction for Supervised Representation Learning

Jian Huang, Yuling Jiao, Xu Liao +2

The goal of supervised representation learning is to construct effective data representations for prediction. Among all the characteristics of an ideal nonparametric representation…

cs.CL2018

Cross-Lingual Cross-Platform Rumor Verification Pivoting on Multimedia Content

Weiming Wen, Songwen Su, Zhou Yu

With the increasing popularity of smart devices, rumors with multimedia content become more and more common on social networks. The multimedia information usually makes rumors look…

cs.CR2025

PAPILLON: Privacy Preservation from Internet-based and Local Language Model Ensembles

Li Siyan, Vethavikashini Chithrra Raghuram, Omar Khattab +2

Users can divulge sensitive information to proprietary LLM providers, raising significant privacy concerns. While open-source models, hosted locally on the user's machine, alleviat…

cs.CL2021

DG2: Data Augmentation Through Document Grounded Dialogue Generation

Qingyang Wu, Song Feng, Derek Chen +3

Collecting data for training dialog systems can be extremely expensive due to the involvement of human participants and need for extensive annotation. Especially in document-ground…

cs.CL2023

Sociocultural Norm Similarities and Differences via Situational Alignment and Explainable Textual Entailment

Sky CH-Wang, Arkadiy Saakyan, Oliver Li +2

Designing systems that can reason across cultures requires that they are grounded in the norms of the contexts in which they operate. However, current research on developing comput…

cs.CL2025

TOD-ProcBench: Benchmarking Complex Instruction-Following in Task-Oriented Dialogues

Sarik Ghazarian, Abhinav Gullapalli, Swair Shah +4

In real-world task-oriented dialogue (TOD) settings, agents are required to strictly adhere to complex instructions while conducting multi-turn conversations with customers. These…

cs.CL2023

Prompt-Based Monte-Carlo Tree Search for Goal-Oriented Dialogue Policy Planning

Xiao Yu, Maximillian Chen, Zhou Yu

Planning for goal-oriented dialogue often requires simulating future dialogue interactions and estimating task progress. Many approaches thus consider training neural networks to p…

cs.CL2022

Towards Socially Intelligent Agents with Mental State Transition and Human Utility

Liang Qiu, Yizhou Zhao, Yuan Liang +4

Building a socially intelligent agent involves many challenges. One of which is to track the agent's mental state transition and teach the agent to make decisions guided by its val…

cs.CL2026

MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes

Maximillian Chen, Xuanming Zhang, Michael Peng +3

The rise of Internet of Things (IoT) devices in the physical world necessitates voice-based interfaces capable of handling complex user experiences. While modern Large Language Mod…

cs.CL2019

Filling Conversation Ellipsis for Better Social Dialog Understanding

Xiyuan Zhang, Chengxi Li, Dian Yu +2

The phenomenon of ellipsis is prevalent in social conversations. Ellipsis increases the difficulty of a series of downstream language understanding tasks, such as dialog act predic…

cs.CL2021

Alternating Recurrent Dialog Model with Large-scale Pre-trained Language Models

Qingyang Wu, Yichi Zhang, Yu Li +1

Existing dialog system models require extensive human annotations and are difficult to generalize to different tasks. The recent success of large pre-trained language models such a…

cs.CL2022

Using Chatbots to Teach Languages

Yu Li, Chun-Yen Chen, Dian Yu +6

This paper reports on progress towards building an online language learning tool to provide learners with conversational experience by using dialog systems as conversation practice…

cs.CV2026

Progressive Video Condensation with MLLM Agent for Long-form Video Understanding

Yufei Yin, Yuchen Xing, Qianke Meng +3

Understanding long videos requires extracting query-relevant information from long sequences under tight compute budgets. Existing text-then-LLM pipelines lose fine-grained visual…

cs.CR2021

Trade or Trick? Detecting and Characterizing Scam Tokens on Uniswap Decentralized Exchange

Pengcheng Xia, Haoyu wang, Bingyu Gao +6

The prosperity of the cryptocurrency ecosystem drives the need for digital asset trading platforms. Beyond centralized exchanges (CEXs), decentralized exchanges (DEXs) are introduc…

cs.CL2024

MultiPA: A Multi-task Speech Pronunciation Assessment Model for Open Response Scenarios

Yu-Wen Chen, Zhou Yu, Julia Hirschberg

Pronunciation assessment models designed for open response scenarios enable users to practice language skills in a manner similar to real-life communication. However, previous open…

cs.AI2026

Orchard: An Open-Source Agentic Modeling Framework

Baolin Peng, Wenlin Yao, Qianhui Wu +11

Agentic modeling aims to transform LLMs into autonomous agents capable of solving complex tasks through planning, reasoning, tool use, and multi-turn interaction with external envi…

cs.CL2022

Robots-Dont-Cry: Understanding Falsely Anthropomorphic Utterances in Dialog Systems

David Gros, Yu Li, Zhou Yu

Dialog systems are often designed or trained to output human-like responses. However, some responses may be impossible for a machine to truthfully say (e.g. "that movie made me cry…

cs.CL2022

Clean or Annotate: How to Spend a Limited Data Collection Budget

Derek Chen, Zhou Yu, Samuel R. Bowman

Crowdsourcing platforms are often used to collect datasets for training machine learning models, despite higher levels of inaccurate labeling compared to expert labeling. There are…

cs.CL2025

How and Where to Translate? The Impact of Translation Strategies in Cross-lingual LLM Prompting

Aman Gupta, Yingying Zhuang, Zhou Yu +2

Despite advances in the multilingual capabilities of Large Language Models (LLMs), their performance varies substantially across different languages and tasks. In multilingual retr…

cond-mat.soft2020

Viscous flow properties and hydrodynamic diameter of phenothiazine-based redox-active molecules in different supporting salt environments

Yilin Wang, Aman Preet Kaur, N. Harsha Attanayake +5

We report viscous flow properties of a redox-active organic molecule, N-(2-(2-methoxyethoxy)ethyl)phenothiazine (MEEPT), a candidate for non-aqueous redox flow batteries, and two o…

cs.CL2024

Effective Unsupervised Constrained Text Generation based on Perturbed Masking

Yingwen Fu, Wenjie Ou, Zhou Yu +1

Unsupervised constrained text generation aims to generate text under a given set of constraints without any supervised data. Current state-of-the-art methods stochastically sample…

cs.CV2019

Deep Modular Co-Attention Networks for Visual Question Answering

Zhou Yu, Jun Yu, Yuhao Cui +2

Visual Question Answering (VQA) requires a fine-grained and simultaneous understanding of both the visual content of images and the textual content of questions. Therefore, designi…

cs.LG2026

A Theoretical Analysis of Memory and Overfitting Phenomena in Stochastic Interpolation Models

Yunchen Li, Shaohui Lin, Zhou Yu

This paper provides a theoretical account of memorization in stochastic interpolation models. By leveraging closed-form expressions for the optimal velocity field and the associate…

cs.CL2022

In-context Learning Distillation: Transferring Few-shot Learning Ability of Pre-trained Language Models

Yukun Huang, Yanda Chen, Zhou Yu +1

Given the success with in-context learning of large pre-trained language models, we introduce in-context learning distillation to transfer in-context few-shot learning ability from…

cs.CV2026

REALM: An MLLM-Agent Framework for Open World 3D Reasoning Segmentation and Editing on Gaussian Splatting

Changyue Shi, Minghao Chen, Yiping Mao +4

Bridging the gap between complex human instructions and precise 3D object grounding remains a significant challenge in vision and robotics. Existing 3D segmentation methods often s…

cs.SE2025

AI Agents for Web Testing: A Case Study in the Wild

Naimeng Ye, Xiao Yu, Ruize Xu +2

Automated web testing plays a critical role in ensuring high-quality user experiences and delivering business value. Traditional approaches primarily focus on code coverage and loa…

cs.CL2022

Seamlessly Integrating Factual Information and Social Content with Persuasive Dialogue

Maximillian Chen, Weiyan Shi, Feifan Yan +4

Complex conversation settings such as persuasion involve communicating changes in attitude or behavior, so users' perspectives need to be addressed, even when not directly related…

cs.CV2020

Weakly-Supervised Multi-Level Attentional Reconstruction Network for Grounding Textual Queries in Videos

Yijun Song, Jingwen Wang, Lin Ma +2

The task of temporally grounding textual queries in videos is to localize one video segment that semantically corresponds to the given query. Most of the existing approaches rely o…

cs.CV2023

ANetQA: A Large-scale Benchmark for Fine-grained Compositional Reasoning over Untrimmed Videos

Zhou Yu, Lixiang Zheng, Zhou Zhao +4

Building benchmarks to systemically analyze different capabilities of video question answering (VideoQA) models is challenging yet crucial. Existing benchmarks often use non-compos…

cs.CV2019

ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering

Zhou Yu, Dejing Xu, Jun Yu +4

Recent developments in modeling language and vision have been successfully applied to image question answering. It is both crucial and natural to extend this research direction to…

cs.CV2022

IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning

Pan Lu, Liang Qiu, Jiaqi Chen +6

Current visual question answering (VQA) tasks mainly consider answering human-annotated questions for natural images. However, aside from natural images, abstract diagrams with sem…

cs.CL2023

End-to-end Task-oriented Dialogue: A Survey of Tasks, Methods, and Future Directions

Libo Qin, Wenbo Pan, Qiguang Chen +5

End-to-end task-oriented dialogue (EToD) can directly generate responses in an end-to-end fashion without modular training, which attracts escalating popularity. The advancement of…

cs.CV2019

Multimodal Unified Attention Networks for Vision-and-Language Interactions

Zhou Yu, Yuhao Cui, Jun Yu +2

Learning an effective attention mechanism for multimodal data is important in many vision-and-language tasks that require a synergic understanding of both the visual and textual co…

cs.CR2026

Beyond Direct Identifiers: Probabilistic Privacy Risk Estimation for Privacy-Conscious LLM Query Delegation

Li Siyan, Zhou Yu, Julia Hirschberg

Recent work on protecting privacy during user-LLM interactions often focuses on direct, explicit identifiers: the personally-identifiable information (PII) captured by standard det…

cs.CL2018

Incorporating Structured Commonsense Knowledge in Story Completion

Jiaao Chen, Jianshu Chen, Zhou Yu

The ability to select an appropriate story ending is the first step towards perfect narrative comprehension. Story ending prediction requires not only the explicit clues within the…

cs.AI2025

Program Synthesis Dialog Agents for Interactive Decision-Making

Matthew Toles, Nikhil Balwani, Rattandeep Singh +2

Many real-world eligibility problems, ranging from medical diagnosis to tax planning, can be mapped to decision problems expressed in natural language, wherein a model must make a…

cs.CL2024

TinyStyler: Efficient Few-Shot Text Style Transfer with Authorship Embeddings

Zachary Horvitz, Ajay Patel, Kanishk Singh +3

The goal of text style transfer is to transform the style of texts while preserving their original meaning, often with only a few examples of the target style. Existing style trans…

stat.ME2018

Sufficient variable screening via directional regression with censored response

Menghao Xu, Zhou Yu, Jun Shao

We in this paper propose a directional regression based approach for ultrahigh dimensional sufficient variable screening with censored responses. The new method is designed in a mo…

physics.flu-dyn2021

Reaction front development from ignition spots in n-heptane/air mixtures: low-temperature chemistry effects induced by ultrafine water droplet evaporation

Zhou Yu, Huangwei Zhang

Effects of low-temperature chemistry induced by ultrafine water droplet evaporation on reaction front development from an ignition spot with temperature gradient are studied in thi…

cs.LG2026

MortarBench: Evaluating Mortgage Loan Origination Agents

Matthew Toles, Yunan Lu, Manav Munjal +6

Loan origination is the process by which a lender creates a new loan, from application and underwriting through approval and funding. This process serves a critical role in evaluat…

cs.IR2026

Rethinking Multi-objective Ranking Ensemble in Recommender System: From Score Fusion to Rank Consistency

Boyang Xia, Zhou Yu, Zhiliang Zhu +5

The industrial recommender systems always pursue more than one business goals. The inherent intensions between objectives pose significant challenges for ranking stage. A popular s…

cs.CV2018

Rethinking Diversified and Discriminative Proposal Generation for Visual Grounding

Zhou Yu, Jun Yu, Chenchao Xiang +3

Visual grounding aims to localize an object in an image referred to by a textual query phrase. Various visual grounding approaches have been proposed, and the problem can be modula…

cs.CL2023

PLACES: Prompting Language Models for Social Conversation Synthesis

Maximillian Chen, Alexandros Papangelis, Chenyang Tao +5

Collecting high quality conversational data can be very expensive for most applications and infeasible for others due to privacy, ethical, or similar concerns. A promising directio…

cs.LG2026

A-LLM: An End-to-end Conversational Audio Avatar Large Language Model

Xiaolin Hu, Hang Yuan, Xinzhu Sang +4

Developing expressive and responsive conversational digital humans is a cornerstone of next-generation human-computer interaction. While large language models (LLMs) have significa…

cs.CL2024

Getting Serious about Humor: Crafting Humor Datasets with Unfunny Large Language Models

Zachary Horvitz, Jingru Chen, Rahul Aditya +4

Humor is a fundamental facet of human cognition and interaction. Yet, despite recent advances in natural language processing, humor detection remains a challenging task that is com…

cs.CL2021

Database Search Results Disambiguation for Task-Oriented Dialog Systems

Kun Qian, Ahmad Beirami, Satwik Kottur +5

As task-oriented dialog systems are becoming increasingly popular in our lives, more realistic tasks have been proposed and explored. However, new practical challenges arise. For i…

stat.ME2025

Sparse Fréchet Sufficient Dimension Reduction with Graphical Structure Among Predictors

Jiaying Weng, Kai Tan, Cheng Wang +1

Fréchet regression has received considerable attention to model metric-space valued responses that are complex and non-Euclidean data, such as probability distributions and vector…

physics.med-ph2023

Comprehensive evaluations of a prototype full field-of-view photon counting CT system through phantom studies

Xiaohui Zhan, Ruoqiao Zhang, Xiaofeng Niu +14

Photon counting CT (PCCT) has been a research focus in the last two decades. Recent studies and advancements have demonstrated that systems using semiconductor-based photon countin…

cs.AI2026

CXRAgent: Director-Orchestrated Multi-Stage Reasoning for Chest X-Ray Interpretation

Jinhui Lou, Yan Yang, Zhou Yu +4

CXRAgent is a director-orchestrated, multi-stage AI agent that coordinates various chest X‑ray analysis tools, plans diagnostics, and integrates expert team insights with visual ev…

#chest x-ray interpretation#multi-stage reasoning#LLM agents#tool orchestration
cs.CL2023

Social Influence Dialogue Systems: A Survey of Datasets and Models For Social Influence Tasks

Kushal Chawla, Weiyan Shi, Jingwen Zhang +3

Dialogue systems capable of social influence such as persuasion, negotiation, and therapy, are essential for extending the use of technology to numerous realistic scenarios. Howeve…

cs.HC2026

JumpStarter: Human-AI Planning with Task-Structured Context Curation

Xuanming Zhang, Sitong Wang, Jenny Ma +3

Human-AI collaboration on complex planning goals is bottlenecked by how LLM interfaces handle context: users must manually curate and re-surface relevant information across long an…

cs.CL2026

Detect, Remask, Repair: Diffusion Editing for Faithful Summarization of Evolving Contexts

Hao Zou, Zachary Horvitz, Chandhru Karthick +2

Summaries of real-world events can become outdated as contexts evolve and new information arrives. A common response is to generate a new summary from the updated context, but full…

cs.HC2026

Invisible Impact of Empathy on Behavioral Change: Isolating the Effect of Empathy in Long-term Physical Activity Coaching Chatbot Interactions

Li Siyan, Kai-Hui Liang, Shopnil Shahriar +7

Current dialogue systems, powered by large language models, often treat empathy as essential without assessing its true impact, especially in behavior change, where motivation and…

cs.RO2024

GarchingSim: An Autonomous Driving Simulator with Photorealistic Scenes and Minimalist Workflow

Liguo Zhou, Yinglei Song, Yichao Gao +10

Conducting real road testing for autonomous driving algorithms can be expensive and sometimes impractical, particularly for small startups and research institutes. Thus, simulation…

cs.CL2022

Weakly Supervised Data Augmentation Through Prompting for Dialogue Understanding

Maximillian Chen, Alexandros Papangelis, Chenyang Tao +5

Dialogue understanding tasks often necessitate abundant annotated data to achieve good performance and that presents challenges in low-resource settings. To alleviate this barrier,…

cs.CL2025

Mind the Gap: Linguistic Divergence and Adaptation Strategies in Human-LLM Assistant vs. Human-Human Interactions

Fulei Zhang, Zhou Yu

As Large Language Models (LLMs) are increasingly deployed in customer-facing applications, a critical yet underexplored question is how users communicate differently with LLM chatb…

cs.CL2022

Dialoging Resonance: How Users Perceive, Reciprocate and React to Chatbot's Self-Disclosure in Conversational Recommendations

Kai-Hui Liang, Weiyan Shi, Yoojung Oh +3

Using chatbots to deliver recommendations is increasingly popular. The design of recommendation chatbots has primarily been taking an information-centric approach by focusing on th…

cs.CV2022

Unsupervised Vision-and-Language Pre-training via Retrieval-based Multi-Granular Alignment

Mingyang Zhou, Licheng Yu, Amanpreet Singh +3

Vision-and-Language (V+L) pre-training models have achieved tremendous success in recent years on various multi-modal benchmarks. However, the majority of existing models require p…

cs.CL2021

Zero-Shot Dialogue State Tracking via Cross-Task Transfer

Zhaojiang Lin, Bing Liu, Andrea Madotto +8

Zero-shot transfer learning for dialogue state tracking (DST) enables us to handle a variety of task-oriented dialogue domains without the expense of collecting in-domain data. In…

cs.CL2020

INSPIRED: Toward Sociable Recommendation Dialog Systems

Shirley Anugrah Hayati, Dongyeop Kang, Qingxiaoyang Zhu +2

In recommendation dialogs, humans commonly disclose their preference and make recommendations in a friendly manner. However, this is a challenge when developing a sociable recommen…

cs.CL2022

Fantastic Questions and Where to Find Them: FairytaleQA -- An Authentic Dataset for Narrative Comprehension

Ying Xu, Dakuo Wang, Mo Yu +15

Question answering (QA) is a fundamental means to facilitate assessment and training of narrative comprehension skills for both machines and young children, yet there is scarcity o…

cs.CR2026

Proactive defense against LLM Jailbreak

Weiliang Zhao, Jinjun Peng, Daniel Ben-Levi +2

The proliferation of powerful large language models (LLMs) has necessitated robust safety alignment, yet these models remain vulnerable to evolving adversarial attacks, including m…

cs.CL2022

ErAConD : Error Annotated Conversational Dialog Dataset for Grammatical Error Correction

Xun Yuan, Derek Pham, Sam Davidson +1

Currently available grammatical error correction (GEC) datasets are compiled using well-formed written text, limiting the applicability of these datasets to other domains such as i…

cs.CL2021

Attribute Alignment: Controlling Text Generation from Pre-trained Language Models

Dian Yu, Zhou Yu, Kenji Sagae

Large language models benefit from training with a large amount of unlabeled text, which gives them increasingly fluent and diverse generation capabilities. However, using these mo…

stat.CO2024

A Trigamma-free Approach for Computing Information Matrices Related to Trigamma Function

Zhou Yu, Niloufar Dousti Mousavi, Jie Yang

Negative binomial related distributions have been widely used in practice. The calculation of the corresponding Fisher information matrices involves the expectation of trigamma fun…

cs.HC2020

Effects of Persuasive Dialogues: Testing Bot Identities and Inquiry Strategies

Weiyan Shi, Xuewei Wang, Yoo Jung Oh +3

Intelligent conversational agents, or chatbots, can take on various identities and are increasingly engaging in more human-centered conversations with persuasive goals. However, li…

cs.CL2019

Importance-Aware Learning for Neural Headline Editing

Qingyang Wu, Lei Li, Hao Zhou +2

Many social media news writers are not professionally trained. Therefore, social media platforms have to hire professional editors to adjust amateur headlines to attract more reade…

cs.AI2026

AutoRPA: Efficient GUI Automation through LLM-Driven Code Synthesis from Interactions

Minghao Chen, Xinyi Hu, Zhou Yu +1

Large Language Model (LLM) based agents have demonstrated proficiency in multi-step interactions with graphical user interfaces (GUIs). While most research focuses on improving sin…

cs.CL2025

Bringing Pedagogy into Focus: Evaluating Virtual Teaching Assistants' Question-Answering in Asynchronous Learning Environments

Li Siyan, Zhen Xu, Vethavikashini Chithrra Raghuram +3

Asynchronous learning environments (ALEs) are widely adopted for formal and informal learning, but timely and personalized support is often limited. In this context, Virtual Teachi…

cs.CL2023

Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations

Yanda Chen, Ruiqi Zhong, Narutatsu Ri +5

Large language models (LLMs) are trained to imitate humans to explain human decisions. However, do LLMs explain themselves? Can they help humans build mental models of how LLMs pro…

cs.CL2023

Sources of Noise in Dialogue and How to Deal with Them

Derek Chen, Zhou Yu

Training dialogue systems often entails dealing with noisy training examples and unexpected user inputs. Despite their prevalence, there currently lacks an accurate survey of dialo…

cs.CL2023

User Adaptive Language Learning Chatbots with a Curriculum

Kun Qian, Ryan Shea, Yu Li +2

Along with the development of systems for natural language understanding and generation, dialog systems have been widely adopted for language learning and practicing. Many current…

cs.CL2023

Building Persona Consistent Dialogue Agents with Offline Reinforcement Learning

Ryan Shea, Zhou Yu

Maintaining a consistent persona is a key quality for any open domain dialogue system. Current state-of-the-art systems do this by training agents with supervised learning or onlin…

cs.CL2019

Task-Oriented Dialog Systems that Consider Multiple Appropriate Responses under the Same Context

Yichi Zhang, Zhijian Ou, Zhou Yu

Conversations have an intrinsic one-to-many property, which means that multiple responses can be appropriate for the same dialog context. In task-oriented dialogs, this property le…

cs.CV2023

Filter Pruning for Efficient CNNs via Knowledge-driven Differential Filter Sampler

Shaohui Lin, Wenxuan Huang, Jiao Xie +5

Filter pruning simultaneously accelerates the computation and reduces the memory overhead of CNNs, which can be effectively applied to edge devices and cloud services. In this pape…

cs.LG2019

Quantifying Intrinsic Uncertainty in Classification via Deep Dirichlet Mixture Networks

Qingyang Wu, He Li, Lexin Li +1

With the widespread success of deep neural networks in science and technology, it is becoming increasingly important to quantify the uncertainty of the predictions produced by deep…

cond-mat.mtrl-sci2023

Simulation Guided Molecular Design of Hydrofluoroether Solvent for High Energy Batteries

Zhou Yu, Zhangxing shi, Sambasiva R. Bheemireddy +9

Electrolyte design is critical for enabling next-generation batteries with higher energy densities. Hydrofluoroether (HFE) solvents have drawn a lot of attention as the electrolyte…

cs.CL2020

A Tailored Pre-Training Model for Task-Oriented Dialog Generation

Jing Gu, Qingyang Wu, Chongruo Wu +2

The recent success of large pre-trained language models such as BERT and GPT-2 has suggested the effectiveness of incorporating language priors in downstream dialog generation task…

cs.CL2020

Matching Text with Deep Mutual Information Estimation

Xixi Zhou, Chengxi Li, Jiajun Bu +4

Text matching is a core natural language processing research problem. How to retain sufficient information on both content and structure information is one important challenge. In…

cs.CL2026

Factorization-Error-Free Discrete Diffusion Language Model via Speculative Decoding

Xun Fang, Yunchen Li, Hang Yuan +1

Discrete diffusion language models improve generation efficiency through parallel token prediction, but standard prediction methods introduce factorization errors by approxim…

cs.CV2021

ROSITA: Enhancing Vision-and-Language Semantic Alignments via Cross- and Intra-modal Knowledge Integration

Yuhao Cui, Zhou Yu, Chunqi Wang +4

Vision-and-language pretraining (VLP) aims to learn generic multimodal representations from massive image-text pairs. While various successful attempts have been proposed, learning…

cs.CV2026

CSD: Content-aware Speculative Decoding for Efficient Image Generation

Mingcheng Wang, Junbo Qiao, Yunchen Li +8

Speculative decoding (SD) has emerged as a key solution to accelerate the inference of autoregressive models. However, in the field of image generation, it faces the challenge of l…