Publications (266)
FormGym: Doing Paperwork with Agents
Matthew Toles, Rattandeep Singh, Isaac Song +1
Completing paperwork is a challenging and time-consuming problem. Form filling is especially challenging in the pure-image domain without access to OCR, typeset PDF text, or a DOM.…
Affective Idiosyncratic Responses to Music
Sky CH-Wang, Evan Li, Oliver Li +2
Affective responses to music are highly personal. Despite consensus that idiosyncratic factors play a key role in regulating how listeners emotionally respond to music, precisely m…
Multiple Consistent 2D-3D Mappings for Robust Zero-Shot 3D Visual Grounding
Yufei Yin, Jie Zheng, Qianke Meng +7
Zero-shot 3D Visual Grounding (3DVG) is a critical capability for open-world embodied AI. However, existing methods are fundamentally bottlenecked by the poor quality of open-vocab…
Revealing Persona Biases in Dialogue Systems
Emily Sheng, Josh Arnold, Zhou Yu +2
Dialogue systems in the form of chatbots and personal assistants are being increasingly integrated into people's lives. Modern dialogue systems may consider adopting anthropomorphi…
Data Annealing for Informal Language Understanding Tasks
Jing Gu, Zhou Yu
There is a huge performance gap between formal and informal language understanding tasks. The recent pre-trained models that improved the performance of formal language understandi…
AutoSpec: An Agentic Framework for Automatically Drafting Patent Specification
Ryan Shea, Zhou Yu
Patents play a critical role in driving technological innovation by granting inventors exclusive rights to their inventions. However the process of drafting a patent application is…
Just Fine-tune Twice: Selective Differential Privacy for Large Language Models
Weiyan Shi, Ryan Shea, Si Chen +3
Protecting large language models from privacy leakage is becoming increasingly crucial with their wide adoption in real-world products. Yet applying differential privacy (DP), a ca…
Perception Score, A Learned Metric for Open-ended Text Generation Evaluation
Jing Gu, Qingyang Wu, Zhou Yu
Automatic evaluation for open-ended natural language generation tasks remains a challenge. Existing metrics such as BLEU show a low correlation with human judgment. We propose a no…
Optimal Model Averaging of Support Vector Machines in Diverging Model Spaces
Chaoxia Yuan, Chao Ying, Zhou Yu +1
Support vector machine (SVM) is a powerful classification method that has achieved great success in many fields. Since its performance can be seriously impaired by redundant covari…
Graph-based Square-Root Estimation for Sparse Linear Regression
Peili Li, Zhuomei Li, Yunhai Xiao +2
Sparse linear regression is one of the classic problems in the field of statistics, which has deep connections and high intersections with optimization, computation, and machine le…
UB-Mesh: a Hierarchically Localized nD-FullMesh Datacenter Network Architecture
Heng Liao, Bingyang Liu, Xianping Chen +31
As the Large-scale Language Models (LLMs) continue to scale, the requisite computational power and bandwidth escalate. To address this, we introduce UB-Mesh, a novel AI datacenter…
VarBench: Robust Language Model Benchmarking Through Dynamic Variable Perturbation
Kun Qian, Shunji Wan, Claudia Tang +4
As large language models achieve impressive scores on traditional benchmarks, an increasing number of researchers are becoming concerned about benchmark data leakage during pre-tra…
End-to-End Trainable Non-Collaborative Dialog System
Yu Li, Kun Qian, Weiyan Shi +1
End-to-end task-oriented dialog models have achieved promising performance on collaborative tasks where users willingly coordinate with the system to complete a given task. While i…
Research on Intelligent Aided Diagnosis System of Medical Image Based on Computer Deep Learning
Jiajie Yuan, Linxiao Wu, Yulu Gong +3
This paper combines Struts and Hibernate two architectures together, using DAO (Data Access Object) to store and access data. Then a set of dual-mode humidity medical image library…
Evaluation of In-Person Counseling Strategies To Develop Physical Activity Chatbot for Women
Kai-Hui Liang, Patrick Lange, Yoo Jung Oh +3
Artificial intelligence chatbots are the vanguard in technology-based intervention to change people's behavior. To develop intervention chatbots, the first step is to understand na…
Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training
Chenlu Ye, Zhou Yu, Ziji Zhang +5
Reinforcement Learning with Verifiable Rewards (RLVR) improves final-answer accuracy on reasoning tasks, but it does not reliably improve reasoning quality. Because outcome rewards…
Multimodal Hierarchical Reinforcement Learning Policy for Task-Oriented Visual Dialog
Jiaping Zhang, Tiancheng Zhao, Zhou Yu
Creating an intelligent conversational system that understands vision and language is one of the ultimate goals in Artificial Intelligence (AI)~\cite{winograd1972understanding}. Ex…
Paraphrase Augmented Task-Oriented Dialog Generation
Silin Gao, Yichi Zhang, Zhijian Ou +1
Neural generative models have achieved promising performance on dialog generation tasks if given a huge data set. However, the lack of high-quality dialog data and the expensive da…
Deep Dimension Reduction for Supervised Representation Learning
Jian Huang, Yuling Jiao, Xu Liao +2
The goal of supervised representation learning is to construct effective data representations for prediction. Among all the characteristics of an ideal nonparametric representation…
Cross-Lingual Cross-Platform Rumor Verification Pivoting on Multimedia Content
Weiming Wen, Songwen Su, Zhou Yu
With the increasing popularity of smart devices, rumors with multimedia content become more and more common on social networks. The multimedia information usually makes rumors look…
PAPILLON: Privacy Preservation from Internet-based and Local Language Model Ensembles
Li Siyan, Vethavikashini Chithrra Raghuram, Omar Khattab +2
Users can divulge sensitive information to proprietary LLM providers, raising significant privacy concerns. While open-source models, hosted locally on the user's machine, alleviat…
DG2: Data Augmentation Through Document Grounded Dialogue Generation
Qingyang Wu, Song Feng, Derek Chen +3
Collecting data for training dialog systems can be extremely expensive due to the involvement of human participants and need for extensive annotation. Especially in document-ground…
Sociocultural Norm Similarities and Differences via Situational Alignment and Explainable Textual Entailment
Sky CH-Wang, Arkadiy Saakyan, Oliver Li +2
Designing systems that can reason across cultures requires that they are grounded in the norms of the contexts in which they operate. However, current research on developing comput…
TOD-ProcBench: Benchmarking Complex Instruction-Following in Task-Oriented Dialogues
Sarik Ghazarian, Abhinav Gullapalli, Swair Shah +4
In real-world task-oriented dialogue (TOD) settings, agents are required to strictly adhere to complex instructions while conducting multi-turn conversations with customers. These…
Prompt-Based Monte-Carlo Tree Search for Goal-Oriented Dialogue Policy Planning
Xiao Yu, Maximillian Chen, Zhou Yu
Planning for goal-oriented dialogue often requires simulating future dialogue interactions and estimating task progress. Many approaches thus consider training neural networks to p…
Towards Socially Intelligent Agents with Mental State Transition and Human Utility
Liang Qiu, Yizhou Zhao, Yuan Liang +4
Building a socially intelligent agent involves many challenges. One of which is to track the agent's mental state transition and teach the agent to make decisions guided by its val…
MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes
Maximillian Chen, Xuanming Zhang, Michael Peng +3
The rise of Internet of Things (IoT) devices in the physical world necessitates voice-based interfaces capable of handling complex user experiences. While modern Large Language Mod…
Filling Conversation Ellipsis for Better Social Dialog Understanding
Xiyuan Zhang, Chengxi Li, Dian Yu +2
The phenomenon of ellipsis is prevalent in social conversations. Ellipsis increases the difficulty of a series of downstream language understanding tasks, such as dialog act predic…
Alternating Recurrent Dialog Model with Large-scale Pre-trained Language Models
Qingyang Wu, Yichi Zhang, Yu Li +1
Existing dialog system models require extensive human annotations and are difficult to generalize to different tasks. The recent success of large pre-trained language models such a…
Using Chatbots to Teach Languages
Yu Li, Chun-Yen Chen, Dian Yu +6
This paper reports on progress towards building an online language learning tool to provide learners with conversational experience by using dialog systems as conversation practice…
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
Yufei Yin, Yuchen Xing, Qianke Meng +3
Understanding long videos requires extracting query-relevant information from long sequences under tight compute budgets. Existing text-then-LLM pipelines lose fine-grained visual…
Trade or Trick? Detecting and Characterizing Scam Tokens on Uniswap Decentralized Exchange
Pengcheng Xia, Haoyu wang, Bingyu Gao +6
The prosperity of the cryptocurrency ecosystem drives the need for digital asset trading platforms. Beyond centralized exchanges (CEXs), decentralized exchanges (DEXs) are introduc…
MultiPA: A Multi-task Speech Pronunciation Assessment Model for Open Response Scenarios
Yu-Wen Chen, Zhou Yu, Julia Hirschberg
Pronunciation assessment models designed for open response scenarios enable users to practice language skills in a manner similar to real-life communication. However, previous open…
Orchard: An Open-Source Agentic Modeling Framework
Baolin Peng, Wenlin Yao, Qianhui Wu +11
Agentic modeling aims to transform LLMs into autonomous agents capable of solving complex tasks through planning, reasoning, tool use, and multi-turn interaction with external envi…
Robots-Dont-Cry: Understanding Falsely Anthropomorphic Utterances in Dialog Systems
David Gros, Yu Li, Zhou Yu
Dialog systems are often designed or trained to output human-like responses. However, some responses may be impossible for a machine to truthfully say (e.g. "that movie made me cry…
Clean or Annotate: How to Spend a Limited Data Collection Budget
Derek Chen, Zhou Yu, Samuel R. Bowman
Crowdsourcing platforms are often used to collect datasets for training machine learning models, despite higher levels of inaccurate labeling compared to expert labeling. There are…
How and Where to Translate? The Impact of Translation Strategies in Cross-lingual LLM Prompting
Aman Gupta, Yingying Zhuang, Zhou Yu +2
Despite advances in the multilingual capabilities of Large Language Models (LLMs), their performance varies substantially across different languages and tasks. In multilingual retr…
Viscous flow properties and hydrodynamic diameter of phenothiazine-based redox-active molecules in different supporting salt environments
Yilin Wang, Aman Preet Kaur, N. Harsha Attanayake +5
We report viscous flow properties of a redox-active organic molecule, N-(2-(2-methoxyethoxy)ethyl)phenothiazine (MEEPT), a candidate for non-aqueous redox flow batteries, and two o…
Effective Unsupervised Constrained Text Generation based on Perturbed Masking
Yingwen Fu, Wenjie Ou, Zhou Yu +1
Unsupervised constrained text generation aims to generate text under a given set of constraints without any supervised data. Current state-of-the-art methods stochastically sample…
Deep Modular Co-Attention Networks for Visual Question Answering
Zhou Yu, Jun Yu, Yuhao Cui +2
Visual Question Answering (VQA) requires a fine-grained and simultaneous understanding of both the visual content of images and the textual content of questions. Therefore, designi…
A Theoretical Analysis of Memory and Overfitting Phenomena in Stochastic Interpolation Models
Yunchen Li, Shaohui Lin, Zhou Yu
This paper provides a theoretical account of memorization in stochastic interpolation models. By leveraging closed-form expressions for the optimal velocity field and the associate…
In-context Learning Distillation: Transferring Few-shot Learning Ability of Pre-trained Language Models
Yukun Huang, Yanda Chen, Zhou Yu +1
Given the success with in-context learning of large pre-trained language models, we introduce in-context learning distillation to transfer in-context few-shot learning ability from…
REALM: An MLLM-Agent Framework for Open World 3D Reasoning Segmentation and Editing on Gaussian Splatting
Changyue Shi, Minghao Chen, Yiping Mao +4
Bridging the gap between complex human instructions and precise 3D object grounding remains a significant challenge in vision and robotics. Existing 3D segmentation methods often s…
AI Agents for Web Testing: A Case Study in the Wild
Naimeng Ye, Xiao Yu, Ruize Xu +2
Automated web testing plays a critical role in ensuring high-quality user experiences and delivering business value. Traditional approaches primarily focus on code coverage and loa…
Seamlessly Integrating Factual Information and Social Content with Persuasive Dialogue
Maximillian Chen, Weiyan Shi, Feifan Yan +4
Complex conversation settings such as persuasion involve communicating changes in attitude or behavior, so users' perspectives need to be addressed, even when not directly related…
Weakly-Supervised Multi-Level Attentional Reconstruction Network for Grounding Textual Queries in Videos
Yijun Song, Jingwen Wang, Lin Ma +2
The task of temporally grounding textual queries in videos is to localize one video segment that semantically corresponds to the given query. Most of the existing approaches rely o…
ANetQA: A Large-scale Benchmark for Fine-grained Compositional Reasoning over Untrimmed Videos
Zhou Yu, Lixiang Zheng, Zhou Zhao +4
Building benchmarks to systemically analyze different capabilities of video question answering (VideoQA) models is challenging yet crucial. Existing benchmarks often use non-compos…
ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering
Zhou Yu, Dejing Xu, Jun Yu +4
Recent developments in modeling language and vision have been successfully applied to image question answering. It is both crucial and natural to extend this research direction to…
IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning
Pan Lu, Liang Qiu, Jiaqi Chen +6
Current visual question answering (VQA) tasks mainly consider answering human-annotated questions for natural images. However, aside from natural images, abstract diagrams with sem…
End-to-end Task-oriented Dialogue: A Survey of Tasks, Methods, and Future Directions
Libo Qin, Wenbo Pan, Qiguang Chen +5
End-to-end task-oriented dialogue (EToD) can directly generate responses in an end-to-end fashion without modular training, which attracts escalating popularity. The advancement of…
Multimodal Unified Attention Networks for Vision-and-Language Interactions
Zhou Yu, Yuhao Cui, Jun Yu +2
Learning an effective attention mechanism for multimodal data is important in many vision-and-language tasks that require a synergic understanding of both the visual and textual co…
Beyond Direct Identifiers: Probabilistic Privacy Risk Estimation for Privacy-Conscious LLM Query Delegation
Li Siyan, Zhou Yu, Julia Hirschberg
Recent work on protecting privacy during user-LLM interactions often focuses on direct, explicit identifiers: the personally-identifiable information (PII) captured by standard det…
Incorporating Structured Commonsense Knowledge in Story Completion
Jiaao Chen, Jianshu Chen, Zhou Yu
The ability to select an appropriate story ending is the first step towards perfect narrative comprehension. Story ending prediction requires not only the explicit clues within the…
Program Synthesis Dialog Agents for Interactive Decision-Making
Matthew Toles, Nikhil Balwani, Rattandeep Singh +2
Many real-world eligibility problems, ranging from medical diagnosis to tax planning, can be mapped to decision problems expressed in natural language, wherein a model must make a…
TinyStyler: Efficient Few-Shot Text Style Transfer with Authorship Embeddings
Zachary Horvitz, Ajay Patel, Kanishk Singh +3
The goal of text style transfer is to transform the style of texts while preserving their original meaning, often with only a few examples of the target style. Existing style trans…
Sufficient variable screening via directional regression with censored response
Menghao Xu, Zhou Yu, Jun Shao
We in this paper propose a directional regression based approach for ultrahigh dimensional sufficient variable screening with censored responses. The new method is designed in a mo…
Reaction front development from ignition spots in n-heptane/air mixtures: low-temperature chemistry effects induced by ultrafine water droplet evaporation
Zhou Yu, Huangwei Zhang
Effects of low-temperature chemistry induced by ultrafine water droplet evaporation on reaction front development from an ignition spot with temperature gradient are studied in thi…
MortarBench: Evaluating Mortgage Loan Origination Agents
Matthew Toles, Yunan Lu, Manav Munjal +6
Loan origination is the process by which a lender creates a new loan, from application and underwriting through approval and funding. This process serves a critical role in evaluat…
Rethinking Multi-objective Ranking Ensemble in Recommender System: From Score Fusion to Rank Consistency
Boyang Xia, Zhou Yu, Zhiliang Zhu +5
The industrial recommender systems always pursue more than one business goals. The inherent intensions between objectives pose significant challenges for ranking stage. A popular s…
Rethinking Diversified and Discriminative Proposal Generation for Visual Grounding
Zhou Yu, Jun Yu, Chenchao Xiang +3
Visual grounding aims to localize an object in an image referred to by a textual query phrase. Various visual grounding approaches have been proposed, and the problem can be modula…
PLACES: Prompting Language Models for Social Conversation Synthesis
Maximillian Chen, Alexandros Papangelis, Chenyang Tao +5
Collecting high quality conversational data can be very expensive for most applications and infeasible for others due to privacy, ethical, or similar concerns. A promising directio…
A-LLM: An End-to-end Conversational Audio Avatar Large Language Model
Xiaolin Hu, Hang Yuan, Xinzhu Sang +4
Developing expressive and responsive conversational digital humans is a cornerstone of next-generation human-computer interaction. While large language models (LLMs) have significa…
Getting Serious about Humor: Crafting Humor Datasets with Unfunny Large Language Models
Zachary Horvitz, Jingru Chen, Rahul Aditya +4
Humor is a fundamental facet of human cognition and interaction. Yet, despite recent advances in natural language processing, humor detection remains a challenging task that is com…
Database Search Results Disambiguation for Task-Oriented Dialog Systems
Kun Qian, Ahmad Beirami, Satwik Kottur +5
As task-oriented dialog systems are becoming increasingly popular in our lives, more realistic tasks have been proposed and explored. However, new practical challenges arise. For i…
Sparse Fréchet Sufficient Dimension Reduction with Graphical Structure Among Predictors
Jiaying Weng, Kai Tan, Cheng Wang +1
Fréchet regression has received considerable attention to model metric-space valued responses that are complex and non-Euclidean data, such as probability distributions and vector…
Comprehensive evaluations of a prototype full field-of-view photon counting CT system through phantom studies
Xiaohui Zhan, Ruoqiao Zhang, Xiaofeng Niu +14
Photon counting CT (PCCT) has been a research focus in the last two decades. Recent studies and advancements have demonstrated that systems using semiconductor-based photon countin…
CXRAgent: Director-Orchestrated Multi-Stage Reasoning for Chest X-Ray Interpretation
Jinhui Lou, Yan Yang, Zhou Yu +4
CXRAgent is a director-orchestrated, multi-stage AI agent that coordinates various chest X‑ray analysis tools, plans diagnostics, and integrates expert team insights with visual ev…
Social Influence Dialogue Systems: A Survey of Datasets and Models For Social Influence Tasks
Kushal Chawla, Weiyan Shi, Jingwen Zhang +3
Dialogue systems capable of social influence such as persuasion, negotiation, and therapy, are essential for extending the use of technology to numerous realistic scenarios. Howeve…
JumpStarter: Human-AI Planning with Task-Structured Context Curation
Xuanming Zhang, Sitong Wang, Jenny Ma +3
Human-AI collaboration on complex planning goals is bottlenecked by how LLM interfaces handle context: users must manually curate and re-surface relevant information across long an…
Detect, Remask, Repair: Diffusion Editing for Faithful Summarization of Evolving Contexts
Hao Zou, Zachary Horvitz, Chandhru Karthick +2
Summaries of real-world events can become outdated as contexts evolve and new information arrives. A common response is to generate a new summary from the updated context, but full…
Invisible Impact of Empathy on Behavioral Change: Isolating the Effect of Empathy in Long-term Physical Activity Coaching Chatbot Interactions
Li Siyan, Kai-Hui Liang, Shopnil Shahriar +7
Current dialogue systems, powered by large language models, often treat empathy as essential without assessing its true impact, especially in behavior change, where motivation and…
GarchingSim: An Autonomous Driving Simulator with Photorealistic Scenes and Minimalist Workflow
Liguo Zhou, Yinglei Song, Yichao Gao +10
Conducting real road testing for autonomous driving algorithms can be expensive and sometimes impractical, particularly for small startups and research institutes. Thus, simulation…
Weakly Supervised Data Augmentation Through Prompting for Dialogue Understanding
Maximillian Chen, Alexandros Papangelis, Chenyang Tao +5
Dialogue understanding tasks often necessitate abundant annotated data to achieve good performance and that presents challenges in low-resource settings. To alleviate this barrier,…
Mind the Gap: Linguistic Divergence and Adaptation Strategies in Human-LLM Assistant vs. Human-Human Interactions
Fulei Zhang, Zhou Yu
As Large Language Models (LLMs) are increasingly deployed in customer-facing applications, a critical yet underexplored question is how users communicate differently with LLM chatb…
Dialoging Resonance: How Users Perceive, Reciprocate and React to Chatbot's Self-Disclosure in Conversational Recommendations
Kai-Hui Liang, Weiyan Shi, Yoojung Oh +3
Using chatbots to deliver recommendations is increasingly popular. The design of recommendation chatbots has primarily been taking an information-centric approach by focusing on th…
Unsupervised Vision-and-Language Pre-training via Retrieval-based Multi-Granular Alignment
Mingyang Zhou, Licheng Yu, Amanpreet Singh +3
Vision-and-Language (V+L) pre-training models have achieved tremendous success in recent years on various multi-modal benchmarks. However, the majority of existing models require p…
Zero-Shot Dialogue State Tracking via Cross-Task Transfer
Zhaojiang Lin, Bing Liu, Andrea Madotto +8
Zero-shot transfer learning for dialogue state tracking (DST) enables us to handle a variety of task-oriented dialogue domains without the expense of collecting in-domain data. In…
INSPIRED: Toward Sociable Recommendation Dialog Systems
Shirley Anugrah Hayati, Dongyeop Kang, Qingxiaoyang Zhu +2
In recommendation dialogs, humans commonly disclose their preference and make recommendations in a friendly manner. However, this is a challenge when developing a sociable recommen…
Fantastic Questions and Where to Find Them: FairytaleQA -- An Authentic Dataset for Narrative Comprehension
Ying Xu, Dakuo Wang, Mo Yu +15
Question answering (QA) is a fundamental means to facilitate assessment and training of narrative comprehension skills for both machines and young children, yet there is scarcity o…
Proactive defense against LLM Jailbreak
Weiliang Zhao, Jinjun Peng, Daniel Ben-Levi +2
The proliferation of powerful large language models (LLMs) has necessitated robust safety alignment, yet these models remain vulnerable to evolving adversarial attacks, including m…
ErAConD : Error Annotated Conversational Dialog Dataset for Grammatical Error Correction
Xun Yuan, Derek Pham, Sam Davidson +1
Currently available grammatical error correction (GEC) datasets are compiled using well-formed written text, limiting the applicability of these datasets to other domains such as i…
Attribute Alignment: Controlling Text Generation from Pre-trained Language Models
Dian Yu, Zhou Yu, Kenji Sagae
Large language models benefit from training with a large amount of unlabeled text, which gives them increasingly fluent and diverse generation capabilities. However, using these mo…
A Trigamma-free Approach for Computing Information Matrices Related to Trigamma Function
Zhou Yu, Niloufar Dousti Mousavi, Jie Yang
Negative binomial related distributions have been widely used in practice. The calculation of the corresponding Fisher information matrices involves the expectation of trigamma fun…
Effects of Persuasive Dialogues: Testing Bot Identities and Inquiry Strategies
Weiyan Shi, Xuewei Wang, Yoo Jung Oh +3
Intelligent conversational agents, or chatbots, can take on various identities and are increasingly engaging in more human-centered conversations with persuasive goals. However, li…
Importance-Aware Learning for Neural Headline Editing
Qingyang Wu, Lei Li, Hao Zhou +2
Many social media news writers are not professionally trained. Therefore, social media platforms have to hire professional editors to adjust amateur headlines to attract more reade…
AutoRPA: Efficient GUI Automation through LLM-Driven Code Synthesis from Interactions
Minghao Chen, Xinyi Hu, Zhou Yu +1
Large Language Model (LLM) based agents have demonstrated proficiency in multi-step interactions with graphical user interfaces (GUIs). While most research focuses on improving sin…
Bringing Pedagogy into Focus: Evaluating Virtual Teaching Assistants' Question-Answering in Asynchronous Learning Environments
Li Siyan, Zhen Xu, Vethavikashini Chithrra Raghuram +3
Asynchronous learning environments (ALEs) are widely adopted for formal and informal learning, but timely and personalized support is often limited. In this context, Virtual Teachi…
Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations
Yanda Chen, Ruiqi Zhong, Narutatsu Ri +5
Large language models (LLMs) are trained to imitate humans to explain human decisions. However, do LLMs explain themselves? Can they help humans build mental models of how LLMs pro…
Sources of Noise in Dialogue and How to Deal with Them
Derek Chen, Zhou Yu
Training dialogue systems often entails dealing with noisy training examples and unexpected user inputs. Despite their prevalence, there currently lacks an accurate survey of dialo…
User Adaptive Language Learning Chatbots with a Curriculum
Kun Qian, Ryan Shea, Yu Li +2
Along with the development of systems for natural language understanding and generation, dialog systems have been widely adopted for language learning and practicing. Many current…
Building Persona Consistent Dialogue Agents with Offline Reinforcement Learning
Ryan Shea, Zhou Yu
Maintaining a consistent persona is a key quality for any open domain dialogue system. Current state-of-the-art systems do this by training agents with supervised learning or onlin…
Task-Oriented Dialog Systems that Consider Multiple Appropriate Responses under the Same Context
Yichi Zhang, Zhijian Ou, Zhou Yu
Conversations have an intrinsic one-to-many property, which means that multiple responses can be appropriate for the same dialog context. In task-oriented dialogs, this property le…
Filter Pruning for Efficient CNNs via Knowledge-driven Differential Filter Sampler
Shaohui Lin, Wenxuan Huang, Jiao Xie +5
Filter pruning simultaneously accelerates the computation and reduces the memory overhead of CNNs, which can be effectively applied to edge devices and cloud services. In this pape…
Quantifying Intrinsic Uncertainty in Classification via Deep Dirichlet Mixture Networks
Qingyang Wu, He Li, Lexin Li +1
With the widespread success of deep neural networks in science and technology, it is becoming increasingly important to quantify the uncertainty of the predictions produced by deep…
Simulation Guided Molecular Design of Hydrofluoroether Solvent for High Energy Batteries
Zhou Yu, Zhangxing shi, Sambasiva R. Bheemireddy +9
Electrolyte design is critical for enabling next-generation batteries with higher energy densities. Hydrofluoroether (HFE) solvents have drawn a lot of attention as the electrolyte…
A Tailored Pre-Training Model for Task-Oriented Dialog Generation
Jing Gu, Qingyang Wu, Chongruo Wu +2
The recent success of large pre-trained language models such as BERT and GPT-2 has suggested the effectiveness of incorporating language priors in downstream dialog generation task…
Matching Text with Deep Mutual Information Estimation
Xixi Zhou, Chengxi Li, Jiajun Bu +4
Text matching is a core natural language processing research problem. How to retain sufficient information on both content and structure information is one important challenge. In…
Factorization-Error-Free Discrete Diffusion Language Model via Speculative Decoding
Xun Fang, Yunchen Li, Hang Yuan +1
Discrete diffusion language models improve generation efficiency through parallel token prediction, but standard prediction methods introduce factorization errors by approxim…
ROSITA: Enhancing Vision-and-Language Semantic Alignments via Cross- and Intra-modal Knowledge Integration
Yuhao Cui, Zhou Yu, Chunqi Wang +4
Vision-and-language pretraining (VLP) aims to learn generic multimodal representations from massive image-text pairs. While various successful attempts have been proposed, learning…
CSD: Content-aware Speculative Decoding for Efficient Image Generation
Mingcheng Wang, Junbo Qiao, Yunchen Li +8
Speculative decoding (SD) has emerged as a key solution to accelerate the inference of autoregressive models. However, in the field of image generation, it faces the challenge of l…