papers

Publications (7)

cs.DB2024

FCBench: Cross-Domain Benchmarking of Lossless Compression for Floating-Point Data

Xinyu Chen, Jiannan Tian, Ian Beaver +4

While both the database and high-performance computing (HPC) communities utilize lossless compression methods to minimize floating-point data size, a disconnect persists between th…

cs.LG2018

Paying Attention to Attention: Highlighting Influential Samples in Sequential Analysis

Cynthia Freeman, Jonathan Merriman, Abhinav Aggarwal +2

In (Yang et al. 2016), a hierarchical attention network (HAN) is created for document classification. The attention layer can be used to visualize text influential in classifying t…

cs.CL2022

A Semi-Supervised Deep Clustering Pipeline for Mining Intentions From Texts

Xinyu Chen, Ian Beaver

Mining the latent intentions from large volumes of natural language inputs is a key step to help data analysts design and refine Intelligent Virtual Assistants (IVAs) for customer…

cs.LG2021

TimeVAE: A Variational Auto-Encoder for Multivariate Time Series Generation

Abhyuday Desai, Cynthia Freeman, Zuhui Wang +1

Recent work in synthetic data generation in the time-series domain has focused on the use of Generative Adversarial Networks. We propose a novel architecture for synthetically gene…

cs.AI2026

When Your LLM Reaches End-of-Life: A Framework for Confident Model Migration in Production Systems

Emma Casey, David Roberts, David Sim +1

We present a framework for migrating production Large Language Model (LLM) based systems when the underlying model reaches end-of-life or requires replacement. The key contribution…

cs.CL2022

An Adaptive Deep Clustering Pipeline to Inform Text Labeling at Scale

Xinyu Chen, Ian Beaver

Mining the latent intentions from large volumes of natural language inputs is a key step to help data analysts design and refine Intelligent Virtual Assistants (IVAs) for customer…

cs.CL2017

An Annotated Corpus of Relational Strategies in Customer Service

Ian Beaver, Cynthia Freeman, Abdullah Mueen

We create and release the first publicly available commercial customer service corpus with annotated relational segments. Human-computer data from three live customer service Intel…