Publications (27)
NOMAD: Non-locking, stOchastic Multi-machine algorithm for Asynchronous and Decentralized matrix completion
Hyokun Yun, Hsiang-Fu Yu, Cho-Jui Hsieh +2
We develop an efficient parallel distributed algorithm for matrix completion, named NOMAD (Non-locking, stOchastic Multi-machine algorithm for Asynchronous and Decentralized matrix…
High-dimensional Time Series Prediction with Missing Values
Hsiang-Fu Yu, Nikhil Rao, Inderjit S. Dhillon
High-dimensional time series prediction is needed in applications as diverse as demand forecasting and climatology. Often, such applications require methods that are both highly sc…
LearningWord Embeddings for Low-resource Languages by PU Learning
Chao Jiang, Hsiang-Fu Yu, Cho-Jui Hsieh +1
Word embedding is a key component in many downstream applications in processing natural languages. Existing approaches often assume the existence of a large collection of text for…
GroupDPO: Memory efficient Group-wise Direct Preference Optimization
Jixuan Leng, Si Si, Hsiang-Fu Yu +2
Preference optimization is widely used to align Large Language Models (LLMs) with preference feedback. However, most existing methods train on a single positive-negative pair per p…
Retrieval-augmented Encoders for Extreme Multi-label Text Classification
Yau-Shian Wang, Wei-Cheng Chang, Jyun-Yu Jiang +3
Extreme multi-label classification (XMC) seeks to find relevant labels from an extremely large label collection for a given text input. To tackle such a vast label space, current s…
Graph DNA: Deep Neighborhood Aware Graph Encoding for Collaborative Filtering
Liwei Wu, Hsiang-Fu Yu, Nikhil Rao +2
In this paper, we consider recommender systems with side information in the form of graphs. Existing collaborative filtering algorithms mainly utilize only immediate neighborhood i…
MinPrompt: Graph-based Minimal Prompt Data Augmentation for Few-shot Question Answering
Xiusi Chen, Jyun-Yu Jiang, Wei-Cheng Chang +3
Recent advances in few-shot question answering (QA) mostly rely on the power of pre-trained large language models (LLMs) and fine-tuning in specific settings. Although the pre-trai…
Extreme Zero-Shot Learning for Extreme Text Classification
Yuanhao Xiong, Wei-Cheng Chang, Cho-Jui Hsieh +2
The eXtreme Multi-label text Classification (XMC) problem concerns finding most relevant labels for an input text instance from a large label set. However, the XMC setup faces two…
Uncertainty in Extreme Multi-label Classification
Jyun-Yu Jiang, Wei-Cheng Chang, Jiong Zhong +2
Uncertainty quantification is one of the most crucial tasks to obtain trustworthy and reliable machine learning models for decision making. However, most research in this domain ha…
Node Feature Extraction by Self-Supervised Multi-scale Neighborhood Prediction
Eli Chien, Wei-Cheng Chang, Cho-Jui Hsieh +4
Learning on graphs has attracted significant attention in the learning community due to numerous real-world applications. In particular, graph neural networks (GNNs), which take nu…
PECOS: Prediction for Enormous and Correlated Output Spaces
Hsiang-Fu Yu, Kai Zhong, Jiong Zhang +2
Many large-scale applications amount to finding relevant results from an enormous output space of potential candidates. For example, finding the best matching product from a large…
Extreme Multi-label Learning for Semantic Matching in Product Search
Wei-Cheng Chang, Daniel Jiang, Hsiang-Fu Yu +9
We consider the problem of semantic matching in product search: given a customer query, retrieve all semantically related products from a huge catalog of size 100 million, or more.…
ELIAS: End-to-End Learning to Index and Search in Large Output Spaces
Nilesh Gupta, Patrick H. Chen, Hsiang-Fu Yu +2
Extreme multi-label classification (XMC) is a popular framework for solving many real-world problems that require accurate prediction from a very large number of potential output c…
PEFA: Parameter-Free Adapters for Large-scale Embedding-based Retrieval Models
Wei-Cheng Chang, Jyun-Yu Jiang, Jiong Zhang +5
Embedding-based Retrieval Models (ERMs) have emerged as a promising framework for large-scale text retrieval problems due to powerful large language models. Nevertheless, fine-tuni…
Label Disentanglement in Partition-based Extreme Multilabel Classification
Xuanqing Liu, Wei-Cheng Chang, Hsiang-Fu Yu +2
Partition-based methods are increasingly-used in extreme multi-label classification (XMC) problems due to their scalability to large output spaces (e.g., millions or more). However…
Taming Pretrained Transformers for Extreme Multi-label Text Classification
Wei-Cheng Chang, Hsiang-Fu Yu, Kai Zhong +2
We consider the extreme multi-label text classification (XMC) problem: given an input text, return the most relevant labels from a large label collection. For example, the input te…
Large-scale Multi-label Learning with Missing Labels
Hsiang-Fu Yu, Prateek Jain, Purushottam Kar +1
The multi-label classification problem has generated significant interest in recent years. However, existing approaches do not adequately address two key challenges: (a) the abilit…
Think Globally, Act Locally: A Deep Neural Network Approach to High-Dimensional Time Series Forecasting
Rajat Sen, Hsiang-Fu Yu, Inderjit Dhillon
Forecasting high-dimensional time series plays a crucial role in many applications such as demand forecasting and financial predictions. Modern datasets can have millions of correl…
PASSCoDe: Parallel ASynchronous Stochastic dual Co-ordinate Descent
Cho-Jui Hsieh, Hsiang-Fu Yu, Inderjit S. Dhillon
Stochastic Dual Coordinate Descent (SDCD) has become one of the most efficient ways to solve the family of -regularized empirical risk minimization problems, including line…
A Scalable Asynchronous Distributed Algorithm for Topic Modeling
Hsiang-Fu Yu, Cho-Jui Hsieh, Hyokun Yun +2
Learning meaningful topic models with massive document collections which contain millions of documents and billions of tokens is challenging because of two reasons: First, one need…
Representer Point Selection for Explaining Regularized High-dimensional Models
Che-Ping Tsai, Jiong Zhang, Eli Chien +3
We introduce a novel class of sample-based explanations we term high-dimensional representers, that can be used to explain the predictions of a regularized high-dimensional model i…
A Greedy Approach for Budgeted Maximum Inner Product Search
Hsiang-Fu Yu, Cho-Jui Hsieh, Qi Lei +1
Maximum Inner Product Search (MIPS) is an important task in many machine learning applications such as the prediction phase of a low-rank matrix factorization model for a recommend…
Learning to Encode Position for Transformer with Continuous Dynamical Model
Xuanqing Liu, Hsiang-Fu Yu, Inderjit Dhillon +1
We introduce a new way of learning to encode position information for non-recurrent models, such as Transformer models. Unlike RNN and LSTM, which contain inductive bias by loading…
PINA: Leveraging Side Information in eXtreme Multi-label Classification via Predicted Instance Neighborhood Aggregation
Eli Chien, Jiong Zhang, Cho-Jui Hsieh +4
The eXtreme Multi-label Classification~(XMC) problem seeks to find relevant labels from an exceptionally large label space. Most of the existing XMC learners focus on the extractio…
Extreme Stochastic Variational Inference: Distributed and Asynchronous
Jiong Zhang, Parameswaran Raman, Shihao Ji +3
Stochastic variational inference (SVI), the state-of-the-art algorithm for scaling variational inference to large-datasets, is inherently serial. Moreover, it requires the paramete…
Enterprise-Scale Search: Accelerating Inference for Sparse Extreme Multi-Label Ranking Trees
Philip A. Etter, Kai Zhong, Hsiang-Fu Yu +2
Tree-based models underpin many modern semantic search engines and recommender systems due to their sub-linear inference times. In industrial applications, these models operate at…
Multiresolution Transformer Networks: Recurrence is Not Essential for Modeling Hierarchical Structure
Vikas K. Garg, Inderjit S. Dhillon, Hsiang-Fu Yu
The architecture of Transformer is based entirely on self-attention, and has been shown to outperform models that employ recurrence on sequence transduction tasks such as machine t…