papers

Publications (17)

cs.CL2026

Interleaved Reasoning for Large Language Models via Reinforcement Learning

Roy Xie, David Qiu, Deepak Gopinath +5

Long chain-of-thought (CoT) significantly enhances the reasoning capabilities of large language models (LLMs). However, extensive reasoning traces lead to inefficiencies and increa…

cs.CV2022

Looking Outside the Window: Wide-Context Transformer for the Semantic Segmentation of High-Resolution Remote Sensing Images

Lei Ding, Dong Lin, Shaofu Lin +5

Long-range contextual information is crucial for the semantic segmentation of High-Resolution (HR) Remote Sensing Images (RSIs). However, image cropping operations, commonly used f…

cs.IR2022

On the Factory Floor: ML Engineering for Industrial-Scale Ads Recommendation Models

Rohan Anil, Sandra Gadanho, Da Huang +9

For industrial-scale advertising systems, prediction of ad click-through rate (CTR) is a central problem. Ad clicks constitute a significant class of user engagements and are often…

cs.IR2020

Learning Multi-granular Quantized Embeddings for Large-Vocab Categorical Features in Recommender Systems

Wang-Cheng Kang, Derek Zhiyuan Cheng, Ting Chen +4

Recommender system models often represent various sparse features like users, items, and categorical features via embeddings. A standard approach is to map each unique feature valu…

cs.LG2022

Dropout Prediction Uncertainty Estimation Using Neuron Activation Strength

Haichao Yu, Zhe Chen, Dong Lin +2

Dropout has been commonly used to quantify prediction uncertainty, i.e, the variations of model predictions on a given input example. However, using dropout in practice can be expe…

cs.LG2021

Understanding and Improving Knowledge Distillation

Jiaxi Tang, Rakesh Shivanna, Zhe Zhao +4

Knowledge Distillation (KD) is a model-agnostic technique to improve model quality while having a fixed capacity budget. It is a commonly used technique for model compression, wher…

cs.LG2020

Beyond Point Estimate: Inferring Ensemble Prediction Variation from Neuron Activation Strength in Recommender Systems

Zhe Chen, Yuyan Wang, Dong Lin +4

Despite deep neural network (DNN)'s impressive prediction performance in various domains, it is well known now that a set of DNN models trained with the same model specification an…

cs.LG2022

Real World Large Scale Recommendation Systems Reproducibility and Smooth Activations

Gil I. Shamir, Dong Lin

Real world recommendation systems influence a constantly growing set of domains. With deep networks, that now drive such systems, recommendations have been more relevant to the use…

cs.IR2020

DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank Systems

Ruoxi Wang, Rakesh Shivanna, Derek Z. Cheng +4

Learning effective feature crosses is the key behind building recommender systems. However, the sparse and large feature space requires exhaustive search to identify effective cros…

cs.IR2023

Learning to Rank when Grades Matter

Le Yan, Zhen Qin, Gil Shamir +3

Graded labels are ubiquitous in real-world learning-to-rank applications, especially in human rated relevance data. Traditional learning-to-rank techniques aim to optimize the rank…

physics.app-ph2021

Analytical Model for Gaussian Disorder Traps in Organic Thin-Film Transistor

Qiusong Chen, Juan E. Sanchez, Dong Lin +2

Structural defects and chemical impurities exist in organic semiconductors acting as trap centers for the excited states. This work presents a novel analytical model to calculate t…

cs.LG2020

Small Towers Make Big Differences

Yuyan Wang, Zhe Zhao, Bo Dai +4

Multi-task learning aims at solving multiple machine learning tasks at the same time. A good solution to a multi-task learning problem should be generalizable in addition to being…

cs.LG2020

Smooth activations and reproducibility in deep networks

Gil I. Shamir, Dong Lin, Lorenzo Coviello

Deep networks are gradually penetrating almost every domain in our lives due to their amazing success. However, with substantive performance accuracy improvements comes the price o…

cs.CV2020

MP-ResNet: Multi-path Residual Network for the Semantic segmentation of High-Resolution PolSAR Images

Lei Ding, Kai Zheng, Dong Lin +4

There are limited studies on the semantic segmentation of high-resolution Polarimetric Synthetic Aperture Radar (PolSAR) images due to the scarcity of training data and the inferen…

cs.LG2026

PAI: Preserving Amplitude Information in Representation-Based Time-Series Anomaly Detection

Kang Zhang, Wei Jian Lau, Shoushou Ren +3

Representation-based time-series anomaly detection algorithms significantly outperform other methods on diverse anomaly detection tasks. However, we notice that they suffer from a…

cs.LG2026

Over-Searching in Search-Augmented Large Language Models

Roy Xie, Deepak Gopinath, David Qiu +4

Search-augmented large language models (LLMs) excel at knowledge-intensive tasks by integrating external retrieval. However, they often over-search -- unnecessarily invoking search…

physics.space-ph2025

Field Aligned Currents and Auroral Precipitation During the Terrestrial Alfven Wing State

Brandon Burkholder, Li-Jen Chen, Kareem Sorathia +3

When sub-Alfvénic (Alfvén Mach number MA < 1) plasmas impact Earth, Alfvén wings (AWs) develop. A Multiscale Atmosphere Geospace Environment (MAGE) simulation of the April 2023…