activity
20242026
collaborators

9 papers

cs.AI2026

Test-Time Coverage: Test-Conditioned Data Curation for Deployment-Aware Learning

Nadine Chang, Maying Shen, Shizhe Diao +6

Deployed AI systems are often trained from broad candidate data pools, necessitating data curation towards the deployment test distribution. However, standard data curation methods…

cs.AI2026

From Modalities to Propositions: A Language-Centric Framework for Multimodal Intelligence

Nadine Chang, Maying Shen, Shizhe Diao +6

We propose a language representation for multimodal data in which any observation, whether image, video, or text, is expressed as a bag of atomic propositions, simple statements ab…

cs.LG2026

Position: Stop Reactively Patching Your Model Every Time and Start Proactive Test-Driven AI Development

Nadine Chang, Maying Shen, Jialiang Wang +2

Many modern AI systems are designed to operate under diverse, open-ended, use-cases. To help generalize deployed systems, many deployed-system maintenance pipelines use a reactive…

cs.LG2026

Scaling-Aware Data Selection for End-to-End Autonomous Driving Systems

Tolga Dimlioglu, Nadine Chang, Maying Shen +2

Large-scale deep learning models for physical AI applications depend on diverse training data collection efforts. These models and correspondingly, the training data, must address…

cs.GT2026

Routing, Cascades, and User Choice for LLMs

Rafid Mahmood

To mitigate the trade-offs between performance and costs, LLM providers route user tasks to different models based on task difficulty and latency. We study the effect of LLM routin…

cs.LG2025

AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs

Feiyang Kang, Yifan Sun, Bingbing Wen +4

Domain reweighting is an emerging research area aimed at adjusting the relative weights of different data sources to improve the effectiveness and efficiency of LLM pre-training. W…