NewEvery arXiv paper, its researchers & institutions — mapped.
papers

Publications (318)

cs.LG2023

OpenDataVal: a Unified Benchmark for Data Valuation

Kevin Fu Jiang, Weixin Liang, James Zou +1

cs.AI2026

Cerebra: A Multidisciplinary AI Board for Multimodal Dementia Characterization and Risk Assessment

Sheng Liu, Long Chen, Zeyun Zhao +16

cs.LG2022

MetaShift: A Dataset of Datasets for Evaluating Contextual Distribution Shifts and Training Conflicts

Weixin Liang, James Zou

cs.LG2023

Data-Driven Subgroup Identification for Linear Regression

Zachary Izzo, Ruishan Liu, James Zou

cs.AI2026

Unlocking LLM Creativity in Science through Analogical Reasoning

Andrew Shen, Shaul Druckmann, James Zou

cs.GT2016

Signal to noise in matching markets

S. Matthew Weinberg, James Zou

cs.CV2025

When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought

Yiyang Zhou, Haoqin Tu, Zijun Wang +11

cs.LG2023

Last-Layer Fairness Fine-tuning is Simple and Effective for Neural Networks

Yuzhen Mao, Zhun Deng, Huaxiu Yao +3

cs.CL2022

Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Weixin Liang, Yuhui Zhang, Yongchan Kwon +2

q-bio.QM2021

CloudPred: Predicting Patient Phenotypes From Single-cell RNA-seq

Bryan He, Matthew Thomson, Meena Subramaniam +3

cs.LG2022

C-Mixup: Improving Generalization in Regression

Huaxiu Yao, Yiping Wang, Linjun Zhang +2

cs.LG2025

Freeze then Train: Towards Provable Representation Learning under Spurious Correlations and Feature Noise

Haotian Ye, James Zou, Linjun Zhang

cs.CL2025

Simple linear attention language models balance the recall-throughput tradeoff

Simran Arora, Sabri Eyuboglu, Michael Zhang +6

cs.AI2026

AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration

Jiaqi Liu, Shi Qiu, Mairui Li +33

cs.AI2025

Optimizing Model Selection for Compound AI Systems

Lingjiao Chen, Jared Quincy Davis, Boris Hanin +4

cs.AI2025

Data Diversification Methods In Alignment Enhance Math Performance In LLMs

Berkan Dokmeci, Qingyang Wu, Ben Athiwaratkun +3

cs.LG2021

Approximate Data Deletion from Machine Learning Models

Zachary Izzo, Mary Anne Smart, Kamalika Chaudhuri +1

cs.CV2022

Development and Clinical Evaluation of an AI Support Tool for Improving Telemedicine Photo Quality

Kailas Vodrahalli, Justin Ko, Albert S. Chiou +7

cs.CL2026

Latent Collaboration in Multi-Agent Systems

Jiaru Zou, Ruizhong Qiu, Gaotang Li +10

cs.LG2023

Discover and Cure: Concept-aware Mitigation of Spurious Correlation

Shirley Wu, Mert Yuksekgonul, Linjun Zhang +1

eess.IV2022

Disparities in Dermatology AI Performance on a Diverse, Curated Clinical Image Set

Roxana Daneshjou, Kailas Vodrahalli, Roberto A Novoa +16

cs.AI2026

Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing Agents

Muyu He, Anand Kumar, Tsach Mackey +3

cs.AI2025

Solving Inequality Proofs with Large Language Models

Pan Lu, Jiayi Sheng, Luna Lyu +4

cs.CL2025

Cartridges: Lightweight and general-purpose long context representations via self-study

Sabri Eyuboglu, Ryan Ehrlich, Simran Arora +8

cs.LG2024

Navigating Dataset Documentations in AI: A Large-Scale Analysis of Dataset Cards on Hugging Face

Xinyu Yang, Weixin Liang, James Zou

cs.LG2024

Are More LLM Calls All You Need? Towards Scaling Laws of Compound Inference Systems

Lingjiao Chen, Jared Quincy Davis, Boris Hanin +4

q-bio.QM2024

Is your data alignable? Principled and interpretable alignability testing and integration of single-cell data

Rong Ma, Eric D. Sun, David Donoho +1

cs.LG2026

Can Small Training Runs Reliably Guide Data Curation? Rethinking Proxy-Model Practice

Jiachen T. Wang, Tong Wu, Kaifeng Lyu +4

stat.ML2019

Knockoffs for the mass: new feature importance statistics with false discovery guarantees

Jaime Roquero Gimenez, Amirata Ghorbani, James Zou

cs.CL2026

Real-Time Voice AI Hears but Does Not Listen

Martijn Bartelds, Federico Bianchi, James Zou

cs.LG2026

Generating readily synthesizable small molecule fluorophore scaffolds with reinforcement learning

Ruhi Sayana, Kate Callon, Jennifer Xu +6

cs.LG2025

ReasonIF: Large Reasoning Models Fail to Follow Instructions During Reasoning

Yongchan Kwon, Shang Zhu, Federico Bianchi +2

cs.CY2023

What Should Data Science Education Do with Large Language Models?

Xinming Tu, James Zou, Weijie J. Su +1

cs.CL2022

GSCLIP : A Framework for Explaining Distribution Shifts in Natural Language

Zhiying Zhu, Weixin Liang, James Zou

cs.CL2025

ClashEval: Quantifying the tug-of-war between an LLM's internal prior and external evidence

Kevin Wu, Eric Wu, James Zou

physics.med-ph2025

Automated radiotherapy treatment planning guided by GPT-4Vision

Sheng Liu, Oscar Pastor-Serrano, Yizheng Chen +10

cs.AI2025

CollabLLM: From Passive Responders to Active Collaborators

Shirley Wu, Michel Galley, Baolin Peng +7

cs.LG2022

Improving Out-of-Distribution Robustness via Selective Augmentation

Huaxiu Yao, Yu Wang, Sai Li +4

cs.LG2022

Domino: Discovering Systematic Errors with Cross-Modal Embeddings

Sabri Eyuboglu, Maya Varma, Khaled Saab +5

stat.ML2018

Stochastic EM for Shuffled Linear Regression

Abubakar Abid, James Zou

cs.CV2023

When and why vision-language models behave like bags-of-words, and what to do about it?

Mert Yuksekgonul, Federico Bianchi, Pratyusha Kalluri +2

cs.LG2025

ExGra-Med: Extended Context Graph Alignment for Medical Vision-Language Models

Duy M. H. Nguyen, Nghiem T. Diep, Trung Q. Nguyen +10

cs.CV2025

How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?

Rahul Thapa, Andrew Li, Qingyang Wu +8

cs.AI2025

Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025

Nitya Thakkar, Mert Yuksekgonul, Jake Silberg +6

cs.AI2026

Graph-of-Agents: A Graph-based Framework for Multi-Agent LLM Collaboration

Sukwon Yun, Jie Peng, Pingzhi Li +5

cs.CL2019

Analyzing Polarization in Social Media: Method and Application to Tweets on 21 Mass Shootings

Dorottya Demszky, Nikhil Garg, Rob Voigt +4

cs.LG2019

Contrastive Variational Autoencoder Enhances Salient Features

Abubakar Abid, James Zou

cs.LG2026

OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning

Pan Lu, Bowen Chen, Sheng Liu +3

cs.CL2025

Efficient and Asymptotically Unbiased Constrained Decoding for Large Language Models

Haotian Ye, Himanshu Jain, Chong You +4

cs.CL2017

Beyond Bilingual: Multi-sense Word Embeddings using Multilingual Context

Shyam Upadhyay, Kai-Wei Chang, Matt Taddy +2

cs.AI2025

Exploring the use of AI authors and reviewers at Agents4Science

Federico Bianchi, Owen Queen, Nitya Thakkar +2

cs.RO2026

Finetuning Vision-Language-Action Models Requires Fewer Layers Than You Think

Gia-Binh Nguyen, Trong-Bao Ho, Thien-Loc Ha +18

cs.AI2026

A Versatile AI Agent for Rare Disease Diagnosis and Risk Gene Prioritization

Tianyu Liu, Wangjie Zheng, Rui Yang +12

cs.AI2026

Combee: Scaling Prompt Learning for Self-Improving Language Model Agents

Hanchen Li, Runyuan He, Qizheng Zhang +11

cs.LG2021

Competing AI: How does competition feedback affect machine learning?

Antonio Ginart, Eva Zhang, Yongchan Kwon +1

cs.AI2025

TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular Reasoning

Jiaru Zou, Soumya Roy, Vinay Kumar Verma +6

cs.LG2021

Explaining medical AI performance disparities across sites with confounder Shapley value analysis

Eric Wu, Kevin Wu, James Zou

cs.AI2026

What LLMs Think When You Don't Tell Them What to Think About?

Yongchan Kwon, James Zou

cs.AI2026

Cost-of-Pass: An Economic Framework for Evaluating Language Models

Mehmet Hamza Erol, Batu El, Mirac Suzgun +2

cs.LG2026

Learning to Discover at Test Time

Mert Yuksekgonul, Daniel Koceja, Xinhao Li +8

cs.GT2024

A Survey on Data Markets

Jiayao Zhang, Yuran Bi, Mengye Cheng +15

cs.AI2026

Advancing AI Research Assistants with Expert-Involved Learning

Tianyu Liu, Simeng Han, Hanchen Wang +27

cs.LG2023

Diagnosing and Rectifying Vision Models using Language

Yuhui Zhang, Jeff Z. HaoChen, Shih-Cheng Huang +3

cs.CL2023

GPT detectors are biased against non-native English writers

Weixin Liang, Mert Yuksekgonul, Yining Mao +2

cs.AI2026

The Agentic Garden of Forking Paths

Jiacheng Miao, Jonathan K Pritchard, James Zou

stat.ML2019

Data Shapley: Equitable Valuation of Data for Machine Learning

Amirata Ghorbani, James Zou

cs.CV2025

4KAgent: Agentic Any Image to 4K Super-Resolution

Yushen Zuo, Qi Zheng, Mingyang Wu +10

cs.LG2024

GraphMETRO: Mitigating Complex Graph Distribution Shifts via Mixture of Aligned Experts

Shirley Wu, Kaidi Cao, Bruno Ribeiro +2

stat.ML2021

Efficient computation and analysis of distributional Shapley values

Yongchan Kwon, Manuel A. Rivas, James Zou

cs.AI2026

AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents

Kunlun Zhu, Xuyan Ye, Zhiguang Han +9

cs.IR2024

STaRK: Benchmarking LLM Retrieval on Textual and Relational Knowledge Bases

Shirley Wu, Shiyu Zhao, Michihiro Yasunaga +7

cs.CL2026

MedArena: Comparing LLMs for Medicine-in-the-Wild Clinician Preferences

Eric Wu, Kevin Wu, Jason Hom +11

cs.CV2025

SMIR: Efficient Synthetic Data Pipeline To Improve Multi-Image Reasoning

Andrew Li, Rahul Thapa, Rahul Chalamala +3

cs.LG2022

Data Budgeting for Machine Learning

Xinyi Zhao, Weixin Liang, James Zou

stat.ML2018

Interpretation of Neural Networks is Fragile

Amirata Ghorbani, Abubakar Abid, James Zou

cs.CL2024

How well do LLMs cite relevant medical references? An evaluation framework and analyses

Kevin Wu, Eric Wu, Ally Cassasola +7

cs.LG2023

DataPerf: Benchmarks for Data-Centric AI Development

Mark Mazumder, Colby Banbury, Xiaozhe Yao +42

cs.LG2021

Clustering Plotted Data by Image Segmentation

Tarek Naous, Srinjay Sarkar, Abubakar Abid +1

cs.LG2025

MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models

Peng Xia, Kangyu Zhu, Haoran Li +6

cs.AI2026

Recursive Multi-Agent Systems

Jiaru Zou, Rui Pan, Ruizhong Qiu +8

The paper proposes RecursiveMAS, a framework that treats a multi-agent system as a recursive latent‑space computation, enabling agents to iteratively refine each other's thoughts a…

#multi-agent systems#recursive computation#language models#collaborative AI
cs.AI2024

Can AI Be as Creative as Humans?

Haonan Wang, James Zou, Michael Mozer +8

cs.LG2018

Minimizing Close-k Aggregate Loss Improves Classification

Bryan He, James Zou

cs.LG2025

A Snapshot of Influence: A Local Data Attribution Framework for Online Reinforcement Learning

Yuzheng Hu, Fan Wu, Haotian Ye +5

cs.LG2023

The Power of Contrast for Feature Learning: A Theoretical Analysis

Wenlong Ji, Zhun Deng, Ryumei Nakada +2

stat.ME2017

NeuralFDR: Learning Discovery Thresholds from Hypothesis Features

Fei Xia, Martin J. Zhang, James Zou +1

cs.LG2025

A Theoretical Framework for Prompt Engineering: Approximating Smooth Functions with Transformer Prompts

Ryumei Nakada, Wenlong Ji, Tianxi Cai +2

cs.SE2022

HAPI: A Large-scale Longitudinal Dataset of Commercial ML API Predictions

Lingjiao Chen, Zhihua Jin, Sabri Eyuboglu +3

cs.LG2026

On the Relationship Between Activation Outliers and Feature Death in Sparse Autoencoders

Elana Simon, Etowah Adams, James Zou

cs.CV2020

TrueImage: A Machine Learning Algorithm to Improve the Quality of Telehealth Photos

Kailas Vodrahalli, Roxana Daneshjou, Roberto A Novoa +3

cs.CL2025

Generative Evaluation of Complex Reasoning in Large Language Models

Haowei Lin, Xiangyu Wang, Ruilin Yan +7

cs.LG2021

How Does Mixup Help With Robustness and Generalization?

Linjun Zhang, Zhun Deng, Kenji Kawaguchi +2

cs.CL2025

Disentangling Reasoning and Knowledge in Medical Large Language Models

Rahul Thapa, Qingyang Wu, Kevin Wu +11

cs.CL2021

Persistent Anti-Muslim Bias in Large Language Models

Abubakar Abid, Maheen Farooqi, James Zou

cs.LG2025

Learning a Canonical Basis of Human Preferences from Binary Ratings

Kailas Vodrahalli, Wei Wei, James Zou

cs.CL2023

New Evaluation Metrics Capture Quality Degradation due to LLM Watermarking

Karanpartap Singh, James Zou

stat.ML2021

Beyond Importance Scores: Interpreting Tabular ML by Visualizing Feature Semantics

Amirata Ghorbani, Dina Berenbaum, Maor Ivgi +2

cs.LG2024

Capturing the Temporal Dependence of Training Data Influence

Jiachen T. Wang, Dawn Song, James Zou +2

stat.ML2022

A Unified f-divergence Framework Generalizing VAE and GAN

Jaime Roquero Gimenez, James Zou

cs.CR2025

AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack Integration

Andy Zhou, Kevin Wu, Francesco Pinto +7

cs.LG2026

Reliable and Responsible Foundation Models: A Comprehensive Survey

Xinyu Yang, Junlin Han, Rishi Bommasani +49