activity
20242026
collaborators

7 papers

cs.LG2026

Towards Anytime-Valid Statistical Watermarking

Baihe Huang, Eric Xu, Kannan Ramchandran +2

The proliferation of Large Language Models (LLMs) necessitates efficient mechanisms to distinguish machine-generated content from human text. While statistical watermarking has eme…

cs.CL2026

Generalization or Hallucination? Understanding Out-of-Context Reasoning in Transformers

Yixiao Huang, Hanlin Zhu, Tianyu Guo +5

Large language models (LLMs) can acquire new knowledge through fine-tuning, but this process exhibits a puzzling duality: models can generalize remarkably from new facts, yet are a…

cs.LG2025

Sample Complexity and Representation Ability of Test-time Scaling Paradigms

Baihe Huang, Shanda Li, Tianhao Wu +5

Test-time scaling paradigms have significantly advanced the capabilities of large language models (LLMs) on complex tasks. Despite their empirical success, theoretical understandin…

cs.CL2025

IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis

Hanyu Li, Haoyu Liu, Tingyu Zhu +4

Large Language Models (LLMs) show promise as data analysis agents, but existing benchmarks overlook the iterative nature of the field, where experts' decisions evolve with deeper i…

cs.CL2025

How Do LLMs Perform Two-Hop Reasoning in Context?

Tianyu Guo, Hanlin Zhu, Ruiqi Zhang +4

``Socrates is human. All humans are mortal. Therefore, Socrates is mortal.'' This form of argument illustrates a typical pattern of two-hop reasoning. Formally, two-hop reasoning r…

stat.ML2025

An Overview of Large Language Models for Statisticians

Wenlong Ji, Weizhe Yuan, Emily Getzen +7

Large Language Models (LLMs) have emerged as transformative tools in artificial intelligence (AI), exhibiting remarkable capabilities across diverse tasks such as text generation,…