collaborators

6 papers

cs.LG2026

FastUMAP: Scalable Dimensionality Reduction via Bipartite Landmark Sampling

Hongmin Li

Exploratory analysis of high-dimensional data rarely stops at a single embedding. In practice, analysts rerun dimensionality reduction after changing preprocessing, subsets, or hyp…

cs.LG2026

The Calibration Turn in AI-Assisted Research: A Conceptual and Methodological Framework for Evidence-Licensed Claims

Hongmin Li

AI-assisted research has entered a stage in which the central question is not only whether systems can generate hypotheses, run experiments, or produce manuscripts, but whether the…

cs.LG2026

Separating Shortcut Transition from Cross-Family OOD Failure in a Minimal Model

Hongmin Li

Shortcut features are often invoked to explain out-of-distribution (OOD) failure, but training correlation, learned shortcut use, and test-time failure need not coincide. We study…

cs.LG2026

Targeted Tests for LLM Reasoning: An Audit-Constrained Protocol

Hongmin Li

Fixed reasoning benchmarks evaluate canonical prompts, but semantically valid changes in presentation can still change model behavior. Studies of prompt variation can reveal such f…

cs.LG2026

A Controlled Counterexample to Strong Proxy-Based Explanations of OOD Performance: in a Fixed Pretraining-and-Probing Setup

Hongmin Li

Task-agnostic structure proxies are often used to interpret why one pretraining corpus transfers better than another, but such explanations require the proxy to track the structure…

cs.AI2026

LLM-guided Semi-Supervised Approaches for Social Media Crisis Data Classification

Jacob Ativo, Bharaneeshwar Balasubramaniyam, Anh Tran +4

Semi-supervised learning approaches have been investigated as a means to enhance the analysis of social media data in disaster management contexts. In this work, we present the fir…