6 papers
FastUMAP: Scalable Dimensionality Reduction via Bipartite Landmark Sampling
Hongmin Li
Exploratory analysis of high-dimensional data rarely stops at a single embedding. In practice, analysts rerun dimensionality reduction after changing preprocessing, subsets, or hyp…
The Calibration Turn in AI-Assisted Research: A Conceptual and Methodological Framework for Evidence-Licensed Claims
Hongmin Li
AI-assisted research has entered a stage in which the central question is not only whether systems can generate hypotheses, run experiments, or produce manuscripts, but whether the…
Separating Shortcut Transition from Cross-Family OOD Failure in a Minimal Model
Hongmin Li
Shortcut features are often invoked to explain out-of-distribution (OOD) failure, but training correlation, learned shortcut use, and test-time failure need not coincide. We study…
Targeted Tests for LLM Reasoning: An Audit-Constrained Protocol
Hongmin Li
Fixed reasoning benchmarks evaluate canonical prompts, but semantically valid changes in presentation can still change model behavior. Studies of prompt variation can reveal such f…
A Controlled Counterexample to Strong Proxy-Based Explanations of OOD Performance: in a Fixed Pretraining-and-Probing Setup
Hongmin Li
Task-agnostic structure proxies are often used to interpret why one pretraining corpus transfers better than another, but such explanations require the proxy to track the structure…
LLM-guided Semi-Supervised Approaches for Social Media Crisis Data Classification
Jacob Ativo, Bharaneeshwar Balasubramaniyam, Anh Tran +4
Semi-supervised learning approaches have been investigated as a means to enhance the analysis of social media data in disaster management contexts. In this work, we present the fir…