collaborators

15 papers

cs.CL2026

Massively Multilingual Joint Segmentation and Glossing

Michael Ginn, Lindia Tjuatja, Enora Rice +5

Automated interlinear gloss prediction with neural networks is a promising approach to accelerate language documentation efforts. However, while state-of-the-art models like GlossL…

cs.CL2026

What do Language Models Learn and When? The Implicit Curriculum Hypothesis

Emmy Liu, Kaiser Sun, Millicent Li +4

Large language models (LLMs) can perform remarkably complex tasks, yet the fine-grained details of how these capabilities emerge during pretraining remain poorly understood. Scalin…

cs.CL2026

IDIOLEX: Unified and Continuous Representations for Idiolectal and Stylistic Variation

Anjali Kantharuban, Aarohi Srivastava, Fahim Faisal +5

Existing sentence representations primarily encode what a sentence says, rather than how it is expressed, even though the latter is important for many applications. In contrast, we…

cs.CV2026

ZINA: Multimodal Fine-grained Hallucination Detection and Editing

Yuiga Wada, Kazuki Matsuda, Komei Sugiura +1

Multimodal Large Language Models (MLLMs) often generate hallucinations, where the output deviates from the visual content. Given that these hallucinations can take diverse forms, d…

cs.LG2026

Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes

Steven Kolawole, Lucio Dery, Jean-François Kagy +3

Structured pruning is a promising approach to create smaller, faster large language models. However, existing methods typically rely on computing the gradient via backward passes,…

cs.CL2025

ClusterFusion: Hybrid Clustering with Embedding Guidance and LLM Adaptation

Yiming Xu, Yuan Yuan, Vijay Viswanathan +1

Text clustering is a fundamental task in natural language processing, yet traditional clustering algorithms with pre-trained embeddings often struggle in domain-specific contexts w…