collaborators

13 papers

cs.CY2026

From Forensics to Ecosystems: Rethinking Watermarks for Generative AI Oversight

Daniel Susser, John Thickstun, Gili Vidan

The arrival of generative AI as a cheap, widely accessible commercial service, and the tidal wave of AI-generated synthetic content it has unleashed, have provoked deep epistemic a…

cs.CL2026

Fidelity-Diversity Metrics for Text

Amanda Wang, Tudor Manole, Florentina Bunea +1

As language modeling technology matures, there is an increasing research focus on the composition and curation of datasets used to train these models. For instance, practitioners c…

cs.LG2026

Benchmark Datasets for Lead-Lag Forecasting on Social Platforms

Kimia Kazemian, Zhenzhen Liu, Yangfanyu Yang +9

Social and collaborative platforms emit multivariate time-series traces in which early interactions -- such as views, likes, or downloads -- are followed, sometimes months or years…

cs.CL2026

Esoteric Language Models: A Family of Any-Order Diffusion LLMs

Subham Sekhar Sahoo, Zhihan Yang, Yash Akhauri +7

Diffusion-based language models offer a compelling alternative to autoregressive (AR) models by enabling parallel and controllable generation. Within this family, Masked Diffusion…

cs.SD2026

Assessing Factual Music Comprehension in Large Audio Language Models

Daniel Chenyu Lin, Michael Freeman, John Thickstun

Large audio language models (LALMs) leverage multimodal representations to generate open-ended answers to natural language queries about audio. In this paper, we (1) provide empiri…

cs.SD2026

Music Transcription with (Almost) No Supervision

Saebyeol Shin, Chao Wan, Zhenzhen Liu +4

Competitive music transcription models require large amounts of paired audio-score data, which is scarce due to collection costs, alignment difficulty, and copyright restrictions.…