7 papers
Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer Review
Sungduk Yu, Man Luo, Avinash Madasu +2
Peer review is a critical process for ensuring the integrity of published scientific research. Confidence in this process is predicated on the assumption that experts in the releva…
A Semantic Parsing Framework for End-to-End Time Normalization
Xin Su, Sungduk Yu, Phillip Howard +1
Time normalization is the task of converting natural language temporal expressions into machine-readable representations. It underpins many downstream applications in information r…
A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled Environment
Raanan Y. Rohekar, Yaniv Gurwicz, Sungduk Yu +2
Are generative pre-trained transformer (GPT) models, trained only to predict the next token, implicitly learning a world model from which sequences are generated one token at a tim…
Probing Semantic Routing in Large Mixture-of-Expert Models
Matthew Lyle Olson, Neale Ratzlaff, Musashi Hinck +4
In the past year, large (>100B parameter) mixture-of-expert (MoE) models have become increasingly common in the open domain. While their advantages are often framed in terms of eff…
ClimDetect: A Benchmark Dataset for Climate Change Detection and Attribution
Sungduk Yu, Brian L. White, Anahita Bhiwandiwalla +6
Detecting and attributing temperature increases driven by climate change is crucial for understanding global warming and informing adaptation strategies. However, distinguishing hu…
LVLM-Compress-Bench: Benchmarking the Broader Impact of Large Vision-Language Model Compression
Souvik Kundu, Anahita Bhiwandiwalla, Sungduk Yu +6
Despite recent efforts in understanding the compression impact on large language models (LLMs) in terms of their downstream task performance and trustworthiness on relatively simpl…