papers

Publications (36)

cs.CL2019

BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model

Alex Wang, Kyunghyun Cho

We show that BERT (Devlin et al., 2018) is a Markov random field language model. This formulation gives way to a natural procedure to sample sentences from BERT. We generate from B…

cs.CL2025

Command A: An Enterprise-Ready Large Language Model

Team Cohere, :, Aakanksha +227

In this report we describe the development of Command A, a powerful large language model purpose-built to excel at real-world enterprise use cases. Command A is an agent-optimised…

physics.ins-det2025

Silicon pinhole strip defects and their impact on ATLAS Inner Tracker HV current measurement

Anthony Affolder, Kirsten Affolder, Emily Duden +11

In preparation for the High-Luminsoity LHC (HL-LHC), the ATLAS detector will undergo major detector upgrades, including the replacement of the current Inner Detector with the new a…

cs.LG2024

OpenChemIE: An Information Extraction Toolkit For Chemistry Literature

Vincent Fan, Yujie Qian, Alex Wang +3

Information extraction from chemistry literature is vital for constructing up-to-date reaction databases for data-driven chemistry. Complete extraction requires combining informati…

cs.AI2024

TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation

Jonathan Cook, Tim Rocktäschel, Jakob Foerster +2

Given the widespread adoption and usage of Large Language Models (LLMs), it is crucial to have flexible and interpretable evaluations of their instruction-following ability. Prefer…

cs.SD2025

JEN-1: Text-Guided Universal Music Generation with Omnidirectional Diffusion Models

Peike Li, Boyu Chen, Yao Yao +3

Music generation has attracted growing interest with the advancement of deep generative models. However, generating music conditioned on textual descriptions, known as text-to-musi…

cs.CL2022

What Do NLP Researchers Believe? Results of the NLP Community Metasurvey

Julian Michael, Ari Holtzman, Alicia Parrish +8

We present the results of the NLP Community Metasurvey. Run from May to June 2022, the survey elicited opinions on controversial issues, including industry influence in the field,…

cs.CL2020

SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems

Alex Wang, Yada Pruksachatkun, Nikita Nangia +5

In the last year, new models and methods for pretraining and transfer learning have driven striking performance improvements across a range of language understanding tasks. The GLU…

cs.CL2019

Can You Tell Me How to Get Past Sesame Street? Sentence-Level Pretraining Beyond Language Modeling

Alex Wang, Jan Hula, Patrick Xia +13

Natural language understanding has recently seen a surge of progress with the use of sentence encoders like ELMo (Peters et al., 2018a) and BERT (Devlin et al., 2019) which are pre…

cs.CL2019

GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Alex Wang, Amanpreet Singh, Julian Michael +3

For natural language understanding (NLU) technology to be maximally useful, both practically and as a scientific object of study, it must be general: it must be able to process lan…

cs.SD2024

JEN-1 Composer: A Unified Framework for High-Fidelity Multi-Track Music Generation

Yao Yao, Peike Li, Boyu Chen +1

With rapid advances in generative artificial intelligence, the text-to-music synthesis task has emerged as a promising direction for music generation. Nevertheless, achieving preci…

cs.LG2024

Pruning for Protection: Increasing Jailbreak Resistance in Aligned LLMs Without Fine-Tuning

Adib Hasan, Ileana Rugina, Alex Wang

This paper investigates the impact of model compression on the way Large Language Models (LLMs) process prompts, particularly concerning jailbreak resistance. We show that moderate…

cs.RO2026

XRZero-G0: Pushing the Frontier of Dexterous Robotic Manipulation with Interfaces, Quality and Ratios

James Wang, Primo Pu, Zephyr Fung +19

The acquisition of high-quality, action-aligned demonstration data remains a fundamental bottleneck in scaling foundation models for dexterous robot manipulation. Although robot-fr…

cs.CL2019

On Measuring Social Biases in Sentence Encoders

Chandler May, Alex Wang, Shikha Bordia +2

The Word Embedding Association Test shows that GloVe and word2vec word embeddings exhibit human-like implicit biases based on gender, race, and other social constructs (Caliskan et…

cs.CL2019

Probing What Different NLP Tasks Teach Machines about Function Word Comprehension

Najoung Kim, Roma Patel, Adam Poliak +9

We introduce a set of nine challenge tasks that test for the understanding of function words. These tasks are created by structurally mutating sentences from existing datasets to t…

math.OC2005

Output Feedback Pole Assignment for Transfer Functions with Symmetries

Uwe Helmke, Joachim Rosenthal, Alex Wang

This paper studies the problem of pole assignment for symmetric and Hamiltonian transfer functions. A necessary and sufficient condition for pole assignment by complex symmetric ou…

cs.AI2026

BigFinanceBench: A Workflow-Grounded Benchmark for Financial-Research Agents

Alex Wang, Georg Meinhardt, Jacob Katz +4

Financial-research answers are decision-relevant only when another analyst can audit how they were produced: which source was chosen, which period and accounting definition were us…

cs.CL2023

When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Max Marion, Ahmet Üstün, Luiza Pozzobon +3

Large volumes of text data have contributed significantly to the development of large language models (LLMs) in recent years. This data is typically acquired by scraping the intern…

cs.CL2019

What do you learn from context? Probing for sentence structure in contextualized word representations

Ian Tenney, Patrick Xia, Berlin Chen +8

Contextualized representation models such as ELMo (Peters et al., 2018a) and BERT (Devlin et al., 2018) have recently achieved state-of-the-art results on a diverse array of downst…

quant-ph2021

Application of Quantum Machine Learning using the Quantum Variational Classifier Method to High Energy Physics Analysis at the LHC on IBM Quantum Computer Simulator and Hardware with 10 qubits

Sau Lan Wu, Jay Chan, Wen Guan +12

One of the major objectives of the experimental programs at the LHC is the discovery of new physics. This requires the identification of rare signals in immense backgrounds. Using…

cs.LG2017

Clustering Stable Instances of Euclidean k-means

Abhratanu Dutta, Aravindan Vijayaraghavan, Alex Wang

The Euclidean k-means problem is arguably the most widely-studied clustering problem in machine learning. While the k-means objective is NP-hard in the worst-case, practitioners ha…

cs.CV2023

GLARE: A Dataset for Traffic Sign Detection in Sun Glare

Nicholas Gray, Megan Moraes, Jiang Bian +6

Real-time machine learning object detection algorithms are often found within autonomous vehicle technology and depend on quality datasets. It is essential that these algorithms wo…

cs.SD2024

JEN-1 DreamStyler: Customized Musical Concept Learning via Pivotal Parameters Tuning

Boyu Chen, Peike Li, Yao Yao +1

Large models for text-to-music generation have achieved significant progress, facilitating the creation of high-quality and varied musical compositions from provided text prompts.…

cs.CL2021

QuestEval: Summarization Asks for Fact-based Evaluation

Thomas Scialom, Paul-Alexis Dray, Patrick Gallinari +4

Summarization evaluation remains an open research problem: current metrics such as ROUGE are known to be limited and to correlate poorly with human judgments. To alleviate this iss…

cond-mat.mtrl-sci2022

A Low-Cost Robot Science Kit for Education with Symbolic Regression for Hypothesis Discovery and Validation

Logan Saar, Haotong Liang, Alex Wang +4

The next generation of physical science involves robot scientists - autonomous physical science systems capable of experimental design, execution, and analysis in a closed loop. Su…

cs.CL2022

SQuALITY: Building a Long-Document Summarization Dataset the Hard Way

Alex Wang, Richard Yuanzhe Pang, Angelica Chen +2

Summarization datasets are often assembled either by scraping naturally occurring public-domain summaries -- which are nearly always in difficult-to-work-with technical domains --…

stat.ME2011

Intent Inference and Syntactic Tracking with GMTI Measurements

Alex Wang, Vikram Krishnamurthy, Bhashyam Balaji

In conventional target tracking systems, human operators use the estimated target tracks to make higher level inference of the target behaviour/intent. This paper develops syntacti…

cs.CL2020

jiant: A Software Toolkit for Research on General-Purpose Text Understanding Models

Yada Pruksachatkun, Phil Yeres, Haokun Liu +5

We introduce jiant, an open source toolkit for conducting multitask and transfer learning experiments on English NLU tasks. jiant enables modular and configuration-driven experimen…

cs.CL2022

GEMv2: Multilingual NLG Benchmarking in a Single Line of Code

Sebastian Gehrmann, Abhik Bhattacharjee, Abinaya Mahendiran +74

Evaluation in machine learning is usually informed by past choices, for example which datasets or metrics to use. This standardization enables the comparison on equal footing using…

cs.LG2020

A Generalized Framework of Sequence Generation with Application to Undirected Sequence Models

Elman Mansimov, Alex Wang, Sean Welleck +1

Undirected neural sequence models such as BERT (Devlin et al., 2019) have received renewed interest due to their success on discriminative natural language understanding tasks such…

eess.IV2019

Deep Learning for the Digital Pathologic Diagnosis of Cholangiocarcinoma and Hepatocellular Carcinoma: Evaluating the Impact of a Web-based Diagnostic Assistant

Bora Uyumazturk, Amirhossein Kiani, Pranav Rajpurkar +17

While artificial intelligence (AI) algorithms continue to rival human performance on a variety of clinical tasks, the question of how best to incorporate these algorithms into clin…

cs.CR2026

How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment

Guang Yang, Fengchen Liu, Alex Wang +2

State-aligned distortion has been documented in China-origin text-based large language models (LLMs), but whether, and in what form, it arises in multimodal systems has not been sy…

cs.AI2025

The Leaderboard Illusion

Shivalika Singh, Yiyang Nan, Alex Wang +10

Measuring progress is fundamental to the advancement of any scientific field. As benchmarks play an increasingly central role, they also grow more susceptible to distortion. Chatbo…

cs.SD2026

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions

Jiashuo Yu, Yao Yao, Boyu Chen +1

We address the challenge of generating high-fidelity, long-form soundtracks that remain coherent across scene transitions. Existing AI music systems are mainly designed for short,…

cond-mat.mtrl-sci2022

Benchmarking Active Learning Strategies for Materials Optimization and Discovery

Alex Wang, Haotong Liang, Austin McDannald +2

Autonomous physical science is revolutionizing materials science. In these systems, machine learning controls experiment design, execution, and analysis in a closed loop. Active le…

cs.CL2020

Asking and Answering Questions to Evaluate the Factual Consistency of Summaries

Alex Wang, Kyunghyun Cho, Mike Lewis

Practical applications of abstractive summarization models are limited by frequent factual inconsistencies with respect to their input. Existing automatic evaluation metrics for su…