papers

Publications (17)

cs.RO2021

Learning Deep Parameterized Skills from Demonstration for Re-targetable Visuomotor Control

Jonathan Chang, Nishanth Kumar, Sean Hastings +6

Robots need to learn skills that can not only generalize across similar problems but also be directed to a specific goal. Previous methods either train a new skill for every differ…

cs.CL2025

A State-of-the-Art SQL Reasoning Model using RLVR

Alnur Ali, Ashutosh Baheti, Jonathan Chang +13

Developing custom reasoning models via Reinforcement Learning (RL) that can incorporate organization-specific knowledge has great potential to address problems faced by enterprise…

cs.AI2026

Laguna M.1/XS.2 Technical Report

Julien Abadji, Marah Abdin, Connor Adams +93

We present Laguna M.1 and Laguna XS.2, two Mixture-of-Experts foundation models built for long-horizon, agentic coding: M.1 has B total parameters (B activated per tok…

cs.CL2023

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

BigScience Workshop, :, Teven Le Scao +391

Large language models (LLMs) have been shown to be able to perform new tasks based on a few demonstrations or natural language instructions. While these capabilities have led to wi…

cs.LG2022

Multitask Prompted Training Enables Zero-Shot Task Generalization

Victor Sanh, Albert Webson, Colin Raffel +38

Large language models have recently been shown to attain reasonable zero-shot generalization on a diverse set of tasks (Brown et al., 2020). It has been hypothesized that this is a…

cs.LG2026

LLMs Can Learn to Reason Via Off-Policy RL

Daniel Ritter, Owen Oertell, Bradley Guo +3

Reinforcement learning (RL) approaches for Large Language Models (LLMs) frequently use on-policy algorithms, such as PPO or GRPO. However, policy lag from distributed training arch…

cs.LG2025

: Provably Optimal Distributional RL for LLM Post-Training

Jin Peng Zhou, Kaiwen Wang, Jonathan Chang +5

Reinforcement learning (RL) post-training is crucial for LLM alignment and reasoning, but existing policy-based methods, such as PPO and DPO, can fall short of fixing shortcuts inh…

cs.CL2017

Learning Representations of Emotional Speech with Deep Convolutional Generative Adversarial Networks

Jonathan Chang, Stefan Scherer

Automatically assessing emotional valence in human speech has historically been a difficult task for machine learning algorithms. The subtle changes in the voice of the speaker tha…

cs.LG2014

Joint Inference of Multiple Label Types in Large Networks

Deepayan Chakrabarti, Stanislav Funiak, Jonathan Chang +1

We tackle the problem of inferring node labels in a partially labeled graph where each node in the graph has multiple label types and each label type has a large number of possible…

cs.LG2025

Value-Guided Search for Efficient Chain-of-Thought Reasoning

Kaiwen Wang, Jin Peng Zhou, Jonathan Chang +4

In this paper, we propose a simple and efficient method for value model training on long-context reasoning traces. Compared to existing process reward models (PRMs), our method doe…

q-bio.PE2012

TAPIR enables high-throughput estimation and comparison of phylogenetic informativeness using locus-specific substitution models

Brant C. Faircloth, Jonathan Chang, Michael E. Alfaro

Massively parallel DNA sequencing techniques are rapidly changing the dynamics of phylogenetic study design by exponentially increasing the discovery of phylogenetically useful loc…

cs.LG2026

Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards

Yiran Shen, Yu Xia, Jonathan Chang +1

Aligning large language models to human preferences is inherently multidimensional, yet most pipelines collapse heterogeneous signals into a single objective. We seek to answer wha…

q-bio.NC2025

Application of an attention-based CNN-BiLSTM framework for in vivo two-photon calcium imaging of neuronal ensembles: decoding complex bilateral forelimb movements from unilateral M1

Ghazal Mirzaee, Jonathan Chang, Shahrzad Latifi

Decoding behavior, such as movement, from multiscale brain networks remains a central objective in neuroscience. Over the past decades, artificial intelligence and machine learning…

physics.soc-ph2016

Competition and extinction explain the evolution of diversity in American automobiles

Erik Gjesfjeld, Jonathan Chang, Daniele Silvestro +2

One of the most remarkable aspects of our species is that while we show surprisingly little genetic diversity, we demonstrate astonishing amounts of cultural diversity. Perhaps mos…

cs.LG2022

MobILE: Model-Based Imitation Learning From Observation Alone

Rahul Kidambi, Jonathan Chang, Wen Sun

This paper studies Imitation Learning from Observations alone (ILFO) where the learner is presented with expert demonstrations that consist only of states visited by an expert (wit…

math.LO2021

2-adjoint equivalences in homotopy type theory

Daniel Carranza, Jonathan Chang, Chris Kapulkin +1

We introduce the notion of (half) 2-adjoint equivalences in Homotopy Type Theory and prove their expected properties. We formalized these results in the Lean Theorem Prover.

stat.AP2010

Hierarchical relational models for document networks

Jonathan Chang, David M. Blei

We develop the relational topic model (RTM), a hierarchical model of both network structure and node attributes. We focus on document networks, where the attributes of each documen…