papers

Publications (14)

eess.SY2023

Optimal Cruise Airspeed for Hybrid-Electric and Electric Aircraft: Applications to Air Mobility

Steven Li, Luis Rodrigues

Electric and hybrid-electric aircraft can help our society transition towards more sustainable aviation and lower greenhouse gas (GHG) emissions. This paper provides solutions to m…

cs.LG2025

R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference

Zhenyu Zhang, Zechun Liu, Yuandong Tian +3

Large Language Models (LLMs), while demonstrating remarkable capabilities across various applications, present significant challenges during inference due to their substantial mode…

astro-ph.IM2020

The Balloon-Borne Large Aperture Submillimeter Telescope Observatory

Ian Lowe, Gabriele Coppi, Peter A. R. Ade +30

The BLAST Observatory is a proposed superpressure balloon-borne polarimeter designed for a future ultra-long duration balloon campaign from Wanaka, New Zealand. To maximize scienti…

cs.LG2026

MoE-Spec: Expert Budgeting for Efficient Speculative Decoding

Bradley McDanel, Steven Li, Sruthikesh Surineni +1

Speculative decoding accelerates Large Language Model (LLM) inference by verifying multiple drafted tokens in parallel. However, for Mixture-of-Experts (MoE) models, this paralleli…

eess.SY2026

Minimum Energy Cruise of All-Electric Aircraft with Applications to Advanced Air Mobility

Steven Li, Luis Rodrigues

Electrified propulsion is expected to play an important role in the sustainable development of Advanced Air Mobility (AAM). However, the limited energy density of batteries motivat…

cs.SD2025

Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction

Advait Gosai, Tyler Vuong, Utkarsh Tyagi +8

End-to-end (E2E) spoken dialogue systems are increasingly replacing cascaded pipelines for voice-based human-AI interaction, processing raw audio directly without intermediate tran…

physics.ins-det2024

Thermal architecture for a cryogenic super-pressure balloon payload: design and development of the Taurus flight cryostat

Simon Tartakovsky, Alexandre E. Adler, Jason E. Austermann +20

We describe the cryogenic system being developed for Taurus: a super-pressure balloon-borne microwave polarimeter scheduled to fly in 2027. The Taurus cryogenic system consists of…

cs.LG2026

JacQuant: STE-Free Quantization-Aware Training via Learned Jacobian Surrogates

Kai Yi, Vignesh Vivekraja, Harshit Khaitan +1

Quantization-aware training (QAT) is widely deployed but typically relies on the Straight-Through Estimator (STE), which passes gradients through non-differentiable quantizers by f…

cs.CV2026

Rethinking Model Efficiency: Multi-Agent Inference with Large Models

Sixun Dong, Juhua Hu, Steven Li +2

Most vision-language models (VLMs) apply a large language model (LLM) as the decoder, where the response tokens are generated sequentially through autoregression. Therefore, the nu…

astro-ph.IM2024

Instrument Overview of Taurus: A Balloon-borne CMB and Dust Polarization Experiment

Jared L. May, Alexandre E. Adler, Jason E. Austermann +20

Taurus is a balloon-borne cosmic microwave background (CMB) experiment optimized to map the E-mode polarization and Galactic foregrounds at the largest angular scales (

cs.LG2026

WinQ: Accelerating Quantization-Aware Training of Language Models Around Saddle Points

Dongyue Li, Zechun Liu, Kai Yi +6

Quantization-aware training (QAT) is widely adopted to quantize language models by training full-precision weights using gradients from the quantized model. The main bottleneck is…

astro-ph.IM2026

The EXoplanet Climate Infrared TElescope (EXCITE): A balloon-borne mission to measure spectroscopic phase curves of transiting hot Jupiters

Timothy D. Rehm, Caitlyn Altermatt, Lee Bernard +26

The EXoplanet Climate Infrared TElescope (EXCITE) is a balloon-borne mission dedicated to measuring spectroscopic phase curves of hot Jupiter-type exoplanets. Phase curve measureme…

cs.LG2026

Quantized Reasoning Models Think They Need to Think Longer, but They Do Not

Sanae Lotfi, Polina Kirichenko, Steven Li +1

Post-training quantization (PTQ) is widely used to deploy large language models efficiently, but its effect on reasoning models is not well understood. Across math, coding, and sci…

cs.CL2026

CLAA: Cross-Layer Attention Aggregation for Accelerating LLM Prefill

Bradley McDanel, Steven Li, Harshit Khaitan

The prefill stage in long-context LLM inference remains a computational bottleneck. Recent token-ranking heuristics accelerate inference by selectively processing a subset of seman…