activity
20242026
collaborators

6 papers

cs.LG2026

StableQAT: Stable Quantization-Aware Training at Ultra-Low Bitwidths

Tianyi Chen, Sihan Chen, Xiaoyi Qu +5

Quantization-aware training (QAT) is essential for deploying large models under strict memory and latency constraints, yet achieving stable and robust optimization at ultra-low bit…

cs.LG2025

Large Language Model Scaling Laws for Neural Quantum States in Quantum Chemistry

Oliver Knitter, Dan Zhao, Stefan Leichenauer +1

Scaling laws have been used to describe how large language model (LLM) performance scales with model size, training data size, or amount of computational resources. Motivated by th…

cs.AI2025

The Challenge of Teaching Reasoning to LLMs Without RL or Distillation

Wei Du, Branislav Kisacanin, George Armstrong +22

Reasoning-capable language models achieve state-of-the-art performance in diverse complex tasks by generating long, explicit Chain-of-Thought (CoT) traces. While recent works show…

cs.LG2025

Layer by Layer: Uncovering Hidden Representations in Language Models

Oscar Skean, Md Rifat Arefin, Dan Zhao +4

From extracting features to generating text, the outputs of large language models (LLMs) typically rely on the final layers, following the conventional wisdom that earlier layers c…

quant-ph2025

Comparative study of the ansätze in quantum language models

Jordi Del Castillo, Dan Zhao, Zongrui Pei

Quantum language models are the alternative to classical language models, which borrow concepts and methods from quantum machine learning and computational linguistics. While sever…

cs.LG2024

Retentive Neural Quantum States: Efficient Ansätze for Ab Initio Quantum Chemistry

Oliver Knitter, Dan Zhao, James Stokes +3

Neural-network quantum states (NQS) has emerged as a powerful application of quantum-inspired deep learning for variational Monte Carlo methods, offering a competitive alternative…