6 papers
StableQAT: Stable Quantization-Aware Training at Ultra-Low Bitwidths
Tianyi Chen, Sihan Chen, Xiaoyi Qu +5
Quantization-aware training (QAT) is essential for deploying large models under strict memory and latency constraints, yet achieving stable and robust optimization at ultra-low bit…
Large Language Model Scaling Laws for Neural Quantum States in Quantum Chemistry
Oliver Knitter, Dan Zhao, Stefan Leichenauer +1
Scaling laws have been used to describe how large language model (LLM) performance scales with model size, training data size, or amount of computational resources. Motivated by th…
The Challenge of Teaching Reasoning to LLMs Without RL or Distillation
Wei Du, Branislav Kisacanin, George Armstrong +22
Reasoning-capable language models achieve state-of-the-art performance in diverse complex tasks by generating long, explicit Chain-of-Thought (CoT) traces. While recent works show…
Layer by Layer: Uncovering Hidden Representations in Language Models
Oscar Skean, Md Rifat Arefin, Dan Zhao +4
From extracting features to generating text, the outputs of large language models (LLMs) typically rely on the final layers, following the conventional wisdom that earlier layers c…
Comparative study of the ansätze in quantum language models
Jordi Del Castillo, Dan Zhao, Zongrui Pei
Quantum language models are the alternative to classical language models, which borrow concepts and methods from quantum machine learning and computational linguistics. While sever…
Retentive Neural Quantum States: Efficient Ansätze for Ab Initio Quantum Chemistry
Oliver Knitter, Dan Zhao, James Stokes +3
Neural-network quantum states (NQS) has emerged as a powerful application of quantum-inspired deep learning for variational Monte Carlo methods, offering a competitive alternative…