#quantization

topicquantization

12 papers · 1 filter

cs.CL2026

CACHE-UK: A Stability-Aware Memory Editor for Sequentially Updated Quantized LLMs in Finance

Anubhav Lakra, Yue Feng

The paper introduces CACHE-UK, a stability-aware memory editing framework for 4-bit quantized large language models used in UK finance, which reduces knowledge degradation during s…

cs.LG2026

Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents

Jiwon Jang, Kisu Yang, Heuiseok Lim +1

The paper evaluates 4-bit post‑training quantization of multi‑turn, tool‑calling LLM agents and finds that while standard scores remain unchanged, quantization substantially increa…

cs.LG2026

HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Models

Hei Yi Mak, Shadan Golestan, Hoang Le +10

The paper introduces HiFloat4, a 4-bit floating-point format and a Rollout Residual Quantization technique that enable end-to-end reinforcement learning post‑training of large lang…

cs.AI2026

How Small Can You Go? A Controlled Study of LoRA Rank, Target Modules, and Quantization Trade-offs for Text-to-SQL on a 60M-Parameter Model

Mahendra Singh Rathor, Anagheem Azzam

The paper investigates how LoRA rank, adapted modules, and low‑bit quantization trade off accuracy and resource usage when fine‑tuning a 60 M‑parameter T5‑small model for the WikiS…

cs.LG2026

ExTernD: Expanded-Rank Ternary Decomposition Ternary LLM PTQ with Accuracy Approaching Any Quantization Level

Chethan Reddy G. P

The paper presents ExTernD, a post‑training factorization that expands the rank of ternary matrix decompositions to correct quantization errors, enabling large language models to a…

cs.LG2026

The Joint Effect of Quantization and Sampling Temperature on LLM Safety Alignment: A Factorial Analysis

Hari Prasad, Ritam Pal

The paper investigates how model quantization and higher sampling temperatures jointly affect the safety alignment of instruction-tuned large language models, finding that quantiza…