#quantization

try —

12 papers match

cs.CL2026

CACHE-UK: A Stability-Aware Memory Editor for Sequentially Updated Quantized LLMs in Finance

Anubhav Lakra, Yue Feng

The paper introduces CACHE-UK, a stability-aware memory editing framework for 4-bit quantized large language models used in UK finance, which reduces knowledge degradation during s…

#large language models#quantization#memory editing#financial domain
cs.LG2026

HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Models

Hei Yi Mak, Shadan Golestan, Hoang Le +10

The paper introduces HiFloat4, a 4-bit floating-point format and a Rollout Residual Quantization technique that enable end-to-end reinforcement learning post‑training of large lang…

#quantization#reinforcement learning#large language models#low-bit precision
cs.LG2026

Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents

Jiwon Jang, Kisu Yang, Heuiseok Lim +1

The paper evaluates 4-bit post‑training quantization of multi‑turn, tool‑calling LLM agents and finds that while standard scores remain unchanged, quantization substantially increa…

#quantization#large language models#tool‑calling agents#evaluation metrics
cs.AI2026

How Small Can You Go? A Controlled Study of LoRA Rank, Target Modules, and Quantization Trade-offs for Text-to-SQL on a 60M-Parameter Model

Mahendra Singh Rathor, Anagheem Azzam

The paper investigates how LoRA rank, adapted modules, and low‑bit quantization trade off accuracy and resource usage when fine‑tuning a 60 M‑parameter T5‑small model for the WikiS…

#parameter-efficient fine-tuning#LoRA#quantization#text-to-sql
cs.CV2026

When W4A4 Breaks Camouflaged Object Detection: Token-Group Dual-Constraint Activation Quantization

Tianqi Li, Wenyu Fang, Xin He +3

The paper proposes a post‑training 4‑bit activation quantization method for transformer‑based camouflaged object detection that mitigates token‑level range domination to preserve s…

#camouflaged object detection#quantization#transformer models#low‑bit inference
cs.LG2026

The Joint Effect of Quantization and Sampling Temperature on LLM Safety Alignment: A Factorial Analysis

Hari Prasad, Ritam Pal

The paper investigates how model quantization and higher sampling temperatures jointly affect the safety alignment of instruction-tuned large language models, finding that quantiza…

#large language models#quantization#sampling temperature#safety alignment
cs.LG2026

ExTernD: Expanded-Rank Ternary Decomposition Ternary LLM PTQ with Accuracy Approaching Any Quantization Level

Chethan Reddy G. P

The paper presents ExTernD, a post‑training factorization that expands the rank of ternary matrix decompositions to correct quantization errors, enabling large language models to a…

#quantization#large language models#ternary decomposition#post-training quantization
cs.CR2026

Silent Alarm: A J-Space Protocol for Comparing Danger Recognition Across Models and Quantization Levels

Roman Prosvirnin, Victor Minchenkov, Alexey Soldatov +1

The paper introduces JADR, a protocol that examines a language model's internal Jacobian representations (J-space) to assess danger recognition before any response is generated, en…

#safety evaluation#large language models#jacobian analysis#quantization
hep-th2026

Symmetries and Conservation Laws in Lie-Poisson Electrodynamics

M. A. Kurkov

The paper studies Lie-Poisson electrodynamics, a non‑Abelian nonlinear deformation of Maxwell theory, showing how a field redefinition maps its dynamics to standard electrodynamics…

#lie-poisson electrodynamics#gauge symmetry#conserved currents#field redefinition
cs.AI2026

Calibrated e-CUSUM Decoding for Quantized Reasoning Models: Why Token Log-Probability Is the Wrong Observable for Decoding Monitors

El Hassane Ettifouri, Ayoub Belfatmi, Mahaman Sanoussi Yahaya Alassan +1

The paper shows that using centered token log‑probability as a monitor for quantized reasoning language models is ineffective, and proposes a training‑free calibrated e‑CUSUM decod…

#quantization#language model decoding#monitoring#e-cusum
cs.DC2026

Decomposing Runtime, Kernel, and Quantization Speedups via a Matched FP16 Intermediate: A Hardware-Conditioned Case Study on Four NVIDIA RTX A5000 GPUs

Weijia Han, Lisha Qu

The paper analyzes how much of the reported speedup from quantized inference on NVIDIA RTX A5000 GPUs comes from reduced runtime versus kernel and quantization changes, using a mat…

#quantization#gpu inference performance#runtime analysis#kernel optimization
cs.LG2026

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs

Zijie Liu, Jie Peng, Jinhao Duan +7

The paper proposes a training‑free method that replicates heavily used experts and quantizes less important ones to rebalance workload in sparse mixture‑of‑experts large language m…

#mixture-of-experts#load balancing#large language models#inference optimization

One search, two signals: results blend meaning (embedding similarity, so papers that never use your words still surface) with keyword matches on titles, abstracts and summaries. Free, no sign-in needed.