3 papers
cs.LG2025
Per-Axis Weight Deltas for Frequent Model Updates
Stefan Kuyumdzhiev, Radostin Cholakov
Serving many task-specialized LLM variants is often limited by the large size of fine-tuned checkpoints and the resulting cold-start latency. Since fine-tuned weights differ from t…
cs.LG2025
Fast Matrix Multiplications for Lookup Table-Quantized LLMs
Han Guo, William Brandon, Radostin Cholakov +3
The deployment of large language models (LLMs) is often constrained by memory bandwidth, where the primary bottleneck is the cost of transferring model parameters from the GPU's gl…
cs.CV2025
ImagiNet: A Multi-Content Benchmark for Synthetic Image Detection
Delyan Boychev, Radostin Cholakov
Recent generative models produce images with a level of authenticity that makes them nearly indistinguishable from real photos and artwork. Potential harmful use cases of these mod…