4 papers
From Sweep to Seam: Interleaved Cross-Block Post-Training Quantization
Achille Jacquemond, Yuma Ichikawa, Akira Sakai
Compressing large language models to two bits or fewer is increasingly feasible through block-wise post-training quantization; cross-block variants reconstruct neighboring Transfor…
LOGOS: A Living Logic for AI Agent Teams That Evolve With Humans
Yuma Ichikawa, Yamato Arai, Kosaku Kimura +2
AI agents are evolving from answer engines into persistent teams that use tools, delegate work, learn from experience, and modify the artifacts that shape their future behavior. Th…
Signs Beat Floats: Low-Rank Double-Binary Adaptation for On-Device Fine-Tuning
Yoshihiko Fujisawa, Yuma Ichikawa, Yudai Fujimoto +2
On-device adaptation of large language models commonly keeps a quantized base model frozen while training and deploying a small, task-specific LoRA adapter. In the unmerged adapter…
OneComp: One-Line Revolution for Generative AI Model Compression
Yuma Ichikawa, Keiji Kimura, Akihiro Yoshida +11
Deploying foundation models is increasingly constrained by memory footprint, latency, and hardware costs. Post-training compression can mitigate these bottlenecks by reducing the p…