8 papers
Interaction-Limited Safe Continuous-Time RL for Dynamical Medical Treatment
Xun Shen, Yuepeng Wang, Akifumi Wachi +13
Dynamic medical treatment requires deciding treatment intensity and intervention timing, while patient states evolve continuously and adverse events may occur between clinical inte…
MedGym:A Unified Continuous-Time Benchmark for Dynamic Medical Treatment Reinforcement Learning
Yuepeng Wang, Ken Kawano, Yongqi Zhou +11
Medical treatment recommendation poses several challenges to reinforcement learning (RL): patient physiology evolves in continuous time, measurements and interventions are performe…
Signs Beat Floats: Low-Rank Double-Binary Adaptation for On-Device Fine-Tuning
Yoshihiko Fujisawa, Yuma Ichikawa, Yudai Fujimoto +2
On-device adaptation of large language models commonly keeps a quantized base model frozen while training and deploying a small, task-specific LoRA adapter. In the unmerged adapter…
Norm Anchors Make Model Edits Last
Mingda Liu, Zhenghan Zhu, Ze'an Miao +1
Sequential Locate-and-Edit (L&E) model editing can fail abruptly after many edits. We identify and formalize this failure as a positive norm-feedback loop, in which solved value ve…
OneComp: One-Line Revolution for Generative AI Model Compression
Yuma Ichikawa, Keiji Kimura, Akihiro Yoshida +11
Deploying foundation models is increasingly constrained by memory footprint, latency, and hardware costs. Post-training compression can mitigate these bottlenecks by reducing the p…
More Than Bits: Multi-Envelope Double Binary Factorization for Extreme Quantization
Yuma Ichikawa, Yoshihiko Fujisawa, Yudai Fujimoto +2
For extreme low-bit quantization of large language models (LLMs), Double Binary Factorization (DBF) is attractive as it enables efficient inference without sacrificing accuracy. Ho…