agent learning 1large language models 1policy optimization 1reinforcement learning 1transition modeling 1
From the 1 of 13 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Learning to Select, Not Relearn: Hard-Routed Mixtures of Reasoning LoRAs
Seyed Alireza Molavi, Zhan Su, Yan Hu +3
Composing independently trained LoRA adapters into a single large language model is useful for multi-domain adaptation, especially when the original training data cannot be shared.…
cs.AI2026
Efficient Data Selection for Multimodal Models via Incremental Optimization Utility
Jinhao Jing, Qiannian Zhao, Chao Huang +1
The scaling of Large Multimodal Models (LMMs) is constrained by the quality-quantity trade-off inherent in synthetic data. Previous approaches, such as LLM-as-a-Judge, have proven…