From the 1 of 17 linked papers with an AI index.
17 papers
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
Changhai Zhou, Kieran Liu, Yuhua Zhou +17
LongStraw introduces an execution framework that enables reinforcement‑learning post‑training on million‑token prompts using a fixed GPU budget by separating prompt evaluation from…
SyncLoop: A Multimodal Dual-Loop Framework for Self-Improving Mathematical Reasoning
Xiuwei Chen, Wentao Hu, Hanhui Li +9
Recent advances in multimodal large language models (MLLMs) have shown impressive reasoning capabilities. However, further enhancing existing MLLMs necessitates high-quality vision…
KEMO: Event-Driven Keyframe Memory for Long-Horizon Robot Manipulation with VLA Policies
Yihan Zeng, Minghao Ye, Yiyuan Chen +4
Long-horizon robot manipulation remains challenging because similar observations may occur at different execution stages, while the appropriate action depends on previously complet…
FedSmoothLoRA: Toward Smoother and Faster Convergence in Federated Low-Rank Adaptation
Zehao Wang, Guanglei Yang, Yihan Zeng +4
Federated fine-tuning of foundation models with Low-Rank Adaptation (LoRA) provides an efficient solution for reducing communication and computation costs while preserving data loc…
AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution
Zehao Wang, Yihan Zeng, Zidong Gong +5
Post-training via Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) is crucial for enhancing reasoning in Multimodal Large Language Models (MLLMs), yet existing paradigm…
OPERA: An Agent for Image Restoration with End-to-End Joint Planning-Execution Optimization
Feng Zhu, Shuyang Xie, Yihan Zeng +2
Real-world image restoration is challenging due to complex and interacting mixed degradations. Recent agent-based approaches address this problem by composing multiple task-specifi…