activity
20242026
collaborators

7 papers

cs.LG2026

MinT: Managed Infrastructure for Training and Serving Millions of LLMs

Mind Lab, :, Song Cao +60

We present MindLab Toolkit (MinT), a managed infrastructure system for Low-Rank Adaptation (LoRA) post-training and online serving. MinT targets a setting where many trained polici…

cs.AI2025

ChestX-Reasoner: Advancing Radiology Foundation Models with Reasoning through Step-by-Step Verification

Ziqing Fan, Cheng Liang, Chaoyi Wu +3

Recent advances in reasoning-enhanced large language models (LLMs) and multimodal LLMs (MLLMs) have significantly improved performance in complex tasks, yet medical AI models often…

cs.LG2025

Combatting Dimensional Collapse in LLM Pre-Training Data via Diversified File Selection

Ziqing Fan, Siyuan Du, Shengchao Hu +5

Selecting high-quality pre-training data for large language models (LLMs) is crucial for enhancing their overall performance under limited computation budget, improving both traini…

cs.LG2024

Continual Task Learning through Adaptive Policy Self-Composition

Shengchao Hu, Yuhang Zhou, Ziqing Fan +4

Training a generalizable agent to continually learn a sequence of tasks from offline trajectories is a natural requirement for long-lived agents, yet remains a significant challeng…

cs.LG2024

Prompt Tuning with Diffusion for Few-Shot Pre-trained Policy Generalization

Shengchao Hu, Wanru Zhao, Weixiong Lin +3

Offline reinforcement learning (RL) methods harness previous experiences to derive an optimal policy, forming the foundation for pre-trained large-scale models (PLMs). When encount…

cs.LG2024

Task-Aware Harmony Multi-Task Decision Transformer for Offline Reinforcement Learning

Ziqing Fan, Shengchao Hu, Yuhang Zhou +4

The purpose of offline multi-task reinforcement learning (MTRL) is to develop a unified policy applicable to diverse tasks without the need for online environmental interaction. Re…