3 papers
cs.CL2026
MegaTrain: Full Precision Training of 100B+ Parameter Large Language Models on a Single GPU
Zhengqing Yuan, Hanchi Sun, Lichao Sun +1
We present MegaTrain, a memory-centric system that efficiently trains 100B+ parameter large language models at full precision on a single GPU. Unlike traditional GPU-centric system…
cs.AI2026
Expert Threshold Routing for Autoregressive Language Modeling with Dynamic Computation Allocation and Load Balancing
Hanchi Sun, Yixin Liu, Yonghui Wu +1
Token-choice Mixture-of-Experts (TC-MoE) routes each token to a fixed number of experts, limiting dynamic computation allocation and requiring auxiliary losses to maintain load bal…
cs.CV2024
Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination
Haojie Zheng, Tianyang Xu, Hanchi Sun +3
Multimodal large language models (MLLMs) have advanced the integration of visual and linguistic modalities, establishing themselves as the dominant paradigm for visual-language tas…