11 papers
HetCCL: Enabling Collective Communication For Mixed-Vendor Heterogeneous Clusters
Yuejie Wang, Tao Chang, Yuanyuan Zhao +10
Training Large Language Models (LLMs) on heterogeneous clusters presents significant challenges for collective communication, as hardware from multiple vendors introduces diverse n…
FreqCache: Accelerating Embodied VLN Models with Adaptive Frequency-Guided Token Caching
Zihao Zheng, Xingyue Zhou, Zhihao Mao +7
Vision-Language-Navigation (VLN) models exhibit excellent navigation accuracy but incur high computational overhead. Token caching has emerged as a promising training-free strategy…
Think While Watching: Online Streaming Segment-Level Memory for Multi-Turn Video Reasoning in Multimodal Large Language Models
Lu Wang, Zhuoran Jin, Yupu Hao +4
Multimodal large language models (MLLMs) have shown strong performance on offline video understanding, but most are limited to offline inference or have weak online reasoning, maki…
RoboBrain 2.5: Depth in Sight, Time in Mind
Huajie Tan, Enshen Zhou, Zhiyu Li +32
We introduce RoboBrain 2.5, a next-generation embodied AI foundation model that advances general perception, spatial reasoning, and temporal modeling through extensive training on…
TrimTokenator-LC: Towards Adaptive Visual Token Pruning for Large Multimodal Models with Long Contexts
Hao Zhang, Mengsi Lyu, Bo Huang +2
Large Multimodal Models (LMMs) have proven effective on various tasks. They typically encode visual inputs into Original Model sequences of tokens, which are then concatenated with…
PDTrim: Targeted Pruning for Prefill-Decode Disaggregation in Inference
Hao Zhang, Mengsi Lyu, Zhuo Chen +3
Large Language Models (LLMs) demonstrate exceptional capabilities across various tasks, but their deployment is constrained by high computational and memory costs. Model pruning pr…