papers

Publications (12)

cs.CV2024

M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining

Qingpei Guo, Furong Xu, Hanxiao Zhang +6

Vision-language foundation models like CLIP have revolutionized the field of artificial intelligence. Nevertheless, VLM models supporting multi-language, e.g., in both Chinese and…

cond-mat.mtrl-sci2020

Two-dimensional Janus van der Waals heterojunctions: a review of recent research progresses

Lin Ju, Mei Bie, Xiwei Zhang +2

Two-dimensional Janus van der Waals (vdW) heterojunctions, referring to the junction containing at least one Janus material, are found to exhibit tuneable electronic structures, wi…

cs.LG2024

Similarity is Not All You Need: Endowing Retrieval Augmented Generation with Multi Layered Thoughts

Chunjing Gan, Dan Yang, Binbin Hu +9

In recent years, large language models (LLMs) have made remarkable achievements in various domains. However, the untimeliness and cost of knowledge updates coupled with hallucinati…

cs.LG2024

G-Meta: Distributed Meta Learning in GPU Clusters for Large-Scale Recommender Systems

Youshao Xiao, Shangchun Zhao, Zhenglei Zhou +5

Recently, a new paradigm, meta learning, has been widely applied to Deep Learning Recommendation Models (DLRM) and significantly improves statistical performance, especially in col…

cs.DC2024

AntDT: A Self-Adaptive Distributed Training Framework for Leader and Straggler Nodes

Youshao Xiao, Lin Ju, Zhenglei Zhou +8

Many distributed training techniques like Parameter Server and AllReduce have been proposed to take advantage of the increasingly large data and rich features. However, stragglers…

cs.LG2020

Trust in AutoML: Exploring Information Needs for Establishing Trust in Automated Machine Learning Systems

Jaimie Drozdal, Justin Weisz, Dakuo Wang +6

We explore trust in a relatively new area of data science: Automated Machine Learning (AutoML). In AutoML, AI methods are used to generate and optimize machine learning models by a…

cs.LG2025

Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs

Ling Team, Binwei Zeng, Chao Huang +71

In this technical report, we tackle the challenges of training large-scale Mixture of Experts (MoE) models, focusing on overcoming cost inefficiency and resource limitations preval…

cs.LG2024

AntBatchInfer: Elastic Batch Inference in the Kubernetes Cluster

Siyuan Li, Youshao Xiao, Fanzhuang Meng +4

Offline batch inference is a common task in the industry for deep learning applications, but it can be challenging to ensure stability and performance when dealing with large amoun…

cs.CV2026

Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks

Qihua Dong, Kuo Yang, Lin Ju +6

Referring Expression Comprehension (REC) links language to region level visual perception. Standard benchmarks (RefCOCO, RefCOCO+, RefCOCOg) have progressed rapidly with multimodal…

cs.LG2024

An Adaptive Placement and Parallelism Framework for Accelerating RLHF Training

Youshao Xiao, Zhenglei Zhou, Fagui Mao +6

Recently, ChatGPT or InstructGPT like large language models (LLM) has made a significant impact in the AI world. Many works have attempted to reproduce the complex InstructGPT's tr…

cs.CL2025

Every Activation Boosted: Scaling General Reasoner to 1 Trillion Open Language Foundation

Ling Team, Ang Li, Ben Liu +138

We introduce Ling 2.0, a series reasoning-oriented language foundation built upon the principle that every activation boosts reasoning capability. Designed to scale from tens of bi…

cs.LG2023

Rethinking Memory and Communication Cost for Efficient Large Language Model Training

Chan Wu, Hanxiao Zhang, Lin Ju +8

Recently, various distributed strategies for large language model training have been proposed. However, these methods provided limited solutions for the trade-off between memory co…