5 papers
Yuan3.0 Ultra: A Trillion-Parameter Enterprise-Oriented MoE LLM
YuanLab. ai, :, Shawn Wu +25
We introduce Yuan3.0 Ultra, an open-source Mixture-of-Experts (MoE) large language model featuring 68.8B activated parameters and 1010B total parameters, specially designed to enha…
Yuan3.0 Flash: An Open Multimodal Large Language Model for Enterprise Applications
YuanLab. ai, :, Shawn Wu +24
We introduce Yuan3.0 Flash, an open-source Mixture-of-Experts (MoE) MultiModal Large Language Model featuring 3.7B activated parameters and 40B total parameters, specifically desig…
Yuan-TecSwin: A text conditioned Diffusion model with Swin-transformer blocks
Shaohua Wu, Tong Yu, Shenling Wang +1
Diffusion models have shown remarkable capacity in image synthesis based on their U-shaped architecture and convolutional neural networks (CNN) as basic blocks. The locality of the…
SQLfuse: Enhancing Text-to-SQL Performance through Comprehensive LLM Synergy
Tingkai Zhang, Chaoyu Chen, Cong Liao +6
Text-to-SQL conversion is a critical innovation, simplifying the transition from complex SQL to intuitive natural language queries, especially significant given SQL's prevalence in…
Yuan 2.0-M32: Mixture of Experts with Attention Router
Shaohua Wu, Jiangang Luo, Xi Chen +12
Yuan 2.0-M32, with a similar base architecture as Yuan-2.0 2B, uses a mixture-of-experts architecture with 32 experts of which 2 experts are active. A new router network, Attention…