7 papers
Phoenix-VL 1.5 Medium Technical Report
Team Phoenix, :, Arka Ray +29
We introduce Phoenix-VL 1.5 Medium, a 123B-parameter natively multimodal and multilingual foundation model, adapted to regional languages and the Singapore context. Developed as a…
Differences in Text Generated by Diffusion and Autoregressive Language Models
Zeyang Zhang, Chengwei Liang, Xingyan Chen +4
Diffusion language models (DLMs) are promising alternatives to autoregressive language models (ARMs), yet the intrinsic differences in their generated text remain underexplored. We…
Towards Multimodal Graph Large Language Model
Xin Wang, Zeyang Zhang, Linxin Xiao +3
Multi-modal graphs, which integrate diverse multi-modal features and relations, are ubiquitous in real-world applications. However, existing multi-modal graph learning methods are…
Revisiting Transformation Invariant Geometric Deep Learning: An Initial Representation Perspective
Ziwei Zhang, Xin Wang, Zeyang Zhang +2
Deep neural networks have achieved great success in the last decade. When designing neural networks to handle the ubiquitous geometric data such as point clouds and graphs, it is c…
Modular Machine Learning: An Indispensable Path towards New-Generation Large Language Models
Xin Wang, Haoyang Li, Haibo Chen +2
Large language models (LLMs) have substantially advanced machine learning research, including natural language processing, computer vision, data mining, etc., yet they still exhibi…
Dynamic Mixture of Curriculum LoRA Experts for Continual Multimodal Instruction Tuning
Chendi Ge, Xin Wang, Zeyang Zhang +5
Continual multimodal instruction tuning is crucial for adapting Multimodal Large Language Models (MLLMs) to evolving tasks. However, most existing methods adopt a fixed architectur…