5 papers · 1 filter
Decoupled Vision-Language System for Multimodal Understanding and Generation
Yifan Xu, Baochen Xiong, Xiaoshan Yang +3
We introduce a new architecture design for multimodal large language models (MLLMs), Libra, capable of both multimodal understanding and generation. Libra architecture contains one…
SAGE: A Visual Language Model for Anomaly Detection via Fact Enhancement and Entropy-aware Alignment
Guoxin Zang, Xue Li, Donglin Di +4
While Vision-Language Models (VLMs) have shown promising progress in general multimodal tasks, they often struggle in industrial anomaly detection and reasoning, particularly in de…
LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural Planning
Shibo Sun, Xue Li, Donglin Di +6
While large language models (LLMs) have advanced procedural planning for embodied AI systems through strong reasoning abilities, the integration of multimodal inputs and counterfac…
Cluster-Driven Expert Pruning for Mixture-of-Experts Large Language Models
Hongcheng Guo, Juntao Yao, Boyang Wang +5
Mixture-of-Experts (MoE) architectures have emerged as a promising paradigm for scaling large language models (LLMs) with sparse activation of task-specific experts. Despite their…
Building Dialogue Understanding Models for Low-resource Language Indonesian from Scratch
Donglin Di, Weinan Zhang, Yue Zhang +1
Making use of off-the-shelf resources of resource-rich languages to transfer knowledge for low-resource languages raises much attention recently. The requirements of enabling the m…