activity
20242026
collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

Decoupled Vision-Language System for Multimodal Understanding and Generation

Yifan Xu, Baochen Xiong, Xiaoshan Yang +3

We introduce a new architecture design for multimodal large language models (MLLMs), Libra, capable of both multimodal understanding and generation. Libra architecture contains one…

cs.CL2025

SAGE: A Visual Language Model for Anomaly Detection via Fact Enhancement and Entropy-aware Alignment

Guoxin Zang, Xue Li, Donglin Di +4

While Vision-Language Models (VLMs) have shown promising progress in general multimodal tasks, they often struggle in industrial anomaly detection and reasoning, particularly in de…

cs.CL2025

LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural Planning

Shibo Sun, Xue Li, Donglin Di +6

While large language models (LLMs) have advanced procedural planning for embodied AI systems through strong reasoning abilities, the integration of multimodal inputs and counterfac…

cs.CL2025

Cluster-Driven Expert Pruning for Mixture-of-Experts Large Language Models

Hongcheng Guo, Juntao Yao, Boyang Wang +5

Mixture-of-Experts (MoE) architectures have emerged as a promising paradigm for scaling large language models (LLMs) with sparse activation of task-specific experts. Despite their…

cs.CL2024

Building Dialogue Understanding Models for Low-resource Language Indonesian from Scratch

Donglin Di, Weinan Zhang, Yue Zhang +1

Making use of off-the-shelf resources of resource-rich languages to transfer knowledge for low-resource languages raises much attention recently. The requirements of enabling the m…