collaborators

5 papers

cs.CV2025

Negative Entity Suppression for Zero-Shot Captioning with Synthetic Images

Zimao Lu, Hui Xu, Bing Liu +1

Text-only training provides an attractive approach to address data scarcity challenges in zero-shot image captioning (ZIC), avoiding the expense of collecting paired image-text ann…

cs.CL2025

Fine-Tuning Large Multimodal Models for Automatic Pronunciation Assessment

Ke Wang, Wenning Wei, Yan Deng +2

Automatic Pronunciation Assessment (APA) is critical for Computer-Assisted Language Learning (CALL), requiring evaluation across multiple granularities and aspects. Large Multimoda…

cs.CL2025

Deep Contrastive Unlearning for Language Models

Estrid He, Tabinda Sarwar, Ibrahim Khalil +2

The past a few years have witnessed the great success of large language models, demonstrating powerful capabilities in comprehending textual data and generating human-like language…

cs.SD2025

Exploring the Potential of Large Multimodal Models as Effective Alternatives for Pronunciation Assessment

Ke Wang, Lei He, Kun Liu +3

Large Multimodal Models (LMMs) have demonstrated exceptional performance across a wide range of domains. This paper explores their potential in pronunciation assessment tasks, with…

cs.CV2024

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information

Ke Wang, Hong Xuan

Multi-modal large language models (MLLMs) utilizing instruction-following data, such as LLaVA, have achieved great progress in the industry. A major limitation in these models is t…