5 papers
NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR
Yuan Xie, Jiaqi Song, Guang Qiu +9
Integrating large language models (LLMs) into automatic speech recognition (ASR) has become a mainstream paradigm in recent years. Although existing LLM-based ASR models demonstrat…
TLoRA: Task-aware Low Rank Adaptation of Large Language Models
Weicheng Lin, Yi Zhang, Jiawei Dang +1
Low-Rank Adaptation (LoRA) has become a widely adopted parameter-efficient fine-tuning method for large language models, with its effectiveness largely influenced by the allocation…
Gaussian Shannon: High-Precision Diffusion Model Watermarking Based on Communication
Yi Zhang, Hongbo Huang, Liang-Jie Zhang
Diffusion models generate high-quality images but pose serious risks like copyright violation and disinformation. Watermarking is a key defense for tracing and authenticating AI-ge…
Training-Free Test-Time Adaptation with Brownian Distance Covariance in Vision-Language Models
Yi Zhang, Chun-Wun Cheng, Angelica I. Aviles-Rivero +2
Vision-language models suffer performance degradation under domain shift, limiting real-world applicability. Existing test-time adaptation methods are computationally intensive, re…
Feature Projection Learning for Better Vision-Language Reasoning
Yi Zhang, Weicheng Lin, Liang-Jie Zhang
Vision-Language Pre-Trained models, notably CLIP, that utilize contrastive learning have proven highly adept at extracting generalizable visual features. To inherit the well-learne…