7 papers
A Novel Automatic Framework for Speaker Drift Detection in Synthesized Speech
Jia-Hong Huang, Seulgi Kim, Yi Chieh Liu +5
Recent diffusion-based text-to-speech (TTS) models achieve high naturalness and expressiveness, yet often suffer from speaker drift, a subtle, gradual shift in perceived speaker id…
MaCP: Minimal yet Mighty Adaptation via Hierarchical Cosine Projection
Yixian Shen, Qi Bi, Jia-Hong Huang +3
We present a new adaptation method MaCP, Minimal yet Mighty adaptive Cosine Projection, that achieves exceptional performance while requiring minimal parameters and memory for fine…
ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding
Shuai Wang, Ivona Najdenkoska, Hongyi Zhu +4
Understanding visual art requires reasoning across multiple perspectives -- cultural, historical, and stylistic -- beyond mere object recognition. While recent multimodal large lan…
SSH: Sparse Spectrum Adaptation via Discrete Hartley Transformation
Yixian Shen, Qi Bi, Jia-Hong Huang +3
Low-rank adaptation (LoRA) has been demonstrated effective in reducing the trainable parameter number when fine-tuning a large foundation model (LLM). However, it still encounters…
Gradient Weight-normalized Low-rank Projection for Efficient LLM Training
Jia-Hong Huang, Yixian Shen, Hongyi Zhu +2
Large Language Models (LLMs) have shown remarkable performance across various tasks, but the escalating demands on computational resources pose significant challenges, particularly…
Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models
Jia-Hong Huang, Hongyi Zhu, Yixian Shen +2
Evaluating the quality of automatically generated image descriptions is a complex task that requires metrics capturing various dimensions, such as grammaticality, coverage, accurac…