7 papers
SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models
Chenshuang Zhang, Kyeong Seon Kim, Chengxin Liu +1
Despite the success of audio-visual large-language models (LLMs), they can produce plausible but ungrounded outputs, termed hallucination. Existing benchmarks focus on environmenta…
Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow
Chengxin Liu, Wonseok Choi, Chenshuang Zhang +1
Vision-Language Models (VLMs) have demonstrated strong capability in a wide range of tasks such as visual recognition, document parsing, and visual grounding. Nevertheless, recent…
Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision
Chenshuang Zhang, Kang Zhang, Joon Son Chung +3
Distinguishing visually similar objects by their motion remains a critical challenge in computer vision. Although supervised trackers show promise, contemporary self-supervised tra…
Text-to-image Diffusion Models in Generative AI: A Survey
Chenshuang Zhang, Chaoning Zhang, Mengchun Zhang +2
This survey reviews the progress of diffusion models in generating images from text, ~\textit{i.e.} text-to-image diffusion models. As a self-contained work, this survey starts wit…
A Survey on Segment Anything Model (SAM): Vision Foundation Model Meets Prompt Engineering
Chaoning Zhang, Joseph Cho, Fachrina Dewi Puspitasari +11
The Segment Anything Model (SAM), developed by Meta AI Research, represents a significant breakthrough in computer vision, offering a robust framework for image and video segmentat…
Towards Understanding Dual BN In Hybrid Adversarial Training
Chenshuang Zhang, Chaoning Zhang, Kang Zhang +3
There is a growing concern about applying batch normalization (BN) in adversarial training (AT), especially when the model is trained on both adversarial samples and clean samples…