5 papers
From Natural Alignment to Conditional Controllability in Multimodal Dialogue
Zeyu Jin, Songtao Zhou, Haoyu Wang +5
The recent advancement of Artificial Intelligence Generated Content (AIGC) has led to significant strides in modeling human interaction, particularly in the context of multimodal d…
Towards Automatic Soccer Commentary Generation with Knowledge-Enhanced Visual Reasoning
Zeyu Jin, Xiaoyu Qin, Songtao Zhou +2
Soccer commentary plays a crucial role in enhancing the soccer game viewing experience for audiences. Previous studies in automatic soccer commentary generation typically adopt an…
ChartHal: A Fine-grained Framework Evaluating Hallucination of Large Vision Language Models in Chart Understanding
Xingqi Wang, Yiming Cui, Xin Yao +3
Large Vision-Language Models (LVLMs) have recently demonstrated remarkable progress, yet hallucination remains a critical barrier, particularly in chart understanding, which requir…
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos
Qixin Wang, Songtao Zhou, Zeyu Jin +3
Automatic video commentary systems are widely used on multimedia social media platforms to extract factual information about video content. However, current systems may overlook es…
Minimal Impact ControlNet: Advancing Multi-ControlNet Integration
Shikun Sun, Min Zhou, Zixuan Wang +7
With the advancement of diffusion models, there is a growing demand for high-quality, controllable image generation, particularly through methods that utilize one or multiple contr…