12 citations · 13 across the 2 of their papers we have counts for
2 papers
cs.CV2024★ 1 cited
LLM-AD: Large Language Model based Audio Description System
Peng Chu, Jiang Wang, Andre Abrantes
The development of Audio Description (AD) has been a pivotal step forward in making video content more accessible and inclusive. Traditionally, AD production has demanded a conside…
cs.CV2023★ 12 cited
Advancing Vision Transformers with Group-Mix Attention
Chongjian Ge, Xiaohan Ding, Zhan Tong +4
Vision Transformers (ViTs) have been shown to enhance visual recognition through modeling long-range dependencies with multi-head self-attention (MHSA), which is typically formulat…