2 papers
cs.CV2024
EVLM: An Efficient Vision-Language Model for Visual Understanding
Kaibing Chen, Dong Shen, Hanwen Zhong +14
In the field of multi-modal language models, the majority of methods are built on an architecture similar to LLaVA. These models use a single-layer ViT feature as a visual prompt,…
cs.MM2024
A Multimodal Transformer for Live Streaming Highlight Prediction
Jiaxin Deng, Shiyao Wang, Dong Shen +4
Recently, live streaming platforms have gained immense popularity. Traditional video highlight detection mainly focuses on visual features and utilizes both past and future content…