3 papers
cs.CV2025
Player-Centric Multimodal Prompt Generation for Large Language Model Based Identity-Aware Basketball Video Captioning
Zeyu Xi, Haoying Sun, Yaofei Wu +5
Existing sports video captioning methods often focus on the action yet overlook player identities, limiting their applicability. Although some methods integrate extra information t…
cs.CV2025
Prohibited Items Segmentation via Occlusion-aware Bilayer Modeling
Yunhan Ren, Ruihuang Li, Lingbo Liu +1
Instance segmentation of prohibited items in security X-ray images is a critical yet challenging task. This is mainly caused by the significant appearance gap between prohibited it…
cs.CV2025
Multimodal LLMs Can Reason about Aesthetics in Zero-Shot
Ruixiang Jiang, Changwen Chen
The rapid technical progress of generative art (GenArt) has democratized the creation of visually appealing imagery. However, achieving genuine artistic impact - the kind that reso…