3 papers
cs.CV2026
StyleGallery: Training-free and Semantic-aware Personalized Style Transfer from Arbitrary Image References
Boyu He, Yunfan Ye, Chang Liu +3
Despite the advancements in diffusion-based image style transfer, existing methods are commonly limited by 1) semantic gap: the style reference could miss proper content semantics,…
cs.CV2025
ALLVB: All-in-One Long Video Understanding Benchmark
Xichen Tan, Yuanjing Luo, Yunfan Ye +2
From image to video understanding, the capabilities of Multi-modal LLMs (MLLMs) are increasingly powerful. However, most existing video understanding benchmarks are relatively shor…
cs.CV2025
RAG-Adapter: A Plug-and-Play RAG-enhanced Framework for Long Video Understanding
Xichen Tan, Yunfan Ye, Yuanjing Luo +3
Multi-modal Large Language Models (MLLMs) capable of video understanding are advancing rapidly. To effectively assess their video comprehension capabilities, long video understandi…