2 papers
cs.CV2025
MegaSR: Mining Customized Semantics and Expressive Guidance for Real-World Image Super-Resolution
Xinrui Li, Jinrong Zhang, Jianlong Wu +3
Text-to-image (T2I) models have ushered in a new era of real-world image super-resolution (Real-ISR) due to their rich internal implicit knowledge for multimodal learning. Although…
cs.CV2025
A Survey on Video Temporal Grounding with Multimodal Large Language Model
Jianlong Wu, Wei Liu, Ye Liu +4
The recent advancement in video temporal grounding (VTG) has significantly enhanced fine-grained video understanding, primarily driven by multimodal large language models (MLLMs).…