2 papers
cs.CV2025
QID: Efficient Query-Informed ViTs in Data-Scarce Regimes for OCR-free Visual Document Understanding
Binh M. Le, Shaoyuan Xu, Jinmiao Fu +6
In Visual Document Understanding (VDU) tasks, fine-tuning a pre-trained Vision-Language Model (VLM) with new datasets often falls short in optimizing the vision encoder to identify…
cs.CV2025
Temporal-Consistent Video Restoration with Pre-trained Diffusion Models
Hengkang Wang, Yang Liu, Huidong Liu +5
Video restoration (VR) aims to recover high-quality videos from degraded ones. Although recent zero-shot VR methods using pre-trained diffusion models (DMs) show good promise, they…