3 papers
cs.CL2025
ToolMem: Enhancing Multimodal Agents with Learnable Tool Capability Memory
Yunzhong Xiao, Yangmin Li, Hewei Wang +2
Agents utilizing tools powered by large language models (LLMs) or vision-language models (VLMs) have demonstrated remarkable progress in diverse tasks across text and visual modali…
cs.CV2025
The Tenth NTIRE 2025 Efficient Super-Resolution Challenge Report
Bin Ren, Hang Guo, Lei Sun +143
This paper presents a comprehensive review of the NTIRE 2025 Challenge on Single-Image Efficient Super-Resolution (ESR). The challenge aimed to advance the development of deep mode…
cs.CV2025
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
Yunlong Tang, Jing Bi, Chao Huang +16
We present CAT-V (Caption AnyThing in Video), a training-free framework for fine-grained object-centric video captioning that enables detailed descriptions of user-selected objects…