3 papers
cs.CV2026
AutoFocus: Uncertainty-Aware Active Visual Search for GUI Grounding
Ruilin Yao, Shegnwu Xiong, Tianyu Zou +2
Vision-Language Models (VLMs) have enabled autonomous GUI agents that translate natural language instructions into executable screen coordinates. However, grounding performance deg…
eess.IV2024
FastHDRNet: A new efficient method for SDR-to-HDR Translation
Siyuan Tian, Hao Wang, Yiren Rong +3
Modern displays nowadays possess the capability to render video content with a high dynamic range (HDR) and an extensive color gamut .However, the majority of available resources a…
cs.CV2024
CRA-PCN: Point Cloud Completion with Intra- and Inter-level Cross-Resolution Transformers
Yi Rong, Haoran Zhou, Lixin Yuan +3
Point cloud completion is an indispensable task for recovering complete point clouds due to incompleteness caused by occlusion, limited sensor resolution, etc. The family of coarse…