10 papers
SkillTrace: Traversing a Query-Skill Graph for Composable LLM Agents
Yue Yao, Shengyuan Wang, Xin Chen +4
Large language model agents increasingly solve complex tasks by composing reusable skills from a library. To address this, the key challenge is not merely to retrieve individually…
DY-LUT: Depth-Aware YCbCr Lookup Tables for Real-Time Underwater Image Enhancement
Cunhao Zhu, Xiangtao Kong, Dongliang Xu +3
Underwater image enhancement is challenged by spatially non-uniform, wavelength-dependent attenuation. Propagation distance and wavelength govern this degradation, while YCbCr sepa…
No Training, Better Flights: Test-Time Scaled VLMs for UAV Navigation
Feinan Cheng, Dongliang Xu, Wenli Nong +4
Test-time scaling offers a promising method to improve the inference performance of Vision-Language Models (VLMs) without additional training. Existing approaches to vision-languag…
Solve the Missing First Step: Can VLMs Standardize Raw Heterogeneous Medical Data?
Xin Chen, Dongliang Xu, Cunhao Zhu +5
As vision-language models (VLMs) are increasingly applied to medical AI, existing benchmarks mainly focus on evaluating their diagnosis ability over given medical images and texts,…
Thought Graph Traversal for Test-time Scaling in Chest X-ray VLLMs
Yue Yao, Zelin Wen, Yan Tong +5
Test-time scaling offers a promising way to improve the reasoning performance of vision-language large models (VLLMs) without additional training. In this paper, we explore a simpl…
Mobile GUI Agent Privacy Personalization with Trajectory Induced Preference Optimization
Zhixin Lin, Jungang Li, Dongliang Xu +5
Mobile GUI agents powered by Multimodal Large Language Models (MLLMs) can execute complex tasks on mobile devices. Despite this progress, most existing systems still optimize task…