4 papers
TAT: Task-Adaptive Transformer for All-in-One Medical Image Restoration
Zhiwen Yang, Jiaju Zhang, Yang Yi +3
Medical image restoration (MedIR) aims to recover high-quality medical images from their low-quality counterparts. Recent advancements in MedIR have focused on All-in-One models ca…
To Trust Or Not To Trust Your Vision-Language Model's Prediction
Hao Dong, Moru Liu, Jian Liang +2
Vision-Language Models (VLMs) have demonstrated strong capabilities in aligning visual and textual modalities, enabling a wide range of applications in multimodal understanding and…
GhostShell: Streaming LLM Function Calls for Concurrent Embodied Programming
Jian Gong, Youwei Huang, Bo Yuan +17
We present GhostShell, a novel approach that leverages Large Language Models (LLMs) for streaming and concurrent behavioral programming in embodied systems. In contrast to predefin…
Adapting Vision-Language Models Without Labels: A Comprehensive Survey
Hao Dong, Lijun Sheng, Jian Liang +3
Vision-Language Models (VLMs) have demonstrated remarkable generalization capabilities across a wide range of tasks. However, their performance often remains suboptimal when direct…