activity
20242026
collaborators

6 papers

cs.CL2026

OralAgent: Integrating Reasoning, Tools, and Knowledge for Interactive Dental Image Analysis

Jing Hao, Siyuan Dai, Yongxin Zhang +13

Dental image analysis plays a pivotal role in supporting accurate diagnosis and treatment planning in oral healthcare. Although recent advances have produced dental AI models for s…

cs.CV2026

Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception

Yanpeng Sun, Jing Hao, Ke Zhu +6

Training Large Multimodality Models (LMMs) relies on descriptive image caption that connects image and language. Existing methods for generating such captions often rely on distill…

cs.CV2025

VRP-SAM: SAM with Visual Reference Prompt

Yanpeng Sun, Jiahui Chen, Shan Zhang +7

In this paper, we propose a novel Visual Reference Prompt (VRP) encoder that empowers the Segment Anything Model (SAM) to utilize annotated reference images as prompts for segmenta…

cs.CV2025

DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation through Loopback Synergy

Ming Dai, Wenxuan Cheng, Jiang-jiang Liu +4

Referring Image Segmentation (RIS) is a challenging task that aims to segment objects in an image based on natural language expressions. While prior studies have predominantly conc…

cs.LG2024

Continual SFT Matches Multimodal RLHF with Negative Supervision

Ke Zhu, Yu Wang, Yanpeng Sun +4

Multimodal RLHF usually happens after supervised finetuning (SFT) stage to continually improve vision-language models' (VLMs) comprehension. Conventional wisdom holds its superiori…

cs.CV2024

Improving Multi-modal Large Language Model through Boosting Vision Capabilities

Yanpeng Sun, Huaxin Zhang, Qiang Chen +5

We focus on improving the visual understanding capability for boosting the vision-language models. We propose \textbf{Arcana}, a multiModal language model, which introduces two cru…