activity
20242026
collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

A unified multi-task framework enables interpretable chest radiograph analysis

Lijian Xu, Ziyu Ni, Xinglong Liu +3

While multimodal deep learning has advanced medical imaging analysis, existing black-box systems \textcolor{black}{may remain confined to isolated tasks, often overlooking} the tru…

cs.CV2025

VBCD: A Voxel-Based Framework for Personalized Dental Crown Design

Linda Wei, Chang Liu, Wenran Zhang +3

The design of restorative dental crowns from intraoral scans is labor-intensive for dental technicians. To address this challenge, we propose a novel voxel-based framework for auto…

cs.CV2025

One Leaf Reveals the Season: Occlusion-Based Contrastive Learning with Semantic-Aware Views for Efficient Visual Representation

Xiaoyu Yang, Lijian Xu, Hongsheng Li +1

This paper proposes a scalable and straightforward pre-training paradigm for efficient visual conceptual representation called occluded image contrastive learning (OCL). Our OCL ap…

cs.CV2024

MedViLaM: A multimodal large language model with advanced generalizability and explainability for medical data understanding and generation

Lijian Xu, Hao Sun, Ziyu Ni +2

Medicine is inherently multimodal and multitask, with diverse data modalities spanning text, imaging. However, most models in medical field are unimodal single tasks and lack good…

cs.CV2024

Pathology-knowledge Enhanced Multi-instance Prompt Learning for Few-shot Whole Slide Image Classification

Linhao Qu, Dingkang Yang, Dan Huang +4

Current multi-instance learning algorithms for pathology image analysis often require a substantial number of Whole Slide Images for effective training but exhibit suboptimal perfo…

cs.CV2024

Enhancing Visual Grounding and Generalization: A Multi-Task Cycle Training Approach for Vision-Language Models

Xiaoyu Yang, Lijian Xu, Hao Sun +2

Visual grounding (VG) occupies a pivotal position in multi-modality vision-language models. In this study, we propose ViLaM, a large multi-modality model, that supports multi-tasks…