activity
20242026
most citedR2GenCSR: Mining Contextual and Residual Information for LLMs-based Radiology Report Generation

2 citations · 2 across the 3 of their papers we have counts for

collaborators

7 papers

cs.CV2026

Dynamic Parsing and Updating Natural Language Specification using VLMs for Robust Vision-Language Tracking

Xiao Wang, Liye Jin, Dan Xu +5

Vision-language tracking guided by natural language specifications leverages high-level semantic cues of target objects to substantially boost tracking accuracy and robustness. Exi…

cs.CV2026

T2I-VeRW: Part-level Fine-grained Perception for Text-to-Image Vehicle Retrieval

Xiao Wang, Ziwen Wang, Weizhe Kong +5

Vehicle Re-identification (Re-ID) aims to retrieve the most similar image to a given query from images captured by non-overlapping cameras. Extending vehicle Re-ID from image-only…

cs.CV20262 cited

R2GenCSR: Mining Contextual and Residual Information for LLMs-based Radiology Report Generation

Xiao Wang, Yuehang Li, Fuling Wang +3

Inspired by the tremendous success of Large Language Models (LLMs), existing Radiology report generation methods attempt to leverage large models to achieve better performance. The…

cs.CV2025

Sign Language Translation using Frame and Event Stream: Benchmark Dataset and Algorithms

Xiao Wang, Yuehang Li, Fuling Wang +5

Accurate sign language understanding serves as a crucial communication channel for individuals with disabilities. Current sign language translation algorithms predominantly rely on…

cs.CV2024

CXPMRG-Bench: Pre-training and Benchmarking for X-ray Medical Report Generation on CheXpert Plus Dataset

Xiao Wang, Fuling Wang, Yuehang Li +5

X-ray image-based medical report generation (MRG) is a pivotal area in artificial intelligence which can significantly reduce diagnostic burdens and patient wait times. Despite sig…

eess.IV2024

Pre-training on High Definition X-ray Images: An Experimental Study

Xiao Wang, Yuehang Li, Wentao Wu +5

Existing X-ray based pre-trained vision models are usually conducted on a relatively small-scale dataset (less than 500k samples) with limited resolution (e.g., 224 224).…