2 citations · 2 across the 3 of their papers we have counts for
7 papers
Dynamic Parsing and Updating Natural Language Specification using VLMs for Robust Vision-Language Tracking
Xiao Wang, Liye Jin, Dan Xu +5
Vision-language tracking guided by natural language specifications leverages high-level semantic cues of target objects to substantially boost tracking accuracy and robustness. Exi…
T2I-VeRW: Part-level Fine-grained Perception for Text-to-Image Vehicle Retrieval
Xiao Wang, Ziwen Wang, Weizhe Kong +5
Vehicle Re-identification (Re-ID) aims to retrieve the most similar image to a given query from images captured by non-overlapping cameras. Extending vehicle Re-ID from image-only…
R2GenCSR: Mining Contextual and Residual Information for LLMs-based Radiology Report Generation
Xiao Wang, Yuehang Li, Fuling Wang +3
Inspired by the tremendous success of Large Language Models (LLMs), existing Radiology report generation methods attempt to leverage large models to achieve better performance. The…
Sign Language Translation using Frame and Event Stream: Benchmark Dataset and Algorithms
Xiao Wang, Yuehang Li, Fuling Wang +5
Accurate sign language understanding serves as a crucial communication channel for individuals with disabilities. Current sign language translation algorithms predominantly rely on…
CXPMRG-Bench: Pre-training and Benchmarking for X-ray Medical Report Generation on CheXpert Plus Dataset
Xiao Wang, Fuling Wang, Yuehang Li +5
X-ray image-based medical report generation (MRG) is a pivotal area in artificial intelligence which can significantly reduce diagnostic burdens and patient wait times. Despite sig…
Pre-training on High Definition X-ray Images: An Experimental Study
Xiao Wang, Yuehang Li, Wentao Wu +5
Existing X-ray based pre-trained vision models are usually conducted on a relatively small-scale dataset (less than 500k samples) with limited resolution (e.g., 224 224).…