6 citations · 10 across the 6 of their papers we have counts for
6 papers
Text-promptable Object Counting via Quantity Awareness Enhancement
Miaojing Shi, Xiaowen Zhang, Zijie Yue +3
Recent advances in large vision-language models (VLMs) have shown remarkable progress in solving the text-promptable object counting problem. Representative methods typically speci…
Bootstrapping Vision-language Models for Self-supervised Remote Physiological Measurement
Zijie Yue, Miaojing Shi, Hanli Wang +3
Facial video-based remote physiological measurement is a promising research area for detecting human vital signs (e.g., heart rate, respiration frequency) in a non-contact way. Con…
Multimodal Large Language Model driven Radiology Report Generation with Clinical Knowledge Enhancement
Miaojing Shi, Tianyu Cen, Zijie Yue +3
Radiology report generation (RRG) has attracted significant attention due to its potential to reduce the workload of radiologists. The performance of current RRG approaches remains…
LoSh: Long-Short Text Joint Prediction Network for Referring Video Object Segmentation
Linfeng Yuan, Miaojing Shi, Zijie Yue +1
Referring video object segmentation (RVOS) aims to segment the target instance referred by a given text expression in a video clip. The text expression normally contains sophistica…
Facial Video-based Remote Physiological Measurement via Self-supervised Learning
Zijie Yue, Miaojing Shi, Shuai Ding
Facial video-based remote physiological measurement aims to estimate remote photoplethysmography (rPPG) signals from human face videos and then measure multiple vital signs (e.g. h…
Enhancing Space-time Video Super-resolution via Spatial-temporal Feature Interaction
Zijie Yue, Miaojing Shi
The target of space-time video super-resolution (STVSR) is to increase both the frame rate (also referred to as the temporal resolution) and the spatial resolution of a given video…