8 citations · 10 across the 5 of their papers we have counts for
8 papers
ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
Hangjie Yuan, Yichen Qian, Zhiwei Tang +21
Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric chall…
CoT-lized Diffusion: Let's Reinforce T2I Generation Step-by-step
Zheyuan Liu, Munan Ning, Qihui Zhang +8
Current text-to-image (T2I) generation models struggle to align spatial composition with the input text, especially in complex scenes. Even layout-based approaches yield suboptimal…
2nd Place Solution to Google Landmark Retrieval 2021
Zhang Yuqi, Xu Xianzhe, Chen Weihua +4
This paper presents the 2nd place solution to the Google Landmark Retrieval 2021 Competition on Kaggle. The solution is based on a baseline with training tricks from person re-iden…
An Empirical Study of Vehicle Re-Identification on the AI City Challenge
Hao Luo, Weihua Chen, Xianzhe Xu +7
This paper introduces our solution for the Track2 in AI City Challenge 2021 (AICITY21). The Track2 is a vehicle re-identification (ReID) task with both the real-world data and synt…
City-Scale Multi-Camera Vehicle Tracking Guided by Crossroad Zones
Chong Liu, Yuqi Zhang, Hao Luo +6
Multi-Target Multi-Camera Tracking has a wide range of applications and is the basis for many advanced inferences and predictions. This paper describes our solution to the Track 3…
Collaborative Multi-Agent Multi-Armed Bandit Learning for Small-Cell Caching
Xianzhe Xu, Meixia Tao, Cong Shen
This paper investigates learning-based caching in small-cell networks (SCNs) when user preference is unknown. The goal is to optimize the cache placement in each small base station…