6 citations · 6 across the 4 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
Images are Worth Variable Length of Representations
Lingjun Mao, Rodolfo Corona, Xin Liang +2
Most existing vision encoders map images into a fixed-length sequence of tokens, overlooking the fact that different images contain varying amounts of information. For example, a v…
cs.CV2024
Will the Inclusion of Generated Data Amplify Bias Across Generations in Future Image Classification Models?
Zeliang Zhang, Xin Liang, Mingqian Feng +2
As the demand for high-quality training data escalates, researchers have increasingly turned to generative models to create synthetic data, addressing data scarcity and enabling co…
cs.CV2024
Biomedical Visual Instruction Tuning with Clinician Preference Alignment
Hejie Cui, Lingjun Mao, Xin Liang +5
Recent advancements in multimodal foundation models have showcased impressive capabilities in understanding and reasoning with visual and textual information. Adapting these founda…