15 citations · 37 across the 9 of their papers we have counts for
14 papers
AeroGround: A Comprehensive Benchmark for Aerial-Ground Collaborative Reasoning
Shenghong Yi, Lin Zhang, Muzian Li +6
Vision-language models (VLMs) have been widely employed in understanding and reasoning tasks for unmanned aerial vehicles (UAVs). Existing UAV benchmarks primarily focus on aerial-…
ReMix: Towards a Unified View of Consistent Character Generation and Editing
Benjia Zhou, Bin Fu, Pei Cheng +3
Recent advances in large-scale text-to-image diffusion models (e.g., FLUX.1) have greatly improved visual fidelity in consistent character generation and editing. However, existing…
CTTS: Collective Test-Time Scaling
Zhende Song, Shengji Tang, Peng Ye +4
Test-time scaling (TTS) has emerged as a promising, training-free approach for enhancing large language model (LLM) performance. However, the efficacy of existing methods, such as…
Decouple and Orthogonalize: A Data-Free Framework for LoRA Merging
Shenghe Zheng, Hongzhi Wang, Chenyu Huang +5
With more open-source models available for diverse tasks, model merging has gained attention by combining models into one, reducing training, storage, and inference costs. Current…
Lightweight Model Pre-training via Language Guided Knowledge Distillation
Mingsheng Li, Lin Zhang, Mingzhen Zhu +4
This paper studies the problem of pre-training for small models, which is essential for many mobile devices. Current state-of-the-art methods on this problem transfer the represent…
DualMamba: A Lightweight Spectral-Spatial Mamba-Convolution Network for Hyperspectral Image Classification
Jiamu Sheng, Jingyi Zhou, Jiong Wang +2
The effectiveness and efficiency of modeling complex spectral-spatial relations are both crucial for Hyperspectral image (HSI) classification. Most existing methods based on CNNs a…