5 papers
A Cloud-Edge System for Multimodal Clinical Screening in Resource-Constrained Rural Settings
Hei Ting, Chan, Chenwei Wu +8
Medical AI has demonstrated specialist-level diagnostic accuracy, yet these capabilities remain largely inaccessible in resource-constrained rural settings where bandwidth is scarc…
Test-Time Scaling for World Action Models via Zero-Shot Geometric Evaluation
Zesen Zhao, Minkyoung Cho, Hui shen +4
Test-time scaling improves foundation-model inference by spending additional computation, but robot control requires deciding whether extra compute is useful before executing an ac…
Dynamic Linear Attention
Xin Wang, Hui Shen, Boyuan Zheng +7
The scalability of Large Language Models (LLMs) to long contexts is fundamentally constrained by the quadratic complexity of standard attention, motivating the adoption of linear a…
CLAP: Contrastive Latent-space Prompt Optimization for End-to-end Autonomous Driving
Ruiyang Zhu, Yuehan He, Boyuan Zheng +4
End-to-end autonomous driving systems powered by Vision-Language-Action (VLA) models achieve strong performance on common driving scenarios, yet remain brittle in rare but safety-c…
MARS: Harmonizing Multimodal Convergence via Adaptive Rank Search
Minkyoung Cho, Insu Jang, Shuowei Jin +5
Fine-tuning Multimodal Large Language Models (MLLMs) with parameter-efficient methods like Low-Rank Adaptation (LoRA) is crucial for task adaptation. However, imbalanced training d…