7 papers
RoadBench: A Vision-Language Foundation Model and Benchmark for Road Damage Understanding
Xi Xiao, Yunbei Zhang, Janet Wang +9
Accurate road damage detection is crucial for timely infrastructure maintenance and public safety, but existing vision-only datasets and models lack the rich contextual understandi…
Sensitivity-LoRA: Low-Load Sensitivity-Based Fine-Tuning for Large Language Models
Hao Zhang, Bo Huang, Zhenjia Li +6
Large Language Models (LLMs) have transformed both everyday life and scientific research. However, adapting LLMs from general-purpose models to specialized tasks remains challengin…
AdsQA: Towards Advertisement Video Understanding
Xinwei Long, Kai Tian, Peng Xu +10
Large language models (LLMs) have taken a great step towards AGI. Meanwhile, an increasing number of domain-specific problems such as math and programming boost these general-purpo…
Multimodal Fusion with Relational Learning for Molecular Property Prediction
Zhengyang Zhou, Yunrui Li, Pengyu Hong +1
Graph based molecular representation learning is essential for accurately predicting molecular properties in drug discovery and materials science; however, it faces significant cha…
Describe Anything in Medical Images
Xi Xiao, Yunbei Zhang, Thanh-Huy Nguyen +10
Localized image captioning has made significant progress with models like the Describe Anything Model (DAM), which can generate detailed region-specific descriptions without explic…
Advancing Drug Discovery with Enhanced Chemical Understanding via Asymmetric Contrastive Multimodal Learning
Yifei Wang, Yunrui Li, Lin Liu +2
The versatility of multimodal deep learning holds tremendous promise for advancing scientific research and practical applications. As this field continues to evolve, the collective…