collaborators

5 papers

cs.LG2026

Balancing Multimodal Learning through Label Space Reshaping

Xiaoyu Ma, Weijie Zhang, Yuanhao Gao +3

Multimodal learning often suffers from modality imbalance, where modalities that converge faster dominate optimization while others remain undertrained. Existing approaches typical…

cs.CV2026

SkyNative: A Native Multimodal Architecture for Remote Sensing Vision-Language Understanding

Xiao Yang, Ronghao Fu, Zhiwen Lin +10

Remote sensing vision-language models (RS-VLMs) commonly employ a pretrained vision encoder and a projection module to map image features into the token space of a large language m…

cs.CV2026

OmniEarth: A Benchmark for Evaluating Vision-Language Models in Geospatial Tasks

Ronghao Fu, Haoran Liu, Weijie Zhang +4

Vision-Language Models (VLMs) have demonstrated effective perception and reasoning capabilities on general-domain tasks, leading to growing interest in their application to Earth o…

cs.MM2025

Nexus: An Omni-Perceptive And -Interactive Model for Language, Audio, And Vision

Che Liu, Yingji Zhang, Dong Zhang +13

This work proposes an industry-level omni-modal large language model (LLM) pipeline that integrates auditory, visual, and linguistic modalities to overcome challenges such as limit…

cs.LG2025

Reflections from the 2024 Large Language Model (LLM) Hackathon for Applications in Materials Science and Chemistry

Yoel Zimmermann, Adib Bazgir, Zartashia Afzal +141

Here, we present the outcomes from the second Large Language Model (LLM) Hackathon for Applications in Materials Science and Chemistry, which engaged participants across global hyb…