activity
20232026
most citedOmniBind: Large-scale Omni Multimodal Representation via Binding Spaces

1 citations · 1 across the 3 of their papers we have counts for

collaborators

6 papers

cs.CL2026

PPL-Factory: Task-Aware and Budget-Aware Data Selection from Language Modeling to Reasoning

Hang Zhang, Warren J. Gross

Not all training samples contribute equally to large language model fine-tuning. Selecting informative training samples can reduce the computational cost while preserving downstrea…

cs.CV20241 cited

OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces

Zehan Wang, Ziang Zhang, Hang Zhang +5

Recently, human-computer interaction with various modalities has shown promising applications, like GPT-4o and Gemini. Given the foundational role of multimodal joint representatio…

cs.CV2024

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Zesen Cheng, Sicong Leng, Hang Zhang +8

In this paper, we present the VideoLLaMA 2, a set of Video Large Language Models (Video-LLMs) designed to enhance spatial-temporal modeling and audio understanding in video and aud…

cs.CL2024

XFormParser: A Simple and Effective Multimodal Multilingual Semi-structured Form Parser

Xianfu Cheng, Hang Zhang, Jian Yang +9

In the domain of Document AI, parsing semi-structured image form is a crucial Key Information Extraction (KIE) task. The advent of pre-trained multimodal models significantly empow…

cs.CL2023

SeaLLMs -- Large Language Models for Southeast Asia

Xuan-Phi Nguyen, Wenxuan Zhang, Xin Li +14

Despite the remarkable achievements of large language models (LLMs) in various tasks, there remains a linguistic bias that favors high-resource languages, such as English, often at…

cs.CV2023

M&M3D: Multi-Dataset Training and Efficient Network for Multi-view 3D Object Detection

Hang Zhang

In this research, I proposed a network structure for multi-view 3D object detection using camera-only data and a Bird's-Eye-View map. My work is based on a current key challenge do…