activity
20222024
most citedSEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

52 citations · 62 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CV2024

S3Editor: A Sparse Semantic-Disentangled Self-Training Framework for Face Video Editing

Guangzhi Wang, Tianyi Chen, Kamran Ghasedi +6

Face attribute editing plays a pivotal role in various applications. However, existing methods encounter challenges in achieving high-quality results while preserving identity, edi…

cs.CL202352 cited

SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Bohao Li, Rui Wang, Guangzhi Wang +3

Based on powerful Large Language Models (LLMs), recent generative Multimodal Large Language Models (MLLMs) have gained prominence as a pivotal research area, exhibiting remarkable…

cs.CV20235 cited

What Makes for Good Visual Tokenizers for Large Language Models?

Guangzhi Wang, Yixiao Ge, Xiaohan Ding +2

We empirically investigate proper pre-training methods to build good visual tokenizers, making Large Language Models (LLMs) powerful Multimodal Large Language Models (MLLMs). In ou…

cs.CV2023

Text to Point Cloud Localization with Relation-Enhanced Transformer

Guangzhi Wang, Hehe Fan, Mohan Kankanhalli

Automatically localizing a position based on a few natural language instructions is essential for future robots to communicate and collaborate with humans. To approach this goal, w…

cs.RO20225 cited

STTAR: Surgical Tool Tracking using off-the-shelf Augmented Reality Head-Mounted Displays

Alejandro Martin-Gomez, Haowei Li, Tianyu Song +6

The use of Augmented Reality (AR) for navigation purposes has shown beneficial in assisting physicians during the performance of surgical procedures. These applications commonly re…

cs.CV2022

Chairs Can be Stood on: Overcoming Object Bias in Human-Object Interaction Detection

Guangzhi Wang, Yangyang Guo, Yongkang Wong +1

Detecting Human-Object Interaction (HOI) in images is an important step towards high-level visual comprehension. Existing work often shed light on improving either human and object…