papers

Publications (11)

cs.CV2022

The Overlooked Classifier in Human-Object Interaction Recognition

Ying Jin, Yinpeng Chen, Lijuan Wang +5

Human-Object Interaction (HOI) recognition is challenging due to two factors: (1) significant imbalance across classes and (2) requiring multiple labels per image. This paper shows…

cs.CV2023

Neural Voting Field for Camera-Space 3D Hand Pose Estimation

Lin Huang, Chung-Ching Lin, Kevin Lin +4

We present a unified framework for camera-space 3D hand pose estimation from a single RGB image based on 3D implicit representation. As opposed to recent works, most of which first…

cs.CV2022

Injecting Semantic Concepts into End-to-End Image Captioning

Zhiyuan Fang, Jianfeng Wang, Xiaowei Hu +5

Tremendous progress has been made in recent years in developing better image captioning models, yet most of them rely on a separate object detector to extract regional features. Re…

stat.ML2018

Generating Realistic Geology Conditioned on Physical Measurements with Generative Adversarial Networks

Emilien Dupont, Tuanfeng Zhang, Peter Tilke +2

An important problem in geostatistics is to build models of the subsurface of the Earth given physical measurements at sparse spatial locations. Typically, this is done using spati…

cs.CV2025

ART: Anonymous Region Transformer for Variable Multi-Layer Transparent Image Generation

Yifan Pu, Yiming Zhao, Zhicong Tang +14

Multi-layer image generation is a fundamental task that enables users to isolate, select, and edit specific image layers, thereby revolutionizing interactions with generative model…

cs.CV2024

Glyph-ByT5-v2: A Strong Aesthetic Baseline for Accurate Multilingual Visual Text Rendering

Zeyu Liu, Weicong Liang, Yiming Zhao +5

Recently, Glyph-ByT5 has achieved highly accurate visual text rendering performance in graphic design images. However, it still focuses solely on English and performs relatively po…

cs.CV2023

MPT: Mesh Pre-Training with Transformers for Human Pose and Mesh Reconstruction

Kevin Lin, Chung-Ching Lin, Lin Liang +2

Traditional methods of reconstructing 3D human pose and mesh from single images rely on paired image-mesh datasets, which can be difficult and expensive to obtain. Due to this limi…

cs.CV2023

MM-VID: Advancing Video Understanding with GPT-4V(ision)

Kevin Lin, Faisal Ahmed, Linjie Li +9

We present MM-VID, an integrated system that harnesses the capabilities of GPT-4V, combined with specialized tools in vision, audio, and speech, to facilitate advanced video unders…

physics.app-ph2021

Sum frequency generation spectroscopy of the attachment disc of a spider

Yue Zhao, Lin Liang, Yanrong Li +3

The pyriform silk of the attachment disc of a spider was studied using infrared-visible vibrational sum frequency generation (SFG) spectroscopy. The spider can attach dragline and…

physics.geo-ph2023

Enhancing Understanding of Hydraulic Fracture Tip Advancement through Inversion of Low-Frequency Distributed Acoustic Sensing Data

Yongzan Liu, Lin Liang, Smaine Zeroug

Characterizing the fluid-driven fracture tip advancing process presents a significant challenge due to the difficulty of replicating real-world conditions in laboratory experiments…

cs.CV2022

The Overlooked Classifier in Human-Object Interaction Recognition

Ying Jin, Yinpeng Chen, Lijuan Wang +5

Human-Object Interaction (HOI) recognition is challenging due to two factors: (1) significant imbalance across classes and (2) requiring multiple labels per image. This paper shows…