Publications (11)
The Overlooked Classifier in Human-Object Interaction Recognition
Ying Jin, Yinpeng Chen, Lijuan Wang +5
Human-Object Interaction (HOI) recognition is challenging due to two factors: (1) significant imbalance across classes and (2) requiring multiple labels per image. This paper shows…
Neural Voting Field for Camera-Space 3D Hand Pose Estimation
Lin Huang, Chung-Ching Lin, Kevin Lin +4
We present a unified framework for camera-space 3D hand pose estimation from a single RGB image based on 3D implicit representation. As opposed to recent works, most of which first…
Injecting Semantic Concepts into End-to-End Image Captioning
Zhiyuan Fang, Jianfeng Wang, Xiaowei Hu +5
Tremendous progress has been made in recent years in developing better image captioning models, yet most of them rely on a separate object detector to extract regional features. Re…
Generating Realistic Geology Conditioned on Physical Measurements with Generative Adversarial Networks
Emilien Dupont, Tuanfeng Zhang, Peter Tilke +2
An important problem in geostatistics is to build models of the subsurface of the Earth given physical measurements at sparse spatial locations. Typically, this is done using spati…
ART: Anonymous Region Transformer for Variable Multi-Layer Transparent Image Generation
Yifan Pu, Yiming Zhao, Zhicong Tang +14
Multi-layer image generation is a fundamental task that enables users to isolate, select, and edit specific image layers, thereby revolutionizing interactions with generative model…
Glyph-ByT5-v2: A Strong Aesthetic Baseline for Accurate Multilingual Visual Text Rendering
Zeyu Liu, Weicong Liang, Yiming Zhao +5
Recently, Glyph-ByT5 has achieved highly accurate visual text rendering performance in graphic design images. However, it still focuses solely on English and performs relatively po…
MPT: Mesh Pre-Training with Transformers for Human Pose and Mesh Reconstruction
Kevin Lin, Chung-Ching Lin, Lin Liang +2
Traditional methods of reconstructing 3D human pose and mesh from single images rely on paired image-mesh datasets, which can be difficult and expensive to obtain. Due to this limi…
MM-VID: Advancing Video Understanding with GPT-4V(ision)
Kevin Lin, Faisal Ahmed, Linjie Li +9
We present MM-VID, an integrated system that harnesses the capabilities of GPT-4V, combined with specialized tools in vision, audio, and speech, to facilitate advanced video unders…
Sum frequency generation spectroscopy of the attachment disc of a spider
Yue Zhao, Lin Liang, Yanrong Li +3
The pyriform silk of the attachment disc of a spider was studied using infrared-visible vibrational sum frequency generation (SFG) spectroscopy. The spider can attach dragline and…
Enhancing Understanding of Hydraulic Fracture Tip Advancement through Inversion of Low-Frequency Distributed Acoustic Sensing Data
Yongzan Liu, Lin Liang, Smaine Zeroug
Characterizing the fluid-driven fracture tip advancing process presents a significant challenge due to the difficulty of replicating real-world conditions in laboratory experiments…
The Overlooked Classifier in Human-Object Interaction Recognition
Ying Jin, Yinpeng Chen, Lijuan Wang +5
Human-Object Interaction (HOI) recognition is challenging due to two factors: (1) significant imbalance across classes and (2) requiring multiple labels per image. This paper shows…