Publications (48)
GenTron: Diffusion Transformers for Image and Video Generation
Shoufa Chen, Mengmeng Xu, Jiawei Ren +7
In this study, we explore Transformer-based diffusion models for image and video generation. Despite the dominance of Transformer architectures in various fields due to their flexi…
Aggregated Sparse Attention for Steering Angle Prediction
Sen He, Dmitry Kangin, Yang Mi +1
In this paper, we apply the attention mechanism to autonomous driving for steering angle prediction. We propose the first model, applying the recently introduced sparse attention m…
Hybrid Graph Neural Networks for Few-Shot Learning
Tianyuan Yu, Sen He, Yi-Zhe Song +1
Graph neural networks (GNNs) have been used to tackle the few-shot learning (FSL) problem and shown great potentials under the transductive setting. However under the inductive set…
Prediction Calibration for Generalized Few-shot Semantic Segmentation
Zhihe Lu, Sen He, Da Li +2
Generalized Few-shot Semantic Segmentation (GFSS) aims to segment each image pixel into either base classes with abundant training examples or novel classes with only a handful of…
UWB-Fat: Non-Intrusive Body Fat Measurement Using Commodity Ultra-Wideband Radar
Haotang Li, Yili Ren, Zhenyu Qi +6
Body fat percentage and its spatial distribution are clinically important health indicators. However, existing measurement methods often impose a tradeoff between accuracy and acce…
UWB-PostureGuard: A Privacy-Preserving RF Sensing System for Continuous Ergonomic Sitting Posture Monitoring
Haotang Li, Zhenyu Qi, Sen He +6
Improper sitting posture during prolonged computer use has become a significant public health concern. Traditional posture monitoring solutions face substantial barriers, including…
Learning Garment DensePose for Robust Warping in Virtual Try-On
Aiyu Cui, Sen He, Tao Xiang +1
Virtual try-on, i.e making people virtually try new garments, is an active research area in computer vision with great commercial applications. Current virtual try-on methods usual…
Image Captioning through Image Transformer
Sen He, Wentong Liao, Hamed R. Tavakoli +3
Automatic captioning of images is a task that combines the challenges of image analysis and text generation. One important aspect in captioning is the notion of attention: How to d…
Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
Haozhe Liu, Ding Liu, Mingchen Zhuge +16
We introduce MoS (Mixture of States), a novel fusion paradigm for multimodal diffusion models that merges modalities using flexible, state-based interactions. The core of MoS is a…
Seamless Interaction: Dyadic Audiovisual Motion Modeling and Large-Scale Dataset
Vasu Agrawal, Akinniyi Akinyemi, Kathryn Alvero +81
Human communication involves a complex interplay of verbal and nonverbal signals, essential for conveying meaning and achieving interpersonal goals. To develop socially intelligent…
Scaling Zero-Shot Reference-to-Video Generation
Zijian Zhou, Shikun Liu, Haozhe Liu +14
Reference-to-video (R2V) generation aims to synthesize videos that align with a text prompt while preserving the subject identity from reference images. However, current R2V method…
Single Stage Multi-Pose Virtual Try-On
Sen He, Yi-Zhe Song, Tao Xiang
Multi-pose virtual try-on (MPVTON) aims to fit a target garment onto a person at a target pose. Compared to traditional virtual try-on (VTON) that fits the garment but keeps the po…
PaceVGGT: Pre-Alternating-Attention Token Pruning for Visual Geometry Transformers
Haotang Li, Zhenyu Qi, Shaohan Henry Wang +5
Visual Geometry Transformer (VGGT) is a strong feed-forward model for multiple 3D tasks, but its Alternating-Attention (AA) stack scales quadratically in the total token count, mak…
MarDini: Masked Autoregressive Diffusion for Video Generation at Scale
Haozhe Liu, Shikun Liu, Zijian Zhou +12
We introduce MarDini, a new family of video diffusion models that integrate the advantages of masked auto-regression (MAR) into a unified diffusion model (DM) framework. Here, MAR…
DynamicLip: Shape-Independent Continuous Authentication via Lip Articulator Dynamics
Huashan Chen, Yifan Xu, Yue Feng +6
Biometrics authentication has become increasingly popular due to its security and convenience; however, traditional biometrics are becoming less desirable in scenarios such as new…
FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing
Yuren Cong, Mengmeng Xu, Christian Simon +7
Text-to-video editing aims to edit the visual appearance of a source video conditional on textual prompts. A major challenge in this task is to ensure that all frames in the edited…
Disentangled Lifespan Face Synthesis
Sen He, Wentong Liao, Michael Ying Yang +3
A lifespan face synthesis (LFS) model aims to generate a set of photo-realistic face images of a person's whole life, given only one snapshot as reference. The generated face image…
Salient Region Segmentation
Sen He, Nicolas Pugeault
Saliency prediction is a well studied problem in computer vision. Early saliency models were based on low-level hand-crafted feature derived from insights gained in neuroscience an…
Learning Flow Fields in Attention for Controllable Person Image Generation
Zijian Zhou, Shikun Liu, Xiao Han +11
Controllable person image generation aims to generate a person image conditioned on reference images, allowing precise control over the person's appearance or pose. However, prior…
TransText: Alpha-as-RGB Representation for Transparent Text Animation
Fei Zhang, Zijian Zhou, Bohao Tang +9
We introduce the first method, to the best of our knowledge, for adapting image-to-video models to layer-aware text (glyph) animation, a capability critical for practical dynamic v…
Adaptive Caching for Faster Video Generation with Diffusion Transformers
Kumara Kahatapitiya, Haozhe Liu, Sen He +5
Generating temporally-consistent high-fidelity videos can be computationally expensive, especially over longer temporal spans. More-recent Diffusion Transformers (DiTs) -- despite…
Context-Aware Layout to Image Generation with Enhanced Object Appearance
Sen He, Wentong Liao, Michael Ying Yang +4
A layout to image (L2I) generation model aims to generate a complicated image containing multiple objects (things) against natural background (stuff), conditioned on a given layout…
Human Attention in Image Captioning: Dataset and Analysis
Sen He, Hamed R. Tavakoli, Ali Borji +1
In this work, we present a novel dataset consisting of eye movements and verbal descriptions recorded synchronously over images. Using this data, we study the differences in human…
An Empirical Study on Virtual Reality Software Security Weaknesses
Yifan Xu, Jinfu Chen, Zhenyu Qi +5
Virtual Reality (VR) has emerged as a transformative technology across industries, yet its security weaknesses, including vulnerabilities, are underinvestigated. This study investi…
Isolating Recurring Execution-Dependent Abnormal Patterns on NISQ Quantum Devices
Zhenyu Qi, Haotang Li, Mominul Islam +3
Quantum devices increasingly expose a fundamental gap between compiler-modeled noise and hardware execution. Today's compilers approximate noise as calibration-derived costs over g…
Self-supervised Graph Transformer with Contrastive Learning for Brain Connectivity Analysis towards Improving Autism Detection
Yicheng Leng, Syed Muhammad Anwar, Islem Rekik +2
Functional Magnetic Resonance Imaging (fMRI) provides useful insights into the brain function both during task or rest. Representing fMRI data using correlation matrices is found t…
VecGlypher: Unified Vector Glyph Generation with Language Models
Xiaoke Huang, Bhavul Gauri, Kam Woh Ng +12
Vector glyphs are the atomic units of digital typography, yet most learning-based pipelines still depend on carefully curated exemplar sheets and raster-to-vector postprocessing, w…
TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
Zhiheng Liu, Weiming Ren, Haozhe Liu +22
Unified multimodal models (UMMs) aim to jointly perform multimodal understanding and generation within a single framework. We present TUNA, a native UMM that builds a unified conti…
Diffused Heads: Diffusion Models Beat GANs on Talking-Face Generation
MichaÅ StypuÅkowski, Konstantinos Vougioukas, Sen He +3
Talking face generation has historically struggled to produce head movements and natural facial expressions without guidance from additional reference videos. Recent developments i…
Style-Based Global Appearance Flow for Virtual Try-On
Sen He, Yi-Zhe Song, Tao Xiang
Image-based virtual try-on aims to fit an in-shop garment into a clothed person image. To achieve this, a key step is garment warping which spatially aligns the target garment with…
What Catches the Eye? Visualizing and Understanding Deep Saliency Models
Sen He, Ali Borji, Yang Mi +1
Deep convolutional neural networks have demonstrated high performances for fixation prediction in recent years. How they achieve this, however, is less explored and they remain to…
PhysDepth: Plug-and-Play Physical Refinement for Monocular Depth Estimation in Challenging Environments
Kebin Peng, Haotang Li, Zhenyu Qi +6
State-of-the-art monocular depth estimation (MDE) models often struggle in challenging environments, primarily because they overlook robust physical information. To demonstrate thi…
A Benchmarking Framework for Interactive 3D Applications in the Cloud
Tianyi Liu, Sen He, Sunzhou Huang +4
With the growing popularity of cloud gaming and cloud virtual reality (VR), interactive 3D applications have become a major type of workloads for the cloud. However, despite their…
Still Camouflage, Moving Illusion: View-Induced Trajectory Manipulation in Autonomous Driving
Shuo Ju, Qingzhao Zhang, Huashan Chen +6
Existing physical adversarial attacks on vision-based autonomous driving induce time-evolving perception errors, including biased object tracking or trajectory prediction, through…
Diagnosing and Resolving Android Applications Building Issues: An Empirical Study
Lakshmi Priya Bodepudi, Yutong Zhao, Ming Quan Fu +3
Building Android applications reliably remains a persistent challenge due to complex dependencies, diverse configurations, and the rapid evolution of the Android ecosystem. This st…
Text-Based Person Search with Limited Data
Xiao Han, Sen He, Li Zhang +1
Text-based person search (TBPS) aims at retrieving a target person from an image gallery with a descriptive text query. Solving such a fine-grained cross-modal retrieval task is ch…
MSSSeg: Learning Multi-Scale Structural Complexity for Self-Supervised Segmentation
Haotang Li, Zhenyu Qi, Hao Qin +4
Self-supervised semantic segmentation methods often suffer from structural errors, including merging distinct objects or fragmenting coherent regions, because they rely primarily o…
OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory
Zhaochong An, Menglin Jia, Haonan Qiu +12
Storytelling in real-world videos often unfolds through multiple shots -- discontinuous yet semantically connected clips that together convey a coherent narrative. However, existin…
Hyper-VolTran: Fast and Generalizable One-Shot Image to 3D Object Structure via HyperNetworks
Christian Simon, Sen He, Juan-Manuel Perez-Rua +3
Solving image-to-3D from a single view is an ill-posed problem, and current neural reconstruction methods addressing it through diffusion models still rely on scene-specific optimi…
Deep saliency: What is learnt by a deep network about saliency?
Sen He, Nicolas Pugeault
Deep convolutional neural networks have achieved impressive performance on a broad range of problems, beating prior art on established benchmarks, but it often remains unclear what…
UIGR: Unified Interactive Garment Retrieval
Xiao Han, Sen He, Li Zhang +2
Interactive garment retrieval (IGR) aims to retrieve a target garment image based on a reference garment image along with user feedback on what to change on the reference garment.…
Unveiling Code Clone Patterns in Open Source VR Software: An Empirical Study
Huashan Chen, Zisheng Huang, Yifan Xu +6
Code cloning is frequently observed in software development, often leading to a variety of maintenance and security issues. While substantial research has been conducted on code cl…
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Zhiheng Liu, Weiming Ren, Xiaoke Huang +12
Unified multimodal models typically rely on pretrained vision encoders and use separate visual representations for understanding and generation, creating misalignment between the t…
Harnessing Large Language Model for Virtual Reality Exploration Testing: A Case Study
Zhenyu Qi, Haotang Li, Hao Qin +3
As the Virtual Reality (VR) industry expands, the need for automated GUI testing is growing rapidly. Large Language Models (LLMs), capable of retaining information long-term and an…
Simpler is Better: Few-shot Semantic Segmentation with Classifier Weight Transformer
Zhihe Lu, Sen He, Xiatian Zhu +3
A few-shot semantic segmentation model is typically composed of a CNN encoder, a CNN decoder and a simple classifier (separating foreground and background pixels). Most existing me…
Understanding and Visualizing Deep Visual Saliency Models
Sen He, Hamed R. Tavakoli, Ali Borji +2
Recently, data-driven deep saliency models have achieved high performance and have outperformed classical saliency models, as demonstrated by results on datasets such as the MIT300…
HiStream: Efficient High-Resolution Video Generation via Redundancy-Eliminated Streaming
Haonan Qiu, Shikun Liu, Zijian Zhou +10
High-resolution video generation, while crucial for digital media and film, is computationally bottlenecked by the quadratic complexity of diffusion models, making practical infere…
Multi-LLM Orchestration for High-Quality Code Generation: Exploiting Complementary Model Strengths
Huashan Chen, Zhenyu Qi, Haotang Li +7
Large Language Models (LLMs) have become central to automated code generation, yet existing approaches operate within a single-LLM paradigm: one model is selected and applied throu…