Publications (28)
Towards Long-Form Spatio-Temporal Video Grounding
Xin Gu, Bing Fan, Jiali Yao +5
In real scenarios, videos can span several minutes or even hours. However, existing research on spatio-temporal video grounding (STVG), given a textual query, mainly focuses on loc…
All You Need is One: Capsule Prompt Tuning with a Single Vector
Yiyang Liu, James C. Liang, Heng Fan +7
Prompt-based learning has emerged as a parameter-efficient finetuning (PEFT) approach to facilitate Large Language Model (LLM) adaptation to downstream tasks by conditioning genera…
SSGA-Net: Stepwise Spatial Global-local Aggregation Networks for for Autonomous Driving
Yiming Cui, Cheng Han, Dongfang Liu
Visual-based perception is the key module for autonomous driving. Among those visual perception tasks, video object detection is a primary yet challenging one because of feature de…
CML-MOTS: Collaborative Multi-task Learning for Multi-Object Tracking and Segmentation
Yiming Cui, Cheng Han, Dongfang Liu
The advancement of computer vision has pushed visual analysis tasks from still images to the video domain. In recent years, video instance segmentation, which aims to track and seg…
Exploring the Adversarial Vulnerabilities of Vision-Language-Action Models in Robotics
Taowen Wang, Cheng Han, James Chenhao Liang +6
Recently in robotics, Vision-Language-Action (VLA) models have emerged as a transformative approach, enabling robots to execute complex tasks by integrating visual and linguistic i…
Facing the Elephant in the Room: Visual Prompt Tuning or Full Finetuning?
Cheng Han, Qifan Wang, Yiming Cui +4
As the scale of vision models continues to grow, the emergence of Visual Prompt Tuning (VPT) as a parameter-efficient transfer learning technique has gained attention due to its su…
Adapting Vision Foundation Models with Cascaded Semantics
Xi Xiao, Xingjian Li, Cheng Han +8
Prompt tuning, a leading parameter-efficient adaptation paradigm in NLP, has recently been extended to computer vision. Visual prompt tuning (VPT) adapts pre-trained vision transfo…
Strongly Modulated Ambipolar Characteristics of Few-layer Black Phosphorus in Oxygen
Cheng Han, Zehua Hu, Jialin Zhang +7
Two-dimensional black phosphorus has been configured as field-effect transistors, showing an intrinsic symmetric ambipolar transport characteristic. Here, we demonstrate the strong…
On-the-Fly VLA Adaptation via Test-Time Reinforcement Learning
Changyu Liu, Yiyang Liu, Taowen Wang +7
Vision-Language-Action models have recently emerged as a powerful paradigm for general-purpose robot learning, enabling agents to map visual observations and natural-language instr…
Visual Recognition with Deep Nearest Centroids
Wenguan Wang, Cheng Han, Tianfei Zhou +1
We devise deep nearest centroids (DNC), a conceptually elegant yet surprisingly effective network for large-scale visual recognition, by revisiting Nearest Centroids, one of the mo…
YOLOPv2: Better, Faster, Stronger for Panoptic Driving Perception
Cheng Han, Qichao Zhao, Shuyi Zhang +3
Over the last decade, multi-tasking learning approaches have achieved promising results in solving panoptic driving perception problems, providing both high-precision and high-effi…
AMD: Automatic Multi-step Distillation of Large-scale Vision Models
Cheng Han, Qifan Wang, Sohail A. Dianat +6
Transformer-based architectures have become the de-facto standard models for diverse vision tasks owing to their superior performance. As the size of the models continues to scale…
TokenSeek: Memory Efficient Fine Tuning via Instance-Aware Token Ditching
Runjia Zeng, Qifan Wang, Qiang Guan +6
Fine tuning has been regarded as a de facto approach for adapting large language models (LLMs) to downstream tasks, but the high training memory consumption inherited from LLMs mak…
MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper
Runjia Zeng, Guangyan Sun, Qifan Wang +8
Considering deep neural networks as manifold mappers, the pretrain-then-fine-tune paradigm can be interpreted as a two-stage process: pretrain establishes a broad knowledge base, a…
Image Translation as Diffusion Visual Programmers
Cheng Han, James C. Liang, Qifan Wang +5
We introduce the novel Diffusion Visual Programmer (DVP), a neuro-symbolic image translation framework. Our proposed DVP seamlessly embeds a condition-flexible diffusion model with…
Visual Fourier Prompt Tuning
Runjia Zeng, Cheng Han, Qifan Wang +5
With the scale of vision Transformer-based models continuing to grow, finetuning these large-scale pretrained models for new tasks has become increasingly parameter-intensive. Visu…
Prototypical Transformer as Unified Motion Learners
Cheng Han, Yawen Lu, Guohao Sun +9
In this work, we introduce the Prototypical Transformer (ProtoFormer), a general and unified framework that approaches various motion tasks from a prototype perspective. ProtoForme…
ProMotion: Prototypes As Motion Learners
Yawen Lu, Dongfang Liu, Qifan Wang +6
In this work, we introduce ProMotion, a unified prototypical framework engineered to model fundamental motion tasks. ProMotion offers a range of compelling attributes that set it a…
Probabilistic Token Alignment for Large Language Model Fusion
Runjia Zeng, James Chenhao Liang, Cheng Han +8
Training large language models (LLMs) from scratch can yield models with unique functionalities and strengths, but it is costly and often leads to redundant capabilities. A more co…
Prompt-based Adaptation in Large-scale Vision Models: A Survey
Xi Xiao, Yunbei Zhang, Lin Zhao +12
In computer vision, Visual Prompting (VP) and Visual Prompt Tuning (VPT) have recently emerged as lightweight and effective alternatives to full fine-tuning for adapting large-scal…
E^2VPT: An Effective and Efficient Approach for Visual Prompt Tuning
Cheng Han, Qifan Wang, Yiming Cui +4
As the size of transformer-based models continues to grow, fine-tuning these large-scale pretrained vision models for new tasks has become increasingly parameter-intensive. Paramet…
Self-supervised Adversarial Training of Monocular Depth Estimation against Physical-World Attacks
Zhiyuan Cheng, Cheng Han, James Liang +3
Monocular Depth Estimation (MDE) plays a vital role in applications such as autonomous driving. However, various attacks target MDE models, with physical attacks posing significant…
Resolving the Ambiguity of Complete-to-Partial Point Cloud Registration for Image-Guided Liver Surgery with Patches-to-Partial Matching
Zixin Yang, Jon S. Heiselman, Cheng Han +3
In image-guided liver surgery, the initial rigid alignment between preoperative and intraoperative data, often represented as point clouds, is crucial for providing sub-surface inf…
Re-Imagining Multimodal Instruction Tuning: A Representation View
Yiyang Liu, James Chenhao Liang, Ruixiang Tang +8
Multimodal instruction tuning has proven to be an effective strategy for achieving zero-shot generalization by fine-tuning pre-trained Large Multimodal Models (LMMs) with instructi…
MPT: Multimodal Prompt Tuning for Zero-shot Instruction Learning
Taowen Wang, Yiyang Liu, James Chenhao Liang +11
Multimodal Large Language Models (MLLMs) demonstrate remarkable performance across a wide range of domains, with increasing emphasis on enhancing their zero-shot generalization cap…
Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards
Xi Li, Shu Zhao, Xiaohan Zou +6
Multi-modal large language models (MLLMs) integrate heterogeneous modalities through modality alignment and fusion, enabling stronger understanding and reasoning. However, this arc…
Layer-Specific Prompt Fusion Discovery via Differentiable Search in Vision Foundation Models
Xi Xiao, Xingjian Li, Yunbei Zhang +7
Visual prompt tuning has emerged as a parameter-efficient fine-tuning approach for adapting large-scale Vision Transformers (ViTs) to downstream tasks. As its learnable prompts are…
A-SelecT: Automatic Timestep Selection for Diffusion Transformer Representation Learning
Changyu Liu, James Chenhao Liang, Wenhao Yang +6
Diffusion models have significantly reshaped the field of generative artificial intelligence and are now increasingly explored for their capacity in discriminative representation l…