Publications (166)
Progressively Guided Alternate Refinement Network for RGB-D Salient Object Detection
Shuhan Chen, Yun Fu
In this paper, we aim to develop an efficient and compact deep network for RGB-D salient object detection, where the depth image provides complementary information to boost perform…
Iterative Soft Shrinkage Learning for Efficient Image Super-Resolution
Jiamian Wang, Huan Wang, Yulun Zhang +2
Image super-resolution (SR) has witnessed extensive neural network designs from CNN to transformer architectures. However, prevailing SR models suffer from prohibitive memory footp…
Recent Advances on Neural Network Pruning at Initialization
Huan Wang, Can Qin, Yue Bai +2
Neural network pruning typically removes connections or neurons from a pretrained converged model; while a new pruning paradigm, pruning at initialization (PaI), attempts to prune…
D-CoDe: Scaling Image-Pretrained VLMs to Video via Dynamic Compression and Question Decomposition
Yiyang Huang, Yizhou Wang, Yun Fu
Video large language models (Vid-LLMs), which excel in diverse video-language tasks, can be effectively constructed by adapting image-pretrained vision-language models (VLMs). Howe…
NoisyCausal: A Benchmark for Evaluating Causal Reasoning Under Structured Noise
Zhi Xu, Yun Fu
Causal reasoning in natural language requires identifying relevant variables, understanding their interactions, and reasoning about effects and interventions, often under noisy or…
Joint Super-Resolution and Alignment of Tiny Faces
Yu Yin, Joseph P. Robinson, Yulun Zhang +1
Super-resolution (SR) and landmark localization of tiny faces are highly correlated tasks. On the one hand, landmark localization could obtain higher accuracy with faces of high-re…
Real-time Memory Efficient Large-pose Face Alignment via Deep Evolutionary Network
Bin Sun, Ming Shao, Siyu Xia +1
There is an urgent need to apply face alignment in a memory-efficient and real-time manner due to the recent explosion of face recognition applications. However, impact factors suc…
SuperFront: From Low-resolution to High-resolution Frontal Face Synthesis
Yu Yin, Joseph P. Robinson, Songyao Jiang +3
Advances in face rotation, along with other face-based generative tasks, are more frequent as we advance further in topics of deep learning. Even as impressive milestones are achie…
Dual Lottery Ticket Hypothesis
Yue Bai, Huan Wang, Zhiqiang Tao +2
Fully exploiting the learning capacity of neural networks requires overparameterized dense networks. On the other side, directly training sparse neural networks typically results i…
Representation Potentials of Foundation Models for Multimodal Alignment: A Survey
Jianglin Lu, Hailing Wang, Yi Xu +3
Foundation models learn highly transferable representations through large-scale pretraining on diverse data. An increasing body of research indicates that these representations exh…
RPCL: A Framework for Improving Cross-Domain Detection with Auxiliary Tasks
Kai Li, Curtis Wigington, Chris Tensmeyer +5
Cross-Domain Detection (XDD) aims to train an object detector using labeled image from a source domain but have good performance in the target domain with only unlabeled images. Ex…
Through the Theory of Mind's Eye: Reading Minds with Multimodal Video Large Language Models
Zhawnen Chen, Tianchun Wang, Yizhou Wang +4
Can large multimodal models have a human-like ability for emotional and social reasoning, and if so, how does it work? Recent research has discovered emergent theory-of-mind (ToM)…
Generative Partial Multi-View Clustering
Qianqian Wang, Zhengming Ding, Zhiqiang Tao +2
Nowadays, with the rapid development of data collection sources and feature extraction methods, multi-view data are getting easy to obtain and have received increasing research att…
Human Action Recognition and Prediction: A Survey
Yu Kong, Yun Fu
Derived from rapid advances in computer vision and machine learning, video analysis tasks have been moving from inferring the present state to predicting the future state. Vision-b…
Neural Sparse Representation for Image Restoration
Yuchen Fan, Jiahui Yu, Yiqun Mei +4
Inspired by the robustness and efficiency of sparse representation in sparse coding based image restoration models, we investigate the sparsity of neurons in deep networks. Our met…
CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning
Qihua Dong, Luis Figueroa, Handong Zhao +5
Referring Expression Comprehension and Segmentation are critical tasks for assessing the integration of language understanding and image comprehension, serving as benchmarks for Mu…
A Trajectory Generator for High-Density Traffic and Diverse Agent-Interaction Scenarios
Ruining Yang, Yi Xu, Yixiao Chen +2
Accurate trajectory prediction is fundamental to autonomous driving, as it underpins safe motion planning and collision avoidance in complex environments. However, existing benchma…
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
Haichao Zhang, Yijiang Li, Shwai He +5
Recent progress in latent world models (e.g., V-JEPA2) has shown promising capability in forecasting future world states from video observations. Nevertheless, dense prediction fro…
Parameter-Efficient Masking Networks
Yue Bai, Huan Wang, Xu Ma +3
A deeper network structure generally handles more complicated non-linearity and performs more competitively. Nowadays, advanced network designs often contain a large number of repe…
CompSRT: Quantization and Pruning for Image Super Resolution Transformers
Dorsa Zeinali, Hailing Wang, Yitian Zhang +1
Model compression has become an important tool for making image super resolution models more efficient. However, the gap between the best compressed models and the full precision m…
Test-time Fourier Style Calibration for Domain Generalization
Xingchen Zhao, Chang Liu, Anthony Sicilia +2
The topic of generalizing machine learning models learned on a collection of source domains to unknown target domains is challenging. While many domain generalization (DG) methods…
Rayleigh fading suppression in one-dimension optical scatters
Shengtao Lin, Zinan Wang, Ji Xiong +6
Highly coherent wave is favorable for applications in which phase retrieval is necessary, yet a high coherent wave is prone to encounter Rayleigh fading phenomenon as it passes thr…
Adversarial Feature Hallucination Networks for Few-Shot Learning
Kai Li, Yulun Zhang, Kunpeng Li +1
The recent flourish of deep learning in various tasks is largely accredited to the rich and accessible labeled data. Nonetheless, massive supervision remains a luxury for many real…
Trajectory Prediction Meets Large Language Models: A Survey
Yi Xu, Ruining Yang, Yitian Zhang +5
Recent advances in large language models (LLMs) have sparked growing interest in integrating language-driven techniques into trajectory prediction. By leveraging their semantic and…
Convergence Analysis and Design of Multi-block ADMM via Switched Control Theory
Jun Li, Hongfu Liu, Yue Wu +1
We consider three challenges in multi-block Alternating Direction Method of Multipliers (ADMM): building convergence conditions for ADMM with any block (variable) sequence, finding…
Balancing Biases and Preserving Privacy on Balanced Faces in the Wild
Joseph P Robinson, Can Qin, Yann Henon +2
There are demographic biases present in current facial recognition (FR) models. To measure these biases across different ethnic and gender subgroups, we introduce our Balanced Face…
Families in the Wild (FIW): Large-Scale Kinship Image Database and Benchmarks
Joseph P. Robinson, Ming Shao, Yue Wu +1
We present the largest kinship recognition dataset to date, Families in the Wild (FIW). Motivated by the lack of a single, unified dataset for kinship recognition, we aim to provid…
Self-Directed Online Machine Learning for Topology Optimization
Changyu Deng, Yizhou Wang, Can Qin +2
Topology optimization by optimally distributing materials in a given domain requires non-gradient optimizers to solve highly complicated problems. However, with hundreds of design…
SoupLM: Model Integration in Large Language and Multi-Modal Models
Yue Bai, Zichen Zhang, Jiasen Lu +1
Training large language models (LLMs) and multimodal LLMs necessitates significant computing resources, and existing publicly available LLMs are typically pre-trained on diverse, p…
Uncovering the Missing Pattern: Unified Framework Towards Trajectory Imputation and Prediction
Yi Xu, Armin Bazarjani, Hyung-gun Chi +2
Trajectory prediction is a crucial undertaking in understanding entity movement or human behavior from observed sequences. However, current methods often assume that the observed s…
Visual Semantic Reasoning for Image-Text Matching
Kunpeng Li, Yulun Zhang, Kai Li +2
Image-text matching has been a hot research topic bridging the vision and language areas. It remains challenging because the current representation of image usually lacks global se…
Zooming SlowMo: An Efficient One-Stage Framework for Space-Time Video Super-Resolution
Xiaoyu Xiang, Yapeng Tian, Yulun Zhang +3
In this paper, we address the space-time video super-resolution, which aims at generating a high-resolution (HR) slow-motion video from a low-resolution (LR) and low frame rate (LF…
LPRNet: Lightweight Deep Network by Low-rank Pointwise Residual Convolution
Bin Sun, Jun Li, Ming Shao +1
Deep learning has become popular in recent years primarily due to the powerful computing device such as GPUs. However, deploying these deep models to end-user devices, smart phones…
Out-of-Sight Embodied Agents: Multimodal Tracking, Sensor Fusion, and Trajectory Forecasting
Haichao Zhang, Yi Xu, Yun Fu
Trajectory prediction is a fundamental problem in computer vision, vision-language-action models, world models, and autonomous systems, with broad impact on autonomous driving, rob…
Don't Judge by the Look: Towards Motion Coherent Video Representation
Yitian Zhang, Yue Bai, Huan Wang +2
Current training pipelines in object recognition neglect Hue Jittering when doing data augmentation as it not only brings appearance changes that are detrimental to classification,…
GlueGen: Plug and Play Multi-modal Encoders for X-to-image Generation
Can Qin, Ning Yu, Chen Xing +6
Text-to-image (T2I) models based on diffusion processes have achieved remarkable success in controllable image generation using user-provided captions. However, the tight coupling…
Slicing Vision Transformer for Flexible Inference
Yitian Zhang, Huseyin Coskun, Xu Ma +6
Vision Transformers (ViT) is known for its scalability. In this work, we target to scale down a ViT to fit in an environment with dynamic-changing resource constraints. We observe…
Collaborative Attention Mechanism for Multi-View Action Recognition
Yue Bai, Zhiqiang Tao, Lichen Wang +3
Multi-view action recognition (MVAR) leverages complementary temporal information from different views to improve the learning performance. Obtaining informative view-specific repr…
Residual Non-local Attention Networks for Image Restoration
Yulun Zhang, Kunpeng Li, Kai Li +2
In this paper, we propose a residual non-local attention network for high-quality image restoration. Without considering the uneven distribution of information in the corrupted ima…
Image as Set of Points
Xu Ma, Yuqian Zhou, Huan Wang +4
What is an image and how to extract latent features? Convolutional Networks (ConvNets) consider an image as organized pixels in a rectangular shape and extract features via convolu…
Exploiting BERT For Multimodal Target Sentiment Classification Through Input Space Translation
Zaid Khan, Yun Fu
Multimodal target/aspect sentiment classification combines multimodal sentiment analysis and aspect/target sentiment classification. The goal of the task is to combine vision and l…
Real-Time Neural Light Field on Mobile Devices
Junli Cao, Huan Wang, Pavlo Chemerys +6
Recent efforts in Neural Rendering Fields (NeRF) have shown impressive results on novel view synthesis by utilizing implicit neural representation to represent 3D scenes. Due to th…
EmoGene: Audio-Driven Emotional 3D Talking-Head Generation
Wenqing Wang, Yun Fu
Audio-driven talking-head generation is a crucial and useful technology for virtual human interaction and film-making. While recent advances have focused on improving image fidelit…
Layout Sequence Prediction From Noisy Mobile Modality
Haichao Zhang, Yi Xu, Hongsheng Lu +2
Trajectory prediction plays a vital role in understanding pedestrian movement for applications such as autonomous driving and robotics. Current trajectory prediction models depend…
VINS: Visual Search for Mobile User Interface Design
Sara Bunian, Kai Li, Chaima Jemmali +3
Searching for relative mobile user interface (UI) design examples can aid interface designers in gaining inspiration and comparing design alternatives. However, finding such design…
Rethinking Zero-Shot Learning: A Conditional Visual Classification Perspective
Kai Li, Martin Renqiang Min, Yun Fu
Zero-shot learning (ZSL) aims to recognize instances of unseen classes solely based on the semantic descriptions of the classes. Existing algorithms usually formulate it as a seman…
Spatially Constrained GAN for Face and Fashion Synthesis
Songyao Jiang, Hongfu Liu, Yue Wu +1
Image generation has raised tremendous attention in both academic and industrial areas, especially for the conditional and target-oriented image generation, such as criminal portra…
Survey on the Analysis and Modeling of Visual Kinship: A Decade in the Making
Joseph P Robinson, Ming Shao, Yun Fu
Kinship recognition is a challenging problem with many practical applications. With much progress and milestones having been reached after ten years - we are now able to survey the…
Boosting Large Language Models with Mask Fine-Tuning
Mingyuan Zhang, Yue Bai, Huan Wang +4
The large language model (LLM) is typically integrated into the mainstream optimization protocol. No work has questioned whether maintaining the model integrity is \textit{indispen…
What Makes a "Good" Data Augmentation in Knowledge Distillation -- A Statistical Perspective
Huan Wang, Suhas Lohit, Mike Jones +1
Knowledge distillation (KD) is a general neural network training approach that uses a teacher model to guide the student model. Existing works mainly study KD from the network outp…
One Label, One Billion Faces: Usage and Consistency of Racial Categories in Computer Vision
Zaid Khan, Yun Fu
Computer vision is widely deployed, has highly visible, society altering applications, and documented problems with bias and representation. Datasets are critical for benchmarking…
Correlative Channel-Aware Fusion for Multi-View Time Series Classification
Yue Bai, Lichen Wang, Zhiqiang Tao +2
Multi-view time series classification (MVTSC) aims to improve the performance by fusing the distinctive temporal information from multiple views. Existing methods mainly focus on f…
Single-Stream Multi-Level Alignment for Vision-Language Pretraining
Zaid Khan, Vijay Kumar BG, Xiang Yu +3
Self-supervised vision-language pretraining from pure images and text with a contrastive loss is effective, but ignores fine-grained alignment due to a dual-stream architecture tha…
A Simple and Efficient Reconstruction Backbone for Snapshot Compressive Imaging
Jiamian Wang, Yulun Zhang, Xin Yuan +2
The emerging technology of snapshot compressive imaging (SCI) enables capturing high dimensional (HD) data in an efficient way. It is generally implemented by two components: an op…
Visual Font Pairing
Shuhui Jiang, Zhaowen Wang, Aaron Hertzmann +2
This paper introduces the problem of automatic font pairing. Font pairing is an important design task that is difficult for novices. Given a font selection for one part of a docume…
Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models
Xu Ma, Peize Sun, Haoyu Ma +22
Autoregressive (AR) models, long dominant in language generation, are increasingly applied to image synthesis but are often considered less competitive than Diffusion-based models.…
Laplace Landmark Localization
Joseph P Robinson, Yuncheng Li, Ning Zhang +2
Landmark localization in images and videos is a classic problem solved in various ways. Nowadays, with deep networks prevailing throughout machine learning, there are revamped inte…
Arbitrary-Scale 3D Gaussian Super-Resolution
Huimin Zeng, Yue Bai, Yun Fu
Existing 3D Gaussian Splatting (3DGS) super-resolution methods typically perform high-resolution (HR) rendering of fixed scale factors, making them impractical for resource-limited…
Accessing Vision Foundation Models via ImageNet-1K
Yitian Zhang, Xu Ma, Yue Bai +2
Vision foundation models are renowned for the generalization ability due to massive training data. Nevertheless, they demand tremendous training resources, and the training data is…
GmNet: Revisiting Gating Mechanisms From A Frequency View
Yifan Wang, Xu Ma, Yitian Zhang +5
Gating mechanisms have emerged as an effective strategy integrated into model designs beyond recurrent neural networks for addressing long-range dependency problems. In a broad und…
Adapting to Length Shift: FlexiLength Network for Trajectory Prediction
Yi Xu, Yun Fu
Trajectory prediction plays an important role in various applications, including autonomous driving, robotics, and scene understanding. Existing approaches mainly focus on developi…
Large Scale Incremental Learning
Yue Wu, Yinpeng Chen, Lijuan Wang +4
Modern machine learning suffers from catastrophic forgetting when learning new classes incrementally. The performance dramatically degrades due to the missing data of old classes.…
Image Super-Resolution Using Very Deep Residual Channel Attention Networks
Yulun Zhang, Kunpeng Li, Kai Li +3
Convolutional neural network (CNN) depth is of crucial importance for image super-resolution (SR). However, we observe that deeper networks for image SR are more difficult to train…
What Will Your Child Look Like? DNA-Net: Age and Gender Aware Kin Face Synthesizer
Pengyu Gao, Siyu Xia, Joseph Robinson +4
Visual kinship recognition aims to identify blood relatives from facial images. Its practical application-- like in law-enforcement, video surveillance, automatic family album mana…
Why is the State of Neural Network Pruning so Confusing? On the Fairness, Comparison Setup, and Trainability in Network Pruning
Huan Wang, Can Qin, Yue Bai +1
The state of neural network pruning has been noticed to be unclear and even confusing for a while, largely due to "a lack of standardized benchmarks and metrics" [3]. To standardiz…
A Close Look at Spatial Modeling: From Attention to Convolution
Xu Ma, Huan Wang, Can Qin +4
Vision Transformers have shown great promise recently for many vision tasks due to the insightful architecture design and attention mechanism. By revisiting the self-attention resp…
Trainability Preserving Neural Pruning
Huan Wang, Yun Fu
Many recent works have shown trainability plays a central role in neural network pruning -- unattended broken trainability can lead to severe under-performance and unintentionally…
MTS-DMAE: Dual-Masked Autoencoder for Unsupervised Multivariate Time Series Representation Learning
Yi Xu, Yitian Zhang, Yun Fu
Unsupervised multivariate time series (MTS) representation learning aims to extract compact and informative representations from raw sequences without relying on labels, enabling e…
BEV-DG: Cross-Modal Learning under Bird's-Eye View for Domain Generalization of 3D Semantic Segmentation
Miaoyu Li, Yachao Zhang, Xu MA +2
Cross-modal Unsupervised Domain Adaptation (UDA) aims to exploit the complementarity of 2D-3D data to overcome the lack of annotation in a new domain. However, UDA methods rely on…
TDAN: Temporally Deformable Alignment Network for Video Super-Resolution
Yapeng Tian, Yulun Zhang, Yun Fu +1
Video super-resolution (VSR) aims to restore a photo-realistic high-resolution (HR) video frame from both its corresponding low-resolution (LR) frame (reference frame) and multiple…
PointDAN: A Multi-Scale 3D Domain Adaption Network for Point Cloud Representation
Can Qin, Haoxuan You, Lichen Wang +2
Domain Adaptation (DA) approaches achieved significant improvements in a wide range of machine learning and computer vision tasks (i.e., classification, detection, and segmentation…
Q: How to Specialize Large Vision-Language Models to Data-Scarce VQA Tasks? A: Self-Train on Unlabeled Images!
Zaid Khan, Vijay Kumar BG, Samuel Schulter +3
Finetuning a large vision language model (VLM) on a target dataset after large scale pretraining is a dominant paradigm in visual question answering (VQA). Datasets for specialized…
RealTalk: Realistic Emotion-Aware Lifelike Talking-Head Synthesis
Wenqing Wang, Yun Fu
Emotion is a critical component of artificial social intelligence. However, while current methods excel in lip synchronization and image quality, they often fail to generate accura…
Unveiling the Unseen: A Comprehensive Survey on Explainable Anomaly Detection in Images and Videos
Yizhou Wang, Dongliang Guo, Sheng Li +2
Anomaly detection and localization in visual data, including images and videos, are crucial in machine learning and real-world applications. Despite rapid advancements in visual an…
EV-Action: Electromyography-Vision Multi-Modal Action Dataset
Lichen Wang, Bin Sun, Joseph Robinson +2
Multi-modal human action analysis is a critical and attractive research topic. However, the majority of the existing datasets only provide visual modalities (i.e., RGB, depth and s…
Den-TP: A Density-Balanced Data Curation and Evaluation Framework for Trajectory Prediction
Ruining Yang, Yi Xu, Yun Fu +1
Trajectory prediction in autonomous driving has traditionally been studied from a model-centric perspective. However, existing datasets exhibit a strong long-tail distribution in s…
Generative One-Shot Face Recognition
Zhengming Ding, Yandong Guo, Lei Zhang +1
One-shot face recognition measures the ability to identify persons with only seeing them at one glance, and is a hallmark of human visual intelligence. It is challenging for conven…
Demystifying When Pruning Works via Representation Hierarchies
Shwai He, Guoheng Sun, Haichao Zhang +2
Network pruning, which removes less important parameters or architectures, is often expected to improve efficiency while preserving performance. However, this expectation does not…
Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks
Qihua Dong, Kuo Yang, Lin Ju +6
Referring Expression Comprehension (REC) links language to region level visual perception. Standard benchmarks (RefCOCO, RefCOCO+, RefCOCOg) have progressed rapidly with multimodal…
Towards Layer-wise Image Vectorization
Xu Ma, Yuqian Zhou, Xingqian Xu +5
Image rasterization is a mature technique in computer graphics, while image vectorization, the reverse path of rasterization, remains a major challenge. Recent advanced deep learni…
Restore-R1: Efficient Image Restoration Agents via Reinforcement Learning with Multimodal LLM Perceptual Feedback
Jianglin Lu, Yuanwei Wu, Ziyi Zhao +4
Complex image restoration aims to recover high-quality images from inputs affected by multiple degradations such as blur, noise, rain, and compression artifacts. Recent restoration…
Fine-T2I: An Open, Large-Scale, and Diverse Dataset for High-Quality T2I Fine-Tuning
Xu Ma, Yitian Zhang, Qihua Dong +1
High-quality and open datasets remain a major bottleneck for text-to-image (T2I) fine-tuning. Despite rapid progress in model architectures and training pipelines, most publicly av…
R2L: Distilling Neural Radiance Field to Neural Light Field for Efficient Novel View Synthesis
Huan Wang, Jian Ren, Zeng Huang +4
Recent research explosion on Neural Radiance Field (NeRF) shows the encouraging potential to represent complex scenes with neural networks. One major drawback of NeRF is its prohib…
Tell Me Where to Look: Guided Attention Inference Network
Kunpeng Li, Ziyan Wu, Kuan-Chuan Peng +2
Weakly supervised learning with only coarse labels can obtain visual explanations of deep neural network such as attention maps by back-propagating gradients. These attention maps…
Exploring Question Decomposition for Zero-Shot VQA
Zaid Khan, Vijay Kumar BG, Samuel Schulter +2
Visual question answering (VQA) has traditionally been treated as a single-step task where each question receives the same amount of effort, unlike natural human question-answering…
Contrastive Alignment of Vision to Language Through Parameter-Efficient Transfer Learning
Zaid Khan, Yun Fu
Contrastive vision-language models (e.g. CLIP) are typically created by updating all the parameters of a vision model and language model through contrastive training. Can such mode…
Rethinking Adam: A Twofold Exponential Moving Average Approach
Yizhou Wang, Yue Kang, Can Qin +4
Adaptive gradient methods, e.g. \textsc{Adam}, have achieved tremendous success in machine learning. Scaling the learning rate element-wisely by a certain form of second moment est…
The 5th Recognizing Families in the Wild Data Challenge: Predicting Kinship from Faces
Joseph P. Robinson, Can Qin, Ming Shao +3
Recognizing Families In the Wild (RFIW), held as a data challenge in conjunction with the 16th IEEE International Conference on Automatic Face and Gesture Recognition (FG), is a la…
HyperSTAR: Task-Aware Hyperparameters for Deep Networks
Gaurav Mittal, Chang Liu, Nikolaos Karianakis +3
While deep neural networks excel in solving visual recognition tasks, they require significant effort to find hyperparameters that make them work optimally. Hyperparameter Optimiza…
LightAvatar: Efficient Head Avatar as Dynamic Neural Light Field
Huan Wang, Feitong Tan, Ziqian Bai +9
Recent works have shown that neural radiance fields (NeRFs) on top of parametric models have reached SOTA quality to build photorealistic head avatars from a monocular video. Howev…
ExpertGen: Training-Free Expert Guidance for Controllable Text-to-Face Generation
Liang Shi, Yun Fu
Recent advances in diffusion models have significantly improved text-to-face generation, but achieving fine-grained control over facial features remains a challenge. Existing metho…
Recognizing Families In the Wild: White Paper for the 4th Edition Data Challenge
Joseph P. Robinson, Yu Yin, Zaid Khan +7
Recognizing Families In the Wild (RFIW): an annual large-scale, multi-track automatic kinship recognition evaluation that supports various visual kin-based problems on scales much…
Efficient Modulation for Vision Networks
Xu Ma, Xiyang Dai, Jianwei Yang +4
In this work, we present efficient modulation, a novel design for efficient vision networks. We revisit the modulation mechanism, which operates input through convolutional context…
Face Recognition: Too Bias, or Not Too Bias?
Joseph P Robinson, Gennady Livitz, Yann Henon +3
We reveal critical insights into problems of bias in state-of-the-art facial recognition (FR) systems using a novel Balanced Faces In the Wild (BFW) dataset: data balanced for gend…
Dynamical Isometry: The Missing Ingredient for Neural Network Pruning
Huan Wang, Can Qin, Yue Bai +1
Several recent works [40, 24] observed an interesting phenomenon in neural network pruning: A larger finetuning learning rate can improve the final performance significantly. Unfor…
Making Reconstruction-based Method Great Again for Video Anomaly Detection
Yizhou Wang, Can Qin, Yue Bai +3
Anomaly detection in videos is a significant yet challenging problem. Previous approaches based on deep neural networks employ either reconstruction-based or prediction-based appro…
Rethinking Fine-Tuning: Unlocking Hidden Capabilities in Vision-Language Models
Mingyuan Zhang, Yue Bai, Yifan Wang +2
Explorations in fine-tuning Vision-Language Models (VLMs), such as Low-Rank Adaptation (LoRA) from Parameter Efficient Fine-Tuning (PEFT), have made impressive progress. However, m…
Contradictory Structure Learning for Semi-supervised Domain Adaptation
Can Qin, Lichen Wang, Qianqian Ma +3
Current adversarial adaptation methods attempt to align the cross-domain features, whereas two challenges remain unsolved: 1) the conditional distribution mismatch and 2) the bias…
Post-Training in End-to-End Autonomous Driving
Ruining Yang, Muxing Wang, Yixiao Chen +8
This survey reviews post‑training methods that refine end‑to‑end autonomous driving models beyond imitation, organizing existing work into four families based on the type of superv…
Frame Flexible Network
Yitian Zhang, Yue Bai, Chang Liu +3
Existing video recognition algorithms always conduct different training pipelines for inputs with different frame numbers, which requires repetitive training operations and multipl…