Publications (105)
GRACE: Boosting Video MLLMs with Grounded Action-Centric Evidence for Viewer Sentiment Prediction
Ruoxuan Yang, Tieyuan Chen, Xiaofeng Huang +6
Viewer sentiment prediction in video advertisements aims to infer the latent affective response evoked in the audience. To bridge the gap between what is shown and what is felt, mo…
Deep Neural Network Compression with Single and Multiple Level Quantization
Yuhui Xu, Yongzhuang Wang, Aojun Zhou +2
Network quantization is an effective solution to compress deep neural networks for practical usage. Existing network quantization methods cannot sufficiently exploit the depth info…
Facial Expression Cloning with Elastic and Muscle Models
Yihao Zhang, Weiyao Lin, Bing Zhou +4
Expression cloning plays an important role in facial expression synthesis. In this paper, a novel algorithm is proposed for facial expression cloning. The proposed algorithm first…
A Fast Sub-Pixel Motion Estimation Algorithm for H.264/AVC Video Coding
Weiyao Lin, Krit Panusopone, David M. Baylon +3
Motion Estimation (ME) is one of the most time-consuming parts in video coding. The use of multiple partition sizes in H.264/AVC makes it even more complicated when compared to ME…
Class-aware Sounding Objects Localization via Audiovisual Correspondence
Di Hu, Yake Wei, Rui Qian +3
Audiovisual scenes are pervasive in our daily life. It is commonplace for humans to discriminatively localize different sounding objects but quite challenging for machines to achie…
Low-Rank Winograd Transformation for 3D Convolutional Neural Networks
Ziran Qin, Mingbao Lin, Weiyao Lin
This paper focuses on Winograd transformation in 3D convolutional neural networks (CNNs) that are more over-parameterized compared with the 2D version. The over-increasing Winograd…
DND: Boosting Large Language Models with Dynamic Nested Depth
Tieyuan Chen, Xiaodong Chen, Haoxing Chen +3
We introduce Dynamic Nested Depth (DND), a novel method that improves performance for off-the-shelf LLMs by selecting critical tokens to reprocess in a nested depth manner. Specifi…
ABC: Adaptive BayesNet Structure Learning for Computational Scalable Multi-task Image Compression
Yufeng Zhang, Wenrui Dai, Hang Yu +4
Neural Image Compression (NIC) has revolutionized image compression with its superior rate-distortion performance and multi-task capabilities, supporting both human visual percepti…
Controllable Augmentations for Video Representation Learning
Rui Qian, Weiyao Lin, John See +1
This paper focuses on self-supervised video representation learning. Most existing approaches follow the contrastive learning pipeline to construct positive and negative pairs by s…
Seeing Through Pixel Motion: Learning Obstacle Avoidance from Optical Flow with One Camera
Yu Hu, Yuang Zhang, Yunlong Song +6
Optical flow captures the motion of pixels in an image sequence over time, providing information about movement, depth, and environmental structure. Flying insects utilize this inf…
A Tube-and-Droplet-based Approach for Representing and Analyzing Motion Trajectories
Weiyao Lin, Yang Zhou, Hongteng Xu +4
Trajectory analysis is essential in many applications. In this paper, we address the problem of representing motion trajectories in a highly informative way, and consequently utili…
A Computation Control Motion Estimation Method for Complexity-Scalable Video Coding
Weiyao Lin, Krit Panusopone, David M. Baylon +1
In this paper, a new Computation-Control Motion Estimation (CCME) method is proposed which can perform Motion Estimation (ME) adaptively under different computation or power budget…
Head-Aware KV Cache Compression for Efficient Visual Autoregressive Modeling
Ziran Qin, Youru Lv, Mingbao Lin +4
Visual Autoregressive (VAR) models adopt a next-scale prediction paradigm, offering high-quality content generation with substantially fewer decoding steps. However, existing VAR m…
Uni-EDEN: Universal Encoder-Decoder Network by Multi-Granular Vision-Language Pre-training
Yehao Li, Jiahao Fan, Yingwei Pan +3
Vision-language pre-training has been an emerging and fast-developing research topic, which transfers multi-modal knowledge from rich-resource pre-training task to limited-resource…
Enhancing Self-supervised Video Representation Learning via Multi-level Feature Optimization
Rui Qian, Yuxi Li, Huabin Liu +5
The crux of self-supervised video representation learning is to build general features from unlabeled videos. However, most recent works have mainly focused on high-level semantics…
A Heat-Map-based Algorithm for Recognizing Group Activities in Videos
Weiyao Lin, Hang Chu, Jianxin Wu +2
In this paper, a new heat-map-based (HMB) algorithm is proposed for group activity recognition. The proposed algorithm first models human trajectories as series of "heat sources" a…
FRIH: Fine-grained Region-aware Image Harmonization
Jinlong Peng, Zekun Luo, Liang Liu +6
Image harmonization aims to generate a more realistic appearance of foreground and background for a composite image. Existing methods perform the same harmonization process for the…
PCGS: Progressive Compression of 3D Gaussian Splatting
Yihang Chen, Mengyao Li, Qianyi Wu +3
3D Gaussian Splatting (3DGS) achieves impressive rendering fidelity and speed for novel view synthesis. However, its substantial data size poses a significant challenge for practic…
BasisFormer: Attention-based Time Series Forecasting with Learnable and Interpretable Basis
Zelin Ni, Hang Yu, Shizhan Liu +2
Bases have become an integral part of modern deep learning-based models for time series forecasting due to their ability to act as feature extractors or future references. To be ef…
Key-Point Sequence Lossless Compression for Intelligent Video Analysis
Weiyao Lin, Xiaoyi He, Wenrui Dai +4
Feature coding has been recently considered to facilitate intelligent video analysis for urban computing. Instead of raw videos, extracted features in the front-end are encoded and…
Massive Activations are the Key to Local Detail Synthesis in Diffusion Transformers
Chaofan Gan, Zicheng Zhao, Yuanpeng Tu +5
Diffusion Transformers (DiTs) have recently emerged as a powerful backbone for visual generation. Recent observations reveal \emph{Massive Activations} (MAs) in their internal feat…
DAC: 2D-3D Retrieval with Noisy Labels via Divide-and-Conquer Alignment and Correction
Chaofan Gan, Yuanpeng Tu, Yuxi Li +1
With the recent burst of 2D and 3D data, cross-modal retrieval has attracted increasing attention recently. However, manual labeling by non-experts will inevitably introduce corrup…
Fractal Dimension Invariant Filtering and Its CNN-based Implementation
Hongteng Xu, Junchi Yan, Nils Persson +2
Fractal analysis has been widely used in computer vision, especially in texture image processing and texture analysis. The key concept of fractal-based image model is the fractal d…
Ensemble of Part Detectors for Simultaneous Classification and Localization
Xiaopeng Zhang, Hongkai Xiong, Weiyao Lin +1
Part-based representation has been proven to be effective for a variety of visual applications. However, automatic discovery of discriminative parts without object/part-level annot…
Improved Image Deblurring based on Salient-region Segmentation
Chongyang Zhang, Weiyao Lin, Wei Li +3
Image deblurring techniques play important roles in many image processing applications. As the blur varies spatially across the image plane, it calls for robust and effective metho…
HAC: Hash-grid Assisted Context for 3D Gaussian Splatting Compression
Yihang Chen, Qianyi Wu, Weiyao Lin +2
3D Gaussian Splatting (3DGS) has emerged as a promising framework for novel view synthesis, boasting rapid rendering speed with high fidelity. However, the substantial Gaussians an…
CFAD: Coarse-to-Fine Action Detector for Spatiotemporal Action Localization
Yuxi Li, Weiyao Lin, John See +4
Most current pipelines for spatio-temporal action localization connect frame-wise or clip-wise detection results to generate action proposals, where only local information is explo…
DNQ: Dynamic Network Quantization
Yuhui Xu, Shuai Zhang, Yingyong Qi +3
Network quantization is an effective method for the deployment of neural networks on memory and energy constrained mobile devices. In this paper, we propose a Dynamic Network Quant…
NeCGS: Neural Compression for 3D Geometry Sets
Siyu Ren, Junhui Hou, Weiyao Lin +1
We present NeCGS, the first neural compression paradigm, which can compress a geometry set encompassing thousands of detailed and diverse 3D mesh models by up to 900 times with hig…
Collaborative Weakly Supervised Video Correlation Learning for Procedure-Aware Instructional Video Analysis
Tianyao He, Huabin Liu, Yuxi Li +4
Video Correlation Learning (VCL), which aims to analyze the relationships between videos, has been widely studied and applied in various general video tasks. However, applying VCL…
Spatio-Temporal Point Process for Multiple Object Tracking
Tao Wang, Kean Chen, Weiyao Lin +4
Multiple Object Tracking (MOT) focuses on modeling the relationship of detected objects among consecutive frames and merge them into different trajectories. MOT remains a challengi…
Macroblock Classification Method for Video Applications Involving Motions
Weiyao Lin, Ming-Ting Sun, Hongxiang Li +3
In this paper, a macroblock classification method is proposed for various video processing applications involving motions. Based on the analysis of the Motion Vector field in the c…
Rodimus*: Breaking the Accuracy-Efficiency Trade-Off with Efficient Attentions
Zhihao He, Hang Yu, Zi Gong +3
Recent advancements in Transformer-based large language models (LLMs) have set new standards in natural language processing. However, the classical softmax attention incurs signifi…
Task-adaptive Spatial-Temporal Video Sampler for Few-shot Action Recognition
Huabin Liu, Weixian Lv, John See +1
A primary challenge faced in few-shot action recognition is inadequate video data for training. To address this issue, current methods in this field mainly focus on devising algori…
Variational Pedestrian Detection
Yuang Zhang, Huanyu He, Jianguo Li +3
Pedestrian detection in a crowd is a challenging task due to a high number of mutually-occluding human instances, which brings ambiguity and optimization difficulties to the curren…
CSTA: Spatial-Temporal Causal Adaptive Learning for Exemplar-Free Video Class-Incremental Learning
Tieyuan Chen, Huabin Liu, Chern Hong Lim +4
Continual learning aims to acquire new knowledge while retaining past information. Class-incremental learning (CIL) presents a challenging scenario where classes are introduced seq…
MCA: 2D-3D Retrieval with Noisy Labels via Multi-level Adaptive Correction and Alignment
Gui Zou, Chaofan Gan, Chern Hong Lim +2
With the increasing availability of 2D and 3D data, significant advancements have been made in the field of cross-modal retrieval. Nevertheless, the existence of imperfect annotati…
CAKE: Cascading and Adaptive KV Cache Eviction with Layer Preferences
Ziran Qin, Yuchen Cao, Mingbao Lin +5
Large language models (LLMs) excel at processing long sequences, boosting demand for key-value (KV) caching. While recent efforts to evict KV cache have alleviated the inference bu…
Fast Feedforward 3D Gaussian Splatting Compression
Yihang Chen, Qianyi Wu, Mengyao Li +3
With 3D Gaussian Splatting (3DGS) advancing real-time and high-fidelity rendering for novel view synthesis, storage requirements pose challenges for their widespread adoption. Alth…
Partition-Aware Adaptive Switching Neural Networks for Post-Processing in HEVC
Weiyao Lin, Xiaoyi He, Xintong Han +5
This paper addresses neural network based post-processing for the state-of-the-art video coding standard, High Efficiency Video Coding (HEVC). We first propose a partition-aware Co…
Finding Action Tubes with a Sparse-to-Dense Framework
Yuxi Li, Weiyao Lin, Tao Wang +5
The task of spatial-temporal action detection has attracted increasing attention among researchers. Existing dominant methods solve this problem by relying on short-term informatio…
Motion Segmentation via Global and Local Sparse Subspace Optimization
Michael Ying Yang, Hanno Ackermann, Weiyao Lin +2
In this paper, we propose a new framework for segmenting feature-based moving objects under affine subspace model. Since the feature trajectories in practice are high-dimensional a…
Region-Based Rate-Control for H.264/AVC for Low Bit-Rate Applications
Hai-Miao Hu, Bo Li, Weiyao Lin +2
Rate-control plays an important role in video coding. However, in the conventional rate-control algorithms, the number and position of Macroblocks (MBs) inside one basic unit for r…
A new network-based algorithm for human activity recognition in video
Weiyao Lin, Yuanzhe Chen, Jianxin Wu +3
In this paper, a new network-transmission-based (NTB) algorithm is proposed for human activity recognition in videos. The proposed NTB algorithm models the entire scene as an error…
Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning
Shijie Li, Yilin Gao, Siyuan Yang +7
Multimodal Large Language Models (MLLMs) are often constrained by a language-space bottleneck, forcing complex visual reasoning into discrete tokens which can lose perceptual nuanc…
MECD+: Unlocking Event-Level Causal Graph Discovery for Video Reasoning
Tieyuan Chen, Huabin Liu, Yi Wang +5
Video causal reasoning aims to achieve a high-level understanding of videos from a causal perspective. However, it exhibits limitations in its scope, primarily executed in a questi…
From Priors to Perception: Grounding Video-LLMs in Physical Reality
Zicheng Zhao, Chaofan Gan, Shijie Li +1
While Video Large Language Models (Video-LLMs) excel in general understanding, they exhibit systematic deficits in fine-grained physical reasoning. Existing interventions not only…
CogStream: Context-guided Streaming Video Question Answering
Zicheng Zhao, Kangyu Wang, Shijie Li +3
Despite advancements in Video Large Language Models (Vid-LLMs) improving multimodal understanding, challenges persist in streaming video reasoning due to its reliance on contextual…
Unsupervised Deep Domain Adaptation for Pedestrian Detection
Lihang Liu, Weiyao Lin, Lisheng Wu +2
This paper addresses the problem of unsupervised domain adaptation on the task of pedestrian detection in crowded scenes. First, we utilize an iterative algorithm to iteratively se…
Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive Activations
Chaofan Gan, Yuanpeng Tu, Xi Chen +4
Pre-trained stable diffusion models (SD) have shown great advances in visual correspondence. In this paper, we investigate the capabilities of Diffusion Transformers (DiTs) for acc…
Few-shot Action Recognition via Intra- and Inter-Video Information Maximization
Huabin Liu, Weiyao Lin, Tieyuan Chen +3
Current few-shot action recognition involves two primary sources of information for classification:(1) intra-video information, determined by frame content within a single video cl…
Tiny-DSOD: Lightweight Object Detection for Resource-Restricted Usages
Yuxi Li, Jiuwei Li, Weiyao Lin +1
Object detection has made great progress in the past few years along with the development of deep learning. However, most current object detection methods are resource hungry, whic…
Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens
Ziran Qin, Youru Lv, Mingbao Lin +4
Autoregressive (AR) visual generation has emerged as a powerful paradigm for image and multimodal synthesis, owing to its scalability and generality. However, existing AR image gen…
The 1st-place Solution for ECCV 2022 Multiple People Tracking in Group Dance Challenge
Yuang Zhang, Tiancai Wang, Weiyao Lin +1
We present our 1st place solution to the Group Dance Multiple People Tracking Challenge. Based on MOTR: End-to-End Multiple-Object Tracking with Transformer, we explore: 1) detect…
Activity Recognition Using A Combination of Category Components And Local Models for Video Surveillance
Weiyao Lin, Ming-Ting Sun, Radha Poovendran +1
This paper presents a novel approach for automatic recognition of human activities for video surveillance applications. We propose to represent an activity by a combination of cate…
Delving into the Cyclic Mechanism in Semi-supervised Video Object Segmentation
Yuxi Li, Ning Xu, Jinlong Peng +2
In this paper, we address several inadequacies of current video object segmentation pipelines. Firstly, a cyclic mechanism is incorporated to the standard semi-supervised process t…
TA2N: Two-Stage Action Alignment Network for Few-shot Action Recognition
Shuyuan Li, Huabin Liu, Rui Qian +5
Few-shot action recognition aims to recognize novel action classes (query) using just a few samples (support). The majority of current approaches follow the metric learning paradig…
Back to Newton's Laws: Learning Vision-based Agile Flight via Differentiable Physics
Yuang Zhang, Yu Hu, Yunlong Song +2
Swarm navigation in cluttered environments is a grand challenge in robotics. This work combines deep learning with first-principle physics through differentiable simulation to enab…
HAC++: Towards 100X Compression of 3D Gaussian Splatting
Yihang Chen, Qianyi Wu, Weiyao Lin +2
3D Gaussian Splatting (3DGS) has emerged as a promising framework for novel view synthesis, boasting rapid rendering speed with high fidelity. However, the substantial Gaussians an…
Network Decoupling: From Regular to Depthwise Separable Convolutions
Jianbo Guo, Yuxi Li, Weiyao Lin +2
Depthwise separable convolution has shown great efficiency in network design, but requires time-consuming training procedure with full training-set available. This paper first anal…
ATRW: A Benchmark for Amur Tiger Re-identification in the Wild
Shuyuan Li, Jianguo Li, Hanlin Tang +2
Monitoring the population and movements of endangered species is an important task to wildlife conversation. Traditional tagging methods do not scale to large populations, while ap…
Multiple Sound Sources Localization from Coarse to Fine
Rui Qian, Di Hu, Heinrich Dinkel +3
How to visually localize multiple sound sources in unconstrained videos is a formidable problem, especially when lack of the pairwise sound-object annotations. To solve this proble…
LSTC: Boosting Atomic Action Detection with Long-Short-Term Context
Yuxi Li, Boshen Zhang, Jian Li +5
In this paper, we place the atomic action detection problem into a Long-Short Term Context (LSTC) to analyze how the temporal reliance among video signals affect the action detecti…
Exploring the Semi-supervised Video Object Segmentation Problem from a Cyclic Perspective
Yuxi Li, Ning Xu, Wenjie Yang +2
Modern video object segmentation (VOS) algorithms have achieved remarkably high performance in a sequential processing order, while most of currently prevailing pipelines still sho…
TRP: Trained Rank Pruning for Efficient Deep Neural Networks
Yuhui Xu, Yuxi Li, Shuai Zhang +6
To enable DNNs on edge devices like mobile phones, low-rank approximation has been widely adopted because of its solid theoretical rationale and efficient implementations. Several…
HACK++: Towards More Effective Head-Aware Key-Value Compression for Efficient Visual Autoregressive Modeling
Ziran Qin, Yuchen Jiang, Mingbao Lin +4
Visual Autoregressive (VAR) models adopt a next-scale prediction paradigm, offering high-quality generation with substantially fewer decoding steps. However, existing VAR models su…
Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation
Chaofan Gan, Zicheng Zhao, Yuanpeng Tu +6
Massive Activations (MAs) have been widely observed in Transformer-based models, yet their structure and functional roles in Diffusion Transformers (DiTs) remain insufficiently und…
Scene Graph Lossless Compression with Adaptive Prediction for Objects and Relations
Yufeng Zhang, Weiyao Lin, Wenrui Dai +2
The scene graph is a new data structure describing objects and their pairwise relationship within image scenes. As the size of scene graph in vision applications grows, how to loss…
Intra-and-Inter-Constraint-based Video Enhancement based on Piecewise Tone Mapping
Yuanzhe Chen, Weiyao Lin, Chongyang Zhang +3
Video enhancement plays an important role in various video applications. In this paper, we propose a new intra-and-inter-constraint-based video enhancement approach aiming to 1) ac…
Group Re-Identification with Multi-grained Matching and Integration
Weiyao Lin, Yuxi Li, Hao Xiao +5
The task of re-identifying groups of people underdifferent camera views is an important yet less-studied problem.Group re-identification (Re-ID) is a very challenging task sinceit…
Object-Attribute-Relation Model Driven Adaptive Hierarchical Transmission for Multimodal Semantic Communication
Chenxing Li, Yiping Duan, Han Jiao +3
Traditional video coding (VVC, HEVC) prioritizes human visual perception, transmitting substantial texture redundancy that severely hinders machine decision-making under constraine…
Enhancing HEVC Compressed Videos with a Partition-masked Convolutional Neural Network
Xiaoyi He, Qiang Hu, Xintong Han +3
In this paper, we propose a partition-masked Convolution Neural Network (CNN) to achieve compressed-video enhancement for the state-of-the-art coding standard, High Efficiency Vide…
Deep Spatial Pyramid: The Devil is Once Again in the Details
Bin-Bin Gao, Xiu-Shen Wei, Jianxin Wu +1
In this paper we show that by carefully making good choices for various detailed but important factors in a visual recognition framework using deep learning features, one can achie…
Learning Correspondence Structures for Person Re-identification
Weiyao Lin, Yang Shen, Junchi Yan +4
This paper addresses the problem of handling spatial misalignments due to camera-view changes or human-pose variations in person re-identification. We first introduce a boosting-ba…
SiamRCR: Reciprocal Classification and Regression for Visual Object Tracking
Jinlong Peng, Zhengkai Jiang, Yueyang Gu +5
Recently, most siamese network based trackers locate targets via object classification and bounding-box regression. Generally, they select the bounding-box with maximum classificat…
Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
Zhihao He, Tianyao He, Yun Xu +5
Despite the prosperity of the video language model, the current pursuit of comprehensive video reasoning is thwarted by the inherent spatio-temporal incompleteness within individua…
A multimodal lossless coding method for skeletons in videos
Mingzhou Liu, Xiaoyi He, Weiyao Lin +4
Nowadays, skeleton information in videos plays an important role in human-centric video analysis but effective coding such massive skeleton information has never been addressed in…
Group Event Detection with a Varying Number of Group Members for Video Surveillance
Weiyao Lin, Ming-Ting Sun, Radha Poovendran +1
This paper presents a novel approach for automatic recognition of group activities for video surveillance applications. We propose to use a group representative to handle the recog…
UnGAP: Uncertainty-Guided Affine Prompting for Real-Time Crack Segmentation
Conghui Li, Huanyu He, Xin Wang +2
Real-time crack segmentation is vital for structural health monitoring but is plagued by aleatoric uncertainties arising from varying lighting, blur, and texture ambiguity. Current…
VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding
Zhihao He, Tieyuan Chen, Kangyu Wang +6
Current Video Large Language Models (Video LLMs) typically encode frames via a vision encoder and employ an autoregressive (AR) LLM for understanding and generation. However, this…
Human in Events: A Large-Scale Benchmark for Human-centric Video Analysis in Complex Events
Weiyao Lin, Huabin Liu, Shizhan Liu +7
Along with the development of modern smart cities, human-centric video analysis has been encountering the challenge of analyzing diverse and complex events in real scenes. A comple…
SPARE-GS: Structural Parsimony and Resource Efficiency for 3D Gaussian Splatting
Zhang Chen, Shuai Wan, Fuzheng Yang +3
3D Gaussian Splatting (3DGS) achieves high-fidelity novel view synthesis in real-time; however its training efficiency and representation compactness are hindered by excessive prim…
Visual Sound Localization in the Wild by Cross-Modal Interference Erasing
Xian Liu, Rui Qian, Hang Zhou +5
The task of audio-visual sound source localization has been well studied under constrained scenes, where the audio recordings are clean. However, in real-world scenarios, audios ar…
Kill Two Birds With One Stone: Boosting Both Object Detection Accuracy and Speed With adaptive Patch-of-Interest Composition
Shihao Zhang, Weiyao Lin, Ping Lu +2
Object detection is an important yet challenging task in video understanding & analysis, where one major challenge lies in the proper balance between two contradictive factors: det…
Tree-based Visualization and Optimization for Image Collection
Xintong Han, Chongyang Zhang, Weiyao Lin +3
The visualization of an image collection is the process of displaying a collection of images on a screen under some specific layout requirements. This paper focuses on an important…
Spatial-Temporal Transformer Networks for Traffic Flow Forecasting
Mingxing Xu, Wenrui Dai, Chunmiao Liu +4
Traffic forecasting has emerged as a core component of intelligent transportation systems. However, timely accurate traffic forecasting, especially long-term forecasting, still rem…
Enriched Long-term Recurrent Convolutional Network for Facial Micro-Expression Recognition
Huai-Qian Khor, John See, Raphael C. W. Phan +1
Facial micro-expression (ME) recognition has posed a huge challenge to researchers for its subtlety in motion and limited databases. Recently, handcrafted techniques have achieved…
AP-Loss for Accurate One-Stage Object Detection
Kean Chen, Weiyao Lin, Jianguo Li +3
One-stage object detectors are trained by optimizing classification-loss and localization-loss simultaneously, with the former suffering much from extreme foreground-background cla…
Towards Accurate One-Stage Object Detection with AP-Loss
Kean Chen, Jianguo Li, Weiyao Lin +6
One-stage object detectors are trained by optimizing classification-loss and localization-loss simultaneously, with the former suffering much from extreme foreground-background cla…
Action Recognition with Coarse-to-Fine Deep Feature Integration and Asynchronous Fusion
Weiyao Lin, Yang Mi, Jianxin Wu +2
Action recognition is an important yet challenging task in computer vision. In this paper, we propose a novel deep-based framework for action recognition, which improves the recogn…
A diffusion and clustering-based approach for finding coherent motions and understanding crowd scenes
Weiyao Lin, Yang Mi, Weiyue Wang +3
This paper addresses the problem of detecting coherent motions in crowd scenes and presents its two applications in crowd scene understanding: semantic region detection and recurre…
CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace Credit
Kangyu Wang, Zhiyun Jiang, Haibo Feng +5
Diffusion large language models (dLLMs) generate text through iterative denoising. In commonly adopted parallel decoding schemes, each step confirms only high-confidence positions…
ThiNet: A Filter Level Pruning Method for Deep Neural Network Compression
Jian-Hao Luo, Jianxin Wu, Weiyao Lin
We propose an efficient and unified framework, namely ThiNet, to simultaneously accelerate and compress CNN models in both training and inference stages. We focus on the filter lev…
MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning
Tieyuan Chen, Huabin Liu, Tianyao He +8
Video causal reasoning aims to achieve a high-level understanding of video content from a causal perspective. However, current video reasoning tasks are limited in scope, primarily…
Trained Rank Pruning for Efficient Deep Neural Networks
Yuhui Xu, Yuxi Li, Shuai Zhang +6
The performance of Deep Neural Networks (DNNs) keeps elevating in recent years with increasing network depth and width. To enable DNNs on edge devices like mobile phones, researche…
An Efficient Coding Method for Coding Region-of-Interest Locations in AVS2
Mingliang Chen, Weiyao Lin, Xiaozhen Zheng
Region-of-Interest (ROI) location information in videos has many practical usages in video coding field, such as video content analysis and user experience improvement. Although RO…
ProgD: Progressive Multi-scale Decoding with Dynamic Graphs for Joint Multi-agent Motion Forecasting
Xing Gao, Zherui Huang, Weiyao Lin +1
Accurate motion prediction of surrounding agents is crucial for the safe planning of autonomous vehicles. Recent advancements have extended prediction techniques from individual ag…
PIoU Loss: Towards Accurate Oriented Object Detection in Complex Environments
Zhiming Chen, Kean Chen, Weiyao Lin +4
Object detection using an oriented bounding box (OBB) can better target rotated objects by reducing the overlap with background areas. Existing OBB approaches are mostly built on h…
Discriminative Sounding Objects Localization via Self-supervised Audiovisual Matching
Di Hu, Rui Qian, Minyue Jiang +5
Discriminatively localizing sounding objects in cocktail-party, i.e., mixed sound scenes, is commonplace for humans, but still challenging for machines. In this paper, we propose a…
Person Re-identification with Correspondence Structure Learning
Yang Shen, Weiyao Lin, Junchi Yan +3
This paper addresses the problem of handling spatial misalignments due to camera-view changes or human-pose variations in person re-identification. We first introduce a boosting-ba…