papers

Publications (105)

cs.CV2026

GRACE: Boosting Video MLLMs with Grounded Action-Centric Evidence for Viewer Sentiment Prediction

Ruoxuan Yang, Tieyuan Chen, Xiaofeng Huang +6

Viewer sentiment prediction in video advertisements aims to infer the latent affective response evoked in the audience. To bridge the gap between what is shown and what is felt, mo…

cs.LG2018

Deep Neural Network Compression with Single and Multiple Level Quantization

Yuhui Xu, Yongzhuang Wang, Aojun Zhou +2

Network quantization is an effective solution to compress deep neural networks for practical usage. Existing network quantization methods cannot sufficiently exploit the depth info…

cs.GR2015

Facial Expression Cloning with Elastic and Muscle Models

Yihao Zhang, Weiyao Lin, Bing Zhou +4

Expression cloning plays an important role in facial expression synthesis. In this paper, a novel algorithm is proposed for facial expression cloning. The proposed algorithm first…

cs.MM2015

A Fast Sub-Pixel Motion Estimation Algorithm for H.264/AVC Video Coding

Weiyao Lin, Krit Panusopone, David M. Baylon +3

Motion Estimation (ME) is one of the most time-consuming parts in video coding. The use of multiple partition sizes in H.264/AVC makes it even more complicated when compared to ME…

cs.CV2021

Class-aware Sounding Objects Localization via Audiovisual Correspondence

Di Hu, Yake Wei, Rui Qian +3

Audiovisual scenes are pervasive in our daily life. It is commonplace for humans to discriminatively localize different sounding objects but quite challenging for machines to achie…

cs.CV2023

Low-Rank Winograd Transformation for 3D Convolutional Neural Networks

Ziran Qin, Mingbao Lin, Weiyao Lin

This paper focuses on Winograd transformation in 3D convolutional neural networks (CNNs) that are more over-parameterized compared with the 2D version. The over-increasing Winograd…

cs.CL2026

DND: Boosting Large Language Models with Dynamic Nested Depth

Tieyuan Chen, Xiaodong Chen, Haoxing Chen +3

We introduce Dynamic Nested Depth (DND), a novel method that improves performance for off-the-shelf LLMs by selecting critical tokens to reprocess in a nested depth manner. Specifi…

eess.IV2025

ABC: Adaptive BayesNet Structure Learning for Computational Scalable Multi-task Image Compression

Yufeng Zhang, Wenrui Dai, Hang Yu +4

Neural Image Compression (NIC) has revolutionized image compression with its superior rate-distortion performance and multi-task capabilities, supporting both human visual percepti…

cs.CV2022

Controllable Augmentations for Video Representation Learning

Rui Qian, Weiyao Lin, John See +1

This paper focuses on self-supervised video representation learning. Most existing approaches follow the contrastive learning pipeline to construct positive and negative pairs by s…

cs.RO2025

Seeing Through Pixel Motion: Learning Obstacle Avoidance from Optical Flow with One Camera

Yu Hu, Yuang Zhang, Yunlong Song +6

Optical flow captures the motion of pixels in an image sequence over time, providing information about movement, depth, and environmental structure. Flying insects utilize this inf…

cs.CV2017

A Tube-and-Droplet-based Approach for Representing and Analyzing Motion Trajectories

Weiyao Lin, Yang Zhou, Hongteng Xu +4

Trajectory analysis is essential in many applications. In this paper, we address the problem of representing motion trajectories in a highly informative way, and consequently utili…

cs.MM2015

A Computation Control Motion Estimation Method for Complexity-Scalable Video Coding

Weiyao Lin, Krit Panusopone, David M. Baylon +1

In this paper, a new Computation-Control Motion Estimation (CCME) method is proposed which can perform Motion Estimation (ME) adaptively under different computation or power budget…

cs.CV2025

Head-Aware KV Cache Compression for Efficient Visual Autoregressive Modeling

Ziran Qin, Youru Lv, Mingbao Lin +4

Visual Autoregressive (VAR) models adopt a next-scale prediction paradigm, offering high-quality content generation with substantially fewer decoding steps. However, existing VAR m…

cs.CV2022

Uni-EDEN: Universal Encoder-Decoder Network by Multi-Granular Vision-Language Pre-training

Yehao Li, Jiahao Fan, Yingwei Pan +3

Vision-language pre-training has been an emerging and fast-developing research topic, which transfers multi-modal knowledge from rich-resource pre-training task to limited-resource…

cs.CV2021

Enhancing Self-supervised Video Representation Learning via Multi-level Feature Optimization

Rui Qian, Yuxi Li, Huabin Liu +5

The crux of self-supervised video representation learning is to build general features from unlabeled videos. However, most recent works have mainly focused on high-level semantics…

cs.CV2015

A Heat-Map-based Algorithm for Recognizing Group Activities in Videos

Weiyao Lin, Hang Chu, Jianxin Wu +2

In this paper, a new heat-map-based (HMB) algorithm is proposed for group activity recognition. The proposed algorithm first models human trajectories as series of "heat sources" a…

cs.CV2022

FRIH: Fine-grained Region-aware Image Harmonization

Jinlong Peng, Zekun Luo, Liang Liu +6

Image harmonization aims to generate a more realistic appearance of foreground and background for a composite image. Existing methods perform the same harmonization process for the…

cs.CV2025

PCGS: Progressive Compression of 3D Gaussian Splatting

Yihang Chen, Mengyao Li, Qianyi Wu +3

3D Gaussian Splatting (3DGS) achieves impressive rendering fidelity and speed for novel view synthesis. However, its substantial data size poses a significant challenge for practic…

cs.LG2024

BasisFormer: Attention-based Time Series Forecasting with Learnable and Interpretable Basis

Zelin Ni, Hang Yu, Shizhan Liu +2

Bases have become an integral part of modern deep learning-based models for time series forecasting due to their ability to act as feature extractors or future references. To be ef…

cs.MM2020

Key-Point Sequence Lossless Compression for Intelligent Video Analysis

Weiyao Lin, Xiaoyi He, Wenrui Dai +4

Feature coding has been recently considered to facilitate intelligent video analysis for urban computing. Instead of raw videos, extracted features in the front-end are encoded and…

cs.CV2025

Massive Activations are the Key to Local Detail Synthesis in Diffusion Transformers

Chaofan Gan, Zicheng Zhao, Yuanpeng Tu +5

Diffusion Transformers (DiTs) have recently emerged as a powerful backbone for visual generation. Recent observations reveal \emph{Massive Activations} (MAs) in their internal feat…

cs.CV2024

DAC: 2D-3D Retrieval with Noisy Labels via Divide-and-Conquer Alignment and Correction

Chaofan Gan, Yuanpeng Tu, Yuxi Li +1

With the recent burst of 2D and 3D data, cross-modal retrieval has attracted increasing attention recently. However, manual labeling by non-experts will inevitably introduce corrup…

cs.CV2017

Fractal Dimension Invariant Filtering and Its CNN-based Implementation

Hongteng Xu, Junchi Yan, Nils Persson +2

Fractal analysis has been widely used in computer vision, especially in texture image processing and texture analysis. The key concept of fractal-based image model is the fractal d…

cs.CV2017

Ensemble of Part Detectors for Simultaneous Classification and Localization

Xiaopeng Zhang, Hongkai Xiong, Weiyao Lin +1

Part-based representation has been proven to be effective for a variety of visual applications. However, automatic discovery of discriminative parts without object/part-level annot…

cs.CV2015

Improved Image Deblurring based on Salient-region Segmentation

Chongyang Zhang, Weiyao Lin, Wei Li +3

Image deblurring techniques play important roles in many image processing applications. As the blur varies spatially across the image plane, it calls for robust and effective metho…

cs.CV2024

HAC: Hash-grid Assisted Context for 3D Gaussian Splatting Compression

Yihang Chen, Qianyi Wu, Weiyao Lin +2

3D Gaussian Splatting (3DGS) has emerged as a promising framework for novel view synthesis, boasting rapid rendering speed with high fidelity. However, the substantial Gaussians an…

cs.CV2020

CFAD: Coarse-to-Fine Action Detector for Spatiotemporal Action Localization

Yuxi Li, Weiyao Lin, John See +4

Most current pipelines for spatio-temporal action localization connect frame-wise or clip-wise detection results to generate action proposals, where only local information is explo…

cs.LG2018

DNQ: Dynamic Network Quantization

Yuhui Xu, Shuai Zhang, Yingyong Qi +3

Network quantization is an effective method for the deployment of neural networks on memory and energy constrained mobile devices. In this paper, we propose a Dynamic Network Quant…

cs.CG2025

NeCGS: Neural Compression for 3D Geometry Sets

Siyu Ren, Junhui Hou, Weiyao Lin +1

We present NeCGS, the first neural compression paradigm, which can compress a geometry set encompassing thousands of detailed and diverse 3D mesh models by up to 900 times with hig…

cs.CV2023

Collaborative Weakly Supervised Video Correlation Learning for Procedure-Aware Instructional Video Analysis

Tianyao He, Huabin Liu, Yuxi Li +4

Video Correlation Learning (VCL), which aims to analyze the relationships between videos, has been widely studied and applied in various general video tasks. However, applying VCL…

cs.CV2023

Spatio-Temporal Point Process for Multiple Object Tracking

Tao Wang, Kean Chen, Weiyao Lin +4

Multiple Object Tracking (MOT) focuses on modeling the relationship of detected objects among consecutive frames and merge them into different trajectories. MOT remains a challengi…

cs.MM2015

Macroblock Classification Method for Video Applications Involving Motions

Weiyao Lin, Ming-Ting Sun, Hongxiang Li +3

In this paper, a macroblock classification method is proposed for various video processing applications involving motions. Based on the analysis of the Motion Vector field in the c…

cs.CL2025

Rodimus*: Breaking the Accuracy-Efficiency Trade-Off with Efficient Attentions

Zhihao He, Hang Yu, Zi Gong +3

Recent advancements in Transformer-based large language models (LLMs) have set new standards in natural language processing. However, the classical softmax attention incurs signifi…

cs.CV2022

Task-adaptive Spatial-Temporal Video Sampler for Few-shot Action Recognition

Huabin Liu, Weixian Lv, John See +1

A primary challenge faced in few-shot action recognition is inadequate video data for training. To address this issue, current methods in this field mainly focus on devising algori…

cs.CV2021

Variational Pedestrian Detection

Yuang Zhang, Huanyu He, Jianguo Li +3

Pedestrian detection in a crowd is a challenging task due to a high number of mutually-occluding human instances, which brings ambiguity and optimization difficulties to the curren…

cs.CV2025

CSTA: Spatial-Temporal Causal Adaptive Learning for Exemplar-Free Video Class-Incremental Learning

Tieyuan Chen, Huabin Liu, Chern Hong Lim +4

Continual learning aims to acquire new knowledge while retaining past information. Class-incremental learning (CIL) presents a challenging scenario where classes are introduced seq…

cs.CV2025

MCA: 2D-3D Retrieval with Noisy Labels via Multi-level Adaptive Correction and Alignment

Gui Zou, Chaofan Gan, Chern Hong Lim +2

With the increasing availability of 2D and 3D data, significant advancements have been made in the field of cross-modal retrieval. Nevertheless, the existence of imperfect annotati…

cs.CL2025

CAKE: Cascading and Adaptive KV Cache Eviction with Layer Preferences

Ziran Qin, Yuchen Cao, Mingbao Lin +5

Large language models (LLMs) excel at processing long sequences, boosting demand for key-value (KV) caching. While recent efforts to evict KV cache have alleviated the inference bu…

cs.CV2025

Fast Feedforward 3D Gaussian Splatting Compression

Yihang Chen, Qianyi Wu, Mengyao Li +3

With 3D Gaussian Splatting (3DGS) advancing real-time and high-fidelity rendering for novel view synthesis, storage requirements pose challenges for their widespread adoption. Alth…

eess.IV2019

Partition-Aware Adaptive Switching Neural Networks for Post-Processing in HEVC

Weiyao Lin, Xiaoyi He, Xintong Han +5

This paper addresses neural network based post-processing for the state-of-the-art video coding standard, High Efficiency Video Coding (HEVC). We first propose a partition-aware Co…

cs.CV2020

Finding Action Tubes with a Sparse-to-Dense Framework

Yuxi Li, Weiyao Lin, Tao Wang +5

The task of spatial-temporal action detection has attracted increasing attention among researchers. Existing dominant methods solve this problem by relying on short-term informatio…

cs.CV2017

Motion Segmentation via Global and Local Sparse Subspace Optimization

Michael Ying Yang, Hanno Ackermann, Weiyao Lin +2

In this paper, we propose a new framework for segmenting feature-based moving objects under affine subspace model. Since the feature trajectories in practice are high-dimensional a…

cs.MM2015

Region-Based Rate-Control for H.264/AVC for Low Bit-Rate Applications

Hai-Miao Hu, Bo Li, Weiyao Lin +2

Rate-control plays an important role in video coding. However, in the conventional rate-control algorithms, the number and position of Macroblocks (MBs) inside one basic unit for r…

cs.CV2015

A new network-based algorithm for human activity recognition in video

Weiyao Lin, Yuanzhe Chen, Jianxin Wu +3

In this paper, a new network-transmission-based (NTB) algorithm is proposed for human activity recognition in videos. The proposed NTB algorithm models the entire scene as an error…

cs.CV2026

Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning

Shijie Li, Yilin Gao, Siyuan Yang +7

Multimodal Large Language Models (MLLMs) are often constrained by a language-space bottleneck, forcing complex visual reasoning into discrete tokens which can lose perceptual nuanc…

cs.CV2025

MECD+: Unlocking Event-Level Causal Graph Discovery for Video Reasoning

Tieyuan Chen, Huabin Liu, Yi Wang +5

Video causal reasoning aims to achieve a high-level understanding of videos from a causal perspective. However, it exhibits limitations in its scope, primarily executed in a questi…

cs.CV2026

From Priors to Perception: Grounding Video-LLMs in Physical Reality

Zicheng Zhao, Chaofan Gan, Shijie Li +1

While Video Large Language Models (Video-LLMs) excel in general understanding, they exhibit systematic deficits in fine-grained physical reasoning. Existing interventions not only…

cs.CV2025

CogStream: Context-guided Streaming Video Question Answering

Zicheng Zhao, Kangyu Wang, Shijie Li +3

Despite advancements in Video Large Language Models (Vid-LLMs) improving multimodal understanding, challenges persist in streaming video reasoning due to its reliance on contextual…

cs.CV2018

Unsupervised Deep Domain Adaptation for Pedestrian Detection

Lihang Liu, Weiyao Lin, Lisheng Wu +2

This paper addresses the problem of unsupervised domain adaptation on the task of pedestrian detection in crowded scenes. First, we utilize an iterative algorithm to iteratively se…

cs.CV2025

Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive Activations

Chaofan Gan, Yuanpeng Tu, Xi Chen +4

Pre-trained stable diffusion models (SD) have shown great advances in visual correspondence. In this paper, we investigate the capabilities of Diffusion Transformers (DiTs) for acc…

cs.CV2023

Few-shot Action Recognition via Intra- and Inter-Video Information Maximization

Huabin Liu, Weiyao Lin, Tieyuan Chen +3

Current few-shot action recognition involves two primary sources of information for classification:(1) intra-video information, determined by frame content within a single video cl…

cs.CV2018

Tiny-DSOD: Lightweight Object Detection for Resource-Restricted Usages

Yuxi Li, Jiuwei Li, Weiyao Lin +1

Object detection has made great progress in the past few years along with the development of deep learning. However, most current object detection methods are resource hungry, whic…

cs.CV2025

Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens

Ziran Qin, Youru Lv, Mingbao Lin +4

Autoregressive (AR) visual generation has emerged as a powerful paradigm for image and multimodal synthesis, owing to its scalability and generality. However, existing AR image gen…

cs.CV2022

The 1st-place Solution for ECCV 2022 Multiple People Tracking in Group Dance Challenge

Yuang Zhang, Tiancai Wang, Weiyao Lin +1

We present our 1st place solution to the Group Dance Multiple People Tracking Challenge. Based on MOTR: End-to-End Multiple-Object Tracking with Transformer, we explore: 1) detect…

cs.CV2015

Activity Recognition Using A Combination of Category Components And Local Models for Video Surveillance

Weiyao Lin, Ming-Ting Sun, Radha Poovendran +1

This paper presents a novel approach for automatic recognition of human activities for video surveillance applications. We propose to represent an activity by a combination of cate…

cs.CV2020

Delving into the Cyclic Mechanism in Semi-supervised Video Object Segmentation

Yuxi Li, Ning Xu, Jinlong Peng +2

In this paper, we address several inadequacies of current video object segmentation pipelines. Firstly, a cyclic mechanism is incorporated to the standard semi-supervised process t…

cs.CV2022

TA2N: Two-Stage Action Alignment Network for Few-shot Action Recognition

Shuyuan Li, Huabin Liu, Rui Qian +5

Few-shot action recognition aims to recognize novel action classes (query) using just a few samples (support). The majority of current approaches follow the metric learning paradig…

cs.RO2024

Back to Newton's Laws: Learning Vision-based Agile Flight via Differentiable Physics

Yuang Zhang, Yu Hu, Yunlong Song +2

Swarm navigation in cluttered environments is a grand challenge in robotics. This work combines deep learning with first-principle physics through differentiable simulation to enab…

cs.CV2025

HAC++: Towards 100X Compression of 3D Gaussian Splatting

Yihang Chen, Qianyi Wu, Weiyao Lin +2

3D Gaussian Splatting (3DGS) has emerged as a promising framework for novel view synthesis, boasting rapid rendering speed with high fidelity. However, the substantial Gaussians an…

cs.CV2018

Network Decoupling: From Regular to Depthwise Separable Convolutions

Jianbo Guo, Yuxi Li, Weiyao Lin +2

Depthwise separable convolution has shown great efficiency in network design, but requires time-consuming training procedure with full training-set available. This paper first anal…

cs.CV2020

ATRW: A Benchmark for Amur Tiger Re-identification in the Wild

Shuyuan Li, Jianguo Li, Hanlin Tang +2

Monitoring the population and movements of endangered species is an important task to wildlife conversation. Traditional tagging methods do not scale to large populations, while ap…

cs.CV2020

Multiple Sound Sources Localization from Coarse to Fine

Rui Qian, Di Hu, Heinrich Dinkel +3

How to visually localize multiple sound sources in unconstrained videos is a formidable problem, especially when lack of the pairwise sound-object annotations. To solve this proble…

cs.CV2021

LSTC: Boosting Atomic Action Detection with Long-Short-Term Context

Yuxi Li, Boshen Zhang, Jian Li +5

In this paper, we place the atomic action detection problem into a Long-Short Term Context (LSTC) to analyze how the temporal reliance among video signals affect the action detecti…

cs.CV2022

Exploring the Semi-supervised Video Object Segmentation Problem from a Cyclic Perspective

Yuxi Li, Ning Xu, Wenjie Yang +2

Modern video object segmentation (VOS) algorithms have achieved remarkably high performance in a sequential processing order, while most of currently prevailing pipelines still sho…

cs.LG2020

TRP: Trained Rank Pruning for Efficient Deep Neural Networks

Yuhui Xu, Yuxi Li, Shuai Zhang +6

To enable DNNs on edge devices like mobile phones, low-rank approximation has been widely adopted because of its solid theoretical rationale and efficient implementations. Several…

cs.CV2026

HACK++: Towards More Effective Head-Aware Key-Value Compression for Efficient Visual Autoregressive Modeling

Ziran Qin, Yuchen Jiang, Mingbao Lin +4

Visual Autoregressive (VAR) models adopt a next-scale prediction paradigm, offering high-quality generation with substantially fewer decoding steps. However, existing VAR models su…

cs.CV2026

Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation

Chaofan Gan, Zicheng Zhao, Yuanpeng Tu +6

Massive Activations (MAs) have been widely observed in Transformer-based models, yet their structure and functional roles in Diffusion Transformers (DiTs) remain insufficiently und…

cs.MM2023

Scene Graph Lossless Compression with Adaptive Prediction for Objects and Relations

Yufeng Zhang, Weiyao Lin, Wenrui Dai +2

The scene graph is a new data structure describing objects and their pairwise relationship within image scenes. As the size of scene graph in vision applications grows, how to loss…

cs.CV2015

Intra-and-Inter-Constraint-based Video Enhancement based on Piecewise Tone Mapping

Yuanzhe Chen, Weiyao Lin, Chongyang Zhang +3

Video enhancement plays an important role in various video applications. In this paper, we propose a new intra-and-inter-constraint-based video enhancement approach aiming to 1) ac…

cs.CV2019

Group Re-Identification with Multi-grained Matching and Integration

Weiyao Lin, Yuxi Li, Hao Xiao +5

The task of re-identifying groups of people underdifferent camera views is an important yet less-studied problem.Group re-identification (Re-ID) is a very challenging task sinceit…

eess.SP2026

Object-Attribute-Relation Model Driven Adaptive Hierarchical Transmission for Multimodal Semantic Communication

Chenxing Li, Yiping Duan, Han Jiao +3

Traditional video coding (VVC, HEVC) prioritizes human visual perception, transmitting substantial texture redundancy that severely hinders machine decision-making under constraine…

cs.MM2018

Enhancing HEVC Compressed Videos with a Partition-masked Convolutional Neural Network

Xiaoyi He, Qiang Hu, Xintong Han +3

In this paper, we propose a partition-masked Convolution Neural Network (CNN) to achieve compressed-video enhancement for the state-of-the-art coding standard, High Efficiency Vide…

cs.CV2015

Deep Spatial Pyramid: The Devil is Once Again in the Details

Bin-Bin Gao, Xiu-Shen Wei, Jianxin Wu +1

In this paper we show that by carefully making good choices for various detailed but important factors in a visual recognition framework using deep learning features, one can achie…

cs.CV2017

Learning Correspondence Structures for Person Re-identification

Weiyao Lin, Yang Shen, Junchi Yan +4

This paper addresses the problem of handling spatial misalignments due to camera-view changes or human-pose variations in person re-identification. We first introduce a boosting-ba…

cs.CV2021

SiamRCR: Reciprocal Classification and Regression for Visual Object Tracking

Jinlong Peng, Zhengkai Jiang, Yueyang Gu +5

Recently, most siamese network based trackers locate targets via object classification and bounding-box regression. Generally, they select the bounding-box with maximum classificat…

cs.CV2025

Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning

Zhihao He, Tianyao He, Yun Xu +5

Despite the prosperity of the video language model, the current pursuit of comprehensive video reasoning is thwarted by the inherent spatio-temporal incompleteness within individua…

cs.MM2019

A multimodal lossless coding method for skeletons in videos

Mingzhou Liu, Xiaoyi He, Weiyao Lin +4

Nowadays, skeleton information in videos plays an important role in human-centric video analysis but effective coding such massive skeleton information has never been addressed in…

cs.CV2015

Group Event Detection with a Varying Number of Group Members for Video Surveillance

Weiyao Lin, Ming-Ting Sun, Radha Poovendran +1

This paper presents a novel approach for automatic recognition of group activities for video surveillance applications. We propose to use a group representative to handle the recog…

cs.CV2026

UnGAP: Uncertainty-Guided Affine Prompting for Real-Time Crack Segmentation

Conghui Li, Huanyu He, Xin Wang +2

Real-time crack segmentation is vital for structural health monitoring but is plagued by aleatoric uncertainties arising from varying lighting, blur, and texture ambiguity. Current…

cs.CV2026

VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding

Zhihao He, Tieyuan Chen, Kangyu Wang +6

Current Video Large Language Models (Video LLMs) typically encode frames via a vision encoder and employ an autoregressive (AR) LLM for understanding and generation. However, this…

cs.CV2023

Human in Events: A Large-Scale Benchmark for Human-centric Video Analysis in Complex Events

Weiyao Lin, Huabin Liu, Shizhan Liu +7

Along with the development of modern smart cities, human-centric video analysis has been encountering the challenge of analyzing diverse and complex events in real scenes. A comple…

cs.CV2026

SPARE-GS: Structural Parsimony and Resource Efficiency for 3D Gaussian Splatting

Zhang Chen, Shuai Wan, Fuzheng Yang +3

3D Gaussian Splatting (3DGS) achieves high-fidelity novel view synthesis in real-time; however its training efficiency and representation compactness are hindered by excessive prim…

cs.CV2022

Visual Sound Localization in the Wild by Cross-Modal Interference Erasing

Xian Liu, Rui Qian, Hang Zhou +5

The task of audio-visual sound source localization has been well studied under constrained scenes, where the audio recordings are clean. However, in real-world scenarios, audios ar…

cs.CV2017

Kill Two Birds With One Stone: Boosting Both Object Detection Accuracy and Speed With adaptive Patch-of-Interest Composition

Shihao Zhang, Weiyao Lin, Ping Lu +2

Object detection is an important yet challenging task in video understanding & analysis, where one major challenge lies in the proper balance between two contradictive factors: det…

cs.MM2015

Tree-based Visualization and Optimization for Image Collection

Xintong Han, Chongyang Zhang, Weiyao Lin +3

The visualization of an image collection is the process of displaying a collection of images on a screen under some specific layout requirements. This paper focuses on an important…

eess.SP2021

Spatial-Temporal Transformer Networks for Traffic Flow Forecasting

Mingxing Xu, Wenrui Dai, Chunmiao Liu +4

Traffic forecasting has emerged as a core component of intelligent transportation systems. However, timely accurate traffic forecasting, especially long-term forecasting, still rem…

cs.CV2018

Enriched Long-term Recurrent Convolutional Network for Facial Micro-Expression Recognition

Huai-Qian Khor, John See, Raphael C. W. Phan +1

Facial micro-expression (ME) recognition has posed a huge challenge to researchers for its subtlety in motion and limited databases. Recently, handcrafted techniques have achieved…

cs.CV2020

AP-Loss for Accurate One-Stage Object Detection

Kean Chen, Weiyao Lin, Jianguo Li +3

One-stage object detectors are trained by optimizing classification-loss and localization-loss simultaneously, with the former suffering much from extreme foreground-background cla…

cs.CV2020

Towards Accurate One-Stage Object Detection with AP-Loss

Kean Chen, Jianguo Li, Weiyao Lin +6

One-stage object detectors are trained by optimizing classification-loss and localization-loss simultaneously, with the former suffering much from extreme foreground-background cla…

cs.CV2017

Action Recognition with Coarse-to-Fine Deep Feature Integration and Asynchronous Fusion

Weiyao Lin, Yang Mi, Jianxin Wu +2

Action recognition is an important yet challenging task in computer vision. In this paper, we propose a novel deep-based framework for action recognition, which improves the recogn…

cs.CV2016

A diffusion and clustering-based approach for finding coherent motions and understanding crowd scenes

Weiyao Lin, Yang Mi, Weiyue Wang +3

This paper addresses the problem of detecting coherent motions in crowd scenes and presents its two applications in crowd scene understanding: semantic region detection and recurre…

cs.CL2026

CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace Credit

Kangyu Wang, Zhiyun Jiang, Haibo Feng +5

Diffusion large language models (dLLMs) generate text through iterative denoising. In commonly adopted parallel decoding schemes, each step confirms only high-confidence positions…

cs.CV2017

ThiNet: A Filter Level Pruning Method for Deep Neural Network Compression

Jian-Hao Luo, Jianxin Wu, Weiyao Lin

We propose an efficient and unified framework, namely ThiNet, to simultaneously accelerate and compress CNN models in both training and inference stages. We focus on the filter lev…

cs.CV2024

MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning

Tieyuan Chen, Huabin Liu, Tianyao He +8

Video causal reasoning aims to achieve a high-level understanding of video content from a causal perspective. However, current video reasoning tasks are limited in scope, primarily…

cs.CV2020

Trained Rank Pruning for Efficient Deep Neural Networks

Yuhui Xu, Yuxi Li, Shuai Zhang +6

The performance of Deep Neural Networks (DNNs) keeps elevating in recent years with increasing network depth and width. To enable DNNs on edge devices like mobile phones, researche…

cs.MM2015

An Efficient Coding Method for Coding Region-of-Interest Locations in AVS2

Mingliang Chen, Weiyao Lin, Xiaozhen Zheng

Region-of-Interest (ROI) location information in videos has many practical usages in video coding field, such as video content analysis and user experience improvement. Although RO…

cs.AI2025

ProgD: Progressive Multi-scale Decoding with Dynamic Graphs for Joint Multi-agent Motion Forecasting

Xing Gao, Zherui Huang, Weiyao Lin +1

Accurate motion prediction of surrounding agents is crucial for the safe planning of autonomous vehicles. Recent advancements have extended prediction techniques from individual ag…

cs.CV2020

PIoU Loss: Towards Accurate Oriented Object Detection in Complex Environments

Zhiming Chen, Kean Chen, Weiyao Lin +4

Object detection using an oriented bounding box (OBB) can better target rotated objects by reducing the overlap with background areas. Existing OBB approaches are mostly built on h…

cs.CV2020

Discriminative Sounding Objects Localization via Self-supervised Audiovisual Matching

Di Hu, Rui Qian, Minyue Jiang +5

Discriminatively localizing sounding objects in cocktail-party, i.e., mixed sound scenes, is commonplace for humans, but still challenging for machines. In this paper, we propose a…

cs.CV2015

Person Re-identification with Correspondence Structure Learning

Yang Shen, Weiyao Lin, Junchi Yan +3

This paper addresses the problem of handling spatial misalignments due to camera-view changes or human-pose variations in person re-identification. We first introduce a boosting-ba…