Publications (31)
ActionHub: A Large-scale Action Video Description Dataset for Zero-shot Action Recognition
Jiaming Zhou, Junwei Liang, Kun-Yu Lin +2
Zero-shot action recognition (ZSAR) aims to learn an alignment model between videos and class descriptions of seen actions that is transferable to unseen actions. The text queries…
DilateFormer: Multi-Scale Dilated Transformer for Visual Recognition
Jiayu Jiao, Yu-Ming Tang, Kun-Yu Lin +4
As a de facto solution, the vanilla Vision Transformers (ViTs) are encouraged to model long-range dependencies between arbitrary image patches while the global attended receptive f…
Towards Weakly Supervised End-to-end Learning for Long-video Action Recognition
Jiaming Zhou, Hanjun Li, Kun-Yu Lin +1
Developing end-to-end action recognition models on long videos is fundamental and crucial for long-video action understanding. Due to the unaffordable cost of end-to-end training o…
Out-of-distribution Detection by Cross-class Vicinity Distribution of In-distribution Data
Zhilin Zhao, Longbing Cao, Kun-Yu Lin
Deep neural networks for image classification only learn to map in-distribution inputs to their corresponding ground truth labels in training without differentiating out-of-distrib…
Revealing the Distributional Vulnerability of Discriminators by Implicit Generators
Zhilin Zhao, Longbing Cao, Kun-Yu Lin
In deep neural learning, a discriminator trained on in-distribution (ID) samples may make high-confidence predictions on out-of-distribution (OOD) samples. This triggers a signific…
Distilling the Unknown to Unveil Certainty
Zhilin Zhao, Longbing Cao, Yixuan Zhang +2
Out-of-distribution (OOD) detection is critical for identifying test samples that deviate from in-distribution (ID) data, ensuring network robustness and reliability. This paper pr…
Diversifying Spatial-Temporal Perception for Video Domain Generalization
Kun-Yu Lin, Jia-Run Du, Yipeng Gao +2
Video domain generalization aims to learn generalizable video classification models for unseen target domains by training in a source domain. A critical challenge of video domain g…
CMT-RAG: Complementary Memory Traces for Multi-turn Multi-hop RAG
Lang Zhou, Yingjian Chen, Shuxuan Li +2
The paper proposes CMT-RAG, a framework that stores structured sub-question reasoning traces as memory to improve multi-turn, multi-hop retrieval‑augmented generation, and introduc…
XOV-Action: Towards Generalizable Open-Vocabulary Action Recognition
Kun-Yu Lin, Henghui Ding, Jia-Run Du +6
Inspired by the impressive success of image-text foundation models, recent works have proposed to adapt these foundation models to video data, leading to efficient and effective vi…
Human-Centric Transformer for Domain Adaptive Action Recognition
Kun-Yu Lin, Jiaming Zhou, Wei-Shi Zheng
We study the domain adaptation task for action recognition, namely domain adaptive action recognition, which aims to effectively transfer action recognition power from a label-suff…
Exploring the Limits of Vision-Language-Action Manipulations in Cross-task Generalization
Jiaming Zhou, Ke Ye, Jiayi Liu +6
The generalization capabilities of vision-language-action (VLA) models to unseen tasks are crucial to achieving general-purpose robotic manipulation in open-world settings. However…
CoopDiff: Anticipating 3D Human-object Interactions via Contact-consistent Decoupled Diffusion
Xiaotong Lin, Tianming Liang, Jian-Fang Hu +5
3D human-object interaction (HOI) anticipation aims to predict the future motion of humans and their manipulated objects, conditioned on the historical context. Generally, the arti…
Towards Completeness: A Generalizable Action Proposal Generator for Zero-Shot Temporal Action Localization
Jia-Run Du, Kun-Yu Lin, Jingke Meng +1
To address the zero-shot temporal action localization (ZSTAL) task, existing works develop models that are generalizable to detect and classify actions from unseen categories. They…
Weakly-Supervised Temporal Action Localization by Progressive Complementary Learning
Jia-Run Du, Jia-Chang Feng, Kun-Yu Lin +5
Weakly Supervised Temporal Action Localization (WSTAL) aims to localize and classify action instances in long untrimmed videos with only video-level category labels. Due to the lac…
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding
Yi-Xing Peng, Qize Yang, Yu-Ming Tang +4
Fine-grained understanding of human actions and poses in videos is essential for human-centric AI applications. In this work, we introduce ActionArt, a fine-grained video-caption d…
Panoptic Captioning: An Equivalence Bridge for Image and Text
Kun-Yu Lin, Hongjun Wang, Weining Ren +1
This work introduces panoptic captioning, a novel task striving to seek the minimum text equivalent of images, which has broad potential applications. We take the first step toward…
ParGo: Bridging Vision-Language with Partial and Global Views
An-Lan Wang, Bin Shan, Wei Shi +7
This work presents ParGo, a novel Partial-Global projector designed to connect the vision and language modalities for Multimodal Large Language Models (MLLMs). Unlike previous work…
Humanoid Whole-Body Manipulation via Active Spatial Brain and Generalizable Action Cerebellum
Zhizhao Liang, Yi-Lin Wei, Xuhang Chen +6
In this paper, we explore spatial-aware humanoid whole-body manipulation task. Compared with tabletop settings, this task poses two key challenges: 1) Spatial understanding is chal…
Event-Guided Procedure Planning from Instructional Videos with Text Supervision
An-Lan Wang, Kun-Yu Lin, Jia-Run Du +2
In this work, we focus on the task of procedure planning from instructional videos with text supervision, where a model aims to predict an action sequence to transform the initial…
Holo-Captioning: Toward the Text Equivalent of 3D Scenes
Kun-Yu Lin, Chengke Bu, Zhenguo Li +1
This work introduces holo-captioning, a novel task that strives to seek the text equivalent of 3D scenes. As the initial step, we formulate holo-captioning as generating a structur…
Mitigating the Human-Robot Domain Discrepancy in Visual Pre-training for Robotic Manipulation
Jiaming Zhou, Teli Ma, Kun-Yu Lin +3
Learning generalizable visual representations across different embodied environments is essential for effective robotic manipulation in real-world scenarios. However, the limited s…
TechCoach: Towards Technical-Point-Aware Descriptive Action Coaching
Yuan-Ming Li, An-Lan Wang, Kun-Yu Lin +4
To guide a learner in mastering action skills, it is crucial for a coach to 1) reason through the learner's action execution and technical points (TechPoints), and 2) provide detai…
From Watch to Imagine: Steering Long-horizon Manipulation via Human Demonstration and Future Envisionment
Ke Ye, Jiaming Zhou, Yuanfeng Qiu +4
Generalizing to long-horizon manipulation tasks in a zero-shot setting remains a central challenge in robotics. Current multimodal foundation based approaches, despite their capabi…
Task-Oriented 6-DoF Grasp Pose Detection in Clutters
An-Lan Wang, Nuo Chen, Kun-Yu Lin +2
In general, humans would grasp an object differently for different tasks, e.g., "grasping the handle of a knife to cut" vs. "grasping the blade to hand over". In the field of robot…
Supervision Adaptation Balancing In-distribution Generalization and Out-of-distribution Detection
Zhilin Zhao, Longbing Cao, Kun-Yu Lin
The discrepancy between in-distribution (ID) and out-of-distribution (OOD) samples can lead to \textit{distributional vulnerability} in deep neural networks, which can subsequently…
Modeling Multiple Normal Action Representations for Error Detection in Procedural Tasks
Wei-Jin Huang, Yuan-Ming Li, Zhi-Wei Xia +4
Error detection in procedural activities is essential for consistent and correct outcomes in AR-assisted and robotic systems. Existing methods often focus on temporal ordering erro…
ProEdit: Inversion-based Editing From Prompts Done Right
Zhi Ouyang, Dian Zheng, Xiao-Ming Wu +4
Inversion-based visual editing provides an effective and training-free way to edit an image or a video based on user instructions. Existing methods typically inject source image in…
EgoTraj-Bench: Towards Robust Trajectory Prediction Under Ego-view Noisy Observations
Jiayi Liu, Jiaming Zhou, Ke Ye +3
Reliable trajectory prediction from an ego-centric perspective is crucial for robotic navigation in human-centric environments. However, existing methods typically assume noiseless…
ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations
Tianming Liang, Kun-Yu Lin, Chaolei Tan +3
Referring video object segmentation (RVOS) aims to segment target objects throughout a video based on a text description. This is challenging as it involves deep vision-language un…
Decoupled Distillation to Erase: A General Unlearning Method for Any Class-centric Tasks
Yu Zhou, Dian Zheng, Qijie Mo +3
In this work, we present DEcoupLEd Distillation To Erase (DELETE), a general and strong unlearning method for any class-centric tasks. To derive this, we first propose a theoretica…
ViSpeak: Visual Instruction Feedback in Streaming Videos
Shenghao Fu, Qize Yang, Yuan-Ming Li +6
Recent advances in Large Multi-modal Models (LMMs) are primarily focused on offline video understanding. Instead, streaming video understanding poses great challenges to recent mod…