papers

Publications (31)

cs.CV2024

ActionHub: A Large-scale Action Video Description Dataset for Zero-shot Action Recognition

Jiaming Zhou, Junwei Liang, Kun-Yu Lin +2

Zero-shot action recognition (ZSAR) aims to learn an alignment model between videos and class descriptions of seen actions that is transferable to unseen actions. The text queries…

cs.CV2023

DilateFormer: Multi-Scale Dilated Transformer for Visual Recognition

Jiayu Jiao, Yu-Ming Tang, Kun-Yu Lin +4

As a de facto solution, the vanilla Vision Transformers (ViTs) are encouraged to model long-range dependencies between arbitrary image patches while the global attended receptive f…

cs.CV2024

Towards Weakly Supervised End-to-end Learning for Long-video Action Recognition

Jiaming Zhou, Hanjun Li, Kun-Yu Lin +1

Developing end-to-end action recognition models on long videos is fundamental and crucial for long-video action understanding. Due to the unaffordable cost of end-to-end training o…

cs.LG2023

Out-of-distribution Detection by Cross-class Vicinity Distribution of In-distribution Data

Zhilin Zhao, Longbing Cao, Kun-Yu Lin

Deep neural networks for image classification only learn to map in-distribution inputs to their corresponding ground truth labels in training without differentiating out-of-distrib…

cs.LG2023

Revealing the Distributional Vulnerability of Discriminators by Implicit Generators

Zhilin Zhao, Longbing Cao, Kun-Yu Lin

In deep neural learning, a discriminator trained on in-distribution (ID) samples may make high-confidence predictions on out-of-distribution (OOD) samples. This triggers a signific…

cs.LG2025

Distilling the Unknown to Unveil Certainty

Zhilin Zhao, Longbing Cao, Yixuan Zhang +2

Out-of-distribution (OOD) detection is critical for identifying test samples that deviate from in-distribution (ID) data, ensuring network robustness and reliability. This paper pr…

cs.CV2023

Diversifying Spatial-Temporal Perception for Video Domain Generalization

Kun-Yu Lin, Jia-Run Du, Yipeng Gao +2

Video domain generalization aims to learn generalizable video classification models for unseen target domains by training in a source domain. A critical challenge of video domain g…

cs.CL2026

CMT-RAG: Complementary Memory Traces for Multi-turn Multi-hop RAG

Lang Zhou, Yingjian Chen, Shuxuan Li +2

The paper proposes CMT-RAG, a framework that stores structured sub-question reasoning traces as memory to improve multi-turn, multi-hop retrieval‑augmented generation, and introduc…

#multi-turn conversation#multi-hop reasoning#retrieval-augmented generation#memory tracing
cs.CV2026

XOV-Action: Towards Generalizable Open-Vocabulary Action Recognition

Kun-Yu Lin, Henghui Ding, Jia-Run Du +6

Inspired by the impressive success of image-text foundation models, recent works have proposed to adapt these foundation models to video data, leading to efficient and effective vi…

cs.CV2024

Human-Centric Transformer for Domain Adaptive Action Recognition

Kun-Yu Lin, Jiaming Zhou, Wei-Shi Zheng

We study the domain adaptation task for action recognition, namely domain adaptive action recognition, which aims to effectively transfer action recognition power from a label-suff…

cs.RO2025

Exploring the Limits of Vision-Language-Action Manipulations in Cross-task Generalization

Jiaming Zhou, Ke Ye, Jiayi Liu +6

The generalization capabilities of vision-language-action (VLA) models to unseen tasks are crucial to achieving general-purpose robotic manipulation in open-world settings. However…

cs.CV2025

CoopDiff: Anticipating 3D Human-object Interactions via Contact-consistent Decoupled Diffusion

Xiaotong Lin, Tianming Liang, Jian-Fang Hu +5

3D human-object interaction (HOI) anticipation aims to predict the future motion of humans and their manipulated objects, conditioned on the historical context. Generally, the arti…

cs.CV2024

Towards Completeness: A Generalizable Action Proposal Generator for Zero-Shot Temporal Action Localization

Jia-Run Du, Kun-Yu Lin, Jingke Meng +1

To address the zero-shot temporal action localization (ZSTAL) task, existing works develop models that are generalizable to detect and classify actions from unseen categories. They…

cs.CV2022

Weakly-Supervised Temporal Action Localization by Progressive Complementary Learning

Jia-Run Du, Jia-Chang Feng, Kun-Yu Lin +5

Weakly Supervised Temporal Action Localization (WSTAL) aims to localize and classify action instances in long untrimmed videos with only video-level category labels. Due to the lac…

cs.CV2025

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding

Yi-Xing Peng, Qize Yang, Yu-Ming Tang +4

Fine-grained understanding of human actions and poses in videos is essential for human-centric AI applications. In this work, we introduce ActionArt, a fine-grained video-caption d…

cs.CV2025

Panoptic Captioning: An Equivalence Bridge for Image and Text

Kun-Yu Lin, Hongjun Wang, Weining Ren +1

This work introduces panoptic captioning, a novel task striving to seek the minimum text equivalent of images, which has broad potential applications. We take the first step toward…

cs.CV2025

ParGo: Bridging Vision-Language with Partial and Global Views

An-Lan Wang, Bin Shan, Wei Shi +7

This work presents ParGo, a novel Partial-Global projector designed to connect the vision and language modalities for Multimodal Large Language Models (MLLMs). Unlike previous work…

cs.RO2026

Humanoid Whole-Body Manipulation via Active Spatial Brain and Generalizable Action Cerebellum

Zhizhao Liang, Yi-Lin Wei, Xuhang Chen +6

In this paper, we explore spatial-aware humanoid whole-body manipulation task. Compared with tabletop settings, this task poses two key challenges: 1) Spatial understanding is chal…

cs.CV2023

Event-Guided Procedure Planning from Instructional Videos with Text Supervision

An-Lan Wang, Kun-Yu Lin, Jia-Run Du +2

In this work, we focus on the task of procedure planning from instructional videos with text supervision, where a model aims to predict an action sequence to transform the initial…

cs.CV2026

Holo-Captioning: Toward the Text Equivalent of 3D Scenes

Kun-Yu Lin, Chengke Bu, Zhenguo Li +1

This work introduces holo-captioning, a novel task that strives to seek the text equivalent of 3D scenes. As the initial step, we formulate holo-captioning as generating a structur…

cs.CV2025

Mitigating the Human-Robot Domain Discrepancy in Visual Pre-training for Robotic Manipulation

Jiaming Zhou, Teli Ma, Kun-Yu Lin +3

Learning generalizable visual representations across different embodied environments is essential for effective robotic manipulation in real-world scenarios. However, the limited s…

cs.CV2025

TechCoach: Towards Technical-Point-Aware Descriptive Action Coaching

Yuan-Ming Li, An-Lan Wang, Kun-Yu Lin +4

To guide a learner in mastering action skills, it is crucial for a coach to 1) reason through the learner's action execution and technical points (TechPoints), and 2) provide detai…

cs.RO2025

From Watch to Imagine: Steering Long-horizon Manipulation via Human Demonstration and Future Envisionment

Ke Ye, Jiaming Zhou, Yuanfeng Qiu +4

Generalizing to long-horizon manipulation tasks in a zero-shot setting remains a central challenge in robotics. Current multimodal foundation based approaches, despite their capabi…

cs.RO2025

Task-Oriented 6-DoF Grasp Pose Detection in Clutters

An-Lan Wang, Nuo Chen, Kun-Yu Lin +2

In general, humans would grasp an object differently for different tasks, e.g., "grasping the handle of a knife to cut" vs. "grasping the blade to hand over". In the field of robot…

cs.LG2023

Supervision Adaptation Balancing In-distribution Generalization and Out-of-distribution Detection

Zhilin Zhao, Longbing Cao, Kun-Yu Lin

The discrepancy between in-distribution (ID) and out-of-distribution (OOD) samples can lead to \textit{distributional vulnerability} in deep neural networks, which can subsequently…

cs.CV2025

Modeling Multiple Normal Action Representations for Error Detection in Procedural Tasks

Wei-Jin Huang, Yuan-Ming Li, Zhi-Wei Xia +4

Error detection in procedural activities is essential for consistent and correct outcomes in AR-assisted and robotic systems. Existing methods often focus on temporal ordering erro…

cs.CV2025

ProEdit: Inversion-based Editing From Prompts Done Right

Zhi Ouyang, Dian Zheng, Xiao-Ming Wu +4

Inversion-based visual editing provides an effective and training-free way to edit an image or a video based on user instructions. Existing methods typically inject source image in…

cs.CV2026

EgoTraj-Bench: Towards Robust Trajectory Prediction Under Ego-view Noisy Observations

Jiayi Liu, Jiaming Zhou, Ke Ye +3

Reliable trajectory prediction from an ego-centric perspective is crucial for robotic navigation in human-centric environments. However, existing methods typically assume noiseless…

cs.CV2025

ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations

Tianming Liang, Kun-Yu Lin, Chaolei Tan +3

Referring video object segmentation (RVOS) aims to segment target objects throughout a video based on a text description. This is challenging as it involves deep vision-language un…

cs.CV2025

Decoupled Distillation to Erase: A General Unlearning Method for Any Class-centric Tasks

Yu Zhou, Dian Zheng, Qijie Mo +3

In this work, we present DEcoupLEd Distillation To Erase (DELETE), a general and strong unlearning method for any class-centric tasks. To derive this, we first propose a theoretica…

cs.CV2025

ViSpeak: Visual Instruction Feedback in Streaming Videos

Shenghao Fu, Qize Yang, Yuan-Ming Li +6

Recent advances in Large Multi-modal Models (LMMs) are primarily focused on offline video understanding. Instead, streaming video understanding poses great challenges to recent mod…