activity
20242026
collaborators

7 papers

cs.CV2026

AdaDINO: Context-Adaptive DINO-Distilled Vision Foundation Models for Efficient Open-Vocabulary Edge Inference

Yiwei Zhao, Yi Zheng, Huapeng Su +9

Always-on contextual AI runs language-aligned vision foundation models (VFMs) on edge devices, where the on-device model is the dominant continuous compute cost under strict latenc…

cs.LG2026

EgoEverything: A Benchmark for Human Behavior Inspired Long Context Egocentric Video Understanding in AR Environment

Qiance Tang, Ziqi Wang, Jieyu Lin +3

Long context egocentric video understanding has recently attracted significant research attention, with augmented reality (AR) highlighted as one of its most important application…

cs.AI2025

PETAH: Parameter Efficient Task Adaptation for Hybrid Transformers in a resource-limited Context

Maximilian Augustin, Syed Shakib Sarwar, Mostafa Elhoushi +3

Following their success in natural language processing (NLP), there has been a shift towards transformer models in computer vision. While transformers perform well and offer promis…

cs.CV2024

Unlocking Visual Secrets: Inverting Features with Diffusion Priors for Image Reconstruction

Sai Qian Zhang, Ziyun Li, Chuan Guo +5

Inverting visual representations within deep neural networks (DNNs) presents a challenging and important problem in the field of security and privacy for deep learning. The main go…

cs.CV2024

GazeGen: Gaze-Driven User Interaction for Visual Content Generation

He-Yen Hsieh, Ziyun Li, Sai Qian Zhang +5

We present GazeGen, a user interaction system that generates visual content (images and videos) for locations indicated by the user's eye gaze. GazeGen allows intuitive manipulatio…

cs.CV2024

x-RAGE: eXtended Reality -- Action & Gesture Events Dataset

Vivek Parmar, Dwijay Bane, Syed Shakib Sarwar +3

With the emergence of the Metaverse and focus on wearable devices in the recent years gesture based human-computer interaction has gained significance. To enable gesture recognitio…