11 papers
Compositional Context Fine-Tuning Vision-Language Model for Complex Assembly Action Understanding from Videos
Hao Zheng, Jinyi Huang, Tiantian Zheng +2
Assembly action understanding is a key enabler for effective human-robot collaborative assembly, yet it remains challenging due to subtle motions and fine-grained hand-object inter…
Polyphony: Diffusion-based Dual-Hand Action Segmentation with Alternating Vision Transformer and Semantic Conditioning
Hao Zheng, Hu Wang, Tiantian Zheng +2
Dual-hand action segmentation, densely predicting actions for both hands from untrimmed videos, is essential for understanding complex bimanual activities. However, it poses severa…
Knowledge distillation through geometry-aware representational alignment
Prajjwal Bhattarai, Mohammad Amjad, Dmytro Zhylko +1
Knowledge distillation is a common paradigm for transferring capabilities from larger models to smaller ones. While traditional distillation methods leverage a probabilistic diverg…
MoENAS: Mixture-of-Expert based Neural Architecture Search for jointly Accurate, Fair, and Robust Edge Deep Neural Networks
Lotfi Abdelkrim Mecharbat, Alberto Marchisio, Muhammad Shafique +2
There has been a surge in optimizing edge Deep Neural Networks (DNNs) for accuracy and efficiency using traditional optimization techniques such as pruning, and more recently, empl…
GLoG-CSUnet: Enhancing Vision Transformers with Adaptable Radiomic Features for Medical Image Segmentation
Niloufar Eghbali, Hassan Bagher-Ebadian, Tuka Alhanai +1
Vision Transformers (ViTs) have shown promise in medical image semantic segmentation (MISS) by capturing long-range correlations. However, ViTs often struggle to model local spatia…
An LSTM Feature Imitation Network for Hand Movement Recognition from sEMG Signals
Chuheng Wu, S. Farokh Atashzar, Mohammad M. Ghassemi +1
Surface Electromyography (sEMG) is a non-invasive signal that is used in the recognition of hand movement patterns, the diagnosis of diseases, and the robust control of prostheses.…