4 papers · 1 filter
Beyond Boxes: Mask-Guided Spatio-Temporal Feature Aggregation for Video Object Detection
Khurram Azeem Hashmi, Talha Uddin Sheikh, Didier Stricker +1
The primary challenge in Video Object Detection (VOD) is effectively exploiting temporal information to enhance object representations. Traditional strategies, such as aggregating…
Text2CAD: Generating Sequential CAD Models from Beginner-to-Expert Level Text Prompts
Mohammad Sadil Khan, Sankalp Sinha, Talha Uddin Sheikh +3
Prototyping complex computer-aided design (CAD) models in modern softwares can be very time-consuming. This is due to the lack of intelligent systems that can quickly generate simp…
UnSupDLA: Towards Unsupervised Document Layout Analysis
Talha Uddin Sheikh, Tahira Shehzadi, Khurram Azeem Hashmi +2
Document layout analysis is a key area in document research, involving techniques like text mining and visual analysis. Despite various methods developed to tackle layout analysis,…
CICA: Content-Injected Contrastive Alignment for Zero-Shot Document Image Classification
Sankalp Sinha, Muhammad Saif Ullah Khan, Talha Uddin Sheikh +2
Zero-shot learning has been extensively investigated in the broader field of visual recognition, attracting significant interest recently. However, the current work on zero-shot le…