2 papers
cs.CV2026
Making Video Models Adhere to User Intent with Minor Adjustments
Daniel Ajisafe, Eric Hedlin, Helge Rhodin +1
With the recent drastic advancements in text-to-video diffusion models, controlling their generations has drawn interest. A popular way for control is through bounding boxes or lay…
eess.AS2025
Towards End-to-End Training of Automatic Speech Recognition for Nigerian Pidgin
Amina Mardiyyah Rufai, Afolabi Abeeb, Esther Oduntan +3
The prevalence of automatic speech recognition (ASR) systems in spoken language applications has increased significantly in recent years. Notably, many African languages lack suffi…