3 papers
cs.CV2026
StrAD: A Streaming Method and Benchmark for Audio Description Generation for Long-form Videos
Julian Spravil, Sebastian Houben, Sven Behnke
Visual content is the dominant medium of communication, yet without audio descriptions (ADs), it remains inaccessible to blind and low-vision people. ADs narrate context-relevant v…
cs.CL2025
Scaling Laws for Conditional Emergence of Multilingual Image Captioning via Generalization from Translation
Julian Spravil, Sebastian Houben, Sven Behnke
Cross-lingual, cross-task transfer is challenged by task-specific data scarcity, which becomes more severe as language support grows and is further amplified in vision-language mod…
cs.CV2024
HyenaPixel: Global Image Context with Convolutions
Julian Spravil, Sebastian Houben, Sven Behnke
In computer vision, a larger effective receptive field (ERF) is associated with better performance. While attention natively supports global context, its quadratic complexity limit…