3 papers
cs.CV2025
GazeNLQ @ Ego4D Natural Language Queries Challenge 2025
Wei-Cheng Lin, Chih-Ming Lien, Chen Lo +1
This report presents our solution to the Ego4D Natural Language Queries (NLQ) Challenge at CVPR 2025. Egocentric video captures the scene from the wearer's perspective, where gaze…
eess.AS2025
Mind the Prompt: Prompting Strategies in Audio Generations for Improving Sound Classification
Francesca Ronchini, Ho-Hsiang Wu, Wei-Cheng Lin +1
This paper investigates the design of effective prompt strategies for generating realistic datasets using Text-To-Audio (TTA) models. We also analyze different techniques for effic…
eess.AS2024
Can Synthetic Data Boost the Training of Deep Acoustic Vehicle Counting Networks?
Stefano Damiano, Luca Bondi, Shabnam Ghaffarzadegan +2
In the design of traffic monitoring solutions for optimizing the urban mobility infrastructure, acoustic vehicle counting models have received attention due to their cost effective…