5 papers
Deriving Benchmarking Datasets from Long-Form Recordings: Challenges and Opportunities
Kaveri K. Sheth, Lawrence Borst, Tarek Kunze +6
Long-form recordings (LFRs) of child-centered audio are ecologically valid sources for studying early language development, but three problems limit their use. First, LFR corpora a…
BabyHuBERT: Multilingual Self-Supervised Learning for Segmenting Speakers in Child-Centered Long-Form Recordings
Théo Charlot, Tarek Kunze, Maxime Poli +3
Child-centered daylong recordings are essential for studying early language development, but existing speech models trained on clean adult data perform poorly due to acoustic and l…
Context-aware child-directed speech detection from long-form recordings
Théo Charlot, Tarek Kunze, Kaveri K. Sheth +2
Automatically distinguishing child-directed speech from adult-directed speech in long-form recordings is key to scalable analyses of children's language environments. Existing appr…
Fifteen Years of Child-Centered Long-Form Recordings: Promises, Resources, and Remaining Challenges to Validity
Loann Peurey, Marvin Lavechin, Tarek Kunze +4
Audio-recordings collected with a child-worn device are a fundamental tool in child language research. Long-form recordings collected over whole days promise to capture children's…
Challenges in Automated Processing of Speech from Child Wearables: The Case of Voice Type Classifier
Tarek Kunze, Marianne Métais, Hadrien Titeux +5
Recordings gathered with child-worn devices promised to revolutionize both fundamental and applied speech sciences by allowing the effortless capture of children's naturalistic spe…