1 citations · 2 across the 5 of their papers we have counts for
6 papers
IDIOLEX: Unified and Continuous Representations for Idiolectal and Stylistic Variation
Anjali Kantharuban, Aarohi Srivastava, Fahim Faisal +5
Existing sentence representations primarily encode what a sentence says, rather than how it is expressed, even though the latter is important for many applications. In contrast, we…
BASS: Benchmarking Audio LMs for Musical Structure and Semantic Reasoning
Min Jang, Orevaoghene Ahia, Nazif Tamer +3
Music understanding is a complex task that often requires reasoning over both structural and semantic elements of audio. We introduce BASS, designed to evaluate music understanding…
Cognitive Foundations for Reasoning and Their Manifestation in LLMs
Priyanka Kargupta, Shuyue Stella Li, Haocheng Wang +9
Large language models (LLMs) solve complex problems yet fail on simpler variants, suggesting they achieve correct outputs through mechanisms fundamentally different from human reas…
Broken Tokens? Your Language Model can Secretly Handle Non-Canonical Tokenizations
Brian Siyuan Zheng, Alisa Liu, Orevaoghene Ahia +3
Modern tokenizers employ deterministic algorithms to map text into a single "canonical" token sequence, yet the same string can be encoded as many non-canonical tokenizations using…
BLAB: Brutally Long Audio Bench
Orevaoghene Ahia, Martijn Bartelds, Kabir Ahuja +13
Developing large audio language models (LMs) capable of understanding diverse spoken interactions is essential for accommodating the multimodal nature of human communication and ca…
AfriWOZ: Corpus for Exploiting Cross-Lingual Transferability for Generation of Dialogues in Low-Resource, African Languages
Tosin Adewumi, Mofetoluwa Adeyemi, Aremu Anuoluwapo +17
Dialogue generation is an important NLP task fraught with many challenges. The challenges become more daunting for low-resource African languages. To enable the creation of dialogu…