Publications (9)
SpidR: Learning Fast and Stable Linguistic Units for Spoken Language Models Without Supervision
Maxime Poli, Mahi Luthra, Youssef Benchekroun +8
The parallel advances in language modeling and speech representation learning have raised the prospect of learning language directly from speech without textual intermediates. This…
Identification of Tasks, Datasets, Evaluation Metrics, and Numeric Scores for Scientific Leaderboards Construction
Yufang Hou, Charles Jochim, Martin Gleize +2
While the fast-paced inception of novel tasks and new datasets helps foster active research in a community towards interesting directions, keeping track of the abundance of researc…
Seamless Interaction: Dyadic Audiovisual Motion Modeling and Large-Scale Dataset
Vasu Agrawal, Akinniyi Akinyemi, Kathryn Alvero +81
Human communication involves a complex interplay of verbal and nonverbal signals, essential for conveying meaning and achieving interpersonal goals. To develop socially intelligent…
Corpus Wide Argument Mining -- a Working Solution
Liat Ein-Dor, Eyal Shnarch, Lena Dankin +10
One of the main tasks in argument mining is the retrieval of argumentative content pertaining to a given topic. Most previous work addressed this task by retrieving a relatively sm…
Are You Convinced? Choosing the More Convincing Evidence with a Siamese Network
Martin Gleize, Eyal Shnarch, Leshem Choshen +4
With the advancement in argument detection, we suggest to pay more attention to the challenging task of identifying the more convincing arguments. Machines capable of responding an…
EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data
Dongyan Lin, Phillip Rust, Angel Villar Corrales +19
Children acquire language grounding with remarkable robustness from limited visuo-linguistic input in ways that surpass today's best large multimodal models. Recent research sugges…
TDMSci: A Specialized Corpus for Scientific Literature Entity Tagging of Tasks Datasets and Metrics
Yufang Hou, Charles Jochim, Martin Gleize +2
Tasks, Datasets and Evaluation Metrics are important concepts for understanding experimental scientific papers. However, most previous work on information extraction for scientific…
SpidR-Adapt: A Universal Speech Representation Model for Few-Shot Adaptation
Mahi Luthra, Jiayi Shen, Maxime Poli +14
Human infants, with only a few hundred hours of speech exposure, acquire basic units of new languages, highlighting a striking efficiency gap compared to the data-hungry self-super…
A Summarization System for Scientific Documents
Shai Erera, Michal Shmueli-Scheuer, Guy Feigenblat +15
We present a novel system providing summaries for Computer Science publications. Through a qualitative user study, we identified the most valuable scenarios for discovery, explorat…