output
20172025
most citedArabic natural language processing: An overview

192 citations

11 papers

cs.DL2025

From data to corpus: semiotic and documentary issues in audiovisual archives

Peter Stockinger

The article examines the theoretical, methodological, and technical foundations of research on audiovisual corpora within the field of digital humanities. It outlines the main tran…

cs.DL2025

Animer une base de connaissance: des ontologies aux mod{è}les d'I.A. g{é}n{é}rative

Peter Stockinger

In a context where the social sciences and humanities are experimenting with non-anthropocentric analytical frames, this article proposes a semiotic (structural) reading of the hyb…

cs.SD2025

Assessing the Impact of Anisotropy in Neural Representations of Speech: A Case Study on Keyword Spotting

Guillaume Wisniewski, Séverine Guillaume, Clara Rosina Fernández

Pretrained speech representations like wav2vec2 and HuBERT exhibit strong anisotropy, leading to high similarity between random embeddings. While widely observed, the impact of thi…

cs.CL2024★ 1 cited

From communities to interpretable network and word embedding: an unified approach

Thibault Prouteau, Nicolas Dugué, Simon Guillot

Modelling information from complex systems such as humans social interaction or words co-occurrences in our languages can help to understand how these systems are organized and fun…

physics.soc-ph2024

The role of spatial structures and social values in shaping local productive systems -New lessons from the wood-furniture cluster of Jepara, Indonesia

Julien Birgi

This paper revisits the well-known wood-furniture cluster of Jepara (Central Java, Indonesia) with new parameters inspired by the theories of industrial districts and clusters. So…

cs.CL2024★ 1 cited

Establishing degrees of closeness between audio recordings along different dimensions using large-scale cross-lingual models

Maxime Fily, Guillaume Wisniewski, Severine Guillaume +2

In the highly constrained context of low-resource language studies, we explore vector representations of speech from a pretrained model to determine their level of abstraction with…