4 papers
CHiPS: Character Histograms and Positional Signals for Lightweight Authorship Attribution in Romanian Texts
Sanda-Maria Avram, George C. Ţurcaş
We propose CHiPS, a lightweight character-level authorship attribution method for Romanian texts. All reported experiments are closed-set: the true author is one of the candidate a…
Language Identification via Compositional Data Analysis: A Linear-Time Classifier Based on Log-Ratio Geometry
Paul-Andrei Pogăcean, Sanda-Maria Avram
Language identification is commonly addressed using either neural architectures or statistical n-gram models. Neural approaches typically require substantial computational resource…
Language Detection by Means of the Minkowski Norm: Identification Through Character Bigrams and Frequency Analysis
Paul-Andrei Pogăcean, Sanda-Maria Avram
The debate surrounding language identification has gained renewed attention in recent years, especially with the rapid evolution of AI-powered language models. However, the non-AI-…
Oldies but Goldies: The Potential of Character N-grams for Romanian Texts
Dana Lupsa, Sanda-Maria Avram, Radu Lupsa
This study addresses the problem of authorship attribution for Romanian texts using the ROST corpus, a standard benchmark in the field. We systematically evaluate six machine learn…