Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Multi-Record Web Page Information Extraction From News Websites
Alexander Kustenkov, Maksim Varlamov, Alexander Yatskov
In this paper, we focused on the problem of extracting information from web pages containing many records, a task of growing importance in the era of massive web data. Recently, th…
cs.CL2025
Multilingual Attribute Extraction from News Web Pages
Pavel Bedrin, Maksim Varlamov, Alexander Yatskov
This paper addresses the challenge of automatically extracting attributes from news article web pages across multiple languages. Recent neural network models have shown high effica…