collaborators

6 papers

cs.SI2026

Form Without Function: Agent Social Behavior in the Moltbook Network

Saber Zerhoudi, Kanishka Ghosh Dastidar, Felix Klement +9

Moltbook is a social network where every participant is an AI agent. We analyze 1,312,238 posts, 6.7~million comments, and over 120,000 agent profiles across 5,400 communities, col…

cs.IR2026

WebFAQ 2.0: A Multilingual QA Dataset with Mined Hard Negatives for Dense Retrieval

Michael Dinzinger, Laura Caspari, Ali Salman +3

We introduce WebFAQ 2.0, a new version of the WebFAQ dataset, containing 198 million FAQ-based natural question-answer pairs across 108 languages. Compared to the previous version,…

cs.HC2026

OwlerLite: Scope- and Freshness-Aware Web Retrieval for LLM Assistants

Saber Zerhoudi, Michael Dinzinger, Michael Granitzer +1

Browser-based language models often use retrieval-augmented generation (RAG) but typically rely on fixed, outdated indices that give users no control over which sources are consult…

cs.IR2026

CoRECT: A Framework for Evaluating Embedding Compression Techniques at Scale

L. Caspari, M. Dinzinger, K. Ghosh Dastidar +3

Dense retrieval systems have proven to be effective across various benchmarks, but require substantial memory to store large search indices. Recent advances in embedding compressio…

cs.LG2025

Compressed Concatenation of Small Embedding Models

Mohamed Ayoub Ben Ayad, Michael Dinzinger, Kanishka Ghosh Dastidar +2

Embedding models are central to dense retrieval, semantic search, and recommendation systems, but their size often makes them impractical to deploy in resource-constrained environm…

cs.CL2025

WebFAQ: A Multilingual Collection of Natural Q&A Datasets for Dense Retrieval

Michael Dinzinger, Laura Caspari, Kanishka Ghosh Dastidar +2

We present WebFAQ, a large-scale collection of open-domain question answering datasets derived from FAQ-style schema.org annotations. In total, the data collection consists of 96 m…