4 papers
Sem-Detect: Semantic Level Detection of AI Generated Peer-Reviews
André V. Duarte, Brian Tufts, Aditya Oke +3
How can we distinguish whether a peer review was written by a human or generated by an AI model? We argue that, in this setting, authorship should not be attributed solely from the…
RECAP: Reproducing Copyrighted Data from LLMs Training with an Agentic Pipeline
André V. Duarte, Xuying li, Bin Zeng +3
If we cannot inspect the training data of a large language model (LLM), how can we ever know what it has seen? We believe the most compelling evidence arises when the model itself…
Leveraging LLMs to Streamline the Review of Public Funding Applications
Joao D. S. Marques, Andre V. Duarte, Andre Carvalho +3
Every year, the European Union and its member states allocate millions of euros to fund various development initiatives. However, the increasing number of applications received for…
DIS-CO: Discovering Copyrighted Content in VLMs Training Data
André V. Duarte, Xuandong Zhao, Arlindo L. Oliveira +1
How can we verify whether copyrighted content was used to train a large vision-language model (VLM) without direct access to its training data? Motivated by the hypothesis that a V…