3 papers
cs.CY2025
Leveraging LLMs to Streamline the Review of Public Funding Applications
Joao D. S. Marques, Andre V. Duarte, Andre Carvalho +3
Every year, the European Union and its member states allocate millions of euros to fund various development initiatives. However, the increasing number of applications received for…
cs.CL2025
RECAP: Reproducing Copyrighted Data from LLMs Training with an Agentic Pipeline
André V. Duarte, Xuying li, Bin Zeng +3
If we cannot inspect the training data of a large language model (LLM), how can we ever know what it has seen? We believe the most compelling evidence arises when the model itself…
cs.CV2025
DIS-CO: Discovering Copyrighted Content in VLMs Training Data
André V. Duarte, Xuandong Zhao, Arlindo L. Oliveira +1
How can we verify whether copyrighted content was used to train a large vision-language model (VLM) without direct access to its training data? Motivated by the hypothesis that a V…