4 papers
Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling
Leonardo Liparulo, Francesco Pierri
We ask whether AI agents powered by locally deployed large language models can reliably automate expert-defined hardware design workflows in an industry-realistic tool-calling sett…
Auditing Exposure to Harmful Content on TikTok using Multimodal Language Models: A Cross-National, Age-Stratified Study
Hamidreza Saffari, Francesco Pierri
Online video platforms can expose young users to harmful content, but independent audits remain difficult because video annotation is costly and moderation judgments vary across la…
Towards an Automated Framework to Audit Youth Safety on TikTok
Linda Xue, Francesco Corso, Nicolo' Fontana +3
This paper investigates the effectiveness of TikTok's enforcement mechanisms for limiting the exposure of harmful content to youth accounts. We collect over 7000 videos, classify t…
Analyzing the Safety of Japanese Large Language Models in Stereotype-Triggering Prompts
Akito Nakanishi, Yukie Sano, Geng Liu +1
In recent years, Large Language Models have attracted growing interest for their significant potential, though concerns have rapidly emerged regarding unsafe behaviors stemming fro…