10 papers
Separating Clicks from Baits: Using Large Language Models to Detect Misleading YouTube Thumbnails
Wajiha Naveed, Muhammad Muneeb Pervez, Zaeem Mohtashim Khan +2
Misleading video thumbnails on platforms like YouTube are a pervasive problem, undermining user trust and platform integrity. This paper proposes a novel multi-modal detection pipe…
Efficient and Adaptable Detection of Malicious LLM Prompts via Bootstrap Aggregation
Shayan Ali Hassan, Tao Ni, Zafar Ayyub Qazi +1
Large Language Models (LLMs) have demonstrated remarkable capabilities in natural language understanding, reasoning, and generation. However, these systems remain susceptible to ma…
MAESTRO: Multi-Agent Evaluation Suite for Testing, Reliability, and Observability
Tie Ma, Yixi Chen, Vaastav Anand +8
We present MAESTRO, an evaluation suite for the testing, reliability, and observability of LLM-based MAS. MAESTRO standardizes MAS configuration and execution through a unified int…
Toward an AI-Native Internet: Rethinking the Web Architecture for Semantic Retrieval
Muhammad Bilal, Zafar Qazi, Marco Canini
The rise of Generative AI Search is fundamentally transforming how users and intelligent systems interact with the Internet. LLMs increasingly act as intermediaries between humans…
DMAS-Forge: A Framework for Transparent Deployment of AI Applications as Distributed Systems
Alessandro Cornacchia, Vaastav Anand, Muhammad Bilal +2
Agentic AI applications increasingly rely on multiple agents with distinct roles, specialized tools, and access to memory layers to solve complex tasks -- closely resembling servic…
Scaling Truth: The Confidence Paradox in AI Fact-Checking
Ihsan A. Qazi, Zohaib Khan, Abdullah Ghani +7
The rise of misinformation underscores the need for scalable and reliable fact-checking solutions. Large language models (LLMs) hold promise in automating fact verification, yet th…