22 papers
MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents
Belal S. Alsinglawi, Weizheng Wang, Junyi Wu +4
Smart-city airspace is transforming Uncrewed Aerial Vehicles (UAVs) from passive sensing platforms into cyber-physical decision makers that must follow operational rules under degr…
ProjQ: Project-and-Quantize for Adapter-Aware LLM Compression
Wenya Yu, Chao Zhang, Li Wang +2
Post-Training Quantization (PTQ) and Low-Rank Adaptation (LoRA) constitute the standard pipeline for efficient Large Language Model (LLM) deployment. However, applying them sequent…
Telecom World Models: Unifying Digital Twins, Foundation Models, and Predictive Planning for 6G
Hang Zou, Yuzhi Yang, Lina Bariah +15
The integration of machine learning tools into telecom networks, has led to two prevailing paradigms, namely, language-based systems, such as Large Language Models (LLMs), and phys…
TelcoAgent-Bench: A Multilingual Benchmark for Telecom AI Agents
Lina Bariah, Brahim Mefgouda, Farbod Tavakkoli +3
The integration of large language model (LLM) agents into telecom networks introduces new challenges, related to intent recognition, tool execution, and resolution generation, whil…
From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review
Mohamed Amine Ferrag, Norbert Tihanyi, Merouane Debbah
Large language models and autonomous AI agents have evolved rapidly, resulting in a diverse array of evaluation benchmarks, frameworks, and collaboration protocols. Driven by the g…
6G-Bench: An Open Benchmark for Semantic Communication and Network-Level Reasoning with Foundation Models in AI-Native 6G Networks
Mohamed Amine Ferrag, Abderrahmane Lakas, Merouane Debbah
This paper introduces 6G-Bench, an open benchmark for evaluating semantic communication and network-level reasoning in AI-native 6G networks. 6G-Bench defines a taxonomy of 30 deci…