Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs
Andrea Sassella, Andrea Chizzola, Tommaso Bianchi +2
This report benchmarks the performance of ENGINEERING Ingegneria Informatica S.p.A.'s EngGPT2MoE-16B-A3B LLM, a 16B parameter Mixture of Experts (MoE) model with 3B active paramete…
cs.CL2025
L1RA: Dynamic Rank Assignment in LoRA Fine-Tuning
Raul Singh, Nicolo Brunello, Vincenzo Scotti +1
The ability of Large Language Models (LLMs) to solve complex tasks has made them crucial in the development of AI-based applications. However, the high computational requirements t…
cs.CL2025
InTraVisTo: Inside Transformer Visualisation Tool
Nicolò Brunello, Davide Rigamonti, Andrea Sassella +2
The reasoning capabilities of Large Language Models (LLMs) have increased greatly over the last few years, as have their size and complexity. Nonetheless, the use of LLMs in produc…