6 citations · 8 across the 7 of their papers we have counts for
3 papers · 1 filter
Can We Trust the Judges? Validation of Factuality Evaluation Methods via Answer Perturbation
Sarra Gharsallah, Adele Robaldo, Mariia Tokareva +5
Evaluating the factual correctness of large language models (LLMs) is vital for many applications. But are our evaluation tools themselves trustworthy? Despite the rise of factuali…
MedScribe: Clinically Grounded CT Reporting through Agentic Workflows
Giuseppe A. Orlando, Paolo Papotti, Maria A. Zuluaga +2
Vision-language models (VLMs) have shown potential for automated radiology report generation, yet existing approaches rely on global embedding compression of volumetric data, often…
Parallel Context-of-Experts Decoding for Retrieval Augmented Generation
Giulio Corallo, Paolo Papotti
Retrieval Augmented Generation faces a trade-off: concatenating documents in a long prompt enables multi-document reasoning but creates prefill bottlenecks, while encoding document…