3 papers
cs.AR2026
The Price of Anarchy in Disaggregated Inference
Athos Georgiou
Disaggregated inference architectures physically separate prefill and decode phases onto distinct GPU pools, creating competing "agents" that share a fixed hardware budget. We prov…
cs.CV2026
Hydra: Unifying Document Retrieval and Generation in a Single Vision-Language Model
Athos Georgiou
Visual document understanding typically requires separate retrieval and generation models, doubling memory and system complexity. We present Hydra, a dual-head approach that provid…
cs.AR2026
Architecture-Aware LLM Inference Optimization on AMD Instinct GPUs: A Comprehensive Benchmark and Deployment Study
Athos Georgiou
We present a cross-architecture evaluation of production LLM inference on AMD Instinct MI325X GPUs, benchmarking four models spanning 235B to 1 trillion parameters across three arc…