2 papers
cs.CL2026
Pitfalls of Evaluating Language Models with Open Benchmarks
Md. Najib Hasan, Md Mahadi Hassan Sibat, Mohammad Fakhruddin Babar +3
Open Large Language Model (LLM) benchmarks, such as HELM and BIG-Bench, provide standardized and transparent evaluation protocols that support comparative analysis, reproducibility…
cs.DC2025
Investigating Timing-Based Information Leakage in Data Flow-Driven Real-Time Systems
Mohammad Fakhruddin Babar, Zain A. H. Hammadeh, Mohammad Hamad +1
Leaking information about the execution behavior of critical real-time tasks may lead to serious consequences, including violations of temporal constraints and even severe failures…