5 papers
What Counts as a Mistake? Annotating Recitation Events in Quran Memorization Transcripts
Mohamad Al Mdfaa, Nursultan Askarbekuly, Ahmed Helaly +2
Checking Quran recitation from an ASR transcript requires distinguishing unresolved mistakes from repetitions, repairs, opening formulas and accepted spelling differences. We repor…
Autoresearch with Coding Agents: Generalizers and Metric-Maximizers on Quran Recitation Data
Nursultan Askarbekuly, Mohamad Al Mdfaa, Ahmed Helaly +2
Coding agents can now be left alone to improve software against a score. In this pattern--recently popularized as "autoresearch"--the agent receives a dataset, an evaluation script…
VL-MemKnG: Hybrid Memory with a Spatio-Temporal Knowledge Graph for Question Answering over Long Egocentric Navigation Trajectories
Svetlana Lukina, Mohamad Al Mdfaa, Gloria Haro +2
Answering navigation-relevant questions over long egocentric videos requires retrieving and organizing evidence distributed across distant temporal moments while maintaining spatia…
Spatiotemporal Knowledge Graphs as Persistent Scene Memory for Embodied Question Answering
Mohamad Al Mdfaa, Svetlana Lukina, Timur Akhtyamov +4
Vision-language models (VLMs) demonstrate strong image-level scene understanding, but reasoning over long egocentric video remains costly: because VLMs maintain no persistent memor…
Mapping the Unseen: Unified Promptable Panoptic Mapping with Dynamic Labeling using Foundation Models
Mohamad Al Mdfaa, Raghad Salameh, Geesara Kulathunga +2
Panoptic maps enable robots to reason about both geometry and semantics. However, open-vocabulary models repeatedly produce closely related labels that split panoptic entities and…