19 papers
On the Robustness of Temporal Vision-Language Models for Surgical Endoscopy Videos
Darakshan Rashid, Raza Imam, Ufaq Khan +9
Temporal vision-language models (TVLMs) offer a reusable, prompt-based interface for surgical video understanding, yet, their robustness under clinically realistic acquisition arti…
Evaluating LLM Personalization via Semantic Constraint Verification
Xuran Li, Guanqin Zhang, Imran Razzak +4
Current evaluation paradigms for Large Language Model (LLM) personalization rely heavily on brittle surface-matching metrics or computationally expensive LLM-as-a-judge protocols,…
Knowledge Graph Enhanced Memory-Augmented Retrieval for Long Context Modeling
Ghadir Alselwi, Basem Suleiman, Hao Xue +4
Long-context language modeling requires not only extending context windows but maintaining coherent understanding of entity states and relationships across thousands of tokens -- a…
DocAtlas: Multilingual Document Understanding Across 80+ Languages
Ahmed Heakl, Youssef Mohamed, Abdullah Sohail +6
Multilingual document understanding remains limited for low-resource languages due to scarce training data and model-based annotation pipelines that perpetuate existing biases. We…
AMO: Adaptive Muon Orthogonalization
Xinlin Zhuang, Panyi Ouyang, Yichen Li +7
Muon has recently emerged as a competitive alternative to AdamW for large-scale pre-training, with orthogonalization via Newton-Schulz (NS) iterations as its core operation. Existi…
TriALS: Triphasic-Aided Liver Lesion Segmentation Benchmark in Non-Contrast CT
Marawan Elbatel, Mohamed Ghonim, Jiaji Mao +62
Automated segmentation of liver lesions on non-contrast computed tomography (NCCT) is clinically important but fundamentally challenging, particularly in low-resource settings acro…