3 papers
cs.CL2026
PragMatch: Separating Pragmatic Incongruity from Cross-Modal Mismatch in Large Vision-Language Models
Zhanna Mukhametsharip, Vera Demberg, Varsha Suresh
Large Vision-Language Models (LVLMs) have demonstrated strong performance on multimodal benchmarks, yet it remains unclear whether they genuinely reason about relationships between…
cs.CV2026
Symb-xMIL: Symbolic Explanations for Multiple Instance Learning in Digital Pathology
Yanqing Luo, Julius Hense, Niklas PreniÃl +4
Explanations of multiple instance learning (MIL) models are widely used for validation and discovery in digital histopathology. Existing methods primarily rely on heatmaps that hig…
cs.CV2026
Towards Reliable Human Evaluations in Gesture Generation: Insights from a Community-Driven State-of-the-Art Benchmark
Rajmund Nagy, Hendric Voss, Thanh Hoang-Minh +18
We review human evaluation practices in automatic, speech-driven 3D gesture generation and find a lack of standardisation and frequent use of flawed experimental setups. This leads…