10 papers
Cooking beyond Frames: A Stereo Event Camera Dataset in the Kitchen
Chengming Feng, Hesam Araghi, Liming Zheng +4
Event cameras, also known as neuromorphic cameras, have gained significant attention in recent years due to their high temporal resolution, high dynamic range, and low power consum…
On the Real-World Generalisability of Optical Flow Models
Petter Reijalt, Sander Gielisse, Rickard Karlsson +1
Real-world deployment of vision models to broadly benefit society is arguably a main research objective. In optical flow, however, the difficulty to obtain the ground truth has foc…
Fine-grained Human Motion Understanding with Language Models
Thomas Markhorst, Zhi-Yi Lin, Jouh Yeong Chew +2
In this work, we propose \methodname, an LLM-based model for fine-grained human motion understanding that represents motion as a sequence of skeletal poses with explicit timestamps…
MuPPet: Multi-person 2D-to-3D Pose Lifting
Thomas Markhorst, Zhi-Yi Lin, Jouh Yeong Chew +2
Multi-person social interactions are inherently built on coherence and relationships among all individuals within the group, making multi-person localization and body pose estimati…
Assessing Situational and Spatial Awareness of VLMs with Synthetically Generated Video
Pascal Benschop, Justin Dauwels, Jan van Gemert
Spatial reasoning in vision language models (VLMs) remains fragile when semantics hinge on subtle temporal or geometric cues. We introduce a synthetic benchmark that probes two com…
LayoutGKN: Graph Similarity Learning of Floor Plans
Casper van Engelenburg, Jan van Gemert, Seyran Khademi
Floor plans depict building layouts and are often represented as graphs to capture the underlying spatial relationships. Comparison of these graphs is critical for applications lik…