13 papers
Crayotter: Learning Long-Horizon Video Editing Agents via Group-Relative Preference Backpropagation
Lecheng Yan, Jianze Lin, Yichong Zhang +5
Long-horizon video editing agents receive final-product feedback only after many interdependent decisions. Yet editing quality is subjective, admits multiple valid solutions, and i…
Decoding EEG Signals to Explore Next-Word Predictability in the Human Brain
Boi Mai Quach, Binh T. Nguyen, Cathal Gurrin +1
Humans invented reading and have passed down this complex skill across generations through language. This study provides empirical evidence of the neural mechanisms underlying bott…
Encoding EEG Signals to Examine Human-Like Next-Word Prediction Behaviour in Language Models
Boi Mai Quach, Binh T. Nguyen, Cathal Gurrin +1
Language models (LMs) are trained to excel at predicting the next word in the sequence given prior context, and humans also share this predictability in reading comprehension. Neur…
HANCLIP: A Family of Hyperbolic Angular Negation Vision Language Models
Hoang-Bao Le, Aiden Durrant, Thai Son Mai +3
Vision-Language Models (VLMs) are typically pre-trained on large-scale image-text datasets to capture semantic correspondences between visual content and natural language. However,…
OpenLifelogQA: An Open-Ended Multi-Modal Lifelog Question-Answering Dataset
Quang-Linh Tran, Hoang-Bao Le, Tuong-Nghiem Diep +3
We introduce OpenLifelogQA, a large-scale open-ended lifelog QA dataset constructed from 18 months of multimodal lifelog data. Lifelogging is the passive collection and analysis of…
The State-of-the-Art in Lifelog Retrieval: A Review of Progress at the ACM Lifelog Search Challenge Workshop 2022-24
Allie Tran, Werner Bailer, Duc-Tien Dang-Nguyen +8
The ACM Lifelog Search Challenge (LSC) is a venue that welcomes and compares systems that support the exploration of lifelog data, and in particular the retrieval of specific infor…