4 papers · 1 filter
Improving Temporal Action Segmentation via Constraint-Aware Decoding
Yeo Keat Ee, Debaditya Roy, Chen Li +2
Temporal action segmentation (TAS) divides untrimmed videos into labeled action segments. While fully supervised methods have advanced the field, challenges such as action variabil…
Mitigating Easy Option Bias in Multiple-Choice Question Answering
Hao Zhang, Chen Li, Basura Fernando
In this early study, we observe an Easy-Options Bias (EOB) issue in some multiple-choice Visual Question Answering (VQA) benchmarks such as MMStar, RealWorldQA, SEED-Bench, Next-QA…
Neuro Symbolic Knowledge Reasoning for Procedural Video Question Answering
Basura Fernando, Thanh-Son Nguyen, Hong Yang +3
In this work we present Knowledge Module Learning (KML) to understand and reason over procedural tasks that requires models to learn structured and compositional procedural knowled…
Training-Free Action Recognition and Goal Inference with Dynamic Frame Selection
Ee Yeo Keat, Zhang Hao, Alexander Matyasko +1
We introduce VidTFS, a Training-free, open-vocabulary video goal and action inference framework that combines the frozen vision foundational model (VFM) and large language model (L…