1 paper
Mehdi Hosseinzadeh, King Hang Wong, Feras Dayoub
We present KITE, a training-free, keyframe-anchored, layout-grounded front-end that converts long robot-execution videos into compact, interpretable tokenized evidence for vision-l…