3 papers
cs.CV2026
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention
Daniel Shalam, Emanuel Ben Baruch, Avi Ben Cohen +1
Multimodal large language models can emit localized predictions, bounding boxes for objects and temporal windows for video and audio events, but they hallucinate these regions prol…
cs.CV2026
TrajLoc: Trajectory-Attention Localization for Multi-Object Motion Control
Omer Sela, Inbar Huberman-Spiegelglas, Michael Rotman +2
Controlling the motion of multiple objects in image-to-video (I2V) generation requires preserving object identities while enforcing adherence to distinct target trajectories. This…
cs.CV2026
Planar-SfM: Camera Pose Estimation via Homography Graph Embeddings
Gabi Pragier, Matan Karklinsky, David Ungarish +1
Structure from Motion (SfM) systems traditionally struggle with planar scenes, where standard epipolar geometry-based methods become degenerate. Rather than viewing planar surfaces…