collaborators

17 papers

cs.CV2026

CoCo-IR: Contextual Composed Image Retrieval

Shengcao Cao, Tanmaya Shekhar Dabral, Zhongli Ding +6

Current instruction-based image retrieval systems are powerful but limited to single-turn interactions, failing to capture the iterative nature of complex, real-world visual search…

cs.CV2026

PPTArena: A Benchmark for PowerPoint Editing

Michael Ofengenden, Yunze Man, Ziqi Pang +2

We introduce PPTArena, a benchmark for PowerPoint editing that evaluates how agents modify real slides from natural-language instructions. Unlike benchmarks that rely on image-PDF…

cs.HC2026

A Two-Validator Web Interface for Structured Geometry Figure Annotation

Sabin-Codrut Badea, Adrian-Marius Dumitran

Annotating geometric figures from scanned documents has long been addressed by adapting generic annotation tools, tools not originally designed for such tasks, to use cases where t…

cs.RO2026

Vesta: A Generalist Embodied Reasoning Model

Johan Bjorck, Zhiqi Li, Yunze Man +29

Robots operating in open-world environments must seamlessly integrate localization, spatial reasoning, navigation, and long-horizon planning. While specialist models excel at indiv…

cs.CV2026

LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight

Yunze Man, Shihao Wang, Guowen Zhang +7

To act in the world, a model must name what it sees and know where it is in 3D. Today's vision-language models (VLMs) excel at open-ended 2D description and grounding, yet multi-ob…

cs.AI2026

BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger Learning

Qiusi Zhan, Hyeonjeong Ha, Rui Yang +7

Recent advances in Vision-Language Models (VLMs) have propelled embodied agents by enabling direct perception, reasoning, and planning task-oriented actions from visual inputs. How…