works on

From the 1 of 7 linked papers with an AI index.

collaborators

7 papers

cs.CV2026

ORCA: ORgan-Centroid Aggregation for Training-Free 3D CT Visual Token Compression

Renjie Liang, Zijian Xu, Jinqian Pan +6

A 3D CT scan entering a vision-language model produces a long sequence of visual tokens, often thousands to tens of thousands per volume, and this sequence must be compressed befor…

cs.CV2026

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles

Shawn Li, Wei Yang, Jike Zhong +11

The paper introduces JigShape, a benchmark of interlocking jigsaw puzzles designed to test visual‑geometric reasoning in vision‑language models, and shows that current zero‑shot an…

cs.RO2026

Artificial Foveated Perception for Mitigating Shortcut Learning in Robotic Foundation Models

Xiatao Sun, Yuan Zhuang, Mateo Sanchez Lopez Negrete +9

Robotic foundation models have recently made substantial progress in multi-task capability, cross-embodiment transfer, and language-conditioned control. Yet robust deployment acros…

cs.AI2026

Agent Safety Is Action Alignment

Shawn Li, Yue Zhao

Large language models increasingly act as agents: they call tools, move money, delete records, and send messages on a user's behalf. To keep them safe, practitioners imported the c…

cs.AI2026

FORTIS: Benchmarking Over-Privilege in Agent Skills

Shawn Li, Chenxiao Yu, Han Wang +8

Large language model agents increasingly operate through an intermediate skill layer that mediates between user intent and concrete task execution. This layer is widely treated as…

cs.AI2026

Geometry over Density: Few-Shot Cross-Domain OOD Detection

Shawn Li, You Qin, Jiate Li +4

Out-of-distribution (OOD) detection identifies test samples that fall outside a model's training distribution, a capability critical for safe deployment in high-stakes applications…