#efficient inference

try —

14 papers match

cs.CV2026

Calibrate Before Reason: Robust Visual Token Reduction against Semantic Drift in VLMs

Jiasheng Li, Zhong Ji, Yan Zhang +1

The paper proposes CaRe, a training‑free method that calibrates compact visual representations before reasoning to keep semantic consistency when reducing visual tokens in large vi…

#visual token reduction#vision-language models#semantic drift#model calibration
cs.CV2026

Finding Change in Satellite Archives from Text: How to Combine Before-and-After Images Efficiently

Simon Roy, Mark Bong, Giovanni Beltrame

The paper studies how to efficiently combine before-and-after satellite images to match natural‑language change queries, comparing attention, state‑space (Mamba), and compressed fu…

#satellite imagery#change detection#text‑image retrieval#fusion architectures
cs.CV2026

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding

Mingkang Dong, Muxin Pu, Jie Li +8

ObjectStream introduces a training‑free method that extracts latent objects from frozen Video‑LLM representations and uses them as persistent memory anchors to improve streaming vi…

#video streaming#object-centric memory#large language models#latent objects
cs.AI2026

Penelope: Localized Latent Recurrence for Efficient Structured Reasoning

Yutong Chen, Shouqian Shi, Xinran Liu +5

Penelope introduces a method that adds a localized recurrent computation within a decoder-only Transformer to perform structured reasoning efficiently, using a latent space instead…

#structured reasoning#latent reasoning#decoder-only transformers#efficient inference
cs.RO2026

GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch

GigaWorld Team, Angen Ye, Angyuan Ma +26

The paper introduces GigaWorld-Policy-0.5, a robot control model that learns from future visual dynamics during training but generates actions only at inference, achieving faster (…

#robot control#world action models#action-conditioned world modeling#efficient inference
cs.CV2026

VideoSEMA: a scalable and efficient Mamba-like attention for video understanding

Nhat Thanh Tran, Fanghui Xue andShuai Zhang, Fanghui Xue +5

The paper introduces VideoSEMA, a split space‑time attention model for video classification that combines a scalable Mamba‑like spatial attention block with softmax temporal attent…

#video classification#attention mechanisms#mamba architecture#spatial-temporal modeling