activity
20242026
collaborators

32 papers

cs.CV2026

GKDT: General Keypoint Detection Transformer

Changsheng Lu, Yuxin Chen, Haokun Gui +5

With the emergence of various pre-trained vision and language models, computer vision is shifting from narrow-domain to open-domain recognition. The construction of a more powerful…

physics.optics2026

Enhancing Speckle Metrology with Diffusion Denoising in Photon-Starved Regimes

Admir Bajraktarevic, Alexander C. Trowbridge, Saba N. Khan +3

Laser speckle is a powerful tool for precision metrology that enables highly sensitive measurements of light sources and subtle environmental perturbations. Many applications requi…

cs.CV2026

LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation

Bo Miao, Weijia Liu, Jun Luo +8

Language-conditioned goal navigation (LGN) requires agents to locate user-specified targets without step-by-step guidance. However, existing benchmarks largely focus on category-le…

cs.CL2026

Procedural Pretraining: Warming Up Language Models with Abstract Data

Liangze Jiang, Zachary Shinnick, Anton van den Hengel +2

Pretraining language models directly on web-scale corpora is the de facto paradigm. We study an alternative where the model is initially exposed to abstract structured data to ease…

cs.CV2026

Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformers

Zachary Shinnick, Liangze Jiang, Hemanth Saratchandran +2

Transformers are remarkably versatile, suggesting the existence of generic inductive biases beneficial across modalities. In this work, we explore a new way to instil such biases i…

cs.LG2026

I Predict Therefore I Am: Is Next Token Prediction Enough to Learn Human-Interpretable Concepts from Data?

Yuhang Liu, Dong Gong, Yichao Cai +6

Recent empirical evidence shows that LLM representations encode human-interpretable concepts. Nevertheless, the mechanisms by which these representations emerge remain largely unex…