5 papers
A Cross-Architecture Audit of Direction-Based Inference-Time Defences in Vision-Language Models
Xiangyu Yin, Tora Bodin, Rohan Menon +1
The paper evaluates five direction‑based inference‑time defenses for vision‑language models across multiple architectures, finding that no single method works best for all models a…
Spatial Reasoning in Multimodal Large Language Models: A Survey of Tasks, Benchmarks and Methods
Weichen Liu, Qiyao Xue, Haoming Wang +3
Spatial reasoning, which requires ability to perceive and manipulate spatial relationships in the 3D world, is a fundamental aspect of human intelligence, yet remains a persistent…
ProGait: A Multi-Purpose Video Dataset and Benchmark for Transfemoral Prosthesis Users
Xiangyu Yin, Boyuan Yang, Weichen Liu +4
Prosthetic legs play a pivotal role in clinical rehabilitation, allowing individuals with lower-limb amputations the ability to regain mobility and improve their quality of life. G…
Never Start from Scratch: Expediting On-Device LLM Personalization via Explainable Model Selection
Haoming Wang, Boyuan Yang, Xiangyu Yin +1
Personalization of Large Language Models (LLMs) is important in practical applications to accommodate the individual needs of different mobile users. Due to data privacy concerns,…
PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation
Qiyao Xue, Xiangyu Yin, Boyuan Yang +1
Text-to-video (T2V) generation has been recently enabled by transformer-based diffusion models, but current T2V models lack capabilities in adhering to the real-world common knowle…