collaborators

5 papers

cs.AI2026

A Cross-Architecture Audit of Direction-Based Inference-Time Defences in Vision-Language Models

Xiangyu Yin, Tora Bodin, Rohan Menon +1

The paper evaluates five direction‑based inference‑time defenses for vision‑language models across multiple architectures, finding that no single method works best for all models a…

cs.AI2025

Spatial Reasoning in Multimodal Large Language Models: A Survey of Tasks, Benchmarks and Methods

Weichen Liu, Qiyao Xue, Haoming Wang +3

Spatial reasoning, which requires ability to perceive and manipulate spatial relationships in the 3D world, is a fundamental aspect of human intelligence, yet remains a persistent…

cs.CV2025

ProGait: A Multi-Purpose Video Dataset and Benchmark for Transfemoral Prosthesis Users

Xiangyu Yin, Boyuan Yang, Weichen Liu +4

Prosthetic legs play a pivotal role in clinical rehabilitation, allowing individuals with lower-limb amputations the ability to regain mobility and improve their quality of life. G…

cs.LG2025

Never Start from Scratch: Expediting On-Device LLM Personalization via Explainable Model Selection

Haoming Wang, Boyuan Yang, Xiangyu Yin +1

Personalization of Large Language Models (LLMs) is important in practical applications to accommodate the individual needs of different mobile users. Due to data privacy concerns,…

cs.CV2025

PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation

Qiyao Xue, Xiangyu Yin, Boyuan Yang +1

Text-to-video (T2V) generation has been recently enabled by transformer-based diffusion models, but current T2V models lack capabilities in adhering to the real-world common knowle…