8 papers
WalkGPT: Grounded Vision-Language Conversation with Depth-Aware Segmentation for Pedestrian Navigation
Rafi Ibn Sultan, Hui Zhu, Xiangyu Zhou +4
Ensuring accessible pedestrian navigation requires reasoning about both semantic and spatial aspects of complex urban scenes, a challenge that existing Large Vision-Language Models…
Not All Tokens Are Meant to Be Forgotten
Xiangyu Zhou, Yao Qiang, Saleh Zare Zade +3
Large Language Models (LLMs), pre-trained on massive text corpora, exhibit remarkable human-level language understanding, reasoning, and decision-making abilities. However, they te…
GeoSAM: Fine-tuning SAM with Multi-Modal Prompts for Mobility Infrastructure Segmentation
Rafi Ibn Sultan, Chengyin Li, Hui Zhu +3
In geographical image segmentation, performance is often constrained by the limited availability of training data and a lack of generalizability, particularly for segmenting mobili…
Learning to Poison Large Language Models for Downstream Manipulation
Xiangyu Zhou, Yao Qiang, Saleh Zare Zade +4
The advent of Large Language Models (LLMs) has marked significant achievements in language processing and reasoning capabilities. Despite their advancements, LLMs face vulnerabilit…
Automatic Calibration for Membership Inference Attack on Large Language Models
Saleh Zare Zade, Yao Qiang, Xiangyu Zhou +4
Membership Inference Attacks (MIAs) have recently been employed to determine whether a specific text was part of the pre-training data of Large Language Models (LLMs). However, exi…
Interpretability-Aware Vision Transformer
Yao Qiang, Chengyin Li, Prashant Khanduri +1
Vision Transformers (ViTs) have become prominent models for solving various vision tasks. However, the interpretability of ViTs has not kept pace with their promising performance.…