collaborators

6 papers

cs.LG2026

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms

Samuel Fernández-Menduiña, Amir Ziashahabi, Eduardo Pavez +2

Long-context LLM decoding reads the key-value (KV) cache at every step. Loading it takes longer than computing attention over it, so throughput is bandwidth-bound. Hence, reducing…

cs.CL2026

ORBIT: Training-Free Multi-Attribute Behavioral Steering via Orthogonal Subspace Rotation

Narges Ghasemi, Amir Ziashahabi, Salman Avestimehr +1

Language models are widely used in assistant settings, where controlling behavioral attributes is often essential. Activation steering modifies hidden-state representations at infe…

cs.CV2025

GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction

Narges Ghasemi, Amir Ziashahabi, Salman Avestimehr +1

Image geolocalization, the task of determining an image's geographic origin, poses significant challenges, largely due to visual similarities across disparate locations and the lar…

cs.LG2025

Reject Only Critical Tokens: Pivot-Aware Speculative Decoding

Amir Ziashahabi, Yavuz Faruk Bakman, Duygu Nur Yaldiz +3

Speculative Decoding (SD) ensures that the output matches the target model's distribution exactly. However, we argue that this distribution matching requirement is too stringent an…

cs.CV2025

OSMGen: Highly Controllable Satellite Image Synthesis using OpenStreetMap Data

Amir Ziashahabi, Narges Ghasemi, Sajjad Shahabi +3

Accurate and up-to-date geospatial data are essential for urban planning, infrastructure monitoring, and environmental management. Yet, automating urban monitoring remains difficul…

cs.LG2025

MobiZO: Enabling Efficient LLM Fine-Tuning at the Edge via Inference Engines

Lei Gao, Amir Ziashahabi, Yue Niu +2

Large Language Models (LLMs) are currently pre-trained and fine-tuned on large cloud servers. The next frontier is LLM personalization, where a foundation model can be fine-tuned w…