11 papers
TopoTuner: Topological Finetuning of Large Language Models
Abdulkadir Erol, Yash Mahajan, Vepaul Hariprashad +4
Full fine-tuning remains a strong way to adapt pretrained LLMs, but it updates all weights and can be expensive. LoRA reduces the number of trainable parameters, but it does not di…
CompLLM: Compression for Long Context Q&A
Gabriele Berton, Jayakrishnan Unnikrishnan, Son Tran +1
Large Language Models (LLMs) face significant computational challenges when processing long contexts due to the quadratic complexity of self-attention. While soft context compressi…
Weakly-Supervised Spatiotemporal Anomaly Detection
Urvi Gianchandani, Praveen Tirupattur, Mubarak Shah
In this paper, we explore a weakly supervised method for anomaly detection. Since annotating videos is time-consuming, we only look at weak video-level labels during training. This…
Dystruct: Dynamically Structured Diffusion Language Model Decoding via Bayesian Inference
Bian Sun, Kevin Zhai, Mubarak Shah +1
Diffusion language models (DLMs) have recently emerged as a promising alternative to autoregressive models, primarily due to their ability to enable parallel decoding. Despite this…
Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach
Utsav Singh, Souradip Chakraborty, Wesley A. Suttle +6
Hierarchical reinforcement learning (HRL) enables agents to solve complex, long-horizon tasks by decomposing them into manageable sub-tasks. However, HRL methods face two fundament…
VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale
Parth Parag Kulkarni, Rohit Gupta, Prakash Chandra Chhipa +1
The task of video geolocalization aims to determine the precise GPS coordinates of a video's origin and map its trajectory; with applications in forensics, social media, and explor…