activity
20242026
collaborators

7 papers

cs.CV2026

Continual Test-Time Adaptation in Computer Vision: Methods, Benchmarks, and Future Directions

Sarthak Kumar Maharana, Shambhavi Mishra, Yunbei Zhang +6

Deep neural nets achieve remarkable performance when training and test data share the same distribution, but this assumption frequently breaks in real-world deployment, where data…

cs.AI2026

AInstein: Can LLMs Solve Research Problems From Parametric Memory Alone?

Shambhavi Mishra, Gaurav Sahu, Marco Pedersoli +3

Can large language models solve AI research problems using only their parametric knowledge, without fine-tuning, retrieval, or other external aids? We introduce AInstein, a framewo…

cs.CV2026

Distilling Specialized Orders for Visual Generation

Rishav Pramanik, Amin Sghaier, Masih Aminbeidokhti +6

Autoregressive (AR) image generators are becoming increasingly popular due to their ability to produce high-quality images and their scalability. Typical AR models are locked onto…

cs.CV2025

Rendering-Aware Reinforcement Learning for Vector Graphics Generation

Juan A. Rodriguez, Haotian Zhang, Abhay Puri +12

Scalable Vector Graphics (SVG) offer a powerful format for representing visual designs as interpretable code. Recent advances in vision-language models (VLMs) have enabled high-qua…

cs.CL2025

AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding

Ahmed Masry, Juan A. Rodriguez, Tianyu Zhang +19

Aligning visual features with language embeddings is a key challenge in vision-language models (VLMs). The performance of such models hinges on having a good connector that maps vi…

cs.CV2024

Unsupervised Object Discovery: A Comprehensive Survey and Unified Taxonomy

José-Fabian Villa-Vásquez, Marco Pedersoli

Unsupervised object discovery is commonly interpreted as the task of localizing and/or categorizing objects in visual data without the need for labeled examples. While current obje…