5 papers
Continual Test-Time Adaptation in Computer Vision: Methods, Benchmarks, and Future Directions
Sarthak Kumar Maharana, Shambhavi Mishra, Yunbei Zhang +6
Deep neural nets achieve remarkable performance when training and test data share the same distribution, but this assumption frequently breaks in real-world deployment, where data…
AInstein: Can LLMs Solve Research Problems From Parametric Memory Alone?
Shambhavi Mishra, Gaurav Sahu, Marco Pedersoli +3
Can large language models solve AI research problems using only their parametric knowledge, without fine-tuning, retrieval, or other external aids? We introduce AInstein, a framewo…
Distilling Specialized Orders for Visual Generation
Rishav Pramanik, Amin Sghaier, Masih Aminbeidokhti +6
Autoregressive (AR) image generators are becoming increasingly popular due to their ability to produce high-quality images and their scalability. Typical AR models are locked onto…
Rendering-Aware Reinforcement Learning for Vector Graphics Generation
Juan A. Rodriguez, Haotian Zhang, Abhay Puri +12
Scalable Vector Graphics (SVG) offer a powerful format for representing visual designs as interpretable code. Recent advances in vision-language models (VLMs) have enabled high-qua…
AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
Ahmed Masry, Juan A. Rodriguez, Tianyu Zhang +19
Aligning visual features with language embeddings is a key challenge in vision-language models (VLMs). The performance of such models hinges on having a good connector that maps vi…