Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Agentic Large Language Models for Training-Free Neuro-Radiological Image Analysis
Ayhan Can Erdur, Daniel Scholz, Jiazhen Pan +3
State-of-the-art large language models (LLMs) show high performance in general visual question answering. However, a fundamental limitation remains: current architectures lack the…
cs.CV2026
VariViT: A Vision Transformer for Variable Image Sizes
Aswathi Varma, Suprosanna Shit, Chinmay Prabhakar +5
Vision Transformers (ViTs) have emerged as the state-of-the-art architecture in representation learning, leveraging self-attention mechanisms to excel in various tasks. ViTs split…