2 papers
cs.LG2025
A Large Recurrent Action Model: xLSTM enables Fast Inference for Robotics Tasks
Thomas Schmied, Thomas Adler, Vihang Patil +6
In recent years, there has been a trend in the field of Reinforcement Learning (RL) towards large action models trained offline on large-scale datasets via sequence modeling. Exist…
cs.CV2025
Linear Alignment of Vision-language Models for Image Captioning
Fabian Paischer, Markus Hofmarcher, Sepp Hochreiter +1
Recently, vision-language models like CLIP have advanced the state of the art in a variety of multi-modal tasks including image captioning and caption evaluation. Many approaches l…