most citedState-Space Large Audio Language Models

1 citations · 1 across the 3 of their papers we have counts for

collaborators

10 papers

cs.CV2025

Activation Reward Models for Few-Shot Model Alignment

Tianning Chai, Chancharik Mitra, Brandon Huang +8

Aligning Large Language Models (LLMs) and Large Multimodal Models (LMMs) to human preferences is a central challenge in improving the quality of the models' generative outputs for…

cs.HC2025

ChartGen: Scaling Chart Understanding Via Code-Guided Synthetic Chart Generation

Jovana Kondic, Pengyuan Li, Dhiraj Joshi +12

Chart-to-code reconstruction -- the task of recovering executable plotting scripts from chart images -- provides important insights into a model's ability to ground data visualizat…

cs.CV2025

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion

Jacob Hansen, Wei Lin, Junmo Kang +6

Visual Instruction Tuning (VisIT) data, commonly available as human-assistant conversations with images interleaved in the human turns, are currently the most widespread vehicle fo…

cs.MM2025

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment

Edson Araujo, Andrew Rouditchenko, Yuan Gong +7

Recent advances in audio-visual learning have shown promising results in learning representations across modalities. However, most approaches rely on global audio representations t…

cs.AI2025

Visualizing Thought: Conceptual Diagrams Enable Robust Planning in LMMs

Nasim Borazjanizadeh, Roei Herzig, Eduard Oks +3

Human reasoning relies on constructing and manipulating mental models -- simplified internal representations of situations used to understand and solve problems. Conceptual diagram…

cs.CV2024

: Bimodal Online Test-Time Adaptation for CLIP

Sarthak Kumar Maharana, Baoming Zhang, Leonid Karlinsky +2

Although open-vocabulary classification models like Contrastive Language Image Pretraining (CLIP) have demonstrated strong zero-shot learning capabilities, their robustness to comm…