collaborators

5 papers

cs.AI2025

DPO Learning with LLMs-Judge Signal for Computer Use Agents

Man Luo, David Cobbley, Xin Su +4

Computer use agents (CUA) are systems that automatically interact with graphical user interfaces (GUIs) to complete tasks. CUA have made significant progress with the advent of lar…

cs.LG2025

Probing Semantic Routing in Large Mixture-of-Expert Models

Matthew Lyle Olson, Neale Ratzlaff, Musashi Hinck +4

In the past year, large (>100B parameter) mixture-of-expert (MoE) models have become increasingly common in the open domain. While their advantages are often framed in terms of eff…

cs.CV2024

Training-Free Mitigation of Language Reasoning Degradation After Multimodal Instruction Tuning

Neale Ratzlaff, Man Luo, Xin Su +2

Multimodal models typically combine a powerful large language model (LLM) with a vision encoder and are then trained on multimodal data via instruction tuning. While this process a…

cs.CV2024

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability

Estelle Aflalo, Gabriela Ben Melech Stan, Tiep Le +5

Large Vision Language Models (LVLMs) have achieved significant progress in integrating visual and textual inputs for multimodal reasoning. However, a recurring challenge is ensurin…

cs.AI2024

FastRM: An efficient and automatic explainability framework for multimodal generative models

Gabriela Ben-Melech Stan, Estelle Aflalo, Man Luo +5

Large Vision Language Models (LVLMs) have demonstrated remarkable reasoning capabilities over textual and visual inputs. However, these models remain prone to generating misinforma…