most citedLLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

4 citations · 4 across the 6 of their papers we have counts for

collaborators

6 papers

cs.RO2025

Data Assessment for Embodied Intelligence

Jiahao Xiao, Bowen Yan, Jianbo Zhang +4

In embodied intelligence, datasets play a pivotal role, serving as both a knowledge repository and a conduit for information transfer. The two most critical attributes of a dataset…

cs.CV2025★ 4 cited

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Xiang An, Yin Xie, Kaicheng Yang +20

We present LLaVA-OneVision-1.5, a novel family of Large Multimodal Models (LMMs) that achieve state-of-the-art performance with significantly reduced computational and financial co…

cs.CV2025

Comp-X: On Defining an Interactive Learned Image Compression Paradigm With Expert-driven LLM Agent

Yixin Gao, Xin Li, Xiaohan Pan +7

We present Comp-X, the first intelligently interactive image compression paradigm empowered by the impressive reasoning capability of large language model (LLM) agent. Notably, com…

cs.CV2025

Static and Plugged: Make Embodied Evaluation Simple

Jiahao Xiao, Jianbo Zhang, BoWen Yan +9

Embodied intelligence is advancing rapidly, driving the need for efficient evaluation. Current benchmarks typically rely on interactive simulated environments or real-world setups,…

cs.CV2025

Image Quality Assessment for Embodied AI

Chunyi Li, Jiaohao Xiao, Jianbo Zhang +8

Embodied AI has developed rapidly in recent years, but it is still mainly deployed in laboratories, with various distortions in the Real-world limiting its application. Traditional…

cs.CV2024

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering

Zheng Cheng, Rendong Wang, Zhicheng Wang

Recently, multi-modal large language models have made significant progress. However, visual information lacking of guidance from the user's intention may lead to redundant computat…