1 citations · 2 across the 5 of their papers we have counts for
7 papers
RLLaVA: An RL-central Framework for Language and Vision Assistants
Lei Zhao, Zihao Ma, Boyu Lin +3
We present an RL-central framework for Language and Vision Assistants (RLLaVA) with its formulation of Markov decision process (MDP). RLLaVA decouples RL algorithmic logic from mod…
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
Xingjian Zhang, Siwei Wen, Wenjun Wu +1
Large Language Models (LLMs) have made remarkable progress in enhancing step-by-step reasoning through reinforcement learning. However, the Group Relative Policy Optimization (GRPO…
CoE-Ops: Collaboration of LLM-based Experts for AIOps Question-Answering
Jinkun Zhao, Yuanshuai Wang, Xingjian Zhang +6
With the rapid evolution of artificial intelligence, AIOps has emerged as a prominent paradigm in DevOps. Lots of work has been proposed to improve the performance of different AIO…
VoQA: Visual-only Question Answering
Jianing An, Luyang Jiang, Jie Luo +2
Visual understanding requires interpreting both natural scenes and the textual information that appears within them, motivating tasks such as Visual Question Answering (VQA). Howev…
TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning
Xingjian Zhang, Siwei Wen, Wenjun Wu +1
Recently, improving the reasoning ability of large multimodal models (LMMs) through reinforcement learning has made great progress. However, most existing works are based on highly…
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler
Xingjian Zhang, Xi Weng, Yihao Yue +3
Video behavior recognition and scene understanding are fundamental tasks in multimodal intelligence, serving as critical building blocks for numerous real-world applications. Throu…