activity
20232025
most citedInstructDiffusion: A Generalist Modeling Interface for Vision Tasks

5 citations · 9 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CV2025

X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again

Zigang Geng, Yibing Wang, Yeyao Ma +10

Numerous efforts have been made to extend the ``next token prediction'' paradigm to visual contents, aiming to create a unified approach for both image generation and understanding…

eess.IV20241 cited

HistoGym: A Reinforcement Learning Environment for Histopathological Image Analysis

Zhi-Bo Liu, Xiaobo Pang, Jizhao Wang +2

In pathological research, education, and clinical practice, the decision-making process based on pathological images is critically important. This significance extends to digital p…

cs.CL20243 cited

Common 7B Language Models Already Possess Strong Math Capabilities

Chen Li, Weiqi Wang, Jingcheng Hu +5

Mathematical capabilities were previously believed to emerge in common language models only at a very large scale or require extensive math-related pre-training. This paper shows t…

cs.LG2023

FP8-LM: Training FP8 Large Language Models

Houwen Peng, Kan Wu, Yixuan Wei +17

In this paper, we explore FP8 low-bit data formats for efficient training of large language models (LLMs). Our key insight is that most variables, such as gradients and optimizer s…

cs.CV20235 cited

InstructDiffusion: A Generalist Modeling Interface for Vision Tasks

Zigang Geng, Binxin Yang, Tiankai Hang +8

We present InstructDiffusion, a unifying and generic framework for aligning computer vision tasks with human instructions. Unlike existing approaches that integrate prior knowledge…