papers

Publications (51)

cs.CV2024

LISA: Reasoning Segmentation via Large Language Model

Xin Lai, Zhuotao Tian, Yukang Chen +4

Although perception systems have made remarkable advancements in recent years, they still rely on explicit human instruction or pre-defined categories to identify the target object…

cs.CV2025

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World

Xiangtai Li, Tao Zhang, Yanwei Li +13

Multimodal Large Language Models (MLLMs) demonstrate a complex understanding of scenes, benefiting from large-scale and high-quality datasets. Most existing caption datasets lack t…

cs.CV2025

Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark

Haobo Yuan, Yueyi Sun, Yanwei Li +7

Recent advances in Multimodal Large Language Models (MLLMs) have significantly improved performance on tasks such as visual grounding and visual question answering. However, the re…

cs.CV2019

FastPose: Towards Real-time Pose Estimation and Tracking via Scale-normalized Multi-task Networks

Jiabin Zhang, Zheng Zhu, Wei Zou +4

Both accuracy and efficiency are significant for pose estimation and tracking in videos. State-of-the-art performance is dominated by two-stages top-down methods. Despite the leadi…

cs.CV2025

Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Tao Zhang, Xiangtai Li, Zilong Huang +6

Multimodal Large Language Models (MLLMs) achieve remarkable performance for fine-grained pixel-level understanding tasks. However, all the works rely heavily on extra components, s…

cs.CV2024

Beyond Pixels: Text Enhances Generalization in Real-World Image Restoration

Haoze Sun, Wenbo Li, Jiayue Liu +7

Generalization has long been a central challenge in real-world image restoration. While recent diffusion-based restoration methods, which leverage generative priors from text-to-im…

eess.IV2023

Democratizing Pathological Image Segmentation with Lay Annotators via Molecular-empowered Learning

Ruining Deng, Yanwei Li, Peize Li +11

Multi-class cell segmentation in high-resolution Giga-pixel whole slide images (WSI) is critical for various clinical applications. Training such an AI model typically requires lab…

cs.CV2021

Dynamic Scale Training for Object Detection

Yukang Chen, Peizhen Zhang, Zeming Li +5

We propose a Dynamic Scale Training paradigm (abbreviated as DST) to mitigate scale variation challenge in object detection. Previous strategies like image pyramid, multi-scale tra…

cs.LG2026

One Loss to Rule Them All: Marked Time-to-Event for Structured EHR Foundation Models

Zilin Jing, Vincent Jeanselme, Yuta Kobayashi +6

Clinical events captured in Electronic Health Records (EHR) are irregularly sampled and may consist of a mixture of discrete events and numerical measurements, such as laboratory v…

cond-mat.soft2021

Unconventional rheological properties in systems of deformable particles

Anshuman Pasupalak, Shawn Khuhan Samidurai, Yanwei Li +3

We demonstrate the existence of unconventional rheological and memory properties in systems of soft-deformable particles whose energy depends on their shape, via numerical simulati…

cs.CV2021

Fully Convolutional Networks for Panoptic Segmentation

Yanwei Li, Hengshuang Zhao, Xiaojuan Qi +4

In this paper, we present a conceptually simple, strong, and efficient framework for panoptic segmentation, called Panoptic FCN. Our approach aims to represent and predict foregrou…

quant-ph2026

High-rate quantum digital signatures over 250 km of optical fiber

Jiemin Lin, Yongqiang Du, Mingxuan Zhang +7

Quantum digital signatures (QDS) offer information-theoretic security for message integrity, authenticity, and non-repudiation, and constitute a fundamental cryptographic primitive…

cs.CV2026

Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs

Haochen Wang, Yuhao Wang, Tao Zhang +13

While Multimodal Large Language Models (MLLMs) excel at holistic understanding, they struggle in capturing the dense world with complex scenes, requiring fine-grained analysis of i…

cs.CV2021

Multi-Scale Aligned Distillation for Low-Resolution Detection

Lu Qi, Jason Kuen, Jiuxiang Gu +5

In instance-level detection tasks (e.g., object detection), reducing input resolution is an easy option to improve runtime efficiency. However, this option traditionally hurts the…

cs.CV2025

Visual Spatial Tuning

Rui Yang, Ziyu Zhu, Yanwei Li +9

Capturing spatial relationships from visual inputs is a cornerstone of human-like general intelligence. Several previous studies have tried to enhance the spatial awareness of Visi…

cs.LG2025

FoMoH: A clinically meaningful foundation model evaluation for structured electronic health records

Chao Pang, Vincent Jeanselme, Young Sang Choi +9

Foundation models (FMs) promise to address core limitations of traditional supervised machine learning: (i) reliance on large amounts of labeled data, (ii) task specificity, and (i…

cond-mat.soft2019

Long-wavelength fluctuations and anomalous dynamics in two-dimensional liquids

Yanwei Li, Chandan K. Mishra, Zhaoyan Sun +4

Long-wavelength Mermin-Wagner fluctuations prevent the existence of translational long-range order, in two-dimensional systems at finite temperature. Their dynamical signature, whi…

cs.CV2019

Attention-guided Unified Network for Panoptic Segmentation

Yanwei Li, Xinze Chen, Zheng Zhu +4

This paper studies panoptic segmentation, a recently proposed task which segments foreground (FG) objects at the instance level as well as background (BG) contents at the semantic…

math.OC2026

Value existence for zero-sum ergodic stochastic differential games

Juan Li, Wenqiang Li, Yanwei Li +1

In this paper we investigate two-player zero-sum stochastic differential games with an ergodic payoff, in which the diffusion coefficient does not need to be non-degenerate. We fir…

cs.CV2025

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Dongzhi Jiang, Renrui Zhang, Ziyu Guo +11

Answering questions with Chain-of-Thought (CoT) has significantly enhanced the reasoning capabilities of Large Language Models (LLMs), yet its impact on Large Multimodal Models (LM…

cs.AI2026

Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond

Meng Chu, Xuan Billy Zhang, Kevin Qinghong Lin +47

As AI systems move from generating text to accomplishing goals through sustained interaction, the ability to model environment dynamics becomes a central bottleneck. Agents that ma…

cs.CV2024

Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition

Zhisheng Zhong, Chengyao Wang, Yuqi Liu +12

As Multi-modal Large Language Models (MLLMs) evolve, expanding beyond single-domain capabilities is essential to meet the demands for more versatile and efficient AI. However, prev…

cs.CV2026

Watch, Remember, Reason: Human-View Video Understanding with MLLMs

Jiahao Meng, Yue Tan, Qi Xu +12

Video understanding is being rapidly transformed by multimodal large language models (MLLMs), as research moves from short clips to long, multimodal, and knowledge-intensive video…

cs.AI2024

RL-GPT: Integrating Reinforcement Learning and Code-as-policy

Shaoteng Liu, Haoqi Yuan, Minda Hu +5

Large Language Models (LLMs) have demonstrated proficiency in utilizing various tools by coding, yet they face limitations in handling intricate logic and precise control. In embod…

cs.CV2023

GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction

Rui Yang, Lin Song, Yanwei Li +4

This paper aims to efficiently enable Large Language Models (LLMs) to use multimodal tools. Advanced proprietary LLMs, such as ChatGPT and GPT-4, have shown great potential for too…

cs.CV2022

Focal Sparse Convolutional Networks for 3D Object Detection

Yukang Chen, Yanwei Li, Xiangyu Zhang +2

Non-uniformed 3D sparse data, e.g., point clouds or voxels in different spatial positions, make contribution to the task of 3D object detection in different ways. Existing basic co…

cs.CV2020

Learning Dynamic Routing for Semantic Segmentation

Yanwei Li, Lin Song, Yukang Chen +4

Recently, numerous handcrafted and searched networks have been applied for semantic segmentation. However, previous works intend to handle inputs with various scales in pre-defined…

quant-ph2026

Field-Trial Quantum Key Distribution with Qubit-Based Frame Synchronization

Rui Guan, Jingchun Yu, Zhaoyun Li +7

Quantum key distribution (QKD) is a cryptographic technique that uses quantum mechanical principles to enable secure key exchange. Practical deployment of QKD requires robust, cost…

cs.CV2024

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Yanwei Li, Yuechen Zhang, Chengyao Wang +5

In this work, we introduce Mini-Gemini, a simple and effective framework enhancing multi-modality Vision Language Models (VLMs). Despite the advancements in VLMs facilitating basic…

cs.CV2022

Fully Convolutional Networks for Panoptic Segmentation with Point-based Supervision

Yanwei Li, Hengshuang Zhao, Xiaojuan Qi +6

In this paper, we present a conceptually simple, strong, and efficient framework for fully- and weakly-supervised panoptic segmentation, called Panoptic FCN. Our approach aims to r…

cs.CV2026

MotionAtlas: Detailed Region Captioning for Motion-Centric Videos

Weisong Liu, Haochen Wang, Kuan Gao +8

We propose MotionAtlas, a system for detailed captioning of motion-centric videos, comprising (1) a dedicated human-annotated benchmark, (2) a scalable, high-quality pipeline to co…

cs.CV2025

Seed1.5-VL Technical Report

Dong Guo, Faming Wu, Feida Zhu +194

We present Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning. Seed1.5-VL is composed with a 532M-parameter v…

math.OC2024

A Global Stochastic Maximum Principle for Mean-Field Forward-Backward Stochastic Control Systems with Quadratic Generators

Rainer Buckdahn, Juan Li, Yanwei Li +1

Our paper is devoted to the study of Peng's stochastic maximum principle (SMP) for a stochastic control problem composed of a controlled forward stochastic differential equation (S…

cs.CV2020

Fine-Grained Dynamic Head for Object Detection

Lin Song, Yanwei Li, Zhengkai Jiang +4

The Feature Pyramid Network (FPN) presents a remarkable approach to alleviate the scale variance in object representation by performing instance-level assignments. Nevertheless, th…

cs.CL2025

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models

Shilin Xu, Yanwei Li, Rui Yang +9

Recent works on large language models (LLMs) have successfully demonstrated the emergence of reasoning capabilities via reinforcement learning (RL). Although recent efforts leverag…

cs.CV2019

Learnable Tree Filter for Structure-preserving Feature Transform

Lin Song, Yanwei Li, Zeming Li +4

Learning discriminative global features plays a vital role in semantic segmentation. And most of the existing methods adopt stacks of local convolutions or non-local blocks to capt…

cs.CV2022

Voxel Field Fusion for 3D Object Detection

Yanwei Li, Xiaojuan Qi, Yukang Chen +4

In this work, we present a conceptually simple yet effective framework for cross-modality 3D object detection, named voxel field fusion. The proposed approach aims to maintain cros…

cs.CY2018

The Governance of Risks in Ridesharing: A Revelatory Case from Singapore

Yanwei Li, Araz Taeihagh, Martin de Jong

Recently we have witnessed the worldwide adoption of many different types of innovative technologies, such as crowdsourcing, ridesharing, open and big data, aiming at delivering pu…

cs.CV2022

Unifying Voxel-based Representation with Transformer for 3D Object Detection

Yanwei Li, Yilun Chen, Xiaojuan Qi +3

In this work, we present a unified framework for multi-modality 3D object detection, named UVTR. The proposed method aims to unify multi-modality representations in the voxel space…

cs.CV2019

State-aware Re-identification Feature for Multi-target Multi-camera Tracking

Peng Li, Jiabin Zhang, Zheng Zhu +3

Multi-target Multi-camera Tracking (MTMCT) aims to extract the trajectories from videos captured by a set of cameras. Recently, the tracking performance of MTMCT is significantly e…

cs.CV2022

Diversified Dynamic Routing for Vision Tasks

Botos Csaba, Adel Bibi, Yanwei Li +2

Deep learning models for vision tasks are trained on large datasets under the assumption that there exists a universal representation that can be used to make predictions for all s…

cs.CV2018

Identity-Enhanced Network for Facial Expression Recognition

Yanwei Li, Xingang Wang, Shilei Zhang +4

Facial expression recognition is a challenging task, arguably because of large intra-class variations and high inter-class similarities. The core drawback of the existing approache…

cs.AI2025

How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective

Songsong Yu, Yuxin Chen, Hao Ju +15

Visual Spatial Reasoning (VSR) is a core human cognitive ability and a critical requirement for advancing embodied intelligence and autonomous systems. Despite recent progress in V…

cs.CV2025

Aligning Effective Tokens with Video Anomaly in Large Language Models

Yingxian Chen, Jiahui Liu, Ruidi Fan +6

Understanding abnormal events in videos is a vital and challenging task that has garnered significant attention in a wide range of applications. Although current video understandin…

cs.CV2026

Semantic Generative Tuning for Unified Multimodal Models

Songsong Yu, Yuxin Chen, Ying Shan +1

Unified multimodal models (UMMs) strive to consolidate visual understanding and visual generation within a single architecture. However, prevailing training paradigms independently…

cs.CV2026

Benchmarking and Evolving Reason-Reflect-Rectify for Reflective Visual Generation

Junjie Wang, Xinghua Lou, Jason Li +8

Text-to-Image (T2I) models and Unified Multimodal Models (UMMs) have achieved remarkable progress in visual generation. However, their reliance on a single-pass generation paradigm…

cs.CV2023

LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Yanwei Li, Chengyao Wang, Jiaya Jia

In this work, we present a novel method to tackle the token generation challenge in Vision Language Models (VLMs) for video and image understanding, called LLaMA-VID. Current VLMs,…

cs.CV2025

Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Chaoyou Fu, Yuhan Dai, Yongdong Luo +18

In the quest for artificial general intelligence, Multi-modal Large Language Models (MLLMs) have emerged as a focal point in recent advancements. However, the predominant focus rem…

cs.CV2020

Rethinking Learnable Tree Filter for Generic Feature Transform

Lin Song, Yanwei Li, Zhengkai Jiang +5

The Learnable Tree Filter presents a remarkable approach to model structure-preserving relations for semantic segmentation. Nevertheless, the intrinsic geometric constraint forces…

cs.CV2024

LLaVA-OneVision: Easy Visual Task Transfer

Bo Li, Yuanhan Zhang, Dong Guo +8

We present LLaVA-OneVision, a family of open large multimodal models (LMMs) developed by consolidating our insights into data, models, and visual representations in the LLaVA-NeXT…

cs.CV2021

Scale-aware Automatic Augmentation for Object Detection

Yukang Chen, Yanwei Li, Tao Kong +4

We propose Scale-aware AutoAug to learn data augmentation policies for object detection. We define a new scale-aware search space, where both image- and box-level augmentations are…