Publications (51)
LISA: Reasoning Segmentation via Large Language Model
Xin Lai, Zhuotao Tian, Yukang Chen +4
Although perception systems have made remarkable advancements in recent years, they still rely on explicit human instruction or pre-defined categories to identify the target object…
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
Xiangtai Li, Tao Zhang, Yanwei Li +13
Multimodal Large Language Models (MLLMs) demonstrate a complex understanding of scenes, benefiting from large-scale and high-quality datasets. Most existing caption datasets lack t…
Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
Haobo Yuan, Yueyi Sun, Yanwei Li +7
Recent advances in Multimodal Large Language Models (MLLMs) have significantly improved performance on tasks such as visual grounding and visual question answering. However, the re…
FastPose: Towards Real-time Pose Estimation and Tracking via Scale-normalized Multi-task Networks
Jiabin Zhang, Zheng Zhu, Wei Zou +4
Both accuracy and efficiency are significant for pose estimation and tracking in videos. State-of-the-art performance is dominated by two-stages top-down methods. Despite the leadi…
Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding
Tao Zhang, Xiangtai Li, Zilong Huang +6
Multimodal Large Language Models (MLLMs) achieve remarkable performance for fine-grained pixel-level understanding tasks. However, all the works rely heavily on extra components, s…
Beyond Pixels: Text Enhances Generalization in Real-World Image Restoration
Haoze Sun, Wenbo Li, Jiayue Liu +7
Generalization has long been a central challenge in real-world image restoration. While recent diffusion-based restoration methods, which leverage generative priors from text-to-im…
Democratizing Pathological Image Segmentation with Lay Annotators via Molecular-empowered Learning
Ruining Deng, Yanwei Li, Peize Li +11
Multi-class cell segmentation in high-resolution Giga-pixel whole slide images (WSI) is critical for various clinical applications. Training such an AI model typically requires lab…
Dynamic Scale Training for Object Detection
Yukang Chen, Peizhen Zhang, Zeming Li +5
We propose a Dynamic Scale Training paradigm (abbreviated as DST) to mitigate scale variation challenge in object detection. Previous strategies like image pyramid, multi-scale tra…
One Loss to Rule Them All: Marked Time-to-Event for Structured EHR Foundation Models
Zilin Jing, Vincent Jeanselme, Yuta Kobayashi +6
Clinical events captured in Electronic Health Records (EHR) are irregularly sampled and may consist of a mixture of discrete events and numerical measurements, such as laboratory v…
Unconventional rheological properties in systems of deformable particles
Anshuman Pasupalak, Shawn Khuhan Samidurai, Yanwei Li +3
We demonstrate the existence of unconventional rheological and memory properties in systems of soft-deformable particles whose energy depends on their shape, via numerical simulati…
Fully Convolutional Networks for Panoptic Segmentation
Yanwei Li, Hengshuang Zhao, Xiaojuan Qi +4
In this paper, we present a conceptually simple, strong, and efficient framework for panoptic segmentation, called Panoptic FCN. Our approach aims to represent and predict foregrou…
High-rate quantum digital signatures over 250 km of optical fiber
Jiemin Lin, Yongqiang Du, Mingxuan Zhang +7
Quantum digital signatures (QDS) offer information-theoretic security for message integrity, authenticity, and non-repudiation, and constitute a fundamental cryptographic primitive…
Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
Haochen Wang, Yuhao Wang, Tao Zhang +13
While Multimodal Large Language Models (MLLMs) excel at holistic understanding, they struggle in capturing the dense world with complex scenes, requiring fine-grained analysis of i…
Multi-Scale Aligned Distillation for Low-Resolution Detection
Lu Qi, Jason Kuen, Jiuxiang Gu +5
In instance-level detection tasks (e.g., object detection), reducing input resolution is an easy option to improve runtime efficiency. However, this option traditionally hurts the…
Visual Spatial Tuning
Rui Yang, Ziyu Zhu, Yanwei Li +9
Capturing spatial relationships from visual inputs is a cornerstone of human-like general intelligence. Several previous studies have tried to enhance the spatial awareness of Visi…
FoMoH: A clinically meaningful foundation model evaluation for structured electronic health records
Chao Pang, Vincent Jeanselme, Young Sang Choi +9
Foundation models (FMs) promise to address core limitations of traditional supervised machine learning: (i) reliance on large amounts of labeled data, (ii) task specificity, and (i…
Long-wavelength fluctuations and anomalous dynamics in two-dimensional liquids
Yanwei Li, Chandan K. Mishra, Zhaoyan Sun +4
Long-wavelength Mermin-Wagner fluctuations prevent the existence of translational long-range order, in two-dimensional systems at finite temperature. Their dynamical signature, whi…
Attention-guided Unified Network for Panoptic Segmentation
Yanwei Li, Xinze Chen, Zheng Zhu +4
This paper studies panoptic segmentation, a recently proposed task which segments foreground (FG) objects at the instance level as well as background (BG) contents at the semantic…
Value existence for zero-sum ergodic stochastic differential games
Juan Li, Wenqiang Li, Yanwei Li +1
In this paper we investigate two-player zero-sum stochastic differential games with an ergodic payoff, in which the diffusion coefficient does not need to be non-degenerate. We fir…
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Dongzhi Jiang, Renrui Zhang, Ziyu Guo +11
Answering questions with Chain-of-Thought (CoT) has significantly enhanced the reasoning capabilities of Large Language Models (LLMs), yet its impact on Large Multimodal Models (LM…
Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond
Meng Chu, Xuan Billy Zhang, Kevin Qinghong Lin +47
As AI systems move from generating text to accomplishing goals through sustained interaction, the ability to model environment dynamics becomes a central bottleneck. Agents that ma…
Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition
Zhisheng Zhong, Chengyao Wang, Yuqi Liu +12
As Multi-modal Large Language Models (MLLMs) evolve, expanding beyond single-domain capabilities is essential to meet the demands for more versatile and efficient AI. However, prev…
Watch, Remember, Reason: Human-View Video Understanding with MLLMs
Jiahao Meng, Yue Tan, Qi Xu +12
Video understanding is being rapidly transformed by multimodal large language models (MLLMs), as research moves from short clips to long, multimodal, and knowledge-intensive video…
RL-GPT: Integrating Reinforcement Learning and Code-as-policy
Shaoteng Liu, Haoqi Yuan, Minda Hu +5
Large Language Models (LLMs) have demonstrated proficiency in utilizing various tools by coding, yet they face limitations in handling intricate logic and precise control. In embod…
GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction
Rui Yang, Lin Song, Yanwei Li +4
This paper aims to efficiently enable Large Language Models (LLMs) to use multimodal tools. Advanced proprietary LLMs, such as ChatGPT and GPT-4, have shown great potential for too…
Focal Sparse Convolutional Networks for 3D Object Detection
Yukang Chen, Yanwei Li, Xiangyu Zhang +2
Non-uniformed 3D sparse data, e.g., point clouds or voxels in different spatial positions, make contribution to the task of 3D object detection in different ways. Existing basic co…
Learning Dynamic Routing for Semantic Segmentation
Yanwei Li, Lin Song, Yukang Chen +4
Recently, numerous handcrafted and searched networks have been applied for semantic segmentation. However, previous works intend to handle inputs with various scales in pre-defined…
Field-Trial Quantum Key Distribution with Qubit-Based Frame Synchronization
Rui Guan, Jingchun Yu, Zhaoyun Li +7
Quantum key distribution (QKD) is a cryptographic technique that uses quantum mechanical principles to enable secure key exchange. Practical deployment of QKD requires robust, cost…
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Yanwei Li, Yuechen Zhang, Chengyao Wang +5
In this work, we introduce Mini-Gemini, a simple and effective framework enhancing multi-modality Vision Language Models (VLMs). Despite the advancements in VLMs facilitating basic…
Fully Convolutional Networks for Panoptic Segmentation with Point-based Supervision
Yanwei Li, Hengshuang Zhao, Xiaojuan Qi +6
In this paper, we present a conceptually simple, strong, and efficient framework for fully- and weakly-supervised panoptic segmentation, called Panoptic FCN. Our approach aims to r…
MotionAtlas: Detailed Region Captioning for Motion-Centric Videos
Weisong Liu, Haochen Wang, Kuan Gao +8
We propose MotionAtlas, a system for detailed captioning of motion-centric videos, comprising (1) a dedicated human-annotated benchmark, (2) a scalable, high-quality pipeline to co…
Seed1.5-VL Technical Report
Dong Guo, Faming Wu, Feida Zhu +194
We present Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning. Seed1.5-VL is composed with a 532M-parameter v…
A Global Stochastic Maximum Principle for Mean-Field Forward-Backward Stochastic Control Systems with Quadratic Generators
Rainer Buckdahn, Juan Li, Yanwei Li +1
Our paper is devoted to the study of Peng's stochastic maximum principle (SMP) for a stochastic control problem composed of a controlled forward stochastic differential equation (S…
Fine-Grained Dynamic Head for Object Detection
Lin Song, Yanwei Li, Zhengkai Jiang +4
The Feature Pyramid Network (FPN) presents a remarkable approach to alleviate the scale variance in object representation by performing instance-level assignments. Nevertheless, th…
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models
Shilin Xu, Yanwei Li, Rui Yang +9
Recent works on large language models (LLMs) have successfully demonstrated the emergence of reasoning capabilities via reinforcement learning (RL). Although recent efforts leverag…
Learnable Tree Filter for Structure-preserving Feature Transform
Lin Song, Yanwei Li, Zeming Li +4
Learning discriminative global features plays a vital role in semantic segmentation. And most of the existing methods adopt stacks of local convolutions or non-local blocks to capt…
Voxel Field Fusion for 3D Object Detection
Yanwei Li, Xiaojuan Qi, Yukang Chen +4
In this work, we present a conceptually simple yet effective framework for cross-modality 3D object detection, named voxel field fusion. The proposed approach aims to maintain cros…
The Governance of Risks in Ridesharing: A Revelatory Case from Singapore
Yanwei Li, Araz Taeihagh, Martin de Jong
Recently we have witnessed the worldwide adoption of many different types of innovative technologies, such as crowdsourcing, ridesharing, open and big data, aiming at delivering pu…
Unifying Voxel-based Representation with Transformer for 3D Object Detection
Yanwei Li, Yilun Chen, Xiaojuan Qi +3
In this work, we present a unified framework for multi-modality 3D object detection, named UVTR. The proposed method aims to unify multi-modality representations in the voxel space…
State-aware Re-identification Feature for Multi-target Multi-camera Tracking
Peng Li, Jiabin Zhang, Zheng Zhu +3
Multi-target Multi-camera Tracking (MTMCT) aims to extract the trajectories from videos captured by a set of cameras. Recently, the tracking performance of MTMCT is significantly e…
Diversified Dynamic Routing for Vision Tasks
Botos Csaba, Adel Bibi, Yanwei Li +2
Deep learning models for vision tasks are trained on large datasets under the assumption that there exists a universal representation that can be used to make predictions for all s…
Identity-Enhanced Network for Facial Expression Recognition
Yanwei Li, Xingang Wang, Shilei Zhang +4
Facial expression recognition is a challenging task, arguably because of large intra-class variations and high inter-class similarities. The core drawback of the existing approache…
How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
Songsong Yu, Yuxin Chen, Hao Ju +15
Visual Spatial Reasoning (VSR) is a core human cognitive ability and a critical requirement for advancing embodied intelligence and autonomous systems. Despite recent progress in V…
Aligning Effective Tokens with Video Anomaly in Large Language Models
Yingxian Chen, Jiahui Liu, Ruidi Fan +6
Understanding abnormal events in videos is a vital and challenging task that has garnered significant attention in a wide range of applications. Although current video understandin…
Semantic Generative Tuning for Unified Multimodal Models
Songsong Yu, Yuxin Chen, Ying Shan +1
Unified multimodal models (UMMs) strive to consolidate visual understanding and visual generation within a single architecture. However, prevailing training paradigms independently…
Benchmarking and Evolving Reason-Reflect-Rectify for Reflective Visual Generation
Junjie Wang, Xinghua Lou, Jason Li +8
Text-to-Image (T2I) models and Unified Multimodal Models (UMMs) have achieved remarkable progress in visual generation. However, their reliance on a single-pass generation paradigm…
LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Yanwei Li, Chengyao Wang, Jiaya Jia
In this work, we present a novel method to tackle the token generation challenge in Vision Language Models (VLMs) for video and image understanding, called LLaMA-VID. Current VLMs,…
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Chaoyou Fu, Yuhan Dai, Yongdong Luo +18
In the quest for artificial general intelligence, Multi-modal Large Language Models (MLLMs) have emerged as a focal point in recent advancements. However, the predominant focus rem…
Rethinking Learnable Tree Filter for Generic Feature Transform
Lin Song, Yanwei Li, Zhengkai Jiang +5
The Learnable Tree Filter presents a remarkable approach to model structure-preserving relations for semantic segmentation. Nevertheless, the intrinsic geometric constraint forces…
LLaVA-OneVision: Easy Visual Task Transfer
Bo Li, Yuanhan Zhang, Dong Guo +8
We present LLaVA-OneVision, a family of open large multimodal models (LMMs) developed by consolidating our insights into data, models, and visual representations in the LLaVA-NeXT…
Scale-aware Automatic Augmentation for Object Detection
Yukang Chen, Yanwei Li, Tao Kong +4
We propose Scale-aware AutoAug to learn data augmentation policies for object detection. We define a new scale-aware search space, where both image- and box-level augmentations are…