papers

Publications (28)

cs.CV2025

Can Atomic Step Decomposition Enhance the Self-structured Reasoning of Multimodal Large Models?

Kun Xiang, Zhili Liu, Zihao Jiang +13

In this paper, we address the challenging task of multimodal mathematical reasoning by incorporating the ability of "slow thinking" into multimodal large language models (MLLMs). O…

cs.RO2021

Does elderly enjoy playing Bingo with a robot? A case study with the humanoid robot Nadine

Nidhi Mishra, Gauri Tulsulkar, Hanhui Li +1

There are considerable advancements in medical health care in recent years, resulting in rising older population. As the workforce for such a population is not keeping pace, there…

eess.IV2018

Multi-column Point-CNN for Sketch Segmentation

Fei Wang, Shujin Lin, Hanhui Li +4

Traditional sketch segmentation methods mainly rely on handcrafted features and complicate models, and their performance is far from satisfactory due to the abstract representation…

cs.RO2021

Towards Complex and Continuous Manipulation: A Gesture Based Anthropomorphic Robotic Hand Design

Li Tian, Hanhui Li, Qifa Wang +5

Most current anthropomorphic robotic hands can realize part of the human hand functions, particularly for object grasping. However, due to the complexity of the human hand, few cur…

cs.CV2024

ConsistentID: Portrait Generation with Multimodal Fine-Grained Identity Preserving

Jiehui Huang, Xiao Dong, Wenhui Song +9

Diffusion-based technologies have made significant strides, particularly in personalized and customized facialgeneration. However, existing methods face challenges in achieving hig…

cs.CV2024

Learning Interaction-aware 3D Gaussian Splatting for One-shot Hand Avatars

Xuan Huang, Hanhui Li, Wanquan Liu +4

In this paper, we propose to create animatable avatars for interacting hands with 3D Gaussian Splatting (GS) and single-image inputs. Existing GS-based methods designed for single…

cs.CV2026

PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models

Bin Hu, Yanwen Ma, Jiehui Huang +14

Recent game world models can synthesize visually plausible, action-conditioned rollouts. However, their interaction behaviors often remain limited to exploratory or wandering traje…

cs.CV2025

AutoStudio: Crafting Consistent Subjects in Multi-turn Interactive Image Generation

Junhao Cheng, Xi Lu, Hanhui Li +5

As cutting-edge Text-to-Image (T2I) generation models already excel at producing remarkable single images, an even more challenging task, i.e., multi-turn interactive image generat…

cs.CV2025

LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation

Wenhui Song, Hanhui Li, Jiehui Huang +5

In this paper, we present LaVieID, a novel \underline{l}ocal \underline{a}utoregressive \underline{vi}d\underline{e}o diffusion framework designed to tackle the challenging \underl…

cs.CV2025

TheaterGen: Character Management with LLM for Consistent Multi-turn Image Generation

Junhao Cheng, Baiqiao Yin, Kaixin Cai +9

Recent advances in diffusion models can generate high-quality and stunning images from text. However, multi-turn image generation, which is of high demand in real-world scenarios,…

cs.CV2018

Beyond Context: Exploring Semantic Similarity for Tiny Face Detection

Yue Xi, Jiangbin Zheng, Xiangjian He +2

Tiny face detection aims to find faces with high degrees of variability in scale, resolution and occlusion in cluttered scenes. Due to the very little information available on tiny…

cs.CV2026

WildGHand: Learning Anti-Perturbation Gaussian Hand Avatars from Monocular In-the-Wild Videos

Hanhui Li, Xuan Huang, Wanquan Liu +5

Despite recent progress in 3D hand reconstruction from monocular videos, most existing methods rely on data captured in well-controlled environments and therefore degrade in real-w…

cs.CV2026

SyncLoop: A Multimodal Dual-Loop Framework for Self-Improving Mathematical Reasoning

Xiuwei Chen, Wentao Hu, Hanhui Li +9

Recent advances in multimodal large language models (MLLMs) have shown impressive reasoning capabilities. However, further enhancing existing MLLMs necessitates high-quality vision…

cs.AI2025

SeePhys: Does Seeing Help Thinking? -- Benchmarking Vision-Based Physics Reasoning

Kun Xiang, Heng Li, Terry Jingchen Zhang +11

We present SeePhys, a large-scale multimodal benchmark for LLM reasoning grounded in physics questions ranging from middle school to PhD qualifying exams. The benchmark covers 7 fu…

cs.CV2024

3D Visibility-aware Generalizable Neural Radiance Fields for Interacting Hands

Xuan Huang, Hanhui Li, Zejun Yang +2

Neural radiance fields (NeRFs) are promising 3D representations for scenes, objects, and humans. However, most existing methods require multi-view inputs and per-scene training, wh…

cs.CV2019

Instance-Aware Representation Learning and Association for Online Multi-Person Tracking

Hefeng Wu, Yafei Hu, Keze Wang +3

Multi-Person Tracking (MPT) is often addressed within the detection-to-association paradigm. In such approaches, human detections are first extracted in every frame and person traj…

cs.CV2026

ProPhy: Progressive Physical Alignment for Dynamic World Simulation

Zijun Wang, Panwen Hu, Jing Wang +7

Recent advances in video generation have shown remarkable potential for constructing world simulators. However, current models still struggle to produce physically consistent resul…

cs.CV2026

Beyond the Eye: Efficient Multimodal Reasoning via Self-Regulated Implicit Visual Tools

Xiuwei Chen, Quanlin Chen, Wentao Hu +8

The paper introduces Beyond the Eye (BEE), an implicit visual‑tool framework for multimodal large language models that learns to self‑regulate when to invoke visual tools, reducing…

#multimodal large language models#implicit visual tools#self‑regulated inference#chain‑of‑thought fine‑tuning
cs.LG2026

Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration

Zhicheng Yang, Zhijiang Guo, Yinya Huang +6

Reinforcement Learning with Verifiable Reward (RLVR) is a powerful method for enhancing the reasoning abilities of Large Language Models, but its full potential is limited by a lac…

cs.RO2020

Fast 3D Modeling of Anthropomorphic Robotic Hands Based on A Multi-layer Deformable Design

Li Tian, Hanhui Li, Muhammad Faaiz Khan Bin Abdul Halil +3

Current anthropomorphic robotic hands mainly focus on improving their dexterity by devising new mechanical structures and actuation systems. However, most of them rely on a single…

cs.CV2026

AtomThink: Multimodal Slow Thinking with Atomic Step Reasoning

Kun Xiang, Zhili Liu, Terry Jingchen Zhang +12

In this paper, we address the challenging task of multimodal reasoning by incorporating the notion of ``slow thinking'' into multimodal large language models (MLLMs). Our core idea…

cs.CV2026

ERGO: Excess-Risk-Guided Optimization for High-Fidelity Monocular 3D Gaussian Splatting

Zehua Ma, Hanhui Li, Zhenyu Xie +4

Generating 3D content from a single image remains a fundamentally challenging and ill-posed problem due to the inherent absence of geometric and textural information in occluded re…

cs.CV2022

Towards Hard-pose Virtual Try-on via 3D-aware Global Correspondence Learning

Zaiyu Huang, Hanhui Li, Zhenyu Xie +3

In this paper, we target image-based person-to-person virtual try-on in the presence of diverse poses and large viewpoint variations. Existing methods are restricted in this settin…

cs.AI2026

Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI

Kun Xiang, Terry Jingchen Zhang, Yinya Huang +13

The rapid advancement of embodied intelligence and world models has intensified efforts to integrate physical laws into AI systems, yet physical perception and symbolic physics rea…

cs.CV2023

Monocular 3D Hand Mesh Recovery via Dual Noise Estimation

Hanhui Li, Xiaojian Lin, Xuan Huang +3

Current parametric models have made notable progress in 3D hand pose and shape estimation. However, due to the fixed hand topology and complex hand poses, current models are hard t…

cs.CV2018

Structured Inhomogeneous Density Map Learning for Crowd Counting

Hanhui Li, Xiangjian He, Hefeng Wu +4

In this paper, we aim at tackling the problem of crowd counting in extremely high-density scenes, which contain hundreds, or even thousands of people. We begin by a comprehensive a…

cs.CV2025

FireEdit: Fine-grained Instruction-based Image Editing via Region-aware Vision Language Model

Jun Zhou, Jiahao Li, Zunnan Xu +6

Currently, instruction-based image editing methods have made significant progress by leveraging the powerful cross-modal understanding capabilities of vision language models (VLMs)…

cs.CV2024

GarmentAligner: Text-to-Garment Generation via Retrieval-augmented Multi-level Corrections

Shiyue Zhang, Zheng Chong, Xujie Zhang +4

General text-to-image models bring revolutionary innovation to the fields of arts, design, and media. However, when applied to garment generation, even the state-of-the-art text-to…