Publications (28)
Can Atomic Step Decomposition Enhance the Self-structured Reasoning of Multimodal Large Models?
Kun Xiang, Zhili Liu, Zihao Jiang +13
In this paper, we address the challenging task of multimodal mathematical reasoning by incorporating the ability of "slow thinking" into multimodal large language models (MLLMs). O…
Does elderly enjoy playing Bingo with a robot? A case study with the humanoid robot Nadine
Nidhi Mishra, Gauri Tulsulkar, Hanhui Li +1
There are considerable advancements in medical health care in recent years, resulting in rising older population. As the workforce for such a population is not keeping pace, there…
Multi-column Point-CNN for Sketch Segmentation
Fei Wang, Shujin Lin, Hanhui Li +4
Traditional sketch segmentation methods mainly rely on handcrafted features and complicate models, and their performance is far from satisfactory due to the abstract representation…
Towards Complex and Continuous Manipulation: A Gesture Based Anthropomorphic Robotic Hand Design
Li Tian, Hanhui Li, Qifa Wang +5
Most current anthropomorphic robotic hands can realize part of the human hand functions, particularly for object grasping. However, due to the complexity of the human hand, few cur…
ConsistentID: Portrait Generation with Multimodal Fine-Grained Identity Preserving
Jiehui Huang, Xiao Dong, Wenhui Song +9
Diffusion-based technologies have made significant strides, particularly in personalized and customized facialgeneration. However, existing methods face challenges in achieving hig…
Learning Interaction-aware 3D Gaussian Splatting for One-shot Hand Avatars
Xuan Huang, Hanhui Li, Wanquan Liu +4
In this paper, we propose to create animatable avatars for interacting hands with 3D Gaussian Splatting (GS) and single-image inputs. Existing GS-based methods designed for single…
PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models
Bin Hu, Yanwen Ma, Jiehui Huang +14
Recent game world models can synthesize visually plausible, action-conditioned rollouts. However, their interaction behaviors often remain limited to exploratory or wandering traje…
AutoStudio: Crafting Consistent Subjects in Multi-turn Interactive Image Generation
Junhao Cheng, Xi Lu, Hanhui Li +5
As cutting-edge Text-to-Image (T2I) generation models already excel at producing remarkable single images, an even more challenging task, i.e., multi-turn interactive image generat…
LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation
Wenhui Song, Hanhui Li, Jiehui Huang +5
In this paper, we present LaVieID, a novel \underline{l}ocal \underline{a}utoregressive \underline{vi}d\underline{e}o diffusion framework designed to tackle the challenging \underl…
TheaterGen: Character Management with LLM for Consistent Multi-turn Image Generation
Junhao Cheng, Baiqiao Yin, Kaixin Cai +9
Recent advances in diffusion models can generate high-quality and stunning images from text. However, multi-turn image generation, which is of high demand in real-world scenarios,…
Beyond Context: Exploring Semantic Similarity for Tiny Face Detection
Yue Xi, Jiangbin Zheng, Xiangjian He +2
Tiny face detection aims to find faces with high degrees of variability in scale, resolution and occlusion in cluttered scenes. Due to the very little information available on tiny…
WildGHand: Learning Anti-Perturbation Gaussian Hand Avatars from Monocular In-the-Wild Videos
Hanhui Li, Xuan Huang, Wanquan Liu +5
Despite recent progress in 3D hand reconstruction from monocular videos, most existing methods rely on data captured in well-controlled environments and therefore degrade in real-w…
SyncLoop: A Multimodal Dual-Loop Framework for Self-Improving Mathematical Reasoning
Xiuwei Chen, Wentao Hu, Hanhui Li +9
Recent advances in multimodal large language models (MLLMs) have shown impressive reasoning capabilities. However, further enhancing existing MLLMs necessitates high-quality vision…
SeePhys: Does Seeing Help Thinking? -- Benchmarking Vision-Based Physics Reasoning
Kun Xiang, Heng Li, Terry Jingchen Zhang +11
We present SeePhys, a large-scale multimodal benchmark for LLM reasoning grounded in physics questions ranging from middle school to PhD qualifying exams. The benchmark covers 7 fu…
3D Visibility-aware Generalizable Neural Radiance Fields for Interacting Hands
Xuan Huang, Hanhui Li, Zejun Yang +2
Neural radiance fields (NeRFs) are promising 3D representations for scenes, objects, and humans. However, most existing methods require multi-view inputs and per-scene training, wh…
Instance-Aware Representation Learning and Association for Online Multi-Person Tracking
Hefeng Wu, Yafei Hu, Keze Wang +3
Multi-Person Tracking (MPT) is often addressed within the detection-to-association paradigm. In such approaches, human detections are first extracted in every frame and person traj…
ProPhy: Progressive Physical Alignment for Dynamic World Simulation
Zijun Wang, Panwen Hu, Jing Wang +7
Recent advances in video generation have shown remarkable potential for constructing world simulators. However, current models still struggle to produce physically consistent resul…
Beyond the Eye: Efficient Multimodal Reasoning via Self-Regulated Implicit Visual Tools
Xiuwei Chen, Quanlin Chen, Wentao Hu +8
The paper introduces Beyond the Eye (BEE), an implicit visual‑tool framework for multimodal large language models that learns to self‑regulate when to invoke visual tools, reducing…
Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration
Zhicheng Yang, Zhijiang Guo, Yinya Huang +6
Reinforcement Learning with Verifiable Reward (RLVR) is a powerful method for enhancing the reasoning abilities of Large Language Models, but its full potential is limited by a lac…
Fast 3D Modeling of Anthropomorphic Robotic Hands Based on A Multi-layer Deformable Design
Li Tian, Hanhui Li, Muhammad Faaiz Khan Bin Abdul Halil +3
Current anthropomorphic robotic hands mainly focus on improving their dexterity by devising new mechanical structures and actuation systems. However, most of them rely on a single…
AtomThink: Multimodal Slow Thinking with Atomic Step Reasoning
Kun Xiang, Zhili Liu, Terry Jingchen Zhang +12
In this paper, we address the challenging task of multimodal reasoning by incorporating the notion of ``slow thinking'' into multimodal large language models (MLLMs). Our core idea…
ERGO: Excess-Risk-Guided Optimization for High-Fidelity Monocular 3D Gaussian Splatting
Zehua Ma, Hanhui Li, Zhenyu Xie +4
Generating 3D content from a single image remains a fundamentally challenging and ill-posed problem due to the inherent absence of geometric and textural information in occluded re…
Towards Hard-pose Virtual Try-on via 3D-aware Global Correspondence Learning
Zaiyu Huang, Hanhui Li, Zhenyu Xie +3
In this paper, we target image-based person-to-person virtual try-on in the presence of diverse poses and large viewpoint variations. Existing methods are restricted in this settin…
Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
Kun Xiang, Terry Jingchen Zhang, Yinya Huang +13
The rapid advancement of embodied intelligence and world models has intensified efforts to integrate physical laws into AI systems, yet physical perception and symbolic physics rea…
Monocular 3D Hand Mesh Recovery via Dual Noise Estimation
Hanhui Li, Xiaojian Lin, Xuan Huang +3
Current parametric models have made notable progress in 3D hand pose and shape estimation. However, due to the fixed hand topology and complex hand poses, current models are hard t…
Structured Inhomogeneous Density Map Learning for Crowd Counting
Hanhui Li, Xiangjian He, Hefeng Wu +4
In this paper, we aim at tackling the problem of crowd counting in extremely high-density scenes, which contain hundreds, or even thousands of people. We begin by a comprehensive a…
FireEdit: Fine-grained Instruction-based Image Editing via Region-aware Vision Language Model
Jun Zhou, Jiahao Li, Zunnan Xu +6
Currently, instruction-based image editing methods have made significant progress by leveraging the powerful cross-modal understanding capabilities of vision language models (VLMs)…
GarmentAligner: Text-to-Garment Generation via Retrieval-augmented Multi-level Corrections
Shiyue Zhang, Zheng Chong, Xujie Zhang +4
General text-to-image models bring revolutionary innovation to the fields of arts, design, and media. However, when applied to garment generation, even the state-of-the-art text-to…