Publications (66)
DreamMat: High-quality PBR Material Generation with Geometry- and Light-aware Diffusion Models
Yuqing Zhang, Yuan Liu, Zhiyu Xie +8
2D diffusion model, which often contains unwanted baked-in shading effects and results in unrealistic rendering effects in the downstream applications. Generating Physically Based…
CrossGen: Learning and Generating Cross Fields for Quad Meshing
Qiujie Dong, Jiepeng Wang, Rui Xu +10
Cross fields play a critical role in various geometry processing tasks, especially for quad mesh generation. Existing methods for cross field generation often struggle to balance c…
3PSDF: Three-Pole Signed Distance Function for Learning Surfaces with Arbitrary Topologies
Weikai Chen, Cheng Lin, Weiyang Li +1
Recent advances in learning 3D shapes using neural implicit functions have achieved impressive results by breaking the previous barrier of resolution and diversity for varying topo…
Align3R: Aligned Monocular Depth Estimation for Dynamic Videos
Jiahao Lu, Tianyu Huang, Peng Li +7
Recent developments in monocular depth estimation methods enable high-quality depth estimation of single-view images but fail to estimate consistent video depth across different fr…
Elastic Trapped States at Dislocation Defects in Scaled Coupling and Hofstadter Models
Yangkai Liu, Cheng Lin, Yuan Liu +4
Elastic topological dislocations provide a pathway for trapping elastic wave energy at internal defects, rather than being confined solely to external boundaries or corners, which…
Surf-D: Generating High-Quality Surfaces of Arbitrary Topologies Using Diffusion Models
Zhengming Yu, Zhiyang Dou, Xiaoxiao Long +9
We present Surf-D, a novel method for generating high-quality 3D shapes as Surfaces with arbitrary topologies using Diffusion models. Previous methods explored shape generation wit…
Wonder3D++: Cross-domain Diffusion for High-fidelity 3D Generation from a Single Image
Yuxiao Yang, Xiao-Xiao Long, Zhiyang Dou +7
In this work, we introduce \textbf{Wonder3D++}, a novel method for efficiently generating high-fidelity textured meshes from single-view images. Recent methods based on Score Disti…
NoRA: Nested Low-Rank Adaptation for Efficient Fine-Tuning Large Models
Cheng Lin, Lujun Li, Dezhi Li +3
In this paper, we introduce Nested Low-Rank Adaptation (NoRA), a novel approach to parameter-efficient fine-tuning that extends the capabilities of Low-Rank Adaptation (LoRA) techn…
TORE: Token Reduction for Efficient Human Mesh Recovery with Transformer
Zhiyang Dou, Qingxuan Wu, Cheng Lin +5
In this paper, we introduce a set of simple yet effective TOken REduction (TORE) strategies for Transformer-based Human Mesh Recovery from monocular images. Current SOTA performanc…
PDT: Point Distribution Transformation with Diffusion Models
Jionghao Wang, Cheng Lin, Yuan Liu +7
Point-based representations have consistently played a vital role in geometric data structures. Most point cloud learning and processing methods typically leverage the unordered an…
Coverage Axis: Inner Point Selection for 3D Shape Skeletonization
Zhiyang Dou, Cheng Lin, Rui Xu +4
In this paper, we present a simple yet effective formulation called Coverage Axis for 3D shape skeletonization. Inspired by the set cover problem, our key idea is to cover all the…
GausSurf: Geometry-Guided 3D Gaussian Splatting for Surface Reconstruction
Jiepeng Wang, Yuan Liu, Peng Wang +5
3D Gaussian Splatting has achieved impressive performance in novel view synthesis with real-time rendering capabilities. However, reconstructing high-quality surfaces with fine det…
Every Step Evolves: Scaling Reinforcement Learning for Trillion-Scale Thinking Model
Ling Team, Anqi Shen, Baihui Li +101
We present Ring-1T, the first open-source, state-of-the-art thinking model with a trillion-scale parameter. It features 1 trillion total parameters and activates approximately 50 b…
Spatiotemporal steering of photoelectron emission in multiphoton above-threshold ionization
Xiaochun Gong, Cheng Lin, Mingming Liu +11
We experimentally demonstrate spatiotemporal steering of photoelectron emission in multiphoton above-threshold single ionization of atoms exposed to a phase-controlled orthogonally…
FlexAM: Flexible Appearance-Motion Decomposition for Versatile Video Generation Control
Mingzhi Sheng, Zekai Gu, Peng Li +4
Effective and generalizable control in video generation remains a significant challenge. While many methods rely on ambiguous or task-specific signals, we argue that a fundamental…
mmTracking: Trajectory Tracking for Uplink mmWave Devices with Multi-Path Doppler Difference of Arrival
Cheng Lin, Chao Yu, Xiaowei Xu +1
This paper presents a method, namely mmTracking, for device trajectory tracking in a millimeter wave (mmWave) communication system. In mmTracking, the base station (BS) relies on o…
Era3D: High-Resolution Multiview Diffusion using Efficient Row-wise Attention
Peng Li, Yuan Liu, Xiaoxiao Long +10
In this paper, we introduce Era3D, a novel multiview diffusion method that generates high-resolution multiview images from a single-view image. Despite significant advancements in…
Fixed and periodic points of the intersection body operators of lower orders
Cheng Lin, Ge Xiong
For the intersection body operator of lower order of a star body in , , we prove that iff is an origin-symmetri…
Floorplan-Jigsaw: Jointly Estimating Scene Layout and Aligning Partial Scans
Cheng Lin, Changjian Li, Wenping Wang
We present a novel approach to align partial 3D reconstructions which may not have substantial overlap. Using floorplan priors, our method jointly predicts a room layout and estima…
Gen6D: Generalizable Model-Free 6-DoF Object Pose Estimation from RGB Images
Yuan Liu, Yilin Wen, Sida Peng +4
In this paper, we present a generalizable model-free 6-DoF object pose estimator called Gen6D. Existing generalizable pose estimators either need high-quality object models or requ…
WonderHuman: Hallucinating Unseen Parts in Dynamic 3D Human Reconstruction
Zilong Wang, Zhiyang Dou, Yuan Liu +7
In this paper, we present WonderHuman to reconstruct dynamic human avatars from a monocular video for high-fidelity novel view synthesis. Previous dynamic human avatar reconstructi…
Learnable Motion Coherence for Correspondence Pruning
Yuan Liu, Lingjie Liu, Cheng Lin +2
Motion coherence is an important clue for distinguishing true correspondences from false ones. Modeling motion coherence on sparse putative correspondences is challenging due to th…
UniRecGen: Unifying Multi-View 3D Reconstruction and Generation
Zhisheng Huang, Jiahao Chen, Cheng Lin +10
Sparse-view 3D modeling represents a fundamental tension between reconstruction fidelity and generative plausibility. While feed-forward reconstruction excels in efficiency and inp…
Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control
Zekai Gu, Rui Yan, Jiahao Lu +9
Diffusion models have demonstrated impressive performance in generating high-quality videos from text prompts or images. However, precise control over the video generation process,…
Topological Rainbow Trapping for Spatial-frequency Demultiplexing of Underwater Acoustic Signals
Cheng Lin, Yangkai Liu, Tuo Liu +3
Efficient separation and localization of multifrequency acoustic waves are essential for underwater target recognition and acoustic energy harvesting. The underwater implementation…
Wonder3D: Single Image to 3D using Cross-Domain Diffusion
Xiaoxiao Long, Yuan-Chen Guo, Cheng Lin +8
In this work, we introduce Wonder3D, a novel method for efficiently generating high-fidelity textured meshes from single-view images.Recent methods based on Score Distillation Samp…
GaussiAnimate: Reconstruct and Rig Animatable Categories with Level of Dynamics
Jiaxin Wang, Dongxin Lyu, Zeyu Cai +4
Free-form bones, that conform closely to the surface, can effectively capture non-rigid deformations, but lack a kinematic structure necessary for intuitive control. Thus, we propo…
MeshMosaic: Scaling Artist Mesh Generation via Local-to-Global Assembly
Rui Xu, Tianyang Xue, Qiujie Dong +9
Scaling artist-designed meshes to high triangle numbers remains challenging for autoregressive generative models. Existing transformer-based methods suffer from long-sequence bottl…
RefAny3D: 3D Asset-Referenced Diffusion Models for Image Generation
Hanzhuo Huang, Qingyang Bao, Zekai Gu +4
In this paper, we propose a 3D asset-referenced diffusion model for image generation, exploring how to integrate 3D assets into image diffusion models. Existing reference-based ima…
Multi-Type Context-Aware Conversational Recommender Systems via Mixture-of-Experts
Jie Zou, Cheng Lin, Weikang Guo +4
Conversational recommender systems enable natural language conversations and thus lead to a more engaging and effective recommendation scenario. As the conversations for recommende…
QuadLink: Autoregressive Quad-Dominant Mesh Generation via Point-Relation Learning
Yiheng Zhang, Zhe Zhu, Tingrui Shen +11
The generation of production-ready quad-dominant meshes is a cornerstone of modern 3D content creation. Generating anisotropic quad-dominant meshes from point clouds is challenging…
Unraveling nonadiabatic ionization and Coulomb potential effects in strong-field photoelectron holography
Xiaohong Song, Cheng Lin, Zhihao Sheng +6
Strong field photoelectron holography has been proposed as a means for interrogating the spatial and temporal information of electrons and ions in a dynamic system. After ionizatio…
Navigating Unmeasured Confounding in Quantitative Sociology: A Sensitivity Framework
Cheng Lin, Jose M. Pena, Adel Daoud
Unmeasured confounding remains a critical challenge in causal inference for the social sciences. This paper proposes a sensitivity analysis framework to systematically evaluate how…
Experimental Realization of Type-II Quadrupole Topological Insulator
Yuan Liu, Yangkai Liu, Cheng Lin +4
The discovery of quadrupole topological insulators (QTIs) has spurred extensive research into higher-order topological phases. Recently proposed type-II QTIs exhibit unconventional…
Momentum mapping of continuum electron wave packet interference
Weifeng Yang, Huatang Zhang, Cheng Lin +5
We analyze the two-dimensional photoelectrons momentum distribution of Ar atom ionized by midinfrared laser pulses and mainly concentrate on the energy range below 2Up. By using a…
Adaptive Surface Normal Constraint for Geometric Estimation from Monocular Images
Xiaoxiao Long, Yuhang Zheng, Yupeng Zheng +6
We introduce a novel approach to learn geometries such as depth and surface normal from images while incorporating geometric context. The difficulty of reliably capturing geometric…
Modeling 3D Shapes by Reinforcement Learning
Cheng Lin, Tingxiang Fan, Wenping Wang +1
We explore how to enable machines to model 3D shapes like human modelers using deep reinforcement learning (RL). In 3D modeling software like Maya, a modeler usually creates a mesh…
SparseNeuS: Fast Generalizable Neural Surface Reconstruction from Sparse Views
Xiaoxiao Long, Cheng Lin, Peng Wang +2
We introduce SparseNeuS, a novel neural rendering based method for the task of surface reconstruction from multi-view images. This task becomes more difficult when only sparse imag…
CelloCut: Constructive Watertight Remeshing via Tetrahedral Cell Cuts
Xuan Yang, Yuhang Zeng, Dinglong Fang +6
Watertight remeshing aims to recover a surface that induces a globally consistent interior--exterior partition of 3D space. However, for meshes with complex topology, single-layer…
Attosecond Interference Induced by Coulomb-Field-Driven Transverse Backward-Scattering Electron Wave-Packets
Xiaohong Song, Peng Liu, Cheng Lin +9
A novel and universal interference structure is found in the photoelectron momentum distribution of atoms in intense infrared laser field. Theoretical analysis shows that this stru…
ABench-Physics: Benchmarking Physical Reasoning in LLMs via High-Difficulty and Dynamic Physics Problems
Yiming Zhang, Yingfan Ma, Yanmei Gu +9
Large Language Models (LLMs) have shown impressive performance in domains such as mathematics and programming, yet their capabilities in physics remain underexplored and poorly und…
SyncDreamer: Generating Multiview-consistent Images from a Single-view Image
Yuan Liu, Cheng Lin, Zijiao Zeng +4
In this paper, we present a novel diffusion model called that generates multiview-consistent images from a single-view image. Using pretrained large-scale 2D diffusion models, rece…
NeuralUDF: Learning Unsigned Distance Fields for Multi-view Reconstruction of Surfaces with Arbitrary Topologies
Xiaoxiao Long, Cheng Lin, Lingjie Liu +5
We present a novel method, called NeuralUDF, for reconstructing surfaces with arbitrary topologies from 2D images via volume rendering. Recent advances in neural rendering based re…
CADDreamer: CAD Object Generation from Single-view Images
Yuan Li, Cheng Lin, Yuan Liu +6
Diffusion-based 3D generation has made remarkable progress in recent years. However, existing 3D generative models often produce overly dense and unstructured meshes, which stand i…
DICE: End-to-end Deformation Capture of Hand-Face Interactions from a Single Image
Qingxuan Wu, Zhiyang Dou, Sirui Xu +11
Reconstructing 3D hand-face interactions with deformations from a single image is a challenging yet crucial task with broad applications in AR, VR, and gaming. The challenges stem…
NASM: Neural Anisotropic Surface Meshing
Hongbo Li, Haikuan Zhu, Sikai Zhong +7
This paper introduces a new learning-based method, NASM, for anisotropic surface meshing. Our key idea is to propose a graph neural network to embed an input mesh into a high-dimen…
NeRO: Neural Geometry and BRDF Reconstruction of Reflective Objects from Multiview Images
Yuan Liu, Peng Wang, Cheng Lin +5
We present a neural rendering-based method called NeRO for reconstructing the geometry and the BRDF of reflective objects from multiview images captured in an unknown environment.…
Low-Energy Interference Structure with Attosecond Temporal Resolution
Xiaohong Song, Wenbin Jia, Xiwang Liu +7
Accessing precisely to the phase variation of electronic wave-packet (EWP) provides unprecedented spatiotemporal information of microworld. A radial interference pattern at near-ze…
Adaptive Compact Attention For Few-shot Video-to-video Translation
Risheng Huang, Li Shen, Xuan Wang +2
This paper proposes an adaptive compact attention model for few-shot video-to-video translation. Existing works in this domain only use features from pixel-wise attention without c…
To What Extent Do Disadvantaged Neighborhoods Mediate Social Assistance Dependency? Evidence from Sweden
Cheng Lin, Adel Daoud, Maria Branden
Occasional social assistance prevents individuals from a range of social ills, particularly unemployment and poverty. It remains unclear, however, how and to what extent continued…
LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging
Zhijian Shu, Cheng Lin, Tao Xie +8
3D vision foundation models like Visual Geometry Grounded Transformer (VGGT) have advanced greatly in geometric perception. However, it is time-consuming and memory-intensive for l…
Adaptive Surface Normal Constraint for Depth Estimation
Xiaoxiao Long, Cheng Lin, Lingjie Liu +4
We present a novel method for single image depth estimation using surface normal constraints. Existing depth estimation methods either suffer from the lack of geometric constraints…
Coverage Axis++: Efficient Inner Point Selection for 3D Shape Skeletonization
Zimeng Wang, Zhiyang Dou, Rui Xu +7
We introduce Coverage Axis++, a novel and efficient approach to 3D shape skeletonization. The current state-of-the-art approaches for this task often rely on the watertightness of…
Seed1.5-VL Technical Report
Dong Guo, Faming Wu, Feida Zhu +194
We present Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning. Seed1.5-VL is composed with a 532M-parameter v…
MOSPA: Human Motion Generation Driven by Spatial Audio
Shuyang Xu, Zhiyang Dou, Mingyi Shi +8
Enabling virtual humans to dynamically and realistically respond to diverse auditory stimuli remains a key challenge in character animation, demanding the integration of perceptual…
HGGT: Robust and Flexible 3D Hand Mesh Reconstruction from Uncalibrated Images
Yumeng Liu, Xiao-Xiao Long, Marc Habermann +6
Recovering high-fidelity 3D hand geometry from images is a critical task in computer vision, holding significant value for domains such as robotics, animation and VR/AR. Crucially,…
SEG-MAT: 3D Shape Segmentation Using Medial Axis Transform
Cheng Lin, Lingjie Liu, Changjian Li +4
Segmenting arbitrary 3D objects into constituent parts that are structurally meaningful is a fundamental problem encountered in a wide range of computer graphics applications. Exis…
PartSAM: A Scalable Promptable Part Segmentation Model Trained on Native 3D Data
Zhe Zhu, Le Wan, Rui Xu +6
Segmenting 3D objects into parts is a long-standing challenge in computer vision. To overcome taxonomy constraints and generalize to unseen 3D objects, recent works turn to open-wo…
Part123: Part-aware 3D Reconstruction from a Single-view Image
Anran Liu, Cheng Lin, Yuan Liu +5
Recently, the emergence of diffusion models has opened up new opportunities for single-view reconstruction. However, all the existing methods represent the target object as a close…
Point2Skeleton: Learning Skeletal Representations from Point Clouds
Cheng Lin, Changjian Li, Yuan Liu +3
We introduce Point2Skeleton, an unsupervised method to learn skeletal representations from point clouds. Existing skeletonization methods are limited to tubular shapes and the stri…
Chiral Landau levels induced by two in-plane pseudomagnetic fields in underwater acoustic metamaterials
Jiao Shen, Zhiyong Chang, Xinzong Wang +6
The paper demonstrates how to create and control chiral zeroth Landau levels in underwater acoustic metamaterials by using two perpendicular in‑plane pseudomagnetic fields, enablin…
Disentangled Clothed Avatar Generation from Text Descriptions
Jionghao Wang, Yuan Liu, Zhiyang Dou +7
In this paper, we introduce a novel text-to-avatar generation method that separately generates the human body and the clothes and allows high-quality animation on the generated ava…
SVGS: Enhancing Gaussian Splatting Using Primitives with Spatially Varying Colors
Rui Xu, Wenyue Chen, Jiepeng Wang +7
Gaussian Splatting demonstrates impressive results in multi-view reconstruction based on Gaussian explicit representations. However, the current Gaussian primitives only have a sin…
DecoRec: Decomposed 3D Scene Reconstruction from Single-View Images via Object-Level Diffusion
Yuhan Ping, Yuan Liu, Xiaoxiao Long +6
In this paper, we introduce \textit{DecoRec}, a novel system designed to elevate single-view 2D images to a decomposed 3D scene mesh. Current methods for single-view scene reconstr…
TrackingWorld: World-centric Monocular 3D Tracking of Almost All Pixels
Jiahao Lu, Weitao Xiong, Jiacheng Deng +6
Monocular 3D tracking aims to capture the long-term motion of pixels in 3D space from a single monocular video and has witnessed rapid progress in recent years. However, we argue t…
Domain-invariant Representation Learning via Segment Anything Model for Blood Cell Classification
Yongcheng Li, Lingcong Cai, Ying Lu +8
Accurate classification of blood cells is of vital significance in the diagnosis of hematological disorders. However, in real-world scenarios, domain shifts caused by the variabili…