Publications (204)
Artemis: Articulated Neural Pets with Appearance and Motion synthesis
Haimin Luo, Teng Xu, Yuheng Jiang +6
We, humans, are entering into a virtual era and indeed want to bring animals to the virtual world as well for companion. Yet, computer-generated (CGI) furry animals are limited by…
Free-Bloom: Zero-Shot Text-to-Video Generator with LLM Director and LDM Animator
Hanzhuo Huang, Yufan Feng, Cheng Shi +3
Text-to-video is a rapidly growing research area that aims to generate a semantic, identical, and temporal coherence sequence of frames that accurately align with the input text pr…
MirrorNeRF: One-shot Neural Portrait Radiance Field from Multi-mirror Catadioptric Imaging
Ziyu Wang, Liao Wang, Fuqiang Zhao +3
Photo-realistic neural reconstruction and rendering of the human portrait are critical for numerous VR/AR applications. Still, existing solutions inherently rely on multi-view capt…
Dissecting and Mitigating Diffusion Bias via Mechanistic Interpretability
Yingdong Shi, Changming Li, Yifan Wang +5
Diffusion models have demonstrated impressive capabilities in synthesizing diverse content. However, despite their high-quality outputs, these models often perpetuate social biases…
HiFi4G: High-Fidelity Human Performance Rendering via Compact Gaussian Splatting
Yuheng Jiang, Zhehao Shen, Penghao Wang +5
We have recently seen tremendous progress in photo-real human modeling and rendering. Yet, efficiently rendering realistic human performance and integrating it into the rasterizati…
Occlusion-Model Guided Anti-Occlusion Depth Estimation in Light Field
Hao Zhu, Qing Wang, Jingyi Yu
Occlusion is one of the most challenging problems in depth estimation. Previous work has modeled the single-occluder occlusion in light field and get good results, however it is st…
Towards Immersive Human-X Interaction: A Real-Time Framework for Physically Plausible Motion Synthesis
Kaiyang Ji, Ye Shi, Zichen Jin +5
Real-time synthesis of physically plausible human interactions remains a critical challenge for immersive VR/AR systems and humanoid robotics. While existing methods demonstrate pr…
OMG: Towards Open-vocabulary Motion Generation via Mixture of Controllers
Han Liang, Jiacheng Bao, Ruichi Zhang +6
We have recently seen tremendous progress in realistic text-to-motion generation. Yet, the existing methods often fail or produce implausible motions with unseen text inputs, which…
Automatic Layer Separation using Light Field Imaging
Qiaosong Wang, Haiting Lin, Yi Ma +2
We propose a novel approach that jointly removes reflection or translucent layer from a scene and estimates scene depth. The input data are captured via light field imaging. The pr…
Disentangling Light Fields for Super-Resolution and Disparity Estimation
Yingqian Wang, Longguang Wang, Gaochang Wu +4
Light field (LF) cameras record both intensity and directions of light rays, and encode 3D scenes into 4D LF images. Recently, many convolutional neural networks (CNNs) have been p…
Neural Residual Radiance Fields for Streamably Free-Viewpoint Videos
Liao Wang, Qiang Hu, Qihan He +5
The success of the Neural Radiance Fields (NeRFs) for modeling and free-view rendering static objects has inspired numerous attempts on dynamic scenes. Current techniques that util…
A Precise Calculation of Delayed Coincidence Selection Efficiency and Accidental Coincidence Rate
Jingyi Yu, Zhe Wang, Shaomin Chen
A model is proposed to address issues on the precise background evaluation due to the complex data structure defined by the delayed coincidence method, which is widely used in reac…
SPREAD: Spatial-Physical REasoning via geometry Aware Diffusion
Minzhang Li, Kuixiang Shao, Xuebing Li +6
Automated 3D scene generation is pivotal for applications spanning virtual reality, digital content creation, and Embodied AI. While computer graphics prioritizes aesthetic layouts…
ManiTwin: Scaling Data-Generation-Ready Digital Object Dataset to 100K
Kaixuan Wang, Tianxing Chen, Jiawei Liu +13
Learning in simulation provides a useful foundation for scaling robotic manipulation capabilities. However, this paradigm often suffers from a lack of data-generation-ready digital…
BANG: Dividing 3D Assets via Generative Exploded Dynamics
Longwen Zhang, Qixuan Zhang, Haoran Jiang +4
3D creation has always been a unique human strength, driven by our ability to deconstruct and reassemble objects using our eyes, mind and hand. However, current 3D design tools str…
Omni-Line-of-Sight Imaging for Holistic Shape Reconstruction
Binbin Huang, Xingyue Peng, Siyuan Shen +8
We introduce Omni-LOS, a neural computational imaging method for conducting holistic shape reconstruction (HSR) of complex objects utilizing a Single-Photon Avalanche Diode (SPAD)-…
IREM: High-Resolution Magnetic Resonance (MR) Image Reconstruction via Implicit Neural Representation
Qing Wu, Yuwei Li, Lan Xu +7
For collecting high-quality high-resolution (HR) MR image, we propose a novel image reconstruction network named IREM, which is trained on multiple low-resolution (LR) MR images an…
BG-Triangle: Bézier Gaussian Triangle for 3D Vectorization and Rendering
Minye Wu, Haizhao Dai, Kaixin Yao +2
Differentiable rendering enables efficient optimization by allowing gradients to be computed through the rendering process, facilitating 3D reconstruction, inverse rendering and ne…
LiDARCap: Long-range Marker-less 3D Human Motion Capture with LiDAR Point Clouds
Jialian Li, Jingyi Zhang, Zhiyong Wang +6
Existing motion capture datasets are largely short-range and cannot yet fit the need of long-range applications. We propose LiDARHuman26M, a new human motion capture dataset captur…
Implicit Swept Volume SDF: Enabling Continuous Collision-Free Trajectory Generation for Arbitrary Shapes
Jingping Wang, Tingrui Zhang, Qixuan Zhang +5
In the field of trajectory generation for objects, ensuring continuous collision-free motion remains a huge challenge, especially for non-convex geometries and complex environments…
Single-pixel p-graded-n junction spectrometers
Jingyi Wang, Beibei Pan, Zi Wang +8
Ultra-compact spectrometers are becoming increasingly popular for their promising applications in biomedical analysis, environmental monitoring, and food safety. In this work, we r…
Neural Video Portrait Relighting in Real-time via Consistency Modeling
Longwen Zhang, Qixuan Zhang, Minye Wu +2
Video portraits relighting is critical in user-facing human photography, especially for immersive VR/AR experience. Recent advances still fail to recover consistent relit result un…
PIANO: A Parametric Hand Bone Model from Magnetic Resonance Imaging
Yuwei Li, Minye Wu, Yuyao Zhang +2
Hand modeling is critical for immersive VR/AR, action understanding, or human healthcare. Existing parametric models account only for hand shape, pose, or texture, without modeling…
A Neural Rendering Framework for Free-Viewpoint Relighting
Zhang Chen, Anpei Chen, Guli Zhang +4
We present a novel Relightable Neural Renderer (RNR) for simultaneous view synthesis and relighting using multi-view image inputs. Existing neural rendering (NR) does not explicitl…
Sparse-view Signal-domain Photoacoustic Tomography Reconstruction Method Based on Neural Representation
Bowei Yao, Yi Zeng, Haizhao Dai +6
Photoacoustic tomography is a hybrid biomedical technology, which combines the advantages of acoustic and optical imaging. However, for the conventional image reconstruction method…
Generic Multiview Visual Tracking
Minye Wu, Haibin Ling, Ning Bi +3
Recent progresses in visual tracking have greatly improved the tracking performance. However, challenges such as occlusion and view change remain obstacles in real world deployment…
A Unified Diffusion Framework for Scene-aware Human Motion Estimation from Sparse Signals
Jiangnan Tang, Jingya Wang, Kaiyang Ji +3
Estimating full-body human motion via sparse tracking signals from head-mounted displays and hand controllers in 3D scenes is crucial to applications in AR/VR. One of the biggest c…
TightCap: 3D Human Shape Capture with Clothing Tightness Field
Xin Chen, Anqi Pang, Yang Wei +2
In this paper, we present TightCap, a data-driven scheme to capture both the human shape and dressed garments accurately with only a single 3D human scan, which enables numerous ap…
SMGDiff: Soccer Motion Generation using diffusion probabilistic models
Hongdi Yang, Chengyang Li, Zhenxuan Wu +5
Soccer is a globally renowned sport with significant applications in video games and VR/AR. However, generating realistic soccer motions remains challenging due to the intricate in…
Editable Free-viewpoint Video Using a Layered Neural Representation
Jiakai Zhang, Xinhang Liu, Xinyi Ye +6
Generating free-viewpoint videos is critical for immersive VR/AR experience but recent neural advances still lack the editing ability to manipulate the visual perception for large…
A Generic Multi-Projection-Center Model and Calibration Method for Light Field Cameras
Qi Zhang, Chunping Zhang, Jinbo Ling +2
Light field cameras can capture both spatial and angular information of light rays, enabling 3D reconstruction by a single exposure. The geometry of 3D reconstruction is affected b…
Robust Guided Image Filtering
Wei Liu, Xiaogang Chen, Chunhua Shen +3
The process of using one image to guide the filtering process of another one is called Guided Image Filtering (GIF). The main challenge of GIF is the structure inconsistency betwee…
PanoVGGT: Feed-Forward 3D Reconstruction from Panoramic Imagery
Yijing Guo, Mengjun Chao, Luo Wang +5
Panoramic imagery offers a full 360° field of view and is increasingly common in consumer devices. However, it introduces non-pinhole distortions that challenge joint pose estimat…
UniDB: A Unified Diffusion Bridge Framework via Stochastic Optimal Control
Kaizhen Zhu, Mokai Pan, Yuexin Ma +4
Recent advances in diffusion bridge models leverage Doob's -transform to establish fixed endpoints between distributions, demonstrating promising results in image translation an…
A scaling law for large-deformation contact in soft materials
Tong Mu, Shizhuo Weng, Changhong Linghu +10
Compression of soft bodies is central to biology, materials science, and robotics, yet existing contact theories break down at large deformations. Here, we develop a general framew…
FUSE: FK-Steered Multi-Modal Flow Matching for Efficient Simulation-Based Posterior Estimation
Weichen Qin, Yufan Xie, Peihao Wang +8
Simulation-Based Inference (SBI) is critical for scientific discovery, with generative models offering a promising path toward efficient inference. However, existing methods strugg…
NeMF: Inverse Volume Rendering with Neural Microflake Field
Youjia Zhang, Teng Xu, Junqing Yu +5
Recovering the physical attributes of an object's appearance from its images captured under an unknown illumination is challenging yet essential for photo-realistic rendering. Rece…
ExFace: Expressive Facial Control for Humanoid Robots with Diffusion Transformers and Bootstrap Training
Dong Zhang, Jingwei Peng, Yuyang Jiao +3
This paper presents a novel Expressive Facial Control (ExFace) method based on Diffusion Transformers, which achieves precise mapping from human facial blendshapes to bionic robot…
IKOL: Inverse kinematics optimization layer for 3D human pose and shape estimation via Gauss-Newton differentiation
Juze Zhang, Ye Shi, Yuexin Ma +3
This paper presents an inverse kinematic optimization layer (IKOL) for 3D human pose and shape estimation that leverages the strength of both optimization- and regression-based met…
TensoRF: Tensorial Radiance Fields
Anpei Chen, Zexiang Xu, Andreas Geiger +2
We present TensoRF, a novel approach to model and reconstruct radiance fields. Unlike NeRF that purely uses MLPs, we model the radiance field of a scene as a 4D tensor, which repre…
CylinderSplat: 3D Gaussian Splatting with Cylindrical Triplanes for Panoramic Novel View Synthesis
Qiwei Wang, Xianghui Ze, Jingyi Yu +1
Feed-forward 3D Gaussian Splatting (3DGS) has shown great promise for real-time novel view synthesis, but its application to panoramic imagery remains challenging. Existing methods…
CLAY: A Controllable Large-scale Generative Model for Creating High-quality 3D Assets
Longwen Zhang, Ziyu Wang, Qixuan Zhang +6
In the realm of digital creativity, our potential to craft intricate 3D worlds from imagination is often hampered by the limitations of existing digital tools, which demand extensi…
HybridCap: Inertia-aid Monocular Capture of Challenging Human Motions
Han Liang, Yannan He, Chengfeng Zhao +4
Monocular 3D motion capture (mocap) is beneficial to many applications. The use of a single camera, however, often fails to handle occlusions of different body parts and hence it i…
AerialGo: Walking-through City View Generation from Aerial Perspectives
Fuqiang Zhao, Yijing Guo, Siyuan Yang +6
High-quality 3D urban reconstruction is essential for applications in urban planning, navigation, and AR/VR. However, capturing detailed ground-level data across cities is both lab…
Relightable Neural Human Assets from Multi-view Gradient Illuminations
Taotao Zhou, Kai He, Di Wu +6
Human modeling and relighting are two fundamental problems in computer vision and graphics, where high-quality datasets can largely facilitate related research. However, most exist…
Resolving Scale Ambiguity Via XSlit Aspect Ratio Analysis
Wei Yang, Haiting Lin, Sing Bing Kang +1
In perspective cameras, images of a frontal-parallel 3D object preserve its aspect ratio invariant to its depth. Such an invariance is useful in photography but is unique to perspe…
DexGrasp Anything: Towards Universal Robotic Dexterous Grasping with Physics Awareness
Yiming Zhong, Qi Jiang, Jingyi Yu +1
A dexterous hand capable of grasping any object is essential for the development of general-purpose embodied intelligent robots. However, due to the high degree of freedom in dexte…
Personalized Saliency and its Prediction
Yanyu Xu, Shenghua Gao, Junru Wu +2
Nearly all existing visual saliency models by far have focused on predicting a universal saliency map across all observers. Yet psychology studies suggest that visual attention of…
V^3: Viewing Volumetric Videos on Mobiles via Streamable 2D Dynamic Gaussians
Penghao Wang, Zhirui Zhang, Liao Wang +5
Experiencing high-fidelity volumetric video as seamlessly as 2D videos is a long-held dream. However, current dynamic 3DGS methods, despite their high rendering quality, face chall…
NeuRBF: A Neural Fields Representation with Adaptive Radial Basis Functions
Zhang Chen, Zhong Li, Liangchen Song +4
We present a novel type of neural fields that uses general radial bases for signal representation. State-of-the-art neural fields typically rely on grid-based representations for s…
DexH2R: A Benchmark for Dynamic Dexterous Grasping in Human-to-Robot Handover
Youzhuo Wang, Jiayi Ye, Chuyang Xiao +6
Handover between a human and a dexterous robotic hand is a fundamental yet challenging task in human-robot collaboration. It requires handling dynamic environments and a wide varie…
SPIDER: Structure-Preferential Implicit Deep Network for Biplanar X-ray Reconstruction
Tianqi Yu, Xuanyu Tian, Jiawen Yang +4
Biplanar X-ray imaging is widely used in health screening, postoperative rehabilitation evaluation of orthopedic diseases, and injury surgery due to its rapid acquisition, low radi…
CADSpotting: Robust Panoptic Symbol Spotting on Large-Scale CAD Drawings
Fuyi Yang, Jiazuo Mu, Yanshun Zhang +7
We introduce CADSpotting, an effective method for panoptic symbol spotting in large-scale architectural CAD drawings. Existing approaches often struggle with symbol diversity, scal…
DreamPrinting: Volumetric Printing Primitives for High-Fidelity 3D Printing
Youjia Wang, Ruixiang Cao, Teng Xu +4
Translating the rich visual fidelity of volumetric rendering techniques into physically realizable 3D prints remains an open challenge. We introduce DreamPrinting, a novel pipeline…
NIMBLE: A Non-rigid Hand Model with Bones and Muscles
Yuwei Li, Longwen Zhang, Zesong Qiu +6
Emerging Metaverse applications demand reliable, accurate, and photorealistic reproductions of human hands to perform sophisticated operations as if in the physical world. While re…
Light Field-Based Underwater 3D Reconstruction Via Angular Resampling
Yuqi Ding, Zhang Chen, Yu Ji +2
Recovering 3D geometry of underwater scenes is challenging because of non-linear refraction of light at the water-air interface caused by the camera housing. We present a light fie…
Improving 2D Diffusion Models for 3D Medical Imaging with Inter-Slice Consistent Stochasticity
Chenhe Du, Qing Wu, Xuanyu Tian +3
3D medical imaging is in high demand and essential for clinical diagnosis and scientific research. Currently, diffusion models (DMs) have become an effective tool for medical imagi…
Guidance with Spherical Gaussian Constraint for Conditional Diffusion
Lingxiao Yang, Shutong Ding, Yifan Cai +3
Recent advances in diffusion models attempt to handle conditional generative tasks by utilizing a differentiable loss function for guidance without the need for additional training…
LiveHPS: LiDAR-based Scene-level Human Pose and Shape Estimation in Free Environment
Yiming Ren, Xiao Han, Chengfeng Zhao +4
For human-centric large-scale scenes, fine-grained modeling for 3D human global pose and shape is significant for scene understanding and can benefit many real-world applications.…
CryoACE: An Atom-centric Framework for Accurate and Automated Model Building in Cryo-EM
Minzhang Li, Mingrui Li, Weichen Qin +5
Protein automodeling from cryo-EM density maps faces unique challenges in enforcing physicochemical validity and managing conformational heterogeneity. Current solvers are often li…
ScalableMap: Scalable Map Learning for Online Long-Range Vectorized HD Map Construction
Jingyi Yu, Zizhao Zhang, Shengfu Xia +1
We propose a novel end-to-end pipeline for online long-range vectorized high-definition (HD) map construction using on-board camera sensors. The vectorized representation of HD map…
CityGo: Lightweight Urban Modeling and Rendering with Proxy Buildings and Residual Gaussians
Weihang Liu, Yuhui Zhong, Yuke Li +8
Accurate and efficient modeling of large-scale urban scenes is critical for applications such as AR navigation, UAV based inspection, and smart city digital twins. While aerial ima…
Capturing the Unseen: Vision-Free Facial Motion Capture Using Inertial Measurement Units
Youjia Wang, Yiwen Wu, Hengan Zhou +9
We present Capturing the Unseen (CAPUS), a novel facial motion capture (MoCap) technique that operates without visual signals. CAPUS leverages miniaturized Inertial Measurement Uni…
Non-line-of-Sight Imaging via Neural Transient Fields
Siyuan Shen, Zi Wang, Ping Liu +5
We present a neural modeling framework for Non-Line-of-Sight (NLOS) imaging. Previous solutions have sought to explicitly recover the 3D geometry (e.g., as point clouds) or voxel d…
Semantic See-Through Rendering on Light Fields
Huangjie Yu, Guli Zhang, Yuanxi Ma +2
We present a novel semantic light field (LF) refocusing technique that can achieve unprecedented see-through quality. Different from prior art, our semantic see-through (SST) diffe…
Causal Mechanism Estimation in Multi-Sensor Systems Across Multiple Domains
Jingyi Yu, Tim Pychynski, Marco F. Huber
To gain deeper insights into a complex sensor system through the lens of causality, we present common and individual causal mechanism estimation (CICME), a novel three-step approac…
Generative Deformable Radiance Fields for Disentangled Image Synthesis of Topology-Varying Objects
Ziyu Wang, Yu Deng, Jiaolong Yang +2
3D-aware generative models have demonstrated their superb performance to generate 3D neural radiance fields (NeRF) from a collection of monocular 2D images even for topology-varyin…
Robust Dual Gaussian Splatting for Immersive Human-centric Volumetric Videos
Yuheng Jiang, Zhehao Shen, Yu Hong +5
Volumetric video represents a transformative advancement in visual media, enabling users to freely navigate immersive virtual experiences and narrowing the gap between digital and…
NeuralHOFusion: Neural Volumetric Rendering under Human-object Interactions
Yuheng Jiang, Suyi Jiang, Guoxing Sun +5
4D modeling of human-object interactions is critical for numerous applications. However, efficient volumetric capture and rendering of complex interaction scenarios, especially fro…
Autoregressive B-Rep Shape Generation with Parametric Surfaces
Dafei Qin, Rui Xu, Zeyu Shen +8
Generative CAD modeling has broad design and application potential. Despite significant advances in Boundary Representation (B-Rep) generation, the dominant representation in CAD,…
BEAM: Bridging Physically-based Rendering and Gaussian Modeling for Relightable Volumetric Video
Yu Hong, Yize Wu, Zhehao Shen +5
Volumetric video enables immersive experiences by capturing dynamic 3D scenes, enabling diverse applications for virtual reality, education, and telepresence. However, traditional…
AffordDP: Generalizable Diffusion Policy with Transferable Affordance
Shijie Wu, Yihang Zhu, Yunao Huang +5
Diffusion-based policies have shown impressive performance in robotic manipulation tasks while struggling with out-of-domain distributions. Recent efforts attempted to enhance gene…
GaussianHair: Hair Modeling and Rendering with Light-aware Gaussians
Haimin Luo, Min Ouyang, Zijun Zhao +6
Hairstyle reflects culture and ethnicity at first glance. In the digital era, various realistic human hairstyles are also critical to high-fidelity digital human assets for beauty…
Hyperspectral Light Field Stereo Matching
Kang Zhu, Yujia Xue, Qiang Fu +3
In this paper, we describe how scene depth can be extracted using a hyperspectral light field capture (H-LF) system. Our H-LF system consists of a 5 x 6 array of cameras, with each…
TAPESTRY: From Geometry to Appearance via Consistent Turntable Videos
Yan Zeng, Haoran Jiang, Kaixin Yao +4
Automatically generating photorealistic and self-consistent appearances for untextured 3D models is a critical challenge in digital content creation. The advancement of large-scale…
CryoGEM: Physics-Informed Generative Cryo-Electron Microscopy
Jiakai Zhang, Qihe Chen, Yan Zeng +4
In the past decade, deep conditional generative models have revolutionized the generation of realistic images, extending their application from entertainment to scientific domains.…
ForeSplat: Optimization-Aware Foresight for Feed-Forward 3D Gaussian Splatting
Yuke Li, Weihang Liu, Cheng Zhang +8
Feed-forward 3D Gaussian Splatting models offer fast single-pass reconstruction,but scaling them to match per-scene optimization quality is fundamentally hindered by the scarcity o…
MotionGPT: Human Motion as a Foreign Language
Biao Jiang, Xin Chen, Wen Liu +3
Though the advancement of pre-trained large language models unfolds, the exploration of building a unified model for language and other multi-modal data, such as motion, remains ch…
LiDAR-aid Inertial Poser: Large-scale Human Motion Capture by Sparse Inertial and LiDAR Sensors
Yiming Ren, Chengfeng Zhao, Yannan He +5
We propose a multi-sensor fusion method for capturing challenging 3D human motions with accurate consecutive local poses and global trajectories in large-scale scenarios, only usin…
Media2Face: Co-speech Facial Animation Generation With Multi-Modality Guidance
Qingcheng Zhao, Pengyu Long, Qixuan Zhang +6
The synthesis of 3D facial animations from speech has garnered considerable attention. Due to the scarcity of high-quality 4D facial data and well-annotated abundant multi-modality…
The Ephemeral Shadow: Hyperreal Beings in Stimulative Performance
Dong Zhang, Yanjun Zhou, Jingyi Yu
The Ephemeral Shadow is an interactive art installation centered on the concept of "simulacrum," focusing on the reconstruction of subjectivity at the intersection of reality and v…
Optical ReLU-like Activation Function Based on a Semiconductor Laser with Optical Injection
Guanting Liu, Yiwei Shen, Ruiqian Li +3
Artificial neural networks usually consist of successive linear multiply-accumulate operations and nonlinear activation functions. However, most optical neural networks only achiev…
3D Face Reconstruction Using Color Photometric Stereo with Uncalibrated Near Point Lights
Zhang Chen, Yu Ji, Mingyuan Zhou +2
We present a new color photometric stereo (CPS) method that recovers high quality, detailed 3D face geometry in a single shot. Our system uses three uncalibrated near point lights…
PartNeXt: A Next-Generation Dataset for Fine-Grained and Hierarchical 3D Part Understanding
Penghao Wang, Yiyang He, Xin Lv +4
Understanding objects at the level of their constituent parts is fundamental to advancing computer vision, graphics, and robotics. While datasets like PartNet have driven progress…
CAST: Component-Aligned 3D Scene Reconstruction from an RGB Image
Kaixin Yao, Longwen Zhang, Xinhao Yan +6
Recovering high-quality 3D scenes from a single RGB image is a challenging task in computer graphics. Current methods often struggle with domain-specific limitations or low-quality…
MouseGPT: A Large-scale Vision-Language Model for Mouse Behavior Analysis
Teng Xu, Taotao Zhou, Youjia Wang +12
Analyzing animal behavior is crucial in advancing neuroscience, yet quantifying and deciphering its intricate dynamics remains a significant challenge. Traditional machine vision a…
Solving Energy-Independent Density for CT Metal Artifact Reduction via Neural Representation
Qing Wu, Xu Guo, Lixuan Chen +8
X-ray CT often suffers from shadowing and streaking artifacts in the presence of metallic materials, which severely degrade imaging quality. Physically, the linear attenuation coef…
DPER: Diffusion Prior Driven Neural Representation for Limited Angle and Sparse View CT Reconstruction
Chenhe Du, Xiyue Lin, Qing Wu +9
Limited-angle and sparse-view computed tomography (LACT and SVCT) are crucial for expanding the scope of X-ray CT applications. However, they face challenges due to incomplete data…
SCULPTOR: Skeleton-Consistent Face Creation Using a Learned Parametric Generator
Zesong Qiu, Yuwei Li, Dongming He +8
Recent years have seen growing interest in 3D human faces modelling due to its wide applications in digital human, character generation and animation. Existing approaches overwhelm…
Weakly Supervised 3D Multi-person Pose Estimation for Large-scale Scenes based on Monocular Camera and Single LiDAR
Peishan Cong, Yiteng Xu, Yiming Ren +5
Depth estimation is usually ill-posed and ambiguous for monocular camera-based 3D multi-person pose estimation. Since LiDAR can capture accurate depth information in long-range sce…
CoTDet: Affordance Knowledge Prompting for Task Driven Object Detection
Jiajin Tang, Ge Zheng, Jingyi Yu +1
Task driven object detection aims to detect object instances suitable for affording a task in an image. Its challenge lies in object categories available for the task being too div…
Evolution-Aware MSA Reasoning for Subsampling via Factor Graphs
Zhangzhi Xiong, Minzhang Li, Haotian Yu +6
Multiple Sequence Alignments (MSAs) provide protein language models with explicit evolutionary context, but their large depth makes subsampling unavoidable under limited token budg…
SCOPE: Sign Language Contextual Processing with Embedding from LLMs
Yuqi Liu, Wenqian Zhang, Sihan Ren +3
Sign languages, used by around 70 million Deaf individuals globally, are visual languages that convey visual and contextual information. Current methods in vision-based sign langua…
UniSteer: Text-Guided Flow Matching in Activation Space for Versatile LLM Steering
Yingdong Shi, Ruiming Zhang, Changming Li +4
Activation-based control steers large language models (LLMs) by intervening on their internal representations during inference, and has emerged as an effective paradigm for control…
Commanding Humanoid by Free-form Language: A Large Language Action Model with Unified Motion Vocabulary
Zhirui Liu, Kaiyang Ji, Ke Yang +4
Enabling humanoid robots to follow free-form natural language commands is a critical step toward seamless human-robot interaction and general-purpose embodied AI. However, existing…
MeshXL: Neural Coordinate Field for Generative 3D Foundation Models
Sijin Chen, Xin Chen, Anqi Pang +11
The polygon mesh representation of 3D data exhibits great flexibility, fast rendering speed, and storage efficiency, which is widely preferred in various applications. However, giv…
Human-centric Scene Understanding for 3D Large-scale Scenarios
Yiteng Xu, Peishan Cong, Yichen Yao +6
Human-centric scene understanding is significant for real-world applications, but it is extremely challenging due to the existence of diverse human poses and actions, complex human…
Fourier PlenOctrees for Dynamic Radiance Field Rendering in Real-time
Liao Wang, Jiakai Zhang, Xinhang Liu +6
Implicit neural representations such as Neural Radiance Field (NeRF) have focused mainly on modeling static objects captured under multi-view settings where real-time rendering can…
EvolvingGrasp: Evolutionary Grasp Generation via Efficient Preference Alignment
Yufei Zhu, Yiming Zhong, Zemin Yang +4
Dexterous robotic hands often struggle to generalize effectively in complex environments due to the limitations of models trained on low-diversity data. However, the real world pre…
Robust 3D Human Motion Reconstruction Via Dynamic Template Construction
Zhong Li, Yu Ji, Wei Yang +2
In multi-view human body capture systems, the recovered 3D geometry or even the acquired imagery data can be heavily corrupted due to occlusions, noise, limited field of- view, etc…