Publications (51)
Photonic unsupervised learning processor for secure and high-throughput optical fiber communication
Yitong Chen, Tiankuang Zhou, Jiamin Wu +4
Following the explosive growth of global data, there is an ever-increasing demand for high-throughput optical fiber communication (OFC) systems to perform massive data transmission…
SPI-Optimizer: an integral-Separated PI Controller for Stochastic Optimization
Dan Wang, Mengqi Ji, Yong Wang +2
To overcome the oscillation problem in the classical momentum-based optimizer, recent work associates it with the proportional-integral (PI) controller, and artificially adds D ter…
Revisiting Light Field Rendering with Deep Anti-Aliasing Neural Network
Gaochang Wu, Yebin Liu, Lu Fang +1
The light field (LF) reconstruction is mainly confronted with two challenges, large disparity and the non-Lambertian effect. Typical approaches either address the large disparity c…
Smart Cameras
David J. Brady, Minghao Hu, Chengyu Wang +6
We review camera architecture in the age of artificial intelligence. Modern cameras use physical components and software to capture, compress and display image data. Over the past…
XScale-NVS: Cross-Scale Novel View Synthesis with Hash Featurized Manifold
Guangyu Wang, Jinzhi Zhang, Fan Wang +2
We propose XScale-NVS for high-fidelity cross-scale novel view synthesis of real-world large-scale scenes. Existing representations based on explicit surface suffer from discretiza…
Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions
Bangsheng Tang, Carl Chengyan Fu, Fei Kou +35
Speculative decoding is a standard method for accelerating the inference speed of large language models. However, scaling it for production environments poses several engineering c…
RobustFusion: Robust Volumetric Performance Reconstruction under Human-object Interactions from Monocular RGBD Stream
Zhuo Su, Lan Xu, Dawei Zhong +4
High-quality 4D reconstruction of human performance with complex interactions to various objects is essential in real-world scenarios, which enables numerous immersive VR/AR applic…
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Adam Paszke, Sam Gross, Francisco Massa +18
Deep learning frameworks have often focused on either usability or speed, but not both. PyTorch is a machine learning library that shows that these two goals are in fact compatible…
TransGI: Real-Time Dynamic Global Illumination With Object-Centric Neural Transfer Model
Yijie Deng, Lei Han, Lu Fang
Neural rendering algorithms have revolutionized computer graphics, yet their impact on real-time rendering under arbitrary lighting conditions remains limited due to strict latency…
AI Pangaea: Unifying Intelligence Islands for Adapting Myriad Tasks
Jianlong Chang, Haixin Wang, Zhiyuan Dang +11
The pursuit of artificial general intelligence continuously demands generalization in one model across myriad tasks, even those not seen before. However, current AI models are isol…
ClickSeg: 3D Instance Segmentation with Click-Level Weak Annotations
Leyao Liu, Tao Kong, Minzhao Zhu +2
3D instance segmentation methods often require fully-annotated dense labels for training, which are costly to obtain. In this paper, we present ClickSeg, a novel click-level weakly…
Spatial-Angular Attention Network for Light Field Reconstruction
Gaochang Wu, Yingqian Wang, Yebin Liu +2
Typical learning-based light field reconstruction methods demand in constructing a large receptive field by deepening the network to capture correspondences between input views. In…
Den-SOFT: Dense Space-Oriented Light Field DataseT for 6-DOF Immersive Experience
Xiaohang Yu, Zhengxian Yang, Shi Pan +8
We have built a custom mobile multi-camera large-space dense light field capture system, which provides a series of high-quality and sufficiently dense light field images for vario…
MulayCap: Multi-layer Human Performance Capture Using A Monocular Video Camera
Zhaoqi Su, Weilin Wan, Tao Yu +4
We introduce MulayCap, a novel human performance capture method using a monocular video camera without the need for pre-scanning. The method uses "multi-layer" representations for…
OmniSeg3D: Omniversal 3D Segmentation via Hierarchical Contrastive Learning
Haiyang Ying, Yixuan Yin, Jinzhi Zhang +4
Towards holistic understanding of 3D scenes, a general 3D segmentation method is needed that can segment diverse objects without restrictions on object quantity or categories, whil…
Crowd3D: Towards Hundreds of People Reconstruction from a Single Image
Hao Wen, Jing Huang, Huili Cui +4
Image-based multi-person reconstruction in wide-field large scenes is critical for crowd analysis and security alert. However, existing methods cannot deal with large scenes contai…
MILD: Multi-Index hashing for Loop closure Detection
Lei Han, Lu Fang
Loop Closure Detection (LCD) has been proved to be extremely useful in global consistent visual Simultaneously Localization and Mapping (SLAM) and appearance-based robot relocaliza…
DANTE-W: Diffuse Albedo Neural Texturing in the Wild
Guangyu Wang, Tianheng Lu, Ruqi Huang +1
Classical mesh texturing techniques blend captured multi-view images directly, which inevitably suffer from baked-in shading and casted shadows that compromise visual fidelity duri…
OccuSeg: Occupancy-aware 3D Instance Segmentation
Lei Han, Tian Zheng, Lan Xu +1
3D instance segmentation, with a variety of applications in robotics and augmented reality, is in large demands these days. Unlike 2D images that are projective observations of the…
SVQNet: Sparse Voxel-Adjacent Query Network for 4D Spatio-Temporal LiDAR Semantic Segmentation
Xuechao Chen, Shuangjie Xu, Xiaoyi Zou +3
LiDAR-based semantic perception tasks are critical yet challenging for autonomous driving. Due to the motion of objects and static/dynamic occlusion, temporal information plays an…
Geo-NI: Geometry-aware Neural Interpolation for Light Field Rendering
Gaochang Wu, Yuemei Zhou, Yebin Liu +2
In this paper, we present a Geometry-aware Neural Interpolation (Geo-NI) framework for light field rendering. Previous learning-based approaches either rely on the capability of ne…
Request-Only Optimization for Recommendation Systems
Liang Guo, Wei Li, Lucy Liao +25
Deep Learning Recommendation Models (DLRMs) represent one of the largest machine learning applications on the planet. Industry-scale DLRMs are trained with petabytes of recommendat…
EventCap: Monocular 3D Capture of High-Speed Human Motions using an Event Camera
Lan Xu, Weipeng Xu, Vladislav Golyanik +3
The high frame rate is a critical requirement for capturing fast human motions. In this setting, existing markerless image-based methods are constrained by the lighting requirement…
LapEPI-Net: A Laplacian Pyramid EPI structure for Learning-based Dense Light Field Reconstruction
Gaochang Wu, Yebin Liu, Lu Fang +1
For dense sampled light field (LF) reconstruction problem, existing approaches focus on a depth-free framework to achieve non-Lambertian performance. However, they trap in the trad…
Generative AI and Sales Productivity: Field Experiments in Online Retail
Lu Fang, Zhe Yuan, Kaifu Zhang +2
We quantify the short-term impact of Generative Artificial Intelligence (GenAI) on sales performance through a series of large-scale randomized field experiments involving millions…
Data-Uncertainty Guided Multi-Phase Learning for Semi-Supervised Object Detection
Zhenyu Wang, Yali Li, Ye Guo +2
In this paper, we delve into semi-supervised object detection where unlabeled images are leveraged to break through the upper bound of fully-supervised object detection models. Pre…
Halftone Image Watermarking by Content Aware Double-sided Embedding Error Diffusion
Yuanfang Guo, Oscar C. Au, Rui Wang +2
In this paper, we carry out a performance analysis from a probabilistic perspective to introduce the EDHVW methods' expected performances and limitations. Then, we propose a new ge…
PARF: Primitive-Aware Radiance Fusion for Indoor Scene Novel View Synthesis
Haiyang Ying, Baowei Jiang, Jinzhi Zhang +4
This paper proposes a method for fast scene radiance field reconstruction with strong novel view synthesis performance and convenient scene editing functionality. The key idea is t…
Large-scale neuromorphic optoelectronic computing with a reconfigurable diffractive processing unit
Tiankuang Zhou, Xing Lin, Jiamin Wu +7
Application-specific optical processors have been considered disruptive technologies for modern computing that can fundamentally accelerate the development of artificial intelligen…
SurfaceNet: An End-to-end 3D Neural Network for Multiview Stereopsis
Mengqi Ji, Juergen Gall, Haitian Zheng +2
This paper proposes an end-to-end learning framework for multiview stereopsis. We term the network SurfaceNet. It takes a set of images and their corresponding camera parameters as…
A^2-FPN: Attention Aggregation based Feature Pyramid Network for Instance Segmentation
Miao Hu, Yali Li, Lu Fang +1
Learning pyramidal feature representations is crucial for recognizing object instances at different scales. Feature Pyramid Network (FPN) is the classic architecture to build a fea…
LEAD: LiDAR Extender for Autonomous Driving
Jianing Zhang, Wei Li, Honggang Gou +2
3D perception using sensors under vehicle industrial standard is the rigid demand in autonomous driving. MEMS LiDAR emerges with irresistible trend due to its lower cost, more robu…
FlyCap: Markerless Motion Capture Using Multiple Autonomous Flying Cameras
Lan Xu, Lu Fang, Wei Cheng +4
Aiming at automatic, convenient and non-instrusive motion capture, this paper presents a new generation markerless motion capture technique, the FlyCap system, to capture surface m…
Learning High-level Prior with Convolutional Neural Networks for Semantic Segmentation
Haitian Zheng, Yebin Liu, Mengqi Ji +2
This paper proposes a convolutional neural network that can fuse high-level prior for semantic image segmentation. Motivated by humans' vision recognition system, our key design is…
RealLiFe: Real-Time Light Field Reconstruction via Hierarchical Sparse Gradient Descent
Yijie Deng, Lei Han, Tianpeng Lin +3
With the rise of Extended Reality (XR) technology, there is a growing need for real-time light field reconstruction from sparse view inputs. Existing methods can be classified into…
Corner2Net: Detecting Objects as Cascade Corners
Chenglong Liu, Jintao Liu, Haorao Wei +4
The corner-based detection paradigm enjoys the potential to produce high-quality boxes. But the development is constrained by three factors: 1) Hard to match corners. Heuristic cor…
SparseFormer: Detecting Objects in HRW Shots via Sparse Vision Transformer
Wenxi Li, Yuchen Guo, Jilai Zheng +4
Recent years have seen an increase in the use of gigapixel-level image and video capture systems and benchmarks with high-resolution wide (HRW) shots. However, unlike close-up shot…
RegNet: Learning the Optimization of Direct Image-to-Image Pose Registration
Lei Han, Mengqi Ji, Lu Fang +1
Direct image-to-image alignment that relies on the optimization of photometric error metrics suffers from limited convergence range and sensitivity to lighting conditions. Deep lea…
PANDA: A Gigapixel-level Human-centric Video Dataset
Xueyang Wang, Xiya Zhang, Yinheng Zhu +8
We present PANDA, the first gigaPixel-level humAN-centric viDeo dAtaset, for large-scale, long-term, and multi-object visual analysis. The videos in PANDA were captured by a gigapi…
StreamVLA: Breaking the Reason-Act Cycle via Completion-State Gating
Tongqing Chen, Hang Wu, Jiasen Wang +2
Long-horizon robotic manipulation requires bridging the gap between high-level planning (System 2) and low-level control (System 1). Current Vision-Language-Action (VLA) models oft…
CrossNet: An End-to-end Reference-based Super Resolution Network using Cross-scale Warping
Haitian Zheng, Mengqi Ji, Haoqian Wang +2
The Reference-based Super-resolution (RefSR) super-resolves a low-resolution (LR) image given an external high-resolution (HR) reference image, where the reference image and LR ima…
Beyond SIFT using Binary features for Loop Closure Detection
Lei Han, Guyue Zhou, Lan Xu +1
In this paper a binary feature based Loop Closure Detection (LCD) method is proposed, which for the first time achieves higher precision-recall (PR) performance compared with state…
Utilizing High-level Visual Feature for Indoor Shopping Mall Navigation
Ziwei Xu, Haitian Zheng, Minjian Pang +4
Towards robust and convenient indoor shopping mall navigation, we propose a novel learning-based scheme to utilize the high-level visual information from the storefront images capt…
Fast 3D cell tracking with wide-field fluorescence microscopy through deep learning
Kan Liu, Hui Qiao, Jiamin Wu +3
Tracking cells in 3D at high speed continues to attract extensive attention for many biomedical applications, such as monitoring immune cell migration and observing tumor metastasi…
SuperSuit: An Isomorphic Bimodal Interface for Scalable Mobile Manipulation
Tongqing Chen, Hang Wu, Jiasen Wang +3
High-quality, long-horizon demonstrations are essential for embodied AI, yet acquiring such data for tightly coupled wheeled mobile manipulators remains a fundamental bottleneck. U…
Roadmap on Neuromorphic Photonics
Daniel Brunner, Bhavin J. Shastri, Mohammed A. Al Qadasi +147
This roadmap consolidates recent advances while exploring emerging applications, reflecting the remarkable diversity of hardware platforms, neuromorphic concepts, and implementatio…
Light Field Reconstruction Using Convolutional Network on EPI and Extended Applications
Gaochang Wu, Yebin Liu, Lu Fang +2
In this paper, a novel convolutional neural network (CNN)-based framework is developed for light field reconstruction from a sparse set of views. We indicate that the reconstructio…
Deep Learning for Surface Material Classification Using Haptic And Visual Information
Haitian Zheng, Lu Fang, Mengqi Ji +3
When a user scratches a hand-held rigid tool across an object surface, an acceleration signal can be captured, which carries relevant information about the surface. More importantl…
DynamicTrack: Advancing Gigapixel Tracking in Crowded Scenes
Yunqi Zhao, Yuchen Guo, Zheng Cao +3
Tracking in gigapixel scenarios holds numerous potential applications in video surveillance and pedestrian analysis. Existing algorithms attempt to perform tracking in crowded scen…
SurfaceNet+: An End-to-end 3D Neural Network for Very Sparse Multi-view Stereopsis
Mengqi Ji, Jinzhi Zhang, Qionghai Dai +1
Multi-view stereopsis (MVS) tries to recover the 3D model from 2D images. As the observations become sparser, the significant 3D information loss makes the MVS problem more challen…
Zoom in to the details of human-centric videos
Guanghan Li, Yaping Zhao, Mengqi Ji +2
Presenting high-resolution (HR) human appearance is always critical for the human-centric videos. However, current imagery equipment can hardly capture HR details all the time. Exi…