Publications (48)
Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Torresani +98
We present Ego-Exo4D, a diverse, large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric…
Bose-Einstein condensation on an atom chip
Bo Yan, Feng Cheng, Min Ke +3
We report an experiment of creating Bose-Einstein condensate (BEC) on an atom chip. The chip based Z-wire current and a homogeneous bias magnetic field create a tight magnetic trap…
Right-sizing Recommendations (RSR): Cloud Workload Conformal Prediction for Virtual Machines in Data Center Operations
Mehryar Majd, Feng Cheng, Ali Pahlevan
Managing cloud infrastructure efficiently, especially in environments of large cloud providers or hyperscalers, requires optimizing the use of physical resources to minimize costs…
When Prompts Become Payloads: A Framework for Mitigating SQL Injection Attacks in Large Language Model-Driven Applications
Farzad Nourmohammadzadeh Motlagh, Mehrdad Hajizadeh, Mehryar Majd +3
Natural language interfaces to structured databases are becoming increasingly common, largely due to advances in large language models (LLMs) that enable users to query data using…
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types
Jiankang Chen, Tianke Zhang, Changyi Liu +8
Multimodal visual language models are gaining prominence in open-world applications, driven by advancements in model architectures, training techniques, and high-quality data. Howe…
DAM: Dynamic Adapter Merging for Continual Video QA Learning
Feng Cheng, Ziyang Wang, Yi-Lin Sung +3
We present a parameter-efficient method for continual video question-answering (VidQA) learning. Our method, named DAM, uses the proposed Dynamic Adapter Merging to (i) mitigate ca…
VindLU: A Recipe for Effective Video-and-Language Pretraining
Feng Cheng, Xizi Wang, Jie Lei +3
The last several years have witnessed remarkable progress in video-and-language (VidL) understanding. However, most modern VidL approaches use complex and specialized model archite…
Unified Coarse-to-Fine Alignment for Video-Text Retrieval
Ziyang Wang, Yi-Lin Sung, Feng Cheng +2
The canonical approach to video-text retrieval leverages a coarse-grained or fine-grained alignment between visual and textual information. However, retrieving the correct video ac…
Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model
Team Seedance, Heyi Chen, Siyan Chen +194
Recent strides in video generation have paved the way for unified audio-visual generation. In this work, we present Seedance 1.5 pro, a foundational model engineered specifically f…
Efficient Deep Learning Pipelines for Accurate Cost Estimations Over Large Scale Query Workload
Johan Kok Zhi Kang, Gaurav, Sien Yi Tan +3
The use of deep learning models for forecasting the resource consumption patterns of SQL queries have recently been a popular area of study. With many companies using cloud platfor…
Large Language Models in Cybersecurity: State-of-the-Art
Farzad Nourmohammadzadeh Motlagh, Mehrdad Hajizadeh, Mehryar Majd +3
The rise of Large Language Models (LLMs) has revolutionized our comprehension of intelligence bringing us closer to Artificial Intelligence. Since their introduction, researchers h…
Stochastic Backpropagation: A Memory Efficient Strategy for Training Video Models
Feng Cheng, Mingze Xu, Yuanjun Xiong +4
We propose a memory efficient method, named Stochastic Backpropagation (SBP), for training deep neural networks on videos. It is based on the finding that gradients from incomplete…
ReGraph: Scaling Graph Processing on HBM-enabled FPGAs with Heterogeneous Pipelines
Xinyu Chen, Yao Chen, Feng Cheng +3
The use of FPGAs for efficient graph processing has attracted significant interest. Recent memory subsystem upgrades including the introduction of HBM in FPGAs promise to further a…
qGDP: Quantum Legalization and Detailed Placement for Superconducting Quantum Computers
Junyao Zhang, Guanglei Zhou, Feng Cheng +6
Noisy Intermediate-Scale Quantum (NISQ) computers are currently limited by their qubit numbers, which hampers progress towards fault-tolerant quantum computing. A major challenge i…
Vanishing viscosity limit of navier-stokes equations in gevrey class
Feng Cheng, Wei-Xi Li, Chao-Jiang Xu
In this paper we consider the inviscid limit for the periodic solutions to Navier-Stokes equation in the framework of Gevrey class. It is shown that the lifespan for the solutions…
Gevrey regularity with weight for incompressible Euler equation in the half plane
Feng Cheng, Wei-Xi Li, Chao-Jiang Xu
In this work we prove the weighted Gevrey regularity of solutions to the incompressible Euler equation with initial data decaying polynomially at infinity. This is motivated by the…
Synthetic Video Enhances Physical Fidelity in Video Synthesis
Qi Zhao, Xingyu Ni, Ziyu Wang +4
We investigate how to enhance the physical fidelity of video generation models by leveraging synthetic videos derived from computer graphics pipelines. These rendered videos respec…
VINCIE: Unlocking In-context Image Editing from Video
Leigang Qu, Feng Cheng, Ziyan Yang +7
In-context image editing aims to modify images based on a contextual sequence comprising text and previously generated images. Existing methods typically depend on task-specific pi…
Learning Directional Feature Maps for Cardiac MRI Segmentation
Feng Cheng, Cheng Chen, Yukang Wang +5
Cardiac MRI segmentation plays a crucial role in clinical diagnosis for evaluating personalized cardiac performance parameters. Due to the indistinct boundaries and heterogeneous i…
Marchenko-Pastur law for tensor powers of exchangeable unconditional vectors
Feng Cheng, Dan Mikulincer
Given an isotropic, exchangeable, and unconditional random vector , we consider the sample covariance matrix constructed from i.i.d. copies of several tensor models of $…
A Survey: Collaborative Hardware and Software Design in the Era of Large Language Models
Cong Guo, Feng Cheng, Zhixu Du +21
The rapid development of large language models (LLMs) has significantly transformed the field of artificial intelligence, demonstrating remarkable capabilities in natural language…
Faster and Non-ergodic O(1/K) Stochastic Alternating Direction Method of Multipliers
Cong Fang, Feng Cheng, Zhouchen Lin
We study stochastic convex optimization subjected to linear equality constraints. Traditional Stochastic Alternating Direction Method of Multipliers and its Nesterov's acceleration…
Weighted gevrey class regularity of euler equation in the whole space
Feng Cheng, Wei-Xi Li, Chao-Jiang Xu
In this paper we study the weighted Gevrey class regularity of Euler equation in the whole space R 3. We first establish the local existence of Euler equation in weighted Sobolev s…
Boosting the Capability of Intelligent Vulnerability Detection by Training in a Human-Learning Manner
Shihan Dou, Yueming Wu, Wenxuan Li +3
Due to its powerful automatic feature extraction, deep learning (DL) has been widely used in source code vulnerability detection. However, although it performs well on artificial d…
LoCoNet: Long-Short Context Network for Active Speaker Detection
Xizi Wang, Feng Cheng, Gedas Bertasius +1
Active Speaker Detection (ASD) aims to identify who is speaking in each frame of a video. ASD reasons from audio and visual information from two contexts: long-term intra-speaker c…
TALLFormer: Temporal Action Localization with a Long-memory Transformer
Feng Cheng, Gedas Bertasius
Most modern approaches in temporal action localization divide this problem into two parts: (i) short-term feature extraction and (ii) long-range temporal boundary localization. Due…
iMOVE: Instance-Motion-Aware Video Understanding
Jiaze Li, Yaya Shi, Zongyang Ma +7
Enhancing the fine-grained instance spatiotemporal motion perception capabilities of Video Large Language Models is crucial for improving their temporal and general video understan…
SkipSR: Faster Super Resolution with Token Skipping
Rohan Choudhury, Shanchuan Lin, Jianyi Wang +6
Diffusion-based super-resolution (SR) is a key component in video generation and video restoration, but is slow and expensive, limiting scalability to higher resolutions and longer…
Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression
Feng Cheng, Cong Guo, Chiyue Wei +7
Large language models (LLMs) have demonstrated transformative capabilities across diverse artificial intelligence applications, yet their deployment is hindered by substantial memo…
TimeRefine: Temporal Grounding with Time Refining Video LLM
Xizi Wang, Feng Cheng, Ziyang Wang +6
Video temporal grounding aims to localize relevant temporal boundaries in a video given a textual prompt. Recent work has focused on enabling Video LLMs to perform video temporal g…
Prosperity: Accelerating Spiking Neural Networks via Product Sparsity
Chiyue Wei, Cong Guo, Feng Cheng +4
Spiking Neural Networks (SNNs) are highly efficient due to their spike-based activation, which inherently produces bit-sparse computation patterns. Existing hardware implementation…
Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix Multiplication
Haoxuan Shan, Cong Guo, Chiyue Wei +4
The rapid scaling of large language models demands more efficient hardware. Quantization offers a promising trade-off between efficiency and performance. With ultra-low-bit quantiz…
AutoRAC: Automated Processing-in-Memory Accelerator Design for Recommender Systems
Feng Cheng, Tunhou Zhang, Junyao Zhang +6
The performance bottleneck of deep-learning-based recommender systems resides in their backbone Deep Neural Networks. By integrating Processing-In-Memory~(PIM) architectures, resea…
Towards Automated Model Design on Recommender Systems
Tunhou Zhang, Dehua Cheng, Yuchen He +10
The increasing popularity of deep learning models has created new opportunities for developing AI-based recommender systems. Designing recommender systems using deep neural network…
Digital Logarithmic Airborne Gamma Ray Spectrometer
GuoQiang Zeng, QingXian Zhang, Chen Li +4
A new digital logarithmic airborne gamma ray spectrometer is designed in this study. The spectrometer adopts a high-speed and high-accuracy logarithmic amplifier (LOG114) to amplif…
An LLM-based Two-Stage Transformer Framework for Cross-Domain Bearing Fault Diagnosis with Limited Data
Jinghan Wang, Feng Cheng, Wentao Wu +3
Bearing fault diagnosis faces critical challenges when dataset heterogeneity, operating condition variations, and limited labeled data occur simultaneously in industrial environmen…
VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
Ziyang Wang, Shoubin Yu, Elias Stengel-Eskin +4
Long-form video understanding is complicated by the high redundancy of video data and the abundance of query-irrelevant information. To tackle these challenges, we propose VideoTre…
Using Single-Trial Representational Similarity Analysis with EEG to track semantic similarity in emotional word processing
Feng Cheng
Electroencephalography (EEG) is a powerful non-invasive brain imaging technique with a high temporal resolution that has seen extensive use across multiple areas of cognitive scien…
VideoAuteur: Towards Long Narrative Video Generation
Junfei Xiao, Feng Cheng, Lu Qi +5
Recent video generation models have shown promising results in producing high-quality video clips lasting several seconds. However, these models face challenges in generating long…
On the gevrey regularity of solutions to the 3d ideal mhd equations
Feng Cheng, Chao-Jiang Xu
In this paper, similar to the incompressible Euler equation, we prove the propagation of the Gevrey regularity of solutions to the three-dimensional incompressible ideal magnetohyd…
Seedance 2.0: Advancing Video Generation for World Complexity
Team Seedance, De Chen, Liyang Chen +168
Seedance 2.0 is a new native multi-modal audio-video generation model, officially released in China in early February 2026. Compared with its predecessors, Seedance 1.0 and 1.5 Pro…
BOASF: A Unified Framework for Speeding up Automatic Machine Learning via Adaptive Successive Filtering
Guanghui Zhu, Xin Fang, Feng Cheng +4
Machine learning has been making great success in many application areas. However, for the non-expert practitioners, it is always very challenging to address a machine learning tas…
Sequence-Based Identification of First-Person Camera Wearers in Third-Person Views
Ziwei Zhao, Xizi Wang, Yuchen Wang +2
The increasing popularity of egocentric cameras has generated growing interest in studying multi-camera interactions in shared environments. Although large-scale datasets such as E…
Block-Wise Mixed-Precision Quantization: Enabling High Efficiency for Practical ReRAM-based DNN Accelerators
Xueying Wu, Edward Hanson, Nansu Wang +9
Resistive random access memory (ReRAM)-based processing-in-memory (PIM) architectures have demonstrated great potential to accelerate Deep Neural Network (DNN) training/inference.…
Probabilistic representation and inverse design of metamaterials based on a deep generative model with semi-supervised learning strategy
Wei Ma, Feng Cheng, Yihao Xu +2
The research of metamaterials has achieved enormous success in the manipulation of light in an artificially prescribed manner using delicately designed sub-wavelength structures, s…
Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model
Team Seawead, Ceyuan Yang, Zhijie Lin +52
This technical report presents a cost-efficient strategy for training a video generation foundation model. We present a mid-sized research model with approximately 7 billion parame…
New vision of convection induced freckle formation theory in Nickel-based superalloys by electron microscopy
Shuai Wang, Yuliang Jia, Yongzhe Wang +7
Freckles, one of the common defects in blades used in heavy duty gas turbines, hugely deteriorates blades mechanical properties and liability under service conditions. Thermal-solu…
On the radius of spatial analyticity for the inviscid Boussinesq equations
Feng Cheng, Chao-Jiang Xu
In this paper, we study the problem of analyticity of smooth solutions of the inviscid Boussinesq equations. If the initial datum is real-analytic, the solution remains real-analyt…