Publications (39)
NeRFTAP: Enhancing Transferability of Adversarial Patches on Face Recognition using Neural Radiance Fields
Xiaoliang Liu, Furao Shen, Feng Han +2
Face recognition (FR) technology plays a crucial role in various applications, but its vulnerability to adversarial attacks poses significant security concerns. Existing research p…
UniReason 1.0: A Unified Reasoning Framework for World Knowledge Aligned Image Generation and Editing
Dianyi Wang, Chaofan Ma, Feng Han +8
Unified multimodal models often struggle with complex synthesis tasks that demand deep reasoning, and typically treat text-to-image generation and image editing as isolated capabil…
Machine Learning-based Online Stability Lobe Diagram Estimation and Chatter Suppression Control in Milling Process
Yi Huang, Feng Han, Wenyi Liu +2
Chatter is a self-excited vibration in milling that degrades surface quality and accelerates tool wear. This paper presents an adaptive process controller that suppresses chatter b…
SASICM A Multi-Task Benchmark For Subtext Recognition
Hua Yan, Feng Han, Junyi An +3
Subtext is a kind of deep semantics which can be acquired after one or more rounds of expression transformation. As a popular way of expressing one's intentions, it is well worth s…
Kwai Keye-VL-2.0 Technical Report
Kwai Keye Team, Bin Wen, Changyi Liu +50
We introduce Kwai Keye-VL-2.0-30B-A3B, an open-source Mixture-of-Experts (MoE) multimodal foundation model designed to advance long-video understanding and agentic intelligence. To…
Physics-informed Multiresolution Wavelet Neural Network Method for Solving Partial Differential Equations
Feng Han, Jianguo Wang, Guoliang Peng +1
In this paper, a physics-informed multiresolution wavelet neural network (PIMWNN) method is proposed for solving partial differential equations (PDEs). This method uses the multire…
Cascaded Nonlinear Control Design for Highly Underactuated Balance Robots
Feng Han, Jingang Yi
This paper presents a nonlinear control design for highly underactuated balance robots, which possess more numbers of unactuated degree-of-freedom (DOF) than actuated ones. To addr…
What do LLMs Know about Financial Markets? A Case Study on Reddit Market Sentiment Analysis
Xiang Deng, Vasilisa Bashlovkina, Feng Han +2
Market sentiment analysis on social media content requires knowledge of both financial markets and social media jargon, which makes it a challenging task for human raters. The resu…
Explaining Model Overfitting in CNNs via GMM Clustering
Hui Dou, Xinyu Mu, Mengjun Yi +3
Convolutional Neural Networks (CNNs) have demonstrated remarkable prowess in the field of computer vision. However, their opaque decision-making processes pose significant challeng…
Coordinated Pose Control of Mobile Manipulation with an Unstable Bikebot Platform
Feng Han, Alborz Jelvani, Jingang Yi +1
Bikebot manipulation has advantages of the single-track robot mobility and manipulation dexterity. We present a coordinated pose control of mobile manipulation with the stationary…
Kling-Omni Technical Report
Kling Team, Jialu Chen, Yuanzheng Ci +65
We present Kling-Omni, a generalist generative framework designed to synthesize high-fidelity videos directly from multimodal visual language inputs. Adopting an end-to-end perspec…
Video Summarization: Towards Entity-Aware Captions
Hammad A. Ayyubi, Tianqi Liu, Arsha Nagrani +7
Existing popular video captioning benchmarks and models deal with generic captions devoid of specific person, place or organization named entities. In contrast, news videos present…
Gaussian Process-Based Learning Control of Underactuated Balance Robots with an External and Internal Convertible Modeling Structure
Feng Han, Jingang Yi
External and internal convertible (EIC) form-based motion control is one of the effective designs of simultaneously trajectory tracking and balance for underactuated balance robots…
VCE: Safe Autoregressive Image Generation via Visual Contrast Exploitation
Feng Han, Chao Gong, Zhipeng Wei +2
Recently, autoregressive image generation models have wowed audiences with their remarkable capability in creating surprisingly realistic images. Models such as GPT-4o and LlamaGen…
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
Chenglin Li, Qianglong Chen, Feng Han +6
Long-form video understanding remains a fundamental challenge for current Video Large Language Models. Most existing models rely on static reasoning over uniformly sampled frames,…
Kartta Labs: Collaborative Time Travel
Sasan Tavakkol, Feng Han, Brandon Mayer +4
We introduce the modular and scalable design of Kartta Labs, an open source, open data, and scalable system for virtually reconstructing cities from historical maps and photos. Kar…
Gaussian Process-Enhanced, External and Internal Convertible (EIC) Form-Based Control of Underactuated Balance Robots
Feng Han, Jingang Yi
External and internal convertible (EIC) form-based motion control (i.e., EIC-based control) is one of the effective approaches for underactuated balance robots. By sequentially con…
DuMo: Dual Encoder Modulation Network for Precise Concept Erasure
Feng Han, Kai Chen, Chao Gong +3
The exceptional generative capability of text-to-image models has raised substantial safety concerns regarding the generation of Not-Safe-For-Work (NSFW) content and potential copy…
Hyperspectral Image Cross-Domain Object Detection Method based on Spectral-Spatial Feature Alignment
Hongqi Zhang, He Sun, Hongmin Gao +4
With consecutive bands in a wide range of wavelengths, hyperspectral images (HSI) have provided a unique tool for object detection task. However, existing HSI object detection meth…
Autonomous Bikebot Control for Crossing Obstacles with Assistive Leg Impulsive Actuation
Feng Han, Xinyan Huang, Zenghao Wang +2
As a single-track mobile platform, bikebot (i.e., bicycle-based robot) has attractive navigation capability to pass through narrow, off-road terrain with high-speed and high-energy…
MapComp: A Secure View-based Collaborative Analytics Framework for Join-Group-Aggregation
Xinyu Peng, Feng Han, Li Peng +8
Join-group-aggregation (JGA) queries are fundamental to data analytics, yet executing them collaboratively across different parties poses significant privacy risks. Secure multi-pa…
A technical solution for the rule of law, peace, security, and evolvability of global cyberspace -- solve the three genetic defects of IP network
Hui Li, Kedan Li, Jiaqing Lv +3
Since its inception in the 1960s, the internet has profoundly transformed human life. However, its original design now struggles to meet the evolving demands of modern society. Thr…
Identity-Driven Multimedia Forgery Detection via Reference Assistance
Junhao Xu, Jingjing Chen, Xue Song +3
Recent advancements in "deepfake" techniques have paved the way for generating various media forgeries. In response to the potential hazards of these media forgeries, many research…
EmbeddingGemma: Powerful and Lightweight Text Representations
Henrique Schechter Vera, Sahil Dua, Biao Zhang +86
We introduce EmbeddingGemma, a new lightweight, open text embedding model based on the Gemma 3 language model family. Our innovative training recipe strategically captures knowledg…
Unified Personalized Reward Model for Vision Generation
Yibin Wang, Yuhang Zang, Feng Han +4
Recent advancements in multimodal reward models (RMs) have significantly propelled the development of visual generation. Existing frameworks typically adopt Bradley-Terry-style pre…
ControlThinker: Unveiling Latent Semantics for Controllable Image Generation through Visual Reasoning
Feng Han, Yang Jiao, Shaoxiang Chen +3
The field of controllable image generation has seen significant advancements, with various architectures improving generation layout consistency with control signals. However, cont…
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431
In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…
A Simple and Efficient Pipeline to Build an End-to-End Spatial-Temporal Action Detector
Lin Sui, Chen-Lin Zhang, Lixin Gu +1
Spatial-temporal action detection is a vital part of video understanding. Current spatial-temporal action detection methods mostly use an object detector to obtain person candidate…
DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing
Dianyi Wang, Ruihang Li, Feng Han +17
Current unified multimodal models for image generation and editing typically rely on massive parameter scales (e.g., >10B), entailing prohibitive training costs and deployment foot…
Scaling Multimodal Pre-Training via Cross-Modality Gradient Harmonization
Junru Wu, Yi Liang, Feng Han +3
Self-supervised pre-training recently demonstrates success on large-scale multimodal data, and state-of-the-art contrastive learning methods often enforce the feature consistency f…
Dual-LoRA and Quality-Enhanced Pseudo Replay for Multimodal Continual Food Learning
Xinlan Wu, Bin Zhu, Feng Han +2
Food analysis has become increasingly critical for health-related tasks such as personalized nutrition and chronic disease prevention. However, existing large multimodal models (LM…
Learning-Based Safe Motion Control of Vehicle Ski-Stunt Maneuvers
Feng Han, Jingang Yi
This paper presents a safety guaranteed control method for an autonomous vehicle ski-stunt maneuver, that is, a vehicle moving with two one-side wheels. To capture the vehicle dyna…
PLaD: Preference-based Large Language Model Distillation with Pseudo-Preference Pairs
Rongzhi Zhang, Jiaming Shen, Tianqi Liu +7
Large Language Models (LLMs) have exhibited impressive capabilities in various tasks, yet their vast parameter sizes restrict their applicability in resource-constrained settings.…
LoMo: Local Modality Substitution for Deeper Vision-Language Fusion
Feng Han, Zhixiong Zhang, Zheming Liang +2
Vision-Language Models (VLMs) have achieved substantial progress across a wide range of understanding and reasoning tasks, driven by large-scale image-text training aimed at multim…
UniREditBench: A Unified Reasoning-based Image Editing Benchmark
Feng Han, Yibin Wang, Chenglin Li +8
Recent advances in multi-modal generative models have driven substantial improvements in image editing. However, current generative models still struggle with handling diverse and…
VideoPro: Adaptive Program Reasoning for Long Video Understanding
Chenglin Li, Feng Han, Yikun Wang +9
Large language models (LLMs) have shown promise in generating program workflows for visual tasks. However, previous approaches often rely on closed-source models, lack systematic r…
Multi-Scale Dilated Convolution Network for Long-Term Time Series Forecasting
Feifei Li, Suhan Guo, Feng Han +2
Accurate forecasting of long-term time series has important applications for decision making and planning. However, it remains challenging to capture the long-term dependencies in…
Gemini Embedding: Generalizable Embeddings from Gemini
Jinhyuk Lee, Feiyang Chen, Sahil Dua +44
In this report, we introduce Gemini Embedding, a state-of-the-art embedding model leveraging the power of Gemini, Google's most capable large language model. Capitalizing on Gemini…
GrOCE:Graph-Guided Online Concept Erasure for Text-to-Image Diffusion Models
Ning Han, Zhenyu Ge, Feng Han +3
Concept erasure aims to remove harmful, inappropriate, or copyrighted content from text-to-image diffusion models while preserving non-target semantics. However, existing methods e…