papers

Publications (43)

cs.SD2024

FGCL: Fine-grained Contrastive Learning For Mandarin Stuttering Event Detection

Han Jiang, Wenyu Wang, Yiquan Zhou +3

This paper presents the T031 team's approach to the StutteringSpeech Challenge in SLT2024. Mandarin Stuttering Event Detection (MSED) aims to detect instances of stuttering events…

cs.CL2025

DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification

Yu Li, Han Jiang, Zhihua Wei

With the widespread adoption of Large Language Models (LLMs), jailbreak attacks have become an increasingly pressing safety concern. While safety-aligned LLMs can effectively defen…

cs.CV2026

Leveraging Latent Visual Reasoning in Silence

Dongyao Zhu, Zhen Wang, Xi Xiao +7

Latent visual reasoning involves visual evidence more directly in multimodal reasoning by inserting continuous latent tokens before textual generation. However, the necessity of th…

cs.LG2025

Trusted Multi-view Learning under Noisy Supervision

Yilin Zhang, Cai Xu, Han Jiang +4

Multi-view learning methods often focus on improving decision accuracy while neglecting the decision uncertainty, which significantly restricts their applications in safety-critica…

cs.CL2026

PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization

Han Jiang, Dongyao Zhu, Xiaoyuan Yi +3

In-Context Learning has shown great potential for aligning Large Language Models (LLMs) with human values, helping reduce harmful outputs and accommodate diverse preferences withou…

cs.CL2025

Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testing

Han Jiang, Xiaoyuan Yi, Zhihua Wei +3

Warning: Contains harmful model outputs. Despite significant advancements, the propensity of Large Language Models (LLMs) to generate harmful and unethical content poses critical c…

cs.AI2026

Can We Trust Item Response Theory for AI Evaluation?

Han Jiang, Sunbeom Kwon, Jinwen Luo +2

The paper investigates how well item response theory (IRT) works for evaluating large language model benchmarks, highlighting challenges when benchmark data differ from traditional…

#item response theory#benchmark evaluation#large language models#simulation study
cs.CV2024

OccSora: 4D Occupancy Generation Models as World Simulators for Autonomous Driving

Lening Wang, Wenzhao Zheng, Yilong Ren +4

Understanding the evolution of 3D scenes is important for effective autonomous driving. While conventional methods mode scene development with the motion of individual instances, w…

cs.CV2025

Visual Self-paced Iterative Learning for Unsupervised Temporal Action Localization

Yupeng Hu, Han Jiang, Hao Liu +3

Recently, temporal action localization (TAL) has garnered significant interest in information retrieval community. However, existing supervised/weakly supervised methods are heavil…

cs.CL2023

You Only Forward Once: Prediction and Rationalization in A Single Forward Pass

Han Jiang, Junwen Duan, Zhe Qu +1

Unsupervised rationale extraction aims to extract concise and contiguous text snippets to support model predictions without any annotated rationale. Previous studies have used a tw…

cs.CL2026

Generative Personality Simulation via Theory-Informed Structured Interview

Pengda Wang, Huiqi Zou, Han Jiang +5

Despite their potential as human proxies, LLMs often fail to generate heterogeneous data with human-like diversity, thereby diminishing their value in advancing social science rese…

physics.flu-dyn2026

Continuum modeling of fluidic and elastic flow during growth-driven wound closure in partial-EMT cell monolayers

Chaozhen Wei, Han Jiang, Yifan Gu +5

Large-scale circular gap closure occurs over a time scale on which cell growth and proliferation become important. Growth is the main driver of the closing process, while cell dyna…

cs.AI2026

AI Evaluation Should Require Standardized Item-Level Data Releases

Han Jiang, Susu Zhang, Dongyao Zhu +6

This position paper argues that standardized item-level benchmark data should become the default infrastructure for AI evaluation. Current evaluations suffer from underspecified it…

cs.CV2023

Registering Neural Radiance Fields as 3D Density Images

Han Jiang, Ruoxuan Li, Haosen Sun +2

No significant work has been done to directly merge two partially overlapping scenes using NeRF representations. Given pre-trained NeRF models of a 3D scene with partial overlappin…

cs.CL2023

SPSQL: Step-by-step Parsing Based Framework for Text-to-SQL Generation

Ran Shen, Gang Sun, Hao Shen +3

Converting text into the structured query language (Text2SQL) is a research hotspot in the field of natural language processing (NLP), which has broad application prospects. In the…

cs.AI2025

The Incomplete Bridge: How AI Research (Mis)Engages with Psychology

Han Jiang, Pengda Wang, Xiaoyuan Yi +2

Social sciences have accumulated a rich body of theories and methodologies for investigating the human mind and behaviors, while offering valuable insights into the design and unde…

cs.RO2025

AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation

Sixiang Chen, Jiaming Liu, Siyuan Qian +9

Recently, mobile manipulation has attracted increasing attention for enabling language-conditioned robotic control in household tasks. However, existing methods still face challeng…

cs.CV2025

PanoWan: Lifting Diffusion Video Generation Models to 360° with Latitude/Longitude-aware Mechanisms

Yifei Xia, Shuchen Weng, Siqi Yang +6

Panoramic video generation enables immersive 360° content creation, valuable in applications that demand scene-consistent world exploration. However, existing panoramic video gene…

cs.CV2026

Lighting-grounded Video Generation with Renderer-based Agent Reasoning

Ziqi Cai, Taoyu Yang, Zheng Chang +4

Diffusion models have achieved remarkable progress in video generation, but their controllability remains a major limitation. Key scene factors such as layout, lighting, and camera…

cs.CV2024

Stag-1: Towards Realistic 4D Driving Simulation with Video Generation Model

Lening Wang, Wenzhao Zheng, Dalong Du +8

4D driving simulation is essential for developing realistic autonomous driving simulators. Despite advancements in existing methods for generating driving scenes, significant chall…

cs.CL2024

DESTEIN: Navigating Detoxification of Language Models via Universal Steering Pairs and Head-wise Activation Fusion

Yu Li, Han Jiang, Chuanyang Gong +1

Despite the remarkable achievements of language models (LMs) across a broad spectrum of tasks, their propensity for generating toxic outputs remains a prevalent concern. Current so…

cs.CE2023

AccidentGPT: Accident Analysis and Prevention from V2X Environmental Perception with Multi-modal Large Model

Lening Wang, Yilong Ren, Han Jiang +9

Traffic accidents, being a significant contributor to both human casualties and property damage, have long been a focal point of research for many scholars in the field of traffic…

cs.AI2026

The Impact of Generative AI on Architectural Conceptual Design: Performance, Creative Self-Efficacy and Cognitive Load

Han Jiang, Yao Xiao, Rachel Hurley +1

Our study examines how generative AI (GenAI) influences performance, creative self-efficacy, and cognitive load in architectural conceptual design tasks. Thirty-six student partici…

cs.CV2023

Inpaint4DNeRF: Promptable Spatio-Temporal NeRF Inpainting with Generative Diffusion Models

Han Jiang, Haosen Sun, Ruoxuan Li +2

Current Neural Radiance Fields (NeRF) can generate photorealistic novel views. For editing 3D scenes represented by NeRF, with the advent of generative models, this paper proposes…

cs.CV2025

MapKD: Unlocking Prior Knowledge with Cross-Modal Distillation for Efficient Online HD Map Construction

Ziyang Yan, Ruikai Li, Zhiyong Cui +7

Online HD map construction is a fundamental task in autonomous driving systems, aiming to acquire semantic information of map elements around the ego vehicle based on real-time sen…

cs.CV2021

GPU-accelerated Faster Mean Shift with euclidean distance metrics

Le You, Han Jiang, Jinyong Hu +4

Handling clustering problems are important in data statistics, pattern recognition and image processing. The mean-shift algorithm, a common unsupervised algorithms, is widely used…

cs.CV2026

Video Generation Models Are Inherent Lighting Estimators

Ziqi Cai, Shuchen Weng, Kaiqi Liu +5

Recovering dynamic environment maps from a single in-the-wild video is crucial for photorealistic rendering, yet remains a challenge. Recent video generation models can produce pho…

cs.CV2025

VIRES: Video Instance Repainting via Sketch and Text Guided Generation

Shuchen Weng, Haojie Zheng, Peixuan Zhang +4

We introduce VIRES, a video instance repainting method with sketch and text guidance, enabling video instance repainting, replacement, generation, and removal. Existing approaches…

cs.CL2023

ToViLaG: Your Visual-Language Generative Model is Also An Evildoer

Xinpeng Wang, Xiaoyuan Yi, Han Jiang +3

Warning: this paper includes model outputs showing offensive content. Recent large-scale Visual-Language Generative Models (VLGMs) have achieved unprecedented improvement in multim…

cs.CV2025

Stability Under Scrutiny: Benchmarking Representation Paradigms for Online HD Mapping

Hao Shan, Ruikai Li, Han Jiang +8

As one of the fundamental modules in autonomous driving, online high-definition (HD) maps have attracted significant attention due to their cost-effectiveness and real-time capabil…

cs.RO2026

PointAction: 3D Points as Universal Action Representations for Robot Control

Mutian Tong, Han Jiang, Qiao Feng +2

Video-Action Models (VAMs) leverage the broad visual dynamics captured by pre-trained video diffusion models, offering a promising path toward generalizable robot manipulation. How…

cs.DS2022

Vertex Sparsifiers for Hyperedge Connectivity

Han Jiang, Shang-En Huang, Thatchaphol Saranurak +1

Recently, Chalermsook et al. [SODA'21(arXiv:2007.07862)] introduces a notion of vertex sparsifiers for -edge connectivity, which has found applications in parameterized algorith…

cs.SD2025

Speech Audio Generation from dynamic MRI via a Knowledge Enhanced Conditional Variational Autoencoder

Yaxuan Li, Han Jiang, Yifei Ma +3

Dynamic Magnetic Resonance Imaging (MRI) of the vocal tract has become an increasingly adopted imaging modality for speech motor studies. Beyond image signals, systematic data loss…

cs.RO2026

Embedding Autonomous Agents in Resource-Constrained Robotic Platforms

Negar Halakou, Juan F. Gutierrez, Ye Sun +4

Many embedded devices operate under resource constraints and in dynamic environments, requiring local decision-making capabilities. Enabling devices to make independent decisions i…

cs.CV2025

AMap: Distilling Future Priors for Ahead-Aware Online HD Map Construction

Ruikai Li, Xinrun Li, Mengwei Xie +12

Online High-Definition (HD) map construction is pivotal for autonomous driving. While recent approaches leverage historical temporal fusion to improve performance, we identify a cr…

cs.CL2024

Dialectical Alignment: Resolving the Tension of 3H and Security Threats of LLMs

Shu Yang, Jiayuan Su, Han Jiang +5

With the rise of large language models (LLMs), ensuring they embody the principles of being helpful, honest, and harmless (3H), known as Human Alignment, becomes crucial. While exi…

cs.CR2026

Intelligent Disruption: Undetectable Attacks on Wireless Autoencoders

Han Jiang, Jifa Zhang, Hu Jin +3

Adversarial attacks can degrade the legitimate decision performance in wireless autoencoder communications. However, in complex scenarios with multiple adversaries, the cumulative…

math.AP2023

Rayleigh-Taylor Instability in Stratified Compressible Fluids with/without the Interfacial Surface Tension

Fei Jiang, Han Jiang, Song Jiang

Guo--Tice formally established in 2011 that the Rayleigh--Taylor instability inevitably occurs within stratified compressible viscous fluids in a slab domain $\mathbb{R}^2\times (h…

cs.CL2023

MHLAT: Multi-hop Label-wise Attention Model for Automatic ICD Coding

Junwen Duan, Han Jiang, Ying Yu

International Classification of Diseases (ICD) coding is the task of assigning ICD diagnosis codes to clinical notes. This can be challenging given the large quantity of labels (ne…

cs.CL2022

CHAE: Fine-Grained Controllable Story Generation with Characters, Actions and Emotions

Xinpeng Wang, Han Jiang, Zhihua Wei +1

Story generation has emerged as an interesting yet challenging NLP task in recent years. Some existing studies aim at generating fluent and coherent stories from keywords and outli…

cs.CL2024

MARE: Multi-Aspect Rationale Extractor on Unsupervised Rationale Extraction

Han Jiang, Junwen Duan, Zhe Qu +1

Unsupervised rationale extraction aims to extract text snippets to support model predictions without explicit rationale annotation. Researchers have made many efforts to solve this…

cs.CL2023

Large-Scale and Multi-Perspective Opinion Summarization with Diverse Review Subsets

Han Jiang, Rui Wang, Zhihua Wei +2

Opinion summarization is expected to digest larger review sets and provide summaries from different perspectives. However, most existing solutions are deficient in epitomizing exte…

cs.AI2026

Enhancing Fitness Intelligence through Domain-Specific LLM Post-Training

Xingtao Zhao, Tian Yang, Han Jiang

Scientific Fitness Coaching (SFC) is typically delivered by human professionals, making it costly and inaccessible to many. While recent advances in Large Language Models (LLMs) sh…