Publications (25)
TS Cache: A Fast Cache with Timing-speculation Mechanism Under Low Supply Voltages
Shan Shen, Tianxiang Shao, Xiaojing Shang +4
To mitigate the ever-worsening Power Wall problem, more and more applications need to expand their power supply to the wide-voltage range including the near-threshold region. Howev…
Spectral Element Simulation of Liquid Metal Magnetohydrodynamics
Yichen Guo, Paul Fischer, Misun Min
A spectral-element-based formulation of incompressible MHD is presented in the context of the open-source fluid-thermal code, Nek5000/RS. The formulation supports magnetic fields i…
SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering
Kai Tang, Jinhao You, Bohua Zhang +6
Large Vision-Language Models (LVLMs) have achieved remarkable progress in visual understanding tasks such as image captioning and visual question answering. However, they remain su…
FORCE: Efficient VLA Reinforcement Fine-Tuning via Value-Calibrated Warm-up and Self-Distillation
Shuyi Zhang, Yunfan Lou, Hongyang Cheng +8
Vision-Language-Action (VLA) models are often constrained by the imitation ceiling imposed by sub-optimal data. While Reinforcement Learning (RL) fine-tuning can surpass this limit…
STAR: Stage-Wise Attention-Guided Token Reduction for Efficient Large Vision-Language Models Inference
Yichen Guo, Hanze Li, Zonghao Zhang +3
Although large vision-language models (LVLMs) leverage rich visual token representations to achieve strong performance on multimodal tasks, these tokens also introduce significant…
TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference
Tinghao Wang, Yichen Guo, Rui Huang +11
Multimodal large language models (MLLMs) have achieved strong multimodal reasoning capabilities, but their efficiency is limited by the large number of visual tokens, which introdu…
Auditing Agent Harness Safety
Chengzhi Liu, Yichen Guo, Yepeng Liu +8
LLM agents increasingly run inside execution harnesses that dispatch tools, allocate resources, and route messages between specialized components. However, a harness can return a c…
FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models
Yichen Guo, Kai Tang, Fenglai Lin +5
Despite the impressive capabilities of Large Vision-Language Models (LVLMs), they remain susceptible to hallucination, generating content inconsistent with the input image. Recent…
Uniform error bound of an exponential wave integrator for the long-time dynamics of the nonlinear Schrödinger equation with wave operator
Yue Feng, Yichen Guo, Yongjun Yuan
We establish the uniform error bound of an exponential wave integrator Fourier pseudospectral (EWI-FP) method for the long-time dynamics of the nonlinear Schrödinger equation with…
DAQE: Enhancing the Quality of Compressed Images by Exploiting the Inherent Characteristic of Defocus
Qunliang Xing, Mai Xu, Xin Deng +1
Image defocus is inherent in the physics of image formation caused by the optical aberration of lenses, providing plentiful information on image quality. Unfortunately, existing qu…
Mitigating Hallucinations in Large Language Models Via Decoder Layer Skipping
Hanze Li, Jinhao You, Yichen Guo +3
Large Language Models (LLMs) have achieved strong performance across diverse natural language tasks, yet their outputs often suffer from hallucinations -- content that is misaligne…
Mitigating Hallucinations via Inter-Layer Consistency Aggregation in Large Vision-Language Models
Kai Tang, Jinhao You, Yichen Guo +8
Despite the impressive capabilities of Large Vision-Language Models (LVLMs), they remain susceptible to hallucinations, where generated content is inconsistent with the input image…
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
Chengzhi Liu, Yuzhe Yang, Sophia Xiao Pu +14
Multimodal large language models are increasingly deployed as long-horizon agents, where memory must do more than recall: it must track an evolving world, revise what has gone stal…
Stopping Criteria for the Conjugate Gradient Algorithm in High-Order Finite Element Methods
Yichen Guo, Eric de Sturler, Tim Warburton
We consider stopping criteria that balance algebraic and discretization errors for the conjugate gradient algorithm applied to high-order finite element discretizations of Poisson…
OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs
Qianqi Yan, Yichen Guo, Ching-Chen Kuo +4
Modern multimodal large language models (MLLMs) generate fluent responses from interleaved text, image, audio, and video inputs. However, identifying which input sources support ea…
SafeDojo: Safe Reinforcement Learning for VLA via Interactive World Model
Kai Tang, Peidong Jia, Zhong Chu +15
Safe control is a prerequisite for real-world embodied intelligence, for which safe reinforcement learning has emerged as a promising paradigm. However, existing safe reinforcement…
Burst Image Quality Assessment: A New Benchmark and Unified Framework for Multiple Downstream Tasks
Xiaoye Liang, Lai Jiang, Minglang Qiao +6
In recent years, the development of burst imaging technology has improved the capture and processing capabilities of visual data, enabling a wide range of applications. However, th…
RegimeVGGT: Layer-Wise Spatially Preserving Redundancy Removal for Visual Geometry Grounded Transformer
Jinhao You, Shuo Lyu, Zhuohang Lyu +5
Visual Geometry Grounded Transformer (VGGT) recovers dense 3D scene structure from multi-view images in one forward pass, but quadratic cross-frame attention limits its scalability…
A Robust Fiber-based Frequency Synchronization System Immune to Dramatic Temperature Fluctuation
Xi Zhu, Bo Wang, Yichen Guo +5
Fiber-based frequency synchronization system is sensitive to temperature change because of the limited isolation and nonlinear effect of RF components in the system. In order to ma…
Envisage: Towards Expressive Visual Graph Querying
Xiaolin Wen, Qishuang Fu, Shuangyue Han +3
Graph querying is the process of retrieving information from graph data using specialized languages (e.g., Cypher), often requiring programming expertise. Visual Graph Querying (VG…
The Impact of Interference Cognition on the Reliability and Capacity of Industrial Wireless Communications
Yichen Guo, Tao Peng, Yujie Zhao +2
Interference significantly impacts the performance of industrial wireless networks, particularly n severe interference environments with dense networks reusing spectrum resources i…
MAP: Mitigating Hallucinations in Large Vision-Language Models with Map-Level Attention Processing
Chenxi Li, Yichen Guo, Benfang Qian +5
Large Vision-Language Models (LVLMs) have achieved impressive performance in multimodal tasks, but they still suffer from hallucinations, i.e., generating content that is grammatic…
Uncertainty Guided Adaptive Warping for Robust and Efficient Stereo Matching
Junpeng Jing, Jiankun Li, Pengfei Xiong +7
Correlation based stereo matching has achieved outstanding performance, which pursues cost volume between two feature maps. Unfortunately, current methods with a fixed model do not…
An Adaptive Mixed Precision and Dynamically Scaled Preconditioned Conjugate Gradient Algorithm
Yichen Guo, Eric de Sturler, Tim Warburton
We propose an adaptive mixed precision and dynamically scaled preconditioned conjugate gradient algorithm (AMP-PCG). It dynamically adjusts the precision for storing vectors and co…
Blind VQA on 360° Video via Progressively Learning from Pixels, Frames and Video
Li Yang, Mai Xu, Shengxi Li +2
Blind visual quality assessment (BVQA) on 360{\textdegree} video plays a key role in optimizing immersive multimedia systems. When assessing the quality of 360{\textdegree} video,…