collaborators

10 papers

cs.CV2025

Xiaoice: Training-Free Video Understanding via Self-Supervised Spatio-Temporal Clustering of Semantic Features

Shihao Ji, Zihui Song

The remarkable zero-shot reasoning capabilities of large-scale Visual Language Models (VLMs) on static images have yet to be fully translated to the video domain. Conventional vide…

cs.CL2025

Credal Transformer: A Principled Approach for Quantifying and Mitigating Hallucinations in Large Language Models

Shihao Ji, Zihui Song, Jiajie Huang

Large Language Models (LLMs) hallucinate, generating factually incorrect yet confident assertions. We argue this stems from the Transformer's Softmax function, which creates "Artif…

cs.LG2025

L-MoE: End-to-End Training of a Lightweight Mixture of Low-Rank Adaptation Experts

Shihao Ji, Zihui Song

The Mixture of Experts (MoE) architecture enables the scaling of Large Language Models (LLMs) to trillions of parameters by activating a sparse subset of weights for each input, ma…

cs.CE2025

RAID-0e: A Resilient Striping Array Architecture for Balanced Performance and Availability

Yanzhao Jia, Zhaobo Wu, Zheyi Cao +3

This paper introduces a novel disk array architecture, designated RAID-0e (Resilient Striping Array), designed to superimpose a low-overhead fault tolerance layer upon traditional…

cs.LG2025

MyGO: Memory Yielding Generative Offline-consolidation for Lifelong Learning Systems

Shihao Ji, Zihui Song

Continual or Lifelong Learning aims to develop models capable of acquiring new knowledge from a sequence of tasks without catastrophically forgetting what has been learned before.…

cs.LG2025

OpenGrok: Enhancing SNS Data Processing with Distilled Knowledge and Mask-like Mechanisms

Lumen AI, Zaozhuang No. 28 Middle School, Shihao Ji +6

This report details Lumen Labs' novel approach to processing Social Networking Service (SNS) data. We leverage knowledge distillation, specifically a simple distillation method ins…