8 papers
Knowledge Distillation Must Account for What It Loses
Wenshuo Wang
This position paper argues that knowledge distillation must account for what it loses: student models should be judged not only by retained task scores, but by whether they preserv…
LLM Reasoning Is Latent, Not the Chain of Thought
Wenshuo Wang
This position paper argues that large language model (LLM) reasoning should be studied as latent-state trajectory formation rather than as faithful surface chain-of-thought (CoT).…
iTAG: Inverse Design for Natural Text Generation with Accurate Causal Graph Annotations
Wenshuo Wang, Boyu Cao, Nan Zhuang +1
A fundamental obstacle to causal discovery from text is the lack of causally annotated text data for use as ground truth, due to high annotation costs. This motivates an important…
Breaking Scale Anchoring: Frequency Representation Learning for Accurate High-Resolution Inference from Low-Resolution Training
Wenshuo Wang, Fan Zhang
Zero-Shot Super-Resolution Spatiotemporal Forecasting requires a deep learning model to be trained on low-resolution data and deployed for inference on high-resolution. Existing st…
Experimentation on Endogenous Graphs
Wenshuo Wang, Edvard Bakhitov, Dominic Coey
We study experimentation under endogenous network interference. Interference patterns are mediated by an endogenous graph, where edges can be formed or eliminated as a result of tr…
S^2-KD: Semantic-Spectral Knowledge Distillation Spatiotemporal Forecasting
Wenshuo Wang, Yaomin Shen, Yingjie Tan +1
Spatiotemporal forecasting often relies on computationally intensive models to capture complex dynamics. Knowledge distillation (KD) has emerged as a key technique for creating lig…