5 papers
SPHINX: First Explain, Then Explore
Nguyen Do, Tue M. Cao, Tien Van Do +3
Generating adversarial driving scenarios is critical for evaluating and improving autonomous vehicle decision-making systems in simulation. Recent approaches rely primarily on the…
Semantic Optimal Transport for Sparse Autoencoder Feature Matching and Circuit Compression
Tue M. Cao, Nguyen Do, My T. Thai
Sparse autoencoders (SAEs) have become a central tool for interpreting language models. However, two key SAE analyses that remain difficult to scale are (1) matching semantically s…
ReSS: Learning Reasoning Models for Tabular Data Prediction via Symbolic Scaffold
Chenlang Yi, Gang Li, Zizhan Xiong +4
Tabular data remains prevalent in high-stakes domains such as healthcare and finance, where predictive models are expected to provide both high accuracy and faithful, human-underst…
Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders
Tue M. Cao, Hoang X. Nhat, Raed Alharbi +2
Learning hierarchical features in Sparse Autoencoders (SAEs) is essential for capturing the structured nature of real-world data and mitigating issues like feature absorption or sp…
NeurFlow: Interpreting Neural Networks through Neuron Groups and Functional Interactions
Tue M. Cao, Nhat X. Hoang, Hieu H. Pham +2
Understanding the inner workings of neural networks is essential for enhancing model performance and interpretability. Current research predominantly focuses on examining the conne…