5 citations · 11 across the 39 of their papers we have counts for
41 papers
Same Outcome, Different Readout: What Does a Steerable Valence Direction in LLMs Represent?
Weihan Li, Xinlei Chen, Yuhan Song +2
Decodability and successful activation steering do not, by themselves, establish what an internal direction represents. This gap is especially consequential for welfare-relevant in…
AgentIdeaBench: Benchmarking Scientific Ideation in the Agent Era
Yunxiang Mo, Tianshi Zheng, Yisen Gao +7
Scientific ideation is the capacity to formulate novel and testable hypotheses from scientific evidence, and autonomous AI scientists depend on it. Existing evaluations largely ass…
MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning
Kawai Chung, Chunkit Chan, Yauwai Yim +12
Multimodal Large Language Models have sparked significant interest due to their potential for social intelligence; however, their ability to perform sequential motivation reasoning…
Finding the Right Evidence: Factor-Guided Coarse-to-Fine Reasoning for Long Videos
Baixuan Xu, Yinyui Xu, Tianshi Zheng +9
While LVLMs rapidly improve, long-video question answering still remains challenging: relevant evidence is sparse, and question-relevant context often fails to provide cues that di…
Grading the Grader: Lessons from Evaluating an Agentic Data Analysis System
Tian Zheng, Kai-Tai Hsu
Agentic data analysis systems produce rich outputs, including code, numerical results, and verbal diagnostics. This makes them more challenging to evaluate than single-turn LLM res…
Target-Aware Linear Regression Under Distribution Shift
Zhewen Hou, Tian Zheng
Distribution shift between training and deployment is a pervasive challenge for modern AI systems. In many cases, the target marginals of covariates and response are known or speci…