7 papers
The Greatness of Science Cannot Be Planned: Agentic Auto-Research is Fuzz Testing
Yifeng He, Jicheng Wang, Yinzhe Zhao +3
Agentic auto-research is emerging, but most systems treat scientific discovery as goal-oriented optimization against a final benchmark. This paradigm rewards a sparse final verdict…
LLM-Based Invariant Testing for Software Functional Bugs
Ruogu Yang, Yifeng He, Yundi Xu +2
Manually writing unit tests to uncover functional bugs in software libraries is not only time-consuming but also requires a deep understanding of the intended semantics of the APIs…
Is Progressive Disclosure All You Need for Long-Context Agents?
Yifeng He, Yinzhe Zhao, Jicheng Wang +1
Long-document question answering usually forces a choice between loading the whole document into the context window and bolting on a separate retriever. Agentic AI suggests a broad…
Content Fuzzing for Escaping Information Cocoons on Digital Social Media
Yifeng He, Ziye Tang, Hao Chen
Information cocoons on social media limit users' exposure to posts with diverse viewpoints. Modern platforms use stance detection as an important signal in recommendation and ranki…
Data-Driven Function Calling Improvements in Large Language Model for Online Financial QA
Xing Tang, Hao Chen, Shiwei Li +7
Large language models (LLMs) have been incorporated into numerous industrial applications. Meanwhile, a vast array of API assets is scattered across various functions in the financ…
Evaluating Program Semantics Reasoning with Type Inference in System F
Yifeng He, Luning Yang, Christopher Castro Gaw Gonzalo +1
Large Language Models (LLMs) are increasingly integrated into the software engineering ecosystem. Their test-time compute (TTC) reasoning capabilities show significant potential fo…