11 papers
When to Retrieve During Reasoning: Adaptive Retrieval for Large Reasoning Models
Dongxin Guo, Jikun Wu, Siu Ming Yiu
Large reasoning models such as DeepSeek-R1 and OpenAI o1 generate extended chains of thought spanning thousands of tokens, yet their integration with retrieval-augmented generation…
FinGround: Detecting and Grounding Financial Hallucinations via Atomic Claim Verification
Dongxin Guo, Jikun Wu, Siu Ming Yiu
Financial AI systems must produce answers grounded in specific regulatory filings, yet current LLMs fabricate metrics, invent citations, and miscalculate derived quantities. These…
ComplianceNLP: Knowledge-Graph-Augmented RAG for Multi-Framework Regulatory Gap Detection
Dongxin Guo, Jikun Wu, Siu Ming Yiu
Financial institutions must track over 60,000 regulatory events annually, overwhelming manual compliance teams; the industry has paid over USD 300 billion in fines and settlements…
AgentEval: DAG-Structured Step-Level Evaluation for Agentic Workflows with Error Propagation Tracking
Dongxin Guo, Jikun Wu, Siu Ming Yiu
Agentic systems that chain reasoning, tool use, and synthesis into multi-step workflows are entering production, yet prevailing evaluation practices like end-to-end outcome checks…
RouteNLP: Closed-Loop LLM Routing with Conformal Cascading and Distillation Co-Optimization
Dongxin Guo, Jikun Wu, Siu Ming Yiu
Serving diverse NLP workloads with large language models is costly: at one enterprise partner, inference costs exceeded $200K/month despite over 70% of queries being routine tasks…
GeoCert: Certified Geometric AI for Reliable Forecasting
Regina Zhang, Zongru Li, Honggang Wen +4
Forecasting systems in science must be accurate, physically consistent, and certifiably reliable. Most existing models address prediction, constraint enforcement, and verification…