9 papers
Agentic Commerce World: An Auditable and Verifiable Environment for Vibe Commerce
Shicheng Fan, Mingdai Yang, Duohao Wang +9
In vibe coding, people describe software in natural language and delegate implementation to AI agents. By analogy, vibe commerce allows people to express buying or selling goals in…
Paying for Honesty Without Knowing the Truth: Reputation-Penalty Design for LLM Marketplace Agents
Mingdai Yang, Shicheng Fan, Kejing Yu +5
The paper proposes CARP, a reputation‑penalty mechanism that discourages LLM agents from fabricating product attributes without needing ground‑truth verification, and shows that co…
ARC-STAR: Auditable Post-Hoc Correction for PDE Foundation Models
Chengze Li, Lingwei Wei, Li Sun +7
Partial differential equation (PDE) foundation models are pretrained networks that forecast how physical fields like velocity and pressure evolve from a single reusable solver. On…
AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery
Guiyao Tie, Jiawen Shi, Dingjie Song +20
Scientific research is being reshaped by AI systems that move beyond isolated assistance toward longer-horizon workflows spanning literature grounding, hypothesis generation, exper…
A Deployment Audit of Release-Side Risk in Conformal Triage under Prevalence Shift
Chengze Li, Xiao Liu, Hanrong Zhang +7
Conformal triage converts predictive scores into deployment actions that either release a case, flag it for urgent attention, or defer it to human review. Under an observed change…
SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language Models
Yiyang Gu, Junwei Yang, Junyu Luo +15
Large language models (LLMs) are increasingly applied to scientific research, yet existing evaluations often fail to reflect the fine-grained capabilities required in practice. Mos…