Can LLMs Catch a Rigged Backtest? A Clean-Control Calibration Benchmark
2609.28090

Authors

Makar Ulesov,Vladislav Smirnov,Omar Ibrahim,Arsenii Bobovnikov

Abstract

Backtest auditing is a calibration problem: high flaw recall is not useful when the model falsely flags matched clean strategies. We build a 96-item paired benchmark in which every flawed backtest has a clean control that holds strategy, dates, code style, labels, and reporting scaffold fixed while changing one methodology detail.

A deterministic scorer separates flaw recall, clean-control false positives, evidence localization, and fix relevance. Over 1440 cached audits from four text endpoints, the primary DeepSeek auditor reaches 100.0% closed and clean-aware code recall, but open prompts over-flag 93.8% of clean code controls, and clean-aware all-three specificity is 87.5% even where recall saturates.

A clean-aware warning drops DeepSeek code false positives from 20.8% (95% CI 11.7--34.3) to 0.0% (0.0--7.4) at unchanged recall, while the budget anchor still flags 38/48 clean controls under the same prompt. Reporting recall alone would rank three of these four models identically; reporting the clean-control rate separates them by 79 points.

Resources

Ray graphicRay graphicRay graphicRay graphic

Stay in the loop

Every AI paper that matters, free in your inbox daily.

Details

  • takara.ai
  • Custom AI and machine learning from the Frontier Research Team.
  • © 2026 takara.ai Ltd
  • Content is sourced from third-party publications.
Ray graphicRay graphicRay graphic