Can You Trust a Backtest? What AI Strategy Tools Won't Tell You in 2026
The most dangerous chart in retail finance is not a meme stock's price history. It is the equity curve of a backtest — that smooth, up-and-to-the-right line that says a strategy would have tripled your money if only you had known. Every algorithmic trading platform can draw one. Almost none of them will tell you how easy that line is to manufacture, and the arrival of AI strategy generators has made manufacturing it roughly a thousand times cheaper.
The overfitting machine was already running
Overfitting is not a new sin. Test enough parameter combinations on the same historical window and something will "work" by pure chance — the multiple-testing problem quants have written about for decades. A moving-average crossover with the periods tuned just so, an RSI threshold nudged until the bad trades disappear: the strategy hasn't learned the market, it has memorized one specific past.
The strategy marketplaces that grew up around retail algo trading made this worse in a predictable way. The strategies you see are the survivors. Nobody publishes the four hundred variants that blew up, so the storefront is a gallery of lucky curves — survivorship bias sold as a product category.
Then AI made iteration free
What changed in the last two years is the cost of a variant. A language model can now write a plausible trading strategy in seconds, and the no-code backtesting tools will happily score it against history just as fast. Iterate that loop a few hundred times — which is exactly what "AI strategy generator" features are built to do — and you have industrialized the thousand-monkeys problem. The monkey that typed Hamlet gets a landing page.
None of this means the underlying platforms are dishonest. The serious quant workbenches ship the right machinery: out-of-sample splits, walk-forward testing, transaction-cost models, realistic fills. But machinery is not discipline. The tools can't force a user — or an AI — to stop iterating when the curve finally looks good, and the incentives of everyone involved point the other way.
Five questions that deflate a fake curve
You don't need a quant degree to audit a backtest. You need five questions.
What exact window was tested, and why that window? Were transaction costs and slippage included, and at what assumptions? How many variants were tried before this one — one, ten, a thousand? Was any of the data held out, untouched, until the strategy was final? And did anyone other than the person selling it check the methodology?
A backtest that survives all five is rare. A backtest that even answers all five is rarer — and the refusal to answer is itself the answer.
The quieter alternative: fewer knobs, more disclosure
Here is where we admit this is the trades.run blog, and that we have a horse in the race — so weigh what follows accordingly.
Our view is that most people reaching for a backtester actually want something a backtest can't give them: a trustworthy answer to a question about how markets behave. "Does buying the dip work on QQQ?" is not a parameter grid — it's a research question, and it deserves a computed, disclosed, graded answer rather than an optimized curve.
That's the shape of what trades.run does. Ask the question in plain English; a frontier AI model writes and runs real analysis code against historical market data — prices, news sentiment, earnings, insider activity — and the report that comes back states its sample window, its statistics, its caveats, and its sample sizes next to its conclusion. An automated methodology reviewer grades every report 1–10 before publication, and the ones that fail are discarded, not dressed up. When you do run a strategy backtest on the platform, the report shows the unglamorous numbers a marketing page would crop out: maximum drawdown, Sharpe and Sortino, profit factor, and performance against simply holding SPY over the same window — because a strategy that loses to the index is a finding, not a failure to hide.
There are fewer knobs to turn on purpose. You cannot iterate a trades.run research report until it flatters you, and we consider that a feature. The overfitting machine runs on frictionless retries; the honest alternative is to make each answer accountable instead of adjustable.
The bottom line
AI has not broken backtesting — it has revealed what was already broken about how backtests get used. The equity curve was always the easiest artifact in finance to fake, and now it can be faked at scale, politely, by a model that never gets tired. So hold every curve to the five questions, ours included. The tools worth trusting in 2026 are not the ones that promise you a beautiful line; they're the ones that show you the window, the costs, the caveats, and the grade — before you ask.