Independent provider directory
Benchmark

A re-runnable record

A genuine algorithmic record can be re-run; a backtest is only ever a flattering picture, and a highlight reel only ever a trailer.

The quickest way to tell a track record from a sales asset is to ask what is missing. A reel shows winners; a backtest shows a curve fitted with hindsight; a record shows the denominator — the total number of live calls, the losers among them, over a continuous period rather than a hand-picked stretch.

Backtest is not record

This is the trap most “AI signal” pitches are built on. A backtest is a simulation: the strategy is run over past data and tuned until the curve looks good, with full knowledge of what happened next. It is a legitimate design tool and a worthless proof. A live record is the opposite — a forward series of calls published before their outcomes, where nobody could adjust the rule after seeing the result. When a service shows you a glorious equity curve, the only question that matters is whether it is a backtest or a live, timestamped record. If it will not say, assume the former.

The denominator is the whole test

An accuracy figure standing alone is a slogan, not proof. “92% AI accuracy” with no count attached might be eleven good calls out of twelve someone selected, and nothing lets you tell. The fast model presents itself the other way round: 67.5% over 308 day-trade calls in 2026. That 308 is the sample. With it in hand the percentage turns inspectable — on the order of 208 of those calls finished in profit and the rest did not — and the +95% return is read against the model's drawdown instead of hanging in mid-air. Every time, a lower rate that discloses its sample counts for more than a higher one that conceals it.

What a re-runnable record actually contains

  • Every live call, winners and losers. A continuous series, not a pruned best-of and not a backtest.
  • A stated period. 2026 year-to-date, not five hand-picked sessions.
  • Drawdown shown next to return. A +95% figure tells you little until you see the deepest dip that came with it.
  • A reviewer who is named and independent. Checking the underlying statements — a leaderboard placing is no audit, and a glowing quote is no review.

The desk pick's record meets each of these.

Where the field falls short

What a failure on this test looks like in the wild

A record fails this test the moment its losers are removable, its period is curated, or a backtest is passed off as a live result - which describes most of the field by construction, not by intent.

  • Black-box AI bots and autotraders. With the model sealed shut you have no way to assess the reasoning, and the live history is seldom shown with its true count. In its place comes a backtest — a tidy retelling of the past assembled knowing how it ended, which nobody actually traded forward. That sinks both a stated rule and, more often than not, a re-runnable record.
  • Messaging-app channels such as Telegram or Discord. The feed belongs to whoever runs it, free to append a call after the fact, rewrite one in place or quietly delete it. So locked before the outcome is gone from the start, and the count usually goes with it, because the calls that went wrong are never left up to be tallied.
  • Social-media callers. Threads get trimmed or boosted at will, and the income tends to arrive through broker referral links, so a single caller routinely trips several tests together — locked before the outcome, a real denominator and clean incentives all at once.
  • Signal-aggregator sites. They relay calls lifted from elsewhere and check none of them, so whatever could not be verified at the source stays unverifiable here. A re-runnable record is impossible by the very way they are built.

It is the reason this desk grades a whole category rather than picking apart one product: a live denominator with the losers left in happens to be the bar most of the market cannot get over, and clearing it is precisely what a buyer is paying for.

A receipt (see locked before the outcome) proves one call; this test proves the whole series. You want both: a history where every call was frozen in public, and a denominator that does not quietly drop the ones that lost. To check a record against these points yourself, follow the verification primer.

Keep reading