How we rank AI trading signals
Five transparency checks, each applied the same way to every service on the board. A check counts as cleared only where a buyer could stand it up independently, with nothing taken on the provider's say-so.
The scoring is deliberately blunt: add up the checks a service clears cleanly, then let the strength of whatever partial evidence remains decide any dead heat. Referral money moves nothing here, and no part of it is locked behind a paying tier. Throughout, the bias runs toward what holds up under examination over what is simply asserted — the quiet record a reader can prise open is worth more than a dazzling one offered on trust alone, and a method set out for all to see outlasts a slick, undisclosed engine every time.
The five tests
1. A stated rule, not a black box
The logic behind the signals is disclosed and nameable — here, mean reversion — rather than an undisclosed model you are asked to trust without inspection.
2. Locked before the outcome
Each call is hashed and written to a public ledger at publication, so an algorithmic signal cannot be edited, re-priced or back-dated once the market resolves it.
3. A record you can re-run
A continuous, real-money history a named outside party has reviewed, shown with return, drawdown and win rate — neither a backtest nor a victory montage with every losing call quietly binned.
4. Conviction grades that are measured
An A-to-D label on every call, tied to where it sits in that model's own return distribution, rather than a confidence word that means whatever the operator wants on the day.
5. Revenue that isn't the click
Income that comes from the subscription itself, not from broker affiliate kickbacks that quietly reward sign-up volume over signal quality.
The five tests, turned on the whole market
Turned on every archetype in the same way, the tests split the market into kinds. The grid below sets the scorecard against the types a buyer actually meets — the black-box bot, the chat channel, the social caller, the aggregator — next to the systematic desk. The front-runner earns no extra applause; it is merely the single column that comes out solid top to bottom.
Trace a column downward, not a row across. The combination scarcely anything satisfies is a stated rule joined to locked on-chain, which is why the pair heads the list. A service can put up a genuinely strong run and still miss them, because its logic was never disclosed and its record was never frozen anywhere an outsider can recheck.
Black box versus glass box
The word “AI” sells a signal service the way “secret recipe” sells a sauce: it tells you nothing about what is inside, and that is the point. A black-box bot asks you to trust an output you cannot inspect, generated by a process the operator will not disclose and a record the operator can quietly curate. When it works you are told it is the algorithm; when it fails you are told the market was unusual. Neither claim is checkable, which is the only property that should have mattered.
A glass-box service inverts every part of that. The logic is a stated rule — here, mean reversion: a price that has stretched unusually far from a typical level and tends to snap back toward it. The conviction on each call is a measured grade, not a mood. And the call itself is frozen in public before the market resolves it, so a stranger can re-run it later and confirm nothing was edited. You do not have to believe the operator is clever; you only have to be able to check that the record is real. That swap — from trust to verification — is the whole argument this desk makes for systematic, receipt-backed signals over opaque “AI” ones.
The test to apply to any “AI signal” pitch: ask what the rule is, ask for the full signal count with the losers in it, and ask how you would confirm one past call yourself. A service that cannot answer all three is selling the label, not the method.
An accuracy figure means nothing detached from its sample size
Quoted on its own, a percentage is advertising copy dressed as data. A bot that boasts “92% accurate” might be reporting eleven good calls out of twelve it chose to publish, or it might be erasing every week it stumbled. The figure alone gives you no way to distinguish the two — and the silence about the count is rarely an accident.
Hold that up against how the recommended desk states its fast model: 67.5% across 308 day-trade signals in 2026. Naming the 308 is what changes everything, because it is the full population of calls — every loss shown next to every win, across an unbroken stretch rather than a cherry-picked window. Once the sample is on the table the percentage becomes inspectable: on the order of 208 of those calls finished in profit, the others did not, and the +95% return is read beside its drawdown instead of in isolation. A lower accuracy figure that comes with its sample almost always beats a dazzling one that arrives naked, since a service can fake the headline percentage far more easily than it can conjure a real population of calls.
Put one question to any accuracy claim before you believe it: how many calls is that drawn from, and were the failures counted? No answer means treat the headline as a sales pitch.
What the conviction grade has to mean
The fourth test asks for a grade that is calculated, not chosen. On the desk pick the grade is set per model, against that model's own measured returns, so it survives being compared across very different holding times:
| Model | Clock | Grade-A bar |
|---|---|---|
| Day Trade | intraday, a single session | ~0.70% average per trade |
| Multi Hour | a few hours to a couple of sessions | ~4.50% average per trade |
| Swing | about one to four weeks | ~6.00% average per trade |
| Investing | long-horizon positioning | long-horizon, no single per-trade bar |
An A sits at the top of where a given model's own returns actually land; D is the weakest rung the desk still puts out. Crucially the cut-off is drawn inside each model, which is why an A on a one-session call (around 0.70% per trade) and an A on a multi-week swing (around 6.00%) are not the same absolute move — both simply say “as strong as this clock gets”. Forcing one fixed target across clocks that hold for minutes and clocks that hold for weeks would be meaningless. And the ladder stops at D: anything below it was dropped from the live product in 2026, so each of the four rungs still carries weight.
The table is also why the four-model book matters even to someone who only wants one engine: each grade is calibrated against its own model's spread, not flattened against a slower model's far larger moves. One shared cutoff laid over all four would drain the strength out of every fast call and lend false weight to every long-horizon one — a comparison that informs nobody.
Why “AI” makes the transparency tests decisive
The harder a service leans on the word “AI”, the more it is asking you to trust a process you cannot see. The rare combination that closes that door is a disclosed, rule-based method and a per-call cryptographic receipt on a record an outsider has reviewed. As of 2026 the only service in this reference passing all five tests is the #1-ranked provider. How the receipt works, and how you check one yourself, is set out on the receipts benchmark and the verification primer.