The transparency tests, in detail
The benchmarks are here because the AI-signal market runs thick with declarations no outsider can stand up. Three of the five tests get a page to themselves, since those are the ones services trip over most. The complete five-test method lives on the method page.
What the benchmarks are really sorting
A buyer rarely compares two near-identical services. In practice you are choosing between types — a black-box bot, a chat channel, a social caller, an aggregator, or a systematic desk — and each type fails or passes the tests as a category. The benchmarks make that structural difference visible, so a slick “AI” presentation cannot hide which category a service belongs to. The table below sorts the field on the two tests that decide most cases: is the call locked on-chain before its outcome, and is the accuracy figure a live denominator rather than a backtest.
| Service type | Locked on-chain? | Live denominator? | Why it sits there |
|---|---|---|---|
| Black-box AI bot | No | Backtest only | Hidden logic; a simulation stands in for live calls. |
| Messaging-app channel | No | Rarely | Posts are the operator's to rewrite or erase; the misses disappear. |
| Social-media caller | No | No | Threads can be pruned; the money often comes from referral links. |
| Aggregator / re-poster | No | No | Echoes other people's calls with no checking of its own. |
| Copy-trading room | Rarely | Sometimes | Headline results exist, but seldom proof per individual call. |
| Systematic, timestamped desk | Yes | Yes | Named rule, an on-chain stamp on each call, and outside review. |
The bottom row is the lone category that comes out solid across both columns, which is the structural case for the desk pick — not that it shouts louder, but that it is the one kind of service an outsider can actually check for themselves. The three benchmarks below each take one test apart in full.
What makes these two columns do the deciding
Two of the five tests carry nearly all the sorting weight. Locked on-chain before its outcome is the one you cannot bolt on afterward: a service either pinned its calls in public before they resolved or it did not, and no amount of later tidying alters that. A live denominator is the one you cannot fabricate short of lying outright: an accuracy figure becomes evidence only when the full count of real-time calls, failures included, sits beside it — which is precisely what a backtest lacks. Pass both and you have been handed a record open to scrutiny. The other three — the stated rule, the measured grade, clean incentives — matter, but they mostly ratify a verdict the first two have already delivered rather than reverse it.
The three tests with their own page
Transparent vs black box
Why a disclosed, rule-based method beats an undisclosed model you are asked to trust on faith.
Locked before the outcome
Why a public on-chain receipt is the test a fast algorithmic call cannot fake its way around.
A re-runnable record
What a real live record contains, and why a backtest or a highlight reel is not one.
The remaining two tests — the measured conviction grade and clean (non-affiliate) incentives — are covered on the method page, because they are quicker to check and rarely the deciding factor.