The whole track record.
Every slice, including misses.
Every prediction — across all markets — scored on the real close-to-close move, broken down by horizon, confidence, sector, region, model and regime. Calibration, risk-adjusted returns and platform discipline included. No cherry-picking.
95% confidence interval: 51.9%–52.5% · All markets
23,290 outcomes (18.1%) moved less than the symbol's noise band and are excluded as indeterminate — not counted for or against.
Top-conviction calls · held 25d
Live (served) predictions only — shadow-challenger calls, which users never saw, are excluded. No stop-loss and no price target: just the direction, held to the horizon. This is a research signal, not investment advice.
The trust funnel
Every slot the system tried, through to the cohort the headline accuracy is computed on — including the ones we declined to call. Hover a stage for its exact definition.
Accuracy over time
Rolling 7-day directional accuracy vs the 50% coin-flip line — All markets.
Recent validation activity
Predictions scored per day, by SETTLE date — the day the outcome actually resolved, not the day the call was made. This is the only series here that reaches today, which makes it the honest answer to “is the scoring engine running right now?”. Every chart below is by generation date and tapers at recent dates because those predictions have not matured yet.
Cumulative accuracy
Running accuracy across every scored prediction through each date — All markets. It converges as the sample grows: early swings shrink as evidence accumulates, so a flat right-hand tail is the sample getting large, not performance freezing. Ends short of today by design, because recent predictions are still maturing.
Daily accuracy vs 7-day rolling
Daily bars are noisy on thin days, so the rolling line is the one to read. Both are by generation date on the canonical cohort — the last stretch of each line is built from only the short-horizon calls that have matured, which is why it wanders more than the rest.
Scored predictions per day
Sample size behind each point of the accuracy charts above, by generation date. Spikes are scheduled batch prediction runs. The last few days read low because those predictions are still maturing — not because generation stopped.
Daily breakdown by maturation date
Each row is the complete cohort of predictions that came due that day, so the counts do not taper the way the generation-date charts above do. Click a row for the per-horizon split.
| Due date | Cohort | Scored | Correct | Pending | Accuracy | Rolling 7d | Status |
|---|---|---|---|---|---|---|---|
| 2026-09-17 | 1,008 | 980 | 348 | 28 | 35.5% | 38.5% | complete |
| 2026-09-16 | 2,965 | 2,868 | 865 | 97 | 30.2% | 39.2% | complete |
| 2026-09-15 | 3,174 | 3,096 | 982 | 78 | 31.7% | 41.0% | complete |
| 2026-09-14 | 4,161 | 4,095 | 1,666 | 66 | 40.7% | 44.0% | complete |
| 2026-09-11 | 4,482 | 4,445 | 1,650 | 37 | 37.1% | 45.6% | complete |
| 2026-09-10 | 3,614 | 3,574 | 1,486 | 40 | 41.6% | 47.5% | complete |
| 2026-09-09 | 3,326 | 3,276 | 1,596 | 50 | 48.7% | 49.5% | complete |
| 2026-09-08 | 3,676 | 3,617 | 1,551 | 59 | 42.9% | 50.1% | complete |
| 2026-09-07 | 1,997 | 1,975 | 945 | 22 | 47.9% | 52.3% | complete |
| 2026-09-04 | 4,542 | 4,485 | 2,303 | 57 | 51.4% | 52.8% | complete |
| 2026-09-03 | 3,320 | 3,286 | 1,712 | 34 | 52.1% | 53.6% | complete |
| 2026-09-02 | 3,050 | 3,006 | 1,429 | 44 | 47.5% | 54.2% | complete |
| 2026-09-01 | 4,005 | 3,974 | 2,145 | 31 | 54.0% | 54.5% | complete |
| 2026-08-31 | 2,913 | 2,896 | 1,564 | 17 | 54.0% | 54.6% | complete |
By horizon
Bullish vs bearish
Walk-forward accuracy
Accuracy on the most recent N days only — is recent live performance holding up?
Do the probabilities mean what they say?
Does confidence mean anything?
Actual directional hit-rate by the model’s own stated confidence bucket — higher confidence should mean more calls come out right. This is a different measurement from the reliability diagram above, which checks whether the calibrated probabilities match realized frequencies. Ranking well here and calibrating badly there is a real, common combination, so read them separately.
| Confidence band | Predictions | We said | We delivered | Delivered − stated |
|---|---|---|---|---|
| Very Low (<50%) | 19,447 | 47.8% | 49.4% | +1.6pp |
| Low (50-60%) | 29,614 | 54.7% | 55.6% | +0.9pp |
| Medium (60-70%) | 8,193 | 60.4% | 53.9% | −6.5pp |
| High (70-80%) | 18,276 | 74.3% | 53.2% | −21.1pp |
| Very High (80-90%) | 10,883 | 86.0% | 47.5% | −38.5pp |
| Extreme (90%+) | 18,930 | 90.0% | 50.6% | −39.4pp |
Calibration by model
lower Brier / ECE is betterThe headline numbers above are the blend of every model that produced a scored call. Split out, they disagree — which is the useful part. Delivered − stated is realized accuracy minus stated confidence, the same convention as the band table above: negative means that model talks a better game than it plays.
| Model | Scored | Avg confidence | Accuracy | Delivered − stated | Brier | ECE |
|---|---|---|---|---|---|---|
| v18 | 42,302 | 57.6% | 46.0% | −11.6pp | 0.267 | 0.116 |
| modular | 6,114 | 58.3% | 40.8% | −17.5pp | 0.276 | 0.175 |
| shadow | 1,551 | 58.8% | 45.3% | −13.6pp | 0.269 | 0.136 |
| besttoo few to judge | 26 | 58.0% | 42.3% | −15.6pp | 0.267 | 0.156 |
| news_causal_v2too few to judge | 7 | 56.2% | 0.0% | −56.2pp | 0.316 | 0.562 |
Risk-adjusted returns
Measured on the disciplined book — the signals we actually commit to after the confidence gate, not every raw signal. This is the money-weighted result.
Computed on all markets · all horizons · last 90 days, 143,727 closed trades.
Tail risk & efficiency
How bad the bad trades get. VaR is the return the worst 1-in-20 (5%) and 1-in-100 (1%) trades beat; CVaR is the average of that tail — the number that matters when the tail actually arrives.
Returns by horizon (holding-period)
How a user actually experiences it: hold from the call to its horizon. Sharpe here is annualized the correct way — per-trade × √(252 / horizon-days), not a calendar-year compound.
| Horizon | n | Avg hold return | Sharpe (holding-period) | Win rate | Max DD |
|---|---|---|---|---|---|
| 1d | 2,868 | -0.271% | -0.97 | 49.0% | -19.43% |
| 5d | 22,336 | +0.206% | 0.25 | 48.3% | -10.39% |
| 7d | 26,534 | +0.452% | 0.45 | 48.2% | -11.22% |
| 10d | 26,015 | +0.547% | 0.36 | 50.5% | -16.05% |
| 15d | 27,394 | +1.080% | 0.54 | 53.4% | -12.10% |
| 20d | 5,532 | +2.150% | 0.78 | 58.4% | -21.49% |
| 25d | 7,169 | +2.143% | 0.66 | 57.0% | -12.57% |
| 30d | 11,362 | +3.659% | 0.91 | 58.7% | -30.47% |
| 45d | 8,782 | +3.816% | 0.75 | 59.7% | -17.54% |
| 60d | 3,131 | +5.380% | 0.78 | 65.3% | -22.35% |
Trade funnel
Strategy returns (gated book)
Cumulative return of the disciplined trade book — the same daily equity series the Sharpe, Calmar and Max-drawdown figures above are built from, not the raw-signal firehose.
By asset class
By region
By sector
By model
Accuracy by market × horizon
Directional accuracy for every market at every horizon, against the 50% coin flip. This is the joint cut, not the two marginal breakdowns above: a market can look strong overall and still be at or below chance on the horizon you actually trade. Cohort: All markets · last 90 days.
| Market | 1d | 5d | 7d | 10d | 15d | 20d | 25d | 30d | 45d | 60d |
|---|---|---|---|---|---|---|---|---|---|---|
| United States | 43.2% n=185 | 42.7% n=5,243 | 38.3% n=2,043 | 48.6% n=2,567 | 50.6% n=2,816 | 54.4% n=182 | 50.5% n=745 | 50.5% n=542 | 46.3% n=257 | — |
| India | 28.6% n=227 | 47.6% n=2,860 | 41.9% n=5,059 | 46.9% n=2,938 | 49.5% n=3,128 | 51.8% n=629 | 49.6% n=716 | 55.7% n=673 | 51.2% n=453 | — |
| Europe | 52.2% n=910 | 49.1% n=4,944 | 48.9% n=5,942 | 53.3% n=3,861 | 54.9% n=4,511 | 65.9% n=478 | 62.1% n=824 | 63.0% n=635 | 60.0% n=350 | — |
| Asia-Pacific | 51.4% n=1,058 | 51.2% n=7,754 | 51.1% n=10,197 | 55.9% n=10,098 | 57.4% n=8,021 | 61.7% n=1,614 | 64.2% n=2,448 | 69.5% n=2,609 | 71.2% n=1,101 | 25.0% n=4 · thin |
| Crypto | 48.2% n=85 | 56.6% n=143 | 55.7% n=334 | 76.9% n=39 | 76.9% n=39 | 91.7% n=12 · thin | 100.0% n=22 · thin | 74.5% n=47 | 85.7% n=7 · thin | — |
| Forex | 53.3% n=167 | 52.4% n=42 | 42.1% n=235 | 59.9% n=132 | 68.9% n=45 | 0.0% n=11 · thin | 54.2% n=24 · thin | 75.0% n=12 · thin | 0.0% n=1 · thin | — |
| Commodities | 53.1% n=145 | 58.3% n=300 | 65.7% n=210 | 70.6% n=204 | 75.8% n=190 | 75.0% n=36 | 90.2% n=41 | 94.3% n=35 | 93.8% n=16 · thin | — |
Green is above the 50% coin flip, red is below; intensity saturates at ±10 points so a thin outlier can't shout down a real one. 9 cells are dimmed for a sample under 30 — those numbers are shown as measured, but they are not yet evidence of anything. A cell with no scored predictions shows an em-dash rather than a zero.
Returns by market × horizon
What following the calls actually returned, for every market at every horizon — the same joint cut as the accuracy grid above, in the unit that pays. These come from the trade ledger, not the scored prediction cohort, so the sample sizes are their own and will not match cell for cell. Cohort: All markets - last 90 days.
| Market | 1d | 5d | 7d | 10d | 15d | 20d | 25d | 30d | 45d | 60d |
|---|---|---|---|---|---|---|---|---|---|---|
| United States | +0.13% n=645 · 43% won | -0.30% n=5,459 · 38% won | -0.19% n=4,194 · 37% won | +0.01% n=3,401 · 43% won | -0.04% n=3,945 · 42% won | +0.22% n=981 · 46% won | +0.31% n=1,638 · 44% won | +1.40% n=2,712 · 52% won | +2.09% n=1,603 · 55% won | +3.07% n=284 · 61% won |
| India | -0.40% n=770 · 32% won | -0.15% n=4,182 · 41% won | -0.19% n=4,312 · 40% won | -0.11% n=4,117 · 42% won | +0.11% n=4,371 · 45% won | +0.81% n=1,615 · 50% won | +0.58% n=943 · 47% won | +1.00% n=2,309 · 50% won | +1.89% n=640 · 49% won | +4.90% n=134 · 66% won |
| Europe | +0.13% n=1,451 · 48% won | +0.01% n=3,943 · 43% won | +0.14% n=5,490 · 45% won | +0.32% n=4,435 · 49% won | +0.39% n=5,017 · 49% won | +1.73% n=1,076 · 59% won | +1.38% n=1,408 · 55% won | +1.33% n=3,043 · 55% won | +1.85% n=1,357 · 55% won | +1.33% n=104 · 49% won |
| Asia-Pacific | +0.07% n=1,094 · 43% won | +0.18% n=7,008 · 44% won | +0.28% n=10,262 · 43% won | +0.49% n=10,595 · 45% won | +0.89% n=11,020 · 50% won | +1.42% n=4,236 · 53% won | +1.87% n=3,738 · 54% won | +2.54% n=7,018 · 56% won | +3.91% n=3,411 · 63% won | +4.90% n=1,121 · 67% won |
| Crypto | -0.85% n=62 · 31% won | +0.43% n=85 · 46% won | +0.49% n=278 · 48% won | +2.20% n=47 · 55% won | +1.24% n=54 · 46% won | +4.82% n=27 · 70% won · thin | +5.63% n=42 · 64% won | +5.47% n=85 · 64% won | +1.27% n=20 · 50% won · thin | -6.50% n=2 · 50% won · thin |
| Forex | +0.00% n=188 · 54% won | +0.07% n=48 · 58% won | +0.07% n=169 · 58% won | +0.28% n=143 · 69% won | +0.11% n=56 · 63% won | -1.03% n=21 · 24% won · thin | +0.08% n=46 · 54% won | +0.14% n=64 · 55% won | +1.32% n=15 · 73% won · thin | -0.21% n=3 · 33% won · thin |
| Commodities | +0.18% n=193 · 50% won | +0.70% n=286 · 52% won | +1.22% n=297 · 57% won | +1.97% n=276 · 62% won | +1.71% n=241 · 58% won | +4.56% n=81 · 73% won | +3.36% n=67 · 69% won | +4.71% n=112 · 71% won | +5.23% n=46 · 70% won | -2.64% n=4 · 25% won · thin |
| INDX | -0.17% n=1 · 0% won · thin | -0.09% n=4 · 25% won · thin | +0.25% n=8 · 38% won · thin | +3.69% n=6 · 100% won · thin | +2.23% n=13 · 69% won · thin | -1.49% n=3 · 33% won · thin | +3.35% n=4 · 100% won · thin | +2.24% n=11 · 91% won · thin | +2.02% n=1 · 100% won · thin | — |
| Latin America | +0.57% n=200 · 56% won | -0.35% n=512 · 37% won | -0.11% n=764 · 42% won | +0.78% n=791 · 50% won | -0.42% n=560 · 40% won | +1.65% n=284 · 52% won | +0.24% n=206 · 48% won | -0.66% n=334 · 43% won | -1.16% n=225 · 44% won | +6.54% n=23 · 70% won · thin |
| Middle East & Africa | — | +0.19% n=71 · 48% won | -0.17% n=714 · 41% won | +0.73% n=241 · 52% won | +1.18% n=204 · 54% won | +1.97% n=114 · 61% won | +0.22% n=125 · 46% won | +1.66% n=109 · 55% won | -4.07% n=59 · 34% won | — |
Green is a profit, red is a loss — the pivot is zero, not the 50% of the accuracy grid, because this is a return and the null result is “made nothing”. Intensity saturates at ±2% per trade, so a thin outlier can't shout down a real one. 17 cells are dimmed for a sample under 30 — those numbers are shown as measured, but they are not yet evidence of anything, and the annualized view magnifies them hardest. A market/horizon pair with no trades shows an em-dash rather than a zero.
Net edge vs benchmark
Portfolio return minus what the benchmark did over the same window, per market, each against its own named local index. Positive means the calls beat simply holding that index; negative means they did not, and it is printed here either way.
Measuring net edge against each market's benchmark…
Performance by market regime
Not just accuracy per regime — the return, Sharpe and win rate of the calls made in each market state. This is where you find out whether an edge is real or is one regime carrying the whole record.
Volatility x trend (ex-ante tags)
The market state recorded at the moment each call was made — knowable in advance, so these numbers are the honest ones.
| Regime | n | Accuracy | Avg return | Sharpe | Win rate | Coverage |
|---|---|---|---|---|---|---|
| high mean reverting | 71,289 | 53.6% | +1.65% | 2.15 | 53.6% | 50.6% |
| medium mean reverting | 68,608 | 51.1% | +0.86% | 2.03 | 51.1% | 48.7% |
Realized move size (post-hoc diagnostic)
| Move size | n | Accuracy | Avg return | Sharpe | Win rate | Coverage |
|---|---|---|---|---|---|---|
| small | 537 | 52.9% | +0.22% | 1.08 | 52.9% | 0.4% |
| large | 152 | 57.9% | +2.54% | 3.96 | 57.9% | 0.1% |
Best tracked
Worst tracked
— shown honestlyDiscipline & coverage
Most attempts never become a call at all — the gate abstains upstream, before a prediction is issued, and those abstentions are counted against us in the two figures on the right. What remains is scored in full. Committed calls here are the published cohort, which applies no confidence floor — every scored prediction is included.
Abstention by horizon
platform-wide 59.1%Share of attempted calls the engine declined to publish, per horizon. Higher means more was withheld.
| Horizon | Attempted | Abstained | Abstention rate |
|---|---|---|---|
| 1d | 109,087 | 79,532 | 72.9% |
| 5d | 225,716 | 123,891 | 54.9% |
| 7d | 161,533 | 69,762 | 43.2% |
| 10d | 250,245 | 138,898 | 55.5% |
| 15d | 301,892 | 173,229 | 57.4% |
| 20d | 233,899 | 133,922 | 57.3% |
| 25d | 241,493 | 141,403 | 58.5% |
| 30d | 354,446 | 225,143 | 63.5% |
| 45d | 309,760 | 192,066 | 62.0% |
| 60d | 310,730 | 193,074 | 62.1% |
| 90d | 303,956 | 186,615 | 61.4% |
The intelligence stack
Which models are serving the predictions scored above, and the rules that govern replacing them.
Currently serving
Cross-reference these against the “By model” breakdown above — a version with a small share of the book is a challenger being measured, not the champion.
DAILY_LIGHT_v20260622_1414_90d_AUTop forecasters
No per-agent forecast records scored yet. This leaderboard fills once the swarm attributes settled outcomes to individual agents — until then there is nothing here to rank, and we would rather show that than an empty table.
Cohort definition & audit trail
canonical_trust_strictExactly which predictions the numbers above are computed over, and what was filtered out to get there. Published so the accuracy figure can be checked rather than taken on faith.
Strict directional accuracy — a prediction is scored correct only when the symbol moved beyond its noise band and the model called that direction. Band-indeterminate outcomes (see indeterminate.band_description for the exact rule in force) are excluded from both the numerator and the denominator — not counted for or against. Cohort also excludes archived, abstained, out-of-universe, and shadow-challenger predictions. No confidence floor is applied: every scored prediction is in the cohort, including the ones the model was least sure of. Short (DOWN) calls are excluded: the platform publishes long calls only, so the track record describes the product actually sold. Short predictions made before 2026-08-23 remain in the database but are not scored here.
Cohort invariants
Every prediction counted above satisfies all of these. Each one removes a population that would otherwise flatter the number.
- 1
archived != True (exclude pre-2026-05-07 broken-pipeline cohort) - 2
abstain != True (exclude model-declined predictions) - 3
in_universe != False (exclude out-of-universe symbols) - 4
is_neutral != True (exclude neutral predictions from directional accuracy) - 5
is_shadow != True (exclude shadow-challenger predictions — the ~75% not shown to users) - 6
confidence >= 0 (no confidence floor — every scored prediction is included, including the lowest-confidence ones) - 7
direction in ("UP","up","Up") (long-only — short calls are neither published nor scored)
Filters applied
The query this snapshot was built with. null means unfiltered.
- days back
- 90
- exchange filter
- —
- model version
- —
- horizons
- —
Excluded from the cohort
Printed exactly as the backend reports it — where a count is not separately tracked it says so rather than showing a zero.
- archived count excluded
- see #202 archive script — 11,153 docs archived 2026-05-16
- abstained count
- 1,657,535
- out of universe count
- not separately surfaced
- neutral count
- 23,290
- low confidence count
- not separately surfaced
Top-conviction ledger
The highest-conviction decile of settled calls, ranked on the calibrated probability — not the stated confidence. Held to the horizon with no stop and no target, which is why these rows reconcile with the hero number above and not with the traded book.
This is the conviction slice itself. These are the individual predictions behind every aggregate number above — one row per call, with the outcome it actually settled to. The aggregates stay public; the row-level record is for signed-in members.
Audit ledger
Every scored prediction in the canonical cohort, newest settlement first. This is the same set of rows the headline accuracy is computed over — the six invariants noted beneath the table are applied identically — so this is the panel to audit against. The recent-settled panel is a wider, non-canonical cohort and will not tie out.
This is the canonical audit trail. These are the individual predictions behind every aggregate number above — one row per call, with the outcome it actually settled to. The aggregates stay public; the row-level record is for signed-in members.
Recent settled predictions
The latest calls that have already settled — pending ones are pulled from the same response and filtered out here, so this is scored rows only. This cohort is not the canonical one. It excludes archived and shadow-challenger rows only — it does not apply the six canonical invariants the headline accuracy above is computed on, so tallying these rows will not tie out to that number. The audit ledger panel is the cohort that does. Every market, because the scope above is set to Global.
This is the recent settled record. These are the individual predictions behind every aggregate number above — one row per call, with the outcome it actually settled to. The aggregates stay public; the row-level record is for signed-in members.
How we measure
Directional
Right if the price moves the way we said by the horizon. Unambiguous, checkable.
Close-to-close
Official closes at the horizon date — not intraday ticks we could pick to flatter.
No look-ahead
Locked before the outcome window opens. Out-of-sample, every time.
Probability, then scored
Every call ships a stated probability and we publish how far it lands from the realised hit-rate — including when that gap is bad.
The record is the proof. Now see the picks.
Today’s high-conviction calls for Forex.