Methodology

How the forecasts are made, how they are scored, and everything we tested and refused to ship.

The event

A forecast asserts the probability of exactly one event: that the dividend- and split-adjusted total return of the stock, from the adjusted close on the price-basis date to the adjusted close on the 21st trading day afterward, is strictly greater than zero.

How it is scored

Resolution is count-based: exactly 21 observed trading days after the basis date, on the exchange calendar derived from the pinned price history — never a wall-clock date. Each forecast is written to an append-only ledger the moment it is issued and is scored automatically at its horizon. Rows are immutable, enforced by the database itself. A correction, if ever needed, is a new superseding row that carries a reason and points at the row it replaces; the original is never altered or removed.

The probability, and its uncertainty

The confidence interval and effective sample size come from a cluster bootstrap: forecasts issued in the same weekly batch share the market's realised path, so they are not independent. Resampling whole issuance-weeks rather than individual forecasts prevents the false precision that treating them as independent would produce. The effective sample size is reported alongside every probability, and when only one batch has resolved it is reported as one — not padded.

What we tested — and refused to ship

Every candidate predictor family is auditioned in a walk-forward harness before it may touch the published probability. Models are refit on resolved-only history over a 2012–2021 development window, S&P 500 point-in-time, 1-month horizon. To clear the bar a family must earn a positive out-of-sample Brier skill score against both climatology and the price/macro baseline, with a calibration slope between 0.8 and 1.25. The 2022-onward holdout is never used for auditions.

A negative skill score means the family did worse than simply quoting the base rate. Every family below is negative. We publish this because a track record of refusals is the honest version of a track record.

#DateCandidate predictor familySkill vs base rateSkill vs baselineVerdict
02026-07-22Price/macro analogues (k-nearest-neighbour)−0.078FAIL
0b2026-07-22Price/macro analogues + walk-forward recalibration
honest but empty; extremes were anti-predictive
−0.008FAIL
12026-07-22Price/macro direct model (baseline)−0.027FAIL
22026-07-22Earnings surprises
worse than baseline despite look-ahead-flattered estimates
−0.029−0.002FAIL
3a2026-07-23Insider transactions — naive (untyped Form 4)−0.028−0.000FAIL
3b2026-07-23Insider transactions — typed conviction gate
no probability skill; a small +28 bp/mo return tilt exists and ships only as a descriptive panel fact
−0.029−0.001FAIL

Where that leaves the product

As of 2026-07-23, across three families — price/macro analogues, earnings surprises, and insider transactions (naive and typed) — none distinguished individual S&P 500 stocks from the roughly 59% one-month base rate by enough to register out-of-sample. So the product ships no stock-specific directional claim: probabilities are the universe base rate (climatology) with an honest confidence interval, alongside per-stock distribution ranges and typed descriptive panels. A future data source graduates to the probability only by earning a positive out-of-sample skill score in this same log.

Families not yet tested (each requires its own data download and audition entry before any product claim): options flow, congressional and institutional trading, point-in-time estimate revisions, and short-interest. Congressional disclosures are shown on stock pages as descriptive facts only — with no historical-tendency number — precisely because that family has no audition entry yet.

Crypto: a separate record, distributions only

For crypto the event is the same shape as for stocks, but on the calendar-day calendar (coins trade every day): the adjusted total return from the close on the basis date to the close 30 calendar days later, strictly greater than zero. Windows are 7 / 30 / 91 calendar days for the 1w / 1m / 3m horizons.

The crypto universe (v1 'majors') is the set of coins with continuous daily price history since 2017-08-01 that rank in the top 10 by current market capitalization. Two limitations we state plainly: this list contains only survivors, so its historical statistics are flattered by construction; and it is ranked by current — not point-in-time — market cap. A point-in-time universe arrives when a historical-rankings source does.

Before any crypto forecast was published, the same walk-forward discipline was run on crypto majors over a development window ending 2021-12 (the 2022-onward data was held out). The result, dated 2026-07-25: the distribution ranges are well-calibrated — the 90% drawdown and rise bands held about 93% of the time at every horizon, inside the 85–95% target. But the directional base rate did not predict the forward up-rate out of sample (calibration slope near zero or negative at every horizon). So crypto ships its distribution ranges and says plainly that its monthly direction is not something our harness can calibrate — no headline probability.

HorizonnCalibration slopeDrawdown coverageRise coverage
1w1,7803.3893.3%92.1%
1m1,750−0.0893.3%93.1%
3m1,660−0.3196.5%92.5%

Crypto records never blend with stocks — separate universe, pool, and scoreboard.

Research tool, not investment advice — historical frequencies, not recommendations. Full terms.