What we measure, how we measure it, and the current state of the measurement — without the "73% accuracy" slogans that typically come without sources. Logged-in users see the per-signal live track inside the dashboard.
Honest disclosure. The full math stack went live April 2026. We've been recording every score every day since 2026-04-09. Short-horizon accuracy (7-day, 14-day, 19-day) is tracked live below — early but real. The 30-day horizon now has at least one complete forward-return window — its IC publishes alongside the shorter horizons as the weekly cron rolls. Every number is either a live diagnostic or a live accuracy measurement — never an in-sample backtest dressed up as out-of-sample performance.
Honest disclosure. Framler's math stack (multi-factor ensemble, BOCPD regime posterior, copula-blended Markowitz, bin-conditional conformal intervals, Kalman dynamic exposures) deployed to production April 2026. signal_history snapshots have accumulated daily since 2026-04-09. Live IC for 7/14/19-day horizons is published below from the weekly accuracy-check cron. The 30-day horizon now has its first complete window of post-snapshot price data and joins the published horizons via the weekly accuracy-check cron. Sharpe / CVaR need ≥3 months of windows for stable estimates. Sector + regime breakdowns shown when sample sizes are sufficient (≥30 paired observations per bucket).
Were we right? — every call, graded
Every ticker page shows a green/red record of what that stock actually did after each published score. This is the same record summed across the whole universe — every score we have ever published whose 7-day outcome is already known, graded by one fixed rule: a “call” is a day the score left the middle zone (above 60 = we said up, below 40 = we said down), and it counts as right only if the stock then moved our way over the following 7 days.
WE MADE A CALL ON 18,893 OF 64,664 GRADED TICKER-DAYS — RIGHT 9,838 · WRONG 9,055 (52.1% right — a coin flip so far)
APR 2026RIGHT 433 · WRONG 394 — 827 calls of 3,025 graded days
MAY 2026RIGHT 2,612 · WRONG 2,288 — 4,900 calls of 18,312 graded days
JUN 2026RIGHT 3,462 · WRONG 3,392 — 6,854 calls of 21,243 graded days
JUL 2026 · STILL MATURINGRIGHT 3,331 · WRONG 2,981 — 6,312 calls of 22,084 graded days
Raw receipts, not a significance test: consecutive days overlap (the same stock scored above 60 five days running produces five correlated calls, not five independent trials), so no p-value is quoted here. The pre-registered pass/fail test with fixed rules and kill dates is the public bet; the per-stock version of this record is on every ticker page. The latest month always looks thinner — its most recent days haven't matured yet.
The skeptic's questions — answered straight
“52.1% right is a coin flip — doesn't that mean the engine has no edge?”
On this record so far, no proven edge — and that is our own headline above, not a critic’s discovery. Binary grading of 7-day moves is also the bluntest possible yardstick: consecutive-day calls are heavily correlated, so the raw percentage carries less information than it seems. The decisive instrument is the pre-registered quintile test on the bet page — fixed rules, fixed deadlines, public verdict either way.
“If the record is a coin flip, why publish it at all?”
Because a forward record only counts if it starts before it looks good. Anyone can show a winning backtest after the fact; a dated, Bitcoin-anchored record published through the unflattering phase is the only evidence a skeptic cannot dismiss — whichever way it ends.
“What would make you admit the engine doesn’t work?”
Fixed dates, already published: if the bet’s confidence interval has not cleared zero by 2027-01-01 (7-day test) or 2027-07-01 (30-day test), the verdict NOT PROVEN goes on the record. The live IC tiles below show how the measurement stands in the meantime — including when it reads zero.
“Stocks get delisted and replaced — isn’t that quiet survivorship bias?”
When a ticker delists or is renamed, its published history stays in the record: the graded days above keep counting it, and the ledger hashes cannot be rewritten. We add a replacement to keep the universe near 1,000. What we never do is fabricate returns for dead tickers or drop losing history.
Current state
Universe
90/1001
90 of 1001 tickers updated in the rolling 24-hour window. Full universe scoring runs Mon-Fri 06:00 UTC; weekend coverage is partial (only the news-sentiment and macro crons fire). The number drops over weekends and rises again Monday.
Regime detection
risk_on
BOCPD posterior as of 2026-08-09. Fires continuously on SPY returns.
Effective breadth
47%
Grinold-Kahn effective count of independent factors out of 13 raw. Low ratio = factors are redundant; high ratio = genuinely different signals.
Implied 30-day move (VIX)
±4.3%
VIX 14.9 as of 2026-08-07. Forward-looking implied σ for the S&P over the next 30 calendar days. Complements the historical SPY drift used by the engine.
Tail-dependence
234 pairs
Max upper-tail alignment: moderate. Non-parametric co-crash probability per factor pair, per regime.
Prediction intervals
Calibrated
53 Mondrian bins calibrated from accumulating residuals.
Factor weights
Literature prior
Prior weights are the Asness-Moskowitz-Pedersen + Novy-Marx + Sloan literature defaults. Activates after forward returns accumulate.
Calibration IC by horizon
Walk-forward cross-validated Information Coefficient produced by the weekly calibrate-weights cron. Different from the live-accuracy block above — that measures the end-to-end score-to-return relationship; this measures the per-horizon IC of the inverse-covariance-shrunk factor stack the engine uses to form the composite. OOS IC is the headline number; train-IC is shown as the second line for shrinkage-leakage sanity-check (large gap = overfitting risk).
30-day OOS IC
pending
Calibrating — 42200 samples accumulated so far. Next refresh: Sunday 12:00 UTC.
90-day OOS IC
pending
Calibrating — 2760 samples accumulated so far. Next refresh: Sunday 12:00 UTC.
Live accuracy — Information Coefficient (IC)
Spearman rank correlation between Framler score on day T and realised price return from T to T+N, averaged across all overlapping windows. Refreshed weekly. 83 days of accumulation as of 2026-08-03. Industry context: a strong multi-factor signal is typically IC 0.03-0.06 out-of-sample.
7-day IC
-0.002
11 windows · hit rate 36.4% (windows where IC > 0).
14-day IC
0.081
5 windows · hit rate 100% (windows where IC > 0).
19-day IC
0.064
4 windows · hit rate 50% (windows where IC > 0).
Rolling IC — last 12 weekly readings
Each tile shows the rolling mean ± 1σ across the last 12 weekly readings. Sparkline traces the actual readings chronologically (oldest left → newest right). Values clip at ±0.20 for the line; the y-axis is centred at zero. Persistent IC ≥ 0 means the engine is empirically predictive over the rolling window, not just on the most-recent snapshot.
7-day rolling IC
-0.014± 0.013
n = 12 weekly readings · 1/12 positive
14-day rolling IC
+0.097± 0.022
n = 12 weekly readings · 12/12 positive
19-day rolling IC
+0.137± 0.049
n = 12 weekly readings · 12/12 positive
Where the engine works — and where it doesn't
Each sector is placed into one of four calibration tiers based on the rank correlation between the composite score and realised forward returns over recent weekly windows. We publish the tier; the magnitude stays internal because raw per-cohort IC is part of the engine moat. Sectors with fewer than 30 paired observations surface as Pending — the weekly cron promotes them once the sample threshold clears.
Communication Services
●POSITIVE
Treat the composite as an early positive read. Cross-check the per-ticker Confluence patterns before acting.
Consumer
●POSITIVE
Treat the composite as an early positive read. Cross-check the per-ticker Confluence patterns before acting.
Consumer Cyclical
●POSITIVE
Treat the composite as an early positive read. Cross-check the per-ticker Confluence patterns before acting.
Healthcare
●POSITIVE
Treat the composite as an early positive read. Cross-check the per-ticker Confluence patterns before acting.
Industrials
●POSITIVE
Treat the composite as an early positive read. Cross-check the per-ticker Confluence patterns before acting.
Technology
●POSITIVE
Treat the composite as an early positive read. Cross-check the per-ticker Confluence patterns before acting.
Utilities
●POSITIVE
Treat the composite as an early positive read. Cross-check the per-ticker Confluence patterns before acting.
Real Estate
◐NOISY
Treat the composite as informational. Lean on Confluence patterns, Insider clustering, and Options flow — sub-signals that retain edge in narrative-driven cohorts.
Semiconductors
◐NOISY
Treat the composite as informational. Lean on Confluence patterns, Insider clustering, and Options flow — sub-signals that retain edge in narrative-driven cohorts.
Basic Materials
⚠WEAK
Do not act on the composite alone. Rely on Confluence patterns, Insider clustering, and Options flow until the next calibration window restores edge.
Consumer Defensive
⚠WEAK
Do not act on the composite alone. Rely on Confluence patterns, Insider clustering, and Options flow until the next calibration window restores edge.
Energy
⚠WEAK
Do not act on the composite alone. Rely on Confluence patterns, Insider clustering, and Options flow until the next calibration window restores edge.
Financial Services
⚠WEAK
Do not act on the composite alone. Rely on Confluence patterns, Insider clustering, and Options flow until the next calibration window restores edge.
By regime — 14-day pooled IC
Regime label comes from the BOCPD posterior dominant state on the snapshot day. Coverage is sparse early — most days so far have been tagged the same regime, so cross-regime comparison needs more time.
risk_off
0.033
n = 1962 score-return pairs in risk_off regime.
risk_on
0.016
n = 49989 score-return pairs in risk_on regime.
By confluence pattern — 14d realised return when fired
Mean realised 14-day return on tickers that triggered each pattern, plus hit rate (fraction of fires with positive return). Bullish patterns with hit rate near 50% or mean return near 0 are calibration candidates — the engine surfaces this honestly rather than hiding underperforming patterns. Patterns with fewer than 5 fires hidden as too noisy.
PRICE_AHEAD_OF_FUNDAMENTALS
+3.37%
hit 57.9% positive returns · n = 214 fires
PEAD_DRIFT
+2.28%
hit 59.5% positive returns · n = 1464 fires
PHARMA_FAILURE_HIGH
+2.06%
hit 63.1% positive returns · n = 483 fires
QUALITY_CRACK
+2.03%
hit 65.5% positive returns · n = 307 fires
NO_EDGE
+2.01%
hit 55.6% positive returns · n = 12362 fires
EARNINGS_VALIDATED
+1.87%
hit 57.6% positive returns · n = 2484 fires
GLAMOUR_UNWIND
+1.85%
hit 64.4% positive returns · n = 205 fires
DEEP_VALUE_PIOTROSKI
+1.84%
hit 58.1% positive returns · n = 2186 fires
SHORT_SQUEEZE_SETUP
+1.82%
hit 55.6% positive returns · n = 3108 fires
CONTRARIAN_BOTTOM
+1.69%
hit 53.9% positive returns · n = 798 fires
EARNINGS_MISS_DRIFT
+1.55%
hit 66.7% positive returns · n = 63 fires
ACCRUALS_RED_FLAG
+1.25%
hit 60.8% positive returns · n = 102 fires
PRICED_FOR_PERFECTION
+1.14%
hit 59.7% positive returns · n = 370 fires
GROWTH_REGIME_ALIGNED
+1.02%
hit 51.6% positive returns · n = 161 fires
QUALITY_COMPOUNDER
+0.83%
hit 54.4% positive returns · n = 3769 fires
SECTOR_BREAKDOWN
+0.42%
hit 57.1% positive returns · n = 254 fires
CONFLICTING_SIGNALS
+0.28%
hit 45.3% positive returns · n = 997 fires
INSIDER_DISTRIBUTION
-0.42%
hit 52.3% positive returns · n = 1171 fires
MOMENTUM_BREAKDOWN
-2.57%
hit 45.6% positive returns · n = 349 fires
PHARMA_CATALYST_NEAR
-2.87%
hit 37.8% positive returns · n = 37 fires
VALUE_TRAP
-3.18%
hit 39.3% positive returns · n = 466 fires
Universe and distribution disclosure
Two structural facts shape what scores you see today. We surface them here because most quant products hide the same limits behind glossy backtests.
Survivorship bias
The universe is roughly 1,000 currently-listed US, European, and Asia-Pacific equities. Companies that delisted, went bankrupt or merged out of existence are not in the snapshot — so the distribution of factor scores skews healthier than the historical true population. Real bearish setups exist but they are under-represented relative to a hypothetical "all listings ever" universe.
Score distribution
Composite scores cluster slightly above the 50 neutral midpoint with a tighter standard deviation than the cross-section literature predicts. The reason is two-fold: the Q×V×M interaction factor compresses to neutral when its three inputs sit near 50, and the confluence pattern library is currently bullish-tilted (more bullish than bearish patterns documented across academic research). We expanded the bearish pattern coverage on 2026-05-22 and have a cross-section z-score recalibration through a shadow pipeline planned for Q3 2026 — neither was hidden, both are published openly here.
Additional metrics on the roadmap
7-day IC is live above. The following longer-horizon and portfolio-level metrics need 30+ days of forward returns or several months of windows before they stabilise. Each is standard in the quant-research literature and will be published per-regime, per-pattern, and aggregated as data accumulates.
Information Coefficient (IC)
Spearman rank correlation between the composite score and realised forward returns. Published per factor, per regime, per horizon. Honest out-of-sample IC for a strong multi-factor signal is typically 0.03-0.06.
Sharpe ratio
Annualised return / annualised volatility of a long/short decile portfolio formed on the composite. Top quint institutional strategies: 0.8-1.2.
Max drawdown
Peak-to-trough decline of the same portfolio. Tracked with and without the conformal-interval-based position-sizing overlay.
Conformal coverage
Fraction of tickers whose realised forward return falls within the published prediction interval. Target matches the stated coverage level.
Per-pattern hit rate
For each pattern in the library, fraction of fires that deliver the expected-direction return over the pattern’s typical horizon. Reported alongside sample size so confidence intervals are legible.
How we will benchmark
Absolute return numbers without a benchmark are meaningless. Every metric above will be reported against three reference portfolios:
SPY — the market baseline. Any strategy below SPY net of effort is failing.
Equal-weighted universe — controls for the universe-selection bias (our 1000+ names are curated).
Fama-French 5-factor portfolio — controls for known factor exposures so any alpha isn't just re-branded beta.
If the composite beats all three over at least a 6-month forward window, the edge is credible. If not, we publish that too.