Skip to main content

Accountability

We keep score.

Every day we record a verdict for the whole universe — a PAS (Academic Score), a Guru Pyramid tier, a best-match guru. This page measures how those cohorts actually performed against the S&P 500 afterward. It is computed mechanically, never hand-picked, and it is a forward, out-of-sample record — not a backtest.

Scorecard vintage September 11, 2026 · cohorts May 16, 2026 Aug 12, 2026 (71 verdict days) · live since May 16, 2026

Top decile vs SPY · 90d

+1.0%

n = 7,155

Top-decile win rate vs SPY

48%

share of names beating SPY

Decile 10 − decile 1 spread

-74.3%

the score's ranking power

Does the score rank the future?

Pooled excess return vs SPY per Pythia decile at the 90-day horizon. A working score slopes upward: low deciles underperform, high deciles outperform.

Current read: the top decile is underperforming the bottom decile by 74.3 pp — the early-window sample is not yet ranking forward returns; we publish it anyway. Thin-coverage names can inflate scores until the coverage gate ships.

Cohort by cohort

Each point is one verdict day's 90-day forward return — the top decile, the bottom decile, and SPY over the identical windows. No pooling, no smoothing.

The full scorecard

By PAS decile

Decile of each cohort's own PAS distribution (10 = highest).

n = names pooled · cohorts = verdict days · P(beat SPY) = per-cohort win rates shrunk toward the family average — estimated in-sample and graded as new cohorts mature

By PAS decile — forward return and win-rate vs SPY. n is names pooled; cohorts are verdict days.
BucketAvg returnSPYExcess (median)Win rateP(beat SPY)n, names pooledCohorts, verdict days
Decile 10 (highest)+2.0%+1.3%-0.8%mean +0.8%46.6%49.1%9% pooled35,14063
Decile 9+3.5%+1.2%-0.2%mean +2.3%47.7%51.8%9% pooled34,77063
Decile 8+2.5%+1.3%-0.3%mean +1.2%47.0%52.6%9% pooled37,91263
Decile 7+6.8%+1.0%+0.3%mean +5.8%49.6%52.8%9% pooled27,10363
Decile 6+4.4%+1.4%-0.5%mean +3.0%45.8%51.8%9% pooled41,79663
Decile 5+6.2%+1.3%-0.7%mean +4.9%46.2%50.1%9% pooled37,56363
Decile 4+15.6%+1.2%-0.6%mean +14.4%47.1%49.5%9% pooled32,38963
Decile 3+11.0%+1.1%-0.7%mean +9.9%46.9%47.8%9% pooled34,07163
Decile 2+30.9%+1.3%-1.4%mean +29.6%44.6%44.2%9% pooled34,28363
Decile 1 (lowest)+37.3%+1.2%-3.0%mean +36.1%41.9%39.7%9% pooled34,22663

By pyramid tier

Canonical pass-count band at the time of the verdict.

n = names pooled · cohorts = verdict days · P(beat SPY) = per-cohort win rates shrunk toward the family average — estimated in-sample and graded as new cohorts mature

By pyramid tier — forward return and win-rate vs SPY. n is names pooled; cohorts are verdict days.
BucketAvg returnSPYExcess (median)Win rateP(beat SPY)n, names pooledCohorts, verdict days
DEEP VALUE+3.1%+1.6%-0.3%mean +1.5%49.0%54.0%15% pooled3,45761
QUALITY VALUE+3.1%+1.3%+0.2%mean +1.8%49.8%55.9%14% pooled37,10063
FAIR VALUE+3.4%+1.1%+0.1%mean +2.2%49.5%53.3%14% pooled112,11063
OVERVALUED+17.4%+1.0%-1.0%mean +16.4%45.2%46.6%14% pooled288,59763

By best-match guru

The guru archetype each company most resembled that day.

n = names pooled · cohorts = verdict days · P(beat SPY) = per-cohort win rates shrunk toward the family average — estimated in-sample and graded as new cohorts mature

By best-match guru — forward return and win-rate vs SPY. n is names pooled; cohorts are verdict days.
BucketAvg returnSPYExcess (median)Win rateP(beat SPY)n, names pooledCohorts, verdict days
Marks Cycle+19.1%+1.0%-0.4%mean +18.1%48.0%48.1%23% pooled115,23163
Pabrai Value+20.9%+1.2%-1.1%mean +19.6%44.3%46.2%23% pooled68,61963
Buffett Moat+12.4%+0.7%-0.6%mean +11.8%46.7%46.1%23% pooled55,61963
Safety First+2.3%+0.8%-0.5%mean +1.5%46.8%50.0%23% pooled36,09263
Greenblatt Magic+4.0%+0.9%+0.7%mean +3.1%52.7%52.4%23% pooled29,96163
Drucker Efficiency+0.9%+1.7%-1.6%mean -0.8%39.9%47.1%29% pooled24,05546
CAPEX Efficiency+4.8%+2.2%-1.8%mean +2.6%41.0%45.0%49% pooled13,23719
Thorndike Outsiders+2.9%+2.5%-1.5%mean +0.3%43.0%43.9%28% pooled9,92048
Value Line Cash+2.1%+2.9%-2.2%mean -0.9%38.3%46.0%40% pooled6,20427
Owner Earnings+3.9%+2.2%-2.0%mean +1.8%41.9%46.1%29% pooled5,31544
Buffettology Growth+1.6%+2.1%-1.3%mean -0.4%42.4%52.3%29% pooled4,68144
Lynch Growth-0.9%+1.3%-1.9%mean -2.2%41.3%45.6%23% pooled67363
Methodology & caveatsShow

Ledger status

as of September 11, 2026 · cohorts May 16, 2026Aug 12, 2026 (71 verdict days) · live since May 16, 2026

How a number gets here

Each daily verdict anchors every company to its adjusted-close price on that date (no look-ahead). Once a horizon of 30, 90, 180, or 365 days has fully elapsed, we measure each company's total return to the next available bar and compare it to SPY over the exact same window. Companies without an anchor price or a matured end bar are excluded — never imputed. Deciles and quintiles are assigned over the full anchored cohort on the verdict day (not just the names that later survive to a matured bar), and companies with identical scores always share a bucket — assignment is deterministic and identical across horizons.

How the table is pooled

Across all matured cohorts we report the n-weighted mean return, SPY return, excess (return − SPY), and win-rate (share of names that beat SPY) for each PAS decile, pyramid tier, best-match guru, and PGS / PCI quintile. A bucket is shown only with at least 5 pooled observations — thinner samples are noise and are dropped. Tables also report cohorts (verdict days that contributed) beside n (names pooled).

Caveats — read before trusting

  • Survivorship. Names that delist or stop trading lose their end bar and leave the cohort, so returns tilt slightly toward survivors. These are price returns of names that kept trading, not a tradable portfolio P&L.
  • Short history. Snapshots began May 16, 2026; samples are small and noisy until cohorts accumulate. This is a forward record that grows in real time — not a backtest.
  • New score families start at zero. PGS and PCI quintile cohorts accrue only from the day their snapshot columns shipped — earlier verdicts have no point-in-time record of those scores, so their history is never reconstructed after the fact.
  • No look-ahead, but adjusted-close mixing. Base anchors and end bars share the adjusted-close basis, so there is no look-ahead. One residual: the base anchor is a true adjusted close while ordinary daily bars store the raw close as adjusted — at most a dividend-yield-scale (≈1–2%/yr) drift on individual names. The SPY benchmark is refreshed on a fully-adjusted basis, so the comparison line is unaffected.
  • Not a recommendation. This measures historical price behavior of scored cohorts. Past performance does not predict future results, and nothing here is investment advice.

Why we grade the process, not just the outcome

Markets are a wicked learning environment: over short windows a sound process loses often and an unsound one wins often, so judging a method purely by its recent outcomes systematically rewards luck (Hogarth, Lejarraga & Soyer 2015, Current Directions in Psychological Science). That is why every claim on this page carries its sample size and a confidence interval, why verdicts refuse when the interval cannot carry them, and why a negative spread is published as plainly as a positive one.

The benchmark is SPY (S&P 500) on a dividend-adjusted basis — a deliberately unforgiving bar. Full methodology: docs/ACCOUNTABILITY_METHODOLOGY.md.

The full record lives in the app

Sign in for the interactive explorer — every horizon, per-bucket cohort histories, the PGS/PCI score-family tables, and each company's own since-first-verdict track record.