Signal Board16 ideas filedidea 2 · data-ready 7 · under test 7 · live 0 · retired 9snapshots as of 2026-09-24 · code unknownledger last measured 2026-09-27
Does anything beat its nulls? The store cannot say; the ledger's dated answer is CANNOT_TELL on each of the 2 graded entries.
The read returned 7 snapshot rows carrying 35 null comparisons: 7 of the 12 grid cells hold at least one, with 35 null comparisons on the newest snapshot of those cells; 0 rows are filed under no ledger entry and 0 are scoped outside the grid. The newest row was written 2026-09-24 at code_version unknown, and no interval is stored beside any of them. This page can show each signal's hit rate against each null's, in hit-rate points, and cannot call a single cell beaten or lost: the interval that decides that is a seeded 2,000-draw date-clustered bootstrap the web cannot reproduce (DECISIONS.md ADR-0023) and grade_report does not persist. The ledger's answers are dated separately — measured on an earlier run, never to be read as the snapshot's date: uw_flow_alert_direction is CANNOT_TELL (the test could not have seen it either way — the verdict is about the sample), measured 2026-09-20: Hit rates 47-52% across horizons. Zero of thirty tests beat any of the five nulls and zero lost to one; the symmetry matters, because neither half is evidence. — could see +-3.9% to +-14.0% in hit-rate points, against a plausible edge of one to three points: this design cannot see one; max_pain_convergence is CANNOT_TELL (the test could not have seen it either way — the verdict is about the sample), measured 2026-09-20: Beats always_short, the climate and a coin at 21 days - and ties always_long (+0.3%) and trend_following, which are the two nulls that also take a side. That is what a long bias in a rising sample looks like. Expiry week alone has 2-3 usable dates and reports insufficient. — could see +-5.1% to +-10.7% in hit-rate points. A sign on a skill figure below is not a verdict.
Ledger: worker/app/core/ledger.py, mirrored and pinned by test. Snapshots: core.calibration_snapshots and core.calibration_skill, append-only by trigger (R11); this page never triggers `grade_report --write`.
Under test — the test has run and could not settle it
7 entries the ledger has measured and not closed. 2 can carry snapshots (a grader subject is filed under them); the other 5 persist nothing, so the verdict is the whole record. Order: under review first, then entries with attached snapshots, then measured_on descending, then ledger order — a declared order, never a statistic (ADR-0022). measured; the ledger has not closed it.
uw_flow_alert_directionvendorCANNOT_TELLunder test · measured 2026-09-20 · not fitted: no parameter here was chosen by optimising an outcome
claim
A flow alert - calls or puts bought at the ask - is followed by a move in that direction.
finding
Hit rates 47-52% across horizons. Zero of thirty tests beat any of the five nulls and zero lost to one; the symmetry matters, because neither half is evidence.
sample
24 sessions (772-782 alerts), one call per ticker per session
could see
+-3.9% to +-14.0% in hit-rate points, against a plausible edge of one to three points: this design cannot see one
method
Directional calls graded against five nulls, date-clustered, family-wise over the sweep.
fitted
not fitted: no parameter here was chosen by optimising an outcome
Live — empty, and that is the finding
graded, beats its nulls, being watched: no verdict maps here.
Nothing is live
Nothing is live. ADR-0005 killed the house directional call, so no house signal is under grading; Owner Decision #2 (DESIGN.md:435) — does any house call survive, demoted and honestly graded — is open; and no ledger verdict admits a signal here, which is itself the record that nothing has. What the graded entries found is on their cards under Under test, in the ledger's words and with the ledger's dates. What would put an entry here: a filed ledger entry whose snapshot series beats always_long AND trend_following at the family-wise level over 6 or more snapshots with a drift state of ok — none of which this page can create: it reads, and db.ts opens every transaction READ ONLY. A signal that reaches this lane leaves it the moment drift reads drifted or stale — shown as 'under review', not removed (DESIGN.md:199 X).
ADR-0005 · DESIGN.md:435 Owner Decision #2
Retired — what was tried and what killed it
NO_EVIDENCE and CANNOT_TELL look alike and are not (ledger.py:16-25): only the first is a finding, which is why the 4 NO_EVIDENCE entries sit here and the 6 CANNOT_TELL entries sit above. The could-see line on every card is what lets a reader check the verdict against the power. Obituaries resurface: `python -m app.core.ledger <word>` finds them by substring, and the ?q= box above is the same lookup.
tested — no evidence at the size claimed 4
dispersion_compressedhouseNO_EVIDENCEretired · tested — no evidence at the size claimed · measured 2026-09-27
claim
When 21-day cross-sectional sector dispersion enters its own bottom decile, the market is trading as one thing and is fragile to a single shock, so what follows is worse than average (DESIGN.md:185, §3.3 S).
finding
Sixty-two episodes clear the report's floor for a 1% move at 5 and 10 sessions (25 and 50), and at both the index did slightly BETTER than baseline, not worse: +0.26% [-0.26%, +0.80%] and +0.17% [-0.70%, +0.95%], up in 69-71% of episodes against 62-67%. A 1% adverse move would have shown, and the point estimates sit on the wrong side of zero for the claim. The 21-session cell is unpowered (needs 103) and reads +0.38% [-0.98%, +1.54%]. The test is on the mean forward return; a claim about the tail of the distribution was not tested.
sample
62 episodes over 2,509 sessions of sector_dispersion_21d, 2016-10-03 to 2026-09-25; outcome SPY close-to-close, entered the session after the reading
Data-ready — described, never graded
described, never graded; not filed in the ledger. Each row names what the call would be and what a grading run would need; the file:line is printed as text and pins nothing on the other page. Exit: a Call loader plus a grading run filed as a ledger entry.
the call: seven deterministic gates, one light per Core single per session
to grade: a light is a gate, not a direction: grade it as an event class — what followed 'kill' vs 'clear' over 5/21 sessions, event_study.base_rate with the 20-date floor
not filed in the ledger — which is different from having passed (ledger.py:361-364).
the call: where each name sits in its own history on one date, across eleven fields
to grade: a named cut as a Call per name per session; rarity_ranked_confluence in Retired is the obituary for the combinatorial version
Ideas — named in DESIGN §3.4 W, no data
named in DESIGN §3.4 W, no collector. Exit: a collector with a declared consumer (R5) moves an idea to data-ready.
analyst_ratings‘Do analyst ratings predict?’
DATA_SOURCES.md:105 names /api/screener/analysts for the Dossier context module only; no collector exists, and R5 needs a named consumer before one is built — this page is that consumer. Moves to data-ready when a collector lands rows with a row floor and a point-in-time class, and a Call loader turns a rating change into a dated direction.
uw_seasonality‘Does UW seasonality hold out of sample?’
No collector for UW's monthly averages, and DESIGN.md:181 (Q) says they are duplicative and must not be the headline (the vendor's own get_market_seasonality reports positive_months_perc 1.1000 for XLRE month 7 — an impossible 110%). calendar_seasonality under Retired is THIS warehouse's price statistic over its own sessions — NO_EVIDENCE — and is not the same claim; grading UW's figure means collecting it first and testing it only on sessions after the collection date.
drift.py reports; this page decides what a state means. drifted or stale demotes to under review. unknown is not ok: an alarm over one measurement is a chart that is green because nothing feeds it.
export restricted — calibration snapshots derived from UW flow alerts and option chains may not be redistributed
subject
snapshots
first / last
n latest (calls)
state
detail
liveness
stale on
power
max_pain:converges [universe:core] @1d
1
2026-09-24 / 2026-09-24
2395
unknown
1 measurements: the baseline needs 5 before a shift can be told from the noise it started with
one measurement: growth cannot be judged
2026-10-15
cannot be evaluated: detectable not stored
max_pain:converges [universe:core] @21d
1
2026-09-24 / 2026-09-24
1663
unknown
1 measurements: the baseline needs 5 before a shift can be told from the noise it started with
one measurement: growth cannot be judged
What this page cannot show, and why
Each of the first five is computed by the grader, printed, and discarded (Skill.ci / .mde / .identical, Grade.n_dates / .detectable; grading.py:98-159). Persisting them is a worker migration and an ADR (ADR-0030, PROPOSED alongside this page), not a page change. The interval itself is a seeded 2,000-draw date-clustered bootstrap the web cannot reproduce (DECISIONS.md:1127-1133), so the only honest path is to store it.
the interval and its half-width (MDE)Skill.ci and Skill.mde are computed, printed and discarded; without them no cell can be called beaten or lost to.
the number of distinct call datesthe 20-date gate is on dates, so n here is calls (Grade.n_dates).
the identity flagSkill.identical is not persisted; derived here from the subject's fixed direction instead.
per-null na null that could not judge a call drops it from that baseline only, and the count it judged is not stored.
detectable per snapshotGrade.detectable is not persisted, so UNPOWERED cannot be evaluated.
a run logan insufficient grade is not written, so a missing cell is ambiguous.
regime scopeck_calib_scope admits it; the grader writes none.
confidence and BrierNULL by design, not by omission: the graded signals state no probability.
window_daysthe price-history length, not the sample window — grade_report.py:276-278 — never printed as a sample size here.
Describes the record and grades nothing itself. Never runs `grade_report --write` (R11 — a snapshot is permanent by trigger). Personal use (ADR-0014): no export, no CSV; every graded row is a statistic over the vendor's personal-use alerts and chain, and a division is not a licence (registry.py:1865).
revive if
A year of alerts. The verdict is about the sample, not the signal, and will stay so until the interval is narrower than the claim.
code
app/core/grading.py
The ledger's verdict was measured on 2026-09-20; the snapshots below were computed on 2026-09-24. A later run of the grader does not update a verdict — the ledger changes only when a new entry is filed.
export restricted — calibration snapshots derived from UW flow alerts and option chains may not be redistributed
uw_flow_alert:call_ask_side @1d
vendorcomputed 2026-09-24hit 50.6% · n = 891 calls (sessions not stored)
price history 3674 days (not the call window — grade_report.py:276-278) · code unknown · confidence not stated · Brier none (the signal states no probability)
always_longthe same call every time: this null IS the signal here
always_short+1.6 ptsnull 49.0%
base_rate_regime-2.2 ptsnull 52.8%
seeded_random-0.1 ptsnull 50.7%
trend_following+1.5 ptsnull 49.1%
Each difference is paired over the calls that null could judge (grading.py:367-373), and that count is not stored: the bar is the cell's hit rate over all its calls, so bar minus rule need not equal the printed points, and the cell's n is not any one null's n.
driftunknown1 measurements: the baseline needs 5 before a shift can be told from the noise it started with
stale from 2026-10-15 (21 days after the write, to the hour) unless the grader writes again · one measurement: growth cannot be judged
uw_flow_alert:call_ask_side @5d
vendorcomputed 2026-09-24hit 51.1% · n = 748 calls (sessions not stored)
price history 3674 days (not the call window — grade_report.py:276-278) · code unknown · confidence not stated · Brier none (the signal states no probability)
always_longthe same call every time: this null IS the signal here
always_short+2.3 ptsnull 48.8%
base_rate_regime-5.0 ptsnull 56.1%
seeded_random+3.7 ptsnull 47.3%
trend_following
uw_flow_alert:call_ask_side @21d
No snapshot for this cell. grade_report --write appends only sufficient grades (20 or more distinct call dates, grade_report.py:358-371), so this is either 'insufficient on the last run' or 'never written' — the store cannot tell the two apart. Last write for any cell: 2026-09-24.
uw_flow_alert:put_ask_side @1d
vendorcomputed 2026-09-24hit 49.9% · n = 880 calls (sessions not stored)
price history 3674 days (not the call window — grade_report.py:276-278) · code unknown · confidence not stated · Brier none (the signal states no probability)
always_long+0.1 ptsnull 49.8%
always_shortthe same call every time: this null IS the signal here
base_rate_regime+2.7 ptsnull 47.2%
seeded_random+1.9 ptsnull 48.0%
trend_following
uw_flow_alert:put_ask_side @5d
vendorcomputed 2026-09-24hit 48.2% · n = 738 calls (sessions not stored)
price history 3674 days (not the call window — grade_report.py:276-278) · code unknown · confidence not stated · Brier none (the signal states no probability)
always_long-3.4 ptsnull 51.6%
always_shortthe same call every time: this null IS the signal here
base_rate_regime+4.3 ptsnull 43.9%
seeded_random-0.5 ptsnull 48.8%
trend_following
uw_flow_alert:put_ask_side @21d
No snapshot for this cell. grade_report --write appends only sufficient grades (20 or more distinct call dates, grade_report.py:358-371), so this is either 'insufficient on the last run' or 'never written' — the store cannot tell the two apart. Last write for any cell: 2026-09-24.
No interval is stored, so nothing here can say a null was beaten or lost to; that claim is the ledger's finding above, measured 2026-09-20, and the family its intervals were narrowed over is the one its own finding and method lines name. The nulls are only useful read together: a signal that beats the coin, the climate and always_short while tying always_long and trend_following is a side, not a signal (grading.py:185-200).
n counts calls, not sessions: forty alerts on one afternoon are one afternoon. The interval behind each snapshot resampled call DATES and the floor is on distinct dates (grading.py:35-37); the date count is not stored. A stored snapshot passed that floor on its write date; the ledger's sample line names the sessions.
The web computes no beaten set and no direction-bias verdict: both need the interval, which is not stored. grading.py:190's DIRECTION_BIAS_NULLS are always_long and trend_following, the two nulls that take a side of their own.
The alarm's UNPOWERED state — the grader's interval wider than the 5-point drift worth catching — cannot be reported here or by drift_report: the half-width is not persisted and drift_report never passes it (drift_report.py:44-53), so a subject whose alarm cannot fire looks like one whose alarm has not fired. The entry's could-see line above is the only record of that width.
Calibration curve — not measurable
uw_flow_alert:call_ask_side: mean_conf and brier are NULL on all 2 newest snapshots read.
uw_flow_alert:put_ask_side: mean_conf and brier are NULL on all 2 newest snapshots read.
A calibration curve plots stated confidence against realised hit rate, per bin, with n per bin. grade_report sets no confidence on the calls it builds: a flow alert says 'this happened', not 'this is 68% likely', and max pain states a direction, not a probability. Inventing a confidence to fill the column would make the Brier a fact about the invention (grading.py:30-33), so the grader leaves it unset and no curve is drawn. What would make it measurable: a signal that states a probability with each call — a house signal, which is Owner Decision #2 (DESIGN.md:435), or a vendor field that is a probability — graded with Call.confidence set on every call; grade() then fills mean_conf and brier (grading.py:290-299), and per-bin rows need a scope this schema does not have. Then this module draws confidence bins against hit rate with n per bin, and skill stays in hit-rate points against every null, never one minus the Brier over its coin value.
Skill by regime — not measured
None of the 4 newest snapshots read for this entry is scoped to a regime: each is universe:core.
grade_report grades one scope, the whole universe, though ck_calib_scope admits a regime scope. What is regime-conditioned already is the NULL: base_rate_regime is the climate's P(up) in the regime in force on each call's own date. What would fill this: grade() per regime_code with 20 distinct call dates in each cell. The ledger's sample line for this entry reads '24 sessions (772-782 alerts), one call per ticker per session', and splitting a sample of that size four ways is how a base rate lies (ledger: resolution_base_rates).
max_pain_convergenceexternalCANNOT_TELLunder test · measured 2026-09-20 · not fitted: no parameter here was chosen by optimising an outcome
claim
Price drifts toward max pain, most visibly in expiry week.
finding
Beats always_short, the climate and a coin at 21 days - and ties always_long (+0.3%) and trend_following, which are the two nulls that also take a side. That is what a long bias in a rising sample looks like. Expiry week alone has 2-3 usable dates and reports insufficient.
sample
43-60 sessions, ~1,550-2,190 calls, since 2026-06-17
could see
+-5.1% to +-10.7% in hit-rate points
method
The stored max_pain_distance_pct as a directional call, acted on the NEXT session's close (the chain is fetched at 19:00), five nulls, family-wise.
fitted
not fitted: no parameter here was chosen by optimising an outcome
revive if
It beats always_long and trend_following, not merely the coin. Until then the reading is a direction bias.
code
app/core/grade_report.py
The ledger's verdict was measured on 2026-09-20; the snapshots below were computed on 2026-09-24. A later run of the grader does not update a verdict — the ledger changes only when a new entry is filed.
export restricted — calibration snapshots derived from UW flow alerts and option chains may not be redistributed
max_pain:converges @1d
externalcomputed 2026-09-24hit 54.5% · n = 2395 calls (sessions not stored)
price history 3674 days (not the call window — grade_report.py:276-278) · code unknown · confidence not stated · Brier none (the signal states no probability)
always_long+4.9 ptsnull 49.6%
always_short+4.3 ptsnull 50.2%
base_rate_regime+5.5 ptsnull 49.0%
seeded_random
max_pain:converges @5d
externalcomputed 2026-09-24hit 54.7% · n = 2258 calls (sessions not stored)
price history 3674 days (not the call window — grade_report.py:276-278) · code unknown · confidence not stated · Brier none (the signal states no probability)
always_long+4.2 ptsnull 50.6%
always_short+5.4 ptsnull 49.4%
base_rate_regime+6.9 ptsnull 47.8%
seeded_random
max_pain:converges @21d
externalcomputed 2026-09-24hit 54.9% · n = 1663 calls (sessions not stored)
price history 3674 days (not the call window — grade_report.py:276-278) · code unknown · confidence not stated · Brier none (the signal states no probability)
always_long-1.2 ptsnull 56.1%
always_short+11.0 ptsnull 43.9%
base_rate_regime+8.7 ptsnull 46.2%
seeded_random
max_pain:converges_in_expiry_week @1d
No snapshot for this cell. grade_report --write appends only sufficient grades (20 or more distinct call dates, grade_report.py:358-371), so this is either 'insufficient on the last run' or 'never written' — the store cannot tell the two apart. Last write for any cell: 2026-09-24. The ledger's max_pain_convergence entry (measured 2026-09-20) records: 'Expiry week alone has 2-3 usable dates and reports insufficient.'
max_pain:converges_in_expiry_week @5d
No snapshot for this cell. grade_report --write appends only sufficient grades (20 or more distinct call dates, grade_report.py:358-371), so this is either 'insufficient on the last run' or 'never written' — the store cannot tell the two apart. Last write for any cell: 2026-09-24. The ledger's max_pain_convergence entry (measured 2026-09-20) records: 'Expiry week alone has 2-3 usable dates and reports insufficient.'
max_pain:converges_in_expiry_week @21d
No snapshot for this cell. grade_report --write appends only sufficient grades (20 or more distinct call dates, grade_report.py:358-371), so this is either 'insufficient on the last run' or 'never written' — the store cannot tell the two apart. Last write for any cell: 2026-09-24. The ledger's max_pain_convergence entry (measured 2026-09-20) records: 'Expiry week alone has 2-3 usable dates and reports insufficient.'
No interval is stored, so nothing here can say a null was beaten or lost to; that claim is the ledger's finding above, measured 2026-09-20, and the family its intervals were narrowed over is the one its own finding and method lines name. The nulls are only useful read together: a signal that beats the coin, the climate and always_short while tying always_long and trend_following is a side, not a signal (grading.py:185-200).
n counts calls, not sessions: forty alerts on one afternoon are one afternoon. The interval behind each snapshot resampled call DATES and the floor is on distinct dates (grading.py:35-37); the date count is not stored. A stored snapshot passed that floor on its write date; the ledger's sample line names the sessions.
The web computes no beaten set and no direction-bias verdict: both need the interval, which is not stored. grading.py:190's DIRECTION_BIAS_NULLS are always_long and trend_following, the two nulls that take a side of their own.
The alarm's UNPOWERED state — the grader's interval wider than the 5-point drift worth catching — cannot be reported here or by drift_report: the half-width is not persisted and drift_report never passes it (drift_report.py:44-53), so a subject whose alarm cannot fire looks like one whose alarm has not fired. The entry's could-see line above is the only record of that width.
Calibration curve — not measurable
max_pain:converges: mean_conf and brier are NULL on all 3 newest snapshots read.
max_pain:converges_in_expiry_week: no snapshot was read, so there is no confidence column to describe.
A calibration curve plots stated confidence against realised hit rate, per bin, with n per bin. grade_report sets no confidence on the calls it builds: a flow alert says 'this happened', not 'this is 68% likely', and max pain states a direction, not a probability. Inventing a confidence to fill the column would make the Brier a fact about the invention (grading.py:30-33), so the grader leaves it unset and no curve is drawn. What would make it measurable: a signal that states a probability with each call — a house signal, which is Owner Decision #2 (DESIGN.md:435), or a vendor field that is a probability — graded with Call.confidence set on every call; grade() then fills mean_conf and brier (grading.py:290-299), and per-bin rows need a scope this schema does not have. Then this module draws confidence bins against hit rate with n per bin, and skill stays in hit-rate points against every null, never one minus the Brier over its coin value.
Skill by regime — not measured
None of the 3 newest snapshots read for this entry is scoped to a regime: each is universe:core.
grade_report grades one scope, the whole universe, though ck_calib_scope admits a regime scope. What is regime-conditioned already is the NULL: base_rate_regime is the climate's P(up) in the regime in force on each call's own date. What would fill this: grade() per regime_code with 20 distinct call dates in each cell. The ledger's sample line for this entry reads '43-60 sessions, ~1,550-2,190 calls, since 2026-06-17', and splitting a sample of that size four ways is how a base rate lies (ledger: resolution_base_rates).
breadth_narrowhouseCANNOT_TELLunder test · measured 2026-09-27 · not fitted: no parameter here was chosen by optimising an outcome
claim
When the advance narrows to three sectors of eight or fewer (regime.NARROW_BREADTH on sector_breadth_200dma), what follows for the index is worse than an average session (DESIGN.md:185, §3.3 S).
finding
Every horizon spans zero, and the sign is the claimed one at every horizon: SPY -1.12% against baseline over 5 sessions [-3.01%, +0.35%], -1.91% over 10 [-5.13%, +0.47%], -1.14% over 21 [-4.47%, +1.70%]; up in 49-51% of episodes against 61-67% for the baseline. Thirty-seven episodes is under the report's own floor for a 1% move at 10 and 21 sessions (50 and 103), and the intervals actually drawn were about twice the printed floors, which are the optimistic number their docstring says they are. The report's previous production run, on 2026-09-22 and before it corrected for its nine cells, read this axis's 5- and 10-session intervals excluding zero with upper bounds of -0.02% and -0.09%: the two marginal exclusions a nine-cell sweep produces on noise. #175 added the 0.05/9 correction for exactly that, and the intervals here, drawn under it three sessions later, include zero at both.
sample
37 episodes (entries into the tail, not sessions in it) over 2,530 sessions of sector_breadth_200dma, 2016-09-01 to 2026-09-25; outcome SPY close-to-close, entered the session after the reading
could see
printed floors +-0.81% at 5 sessions, +-1.15% at 10, +-1.67% at 21, at an uncorrected 95%; the intervals drawn are family-wise at 0.05/9 and their half-widths were 1.7%, 2.8% and 3.1%, all wider than a 1% move
method
Level rule on the registered metric, episodes by breadth.entries(), re-dated by tradeable_on(), event_study.base_rate against a same-dates baseline, nine intervals Bonferroni-corrected (breadth_report.py).
fitted
not fitted: no parameter here was chosen by optimising an outcome
revive if
Fifty episodes for the 10-session horizon and 103 for 21 - at the observed rate of about four a year, 2030 and decades later - or an interval at any horizon that excludes zero at the family-wise 0.05/9 (an uncorrected one did on 2026-09-22, which is why the correction exists). The sign is worth remembering and not worth acting on.
code
app/core/breadth.py
No snapshots: app/core/breadth.py persists nothing — it prints — so the verdict above is the whole record. What would attach a trace: a --write path for this study, which is a worker change and a decision (R11).
correlation_spikehouseCANNOT_TELLunder test · measured 2026-09-27 · not fitted: no parameter here was chosen by optimising an outcome
claim
When the 63-day average pairwise sector correlation enters its own top decile, diversification has already stopped working and drawdowns follow (DESIGN.md:185, §3.3 S).
finding
Fifteen episodes in ten years, and the base rate needs twenty distinct event dates before it prints one, so no interval was drawn at any horizon. A 63-day correlation moves slowly and a spike lasts months, so entries are rare by construction: about one and a half a year.
sample
15 episodes over 2,468 sessions of sector_corr_63d, 2016-11-30 to 2026-09-25; below the 20-date floor at every horizon
could see
nothing was tested; the printed floors for 15 episodes are +-1.28% at 5 sessions, +-1.81% at 10 and +-2.62% at 21, against a 1% move that needs 25, 50 and 103
method
Point-in-time expanding percentile (20 prior readings before any is formed), top-decile rule, episodes by entries(), re-dated by tradeable_on(); event_study.base_rate refused at its 20-date floor.
fitted
not fitted: no parameter here was chosen by optimising an outcome
revive if
Five more episodes reach the 20-date floor, about 2030 at the observed rate - and reaching it buys an interval, not power: 25 episodes is the floor for a 1% move at 5 sessions and nothing shorter.
code
app/core/breadth.py
No snapshots: app/core/breadth.py persists nothing — it prints — so the verdict above is the whole record. What would attach a trace: a --write path for this study, which is a worker change and a decision (R11).
market_divergenceshouseCANNOT_TELLunder test · measured 2026-09-21 · not fitted: no parameter here was chosen by optimising an outcome
claim
When price and breadth, credit, correlation, implied vol or dealer gamma disagree sharply, the disagreement resolves - one side gives (DESIGN M).
finding
Of five pairs, three have produced events at all and exactly one cell clears the 20-date floor: price_vs_correlation with the index extended, which was followed by -0.50% over 5 sessions [-1.48%, +0.37%] and +0.14% over 21 [-1.28%, +1.53%] - both spanning zero. Every other cell reports insufficient (2 to 16 event dates), and implied-vs-realised and price-vs-gamma have never reached the threshold since both their legs had enough history to be percentiled.
sample
21 event dates in the only measurable cell; 2-16 in the rest
could see
about +-0.9% at 5 sessions and +-1.4% at 21, in the one cell that ran
method
Both legs as expanding point-in-time percentiles, the gap in percentile points, an event on ARRIVAL at a 60-point gap, the two tails graded apart, then event_study.base_rate with its 20-date floor and date-clustered interval.
fitted
not fitted: no parameter here was chosen by optimising an outcome
revive if
A pair accumulates 20+ event dates in a tail. The board is honest today precisely because it mostly reads insufficient: a version that drew the gaps without the base rates beside them would look far more impressive and say considerably less.
code
app/core/divergence.py
No snapshots: app/core/divergence.py persists nothing — it prints — so the verdict above is the whole record. What would attach a trace: a --write path for this study, which is a worker change and a decision (R11).
dealer_gamma_dampinghouseCONSISTENT_WITHunder test · measured 2026-09-16 · not fitted: no parameter here was chosen by optimising an outcome
claim
High dealer net gamma precedes a smaller next-session move (DESIGN L).
finding
Controlled for price level and trailing vol: -0.149 sigma [-0.271, -0.049] at lag 0, essentially unchanged from -0.146 at lag 1. The price placebo now spans zero; the trailing-vol placebo is larger than the effect, which is why the controlled line is the answer.
sample
1,599 observations on 41 dates, 39 tickers - about two months, one regime
could see
about +-0.11 sigma
method
Within-ticker expanding median, moves in own sigmas, date-clustered CI, stratified by (price high?, vol high?), with placebos.
fitted
not fitted: no parameter here was chosen by optimising an outcome
revive if
It holds over a second regime. Two months of one regime is not a base rate, and the honest label stays 'consistent with' until it is.
code
app/core/gex_study.py
No snapshots: app/core/gex_study.py persists nothing — it prints — so the verdict above is the whole record. What would attach a trace: a --write path for this study, which is a worker change and a decision (R11).
resolution_base_rateshouseCANNOT_TELLunder test · measured 2026-09-14 · not fitted: no parameter here was chosen by optimising an outcome
claim
A backwardation flip or a 90th-percentile skew steepening is followed by something different from an ordinary day (DESIGN H).
finding
Every cell either spans zero or reports insufficient, and every regime-conditioned cell is insufficient - splitting a thin sample four ways is how a base rate lies.
sample
Fewer than 20 distinct event dates in most cells
could see
not reported per cell at the time; the cells that ran were far wider than any plausible effect
method
Event study with a same-names same-dates baseline, look-ahead gate, date-clustered bootstrap, 20-date floor.
fitted
not fitted: no parameter here was chosen by optimising an outcome
revive if
The event classes accumulate 20+ distinct dates per regime cell.
code
app/core/event_study.py
No snapshots: app/core/event_study.py persists nothing — it prints — so the verdict above is the whole record. What would attach a trace: a --write path for this study, which is a worker change and a decision (R11).
worker/app/core/ledger.py · core.calibration_snapshots · grade_report --write, last 2026-09-24 (manual)
could see
printed floors +-0.63% at 5 sessions, +-0.89% at 10, +-1.29% at 21 at an uncorrected 95%; the family-wise intervals actually drawn had half-widths of 0.53%, 0.83% and 1.26%, so a 1% move was visible at 5 and 10 sessions and not at 21
method
Point-in-time expanding percentile, bottom-decile rule, episodes by entries(), re-dated by tradeable_on(), event_study.base_rate against a same-dates baseline, nine intervals Bonferroni-corrected (breadth_report.py).
fitted
not fitted: no parameter here was chosen by optimising an outcome
revive if
A test of the tail rather than the mean - compressed dispersion could precede a wider spread of outcomes without a worse average - or a second regime in which the 5- and 10-session intervals move below zero. The 21-session cell needs 103 episodes, about 2033.
code
app/core/breadth.py
earnings_implied_vs_realisedhouseNO_EVIDENCEretired · tested — no evidence at the size claimed · measured 2026-09-20
claim
The options market overprices single-name earnings moves.
finding
49% of reports came in under the implied move [45%, 54%]; median realised / implied 1.01. No evidence it is rich, and none that it is cheap.
sample
460 reactions on 312 dates, 25 names, 2019-10 to 2026-09
could see
about +-4.5 percentage points on the share
method
Vendor expected_move_perc against our own close-to-close move, session convention by report_time, date-clustered.
fitted
not fitted: no parameter here was chosen by optimising an outcome
revive if
The live series - reports whose implied figure we captured BEFORE the event - reaches the 20-date floor, expected in 2027. Every reaction in this sample is the vendor's post-hoc figure.
code
app/core/earnings_study.py
component_cross_sectional_ichouseNO_EVIDENCEretired · tested — no evidence at the size claimed · measured 2026-09-20
claim
A registered metric ranks the Core names in a way the next month agrees with.
finding
Zero of 51 component-horizons show an IC this test could call real. The largest honest readings are px_vs_200dma +0.052 and realized_vol_21d +0.067 at 21 days, both inside their own detection thresholds.
sample
Up to 2,524 sessions x 35-39 names for the price metrics; 30-64 sessions for the chain and risk-reversal families
could see
0.018 to 0.29 depending on the component's history
method
Spearman cross-sectional IC, Newey-West SE at lag = horizon - 1, ranked on the first close the value could have been ranked on, leak control injected.
fitted
not fitted: no parameter here was chosen by optimising an outcome
revive if
A component clears its own detectable threshold without being flagged as a leak. Seven young series read past 0.10 on samples too thin to tell; they are worth re-running in 2027, not believing now.
code
app/core/ic.py
calendar_seasonalityhouseNO_EVIDENCEretired · tested — no evidence at the size claimed · measured 2026-09-19
claim
Some calendar buckets - weekdays, months, turn of month, opex week - carry returns different from the rest of the calendar.
finding
Zero of nineteen buckets exclude zero at a family-wise 95%. The largest are November +0.112 sigma [-0.037, +0.263] and Monday +0.060 [-0.049, +0.170]; unadjusted, both would have read as findings.
sample
2,524 sessions, 2016-09 to 2026-09-18, 39 names
could see
about +-0.15 sigma per bucket after the nineteen-test adjustment
method
Cross-sectional mean return in each name's own sigmas, scaled by the PRIOR session's realised vol; date-clustered bootstrap; Bonferroni over 19 tests.
fitted
not fitted: no parameter here was chosen by optimising an outcome
revive if
A bucket survives the family-wise interval on a sample that includes a bear market. The whole window is one bull decade.
code
app/core/seasonality.py
cut before building 4
rarity_ranked_confluencehouseCUTretired · cut before building · measured never built
claim
Rare combinations of boolean signals identify unusual setups.
finding
Cut before building (DESIGN.md §3.5). ~80 signals over ~139 symbols and two years is ~70k symbol-days, so any 4-signal combination has n = 0-6 by combinatorics alone. 'This has happened 6 times' is an artefact of the state space wearing a confident label - the new bias digit.
sample
not run: the arithmetic decides it in advance
could see
nothing, at any sample size reachable here
method
none - refused on combinatorial grounds
fitted
not fitted: no parameter here was chosen by optimising an outcome
revive if
Never on this universe. The divergence board and the unusualness ranking deliver the honest half of what it promised.
code
inherited
category_accuracy_30dhouseCUTretired · cut before building · measured never built
claim
A per-category accuracy score summarises how the system is doing.
finding
Cut (DESIGN.md §3.5). Its denominator included every neutral call, so with 11 of 14 analyzers abstaining on a typical day the ceiling was ~21% and every category read below the 'red flag' threshold permanently. The system's only data-driven input was a constant.
sample
not run
could see
nothing: the metric could not reach its own threshold
method
none - refused on definitional grounds
fitted
not fitted: no parameter here was chosen by optimising an outcome
revive if
Never in this form. A denominator that includes abstentions is not an accuracy.
code
inherited
reconstructed_advance_declinehouseCUTretired · cut before building · measured never built
claim
An advance/decline line, net new highs minus lows, and a McClellan oscillator can be reconstructed from the tracked universe to give the breadth measures DESIGN.md:185 (§3.3 S) asks for.
finding
Cut before building, and DESIGN cuts it in the same sentence that asks for it: 'consume published breadth indices where they exist rather than reconstructing them from constituents - the reconstruction is the worst novelty-to-effort ratio in the catalogue.' The objection is not effort, it is that the result is a DIFFERENT STATISTIC wearing the same name. A/D is an exchange-wide count of issues; computed over 39 tickers its level is a fact about the universe's size, and it would move when a collector was added. NH-NL is worse: a universe assembled in 2026 holds few names with the history to print a 52-week extreme in it, so the series would read as persistently healthy for a reason that has nothing to do with the market. McClellan inherits both because it is computed from A/D.
sample
not run: the construction decides it in advance
could see
nothing - a reconstructed series has no true value to be compared with
method
none - refused on construction grounds (app/core/breadth.py REFUSED)
fitted
not fitted: no parameter here was chosen by optimising an outcome
revive if
A PUBLISHED A/D or NH-NL series is collected. Unusual Whales market-tide and sector-tide are specified in DATA_SOURCES.md:128 and budgeted at 13 calls/day; no collector exists yet. Consuming one of those is the revival, not rebuilding the count.
breadth_decile_thresholdhouseCUTretired · cut before building · measured never built
claim
Narrow breadth can be defined as sector_breadth_200dma entering its own bottom decile, the way the other percentile markers in this repo work.
finding
Cut on inspection, and it was one line from shipping. sector_breadth_200dma is the share of EIGHT sector funds above their 200-day average, so it takes nine values and no others. A decile threshold cannot land at 10% on a nine-valued series: the rule silently becomes 'at or below whichever eighth sits nearest the boundary', whose frequency can be 4% of sessions in one year and 30% in the next with no change in the market. It would have been reading the quantisation. The built axis uses regime.NARROW_BREADTH - three sectors of eight or fewer, a threshold already declared and cited - and breadth.percentile_is_meaningful() refuses a percentile rule on any series too coarse to carry one.
sample
not run: nine distinct values decides it in advance
could see
nothing - the threshold's meaning would vary with the sample
method
none - refused on quantisation grounds (app/core/breadth.py)
fitted
not fitted: no parameter here was chosen by optimising an outcome
revive if
Breadth is computed over a wide enough cross-section to be near continuous - which needs DESIGN.md:433 Owner Decision #1 (the universe) resolved, or a published index. On eight funds, never.
MACD crossovers carry cross-sectional information (inherited from v1).
finding
Dropped at IC -0.026. DESIGN.md:201 calls this the best-documented decision in the predecessor codebase precisely because someone wrote down why - and this ledger exists so that stops being remarkable.
sample
v1's cross-sectional pipeline; the exact window is not recorded here because it was not recorded there
could see
unknown - v1 reported no interval, which is part of why it is here
method
Cross-sectional IC, v1's app/cross_sectional/.
fitted
not fitted: no parameter here was chosen by optimising an outcome
revive if
Somebody re-proposes it. The obituary is the answer: it was measured, it was negative, and the measurement is older than this warehouse.
code
inherited
worker/app/core/ledger.py
not filed in the ledger — which is different from having passed (ledger.py:361-364).
the call: each fund's quadrant and whether its short window disagrees with it
to grade: turning is a categorical call per fund per session
not filed in the ledger — which is different from having passed (ledger.py:361-364).
page constant DATA_READY · lib/signals.ts
2026-10-15
cannot be evaluated: detectable not stored
max_pain:converges [universe:core] @5d
1
2026-09-24 / 2026-09-24
2258
unknown
1 measurements: the baseline needs 5 before a shift can be told from the noise it started with
one measurement: growth cannot be judged
2026-10-15
cannot be evaluated: detectable not stored
uw_flow_alert:call_ask_side [universe:core] @1d
1
2026-09-24 / 2026-09-24
891
unknown
1 measurements: the baseline needs 5 before a shift can be told from the noise it started with
one measurement: growth cannot be judged
2026-10-15
cannot be evaluated: detectable not stored
uw_flow_alert:call_ask_side [universe:core] @5d
1
2026-09-24 / 2026-09-24
748
unknown
1 measurements: the baseline needs 5 before a shift can be told from the noise it started with
one measurement: growth cannot be judged
2026-10-15
cannot be evaluated: detectable not stored
uw_flow_alert:put_ask_side [universe:core] @1d
1
2026-09-24 / 2026-09-24
880
unknown
1 measurements: the baseline needs 5 before a shift can be told from the noise it started with
one measurement: growth cannot be judged
2026-10-15
cannot be evaluated: detectable not stored
uw_flow_alert:put_ask_side [universe:core] @5d
1
2026-09-24 / 2026-09-24
738
unknown
1 measurements: the baseline needs 5 before a shift can be told from the noise it started with
one measurement: growth cannot be judged
2026-10-15
cannot be evaluated: detectable not stored
app/core/drift.py · drift_report.py · assessed at render from the trace
+1.4 ptsnull 49.0%
Each difference is paired over the calls that null could judge (grading.py:367-373), and that count is not stored: the bar is the cell's hit rate over all its calls, so bar minus rule need not equal the printed points, and the cell's n is not any one null's n.
driftunknown1 measurements: the baseline needs 5 before a shift can be told from the noise it started with
stale from 2026-10-15 (21 days after the write, to the hour) unless the grader writes again · one measurement: growth cannot be judged
+1.2 ptsnull 48.7%
Each difference is paired over the calls that null could judge (grading.py:367-373), and that count is not stored: the bar is the cell's hit rate over all its calls, so bar minus rule need not equal the printed points, and the cell's n is not any one null's n.
driftunknown1 measurements: the baseline needs 5 before a shift can be told from the noise it started with
stale from 2026-10-15 (21 days after the write, to the hour) unless the grader writes again · one measurement: growth cannot be judged
-1.3 ptsnull 50.3%
Each difference is paired over the calls that null could judge (grading.py:367-373), and that count is not stored: the bar is the cell's hit rate over all its calls, so bar minus rule need not equal the printed points, and the cell's n is not any one null's n.
driftunknown1 measurements: the baseline needs 5 before a shift can be told from the noise it started with
stale from 2026-10-15 (21 days after the write, to the hour) unless the grader writes again · one measurement: growth cannot be judged
+5.1 pts
null 49.4%
trend_following+5.5 ptsnull 49.4%
Each difference is paired over the calls that null could judge (grading.py:367-373), and that count is not stored: the bar is the cell's hit rate over all its calls, so bar minus rule need not equal the printed points, and the cell's n is not any one null's n.
driftunknown1 measurements: the baseline needs 5 before a shift can be told from the noise it started with
stale from 2026-10-15 (21 days after the write, to the hour) unless the grader writes again · one measurement: growth cannot be judged
+3.9 pts
null 50.8%
trend_following+5.4 ptsnull 49.4%
Each difference is paired over the calls that null could judge (grading.py:367-373), and that count is not stored: the bar is the cell's hit rate over all its calls, so bar minus rule need not equal the printed points, and the cell's n is not any one null's n.
driftunknown1 measurements: the baseline needs 5 before a shift can be told from the noise it started with
stale from 2026-10-15 (21 days after the write, to the hour) unless the grader writes again · one measurement: growth cannot be judged
+6.0 pts
null 48.9%
trend_following+4.5 ptsnull 50.3%
Each difference is paired over the calls that null could judge (grading.py:367-373), and that count is not stored: the bar is the cell's hit rate over all its calls, so bar minus rule need not equal the printed points, and the cell's n is not any one null's n.
driftunknown1 measurements: the baseline needs 5 before a shift can be told from the noise it started with
stale from 2026-10-15 (21 days after the write, to the hour) unless the grader writes again · one measurement: growth cannot be judged