Methodology
Status: pre-publication. The definitions below are the constructs the platform is building toward. No metric has passed its evidence gates yet; when one does, this page gains the exact model identity, validation results, and known limitations for that release. A definition here is a promise of construction, not a claim of validity.
▸ CORE — Contextual On-Field Rating Estimate
A retrospective, context-adjusted estimate of a player's contribution to scoring value per standardized opportunity, expressed in expected points above average per 100 qualifying snaps. CORE is model-conditional: it estimates contribution after accounting for teammates, opponents, situation, coaching, package, and role — it is not a claim about metaphysical football ability, career value, or the future.
CORE combines three evidence channels: context-adjusted outcome evidence (collinear but broad), observable action credit (narrow but direct), and role-specific statistical priors (regularizing, never trained on proprietary all-in-one metrics). Team, unit, opponent, coach, and package effects are modeled explicitly so a player coefficient cannot silently absorb system quality.
▸ FORGE — Football Overall Replacement-adjusted Game Equity
Cumulative value above a role-specific replacement baseline, in expected points first and wins only after an empirically calibrated points-to-wins conversion with published error bars. Replacement cohorts are defined result-blind from labor-market evidence (roster-fringe, practice-squad promotions, backup stratification), and every team-season publishes a reconciliation ledger — assigned player value, coaching/scheme, interactions, and a visible unassigned residual. A visible residual is more honest than false precision.
▸ PULSE — Performance Under Latest Sample Evidence
A rolling recent-form descriptor with its exact evidence window always visible, shown beside full-season CORE, never in place of it. Window and half-life parameters are chosen on historical out-of-time tasks, not on making a preferred player look hot. PULSE is not automatically a forecast, and it never generates causal narratives from a trend.
▸ SIGNAL — the evidence profile
Six axes, each 0–100 with its raw evidence attached: Sample (enough relevant opportunity?), Identifiability (can this player be separated from stable teammates?), Granularity (how directly does data observe the claimed role?), Noise (how unstable is the estimate?), Availability (is the period complete?), Lineage (is every input reproducible?). The overall band is conservative — a huge sample can never conceal missing granularity, and a pristine pipeline cannot create missing assignment data. SIGNAL is never a percentage chance that a rank is "correct."
▸ Evidence tiers by position
| Tier | Typical roles | Visibility |
|---|---|---|
| A — directly observed | QB, kicker, punter, returner | Strong event and opportunity evidence on nearly every snap |
| B — partially observed | RB, WR, TE, edge, targeted DB | Participation plus direct events; non-target/non-pressure snaps less visible |
| C — presence-dominant | OL, interior DL, off-ball LB, non-target safety | Presence and unit outcomes; assignment unobserved — wider intervals required |
| D — assignment-blocked | Roles without reliable participation evidence | Counting facts and team context only; estimates render unavailable, never guessed |
▸ The trench program
Evidence tiers gate publication, not research. For linemen, off-ball linebackers, and rarely-targeted coverage defenders, Gridiron Signal builds complete individual and unit pipelines with three to five competing model families — participation residuals, observable-event allocation, unit decomposition with shrinkage, role-prior hierarchical models, and (privately) charting comparators. How much the families agree becomes part of SIGNAL. Estimates that remain unstable are published as research with wide intervals and identifiability scores — the instability itself is a finding.
▸ Special teams
Kicking points above expected (distance and environment adjusted), punting field-position value, return value above expectation, and unit-level coverage value — the cleanest attribution problems in football, and an early proving ground for the platform's honesty machinery.
▸ The claim ladder
Every metric occupies exactly one status:
concept (named question, no values) →
scaffold (contracts, fixtures, and gates exist) →
research (real data, research lab only) →
provisional (public with required disclosures) →
validated (passed construct and out-of-time gates) —
or retired / blocked, both published with reasons.
No metric goes live because the site needs another card.