The Strategic Preparedness Index (SPI)
This page documents, in full, how the headline score on this site is actually computed, sourced, and gated — and it is deliberately written to be checked. Nothing below is marketing copy: every claim matches the code that runs in production, and every limitation is disclosed rather than hidden. If a sentence here couldn't survive an expert clicking through to verify it, it doesn't belong on this page.
1. What the Index measures
The Strategic Preparedness Index is a composite 0–100 score across eight interdependent pillars of Australian strategic preparedness, drawn from the argument of Unprepared: Australia in an Age of Chinese Power. Each pillar tracks real, sourced variables against the book's programme targets.
Australia captures ~35% of mining operating profits versus Norway's 70-80%. No sovereign wealth fund exists. LNG pays near-zero PRRT. The gap is A$50-80B/year.
Defence spending at 2.05% GDP against a 5% target. Gross Commonwealth debt at $993B growing at ~8%/year. NDIS at ~$49B/year versus its original design of $22-25B.
Australia imports ~80% of major defence equipment. Manufacturing is 5.1% of GDP — lowest in OECD. GWEO targets 4,000 GMLRS rounds/year by 2029. Ghost Shark and Ghost Bat are on track.
Fertility rate 1.48 — lowest in a century. No national service programme. 73% of population growth from migration, not births. Military-age cohort too small to sustain 140,000 ADF.
China takes 29% of Australian exports — deep dependence. But the iron-ore asymmetry runs the OTHER way: at 24.5:1 China cannot replace ~60% of its ore (a ~A$150B hit vs ~A$6.2B to Australia), which is why it never sanctioned it. A broad trade war, though, cuts against Australia (~6% of GDP vs ~0.5% for China).
61,189 ADF personnel against 140,000 target. ~2-3 operational submarines against a 230+ platform target. Zero coastal missile batteries. 34 days diesel reserves.
AUKUS faces severe production risk: US Virginia class running at 1.13 boats/year against 2.33 required. Five Eyes integration strong. But an indispensable alliance is not a sufficient strategy.
Manus Island remains a patrol boat facility. Two Pacific Island nations have signed Chinese security agreements. Indonesia underdeveloped. Pacific Compact not established.
2. How the score is computed
Every score on this platform flows through one canonical formula — there is a single source of truth in the codebase, not a separate calculation per page. Here is exactly what it does:
Step 1 — variable progress (0–100)
Each tracked variable has a fixed baseline (its value when tracking began) and a programme target. Progress is the distance travelled from baseline toward target, clamped to 0–100:
- Higher-is-better variables (e.g. defence spending % GDP): 0% at baseline, 100% at target.
- Lower-is-better variables (e.g. gross debt, China export share): 0% at baseline, 100% at target — symmetric with the above, not a 50% midpoint. A metric sitting exactly at its baseline reads 0% (“no progress yet”), never a misleading halfway score.
The baseline is fixed at the point tracking began and never moves — this matters because an earlier version of the formula anchored to the live value instead, which silently ratcheted scores toward zero as ordinary reporting noise drifted the number. That bug was fixed; the fixed baseline is what makes the score honest rather than a moving target.
Step 2 — pillar score
Each pillar score is the weighted average of its variables' progress. Every variable gets full weight (1.0) exceptthose flagged “estimate”-confidence (author-assessed rather than sourced from an official dataset), which are down-weighted to 80% — so a shakier figure can still move a pillar, just less than a hard number.
Step 3 — overall score + the pessimism nudge
The overall score is a weighted average of all eight pillar scores. The weights (foundational pillars count more):
Weights sum to 8.9; the average is normalised by that total.
That weighted average is the calculated score — it is always derived from live variable data, never manually set. An administrator can apply one further, deliberately limited adjustment: a score nudge that can only move the published score downward(it is clamped to ≤0) and never up. It exists as an explicit, disclosed pessimism correction for known gaps in what's currently modelled — not a lever to make the headline number look better. The overall score is floored at 1 (it never displays 0, to avoid implying total, literal collapse).
3. FACT vs PRESCRIPTION
Every number on this platform splits cleanly into two categories, and conflating them would be the single easiest way to mislead a reader:
FACT
The measured, present-day state of the world — ABS statistics, Defence budget papers, Pentagon reporting, DFAT trade data, GDELT press monitoring. These figures are refreshed on a stated cadence and change as reality changes.
PRESCRIPTION
The book's targets — what the author argues Australia should build. These are fixed to the book's argument and do not move with the news. “The Programme” throughout this site refers to this prescription, not government policy.
Progress is the gap between the two: how far the measured FACT has moved toward the fixed PRESCRIPTION. Where a target figure is a synthesised, book-consistent estimate rather than an explicit number from the text, the variable is flagged “estimate” confidence (see above) and down-weighted accordingly.
4. Data sources & refresh
Every tracked variable carries a named source, an update frequency, and a confidence tier (high / medium / low / estimate) — all visible on the pillar and aboutpages. Core sources include the ABS, Defence Budget Papers, ANAO Major Projects Reports, the Pentagon's Annual Report on Chinese Military Power, DFAT trade statistics, SIPRI and IISS. Live press-coverage and conflict-event monitoring via GDELT (BigQuery) is described in full on the Signals page. Every change made to a live variable — approvals, corrections, source updates — is logged on the Updates page.
5. Signal fusion — Strategic Tempo
The Signalspage shows each GDELT cluster's own intensity — recent volume against its own historical median — one card at a time. Strategic Tempois a single gauge that blends those already-published measures into one number, so a reader doesn't have to scan every card to get a sense of the overall picture. It is a composite of existing descriptive data, not a new data source and not a forecast.
The exact recipe
- For every enabled cluster, compute its own ratio: mean(last 3 days) ÷ median(last ~90 days) of its daily volume — the identical ratio already shown per card on the Signals page.
- Average the ratios within two categories: coverage clusters (press attention) and events clusters (CAMEO conflict-event counts).
- Blend the two category averages with equal weight — 50% coverage / 50% events — into one combined ratio. If an entire category has no cluster with enough history, the gauge falls back to the other category alone rather than assuming the missing one is at a neutral “normal” ratio.
- Map the combined ratio onto a 0–100 gauge: a ratio of 1 (both categories exactly at their own historical normal) reads 50; a ratio of 2 or more (double normal) caps the gauge at 100; a ratio of 0 floors it at 0.
Because it is built entirely from the same descriptive GDELT measures documented above and on Signals, it carries the identical limitation: a rise in Strategic Tempo means press and event activity are more intense than their own recent normal — it is not a probability, and it predicts nothing.
6. Human-in-the-loop — the trust core
No score moves without a human approving it. Ever.
Claude reads incoming news (RSS, Reddit, Google News) and GDELT signal data, classifies relevance, and proposes changes into an admin approval queue. It never writes directly to a live variable, pillar, or overall score. A human operator reviews every proposal and explicitly approves, edits, or rejects it before anything public changes. This is the single hardest rail in the whole system, and it applies uniformly — including to the private forecast engine described below, which cannot touch a public score under any circumstance.
7. The forecast engine — an honest account
This is private, unpublished, and not proven to work.
Separately from the public score, the platform runs a private paper-trading forecast engine: it logs dated, falsifiable predictions — a calibrated 90-day probability for each tracked scenario, and a 14-day directional call on GDELT coverage volume per monitored cluster — into an append-only table with no public read access. Nothing it produces is displayed on this site, and it makes no claim of predictive accuracy.
Why paper-trading instead of a backtest: an earlier historical backtest of the coverage-direction approach ran on a held-out sample small enough (n≈2) that its result — discouraging — couldn't be trusted either way. Rather than lean on that, the engine now makes real, dated calls prospectively and grades itself only once each call's horizon actually passes. That is a slower, more honest test, and it is still running.
Self-scoring
Coverage-direction calls resolve automatically from later GDELT data (did the volume actually go up or down); scenario probabilities resolve by human judgement at the 90-day horizon and are graded with a Brier score (a standard proper scoring rule for probabilistic forecasts — lower is better).
Ensemble forecasting — independent personas, aggregated
Each scenario probability is produced by four independent analyst personas — a base-rate statistician (anchors on historical frequency), an escalation hawk (weights current tension signals), a status-quo dove (expects regression to the mean), and an economic-structural realist (weights trade and fiscal constraints) — each answering the identical question independently, on a cheap model. The four numbers are combined with a trimmed mean(the highest and lowest replies are dropped before averaging, so one outlier persona can't dominate), and the spreadbetween the personas — how much they disagreed — is recorded alongside the result as an honest signal of how uncertain the call really is. Every forecast row's provenance records which model actually produced it, the aggregation method, the spread, and each persona's individual number and reasoning, so nothing here is a black box. If the ensemble can't be formed (an API outage, a parsing failure), the engine falls back to a single call rather than skip the week's forecast — and records honestly which path produced the row.
Two self-improvement mechanisms, both gated
Statistical calibration:a conservative correction derived from the engine's own resolved track record. Below 20 resolved scenario calls it is the identity — it changes nothing. Above that gate, it applies a heavily shrunk (weight 0.25) nudge toward the engine's observed historical accuracy, and every forecast row records whether the correction was active and by how much.
Reasoned lessons: after each forecast resolves, a separate process asks Claude to write a short post-mortem — a verdict (hit / miss / partial) plus a transferable lesson about why — and the accumulated, human-prunable lessons are injected into future forecast prompts. An operator can deactivate any lesson that turns out to be spurious. This complements calibration: calibration adjusts confidence, lessons adjust reasoning.
Before any forecast track record could even be considered for disclosure, it would need to clear stated minimum sample sizes on both call types — and even then, only ever as a labelled record with its sample size attached, never as an accuracy claim. There is no schedule for if or when that happens. The first scenario-probability forecasts resolve in mid-September 2026; until a meaningful number of calls have resolved, treat this entire section as a description of a research instrument, not a working forecaster.
8. Confidence & limitations
- Scenario probabilities editorially shown elsewhere on the site are qualitative estimates, not statistically derived forecasts.
- Several Pillar 5 and Pillar 8 target figures are book-consistent syntheses rather than literal numbers from the text, and are flagged “estimate” confidence accordingly.
- Chinese military data is inherently uncertain; PLA figures use Pentagon estimates as a baseline.
- The fifty-year trajectory model uses simplified, illustrative assumptions — it is a model, not a prediction.
- Sample sizes in the private forecast engine are, and will remain for some time, small. Small-sample statistics are noisy; the calibration mechanism above is deliberately gated and shrunk specifically to avoid overfitting to that noise.
- This platform advances a partisan argument as well as an analytical one. The programme targets reflect the author's judgement about what Australia should do — they are a prescription, not a consensus forecast.
9. Why this is different
Most public arguments about national strategy are static: a book, a report, a PDF, published once and left to age. This platform is a measured instrument instead — a score computed from a fixed, disclosed formula over sourced, dated data; refreshed on a stated cadence; gated so no automated system can move a public number without a human approving it; and honest enough to publish its own weak points, including an experimental forecasting module that plainly hasn't proven itself yet. That is the whole claim: not that the system is infallible, but that every part of it can be checked.