It shows the data, the interface, and how forecasts get validated, using public ABS data and a simple damped-trend forecaster as a baseline. The full model does more: a cohort-component core, reconciled to Victoria in Future, forecasting government-school enrolments against capacity. The gaps are listed below. The baseline is here so the production model’s accuracy can be compared against it rather than just claimed.
| Aspect | This public-data demonstrator | Full production model |
|---|---|---|
| Forecast method | Damped log-linear trend per SA2, a transparent baseline | Hamilton–Perry cohort-component core (ages real cohorts forward), reconciled to Victoria in Future, with a capture-rate sub-model and the housing-unit method for greenfield |
| Data | Public ABS ERP (school-age 5–14) only | The above plus school enrolment and catchment data, births, kindergarten, and PSP/UDP land supply |
| Unit & geography | SA2 (462 statewide), one aggregate 5–14 cohort | SA1 → catchment → network → state, MinT-reconciled; single-year ages / grade cohorts; Prep intake; specialist-school & kindergarten disaggregation |
| What it forecasts | Population demand, with back-tested bands | Government-school enrolment demand vs capacity → shortfall / surplus by area and year |
| Validation | Rolling-origin back-test of the trend model, to a 10-year horizon | Standing back-test every cycle to full horizon, with assumption-invalidation alerts |
| Horizon | Illustrated to 2031 (10 yr) | 20-year statutory horizon, with bands that widen with distance |
| Decision layer | None (demand only) | Capacity overlay, demand-vs-capacity, zones and travel-time; the build/hold decision itself |
The demonstrator’s damped-trend forecaster is back-tested by rolling origin on the real ABS series (2001–2021). Median absolute percentage error grows with horizon and is far higher in growth corridors than established areas:
| Growth class | 1 yr | 3 yr | 5 yr | 10 yr |
|---|---|---|---|---|
| High growth | 3.1% | 8.7% | 14.5% | 28.2% |
| Moderate growth | 1.7% | 4.7% | 7.8% | 14.1% |
| Stable / declining | 1.5% | 3.9% | 5.7% | 8.8% |
The trend systematically under-forecasts high-growth SA2s (median bias about −26% at 10 years): it cannot see greenfield acceleration. That is the evidence-based case for the proposed cohort-component and housing-unit methods — the full model’s job, not demonstrated here. The back-test is reproducible from public data: school-demand/analysis/backtest.py.
Detailed methodology for this demonstrator. Everything below is either demonstrated on real public ABS data (see the pilot, cohort and leading-indicator appendices) or is the full model's design, cited to the literature. References at the end; only primary-source-verified claims are cited.
The planning question is not "what will the population be", it is "where and when do we build or expand schools, and where do we hold off?" The forecast is instrumental to that decision:
forecast demand (by small area, cohort, year) → compare to current + planned capacity → shortfall / surplus by area and year → infrastructure & enrolment-management response (new school pipeline, expansions, modular classrooms, zone changes).
So the model is judged not by elegance but by whether it makes the build decision right, and by how much it reduces the cost of being wrong. That framing drives every choice below.
| Component | What it does | Status / evidence on public data |
|---|---|---|
| Cohort-component core (Hamilton–Perry cohort-change ratios) | Ages each cohort forward from people already observed (5–9 in 5 yrs = today's 0–4) rather than extrapolating a line | Design (full model); to be validated against the trend once enrolment data by age is available |
| Reconciliation to Victoria in Future | Constrains small-area forecasts to the state's official control totals | Design (VIF is public; wired in for the full model) |
| MinT hierarchical reconciliation | Makes SA1 → zone → region → state forecasts coherent (sum correctly) | Design; method verified (Wickramasuriya et al. 2019) |
| Capture-rate sub-model | Government share of enrolments by area & level, the biggest swing factor | Design; calibrated on school enrolment data |
| Housing-unit method (dwelling approvals × typed yield) | Forward demand in greenfield the history cannot see | Design; needs building-approvals data to calibrate typed student-yield rates |
| Damped trend | Transparent baseline & benchmark | Built: the demonstrator tool's forecaster |
| Evidence-gated experimental track | Benchmarks ML / reconciliation / foundation models; promotes only what beats the core on back-test | Design (Grossman & Wilson 2022) |
| Uncertainty quantification | Rolling-origin back-test → empirical error bands by area class × horizon | Built: bands drive the tool (§4) |
The core is deliberately a simplified cohort-component method: the small-area literature is consistent that simple, structurally-grounded methods constrained to reliable totals are hard to beat (Wilson 2015; Wilson et al. 2022), and my own back-test confirms it here.
Note on the demonstrator's age range. The public ABS ERP series used in the demonstrator publishes SA2 population only in 5-year bands, so it uses ages 5–14 (the 5–9 and 10–14 bands) as a clean school-age proxy and excludes the 15–19 band, which mixes senior-secondary students with post-school 18–19-year-olds. The production model works from school enrolment data by single year of age across Prep to Year 12.
The demonstrator’s damped-trend forecaster is back-tested by rolling origin on the real ABS ERP series (2001–2021); the code is public: school-demand/analysis/backtest.py.
Not demonstrated here (left for the full model): the cohort-component-vs-trend comparison and the dwelling-approvals leading indicator. Validating them needs school enrolment data (or ABS age-band and building-approvals series), so they are labelled design, not claimed as results on this public demonstrator. Being explicit about what is back-tested versus proposed is itself how the assurance gap the Auditor-General identified (VAGO 2017) is closed.
The back-tests locate it precisely:
The critical asymmetry: the highest-uncertainty areas, the outer growth corridors, are exactly where the biggest, most expensive, least-reversible infrastructure decisions are made (new-school pipeline). The widest error bands fall where the largest capital decisions are made, so those areas get interval forecasts and the closest monitoring.
Concrete, and mostly already evidenced above:
Growth-corridor forecasts carry the most error. I report that error, use approvals as an early signal, update those areas most often, and recommend staged or relocatable capacity where the band is wide. That is what makes the forecast safe to build against.
Primary-source-verified (Tier A) and well-established (Tier B) from the literature review; full list in the literature review (Appendix E).
Note on tiers: every claim cited above is verified against a primary source or is a standard, well- established reference. Claims that could not be verified against a primary source are excluded.