Victoria: School-Age (5–14) Demand Explorer · demonstrator

Real ABS Estimated Resident Population (dataflow ERP_ASGS2016) for all 462 Victorian SA2s. Forecast: damped-trend projection with back-tested (rolling-origin) median-error bands. Choose an area on the map, from the search box, or from the data table. Public-data demonstrator of a school provision-planning model.
Update, October 2026: a newer version of this analysis, with a cohort-survival back-test across 505 areas, is at Deeper Than Data.
Colour by
Scroll to zoom · drag to pan · use the + / − / ↻ buttons

Select an area

Pick a region on the map, from the “Find an area” box, or a row in the data table to see its school-age population history and 10-year forecast with an uncertainty band.

Methodology, validation & scope

This is a public-data demonstrator, not the full model

It shows the data, the interface, and how forecasts get validated, using public ABS data and a simple damped-trend forecaster as a baseline. The full model does more: a cohort-component core, reconciled to Victoria in Future, forecasting government-school enrolments against capacity. The gaps are listed below. The baseline is here so the production model’s accuracy can be compared against it rather than just claimed.

Demonstrator vs the production model

AspectThis public-data demonstratorFull production model
Forecast methodDamped log-linear trend per SA2, a transparent baselineHamilton–Perry cohort-component core (ages real cohorts forward), reconciled to Victoria in Future, with a capture-rate sub-model and the housing-unit method for greenfield
DataPublic ABS ERP (school-age 5–14) onlyThe above plus school enrolment and catchment data, births, kindergarten, and PSP/UDP land supply
Unit & geographySA2 (462 statewide), one aggregate 5–14 cohortSA1 → catchment → network → state, MinT-reconciled; single-year ages / grade cohorts; Prep intake; specialist-school & kindergarten disaggregation
What it forecastsPopulation demand, with back-tested bandsGovernment-school enrolment demand vs capacity → shortfall / surplus by area and year
ValidationRolling-origin back-test of the trend model, to a 10-year horizonStanding back-test every cycle to full horizon, with assumption-invalidation alerts
HorizonIllustrated to 2031 (10 yr)20-year statutory horizon, with bands that widen with distance
Decision layerNone (demand only)Capacity overlay, demand-vs-capacity, zones and travel-time; the build/hold decision itself

Known gaps in this demonstrator

  • It forecasts population, not government-school enrolments. The capture-rate step (the government share of enrolments, and the largest source of variation) needs school enrolment data and is not in this demonstrator.
  • It uses a trend method, not the cohort-component core, so it is weakest in fast-growing greenfield areas: the back-test shows the trend under-forecasts high-growth SA2s by a median of about 26% at a 10-year horizon. Closing that bias is the job of the proposed cohort-component and housing-unit methods.
  • Boundaries are ASGS 2016 (the matching public release); production runs on the current ASGS.
  • There is no capacity, zone or travel-time layer; this shows the demand and uncertainty views only.
  • Rural SA2 boundaries are simplified so the map stays fast in a browser; production uses full-resolution boundaries.

Forward work: how this becomes the production system

  1. Ingest school enrolment, capacity and catchment data; calibrate government capture rates by area and level.
  2. Swap the trend baseline for the Hamilton–Perry cohort-component core, reconciled to Victoria in Future and made hierarchically coherent with MinT.
  3. Wire in the housing-unit method (dwelling approvals × typed student-yield) and PSP/UDP land supply as forward drivers of greenfield demand.
  4. Add the capacity overlay, demand-vs-capacity comparison, zone and travel-time tools: the actual provision-planning decision surface.
  5. Extend to single-year ages, Prep intake, and specialist-school & kindergarten disaggregation.
  6. Stand up the monitoring, assumption-invalidation and champion–challenger loop so the model is checked against outcomes and updated as new evidence arrives.
  7. Run to the full 20-year horizon, statewide, WCAG 2.1 AA.

Validation: rolling-origin back-test of the trend model

The demonstrator’s damped-trend forecaster is back-tested by rolling origin on the real ABS series (2001–2021). Median absolute percentage error grows with horizon and is far higher in growth corridors than established areas:

Growth class1 yr3 yr5 yr10 yr
High growth3.1%8.7%14.5%28.2%
Moderate growth1.7%4.7%7.8%14.1%
Stable / declining1.5%3.9%5.7%8.8%

The trend systematically under-forecasts high-growth SA2s (median bias about −26% at 10 years): it cannot see greenfield acceleration. That is the evidence-based case for the proposed cohort-component and housing-unit methods — the full model’s job, not demonstrated here. The back-test is reproducible from public data: school-demand/analysis/backtest.py.

Full methodology write-up: model stack, the trend back-test, where uncertainty concentrates, de-risking the build decision, model monitoring & references

Appendix: Methodology, validation & managing uncertainty for the infrastructure decision

Detailed methodology for this demonstrator. Everything below is either demonstrated on real public ABS data (see the pilot, cohort and leading-indicator appendices) or is the full model's design, cited to the literature. References at the end; only primary-source-verified claims are cited.

1. The decision this actually serves

The planning question is not "what will the population be", it is "where and when do we build or expand schools, and where do we hold off?" The forecast is instrumental to that decision:

forecast demand (by small area, cohort, year) → compare to current + planned capacity → shortfall / surplus by area and year → infrastructure & enrolment-management response (new school pipeline, expansions, modular classrooms, zone changes).

So the model is judged not by elegance but by whether it makes the build decision right, and by how much it reduces the cost of being wrong. That framing drives every choice below.

2. Methodology: the model stack

Component What it does Status / evidence on public data
Cohort-component core (Hamilton–Perry cohort-change ratios) Ages each cohort forward from people already observed (5–9 in 5 yrs = today's 0–4) rather than extrapolating a line Design (full model); to be validated against the trend once enrolment data by age is available
Reconciliation to Victoria in Future Constrains small-area forecasts to the state's official control totals Design (VIF is public; wired in for the full model)
MinT hierarchical reconciliation Makes SA1 → zone → region → state forecasts coherent (sum correctly) Design; method verified (Wickramasuriya et al. 2019)
Capture-rate sub-model Government share of enrolments by area & level, the biggest swing factor Design; calibrated on school enrolment data
Housing-unit method (dwelling approvals × typed yield) Forward demand in greenfield the history cannot see Design; needs building-approvals data to calibrate typed student-yield rates
Damped trend Transparent baseline & benchmark Built: the demonstrator tool's forecaster
Evidence-gated experimental track Benchmarks ML / reconciliation / foundation models; promotes only what beats the core on back-test Design (Grossman & Wilson 2022)
Uncertainty quantification Rolling-origin back-test → empirical error bands by area class × horizon Built: bands drive the tool (§4)

The core is deliberately a simplified cohort-component method: the small-area literature is consistent that simple, structurally-grounded methods constrained to reliable totals are hard to beat (Wilson 2015; Wilson et al. 2022), and my own back-test confirms it here.

Note on the demonstrator's age range. The public ABS ERP series used in the demonstrator publishes SA2 population only in 5-year bands, so it uses ages 5–14 (the 5–9 and 10–14 bands) as a clean school-age proxy and excludes the 15–19 band, which mixes senior-secondary students with post-school 18–19-year-olds. The production model works from school enrolment data by single year of age across Prep to Year 12.

3. Validation: rolling-origin back-test (trend model)

The demonstrator’s damped-trend forecaster is back-tested by rolling origin on the real ABS ERP series (2001–2021); the code is public: school-demand/analysis/backtest.py.

  1. Trend error grows with horizon and area volatility. Median APE at 5 years: high-growth 14.5%, moderate 7.8%, stable/declining 5.7%; at 10 years, high-growth 28.2%. The trend beats a naive last-value baseline overall.
  2. The trend systematically under-forecasts growth corridors — median signed error about −26% at 10 years in high-growth SA2s, because a trend cannot see greenfield acceleration. This is the measured, evidence-based case for the cohort-component and housing-unit methods.

Not demonstrated here (left for the full model): the cohort-component-vs-trend comparison and the dwelling-approvals leading indicator. Validating them needs school enrolment data (or ABS age-band and building-approvals series), so they are labelled design, not claimed as results on this public demonstrator. Being explicit about what is back-tested versus proposed is itself how the assurance gap the Auditor-General identified (VAGO 2017) is closed.

4. Where the uncertainty concentrates

The back-tests locate it precisely:

  • By horizon: error compounds; bands widen with years out (this is intrinsic).
  • By area type: high-growth/greenfield SA2s are 2–4× harder than established areas (trend median APE 14.5% vs 5.7% at 5 yr; up to ~28% at 10 yr in growth corridors).
  • Why: greenfield demand is driven by development timing (PSP roll-out, dwelling completions) and family in-migration that historical population series cannot see; and small starting denominators amplify percentage error.

The critical asymmetry: the highest-uncertainty areas, the outer growth corridors, are exactly where the biggest, most expensive, least-reversible infrastructure decisions are made (new-school pipeline). The widest error bands fall where the largest capital decisions are made, so those areas get interval forecasts and the closest monitoring.

5. Managing it: de-risking the build decision

Concrete, and mostly already evidenced above:

  1. Know the confidence per area. Every area ships with its back-tested error band, so planners weight decisions by reliability. The demand explorer surfaces this per SA2.
  2. Lead, don't lag, in growth areas. Use dwelling approvals + PSP/UDP land supply as early-warning signals: they lead occupancy by 2–5 years, flagging a corridor before the children arrive, when there's still time to acquire land and plan a school.
  3. Monitor high-uncertainty areas more often. Concentrate effort where error and stakes are highest: quarterly / trigger-based updates in growth corridors vs annual elsewhere.
  4. Plan robust to the range, not the point. Base build/no-build on the forecast interval against capacity: build where even the low case exceeds capacity; stage where only the high case does.
  5. Hedge with flexible capacity. Where uncertainty is high, prefer relocatable/modular capacity and land-banking (reversible, low-regret) over committing permanent capital until the signal firms; commit permanent build where confidence is high. This is real-options logic applied to school infrastructure. The forecast's uncertainty band is the input that tells you which lever.
  6. Triage by decision value. Prioritise accuracy where a wrong call is most costly (large greenfield precincts and capacity-constrained established schools), not uniformly across 300 areas.
  7. Back-test continuously, with assumption-invalidation alerts. Every cycle scores the prior forecast against new data; every material assumption (capture rate, typed yield rates, migration, damping, completion timing) sits in an assumption register with tolerances, and a breach (of a tolerance, of a class's back-tested error band, or a persistent bias) raises a drift flag naming the area and the assumption at fault. Flags drive targeted re-baselining, not a blind refresh; and new techniques run champion/challenger in shadow, promoted only when they beat the incumbent on the standing back-test. The accuracy/assurance dashboard makes the whole system self-correcting, improving on evidence over time, and auditable, with no silent method swaps.

Growth-corridor forecasts carry the most error. I report that error, use approvals as an early signal, update those areas most often, and recommend staged or relocatable capacity where the band is wide. That is what makes the forecast safe to build against.

References

Primary-source-verified (Tier A) and well-established (Tier B) from the literature review; full list in the literature review (Appendix E).

  • Hamilton, C.H. & Perry, J. (1962). A short method for projecting population by age from one decennial census to another. Social Forces 41(2). (the cohort-change-ratio method used here)
  • Wilson, T., Grossman, I., Alexander, M., Rees, P. & Temple, J. (2022). Methods for small area population forecasts: state-of-the-art and research needs. Population Research and Policy Review.
  • Wilson, T. (2015). New evaluations of simple models for small area population forecasts. Population, Space and Place.
  • Smith, S.K., Tayman, J. & Swanson, D.A. (2013). A Practitioner's Guide to State and Local Population Projections. Springer.
  • Grossman, I. & Wilson, T. (2022). Can machine learning improve small area population forecasts? A forecast combination approach. Computers, Environment and Urban Systems.
  • Wickramasuriya, S., Athanasopoulos, G. & Hyndman, R.J. (2019). Optimal forecast reconciliation for hierarchical and grouped time series (MinT). JASA 114(526).
  • Tayman, J. (2011) and Smith, S.K. & Sincich, T. (1990/1991), small-area forecast error and its growth with horizon and area volatility.
  • Victorian Auditor-General's Office (2017). Managing School Infrastructure (PP No. 253), the finding that demand forecasts were not routinely back-tested (Recommendation 4).

Note on tiers: every claim cited above is verified against a primary source or is a standard, well- established reference. Claims that could not be verified against a primary source are excluded.