Platform team · Jane's Weather / ACM (Australia)

Weather-Data Platform

Lead engineer for Jane's Weather, an Australian weather-data platform now part of ACM: from building its weather map to running the whole platform with the Territorial team.

Architecture

  1. Weather providers
    NWP models, BoM radar & satellite, warnings
  2. Argo Workflows
    Kubernetes · GRIB2 → Zarr · PMTiles
  3. Storage
    Cloudflare R2 · Supabase · OVH S3
  4. Edge Workers
    KV · D1 · R2-first delivery
  5. Weather map & clients
    MapLibre runtime · Vercel apps · APIs

I joined Jane's Weather as its GIS developer and built the weather map runtime and the data pipelines behind it. When the platform became part of ACM, the engagement grew: today I lead the Territorial team responsible for development, maintenance, DevOps, monitoring and documentation across the entire system.

At a glance

  • Built the weather map runtime (React, MapLibre, PMTiles) and its Argo data-processing workflows: forecast isobands, radar, satellite, warnings and basemaps
  • Argo Workflows on Kubernetes ingest seven forecast models (GFS, ACCESS-G, GEM, three ECMWF variants, CAMS) plus radar, satellite and warnings feeds
  • Cloudflare Workers, R2, KV and D1 serve tiles and forecast data at the edge, with Kubernetes APIs behind Kong as fallback
  • 63 synthetic tests in Sentinel verify freshness and health from outside the cluster; a Diátaxis-structured documentation site sourced from code, cluster inventories and a 130-page Confluence export

The platform

Jane's Weather produces forecasts, radar and satellite imagery, and warnings for consumer apps and for clients such as ACM's FarmOnline Weather. Data production runs as Argo Workflows and Kubernetes CronJobs on an OVH cluster in Sydney: GRIB2 model runs are converted to Zarr, radar and Himawari imagery become PMTiles and tile pyramids every five minutes, Bureau of Meteorology warnings are ingested into PostGIS. Delivery is storage-first: frontends resolve versioned artifacts in Cloudflare R2 through Workers, and a data-provider Worker uses D1 and KV for catalog and auth before falling back to the Kubernetes APIs.

The map

My first responsibility was the map. I designed and built the weather map runtime that is now the canonical map experience for the platform and for embedders such as FarmOnline Weather: a React application on MapLibre with PMTiles for basemaps and weather layers, two modes (forecast, timeline-first and model-driven; now, radar-prioritised with satellite, observations and warnings), a shared time-enabled layer abstraction with preloading, and a URL-first embed contract so partner sites can drive mode, model, layers and viewport from query parameters.

The map is only half of it. I also wrote the data-processing side that feeds it: the layer configuration model, the Python scripts that turn Zarr gridded forecasts into PMTiles isobands, and the Argo workflows that publish forecast, radar, Himawari, warnings and basemap artifacts to R2 on schedule, plus the headless capture tool that renders static forecast images for the main site.

Running the whole platform

After the platform moved to ACM and between clouds, the knowledge of how it fit together was spread across repositories, a stale Confluence space and people's heads. The Territorial team took on the whole platform, and I lead that work. We built the documentation program first: a VitePress site organised by system boundary, with every claim sourced from code or live configuration and every gap stated explicitly. That surfaced concrete risks (the development cluster serving production traffic, an ingestion cronjob whose deployed source was a different repository than everyone assumed, credentials with no rotation path) and turned them into tracked issues.

Then monitoring. Argo exit handlers only told us when a pipeline said it failed; nothing verified that artifacts were actually fresh and reachable. We deployed Sentinel with a suite that now covers the edge Workers, the Kubernetes APIs, the BoM proxy, upstream model sources and all three frontends, using independently computed staleness thresholds derived from observed model arrival times rather than the platform's own flags. Day to day the team handles development and maintenance across the APIs and frontends, Kubernetes secrets and rollouts, incident investigation and postmortems (an ECMWF rate limit, a GEM URL migration, a Vercel spend spike, a NaN-corruption crash loop in the consumer API).

More projects

All projects