Technical programming course

Advanced Programming for Climate Data Analysis

Develop practical capability for large-scale climate-data analysis using modern Python tools, cloud-native data access, geospatial workflows, non-stationary statistics and machine-learning methods.

Course purpose

Move from climate-data access to scalable analytical workflows.

This course is designed for participants who need to work confidently with modern climate datasets, from multi-dimensional arrays and out-of-core computing through geospatial analysis, trend detection, downscaling and extreme-event detection.

Handle modern datasets

Work with Xarray, NetCDF, Zarr and cloud-native climate archives rather than flat tables alone.

Scale analysis effectively

Use Dask, chunking strategies and memory-aware pipelines for datasets larger than local RAM.

Build analytical workflows

Apply geospatial methods, time-series analysis, spectral methods, machine learning and extreme-event detection.

Target participants

Built for technical climate-data practitioners.

The course is well suited to analysts, researchers and environmental data practitioners who want to strengthen applied programming skills for climate and geospatial datasets.

Climate-data analysts Atmospheric scientists Environmental engineers Geospatial practitioners Researchers and graduate trainees Climate and adaptation teams Data scientists working with environmental data
Curriculum

Eight live modules with concept instruction and coding labs.

Each day combines a 60-minute live concept lecture with a 60-minute live data or coding lab built around realistic climate-data tasks.

Day 1 Multi-Dimensional Climate Arrays (Xarray & NetCDF)2 hours

Live concept

Master the standard data structures of climate science by moving beyond flat tables into spatio-temporal data cubes using Xarray.

Analysis focus

  • Dimensions such as time, latitude, longitude and pressure levels
  • Metadata attributes, label-based indexing and coordinate alignment
Live data labSlicing global reanalysis data such as ERA5 by loading a large dataset, extracting a localized region and subsetting pressure levels without loading the full array into memory.
Day 2 Scaling Analysis with Out-of-Core Computing (Dask)2 hours

Live concept

Overcome memory limits by analyzing climate datasets larger than your RAM using lazy evaluation and parallel task graphs.

Analysis focus

  • Optimizing array chunking across spatial and temporal axes
  • Monitoring memory footprints using the Dask dashboard
Live data labCalculating long-term anomalies through a memory-efficient pipeline that computes a 30-year daily climatology and extracts daily temperature anomalies from a 200 GB global dataset.
Day 3 Cloud-Native Architecture & Planetary-Scale Access2 hours

Live concept

Shift from “download then analyze” to “analyze in the cloud” by interfacing with optimized cloud-hosted formats and metadata catalogues.

Analysis focus

  • Querying SpatioTemporal Asset Catalogs (STAC)
  • Using cloud object storage such as AWS S3 and Google Cloud Storage
  • Understanding the advantages of Zarr over NetCDF in cloud workflows
Live data labQuerying a CMIP6 climate-model projection archive in the cloud, filtering by experiment scenario such as SSP5-8.5 and lazily streaming targeted spatial subsets into your environment.
Day 4 Advanced Geospatial Aggregations & Masking2 hours

Live concept

Merge vector-based geographic boundaries such as political borders, watersheds and ecological zones with gridded climate datasets.

Analysis focus

  • Spatial clipping, rasterization and spatial weighted averages
  • Accounting for latitude-based grid-cell area distortion
  • Optimizing Rioxarray and GeoPandas pipelines
Live data labCalculating watershed-specific precipitation trends by masking a gridded satellite rainfall dataset with a global river-basin shapefile and extracting area-weighted monthly timeseries.
Day 5 Time-Series Analysis & Non-Stationary Statistics2 hours

Live concept

Apply advanced temporal analytics to a changing climate in which historical statistical baselines are no longer constant.

Analysis focus

  • Rolling-window statistics and handling missing or unevenly sampled data
  • Monotonic trend detection with the Mann–Kendall test and Sen’s slope
  • Seasonal decomposition
Live data labIsolating the climate-change signal by detrending a long station or tidal-gauge record, patching observation gaps and separating long-term change from seasonal oscillations.
Day 6 Frequency Domain & Spectral Analysis2 hours

Live concept

Move beyond the time domain to analyze climate cycles, oscillations and return-period behaviour for extreme events.

Analysis focus

  • Fast Fourier Transforms, power spectral density estimation and wavelet analysis
  • Localized time-frequency variations and filtering high-frequency noise from low-frequency climate trends
Live data labIdentifying cyclical climate signals by analyzing a multi-decade sea-surface-temperature dataset to detect and isolate the periodic signature of El Niño–Southern Oscillation.
Day 7 Statistical Downscaling & Machine Learning Workflows2 hours

Live concept

Use data-driven methods to translate coarse climate-model outputs into high-resolution, actionable local insights.

Analysis focus

  • Feature engineering using physical variables such as elevation and wind vectors
  • Spatial and temporal cross-validation to avoid data leakage
  • Training regression trees or neural-network-based approaches on spatial blocks
Live data labBuilding a predictive temperature-downscaling pipeline with Scikit-Learn or XGBoost using future projections and high-resolution digital elevation models.
Day 8 Extreme Event Detection & Vectorized Thresholding2 hours

Live concept

Design robust algorithms to quantify climate extremes rather than simply tracking average changes.

Analysis focus

  • Programming custom rolling thresholds and run-length encoding
  • Multi-variable compound-extreme analysis
Live data labBuilding a highly optimized vectorized extreme-event detector that scans a 4D climate array to flag, count and map the frequency and duration of compound drought and heatwave events.
Key data packages and infrastructure

Core tools taught in the course.

  • Core array analysis: Xarray, NumPy, SciPy
  • Parallel computing: Dask
  • Spatial analysis: GeoPandas, Shapely, Rioxarray, Rasterio
  • Data access: STAC API, fsspec, Zarr
  • Statistics and machine learning: Statsmodels, Scikit-Learn
Technical enrolment CAD 1,800

Per participant, inclusive of taxes. Contact NeoCare to discuss schedules, cohort planning or course questions.

Strengthen your technical climate-data capability.

Request the next available schedule or discuss whether this course fits your team’s analytical workflow needs.