AIFL: advancing global streamflow prediction with operational machine learning

27 August 2026
Maria Luisa Taccari
Kenza Tazi
Matthew Chantry
Christel Prudhomme

Every year, floods affect more people worldwide than any other natural hazard. Accurate streamflow forecasts, predicting how much water will flow through a river in the coming days, are essential for issuing timely warnings, managing reservoirs, and protecting lives.

For over 15 years, ECMWF has provided the computational backbone for the European and Global Flood Awareness Systems (EFAS and GloFAS), coupling weather forecasts with process-based hydrological models, which are physics-driven simulations that represent the movement of water through the landscape using mathematical equations. Now, we are exploring how machine learning can complement these capabilities.

In a new study, we introduce Artificial Intelligence for Floods (AIFL), a global streamflow forecasting model that learns directly from data and, crucially, learns to cope with the realities of operational weather forecasting.

This work contributed to the Destination Earth (DestinE) initiative of the European Commission. As part of DestinE, ECMWF together with many partners is building data-driven models across the Earth system.

AIFL targets the hydrology component, complementing parallel efforts on the atmosphere, ocean, sea ice, waves, and land, and helping lay the groundwork for a fully coupled machine learning Earth system model. 

The challenge: training on the past, forecasting in real time 

Machine learning models for hydrology are typically trained and forced with atmospheric reanalysis data, which provide gridded reconstructions of past weather that are temporally consistent and grounded in observations via data assimilation.

This is a problem as when these models are deployed operationally, they are driven not by reanalysis, but by weather forecasts from systems like ECMWF’s Integrated Forecasting System (IFS). These systems have a different climatology as well as forecast uncertainty.

This mismatch, called the reanalysis-to-forecast domain shift, can significantly degrade a model's performance in the real world. A model that looks excellent in research papers may struggle when it matters most. 

A two-stage learning strategy 

AIFL addresses the gap between reanalysis data and operational forecasts with a straightforward but effective two-stage training approach.

In the first stage, we pre-train a Long Short-Term Memory (LSTM) neural network on 40 years of ERA5-Land reanalysis data (1980–2019). The LSTM is a type of deep learning model specifically designed to learn patterns in sequential data, making it well suited to time series like rainfall and river flow. 

The training covers over 18,000 river basins worldwide curated from the open-source CARAVAN dataset, a publicly available community-curated collection of streamflow records and catchment attributes from across the globe. This stage teaches the model robust rainfall–runoff relationships across a huge diversity of climates, landscapes and river systems. 

In the second stage, we fine-tune the model on operational IFS control forecasts from 2016 to 2019. This step exposes the network to the specific error structures and biases of real-time weather predictions, allowing it to adapt its internal representations accordingly. The result is a model that has learned hydrology from decades of consistent data but has also been calibrated to perform under operational conditions. 

How does AIFL perform? 

We evaluated AIFL over an independent test period from 2021 to 2024 across 2,000 gauged river basins. The model achieves a median modified Kling–Gupta Efficiency of 0.66 – above the threshold generally considered satisfactory in the hydrological literature (Figure 1) – and a near-perfect median discharge bias ratio of 1.00, indicating that it neither systematically overestimates nor underestimates river flow volumes. 

Four regional maps (Europe, North America, South Africa, and Australia & New Zealand) showing colored points at monitoring sites, with colours indicating KGE scores on a scale from −1 to 1.

Figure 1: A global view of forecasting skill. AIFL has extensive coverage, spanning thousands of gauge stations across the globe. Our analysis focused on four key regions: Europe, North America, South Africa, and Australia and New Zealand. In these areas, the model consistently demonstrates high performance, with many stations achieving Kling–Gupta Efficiency (KGE′) scores well above the 0.5 "skillful" benchmark, proving its reliability across diverse climates and terrains. 

We also benchmarked AIFL against the Google global flood model across more than 1,200 shared stations. AIFL matches or exceeds Google's skill at 43% of locations, showing that AIFL is competitive but also has room for further improvement. We found that performance varies with basin size: AIFL is particularly competitive in smaller catchments and provides consistent, stable predictions across all scales. 

For flood event detection, the model is notably conservative. When AIFL triggers a flood alert, it is almost certain to correspond to a real event – we found zero false alarms across all return periods up to 50 years. This high reliability comes at the cost of missing some events, particularly the rarest ones. Improving the detection of extreme floods without introducing false alarms is a clear priority for future work. 

A real-world test: Storm Henk 

In January 2024, Storm Henk brought severe flooding to parts of Belgium. At the Straimont station on a 182 km² catchment, AIFL consistently predicted a clear flood signal up to six days before the observed peak – an approximately 1-in-20-year event (Figure 2).

While forecast magnitudes varied slightly across lead times, the event was robustly detected well in advance, demonstrating the practical value of the model's conservative but reliable detection strategy. 

Map showing the location of a European streamflow station and its upstream catchment, alongside a time series comparing observed and predicted streamflow during a peak flow event, with forecast lead times indicated by colour.

Figure 2: Flood forecasting case study for Storm Henk (January 2024) at a Belgian river station. Observed streamflow (orange dash) is shown alongside AIFL forecasts at different lead times (blue shades; darker colours indicate shorter lead times). The model detected the approximately 1-in-20-year flood signal up to six days in advance. 

Why AIFL matters 

AIFL is designed from the outset for integration into real-time forecasting workflows at ECMWF. By solving the reanalysis-to-forecast domain shift through a simple and reproducible training strategy, it provides a transparent baseline that complements existing systems like GloFAS. The model was developed using the Neural Hydrology Python library and trained on the CARAVAN dataset, ensuring a robust and reproducible framework for the broader research community.  

As we look ahead, we plan to extend AIFL with probabilistic outputs for uncertainty quantification, explore multi-source precipitation products to improve event detection, and conduct systematic comparisons with GloFAS. By combining the strengths of machine learning with the operational infrastructure of ECMWF, we aim to push the boundaries of what is possible in global flood forecasting. 

Read the paper

The full paper is available to read in the Journal of Hydrology.

 

Top banner image: © Viktoriia Voievodina / iStock / Getty Images Plus

DOI
10.21957/9432d75dc2