Work / Independent research

United States & U.S. territories

Airport weather & disruption

When does weather disrupt an airport, and how early can we know?

Read the research design
Aviation view

Connecting to the weather snapshot…

Both views are experimental. Compare the selected airport’s backtest below.
Look ahead
All times are UTC. Airport-level context.

Explore the U.S. outlook

Where is disruption more likely?

Select an airport on the map or use search. Drag to pan; use + / − to zoom.

<15%15–30%30–45%45%+No estimateLow dataLimited evidence

Colors show predicted delay or cancellation share (plus diversion for arrivals). Dashed rings mark limited model evidence or a band below its coverage target. Gray means no estimate, never “low risk.” Shading applies at airports, not the airspace between them.

Loading the latest downloaded weather for Dallas–Fort Worth…

The airport’s data story

What supports this estimate?

Loading the held-out evaluation…

Coverage, not a promise

U.S. weather. Airport-specific evidence.

The catalog includes large and medium airports in the United States and U.S. territories. Available observations and forecasts vary by airport.

The study assesses 100 major U.S. airports. 97 have at least one tested direction and forecast horizon. Each additional airport is fitted and calibrated using its own flight outcomes. Horizons without enough evidence show a reason, with weather context still available.

Major U.S. airports are the modeling target. Medium airports remain searchable for weather and private-aviation context. The flight-outcome population is reporting carriers’ scheduled domestic flights, not every flight using an airport.

The research question

How much warning is useful?

I’m applying the same discipline I use in commercial forecasting: establish a baseline, measure what new information adds, and make uncertainty visible.

T−24

A day to plan

Can tomorrow’s weather distinguish an unusually disruptive airport window from a normal one?

T−12

A forecast revision

Does updated timing or a change in confidence improve the prediction?

T−6

Closer to the operation

How much additional signal is available as the airport’s arrival or departure window approaches?

Define the outcome

Estimate the share of scheduled domestic airline flights departing late or cancelled, and arriving late, cancelled, or diverted, within a two-hour window. A delay is at least 15 minutes. These are mutually exclusive outcomes, with an on-time remainder.

Reconstruct what was known

Join each window to forecasts available at the cutoff. Keep original issues, amendments, and retrieval times. Later weather observations can verify a forecast, but cannot become an earlier prediction’s inputs.

Measure what each input adds

Compare a calendar baseline with the same learner plus airport forecasts. The original 25 models retain their seasonal development experiment. Added airports use fixed regularized logistic models trained only on their own outcomes, with separate probability and uncertainty calibration. Wind, gusts, visibility, ceiling, storms, precipitation, and forecast age enter alongside airport, local time, day-of-week, and seasonal patterns. Flight counts weight fitting and evaluation; actual arrival/departure volume and live congestion are not predictors.

Publish the evidence

Use chronological splits, probability calibration, Brier score, and log loss. Compare the same eligible windows at each horizon, report missing coverage, and assess uncertainty by day or storm rather than treating every flight as independent.

A test that can fail

Fit: 2023-01-03 to 2024-06-27. Probability calibration: 2024-07-04 to 2024-12-27. Interval calibration: 2025-01-04 to 2025-09-27. Winter backtest: 2026-01-03 to 2026-02-26. New airports retain the available windows at each horizon; matched-window scores separately compare T−24, T−12, and T−6 where all three forecasts exist. These dates and some connected flight outcomes were inspected in prior experiments, so this is retrospective testing. March–June 2026 provides a separate stress re-evaluation. Prior versions remain downloadable.

Download the protocol, exclusions, sources, and evaluation (JSON)

FAA history, tested nationally

Does an operations model improve the outlook?

FAA advisory history adds a small average gain over matched local weather models. It does not beat the currently served weather model across both seasons, so weather remains the default. Choose “Operations” above to explore the alternative and its local evidence.

100 major airports assessed · 97 with sufficient outcomes · 478 arrival/departure/horizon comparisons per period. No probabilities transfer to airports without outcomes.

Winter · January–February 2026

Mean error ↓Current weather modelMatched local weatherOperations
Brier error0.167910.167760.16746
Log loss0.517120.516960.51588

161 of 478 operations comparisons lowered Brier error versus matched weather; 5 retained positive improvement after the approximate multiple-comparison adjustment. 1 also met the sample, calibration, and uncertainty criteria.

Operations’ nominal 80% outcome-rate bands covered 75.5% of test windows on average. Lower probability error does not guarantee reliable uncertainty bands.

Spring · March–June 2026

Mean error ↓Current weather modelMatched local weatherOperations
Brier error0.173130.174580.17421
Log loss0.525500.529240.52844

158 of 478 operations comparisons lowered Brier error versus matched weather; 27 retained positive improvement after the approximate multiple-comparison adjustment. 7 also met the sample, calibration, and uncertainty criteria.

Operations’ nominal 80% outcome-rate bands covered 78.3% of test windows on average. Lower probability error does not guarantee reliable uncertainty bands.

What FAA history means—and how the default was chosen

The model counts airport-specific ground-stop and ground-delay-program messages in the previous 6 and 24 hours, including revisions and cancellations. This is a signal of recent activity, not a complete inventory of active restrictions, traffic volume, runway capacity, or aircraft rotations. Carrier-specific notices are not restrictions on every flight.

903 cached national daily listings passed validation; 30 dates remained missing or ambiguous. Every comparison uses identical eligible windows. Publication timestamps and the hourly collection phase determine what was known; missing history never becomes a zero-risk signal. The same FAA feature code is used in training and live scoring.

The three local variants share fixed learners and separate fitting, probability-calibration, and uncertainty-calibration periods. Both 2026 evaluations are retrospective. Tables average airport/direction/horizon slices equally. The original production comparator was scored on the same windows, preserving its existing training and calibration.

186 comparisons had no FAA activity during training, so they cannot establish a learned FAA effect. Tiny numerical differences from fitting were excluded from improvement tests before release; raw scores and model weights were preserved.

Operations would become the default only if average Brier error and log loss beat both matched weather and the currently served model in both periods. It did not meet that rule. Local evidence, including regressions, remains visible when you select an airport.

Protocol, national results, and release decision · Full reproducibility report

A traffic proxy, tested

Does traffic volume improve predictions?

Adding historical arrival and departure volume did not reliably improve predictions in this experiment. Average probability error rose slightly in winter and fell slightly in spring; average calibration error increased in both. The live model has not been replaced.

100 airports assessed · 97 with sufficient evidence · 478 airport/direction/horizon comparisons per period. Every variant uses the same test windows within each comparison.

Winter · January–February 2026

Mean metricCalendar + weather+ historical traffic
Brier error ↓0.166600.16675
Calibration gap ↓4.36 pp4.52 pp
Band coverage · 80% target75.9%75.5%

250 of 478 comparisons improved on Brier error; 38 retained a positive improvement after the approximate multiple-comparison adjustment. Local gains do not establish a general improvement.

Spring · March–June 2026

Mean metricCalendar + weather+ historical traffic
Brier error ↓0.175230.17514
Calibration gap ↓3.91 pp3.99 pp
Band coverage · 80% target78.2%78.0%

266 of 478 comparisons improved on Brier error; 114 retained a positive improvement after the approximate multiple-comparison adjustment. Local gains do not establish a general improvement.

How the traffic test avoids looking ahead

At each airport, older flight counts estimate typical arrivals and departures by local weekday and two-hour block. The profile uses up to 12 source months with a conservative 120-day publication lag. Missing source months are excluded, not counted as zero. Actual flight counts in the predicted window never become a predictor.

Calendar-only, calendar plus weather, and calendar plus weather plus traffic use the same fixed airport-local logistic learner, with separate probability and uncertainty calibration. Traffic features include typical demand, demand relative to historical peak periods, adjacent periods, and weather interactions. Historical peaks are not measured runway capacity.

Model fitting uses September 2023–June 2024, probability calibration July–December 2024, and uncertainty calibration January–September 2025. Both 2026 evaluation periods are retrospective; their dates and connected outcomes have been seen before. The publication lag is assumed rather than reconstructed from historical data releases.

The tables average the matched airport/direction/horizon comparisons equally. Tests across horizons may contain different windows. The downloadable report also provides flight-weighted comparisons with shared calendar-block uncertainty. This experiment has different fitting dates and local comparators from the live model, so it is not a direct before/after production comparison.

Download protocol and airport results (JSON) · Full reproducibility report (compressed JSON)

Historical replay · DFW · May 2025

What did the forecast say beforehand?

Choose a two-hour airport window. Compare the archived forecasts at three cutoffs with the flight outcomes recorded afterward.

One month is a test of the data pipeline, not a representative annual baseline.

Reading this tool

Freshness and missing data

A shared AWS collector checks the weather feeds hourly and publishes a complete snapshot for this page. Reports themselves have different issue schedules. Each result shows report and download times; an old or missing report is not evidence of good weather. Probabilities are withheld when the snapshot is more than 90 minutes old.

Airport outlooks and individual flights

This project does not look up a flight or monitor its current status. Estimates describe reporting carriers’ scheduled US domestic flights in an airport window, including cancellations. They do not condition on a flight still being active at the cutoff. Crew, maintenance, traffic restrictions, and disruptions elsewhere can change an individual flight’s outcome. Weather associations do not establish weather as the cause of every delay.

Traffic volume and factors beyond weather

The live model includes the airport, local hour, day of week, and season, with separate arrival and departure models. These capture typical operating patterns, but do not measure that day’s demand or congestion. Historical flight counts determine each window’s weight and evidence support; they are not traffic-volume inputs to the live model. The optional operations model adds recent FAA ground-stop/GDP advisory activity. Live arrival/departure demand, actual runway capacity, complete restriction scope, delay buildup, airline mix, staffing, maintenance, and aircraft rotations are not modeled.

The airport panel adds hourly FAA airport notices and a monthly runway inventory as separate context. The current-notice panel and runway inventory are separate context. The operations model uses recent advisory-message history rather than treating these displayed notices as universal restrictions. FAA status describes its report time, not the selected future departure or arrival; a missing notice is not evidence of normal operations. Runway counts and dimensions describe layout, not current usable capacity. Closure notices may affect only certain flights or include exceptions.

We tested a lagged historical traffic proxy against calendar and weather models on the same chronological windows. It did not establish a reliable general improvement, so the live model remains unchanged. Read the experiment and its limitations.

What confidence means here

The 80% target band describes variation in the observed disruption share across two-hour windows. Nine months of errors calibrate separate upper and lower limits, by airport and predicted-risk band where enough samples exist. Confidence intervals resample blocks of up to three days. Moderate evidence also requires a multiple-comparison adjustment across all tested slices, at least 45 test days and 5,000 flights, calibration error below five percentage points, and interval coverage of at least 75%. Temporal dependence means target coverage is not guaranteed. Actual coverage and band width are displayed; live confidence uses the winter backtest. There are 478 tested airport, direction, and horizon combinations in that period.

Backtest limitations and next experiments

Historical forecast availability uses the later of issue and archive product time, plus 10 minutes. A 60-minute delivery sensitivity check is published. Final-vintage flight outcomes are used; data-release revisions and cancellation announcement times are not reconstructed. Low-volume windows, unknown outcomes, unsupported forecasts, month edges, and incomplete coverage are excluded. Backtests use fixed local two-hour blocks; the live view uses rolling windows. Live windows touching a local clock block with fewer than 30 eligible training windows show “low data.” The winter and spring evaluations are retrospective. Additional-airport settings were fixed before fitting their models; training, probability calibration, uncertainty calibration, and test dates remain separate. Adjusted significance is approximate and does not remove dependence or selection uncertainty. Forecast error, upstream disruption, aircraft rotation, and weather beyond the airport are not yet modeled.

Private aviation and shared data

Choose one airport, arrivals or departures, and the airport time. Results populate automatically, with forecast coverage at +6, +12, and +24 hours. Both aviation views use the same NOAA weather reports, FAA airport notices, and runway inventory. The arrival/departure selector identifies the operation at that airport; it does not change the underlying weather forecast or filter FAA notices into a complete restriction briefing.

The private-jet view is independent of airline-model coverage, so medium airports remain available wherever weather is reported. Map colors describe prevailing ceiling and visibility, with temporary and probability groups flagged separately. A missing or expired forecast stays unavailable; observations do not replace future forecasts.

Airline probabilities additionally use historical scheduled-airline outcomes. Those delay rates and backtest confidence have not been validated for private jets. Weather categories are not aircraft operating minima; this view does not assess aircraft capability, route conditions, or whether a flight is safe.

Conditional weather forecasts

Temporary, transition, and probability groups remain separate from prevailing conditions. A “30% weather group” concerns the specified weather, not a 30% chance of delay or cancellation. Missing fields in a change group do not mean calm winds or clear skies. Read the source report for its full context.

Cost and operating limits

This version uses public data without paid API subscriptions. Visitors read shared snapshots; opening the map never starts a data collection or training job.

A shared AWS job refreshes weather hourly. It covers the airport catalog; airline probabilities are available only for the airports supported by the published model. A monthly job can evaluate a weather-model candidate after enough new flight outcomes arrive. The operations option currently uses the completed national backtest; its FAA inputs refresh hourly and are archived for future evaluation, while its fitted weights remain fixed. Candidate models require review before replacing an active model.

The service currently runs on an AWS Free account plan, with no charges while that plan is active. The trial ends March 11, 2027 or when credits run out, whichever comes first. Continued service then requires a separate cost review. Collection and training have run limits, size limits, an early budget stop, and a renewable operating period of at most 31 days. Missing, stale, or unverified data can interrupt estimates. No automatic paid upgrade is enabled.

Sources and coverage

Weather comes from the NOAA Aviation Weather Center. The collector uses its public bulk feeds and retains only weather for the U.S. airport catalog. Airport and runway metadata come from OurAirports, using large and medium airports with four-character identifiers in the United States, Puerto Rico, the U.S. Virgin Islands, Guam, the Northern Mariana Islands, and American Samoa. Airport notices come from the FAA’s public airport-status feed; en-route airspace programs are not mapped to individual airports. The operations option uses dated FAA advisory listings for recent message activity. The study uses BTS airline on-time records and Iowa Environmental Mesonet forecast archives. Geography is from Natural Earth. Private aviation receives weather and airport context, without airline outcome probabilities.