the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Efficacy of Scalable Airline-led Contrail Avoidance
Abstract. Contrails account for a large portion of aviation's contribution to anthropogenic climate change. Navigational contrail avoidance is a promising solution to mitigate the warming caused by contrails. Prior trials testing navigational contrail avoidance have relied on bespoke integrations of contrail forecasts into airline operations. Here, we use a randomized control trial to test the feasibility of dispatcher-led contrail avoidance integrated into standard flight planning operations using a workflow that scales to an airline's entire network. We validated the efficacy of this intervention using satellite imagery and an automated flight-contrail attribution algorithm. Using this system, we observed an 11.6 % reduction in contrail formation rate for the 1232 flights marked as eligible for contrail avoidance (intent-to-treat) relative to the flights in the control group (p = 0.011). In the 112 flights that flew contrail avoidance as planned (per-protocol flights), we observed a 62.0 % lower contrail formation rate relative to the flights in the control group (p < 0.001). No statistically significant difference in fuel usage was observed between the two groups.
Competing interests: As denoted by their affiliations, some authors are employed by Google LLC, Flightkeys GmbH, and American Airlines. Contrails.org is operated by Breakthrough Energy, a family of organizations and activities committed to transitioning the world to net zero by 2050. All other authors declare no competing interests.
Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.- Preprint
(1578 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
-
RC1: 'Comment on jecats-2026-4', Anonymous Referee #1, 25 May 2026
-
AC1: 'Reply on RC1', Tharun Sankar, 22 Jul 2026
Dear reviewer,
Thank you for your thorough review of our manuscript. Please see our responses inline below.
General comments
Contrail avoidance is often seen as a possibility to quickly lessen the rate at which aviation's climate warming contribution rises. For this reason it is important to conduct tests that occur not only in computer simulations but as well in the operational environment. A few such trials did occur in the recent past but they involved only few flights to evaluate, and the planning and evaluation was done manually which takes too much time for an operational system. The trial described in this manuscript intends to overcome some of these problems, the much manual work and the low number of flights for evaluation which is a step forward. But the paper also demonstrates that there are still difficulties and room for improvements, issues that must be tackled before contrail-avoidance flight-planning can actually become operational. These issues are discussed in Sect. 4, and this is the best part of the manuscript.
Otherwise, I admit that it was difficult for me to understand how this trial was operated although the authors tried to explain it. Things and notions are not clear to me. Thus I recommend to work especially on an exact description of what has been done, explaining all the steps.
I have also a problem with the climatological evaluation of the trial. The evaluation of how much contrail coverage was avoided is ok, but the evaluation of how much forcing or climate impact was avoided is not. Since this evaluation is based on a climatology, I believe that the results are more or less random, and that therefore the statistical evaluation ("statistically significant...") does not say anything for the real world. Please consider that the actual weather is the most dominant influencing factor on contrail radiative impacts. For this reason, I think, the ideas of conducting contrail-avoidance in a "climatological sense" have been abandoned about 20 years ago. I hope that this is acknowledged in the revised version.
In retrospect, we regret using the term “climatology” throughout this manuscript, as it does not accurately describe our methods. We have replaced this term with “observation-informed average forcing,” which is a more accurate description of our methods. The forcing estimate in this method is spatio-temporally averaged CoCiP warming. Importantly, the forcing estimates are always masked by whether or not a contrail was actually observed, which means that the actual weather is always included in the evaluation. We have included a more detailed response to the reviewer’s major comment below to address this concern. We have also included a second estimate of contrail forcing to the impact evaluation of the trial, which will hopefully address some of the concerns with this averaged evaluation.
Major comments:
1) I had problems to understand how this trial had been performed. I have a lot of questions for clarification:
We have made changes to Sec. 2.4 (Operational Workflow) to clarify the questions raised here by the reviewer, including moving some details from Sec. 2.2 to Sec. 2.4 to improve the readability of the manuscript. For each individual question, we provide an answer inline below for additional clarification.
1.1) There is a control group and an intent-to-treat group. There are also "non-avoidance" plans and "contrail-optimized" plans. It is unclear how these two pairs of notions are related. Is the control group equivalent to the group with the "non-avoidane" plans? Still another expression is "per-protocol flights". I am confuesed by this multiplicity of notions.
The control and treatment (intent-to-treat) groups are the sets of flights determined by the city pair randomization procedure. For a given flight, multiple candidate flight plans may be presented to the dispatcher. All flights, whether in the control or treatment group, will have at least one non-avoidance flight plan. Treatment group flights also have a contrail-optimized flight plan. This distinction is hopefully more clear in the updated Sec. 2.4.
“Per-protocol” is a term from the clinical trial literature which denotes participants in the study which strictly adhered to the assigned treatment, in this case contrail avoidance. We include a definition in the abstract and expanded the definition in Sec. 2.7. We additionally added a citation which defines the terms “intent-to-treat” and “per-protocol.”
1.2) Flights with a "non-avoidance" plan were re-optimized including the contrail cost term. Why and for what purpose?
By default, all flight plans in AAL’s system are “non-avoidance” plans, since contrails are not considered in their standard operations. Therefore, when the original non-avoidance plan met the 10 t contrail warming threshold, we triggered the re-optimization to create a new flight plan which does consider contrails. This new “contrail-optimized” flight plan was presented to the dispatcher alongside the original non-avoidance plan, so that the dispatcher could use their discretion to decide whether to work on the contrail-optimized plan. We hope the changes to Sec 2.2 and 2.4 have made this more clear.
1.3) Why and for what purpose were dispatchers presented with one or more non-avoidance plans? Is it a common process to generate several flight plans for the dispatcher to select one?
Yes, this is a common process. In the majority of cases, a dispatcher will only be looking at one non-avoidance plan. However, in some cases, they will consider multiple plans. For example, a dispatcher might trigger the generation of one non-avoidance plan per NAT track to assess which track in the North Atlantic system is optimal for fuel and turbulence. We added a sentence clarifying this to Sec. 2.4.
1.4) Initially I thought Flightkeys computes an optimal route, either cost or contrail-optimized, and that this is given to the dispatcher. I thought that neither the dispatcher nor the pilot knows on what kind of optimization the flight was planned. But this impression later turns out wrong.
This is hopefully clarified with the reordering of sections and consolidation of the operational workflow in Sec. 2.4. Voluntary participation in the trial was an important operational constraint. In order to achieve this, in the treatment group, dispatchers and pilots were informed which plan was a contrail-optimized plan, so that they could choose not to participate if they desired. In the control group, operations proceed as normal with non-avoidance plans, so dispatchers and pilots are not informed that they are participating in the trial.
1.5) Please define the L1,2,3 groups with more details. My questions were: Does the L1 group contain routes for which no (strong) contrail was forecasted? If so, then the L2 group is the one where contrails were forecasts, but not all of them have been avoided when the dispatcher had other priorities? And the, L3 is the group where contrails have been avoided acitvely by dispatcher and pilots. Is this correct?
We have added some clarifications to Sec. 2.7 to address these comments. To answer your specific questions:
- No, the L1 group does not contain routes for which no strong contrail was forecasted. This is achieved by applying the 10 ton threshold on contrail warming of the original non-avoidance plan. This is mentioned in the first sentence of Sec. 2.7 (Randomized Controlled Trial). We consider this “intent-to-treat” since in the L1 treatment group, a contrail-optimized plan is proposed to the dispatcher.
- The L2 group is the set of flights where contrails were forecasted and the dispatcher considered it safe and flyable to use the contrail-optimized plan. The reason that this is different from L3 is that the pilot may not follow the flight plan exactly.
- Yes, the L3 group is where contrails have been avoided actively by dispatchers and pilots.
2.) As stated above, I find the evaluation of the contrail impact (or the avoided impact), based on climatological fields, insufficient. I wonder whether the derived values have any significance. You can say, if we assume that the actual weather were like the climatology, the results were x and y, and the difference were statistically significant. BUT, the weather is generally not like the climatology on a certain day, and thus we must expect large deviations of the actual forcings from the obtained ones. As these values are not known, their differences cannot be determined. To my opinion, the corresponding tables and statements should be removed from the paper in order not no rise wrong expectations in uncautious readership.
As mentioned at the beginning of our response, when reporting forcing, if a contrail is not observed we always report 0 forcing. Therefore, the reported result is strongly influenced by the observed reduction in contrail formation (e.g. had we observed no reduction in contrail formation, we would be reporting no reduction in forcing). After this masking, the question is how much warming to attribute to the contrails that we do observe. In the updated manuscript, we use two different methods for estimating the warming, and we have renamed the “climatological” estimate to the “observation-informed average forcing” estimate to clarify the methods used to calculate it.
Variability in contrail forcing between flights can be broken down into two components: the variability of whether a persistent contrail formed, and the variability in the warming between persistent contrails. The variability of the first component is very large, with ~90% of flights not intersecting regions of persistent contrail formation (ice supersaturated regions). The variability of the second component, though nonzero, is not as large. Every nighttime contrail which persists for the 20-30 minutes required to be visible from a geostationary satellite has nontrivial radiative forcing. Our ‘observation-informed average forcing’ approach is ignoring only the variability between persistent contrails (and in fact only a subset of it since we take into account time of day, season and latitude). For example, Platt et al (2024) shows that using a climatological forecast to choose 20% of persistent contrails for avoidance leads to avoiding 40% of the most warming contrails and all contrails avoided have a positive (warming) energy forcing.
In response to this comment, we have also added another way to estimate the warming of the contrails we observe. For each observed contrail, we ran CoCiP 10 times using as input 10 different IFS ensembles, and we report the average of all the ensembles where contrails are predicted. Platt et al (2024) also evaluated this method for predicting the energy forcing of formed contrails and found it performs well.
Both methods of estimating warming find that the warming reduction is similar to the reduction in observed contrail distance. For example, the central values of the reductions in the L3 group are 62.0%, 69.3%, and 83.0% for the observed contrail distance, observation-informed average forcing, and observation-informed ensemble forcing endpoints respectively. In other words, our conclusion that the warming of contrails is reduced is based primarily on the reduction in observed contrail distance, not on simulations claiming that the observed contrails in the treatment group are less warming than those in the control group.
Minor comments:
1.) I suggest to spend a bit more time to explain a bit how the ML-contrail forecast works. A short overview is sufficient. This is for the convenience of the reader, who otherwise has to change to the paper by Sonabend-W (and in fact, I did not remember that a description of the ML-based contrail forecast was given there, but that may be my fault). I wonder how satellite images of contrails or series of images can be used for contrail forecasting. The contrails in the images already exist and stem from past flights where contrail formation has not been avoided. Please help the reader with a short text to an imagination how this system works.
We have added a new section (Appendix A) which describes the ML contrail forecast in much more detail.
2.) The trial took place mainly over the western part of the Atlantic ocean where flight density is probably low. Do the authors think that the methods investigated in this trial can be used in similar way in more congested regions like continental US or even Europe?
This is an important question and one we hope to answer in future research, but investigating the applicability of this intervention to other operational environments is out of scope for this work. We added a bit more detail to the last paragraph of the Conclusions outlining this future work.
3.) "The results provide strong evidence... physically effective and operationally feasible." This is an unclear statement. I believe that contrail avoidance is operationally feasible, under certain circumstances, as you show in the paper, namely when the dispatcher does not have other problems to solve. Whether it is feasible in congested air-spaces remains doubtful. What do you mean with "physically effective"? If you mean the reduction of contrail coverage, please say so. If you however mean the calculated reduction of warming impact, I do not agree unless you change your method from a climatologically based one to a actual-weather based one. The expression "climatological warming" makes no sense for single flights.
We have edited this sentence to say “The results provide strong evidence that navigational contrail avoidance is effective in reducing observable contrail formation and operationally feasible in realistic circumstances.” We have also included the additional warming estimate in the results.
Miscellaneous:
L 29-31: One should remark here that at least in the paper by Frias et al. a perfect weather forecast was assumed which renders results doubtful.
We have edited two sentences after this to say “Additionally, several studies have called into question how well numerical weather prediction (NWP) models and their reanalysis products match real-world observations (Agarwal 2022, Gierens 2020), which introduces additional uncertainties into simulation-based studies.”
L 51 ff: "Per-contrail estimates": what do you mean, estimates of what?
We have clarified this to say “Per-contrail estimates of contrail warming.”
L 85 ff: it is unclear where in this conversion the ERF/RF ratio applies. RF and ERF have units w/m², but if the conversion is J/tonne, I don't see how this matches. I have a conjecture, but please write it down clearly. Further: The (quite uncertain) ERF/RF ratio is a climatological quantity and cannot be applied to evaluate the impact of single flights. I would find it acceptable if the impact evaluation would have been perfomed on an actual instead of a climatological basis, since then it is just a change of units and is transparent for statistical analyses.
The full order of operations is instantaneous radiative forcing -> integration over contrail area and lifetime -> instantaneous energy forcing -> application of RF/ERF factor -> effective energy forcing. The RF/ERF was derived for radiative forcings, but it is unitless and applying it after the conversion from radiative forcing to energy forcing gives the same result as if we had first converted instantaneous radiative forcing to effective radiative forcing, and then converted to effective energy forcing. We have modified the text to make this more explicit.
The conversion is just a change of units, and to make this clear we have added the statement “everywhere units of CO$_2$e are reported in this work they can be converted back to joules by multiplying by the provided factor if desired. “
We have modified the text to make it clearer that the units of CO2e are used only in the forecast and not in the impact evaluation. We have modified the impact evaluation to report instantaneous energy forcings (in joules) rather than effective energy forcings, so there is now no ERF/RF ratio being applied to evaluate the impact of single flights.
Sect. 2.2 first par: as the trial was restricted to eastbound flights which take place predominantly at local night and whose contrails therefore are probably warming, would you recommend to make such a system operational on eastbound flights only, or perhaps to begin the operational phase with eastbound flights?
This is an important question for future work. In the Conclusions section (Sec. 4), we highlight global expansion as a direction for future work.
L 102: "forecasted contrail EF". To make it clear, please state whether this prediction was done with the actual weather or with the climatology mentioned before.
We added a clarification that this is using the forecast described previously in the manuscript. We hope that the responses to the other comments in this review clarify that the prediction was done with the actual weather.
L 380: Please note that a fuel penalty is accompanied whith higher emissions of CO2 and non-CO2 gases and aerosols, and with their corresponding impacts.
We added the following sentence: “Understanding the fuel penalty is critical for assessing the climate impact of contrail avoidance, since additional fuel usage is accompanied with higher emissions and climate impacts of CO2 and non-CO2 gases and aerosols.”.
Appendix B1, L 419: 789 m is more than 2000 ft, that is more than 2 flight levels. Is this altitude definition good enough?
See response below.
Appendix B2, L 432: 1615 m is even more than 5 flight levels. So, same question.
The flight attribution algorithm can work even without any information about the altitude of the contrail, just by comparing the latitudes/longitudes and orientation of the contrail and advected flight path. But in congested airspaces there can be many advected flight paths, so determining which one best matches the observed contrail is challenging. Here we use altitude only to filter out flights whose altitude is very different from the observed contrail. As the reviewer points out, our uncertainty about the altitude is large enough that this doesn’t by itself determine exactly which flight to attribute to the contrail, but it narrows it down to make the task of comparing locations and orientations easier. We have added details to Appendix B clarifying these points.
Citation: https://doi.org/10.5194/jecats-2026-4-AC1
-
AC1: 'Reply on RC1', Tharun Sankar, 22 Jul 2026
-
RC2: 'Comment on jecats-2026-4', Anonymous Referee #2, 30 Jun 2026
General Comments
This manuscript presents a randomized controlled trial evaluating the operational feasibility of dispatcher-led contrail avoidance integrated into routine airline flight planning. The proposed workflow is designed to scale across an airline network and investigates whether contrail avoidance can be implemented within existing operational processes while maintaining operational applicability.
The topic is highly relevant at the intersection of aviation operations and climate impact mitigation. Demonstrating contrail avoidance in an operational airline environment is an important step beyond purely simulation-based studies, and the randomized controlled trial design represents a strong methodological foundation.
However, the manuscript in its current form leaves several critical methodological and scientific aspects insufficiently developed. Key components of the framework—particularly contrail forecasting, satellite-based validation, and aspects of the statistical interpretation—are treated in a way that limits transparency and reproducibility. In several instances, essential model assumptions, uncertainties, and boundary conditions are not sufficiently discussed, making it difficult to fully assess robustness and generalizability of the results. As a consequence, while the operational concept is promising, the scientific underpinning requires substantial strengthening.
Specific Comments
- State of the art and research gap
The literature review largely summarizes prior work without clearly identifying a specific research gap or articulating how the present study advances beyond existing operational or modelling studies. Statements such as the role of warming contrail clusters in climate forcing (e.g. Burkhardt et al., 2018) are correct but remain too general and do not motivate the necessity of the proposed operational trial.
- Contrail forecasting methodology
The contrail forecasting and climate impact modelling approaches are insufficiently described and effectively treated as black-box components. Critical information is missing regarding:
- model structure and formulation
- underlying assumptions
- validation status and known limitations
- applicability to operational trajectory optimization
- expected predictive accuracy
Given that the optimization depends strongly on these forecasts, a more transparent methodological description is essential.
- Atmospheric humidity uncertainty
The manuscript does not adequately address uncertainties in atmospheric humidity fields, which are a dominant source of uncertainty in contrail prediction. Although optimization is performed up to ~24 hours before departure, the limited predictability of ice-supersaturated regions at this lead time is not discussed. This has direct implications for both contrail prediction accuracy and operational decision quality.
- Model assumptions and boundary conditions
Important model assumptions, boundary conditions, operational constraints, and sensitivity of results to these assumptions are not sufficiently described. Without this information, the robustness and transferability of the proposed workflow cannot be properly assessed.
- Spatial resolution and grid effects
The justification for the chosen spatial resolution (“balance between accuracy and compute resources”) is qualitative only. No quantitative analysis is provided regarding the impact of the coarse satellite-based grid on:
- contrail prediction accuracy
- trajectory optimization outcomes
- estimated climate benefits
- overall model sensitivity
A sensitivity analysis would be necessary to support this design choice.
- Validation of satellite-based contrail attribution
The manuscript does not report quantitative validation metrics for the contrail detection and attribution system, nor does it adequately discuss limitations of detecting narrow contrails using coarse-resolution geostationary satellite imagery (GOES). As a result, the reliability and robustness of the observed contrail identification cannot be fully assessed.
- Dependence on underlying model accuracy
Across Sections 2.6 and 2.7, the statistical analysis assumes sufficient accuracy of both contrail forecasting and satellite-based attribution, yet no quantitative assessment of these underlying models is provided. Consequently, the reported statistical significance cannot be interpreted independently of the uncertainties in contrail prediction, meteorological inputs, and detection algorithms.
- Fuel consumption interpretation
The fuel analysis requires a more balanced interpretation. Large differences in unadjusted fuel usage (7–11%) highlight strong aircraft-type confounding and show that treatment and control groups are not directly comparable without adjustment. After correction, differences become very small (<1%) and are not statistically significant for L2 and L3. Therefore, strong conclusions regarding fuel savings should be avoided.
- Confounding factors
Residual confounders such as city pair, weather conditions, and load factor are acknowledged for fuel consumption but are equally relevant for contrail formation. Their potential influence on the primary endpoint should be discussed more systematically.
- Uncertainty in reported reductions
While statistically significant reductions are reported across all treatment groups, confidence intervals—particularly for L1 (2–20%)—remain relatively wide. This indicates substantial uncertainty in the achievable operational reduction and should be more clearly emphasized.
- Interpretation of climate relevance
The manuscript reports reductions in observable contrail distance. However, observable contrail length is only an intermediate proxy for climate impact. Differences in contrail persistence, optical properties, and radiative forcing are not accounted for, and the manuscript should avoid implying direct proportionality between contrail length reduction and climate impact reduction without further justification.
- Forecast calibration claims
The claim that the ML contrail forecast is “well calibrated” is not sufficiently supported. Only qualitative agreement at aggregated level is shown, while no quantitative calibration metrics (e.g., Brier score, ROC/AUC, reliability diagrams) are provided. This weakens the strength of the stated conclusion.
- Causality interpretation of results
The interpretation that overlapping confidence intervals between observed and expected reductions confirm causality is not fully justified. While results are consistent with the hypothesis of contrail avoidance effectiveness, overlapping confidence intervals alone do not demonstrate causal attribution. The results instead indicate consistency within the uncertainties of the modelling framework.
Technical Corrections
- Minor typographical and formatting inconsistencies appear throughout (e.g., spacing and hyphenation in compound terms such as “city-pair”, “contrail-optimized”, and similar variants).
- Some figures and sections referenced in the text (e.g., Appendix A/D/E) should be checked for consistency in labeling and cross-referencing.
- The notation for statistical quantities (e.g., p-values, confidence intervals) should be standardized throughout the manuscript.
- Several instances of informal phrasing (e.g., “black box”, “meets threshold”) should be replaced with more formal scientific terminology.
Citation: https://doi.org/10.5194/jecats-2026-4-RC2 -
AC2: 'Reply on RC2', Tharun Sankar, 22 Jul 2026
Dear reviewer,
Thank you for your review of our manuscript. Please see our responses inline below.
General Comments
This manuscript presents a randomized controlled trial evaluating the operational feasibility of dispatcher-led contrail avoidance integrated into routine airline flight planning. The proposed workflow is designed to scale across an airline network and investigates whether contrail avoidance can be implemented within existing operational processes while maintaining operational applicability.
The topic is highly relevant at the intersection of aviation operations and climate impact mitigation. Demonstrating contrail avoidance in an operational airline environment is an important step beyond purely simulation-based studies, and the randomized controlled trial design represents a strong methodological foundation.
However, the manuscript in its current form leaves several critical methodological and scientific aspects insufficiently developed. Key components of the framework—particularly contrail forecasting, satellite-based validation, and aspects of the statistical interpretation—are treated in a way that limits transparency and reproducibility. In several instances, essential model assumptions, uncertainties, and boundary conditions are not sufficiently discussed, making it difficult to fully assess robustness and generalizability of the results. As a consequence, while the operational concept is promising, the scientific underpinning requires substantial strengthening.
Specific Comments
- State of the art and research gap
The literature review largely summarizes prior work without clearly identifying a specific research gap or articulating how the present study advances beyond existing operational or modelling studies. Statements such as the role of warming contrail clusters in climate forcing (e.g. Burkhardt et al., 2018) are correct but remain too general and do not motivate the necessity of the proposed operational trial.
We respectfully disagree with the reviewer's assertion that the literature review does not identify a research gap or motivate the necessity of the proposed operational trial. The first paragraph in Section 1 discusses the impact of contrails. The next paragraph introduces the idea and merits of navigational avoidance. The third paragraph identifies challenges identified in simulations and the limits of simulations. The fourth paragraph identifies the gaps in previous operational trials. The fifth paragraph states how this trial aims to address many of those gaps.
- Contrail forecasting methodology
The contrail forecasting and climate impact modelling approaches are insufficiently described and effectively treated as black-box components. Critical information is missing regarding:
- model structure and formulation
- underlying assumptions
- validation status and known limitations
- applicability to operational trajectory optimization
- expected predictive accuracy
Given that the optimization depends strongly on these forecasts, a more transparent methodological description is essential.
We have edited the manuscript to provide more details on each of these topics, as described in more detail below. The system studied in this work combines contrail forecasting, operational implementation and validation. The end-to-end performance of the system is the primary result of the work. Though it is interesting to compare the end-to-end performance with the performance of each component individually (and we have done so where possible), measuring the end-to-end performance does not require a detailed evaluation of each component individually. We have added language to the introduction to make it more clear that the primary goal of the study is to measure this end-to-end performance.
- Atmospheric humidity uncertainty
The manuscript does not adequately address uncertainties in atmospheric humidity fields, which are a dominant source of uncertainty in contrail prediction. Although optimization is performed up to ~24 hours before departure, the limited predictability of ice-supersaturated regions at this lead time is not discussed. This has direct implications for both contrail prediction accuracy and operational decision quality.
We have included a new section in the Appendix (Appendix A) with a detailed description of the ML forecast model and its performance to address this comment and the previous one. It includes details about the model construction, performance, and discussion of uncertainties in atmospheric humidity fields and lead time. The ML forecast is also described in Sonabend-W 2024.
- Model assumptions and boundary conditions
Important model assumptions, boundary conditions, operational constraints, and sensitivity of results to these assumptions are not sufficiently described. Without this information, the robustness and transferability of the proposed workflow cannot be properly assessed.
The added Appendix A (on the forecast construction and performance) may address these concerns. Apart from that, it is not clear to us which model assumptions, boundary conditions or operational constraints the referee is talking about here, without more specificity it is not possible to respond in more detail.
Regarding boundary conditions, our ML model makes local adjustments to the ECMWF IFS weather forecast. Therefore the concept of boundary conditions does not really apply to it. The concept of boundary conditions does apply to the underlying weather forecast, however our understanding is that the IFS forecast is a global model. We hope that Appendix A makes this more clear.
- Spatial resolution and grid effects
The justification for the chosen spatial resolution (“balance between accuracy and compute resources”) is qualitative only. No quantitative analysis is provided regarding the impact of the coarse satellite-based grid on:
- contrail prediction accuracy
- trajectory optimization outcomes
- estimated climate benefits
- overall model sensitivity
A sensitivity analysis would be necessary to support this design choice.
We added a citation to Martin Frias 2025 which outlines the software configuration of the grid-based Flightkeys optimizer. Sec. 2.2.1 in that paper contains a discussion of the chosen resolution, which is built into the software and is unadjustable for potential sensitivity analyses.
The purpose of this paper is to assess and report on the end-to-end effectiveness of a given forecast and optimization system within the operational environment. Sensitivity analyses of the parameters influencing the forecast and optimization are outside the scope of this work.
- Validation of satellite-based contrail attribution
The manuscript does not report quantitative validation metrics for the contrail detection and attribution system, nor does it adequately discuss limitations of detecting narrow contrails using coarse-resolution geostationary satellite imagery (GOES). As a result, the reliability and robustness of the observed contrail identification cannot be fully assessed.
A discussion of these limitations is already included in the “Conclusions” section, specifically in the third paragraph. To make it more explicitly, we have added a sentence: “ In particular, the GOES ABI images have a nadir resolution of 2km, which makes it difficult to detect young, narrow contrails.” As mentioned above, the purpose of this work is to assess the end-to-end performance of the forecast, operational, and validation systems. We have added sentences to the introduction and Sec. 2.5 to clarify this intent. We have also added citations to literature which describes the performance of the attribution system in more depth.
- Dependence on underlying model accuracy
Across Sections 2.6 and 2.7, the statistical analysis assumes sufficient accuracy of both contrail forecasting and satellite-based attribution, yet no quantitative assessment of these underlying models is provided. Consequently, the reported statistical significance cannot be interpreted independently of the uncertainties in contrail prediction, meteorological inputs, and detection algorithms.
The statistical analysis in this work does not assume anything about the accuracy of either the contrail forecast or the satellite-based attribution. Rather, the analysis is measuring the accuracy of the combined system. It is true that measuring the combined system does not tell you the performance of the individual components, but the performance of the combined system is an important result in its own right. Measuring the performance of the individual components is out of scope for this work, but to better put our results in context, we have included details on the forecast performance in Appendix A and details on the attribution performance in Sec 2.5.
We additionally added a sentence to the introduction which clarifies that the primary goal of this trial is to test the end-to-end performance of the forecast system, operational implementation, and the validation system.
- Fuel consumption interpretation
The fuel analysis requires a more balanced interpretation. Large differences in unadjusted fuel usage (7–11%) highlight strong aircraft-type confounding and show that treatment and control groups are not directly comparable without adjustment. After correction, differences become very small (<1%) and are not statistically significant for L2 and L3. Therefore, strong conclusions regarding fuel savings should be avoided.
We would appreciate more specificity on which claims about fuel consumption are concerning to the reviewer. The Conclusions section in the second-to-last paragraph contains a discussion of the difficulty encountered in isolating the fuel penalty. We are careful throughout the manuscript to avoid claiming that contrail avoidance causes no difference in fuel consumption; the claim we do make is that we are not able to reject the null hypothesis based on the current sample size.
- Confounding factors
Residual confounders such as city pair, weather conditions, and load factor are acknowledged for fuel consumption but are equally relevant for contrail formation. Their potential influence on the primary endpoint should be discussed more systematically.
These issues absolutely could also apply to contrail formation. For this reason, we did the counterfactual analysis of Sec 3.4. Because this analysis is using the same flights, confounders like city pairs, weather conditions and load factors are controlled for. We see that the counterfactual analysis gives results which are slightly smaller than the treatment-controlled analysis, though the difference is within our uncertainty estimates. This suggests that the confounders may also be impacting contrail formation, but their effect is very small compared to the observed differences between the treatment and control groups.
We have added a new paragraph to the start of Sec 3.4 to make this more clear.
- Uncertainty in reported reductions
While statistically significant reductions are reported across all treatment groups, confidence intervals—particularly for L1 (2–20%)—remain relatively wide. This indicates substantial uncertainty in the achievable operational reduction and should be more clearly emphasized.
We added a language to the L1 paragraph in the Conclusions section to emphasize this point.
- Interpretation of climate relevance
The manuscript reports reductions in observable contrail distance. However, observable contrail length is only an intermediate proxy for climate impact. Differences in contrail persistence, optical properties, and radiative forcing are not accounted for, and the manuscript should avoid implying direct proportionality between contrail length reduction and climate impact reduction without further justification.
We have included an additional estimate of the climate impact in the results to address this concern, which strengthens the claims that the trial also likely achieved a climate impact reduction.
- Forecast calibration claims
The claim that the ML contrail forecast is “well calibrated” is not sufficiently supported. Only qualitative agreement at aggregated level is shown, while no quantitative calibration metrics (e.g., Brier score, ROC/AUC, reliability diagrams) are provided. This weakens the strength of the stated conclusion.
We updated the language to use “agrees with” instead of “calibrated,” as our goal is simply to measure that the average contrail distance is the same between the forecast and observations when averaged over enough flights.
- Causality interpretation of results
The interpretation that overlapping confidence intervals between observed and expected reductions confirm causality is not fully justified. While results are consistent with the hypothesis of contrail avoidance effectiveness, overlapping confidence intervals alone do not demonstrate causal attribution. The results instead indicate consistency within the uncertainties of the modelling framework.
We do not see where the reviewer finds us interpreting overlapping confidence intervals as confirming causality. In subsection 3.4, we state that the error bars in the counterfactual reduction and the observed reduction overlap, as a way to quantify precision and uncertainty between the two methods. However, this is not asserting that the overlap confirms causality. We have adjusted some language in subsection 3.4 to clarify that we are only looking for consistency of evidence, rather than claiming causality.
Technical Corrections
- Minor typographical and formatting inconsistencies appear throughout (e.g., spacing and hyphenation in compound terms such as “city-pair”, “contrail-optimized”, and similar variants).
We did not find any such inconsistencies in these compound terms (“city pair”, “contrail-optimized”, “non-avoidance”).
- Some figures and sections referenced in the text (e.g., Appendix A/D/E) should be checked for consistency in labeling and cross-referencing.
We performed a pass through all of the references for consistency.
- The notation for statistical quantities (e.g., p-values, confidence intervals) should be standardized throughout the manuscript.
We performed a pass through all of the notation for consistency.
- Several instances of informal phrasing (e.g., “black box”, “meets threshold”) should be replaced with more formal scientific terminology.
We did not find any instances of “black box” or “black-box” in the manuscript, this term was only introduced earlier in the reviewer’s comments. We have replaced “meets the threshold for statistical significance” with “achieves statistical significance” in the instances where it appears.
Citation: https://doi.org/10.5194/jecats-2026-4-AC2
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 214 | 73 | 31 | 318 | 32 | 42 |
- HTML: 214
- PDF: 73
- XML: 31
- Total: 318
- BibTeX: 32
- EndNote: 42
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Review of
Efficacy of scalable airline-led contrail avoidance
by T. Sankar et al., JECATS-2026-4
General comments
Contrail avoidance is often seen as a possibility to quickly lessen the rate at which aviation's climate warming contribution rises. For this reason it is important to conduct tests that occur not only in computer simulations but as well in the operational environment. A few such trials did occur in the recent past but they involved only few flights to evaluate, and the planning and evaluation was done manually which takes too much time for an operational system. The trial described in this manuscript intends to overcome some of these problems, the much manual work and the low number of flights for evaluation which is a step forward. But the paper also demonstrates that there are still difficulties and room for improvements, issues that must be tackled before contrail-avoidance flight-planning can actually become operational. These issues are discussed in Sect. 4, and this is the best part of the manuscript.
Otherwise, I admit that it was difficult for me to understand how this trial was operated although the authors tried to explain it. Things and notions are not clear to me. Thus I recommend to work especially on an exact description of what has been done, explaining all the steps.
I have also a problem with the climatological evaluation of the trial. The evaluation of how much contrail coverage was avoided is ok, but the evaluation of how much forcing or climate impact was avoided is not. Since this evaluation is based on a climatology, I believe that the results are more or less random, and that therefore the statistical evaluation ("statistically significant...") does not say anything for the real world. Please consider that the actual weather is the most dominant influencing factor on contrail radiative impacts. For this reason, I think, the ideas of conducting contrail-avoidance in a "climatological sense" have been abandoned about 20 years ago. I hope that this is acknowledged in the revised version.
Major comments:
1) I had problems to understand how this trial had been performed. I have a lot of questions for clarification:
1.1) There is a control group and an intent-to-treat group. There are also "non-avoidance" plans and "contrail-optimized" plans. It is unclear how these two pairs of notions are related. Is the control group equivalent to the group with the "non-avoidane" plans? Still another expression is "per-protocol flights". I am confuesed by this multiplicity of notions.
1.2) Flights with a "non-avoidance" plan were re-optimized including the contrail cost term. Why and for what purpose?
1.3) Why and for what purpose were dispatchers presented with one or more non-avoidance plans? Is it a common process to generate several flight plans for the dispatcher to select one?
1.4) Initially I thought Flightkeys computes an optimal route, either cost or contrail-optimized, and that this is given to the dispatcher. I thought that neither the dispatcher nor the pilot knows on what kind of optimization the flight was planned. But this impression later turns out wrong.
1.5) Please define the L1,2,3 groups with more details. My questions were: Does the L1 group contain routes for which no (strong) contrail was forecasted? If so, then the L2 group is the one where contrails were forecasts, but not all of them have been avoided when the dispatcher had other priorities? And the, L3 is the group where contrails have been avoided acitvely by dispatcher and pilots. Is this correct?
2.) As stated above, I find the evaluation of the contrail impact (or the avoided impact), based on climatological fields, insufficient. I wonder whether the derived values have any significance. You can say, if we assume that the actual weather were like the climatology, the results were x and y, and the difference were statistically significant. BUT, the weather is generally not like the climatology on a certain day, and thus we must expect large deviations of the actual forcings from the obtained ones. As these values are not known, their differences cannot be determined. To my opinion, the corresponding tables and statements should be removed from the paper in order not no rise wrong expectations in uncautious readership.
Minor comments:
1.) I suggest to spend a bit more time to explain a bit how the ML-contrail forecast works. A short overview is sufficient. This is for the convenience of the reader, who otherwise has to change to the paper by Sonabend-W (and in fact, I did not remember that a description of the ML-based contrail forecast was given there, but that may be my fault). I wonder how satellite images of contrails or series of images can be used for contrail forecasting. The contrails in the images already exist and stem from past flights where contrail formation has not been avoided. Please help the reader with a short text to an imagination how this system works.
2.) The trial took place mainly over the western part of the Atlantic ocean where flight density is probably low. Do the authors think that the methods investigated in this trial can be used in similar way in more congested regions like continental US or even Europe?
3.) "The results provide strong evidence... physically effective and operationally feasible." This is an unclear statement. I believe that contrail avoidance is operationally feasible, under certain circumstances, as you show in the paper, namely when the dispatcher does not have other problems to solve. Whether it is feasible in congested air-spaces remains doubtful. What do you mean with "physically effective"? If you mean the reduction of contrail coverage, please say so. If you however mean the calculated reduction of warming impact, I do not agree unless you change your method from a climatologically based one to a actual-weather based one. The expression "climatological warming" makes no sense for single flights.
Miscellaneous:
L 29-31: One should remark here that at least in the paper by Frias et al. a perfect weather forecast was assumed which renders results doubtful.
L 51 ff: "Per-contrail estimates": what do you mean, estimates of what?
L 85 ff: it is unclear where in this conversion the ERF/RF ratio applies. RF and ERF have units w/m², but if the conversion is J/tonne, I don't see how this matches. I have a conjecture, but please write it down clearly. Further: The (quite uncertain) ERF/RF ratio is a climatological quantity and cannot be applied to evaluate the impact of single flights. I would find it acceptable if the impact evaluation would have been perfomed on an actual instead of a climatological basis, since then it is just a change of units and is transparent for statistical analyses.
Sect. 2.2 first par: as the trial was restricted to eastbound flights which take place predominantly at local night and whose contrails therefore are probably warming, would you recommend to make such a system operational on eastbound flights only, or perhaps to begin the operational phase with eastbound flights?
L 102: "forecasted contrail EF". To make it clear, please state whether this prediction was done with the actual weather or with the climatology mentioned before.
L 380: Please note that a fuel penalty is accompanied whith higher emissions of CO2 and non-CO2 gases and aerosols, and with their corresponding impacts.
Appendix B1, L 419: 789 m is more than 2000 ft, that is more than 2 flight levels. Is this altitude definition good enough?
Appendix B2, L 432: 1615 m is even more than 5 flight levels. So, same question.