Aviation Safety Paper on Operations Under Meteorological Hazard with Early Detection
10218 IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, VOL. 57, NO. 12, DECEMBER 2019
Applying Deep Learning to Hail Detection: A Case Study
Melinda Pullman , Iksha Gurung, Manil Maskey, Member, IEEE, Rahul Ramachandran, Senior Member, IEEE, and Sundar A. Christopher
Abstract— Deep learning is a subset of machine learning that uses deep neural networks (DNNs) capable of learning representations and extracting valuable information from vast data sets. Similarly, weather phenomena are often identified by patterns in data that serve as precursor signatures. Therefore, deep learning networks can be used to identify signatures of the weather phenomena, or possibly signatures not yet established by forecasters in addition to aiding forecasters in synthesizing the growing amount of meteorological observations. In this article, we demonstrate the value of deep learning for atmospheric science applications by providing a proof of concept, using deep learning for the detection of hail-bearing storms as a test case study. The deep learning network presented in this article obtains a higher precision when presented with multisource data and is able to identify a common feature associated with hail storms— decreased infrared brightness temperatures. This network and case study illustrate the capability of deep networks for the detection of weather phenomena and contribute to the growing awareness of deep learning among atmospheric scientists.
Index Terms— Artificial intelligence, event detection, neural networks.
I. INTRODUCTION
ADIGITAL era has dawned; one in which data are now theworld’s most valuable resource and serve as the engine for artificial intelligence techniques that allow computers to mimic the human learning process and extract valuable infor- mation from vast amounts of data. Technology giants like Amazon and Google increase their revenues by using artificial intelligence techniques to improve their services and products. In fact, these companies and other technology titans have used information obtained through artificial intelligence to improve their digital advertising, and as a result, have amassed a net profit of over U.S $25 billion in the first quarter of 2017 [1]. In recent years, artificial intelligence techniques have evolved
Manuscript received May 31, 2018; revised April 1, 2019; accepted June 17, 2019. Date of publication August 28, 2019; date of current version November 25, 2019. This work was supported in part by NASA under Grant NNM11AA01A. (Corresponding author: Melinda Pullman.)
M. Pullman is with the U.S. Army Corps of Engineers Vicksburg District, Vicksburg, MS 39183 USA (e-mail: [email protected]).
I. Gurung is with Inter Agency Implementation and Advanced Concepts (IMPACT), NASA Marshall Space Flight Center, Huntsville, AL 35812 USA, and also with the Earth System Science Center, University of Alabama in Huntsville, Huntsville, AL 35805 USA (e-mail: [email protected]).
M. Maskey and R. Ramachandran are with the NASA Marshall Space Flight Center, Huntsville, AL 35812 USA (e-mail: [email protected]; [email protected]).
S. A. Christopher is with the Department of Atmospheric Science, Uni- versity of Alabama in Huntsville, Huntsville, AL 35805 USA (e-mail: [email protected]).
Digital Object Identifier 10.1109/TGRS.2019.2931944
to incorporate machine learning and deep learning networks. Machine learning utilizes statistical methods that enable machines or computers to learn when iteratively presented with more data or experiences. On the other hand, deep learning is a subset of machine learning that uses neural networks capable of learning representations or patterns within data sets to add value to large amounts of data [2].
The value and use of deep learning for the technology indus- try is unmistakable, but why should we explore deep learning for weather prediction? First and foremost, forecasters recog- nize and rely upon patterns within meteorological observations to serve as indicators of impending weather phenomena [3]. Similarly, deep learning networks learn data representations and patterns, and, thus, are capable of learning precursor signatures of weather patterns and possibly signatures not yet established by forecasters. Second, weather prediction has become a big data task as meteorological observations are gathered more frequently from the growing number of in situ and remote sensors. For instance, numerical weather prediction models and state-of-the-art satellites, such as the Geostationary Operational Environmental Satellite (GOES) 16, are now capable of operating at higher spatial and temporal resolutions, multiplying the amount of available data that forecasters must synthesize [4]. Utilizing these deep learning networks that can rapidly extract valuable information from large amounts of data would aid forecasters in examining the growing amount of meteorological observations. Finally, and possibly, more importantly, using deep learning networks for weather prediction would benefit regions across the world that may not have the resources to develop sound weather prediction systems with trained professionals.
Fortunately, recent advances in deep learning have demon- strated the application of atmospheric science. For instance, deep networks have been applied to various areas of weather detection and forecasting, including improving short-range weather prediction from radar images [5]; detecting extreme weather events in climate reanalysis data sets [6]; and identify- ing convection initiation in radar and reanalysis data sets [7]. Although deep learning has been used to identify various types of severe weather, deep learning applications for the detection of other weather phenomena, such as hail, are sparse or absent from the literature altogether.
Hail storms have traditionally been identified by criti- cal radar reflectivities [8] or decrease in satellite infrared brightness temperatures [9], both of which are patterns based on radiative transfer analysis, which, machines are capable
0196-2892 © 2019 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission. See http://www.ieee.org/publications_standards/publications/rights/index.html for more information.
PULLMAN et al.: APPLYING DEEP LEARNING TO HAIL DETECTION 10219
of learning. In fact, hail-bearing storms have been identi- fied with other machine learning techniques. For instance, Marzban and Witt [7] constructed two Bayesian neural networks from radar imagery that applied both a regres- sion and classification scheme to estimate hail sizes. The Bayesian neural networks were used to predict hail sizes from radar imagery that corresponded to severe-hail reports from 81storms across the United States. Lu et al. [10], [11] used image mining and time series association rules to detect hail storms from radar images. Similarly, Gagne et al. [12] employed machine learning models to identify potential hail storms and predict hail sizes from numerical weather predic- tion model outputs. Hail probabilities and sizes were forecast for more specific times and areas and were verified using hail reports for 12 hail days over the Midwestern United States. Although these studies concluded that the machine learning techniques were able to produce more accurate and reliable hail forecasts, none of them aimed to detect hail storms at a global or even continental scale. Lu et al. [10], [11] utilized radar imagery to identify hail storms and predict hail sizes. Using only radar imagery to detect hail storms limits the amount of data deep networks can learn by providing meteorological observations at smaller spatial coverages com- pared to global observations. In addition, Gagne et al. [12] investigated how numerical weather prediction model output can be combined with machine learning models to predict hail potential over more specific areas.
Leveraging meteorological observations available on a global or continental scale, such as satellite imagery, would provide hail detection capabilities in areas that do not have extensive radar networks established. These observations are also capable of providing hail detection capabilities regardless of the time of day. However, detecting hail through the visual interpretation of satellite images is a big data task because satellites provide tens of terabytes of continuous global data annually [6]. Because deep learning networks can easily identify representations of hail given massive amounts of data, it would be advantageous to utilize these types of algorithms for hail detection. Thus, this article will explore the novel approach of using deep learning for the detec- tion of hail-bearing storms on a continental scale. Our goal is to illustrate the value of deep learning for atmospheric science applications, by providing a proof of concept, with the identification of hail storms as a test case study. The test case study will explore hail detection given multisource meteorological observations through the use of a deep network that concatenates a convolutional neural network (CNN) and deep neural network (DNN). The CNN will automate the visual interpretation of GOES satellite imagery; whereas the DNN will receive 2-D Modern-Era Retrospective analysis for Research and Applications, Version 2 (MERRA-2) reanalysis parameters. A merged architecture will be utilized because it effectively integrates data of varying dimensions and illustrates how well computers learn data representations when presented with multisource data. This case study will also examine how well the CNN and DNN perform operating on an individual basis to determine if deep networks can improve their perfor- mance when presented with data from multiple sources.
II. TEST CASE
Our test case uses deep learning to add value or confidence to meteorological observations for the detection of hail storms. The proof of concept will explore hail detection from single- source and multi-source data and will utilize various deep network architectures to learn: 1) the spatial features and patterns of GOES satellite infrared brightness temperatures for hail-bearing storms and 2) MERRA-2 surface and mid- level atmospheric properties of hail storms. The deep networks will learn hail storm signatures from January 2011 through October 2017 and will examine hail-bearing storms across the contiguous United States to generalize the model for detecting hail on a continental scale.
A. GOES Satellite Imagery
GOES-13 and GOES-15 satellite composites for the contiguous United States from January 2011 through October 2017 are collected from the Iowa Environmental Mesonet. These pseudocolor composites are obtained from the GOES Imager, a multispectral sensor containing five channels that provide data across a range of wavelengths in the electromagnetic spectrum. For our purpose, we utilized the composites of channel 3 (6.5–7.0 µm) and channel 4 (10.2–11.2 µm), which are two infrared channels that acquire data regardless of the time of day and can advantageously identify hailstorms even during the nighttime hours. However, these channels have a coarser spatial resolution of 4 km, compared to the visible band on the GOES Imager with a 1-km spatial resolution, reducing the amount of informa- tion that our deep networks can learn from these images. Despite the coarse spatial resolutions, geostationary satellites provide coverage across the entire globe and make it pos- sible to develop a hail detection algorithm for areas lack- ing hail detection resources and personnel. Thus, these two channels were selected to serve as input data for our deep networks because they are commonly used by meteorologists for examining midtropospheric water vapor and cloud top temperatures of convective storms, both of which are also important for hail storm studies [13]. Overall, we obtained a total of 37 236 GOES channel 3 and channel 4 images that were used for training, validating, and testing our deep networks.
B. MERRA-2 Atmospheric Parameters
MERRA-2 provides NASA atmospheric reanalysis data across a global spatially interpolated latitude-longitude grid and 72 vertical atmospheric layers [14]. The MERRA-2 data collection M2I3NVASM (DOI: 10.5067/WWQSXQ8IVFW8) contains instantaneous 3-D assimilated meteorological fields every 3 h and will be used to derive parameters that demon- strate atmospheric conditions at the time of the correspond- ing hail reports. The atmospheric parameters derived from MERRA-2 for January 2011 through October 2017 are con- vective available potential energy (CAPE), 0–6 km wind shear, freezing level, and the 700–500 hectopascals (hPa) lapse rate. These four atmospheric parameters were included in our networks because instability measures and wind profile charac- teristics play an important role in hail-producing environments
10220 IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, VOL. 57, NO. 12, DECEMBER 2019
and have been frequently utilized to identify environments supportive of deep moist convection and hail. These four parameters and their influence on hailstorms will be described in detail in the following paragraphs.
CAPE is the amount of thermodynamic energy that the atmosphere has to lift a parcel and is an indicator of updraft strength or convective instability, where higher CAPE values suggest stronger convection and greater updraft velocities [15]. Larger CAPE values and greater updraft velocities are more favorable for hail development because stronger updrafts can suspend ice particles long enough for hail growth to occur [16]–[19]. Johnson and Sugden [19] examined how thermodynamic parameters, such as CAPE, relate to hail size. They discovered that CAPE tends to increase, as binned hail size also increases.
In addition to CAPE, variations in storm-scale wind structures, or wind shear is another sufficient indicator for hail development [18], [19]. Low-level vertical wind shear, in particular, determines the longevity of convective updrafts. When greater low-level vertical wind shear exists, a storm’s structure becomes tilted, and the updraft is sep- arated from the downdraft, allowing the storm to have a longer life cycle. For instance, a single-cell thunderstorm lasts roughly 20–30 min, but a supercell thunderstorm, which has a pronounced separated updraft and downdraft, can last up to 60 min and has a higher potential for bearing hail [16].
The freezing level is another predictor for hail potential [20]. Billet et al. [20] used logistic regression equations to relate the freezing level and the probability of large hail. They concluded that the logistic regression equation was capable of predicting the probability of large hail and obtained a critical success index (CSI) of 0.59. Kitzmiller and Breidenbach [21] found the height of the freezing level is inversely correlated with hail potential. In other words, when the freezing level is lowered, exposure to melting would be reduced prior to hail reaching the surface. Thus, the probability of hail occurring would increase with lower freezing levels [15].
Finally, the 700–500-hPa lapse rate will be derived from the MERRA-2 data set. The lapse rate describes the change in temperature with a height between 700 mb and 500 hPa. It has been observed that steeper lapse rates, with rapid declines in temperature, are correlated with significant severe weather events due to the enhanced convection from the surface [19].
Overall, these four atmospheric parameters were included in our network because instability measures and wind profile characteristics play an important role in hail-producing envi- ronments. In addition, previous research and the Storm Pre- diction Center (SPC) have frequently utilized these parameters as indicators or tools for hail detection [20], [19]. However, while the aforementioned variables are commonly used by operational forecasters and researchers alike to characterize hail-producing environments, they do not guarantee that hail will be observed at the surface. Microphysical processes within the cloud also play an important role but are not observable by forecasters in near-real-time. Instead, this article focuses on the spatial patterns of mesoscale processes that occur on a spatial scale of 10–1000 km and can be resolved at the spatial
and temporal resolution of remotely sensed and reanalysis data sets.
C. Hail Storm Reports
In order for a deep network to learn the patterns associated with hail, a large labeled data set of hail storms must be created. Therefore, reports of hail events obtained from the National Centers for Environmental Information (NCEI) Storm Events Database, for January 2011 through October 2017, will serve as the ground truth data used for classifying, or labeling, satellite imagery, and MERRA-2 parameters in our deep network. The Storm Events Database contains reports entered by the National Weather Service that provide the occurrence time and geographic location of various storm events. Hail reports are generated when hail sizes are equal to or greater than three-fourth of an inch. In addition, hail accumulation of smaller sizes will be entered when property or crop damage or casualties are reported [22].
D. Data Preprocessing
Before obtaining the GOES imagery or MERRA-2 para- meters that correspond to hail reports, duplicate hail reports were removed. Duplicate hail reports were defined as reports that occurred within 5 minutes (± 5) and within about 0.2◦ of latitude and longitude of other hail reports. Then GOES images were downloaded for non-duplicate hail reports, color-enhanced to aid in satellite interpretation of features of interest within the imagery, and cropped to a size of 256 pixels × 256 pixels (22 km × 22 km), with the center of the cropped images corresponding to the location of the hail report. Finally, the GOES satellite imageries for each of the infrared channels are all merged. Similarly, the MERRA-2 atmospheric parameters were derived using the hail reports once duplicate hail reports were removed. After pre-processing the data, we obtained a total of 18 618 merged GOES satellite images and corresponding MERRA-2 parameters that were then randomly divided into three data sets with a ratio of 7:2:1 for training, validation, and testing, respectively.
E. Deep Learning Networks
As previously stated, deep learning refers to a subset of machine learning algorithms for artificial intelligence that consist of multiple, deep layers, allowing these networks to model complex, nonlinear, dynamical phenomenon. There are two types of networks that are generally classified as deep learning networks: DNNs and CNNs. DNNs are biologically inspired models that learn patterns from observational data by receiving input data and transforming it with nonlinear functions through a series of hidden layers before producing class scores and a loss function as an output [23]. The class scores indicate the probability that the input data belong to each class; whereas, the loss function is the error associated with the model predictions. The loss function is derived by comparing a DNN’s prediction with the actual outcome and will be high if the prediction is a poor representation of the true data. The overall objective of learning for the DNN is to minimize the loss function or the error of the model, and
PULLMAN et al.: APPLYING DEEP LEARNING TO HAIL DETECTION 10221
thus increase the model performance [3]. Reducing the loss function can be achieved by updating weights in the DNN, which interconnect the nodes, or neurons, of the hidden layers.
Compared to traditional artificial neural networks, CNNs are ideal for computer visualization tasks because they constrain their architecture to mirror the 3-D spatial structure of images, consisting of width, height, and depth [23]. This allows CNNs to preserve spatial relationships that are critical for learning features in images. CNNs are capable of learning spatial features through the use of several hierarchical convolutional layers that extract features for learning with the use of filters, where early convolutional layers learn simple features, such as edges, and later convolutional layers learn more com- plex features, such as shapes. The convolutional layers are then followed by a small amount of fully connected layers that compute the class scores, corresponding to the various classes [6]. Thus, a CNN essentially performs two major tasks: finding patterns in images with feature learning and classifying the images based on the visual patterns detected. Similar to DNNs, CNNs will also output a loss function associated with its computed class scores and update its weights in order to reduce the loss function.
Because CNNs are ideal for computer vision tasks with 3-D data, this case study will utilize a CNN to explore hail detection with GOES satellite imagery. The architecture of the CNN used for this study is loosely inspired by the AlexNet developed by [24]. Our design consists of four convolutional layers; each proceeded by a zero-padding layer and followed by a max-pooling layer. The zero-padding layer helps to preserve information with the original input dimensions so that convolution layers can extract as many possible features with the use of filters [23]. Each convolutional layer contains rectified linear unit (ReLU) activation functions that introduce non-linearities into the deep CNN, allowing it to model non- linear features and perform faster learning [6]. Then, pooling layers downscale the size of the convolutional output feature maps to help control overfitting of the CNN and improve the overall computational performance [23]. The second convolu- tional layer is followed by a batch normalization layer, which further normalizes the CNN’s output and speeds up the training of the CNN. Then, one normalization layer is utilized after the second convolutional layer to further reduce overfitting of the model by standardizing the neuron values. The convolutional layers are also followed by three fully connected layers that return probability units, or class scores, for each classifica- tion category, “Hail” and “No Hail.” The first, second, and third fully connected layers have a total of 3072, 1024, and 20 neurons, respectively.
Whereas CNNs are advantageous when classifying 3-D data, such as imagery, DNNs can rapidly learn non-linear data representations for classifying 2-D data, such as atmospheric parameters. Therefore, a simple DNN architecture, consisting of four fully connected layers, was used to detect hail when presented with MERRA-2 parameters. The fully connected layers use mathematical functions and weights to translate the input MERRA-2 data into class scores. Similar to the CNN, each fully connected layer in the DNN has ReLU activation functions due to its ability to result in faster learning [24].
Fig. 1. Merged deep learning network, where F.C. represents fully connected layers.
As a result of all four fully connected layers, the DNN has a total of 220 neurons that aid in returning two class scores: “Hail” and “No Hail.”
To develop a network that can detect hail from multi-source data, we coupled the DNN and CNN described above. As the merged GOES imagery are fed into the CNN, the MERRA-2 data are introduced into the four-layer neural network. Then, the third fully connected layer from the CNN and the fourth hidden layer from the DNN are merged using the “concatenate” method in Keras [25]. The “concatenate” method takes in the two output arrays from the CNN and DNN and merges them, where the first element of the second input array follows the last element of the first input array. Finally, two fully connected layers are introduced into the merged network (Fig. 1). The first fully connected layer uses a ReLU activation function, similar to the CNN. However, the second fully connected layer uses a Softmax function, which is used for classification purposes as it calculates the probability for every class [23]. Because our goal is to detect hail, there will be two classes: “Hail” and “No Hail” returned from the merged CNN and DNN.
A merged network was chosen because it effectively inte- grates data of varying dimensions and can be used to assess if networks perform better when presented with data from mul- tiple sources. For instance, in this case study, the merged net- work receives both the GOES satellite imagery and MERRA-2 atmospheric parameters that are capable of illustrating sur- face, mid-level, and upper-level characteristics of hail storms. By providing the merged network with various atmospheric
10222 IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, VOL. 57, NO. 12, DECEMBER 2019
Fig. 2. Three phases of deep learning that illustrate how a deep network learns features through training and validation before being tested on new data.
properties, the network can potentially develop a more com- prehensive understanding of hail storms that are needed to improve its performance.
F. Three Experimental Phases of Deep Learning
For developing a deep network capable of detecting hail storms within satellite imagery and MERRA-2 reanalysis parameters, the network must go through three experimental phases, consisting of training, validation, and testing (Fig. 2). Each of the three experimental phases is assigned a subset of the original satellite imagery and MERRA-2 data. During training, the network was presented the training data and the associated labels “Hail” and “No Hail” to learn the characteristics of hail-bearing storms from satellite images and MERRA-2 atmospheric parameters. The training input data were passed through the network during a forward pass, and the predicted class scores are calculated. The network then used the “Hail” or “No Hail” labels to calculate the loss and accuracy associated with its predicted scores. Finally, backward passes were used to update the weights within the network until the loss associated with the predicted scores was minimized.
During validation, the network received input data from the validation data set to verify whether or not the network has learned. Unlike the training phase, the network was not given immediate access to the hail reports. Instead, the network was required to make a prediction without the help of labels, which illustrated its progress of learning. Validation accuracies were calculated each time the network completed a pass, or epoch, over the validation data set. The validation accuracies were then monitored over epochs and compared to the training accuracies. If the training accuracies were to increase and the validation accuracies were to decrease or remain unchanged with epochs, then training was ceased to prevent overfitting. For reducing the overfitting of the deep network, optimization was manually conducted by tuning hyperparameters, or net- work parameters that are set before training begins, between validation epochs. Hyperparameters that were altered were the number of hidden layers and filters used in the network and the learning rate. The network then continued through the validation epochs, with the altered hyperparameters, until the
Fig. 3. Change in training and validation accuracies with each epoch.
validation accuracies were similar to the training accuracies and the issue of overfitting in the model was reduced.
After the learnable parameters and hyperparameters are validated, the model with the highest validation accuracy is tested on a subset of input data that are different from the data used during training or validation. The purpose of testing the network on unseen data is to evaluate the accuracy and performance of the network on new data. Thus, the evaluation results in this article were generated using the test data set.
G. Results
Before assessing the overall test performance of our approach, it is vital to determine the effectiveness of our network’s learning. In this article, we concluded the training and validation of our network after 74 epochs, where epochs are training cycles marking each time the weights of the deep network are updated. Fig. 3 illustrates the change in accuracy with each epoch for both the training and validation of our network. At the 74th epoch, the final training accuracy was 0.754, the validation accuracy was 0.734, and the difference between the two accuracies was 0.02. Because the validation accuracy is relatively close to the training accuracy, the model did not depend solely upon the training data set for learning; and thus, does not suffer from overfitting as it is capable of generalizing to both data sets. In addition, while the training and validation accuracies increase, the loss of the network is reduced over epochs, suggesting an improvement in model performance (Fig. 3).
Verification metrics are commonly used that determine the correspondence between forecasts and observations to assess the accuracy of the weather forecasts. Verification metrics, such as the probability of detection (POD), false alarm ratio (FAR), and CSI have been previously implemented in severe weather literature to verify the occurrence of severe weather, such as hail [7], [8], [16], [27]. These verification metrics are computed from contingency tables based upon possible prediction or forecast outcomes and are described in full in [26]. Contingency tables display four types of outcomes: test images were correctly classified when hail was present [true-positive (TP)]; test images were correctly classified when hail was not present [true-negative (TN)];
PULLMAN et al.: APPLYING DEEP LEARNING TO HAIL DETECTION 10223
test images were incorrectly classified when hail was present [false-negative (FN)]; and test images were incorrectly classi- fied when hail was not present [false-positive (FP)]. Verifica- tion metrics are then derived using the four possible outcomes displayed in the contingency table. For the assessment of our networks’ performance, the following verification metrics are derived: precision, POD, FAR, and CSI.
The precision of the network is expressed as
Precision = TP TP + FP . (1)
The precision represents the ratio of correctly classified hail events over the total correct classifications. A precision closer to 1 indicates that every hail event was correctly classified by our network.
The POD is defined by
POD = TP TP + FN . (2)
The POD represents the ratio of correctly classified hail events to the total number of hail events observed, indicating the fre- quency of hail being forecast when the hail was observed [26]. If test images containing hail were always correctly classified, the POD will be equal to 1. On the other hand, if all test images containing hail are incorrectly classified, the POD will be equal to 0.
The FAR is the fraction of incorrect hail classifications over the total number of hail events forecast. The FAR, in other words, is a measure of the reliability of the networks and is represented as
FAR = FP TP + FP . (3)
FAR values closer to zero indicate less false alarms, or hail events forecast when the hail was not observed and suggested greater model performance.
Finally, the CSI combines the POD and FAR into one score and implies the confidence level of using the trained network to detect hail storms [26]
CSI = TP TP + FP + FN . (4)
The CSI is the ratio of hail events correctly classified given that the event was either forecast, observed, or both [26]. A perfect forecasting model will have a CSI of 1, meaning that all events are forecast correctly.
If the deep network is successfully detecting hail-bearing storms, the network will have a precision, POD, and CSI closer to 1 and a FAR closer to 0. From our case study, the merged, or multisource, deep network achieved a precision, POD, FAR, and CSI of 0.732, 0.665, 0.268, and 0.536, respectively, and exhibits relatively good precision, indicating its capability of detecting hail storms (Table I). Both the CNN and DNN obtained almost identical verification metrics but have relatively low precisions compared to the merged network (Table I). This indicates that 1) the merged network is relying equally upon the CNN and DNN during the learning process and 2) the merged network outperforms the individual net- works, suggesting that integrating data from multiple sources can be advantageous for detecting hail-bearing storms.
TABLE I
VERI FICATI ON METRI CS F OR DNNS
Fig. 4. Sample GOES imager channel 3 (a) report 1, (b) report 2, (c) report 3 and channel 4 (d) report 1, (e) report 2, (f) report 3 test images, for three hail reports, correctly classified as Hail from our trained deep network.
For meteorologists, it is often advantageous to understand the physical process or mechanisms relating precursor signa- tures to weather phenomena. Therefore, understanding what deep networks are learning and why it would be valuable. In fact, we can try to identify the particular features that our merged deep network has learned by comparing the network’s final predictions with the original input data. Fig. 4 shows images correctly classified as “Hail.” Given Fig. 4, it is evident that the network is recognizing lower infrared brightness temperatures (<220 K) for both GOES channel 3 and 4, as a potential indicator of hail. Despite this attempt to identify the features that the deep network is learning for hail storms, it is impossible to determine the patterns that machines are learning with absolute certainty. Deep learning networks, such as DNNs and CNNs are complex structures that are capable of learning millions of parameters. These networks receive input data and rely upon actual outcomes to formulate a function that relates the two. However, these networks are not explicitly programmed, and the nonlinear function they develop is not inherently clear. For previous atmospheric science applications, these networks have been likened to “black boxes” [27]. Instead, the acceptance of these deep networks and their performance come from how well they have learned patterns associated with weather phenomena [3].
10224 IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, VOL. 57, NO. 12, DECEMBER 2019
III. CONCLUSION
This article contributes to the growing awareness of deep learning among atmospheric scientists by presenting an appli- cation to near real-time hail detection. Deep learning tech- niques operate with data as their engine and can extract valuable information from vast amounts of data. The dispo- sition to avoid such techniques in atmospheric science may result from the existing stigma that forecasters should fully understand the physical processes and mechanisms leading to weather phenomena. It takes years or even decades for one to understand how the atmosphere behaves and exactly why it behaves the way it does. Trusting a computer to rapidly learn patterns without being able to communicate why or how it is learning those patterns can be off-putting to forecasters. However, previous applications of deep learning networks in atmospheric science have shown these networks can recognize features commonly associated with weather events. Even in the case study presented in this article, the deep network was able to identify a commonly used precursor for hail storms, decreased infrared brightness temperatures, and obtains greater performance metrics when presented with data from multiple sources. Results that indicate deep networks can learn indicators of weather phenomena are promising, and as technology continues to advance and weather observations become more frequent, the value of using such techniques can no longer be ignored.
REFERENCES
[1] The Economist. (May 2017). The World’s Most Valuable Resource Is no Longer Oil, But Data. Accessed: Jan. 11, 2018. [Online]. Available: https://www.economist.com/news/leaders/21721656-data-economy- demands-new-approach-antitrust-rules-worlds-most-valuable-resource
[2] A. Mcgovern et al., “Using artificial intelligence to improve real-time decision-making for high-impact weather,” Bull. Amer. Meteorol. Soc., vol. 98, no. 10, pp. 2073–2090, Oct. 2017.
[3] M. W. Gardner and S. R. Dorling, “Artificial neural networks (the multilayer perceptron)—A review of applications in the atmospheric sci- ences,” Atmos. Environ., vol. 32, nos. 14–15, pp. 2627–2636, Aug. 1998.
[4] T. J. Schmit, P. Griffith, M. M. Gunshor, J. M. Daniels, S. J. Goodman, and W. J. Lebair, “A closer look at the ABI on the GOES-R series,” Bull. Amer. Meteorol. Soc., vol. 98, no. 4, pp. 681–698, 2017.
[5] B. Klein, L. Wolf, and Y. Afek, “A dynamic convolutional layer for short range weather prediction,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., Jun. 2015, pp. 4840–4848.
[6] Y. Liu et al., “Application of deep convolutional neural networks for detecting extreme weather in climate datasets,” in Proc. 3rd Int. Conf. Adv. Big Data Anal., 2016, pp. 81–88.
[7] C. Marzban and A. Witt, “A Bayesian neural network for severe-hail size prediction,” Weather Forecasting, vol. 16, no. 5, pp. 600–610, Oct. 2001.
[8] A. H. Auer, “Hail recognition through the combined use of radar reflectivity and cloud-top temperatures,” Monthly Weather Rev., vol. 122, no. 9, pp. 2218–2221, Sep. 1994.
[9] A. Merino, L. López, J. L. Sánchez, E. García-Ortega, E. Cattani, and V. Levizzani, “Day-time identification of summer hailstorm cells from MSG data,” Natural Hazards Earth Syst. Sci. Discuss., vol. 1, no. 5, pp. 5453–5498, Apr. 2014.
[10] Z. Lu, H. Zhang, H. Ma, and H. Jia, “Hailstone detection based on time series association rules,” in Proc. 7th Int. Conf. Fuzzy Syst. Knowl. Discovery, Aug. 2010, pp. 2143–2146.
[11] Z. Lu, L. Wang, H. Ma, Q. Zhang, and H. Jia, “Hailstone detection based on image mining,” in Proc. 5th Int. Conf. Fuzzy Syst. Knowl. Discovery, 2008.
[12] J. G. Gagne, A. McGovern, J. Brotzge, M. Coniglio, J. Correia, and M. Xue, “Day-ahead hail prediction integrating machine learning with storm-scale numerical weather models,” in Proc. 29th AAAI Conf. Artif. Intell., Jan. 2015, pp. 3954–3960.
[13] NOAA SIS. (Mar. 2013). GOES Imager Instrument. Accessed: Jul. 13, 2017. [Online]. Available: http://noaasis.noaa.gov/NOAASIS/ml/ imager.html
[14] NASA GMAO. (Mar. 2016). Merra-2: File Specification. Accessed: Dec. 15, 2017. [Online]. Available: https://gmao.gsfc.nasa.gov/pubs/ docs/Bosilovich785.pdf
[15] R. Edwards and R. L. Thompson, “Nationwide comparisons of hail size with WSR-88D vertically integrated liquid water and derived thermodynamic sounding data,” Weather Forecasting, vol. 13, no. 2, pp. 277–285, 1998.
[16] J. C. Brimelow, G. W. Reuter, and E. R. Poolman, “Modeling maximum hail size in Alberta thunderstorms,” Weather Forecasting, vol. 17, no. 5, pp. 1048–1062, 2002.
[17] G. B. Foote, “A study of hail growth utilizing observed storm condi- tions,” J. Climate Appl. Meteorol., vol. 23, no. 1, pp. 84–101, 1984.
[18] R. Johns and C. Doswell, “Severe local storms forecasting,” Weather Forecasting, vol. 7, no. 4, pp. 588–612, 1992.
[19] A. W. Johnson and K. E. Sugden, “Evaluation of sounding-derived thermodynamic and wind-related parameters associated with large hail events,” E-J. Severe Storms Meteorol., vol. 9, no. 5, pp. 1–42, 2014.
[20] J. Billet, M. DeLisi, B. Smith, and C. Gates, “Use of regression techniques to predict hail size and the probability of large hail,” Weather Forecasting, vol. 12, no. 1, pp. 154–164, 1997.
[21] D. H. Kitzmiller and J. P. Breidenbach, “Detection of severe local storm phenomena by automated interpretation of radar and storm environment,” NOAA Tech. Memorandum NWS TDL, vol. 82, no. 33, pp. 141–159, 1995.
[22] NWS. (Mar. 2016). Storm Data Preparation. Accessed: Jul. 10, 2017. [Online]. Available: http://www.nws.noaa.gov/directives/sym/ pd01016005curr.pdf
[23] Stanford University. (2017). CS231n Convolutional Neural Networks for Visual Recognition. Accessed: May 14, 2017. [Online]. Available: http://cs231n.github.io/convolutional-networks/
[24] A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Proc. Adv. Neural Inf. Process. Syst., 2012, pp. 1097–1105.
[25] Merge Layers—Keras Documentation. Accessed: Oct. 16, 2018. [Online]. Available: https://keras.io/layers/merge/
[26] I. T. Jolliffe and D. B. Stephenson, Forecast Verification: A Practitioner’s Guide in Atmospheric Science. Oxford, U.K.: Wiley-Blackwell, 2012.
[27] D. W. Mccann, “A neural network short-term forecast of significant thun- derstorms,” Weather Forecasting, vol. 7, no. 3, pp. 525–534, Sep. 1992.
Melinda Pullman received the B.S. degree in meteorology from Mississippi State University, Starkville, MS, USA, in 2016, and the M.S. degree in earth system science from the University of Alabama in Huntsville, Huntsville, AL, USA, in 2018.
She is currently a Hydrologist with the U.S. Army Corps of Engineers (USACE) Vicksburg, Vicksburg, MS, USA.
Iksha Gurung received the B.S. degree in computer engineering from Kathmandu University, Dhulikhel, Nepal, in 2013, and the M.S. degree in computer science from the University of Alabama in Huntsville, Huntsville, AL, USA, in 2017.
He is currently a Research Associate with Inter Agency Implementation and Advanced Concepts (IMPACT) project, National Aeronautics and Space Administration (NASA) Marshall Space Flight Center, Huntsville, AL, USA. He is also with the Earth System Science Center, University of Alabama in Huntsville.
Manil Maskey (M’09) is currently pursuing the Ph.D. degree with the Department of Computer Science, The University of Alabama in Huntsville, Huntsville, AL, USA.
He is currently a Research Scientist with the National Aeronautics and Space Administration (NASA) Marshall Space Flight Center (MSFC), Huntsville, AL, USA. He also leads the Advanced Concepts team, Inter Agency Implementation and Advanced Concepts (IMPACT) project. His research interests include computer vision, visualization, and data analytics.
PULLMAN et al.: APPLYING DEEP LEARNING TO HAIL DETECTION 10225
Rahul Ramachandran (SM’09) is currently a Senior Research Scientist with the National Aero- nautics and Space Administration (NASA) Marshall Space Flight Center (MSFC), Huntsville, AL, USA. He also leads the Inter Agency Implementation and Advanced Concepts (IMPACT) Team, NASA MSFC. The IMPACT Team seeks to infuse NASA’s Earth Science data into other agencies and organi- zations application workflows. The IMPACT Team monitors trends across the informatics, data science, and information technology fields to inform strategy
and develop effective new solutions for Earth Science data management and dissemination. He has authored or coauthored more than 75 peer-reviewed publications, including 4 book chapters, and more than 150 other scientific publications including workshop reports. His research interests include earth science informatics and data science, especially the application of novel computational methods and information technology to the acquisition, storage, processing, discovery, interchange, analysis, and visualization of earth science data and information.
Dr. Ramachandran was a recipient of the Presidential Early Career Award for Scientists and Engineers (PECASE) Award in 2009 and the NASA Exceptional Achievement Medal in 2018. He has served as the Deputy Editor for the Earth Science Informatics Journal (Springer) and the Guest Editor for the Computer and Geosciences Journal (Elsevier).
Sundar A. Christopher is currently a Professor with the Department of Atmospheric and Earth Science, University of Alabama in Huntsville, Huntsville, AL, USA. He has authored or coauthored more than 100 peer- reviewed publications, several book chapters, and 2 books. His research interests include the usage of multi-sensor satellite data sets to study the role of aerosols on air quality and climate.
Dr. Christopher is a member of the AGU, AMS, and AAAS. He has served as a Principal Investigator on various satellite science teams.