Review on Energy Resilience

profileharsh55
Aquantitativemethodforassessingresilienceofinterdependentinfrastructures.pdf

Reliability Engineering and System Safety 157 (2017) 35–53

Contents lists available at ScienceDirect

Reliability Engineering and System Safety

http://d 0951-83

Abbre availabi commo method served; field lev lience; H operato tion tec work; M MTTR, m system; RAPI, ra unit; RT SD, syst average

n Corr E-m

journal homepage: www.elsevier.com/locate/ress

A quantitative method for assessing resilience of interdependent infrastructures

Cen Nan, Giovanni Sansavini n

Reliability and Risk Engineering Laboratory, ETH Zürich, Zürich, Switzerland

a r t i c l e i n f o

Article history: Received 19 January 2015 Received in revised form 17 June 2015 Accepted 15 August 2016 Available online 24 August 2016

Keywords: Interdependent critical infrastructure Resilience Reliability Agent-based modeling Interdependency

x.doi.org/10.1016/j.ress.2016.08.013 20/& 2016 Elsevier Ltd. All rights reserved.

viations: ABM, agent based modeling; ASSA lity index; CI, critical infrastructure; CNT, com n performance condition; CREAM, cognitive r ; CU, communication unit; CV, coefficient var EPSS, electric power supply system; FCD, fiel el instrumentation device; FIS, fuzzy inferenc EP, human error probability; HLA, high leve

r level; HRA, human reliability analysis; ICT, i hnology; IIM, input–output inoperability mod MI, man-made machine interface; MOP, mea ean time to repair; MTU, master terminal un PL, performance loss; PN, Petri-Net; R, robus pidity; RBC, RTU battery capacity; RL, resilienc I, run time infrastructure; SCADA, supervisory em dynamic; SoS, system of systems; SUC, syst d performance loss esponding author. ail address: [email protected] (G. Sansavini).

a b s t r a c t

The importance of understanding system resilience and identifying ways to enhance it, especially for interdependent infrastructures our daily life depends on, has been recognized not only by academics, but also by the corporate and public sectors. During recent years, several methods and frameworks have been proposed and developed to explore applicable techniques to assess and analyze system resilience in a comprehensive way. However, they are often tailored to specific disruptive hazards/events, or fail to properly include all the phases such as absorption, adaptation, and recovery. In this paper, a quantitative method for the assessment of the system resilience is proposed. The method consists of two compo- nents: an integrated metric for system resilience quantification and a hybrid modeling approach for representing the failure behavior of infrastructure systems. The feasibility and applicability of the pro- posed method are tested using an electric power supply system as the exemplary infrastructure. Si- mulation results highlight that the method proves effective in designing, engineering and improving the resilience of infrastructures. Finally, system resilience is proposed as a proxy to quantify the coupling strength between interdependent infrastructures.

& 2016 Elsevier Ltd. All rights reserved.

1. Introduction

1.1. Critical infrastructures and resilience

The welfare and security of each nation rely on a continuous flow of essential goods (such as energy, food, water) and services (such as banking, health care, public administration) provided by a set of systems called critical infrastructures (CI). Their incapacity or destruction would have a debilitating impact on the health,

I, average substation service plex network theory; CPC, eliability error analysis iation; ENS, energy not d level control device; FID, e system; GR, general resi- l architecture; HOL, human nformation and communica- eling; LAN, local area net- surement of performance; it; OCS, operational control tness; RA, recovery ability; e loss; RTU, remote terminal control and data acquisition; em under control; TAPL, time

safety, security, economics and social well-being [1,2]. They have always been “complicated”, but in recent years, they have wit- nessed higher integration and growing interconnectedness, which have turned them into a complex “System-of-Systems” (SoS) [3]. Interdependent infrastructures can correspond to one infra- structure (internal interdependency) or more infrastructures (ex- ternal interdependency). In [4], interdependency is characterized as different types: physical, geospatial, informational, logical [4]. These interdependencies may provide the tolerance to attacks and fail- ures if well managed (positive impact). For instance, technical failures such as abnormal disconnection of a transmission line can be detected by remotely installed devices at a substation and corresponding alarms can be transmitted to a control center via services provided by coupled telecommunication systems in order to prevent further failure propagations. However, these inter- dependencies might also be a source of threat generating risks, e.g., the risk of cascading failures, which make infrastructures more vulnerable (negative impact). In power blackout events, service disruptions further propagate to other infrastructures (transportation, telecommunication and water supply) and worsen the overall negative impacts [5]. Even though not preventable, these damages may be minimized, if the capabilities of both direct and indirect affected infrastructures are strengthened and effects of interdependencies are recognized [6–8]. Engineering the cou- pling among infrastructures, e.g., loosening it by adding buffer capacity, “slack” resources, redundancy or structure modularity,

C. Nan, G. Sansavini / Reliability Engineering and System Safety 157 (2017) 35–5336

has been suggested in order to decrease the impact of inter- dependency [9]. It is important to understand if interdependencies are essential, e.g. the dependency of power grids on the control system, or “parasitic”, e.g. the dependency of the control system on the controlled grid. The latter can be removed or redesigned. Moreover, the human/social element is recognized to play a key role in the operations of infrastructures.

To better understand the performance of infrastructures, especially their behavior during and after the occurrence of dis- turbances (e.g., natural hazards or technical failures), resilience analysis [10–14] has grown as a proactive approach to enhance the ability of infrastructures to prevent damage before disturbance events, mitigate losses during the events and improve recovery capability after the events, by extending the concept of pure pre- vention and hardening. The term resilience is still evolving and has been developing in various fields. The first definition is given by an ecologist, Holling, who described resilience as “a measure of the persistence of systems and of their ability to absorb change and disturbance and still maintain the same relationships between po- pulations or state variables” [15]. Since then, others have put for- ward different domain-specific resilience definitions [16,17]. The concept of resilience is also related to technical systems. In recent years, the ability to assess and engineer resilience of these systems has emerged as a fundamental concern for researchers [18–22]. Further development of this term should include endogenous and exogenous events as well as recovery efforts. In order to ade- quately consider these factors, a definition is proposed, which describes resilience as “the ability of a system1 to resist the effects of a disruptive force and to reduce performance deviations”.

1.2. Resilience capabilities

The proposed resilience definition can be further interpreted as the ability of the system to withstand a change or a disruptive event by reducing the initial negative impacts (absorptive cap- ability), by adapting itself to them (adaptive capability) and by recovering from them (restorative capability). Enhancing any of these features will enhance system resilience. It is important to understand and quantify these capabilities that contribute to the characterization of system resilience [23]. Absorptive capability refers to an endogenous ability of the system to reduce the ne- gative impacts caused by disruptive events and minimize con- sequences. In order to quantify this capability, robustness (R) can be used, which is defined as the strength of the system to resist disruption [24,25]. This capability can be enhanced by improving system redundancy, which provides an alternative way for the system to operate. Adaptive capability refers to an endogenous ability of the system to adapt to disruptive events through self- organization in order to minimize consequences. Emergency sys- tems can be used to enhance adaptive capability. Restorative capability refers to an ability of the system to be repaired. For example, installing real-time automated monitoring systems, e.g., Supervisory Control and Data Acquisition system (SCADA), is able to enhance the restorative capability since protection relays can be programmed to bypass faulty sections to transfer loads to sound areas or lines for load shedding [26]. The effects of adaptive and restorative capacities overlap and therefore, their combined effects on the system performance are quantified by rapidity (RAPI) and performance loss (PL).

Fig. 1 provides a general illustration of the essential resilience capabilities. The y-axis represents the measurement of performance (MOP). Examples of MOP include the availability of critical facil- ities, the number of customers served, connectivity of a network,

1 The system can be a single system, an infrastructure, or even a SoS.

or level of functionality and economic activities. The selection of the appropriate MOP depends on the specific service provided by the system under analysis. It is assumed that the disruptive event occurs at t0, and that the MOP values drops at t1. It should be noted that in many cases t0 might not be equal to t1 and the t1–t0 delay depends on the selection of the MOP and on the disruptive event. For instance, it could take several hours for customers to lose electricity services due to maintenance errors, while it might only take seconds for same customers to lose services due to natural hazards such as earthquakes or hurricanes. System susceptibility, defined as “the inability of a system to avoid being hit by a threat mechanism” [27], can be used to characterize the system perfor- mance between t0 and t1. The focus of this paper is related to re- silience quantification after the appearance of the negative effects, and therefore, susceptibility is not considered.

The system capabilities that have the effect on system resi- lience can be exemplified with respect to system performance variations following a disruptive event. As shown in Fig. 1, sys- tem#1 performance returns to its original steady level after re- covering from the lowest level at t2. System#2 performance reaches a new steady level, which is lower than the original steady level. The new performance level could also be higher than the original level, e.g. when the system is retrofitted. System#3 per- formance drops significantly and finally collapses to zero. Sys- tem#1 and system#2 outperform system#3 with respect to the three essential resilience capabilities. Therefore, system#3 can be considered the least resilient one. On the other hand, system#1 seems more robust against the disruptive event than system#2, i.e. the lowest performance level is higher for system#1 than for system#2. However, it takes more time for system#1 to reach the new steady level. Determining which system is more adaptive and restorative is not clear-cut. To this aim, an approach to quantify resilience capacities and to compare them using a unique resi- lience index is expedient.

1.3. Frameworks for assessing resilience

During the last decade, researchers have proposed different methods for quantifying resilience. In 2003, the first conceptual framework was proposed by Bruneau et al. in [28], aiming to measure the seismic resilience of a community to an earthquake by estimating the expected degradation regarding the quality of community infrastructures. In this pioneering research work, the concept of Resilience Loss (RL), later also referred to as “resilience triangle”, is introduced, which has been widely used afterward as a fundamental guidance for resilience assessment. In 2004, Chang and Shinozuka proposed a probabilistic approach for measuring seismic resilience after earthquake events [29].

In recent years, the importance of improving the resilience of interdependent infrastructures or minimizing negative impacts caused by unexpected disruptive events has been recognized and broadly accepted [30]. Therefore, a variety of research work has been developed targeting interdependent infrastructures. In 2008, McDaniels et al. developed a knowledge-based approach using decision flow diagrams to improve the understanding of two di- mensions (robustness and rapidity) of the resilience of infra- structures [22]. A similar approach was also proposed by Argonne National Lab, in which a resilience index is designed and devel- oped by assigning relative weights to major capacities of system resilience such as robustness and recovery in order to provide information to system owners and operators about ways to im- prove resilience [11]. However, this knowledge-based approach is purely data-driven: the quality of the collected information could have significant effects on the accuracy of final results. To over- come these limitations, more approaches have been developed with the help of advanced modeling techniques. In [31], System

Fig. 1. Illustration of essential resilience capabilities.

C. Nan, G. Sansavini / Reliability Engineering and System Safety 157 (2017) 35–53 37

Dynamics was applied to assess the degree of socio-ecological resilience. This approach combined with Complex Network Theory was later applied by Filippini and Silva as part of the framework for qualitative resilience analysis of infrastructures [32]. However, most of the existing methods for resilience quantification lack the ability to cover all phases, and to include all resilience capabilities. Furthermore, they often overlap with other concepts such as ro- bustness, vulnerability, and fragility [33]. Some quantitative methods for resilience estimation are not consistent with the underlying concept of resilience. For instance, several methods mainly focus on the evaluation of performance loss without con- sidering rapidity and robustness [34]. Furthermore, these methods rely on modeling approaches that partially capture the complex behavior of interdependent infrastructures. Therefore, there is the pressing need to develop a unifying methodology for the assess- ment of the resilience of infrastructures in response to various disturbances. The research in resilience quantification of inter- dependent infrastructures is still at an early stage. Currently, a comprehensive method aiming at improving our understanding of the system resilience and at analyzing the resilience by performing in-depth experiments is still missing.

This paper aims at (I) proposing a methodology to analyze dynamics of infrastructures, (II) developing time-dependent me- trics for resilience quantification in infrastructures and compared it with performability metrics, (III) providing insights to decision makers for strengthening infrastructure resilience, e.g. rank the benefits of resilience-improving strategies (improving recovery or battery capacity of remote terminal unit), trade-off between in- terdependencies and resilience.

In order to achieve these goals, a generic quantitative method for the assessment of resilience of interdependent infrastructures is presented in this paper. The method consists of two components: (1) an integrated metric for resilience quantification with capabilities of incorporating different performance measures and characterizing resilience as a sys- tem attribute, (2) a multi-layer hybrid modeling approach to achieve a close representation of the specific systems and analyze their dynamics during and after the occurrence of disruptive events.

This paper presents the continuation of the research work in [35]. The major contributions are summarized below:

� a systematic procedure for the application of the proposed multi-layer hybrid modeling approach is developed;

� the application of the proposed method is extended to multiple infrastructures;

� the resilience metric is extended to account also for system functionality as well as for system topology;

� the impact of interdependencies among systems on the resi- lience is studied and discussed;

� resilience and performability metrics are compared and discussed.

The paper is organized as follows: Sections 2 and 3 introduce the two components of the proposed framework, i.e. the in- tegrated metric for resilience quantification and the multi-layer hybrid modeling approach respectively. Section 4 considers the Swiss electric power supply system (EPSS) as an exemplary system to apply the proposed method for resilience assessment; the de- sign of simulation experiments is also presented. The discussion of the simulation results is included in Section 5. Finally, Section 6 provides conclusions and outlines directions for future work.

2. Method component 1: development of an integrated resi- lience metric

2.1. Quantification of resilience capabilities

Resilience cannot be adequately addressed considering one single system capability. Therefore, the measures of the resilience capabilities (i.e. absorptive, adaptive and restorative capability) in different phases (i.e. original steady, disruptive, recovery and new steady phase) are identified and integrated into a unique resilience metric. In order to guide the development of the integrated resi- lience metric, the performance of the system#1 in Fig. 1 is further analyzed and characterized by four phases and three transitions in MOP value in Fig. 2.

The selection of the appropriate MOP depends on the specific service provided by the infrastructure under analysis. For gen- erality, it is assumed that the value of MOP is normalized between 0 and 1 where 0 is total loss of operation and 1 is the target MOP value in the steady phase. As illustrated in Fig. 2, the first phase is the original steady phase (totd), in which the system performance assumes its target value. The second phase is the disruptive phase (tdrtotr), in which the system performance starts dropping until

Fig. 2. System resilience transitions and phases.

Table 1 Summary of different resilience phases.

Phases Time scope Transition point

Capabilities (features)

Measurements

Original steady phase

totD Susceptibility Susceptibility

Disruptive phase

tDrtotR TRNS(D) Absorptive capability

R RAPIDP PLDP

Recovery phase

tRrtotNS TRNS(R) Adaptive cap- ability Re- storative capability

RAPIRP PLRP

New steady phase

tZtNS TRNS(NS) Recovery capability

RA

R: robustness; RAPIDP: rapidity in disruptive phase; PLDP: performance Loss in disruptive phase; RAPIRP: rapidity in recovery phase; PLRP: performance Loss in recovery phase; RA: recovery ability.

C. Nan, G. Sansavini / Reliability Engineering and System Safety 157 (2017) 35–5338

reaching the lowest level at time tr. During this phase, the system absorptive capability can be assessed by identifying appropriate measures. As discussed in Section 1.2, robustness (R) is a measure to assess this capability, which quantifies the minimum MOP value between td and tns:

)( ) ({ }= ≤ ≤ ( )R MOP t td t tnsmin for 1 where td represents the time when the system is in disruptive phase and tns represents the time when the system reaches the new steady phase. This measure is able to identify the maximum impact of disruptive events; however, it is not sufficient to reflect the ability of the system to absorb the impact. Two additional complementary measures are further developed: Rapidity (RAPIDP) and Performance Loss (PLDP) during the disruptive phase. The measure Rapidity can be approximated by the average slope of the MOP function.

= ( )− ( )

− ( ) RAPI

MOP t MOP t t t 2

DP d r

r d

To improve the accuracy of the estimation of RAPI, ramp de- tection is applied to quantify the average slope [36]. According to [37,38], a ramp is assumed to occur if the difference between the measured value at the initial and final points of a time interval Δt is greater than a predefined ramping threshold value:

( )+∆ − ( ) ∆

>∆ ( )

MOP t t MOP t

t X

3ramp

where ∆Xramp represents the predefined ramping threshold value. The system rapidity can then be calculated as the average of slope of each ramp:

( )

= ∑

( ) =

− ( − ∆ ) ∆

K RAPI

4 i K MOP t MOP t t

t1 i i

where K represents the number of detected ramps and MOP(ti) represents the MOP value at the i-th detected ramp. Compared to (2), this method better captures the speed of change in the system performance during disruption and recovery phases. According to this approach, the rapidity during disruptive phase can be calcu- lated as:

( )

= ∑

( ≤ < ) ( )

= − ( − ∆ )

∆ RAPI

K t t tfor i

5 DP

i K MOP t MOP t t

t

Dp d r

1 DP i i

where KDP represents the number of detected ramps during the disruptive phase.

The performance loss in the disruptive phase (PLDP), using the system illustrated in Fig. 2 as an example, can be quantified as the area of the region bounded by the MOP curve before and after occurrence of the negative effects caused by the disruptive events, i.e. between td and tr which is referred to as the system impact area:

( )∫ ( ) ( )= − ( )

MOP t MOP t dtPL 6t

t

DP o d

r

Where to represents the time when the system is in original steady phase. A new measure, i.e. the time averaged performance loss (TAPL), is introduced. Compared to PL, TAPL considers the time of appearance of negative effects due to disruptive events up to full system recovery and provides a time-independent indication of both adaptive and restorative capabilities as responses to the disruptive events. TAPLDP in the disruptive phase (tdrtotr) can be calculated as:

1

2

3

1

2

3

2 In this paper, the term “system” is used if individually introduced and dis- cussed, while term “subsystem” is used if discussed as part of a system.

C. Nan, G. Sansavini / Reliability Engineering and System Safety 157 (2017) 35–53 39

( )( ) ( )∫ =

− ( ) TAPL

MOP MOP t dt

t t

t

7 DP

t

t

r d

o d

r

The third phase is the recovery phase (trrtotns), in which the system performance increases until the new steady level. During this phase, the system adaptive and restorative capability can be assessed by developing appropriate measures: rapidity (RAPIRP), performance loss (PLRP) and time average performance loss (TAPLRP).

( ) )( )

= ∑

≤ < ( )

= − ( − ∆

∆ RAPI

K t tfor ti

8 RP

i K MOP t MOP t t

t

Rp r ns

1 RP i

where KRP represents the number of detected ramps in the re- covery phase.

( )∫ ( ) ( )= − ( )

PLRP MOP t MOP t dt 9t

t

0 r

ns

( )( ) ( )∫ =

− ( ) TAPL

MOP t MOP t dt

t t 10 RP

t

t

ns r

0 r

ns

The fourth phase is the new steady state (tZtns), in which system performance reaches and maintains a new steady level. As seen in Fig. 1, the newly attained steady level may equal the pre- vious steady level or reach a lower level. It should be noted that the new steady state may even be at a higher level than the ori- ginal one. In order to take this situation into consideration, a simple quantitative measure Recovery Ability (RA) is developed:

( ) ( )

= ( )− ( )− ( )

RA MOP t MOP t

MOP t MOP t 11

ns r

o r

Different system phases and related system capabilities are summarized in Table 1.

2.2. The integrated resilience metric

Although the measurements introduced and discussed in Sec- tion 2. A are useful in assessing the system behavior during and after disruptive events, an integrated metric with the ability to combine these capabilities is needed in order to assess system resilience with an overall perspective and to allow comparisons among different systems and system configurations. The basic idea of incorporating various resilience capacities into one metric has been proposed by Francis and Bekera to develop resilience factor [39]. The idea is also supported by [22]. In this paper, a resilience metric (GR) is proposed, which integrates the previous measures:

( )= = ×( ) × ( ) × ( )−

GR f R RAPI RAPI TAPL RA R RAPI RAPI

TAPL RA

, , , ,

12

DP RP RP

DP

1

where TAPLDP and TAPLRP have been combined into one TAPL

measure ( ( ) ( )∫ − ]

⎡⎣ t t t t

MOP MOP dt td

tns

ns d

0 ) in order to incorporate effects of

total performance loss during disruptive and recovery phases. The functional form of the proposed resilience metric assumes

that robustness R, recovery speed RAPIRP and recovery ability RA have a positive effect on resilience, and, conversely, performance loss TAPL and loss speed RAPIDP have a negative effect. To compile the in- tegrated metric (12), no weighting factor is assigned to the measures so that no bias is introduced, i.e. they contribute equally to resilience. GR is consistent with the definition proposed in Section 1.1:

) If the system is more capable of resisting a disruptive event or force (large R, small RAPIDP), the system is more resilient (large GR).

) If the system is more capable of reducing the magnitude and duration of deviation of its performance level between original state and new steady state (small TAPL, large RAPIRP), the system is more resilient (large GR).

) Additionally, GR also incorporates the possibility of improve- ment of the system performance after the occurrence of the disruptive event. If the new performance level is larger than the original (large RA), the system is more resilient (large GR).

GR is a non-negative metric and its value equals zero in the following relevant cases:

) System performance level drops to zero after the disturbance (R¼0).

) After the disturbance, system performance immediately drops to its lowest level (RAPIDP - 1, i.e. no absorptive capability).

) System performance never increases past the lower level, R, which is the new steady phase (RAPIDP¼0, i.e. no adaptive and restorative capability).

GR is dimensionless and is most useful in a comparative man- ner, i.e. to compare the resilience of various systems to the same disruptive event, or to compare resilience of the same system under different disruptive events. This approach of measuring system resilience is neither model nor domain specific. For in- stance, historical data can also be used for the resilience analysis. It only requires the time series that represents system output during the whole time period. In this respect, the selection of the MOP is very important.

3. Method component 2: multi-layer hybrid modeling approach

A multi-layer hybrid modeling approach is proposed, in which an infrastructure can be represented by the coupling of interacting subsystems characterized by layers, i.e. system under control, op- erational control system, and human operator level system.2 Using this modeling approach, different types of subsystem-specific models can be developed at each layer, and integrated to include interactions among layers. This approach is also flexible and adaptable to modeling diverse types of infrastructures. The deci- sion whether to model the infrastructure as a whole, i.e. addres- sing small problems in isolation and combining the solution to- gether in a bigger picture, or having the global picture at a reduced resolution, guides the model development and should be an- swered by the system analyst beforehand.

The implementation of the multi-layer hybrid modeling ap- proach can be divided into 3 working steps.

� Step 1 – Screening analysis: given the infrastructure under analysis, the main purpose of this step is to: (a) decompose the infrastructure into the relevant subsystems which influence the characteristics and behaviors to be modeled, (b) for each sub- system, select the appropriate modeling approach that captures the characteristics and behaviors. Modeling infrastructures is often challenged by their inherent characteristics, e.g., dynamic and nonlinear behaviors, complex interdependencies as the results of their openness and the high degree of inter- connectedness [40,41]. Currently, a variety of modeling ap- proaches have been developed and applied, e.g., Input–output Inoperability Modeling (IIM), Complex Network Theory (CNT),

C. Nan, G. Sansavini / Reliability Engineering and System Safety 157 (2017) 35–5340

Agent-based Modeling (ABM), System Dynamic (SD), and Petri- Net (PN) (see [1,42,43] for details about these approaches). Their abilities to capture specific system characteristics vary. Some approaches are only capable of analyzing the structure/ topology level, e.g., CNT, while some approaches capture dy- namic behaviors at the functional level, e.g., ABM. Furthermore, it is also important to take the costs for modeling and simula- tion into consideration when selecting appropriate approaches, e.g., the necessary expertise (e.g., how modeling languages and math formalisms are handled), data retrieve to feed model, overhead in updating/modifying model, and interpretation of results.

� Step 2 – Individual model development: in this step, sub- system-specific models are individually developed using the modeling approaches selected in Step1. Each model captures the characteristics of the related system and can be developed independently of other peer models. This makes the reuse of existing models possible.

� Step 3 – Model interaction: in this step, variables of each model that are an input to other peer models are determined. These variables identify the interdependency pattern among the subsystems and guide the assembly of the models of those subsystems. In order to facilitate interoperability of the models and ensure efficient data exchange, appropriate simulation standards must be used. Currently, no simulation standard ex- ists, which is specifically able to represent the interactions across sector-specific models. Several standards do exist which enable efficient data exchange among sector-specific models. Among these standards, the High Level Architecture (HLA) is the most implemented and applicable and it supports simulations composed of coupled models [44].

3.1. Modeling the EPSS system using the hybrid multi-layer approach

Electric power supply systems (EPSS) are the most common example of CI and the need for their reliable and resilient perfor- mance following disruptive events is essential [45,46]. In this

Fig. 3. The representation of the EP

paper, the high-voltage EPSS is selected as an exemplary infra- structure to apply the proposed quantitative method for resilience assessment. The development of EPSS model using the proposed multi-layer hybrid modeling approach can be conducted following the three aforementioned steps:

Step 1: screening analysis EPSS consists of 3 interdependent subsystems arranged in

3 different layers, i.e. System Under Control (SUC), Operational Control System (OCS), and human operator level system (HOL). SUC represents a technical part of the EPSS. Its components in- clude transmission lines, generators, busbars and relays. It is a time-stepped system, i.e. the time scale has a strong influence on its functionalities. OCS also represents the technical part of the EPSS. Its major responsibility is to control and monitor the coupled SUC. Compared to the SUC, OCS is an event-driven system, i.e. its functionalities are mainly influenced by events rather than by the time scale. In this paper, the Supervisory Control and Data Ac- quisition (SCADA) system is used as an example of the OCS. Components of the SCADA include field instrumentation and control devices (FIDs and FCDs), remote terminal units (RTUs), communication units (CUs), and master terminal unit (MTU). HOL represents a non-technical part of the EPSS, which is related to human and organizational factors influencing the overall system performance. In this paper, only the human factor is considered and organizational factors are ignored. Generally, the HOL is re- sponsible for monitoring/processing generated alarms, switching off components at remote substations and sending commands to remote substations.

Fig. 3 shows a simplified multi-layer representation of the EPSS with three interacting subsystems and associated elements (components): parallel planes represent different subsystems in corresponding layers, and nodes represent various elements to- gether with the interconnections among them. The elements of the various layers depend on each other, depicted by various horizontal (inside a layer) and vertical (between layers) links. These vertical links are interdependencies among subsystems and

SS in three subsystems (layers).

C. Nan, G. Sansavini / Reliability Engineering and System Safety 157 (2017) 35–53 41

a failure of an element can cause cascading failures to other layers that could affect the operations of the whole infrastructure.

� SUC and OCS: In order to achieve a high-fidelity modeling of these two technical parts, both functionality (physical laws) and structure (topology) should be considered. Furthermore, the developed model for OCS needs to be able to process messages among components. ABM is selected to model these systems. This approach intends to represent a whole system by dividing it into interacting agents. Each agent is capable of modifying its internal data, behaviors and adapts itself to environmental changes. ABM is a bottom-up approach and each component is represented as an agent [47,48].

� HOL: The developed model for HOL should be able to quantify the effects of human performance. The method of HRA (Human Reliability Analysis) is suitable to this aim. HRA provides a way to assess human performance in either qualitative or quantita- tive ways. Qualitative methods focus on the identification of events or errors, and quantitative methods focus on translating identified events/errors into Human Error Probability (HEP) [49].

Step 2: individual model development

� SUC: The ABM approach integrates the physical relationships among the components of SUC and the stochastic time-depen- dent factors. The physical and operational processes are mod- eled by means of DC power flow calculations, and each com- ponent (e.g., lines, bus bars, protection relays) is mapped to one agent. Specific behavioral rules are assigned to each agent, in- cluding both deterministic and stochastic time-dependent processes, triggered by time or inputs from other agents. For instance, a deterministic process is the power overload of a transmission line, and a stochastic process is the triggering of a

Fig. 4. Overview of interactions among agents of the SCADA model. (For inter- pretation of the references to color in this figure, the reader is referred to the web version of this article.)

component failure mode, e.g., the unplanned outage of a gen- erator. Within this model, different types of agents interact. During the simulation, their conditions are updated according to pre-defined physical laws and interaction rules. The new conditions are then sent back to the agents to update their behaviors. For example, the power flow of the agent of a transmission line (Pik) can be computed using following formula [50]:

θ θ= ( − ) ( )

P x 1

13 ik

ik i k

Where xik represents the line reactance (ohms) between the agent of busbar(i) and busbar(k), θi and θk represent the phase angle (rad) of the agent of busbar(i) and busbar(k) relative to the reference busbar. The use of the DC power flow calculation results in the partial representation of the SUC because it neglects voltage magnitude and reactive power flows. Further- more, the impact of protection systems, i.e. automatic circuit breakers, is considered. Including other operational procedures and protection systems, e.g. voltage regulations, reserves for dynamic transients and frequency control, is affordable but increases the computational burden for the simulation to a relevant extent. Nonetheless, the main goal of this paper is to ascertain the feasibility of coupling several models of infra- structures and of using the hybrid multi-layer model to quantify system resilience. In order to obtain more accurate results for the power network, the DC power flow assumption can be dropped and other protection systems and operational proce- dures can be included in a later stage. The development of the SUC model is detailed in [51].

� OCS: An agent-based model is applied for accounting nominal behavior and misbehavior (induced by failures) of the SCADA. Each hardware device is modeled as an agent, which maps the hardware status including operational and failures modes. For example, failure to open and failure to close are two examples of failure modes defined for the FCDs (see complete failure modes of the SCADA devices in [52]). Interactions among the agents of the SCADA system are briefly demonstrated in Fig. 4.

As shown in Fig. 4, FCD, FID and MTU agents are the interface between SCADA and other subsystems. Variables of the SUC model which are the input of the SCADA model are handled through FCD and FID agents. Status variables, e.g., line connected, line dis- connected and numerical variables, e.g., power load of carried by transmission lines, are sent to FCD and FID agents, respectively. The RTU agents acquire these variables from the linked FID and FCD agents and forward them to the MTU agent whenever re- quired. The decision about whether or not an alarm should be generated is determined by RTU agents based on the analysis of these variables. If an alarm is generated by an RTU, it is forwarded to the MTU agent. In case an alarm is received by the MTU agent and further interpreted by the operator (represented by the HOL model) successfully, a command for related corrective actions will be issued by the operator and sent to the MTU agent. The MTU then forwards the command to the RTU which initializes the corresponding alarm. If the corrective action requires the partici- pation of field devices, e.g., disconnection of a transmission line, the command will be forwarded to an FCD agent. The SCADA model also includes a number of objects, represented by green circles in Fig. 4, i.e. commands, alarms, and monitors, whose aim is to transmit data among agents. While an agent is a decision- making entity, an object is a data structure consisting of data fields and methods. The development of the SCADA model is detailed in [52].

Fig. 5. The structure of the hierarchy modeling approach of the HOL.

Table 2 List of EPSS variables and the interactions among the models.

EPSS variables Role of EPSS models

Name Type SUC OCS (SCADA)

HOL (Human operator)

Power flow (transmis- sion line)

Real number Output Input Input

Status (transmission line)

Boolean Input Output Input

Status (load) Boolean Output Input Input Actual load (load) Real number Output Input Input Demand (load) Real number Output Input Input Status (generator) Boolean Output Input Input Actual power (generator)

Real number Output Input Input

Alarm Real number N/A Output Input String Boolean

Command Real number N/A Input Output String Boolean

C. Nan, G. Sansavini / Reliability Engineering and System Safety 157 (2017) 35–5342

� HOL: An agent-based hierarchal model is developed to capture the human trends of EPSS (Fig. 5).

The model describes the process of information acquisition, analysis and decision making of the human operator, and evalu- ates the performance of this layer. The components “sensor” and “perception” are the input of the model. The components of “physical status” and “emotion status” contain parameters, state variables, rules, which have influences on the “sensor”, “percep- tion” and “cognition” components. The “behavior” and “actor” components are the output of the model. The component “beha- vior” determines the order of action execution and the component “actor” executes them. The information that has the influence on other agents will be sent out by the component “actor”. When an alarm is generated by the OCS, it is sent to the “sensor” component in the HOL. The output of the “perception” component is the op- erator situation awareness, e.g., the severity of an overload con- dition on a power line, which is used by the “cognition” compo- nent to elaborate the situation and create tasks. A task is handled by dividing it into sequential subtasks. For example, the task of

alarm handling can be divided into following subtasks. Subtask 1: the operator needs to check whether the alarm system is working properly; subtask 2: the operator acknowledges the alarm; subtask 3: the operator issues the control command. Finally, the subtasks are sent to the “behavior” component to determine possible error modes and causes during their execution. To this aim, Common Performance Conditions (CPCs) are assessed based on the Cogni- tive Reliability Error Analysis Method (CREAM) [53]. The examples of CPCs are adequacy of organization, working conditions, ade- quacy of man–machine interface and operational support, avail- ability of procedures/plans, number of simultaneous goals, avail- able time, time of day, and crew collaboration quality. Various levels are assigned to each CPC. For instance, three levels are as- signed to the CPC “working conditions”: advantageous, compatible, and incompatible. The assigned levels of CPCs determine weight- ing factors for each subtask, which have influence on the value of the final Human Error Probability (HEP):

( )( )= × = … ( )⎡⎣ ⎤⎦iHEP Max HEP WF , 1, 2 .n 0, 1 14i i where HEPi and WFi represent the nominal HEP value and weighting factor of the i-th predefined subtask, and n is the number of subtasks. Small HEP value indicates better human performance. To assess whether or not there is an error by the human operator, it is ne- cessary to set a threshold value (HEPA) representing the maximum acceptable HEP value. If a calculated HEP is more than HEPA, then a human error occurs (the human action fails to perform); large HEPA indicates that the operator is less error prone. The HEPA value can be determined based on several factors related to operators such as adequacy of experience and training. The development of the HOL model is detailed in [54].

Step 3: model interaction The variables that define the interactions among the three

subsystems are identified. They act as coupling points among the three models. For example, the OCS model requires numerical variables such as power flow of transmission lines, which are the output of the SUC model, and, on the other hand, the SUC model requires the most recent operating status of field devices (pro- tection relays), which are the output of the OCS model. The OCS model requires the action command in order to maintain its nor- mal operation, which is the output of the HOL model. Table 2 lists all the variables related to the interactions among three models including their roles about how they handle those variables.

Fig. 6. The high voltage Swiss electric power transmission system [55].

C. Nan, G. Sansavini / Reliability Engineering and System Safety 157 (2017) 35–53 43

In order to exchange these variables with other models via the HLA simulation standard, the local interface for each model is developed, which contains the HLA-related classes and methods responsible for sending and receiving data with other models. All three subsystem-specific models are connected over a Local Area Network (LAN). Simulation monitor system is a real-time tool, through which the simulation of three subsystem-specific models is observed.

4. Case study for resilience assessment

4.1. Design of the experiment

The Swiss high-voltage EPSS is selected as an exemplary ap- plication to demonstrate the feasibility and applicability of the proposed analytical method for resilience assessment. The system is illustrated in Fig. 6, which consists of 219 transmission lines and 129 substations.

It is assumed that this transmission system is a stand-alone system and the energy exchange with the neighboring countries is regarded as independent positive and negative power injections at the respective boundary substations. The multi-layer hybrid modeling approach introduced in Section 3, is applied to model the Swiss high-voltage EPSS. In total, 587 agents are created to represent components of SUC, i.e., transmission lines, generators, busbars, while 588 agents are created to represent components of OCS (SCADA), i.e., FCDs, FIDs, RTUs, and MTU.

Historical records reveal that hazards such as earthquakes and winter storms were the cause of significant damage in at least 9 events over the past 1000 years in Switzerland [56]. According to [57], the estimated frequency of natural hazards, i.e. winter storms, which have the potential of resulting in the simultaneous

disconnection of 20 transmission lines is in the range of 6 � 10�4– 7 � 10�4 per year. The impact of this disruptive event will result in large negative effects, and should not be underestimated even if the frequency of occurrence is relatively low. In January 2014, Slovenia was hit by the heaviest ice rain in the history and many power lines were affected [58,59]. As one of the consequences, many areas were cut off electric power for days. Motivated by the severe consequence of these hazardous events, in this experiment, it is assumed that a natural hazard, i.e. winter storm or ice rain, impacts the central region of Switzerland (highlighted in the circle of Fig. 6), where power transmission lines are located; as a result, about 17 power transmission lines are disconnected. The selected region was affected by Lothar winter storm in 1999 [60].

In this case study, the total power demand in a snapshot of the Swiss transmission grid on a day in winter is used. It is assumed that the disruptive event occurs at time 3 h. Before this time, all the modeled systems are in the original steady state and operate under normal conditions. At t¼3 h, the disconnection of the 17 transmission lines in the selected region is triggered. In order to model the quasi-simultaneous disconnection, it is assumed that the interval between disconnections of two lines follows a normal distribution N(35 s,3 s2). The sudden disconnection of a transmis- sion line will be detected by the corresponding RTU component of the SCADA system and the “abnormal line disconnection alarm” will be sent to the control center (MTU). After receiving the alarm by the control center, the operator activates repair actions in order to restore the disconnected lines, e.g., assessing network status, attempting line reclose, or decide whether repair crew should be sent to the field. It is assumed that the response time (ResponseG) for this type of alarm (i.e. time to make decisions after acknowl- edging the alarm) follows a normal distribution N(80 s,5 s2). Nonetheless, the sudden disconnections of many transmission lines may overwhelm the operator, possibly delaying the response

C. Nan, G. Sansavini / Reliability Engineering and System Safety 157 (2017) 35–5344

and repair actions. In order to simulate this situation, the formula below is used to calculate actual response time (ResponseA):

( ) ( )

= *

= * − + ≥

<⎪ ⎪⎧⎨ ⎩ 15

if HEP HEP

if HEP HEP

Response Delay factor Response

Delay factor weighting factor HEP HEP 1

1 A

A

A G

A

The actual response time (ResponseA) determined by the delay factor must lie in a sensible range. For simplicity but with no loss of generality, the weighting factor is assumed to be 100 in this experiment. It is assumed that repair actions for the abnormally disconnected lines are always performed successfully and the re- pair time is assumed to follow exponential distribution with mean value equal to the mean time to repair (MTTR).

The sudden disconnections of transmission lines could also overload other transmission lines, especially the neighboring ones [61], and have the potential for knock-on effects with cascading consequences. If a transmission line becomes overloaded, the sec- ond type of alarm, i.e. the “overload alarm”, is triggered and sent to the operator in the control center (MTU) by the corresponding RTU component. If the operator recognizes this alarm and handles it successfully (HEPoHEPA), the corrective actions will be performed, i.e. generation re-dispatch. However, if no action is taken after a certain time (i.e. 20 min) after the “overload alarm”, it is considered that the operator has failed to react (HEPZHEPA), and the protec- tion devices, e.g., circuit breakers, will automatically disconnect the overloaded line to prevent permanent damages to the EPSS, which could even trigger more line disconnections. It should be noted that the procedure for handling a power line “overload alarm” in real practice is complicated and that multiple factors should also be considered. In order to simplify this problem, two factors are con- sidered, i.e. the operator and the protection devices [44]. In general, “abnormal line disconnection alarms” are generated when the sys- tem is in the disruptive phase (tdrtotr) and the handling of this alarm completes when the system is in the recovery phase (trrtotns). Conversely, “overload alarm” can be generated either at the disruptive phase or the recovery phase. The consequence of the disconnections of transmission lines is the disruption of power supply to customers, which is estimated by calculating the actual power demand served. The disconnection of transmission lines of SUC also has negative effects on its interlinked SCADA system, i.e. service interruption of RTUs. In this experiment, it is assumed that an RTU component of the SCADA system has two power supply sources: the monitored/control transmission lines and the battery. An RTU component is out of power if its connected transmission lines are all disconnected and its battery is fully depleted. The si- mulation stops when the system reaches the final steady state. Then, all the performance measurements of SUC and SCADA system, including the GR value, are evaluated.

To simulate the performance response of the EPSS subjected to the disruptive event, all three models (SUC, SCADA, HOL) are in- cluded in the experiment. Two different MOPs are selected to analyze the behavior of the SUC of the EPSS for the purpose of validating the usefulness of all the resilience related measures.

Table 3 Summary of the three resilience improvement strategies.

# Strategy Target system R

1 Improvement of efficiency of line reparation SUC M 2 Improvement of human operator performance SUC H 3 Increasing RTU battery capacity SCADA RB

SUC: system under control; SCADA: supervisory control and data acquisition; MTTR: me repair time; GR: general resilience; CV: coefficient of variation.

These MOPs focus on different characteristics of the system: (1) the number of available transmission lines (topology related) and (2) actual power demand served (functionality related). One MOP is selected for the SCADA: the number of available RTUs (topology related). In order to allow for comparison, the MOPs are normalized in the range [0,1]:

= ∈⎡⎣ ⎤⎦MOP number of available transmission lines number of total transmission lines

0, 1SUC1

) )

= (

( ∈⎡⎣ ⎤⎦MW

MW MOP

actual power demand served

total power demand 0, 1SUC2

= ∈⎡⎣ ⎤⎦MOP number of available RTU devices number of total RTU devices

0, 1SCADA1

In order to compare the resilience metric GR with competing solutions, two performability metrics are also calculated. The Average Substation Service Availability Index (ASSAI) is the ratio of the total number of hours that service is provided by all available substations to the total demanded hours:

( ) ( )

= × − ∑

× ( ) =ASSAI

N number of hours Res

N number of hours 16 S

S

i 1 N

i S

where Resi represents the restoration time for the i-th substation if service interruption exists and NS represents the total number of substations. The Energy Not Served (ENS) quantifies the total en- ergy that could not be supplied to the customers due to line dis- connections:

∑= ( −

) × ∆ [ ]

( ) ∆

⎡ ⎣ ⎢ ⎢

⎤ ⎦ ⎥ ⎥ENS

total power demand

actual power demand served t MWh

17i t

i

i i

where Δti are the time intervals during which some power could not be supplied.

4.2. Development of Resilience Improvement Strategies

Different strategies are developed and tested, each of which mainly focuses on the enhancement of a specific system resilience capability. The improvement of the efficiency of line reparation (strategy 1) enhances the restorability capability during the re- covery phase, the improvement of the human operator perfor- mance (strategy 2) enhances the adaptive capability during the recovery phase, and the improvement of RTU battery capacity (RBC) (strategy 3) enhances the absorptive capability during the disruptive phase. The target system for each strategy also varies: SUC is the target system for strategy 1 and 2, and SCADA is the target system for strategy 3.

In order to demonstrate the feasibility of the proposed method and the applicability of the resilience improvement strategies, various simulation experiments are developed. The summary of the adopted resilience improvement strategies and the system parameters which are varied in each strategy are shown in Table 3. In total, 18 experiments are set for strategy 1 and 2, resulting from the combination of varying parameters in the two strategies

elated parameters Simulation scenarios Simulation stops criteria

TTR (h) 0.5,1,1.5,2,2.5,3 CV(GRSUC)r0.13 EPA 0.03, 0.3, 1 CV(GRSUC)r0.13 C (min) 5,10,15,20,25,30 CV(GRSCADA)r0.13

an time to repair; HEPA: maximum acceptable human error probability; RBC: RTU

Fig. 7. The average performance of SUC and SCADA (MTTR¼2.5 h, HEPA¼0.03, RBC¼10 min, N¼16). The y-axis denotes the MOPSUC1 (i.e. the number of connected lines), MOPSUC2 (i.e. actual power demand served), and MOPSCADA1 (i.e. the number of available RTU devices); the x-axis denotes time.

C. Nan, G. Sansavini / Reliability Engineering and System Safety 157 (2017) 35–53 45

targeting SUC; 6 experiments are set for strategy 3. The number of simulation runs (N) for each experiment is determined by the coefficient of variation (CV) of the resilience measure of the target system. CV is defined as the ratio of the standard deviation s to the mean m [62]. In this case study, CV(GRTargetSystem)r0.13 is the stopping criteria to determine the number of runs for each

Fig. 8. The average system performance under different experiments with varying MTTR axis denotes the MOPSUC1 (i.e. the number of connected lines) and the x-axis denotes ti

computer experiment to balance computational efforts and accu- racy of results. With no loss of generality, RBC is set to 10 min for the simulations of strategy 1 and 2, and MTTR is set to 2 h and HEPA is set to 0.03 for the simulations of strategy 3. A total of N¼485 (E240 h) is performed.

(HEPA¼0.3, RBC¼10 min, N¼{18,10,11,10,13,11} for MTTR from 0.5 h to 3 h). The y- me.

Fig. 9. The average system performance under different experiments with varying MTTR (HEPA¼0.3, RBC¼10 min, N¼{15,11,11,10,13,11} for MTTR from 0.5 h to 3 h). The y-axis denotes the MOPSUC2 (i.e. actual power demand served) and the x-axis denotes time.

Fig. 10. The average system performance under different experiments with varying MTTR (HEPA¼0.3, RBC¼10 min, N¼{15,11,11,10,13,11}for MTTR from 0.5 h to 3 h). The y-axis denotes the power load (MW) and the x-axis denotes time.

C. Nan, G. Sansavini / Reliability Engineering and System Safety 157 (2017) 35–5346

5. Simulation results

Fig. 7 shows the performance loss and recovery of SUC and SCADA following the disruptive event defined in Section 4.1 under the simulation scenario, in which MTTR is set to 2.5 h, HEPA is set to 0.03, and RBC is set to 10 min. At t¼3 h, the disruptive event is triggered. The negative effects appear immediately, i.e. the MOPSUC1 value drops after the first line disconnection. As custo- mers are disconnected from the network, the actual power de- mand served also drops, as shown by the decrease of the value of the MOPSUC2. Due to power shortage from the disconnection of transmission lines and RTU batteries consumption, the MOPSCADA1 value drops as well after a certain time, i.e. 9 min in this case. This period is also referred as the delay of dependency failures, which is an important factor for minimizing negative effects caused by in- terdependencies among systems [43]. About 30 min after the shock, MOPSUC1 reaches the lowest level RSUC(MOP1)¼91.6%, i.e. 19 lines are disconnected. The further disconnection of lines beyond

the initiating scenario is due to cascading failures in the power grid as mentioned in Section 4.1. About 40 min after the shock, the MOPSCADA1 value also reaches its lowest level, i.e. RSCADA (MOP1)¼94.2%, due to disconnection of corresponding transmission lines and battery depletion. Both systems start recovering after reaching their lowest performance levels as the result of the re- pairing actions and reach the new steady state after a certain period.

As demonstrated in Fig. 7, the performances of SUC and SCADA show similar trends, indicating how tightly they are inter- dependent with each other. Interdependencies among two sys- tems may have direct impacts on resilience capabilities. In this section, simulation results related to the performance of each system under corresponding improvement strategies will be firstly analyzed with the aim of quantifying their effects on system re- silience. Furthermore, the effects of interdependencies on resi- lience capabilities will also be discussed. Finally, resilience and performability metrics will be compared.

Table 4 Summary of the overall simulation results for different computer experiments re- lated to strategy 1 and 2 using MOPSUC1.

MTTR HEPA GR (SUC) ASSAI Disruptive phase Recovery phase

R PLDP RAPIDP (1/h)

PLRP RAPIRP (1/h)

0.5 h 0.03 16.33 0.993 0.916 0.026 0.44 0.09 0.29 0.3 19.15 0.995 0.921 0.0076 0.45 0.079 0.31 1 20.82 0.995 0.921 0.0075 0.45 0.083 0.32

1 h 0.03 15.90 0.986 0.916 0.027 0.44 0.14 0.27 0.3 18.67 0.990 0.921 0.0073 0.46 0.12 0.30 1 20.40 0.991 0.921 0.0073 0.46 0.12 0.30

1.5 h 0.03 14.98 0.982 0.916 0.026 0.45 0.18 0.28 0.3 16.29 0.984 0.921 0.0072 0.47 0.15 0.29 1 19.26 0.985 0.921 0.0072 0.45 0.16 0.29

2 h 0.03 14.04 0.977 0.916 0.026 0.44 0.22 0.26 0.3 15.98 0.984 0.921 0.0072 0.47 0.18 0.29 1 18.03 0.985 0.921 0.0072 0.47 0.17 0.29

2.5 h 0.03 13.76 0.976 0.916 0.026 0.45 0.24 0.27 0.3 15.74 0.980 0.921 0.0075 0.46 0.20 0.28 1 17.16 0.981 0.921 0.0072 0.47 0.23 0.28

3 h 0.03 13.73 0.973 0.916 0.026 0.45 0.25 0.25 0.3 14.86 0.975 0.921 0.0072 0.47 0.27 0.26 1 16.38 0.975 0.921 0.0075 0.45 0.27 0.27

MTTR: mean time to repair; HEPA: maximum acceptable human error probability; GR: general resilience; ASSAI: average substation service availability index R: ro- bustness PLDP: performance loss in disruptive phase; PLRP: performance loss in recovery phase RAPIDP: rapidity in disruptive phase RAPIRP: rapidity in recovery phase.

Table 5 Summary of the overall simulation results for different computer experiments re- lated to strategy 1 and 2 using MOPSUC2.

MTTR HEPA GR (SUC) ENS (MWH)

Disruptive phase Recovery phase

R PLDP RAPIDP (1/h)

PLRP RAPIRP (1/h)

0.5 h 0.03 3.47 1399 0.855 0.053 0.61 0.13 0.16 0.3 4.23 934 0.868 0.035 0.77 0.087 0.21 1 4.97 850 0.868 0.040 0.71 0.07 0.24

1 h 0.03 2.65 2066 0.851 0.057 0.55 0.20 0.11 0.3 3.02 1454 0.866 0.048 0.66 0.14 0.13 1 3.38 1463 0.867 0.056 0.55 0.13 0.13

1.5 h 0.03 2.14 2630 0.852 0.057 0.55 0.27 0.090 0.3 3.02 1669 0.867 0.054 0.56 0.16 0.11 1 3.18 1675 0.867 0.057 0.52 0.16 0.13

2 h 0.03 1.83 2869 0.850 0.059 0.56 0.30 0.071 0.3 2.23 2007 0.868 0.045 0.60 0.21 0.081 1 2.41 1832 0.868 0.041 0.64 0.19 0.086

2.5 h 0.03 1.79 3071 0.851 0.055 0.57 0.32 0.077 0.3 2.10 2424 0.867 0.053 0.58 0.23 0.073 1 2.17 2499 0.866 0.049 0.56 0.27 0.065

3 h 0.03 1.63 3135 0.852 0.056 0.59 0.33 0.066 0.3 1.98 3032 0.867 0.071 0.52 0.30 0.058 1 2.02 3074 0.867 0.053 0.53 0.31 0.049

MTTR: mean time to repair; HEPA: maximum acceptable human error probability; GR: general resilience; ENS: energy not served; ASSAI: average substation service availability index; R: robustness; PLDP: performance loss in disruptive phase; PLRP: performance loss in recovery phase; RAPIDP: rapidity in disruptive phase; RAPIRP: rapidity in recovery phase.

3 In order to demonstrate the statistical significance of these comparisons, the p–value is calculated for the two-sample (unequal variance) t-test, where H0: mean1¼mean2 and H1: mean14mean2.

C. Nan, G. Sansavini / Reliability Engineering and System Safety 157 (2017) 35–53 47

5.1. SUC

Fig. 8 shows the performance of SUC following the disruptive event under various simulation scenarios related to strategy 1 using MOPSUC1, in which MTTR varies from 0.5 h to 3 h.

About 12 min after the shock, the MOP value reaches its lowest level, i.e. RSUC¼92.1%. The MOP value then increases as the result of the repair actions. As shown in Fig. 8, as MTTR increases, the time for the SUC to reach new steady state also increases. For instance, tns¼6.4 h when MTTR is set to 0.5 h, while tns¼11.4 h when MTTR is set to 3 h. Strategy 1 has a strong impact on the restorability of SUC during the recovery phase. Fig. 9 illustrates the performance of SUC under various simulation scenarios related to strategy 1 using the functionality-related MOPSUC2. Compared to Fig. 8, similar patterns are observed. For instance, as MTTR increases, the time for the SUC to reach new steady state also increases.

Fig. 10 shows how the amount of actual demand served is af- fected due to the variation of MTTR. As shown in Fig. 10, the dis- ruptive event has the influence on the functionality (actual de- mand served) of the system. As MTTR increases, larger deviations between total power demand and actual demand served can be observed, indicating that disruption of the power supply also in- creases. The value of energy not served (ENS) increases from 934 MW h to 3032 MW h as MTTR increases from 0.5 h to 3 h.

The results for strategy 1 and 2 using MOPSUC1 and MOPSUC2 are summarized in Tables 4 and 5, respectively. Strategy 2 has a stronger impact on the absorptive capability of the target system, i.e. SUC. As HEPA increases (from 0.03 to 1), system robustness R increase according to both MOPs. For example, RSUC(MOP1) ¼ 91.6% when HEPA is set to 0.03 and increases to 92.1% when HEPA is set to 1, while the RSUC(MOP2) increases from 85.5% to 86.8%. In addition, the value of ENS also decreases due to the improvement of human performance, as shown in Fig. 11.

Strategy 1 has little impact on the absorptive capability of the target system, i.e. SUC. As shown in Tables 4 and 5, all the mea- surements related to this capability, i.e., R, PLDP, RAPIDP, do not vary significantly for results of experiments with same HEPA value but different MTTR value. On the other hand, strategy 1 has a positive impact on the restorative and adaptive capabilities in the recovery phase. For instance,3 PLRP (MOPSUC1) decreases from 0.25 to 0.09 (p-value¼10�8) and RAPIRP (MOPSUC1) increases from 0.25 to 0.29 h�1 (p-value¼10�6), as MTTR decreases from 3 h to 0.5 h (HEPA¼ 0.03); PLRP (MOPSUC2) decreases from 0.33 to 0.13 (p- value¼10�8) and RAPIRP (MOPSUC2) increases from 0.066 to 0.16 h�1 (p-value¼10�10) in the same scenario. These results quantify the positive effects on resilience stemming from im- provements in repair efficiency. Strategy 2 also has a positive impact on the adaptive and restorative capabilities during re- covery phase when the operator performance improves from poor level (HEPA¼0.03) to average/acceptable level (HEPA¼0.3). This positive impact seems not clear when the operator performance continuously improves after reaching to average/acceptable level. Compared to strategy 1, strategy 2 enhances the absorptive cap- ability during disruptive phase more significantly. For example, if HEPA¼1, PLDP(MOPSUC1) drops to the range [0.0072, 0.0075] and PLDP(MOPSUC2) drops to the range [0.04, 0.057], indicating that performance of SUC is less impacted during disruptive phase if the human operator performance is improved.

In order to assess whether these strategies are able to enhance the overall resilience capability, the integrated resilience metric, GR, must be quantified. Figs. 12 and 13 illustrate the value of GRSUC, i.e. the resilience of SUC to the disruptive event, with respect to the strategy 1 and 2, using MOPSUC1 and MOPSUC2 respectively. When both strategies are implemented simultaneously, the resilience of SUC can be enhanced significantly. The results confirm the con- sistency of the model and of the resilience metric. Furthermore,

Fig. 11. The average system performance under different experiments with varying HEPA (MTTR¼2.5 h, RBC¼10 min, N¼{12,12,10} for HEPA from 0.03 to 1). The y-axis denotes the power load (MW) and the x-axis denotes time.

Fig. 12. The GRSUC value based on MOPSUC1 under different simulation scenarios for strategy 1 and 2.

Fig. 13. The GRSUC value based on MOPSUC2 under different simulation scenarios for strategy 1 and 2.

C. Nan, G. Sansavini / Reliability Engineering and System Safety 157 (2017) 35–5348

they allow comparing the relative benefits of different improve- ment strategies. Using the functionality related MOP as an ex- ample (Fig. 13), GRSUC(MOPSUC2)¼3.02 when MTTR¼1 h and HEPA¼0.3. At this point, if the efficiency of reparation is further improved (MTTR decreases to 0.5 h), GR metric indicates 40%

resilience increase (GRSUC(MOPSUC2)¼4.23). On the other hand, if the human operator performance is further improved (HEPA in- creases to 1), GR metric indicates 12% resilience increase (GRSUC(MOPSUC2)¼3.38). If both strategies are implemented, GR metric indicates 64% resilience increase (GRSUC(MOPSUC2)¼4.97).

Fig. 14. SCADA performance under different experiments with varying RBC (MTTR¼2 h, HEPA¼0.3, N¼{10,10,12,13,10,12} for RBC from 5 to 30 min). The y-axis denotes the MOPSCADA1 (i.e. the number of available RTU devices) and the x-axis denotes the time.

Table 6 Summary of the overall simulation results of strategy 3.

RBC (min)

GR (SCADA) ASSAI Disruptive phase Recovery phase

R PLDP RAPIDP (1/ h)

PLRP RAPIRP (1/ h)

5 28.65 0.979 0.940 0.017 0.48 0.41 0.14 10 29.97 0.977 0.940 0.017 0.50 0.40 0.13 15 30.80 0.978 0.941 0.014 0.51 0.39 0.12 20 32.45 0.980 0.942 0.013 0.50 0.38 0.10 25 35.00 0.981 0.947 0.0040 0.45 0.40 0.11 30 38.85 0.980 0.948 0.0040 0.47 0.39 0.12

RBC: RTU battery capacity; GR: general resilience; CV: coefficient of variation; R: robustness; PLDP: performance loss in disruptive phase; PLRP: performance loss in recovery phase; RAPIDP: rapidity in disruptive phase; RAPIRP: rapidity in recovery phase.

C. Nan, G. Sansavini / Reliability Engineering and System Safety 157 (2017) 35–53 49

The best selection of improvement strategies can be determined based on GR and on the implementation costs.

The similarities in the classification of system resilience are assessed to verify the consistency of MOPSUC1 and MOPSUC2 with respect to the different simulation scenarios. To this aim, 18 items of the rankings are grouped in three batches [6, 6, 6] of decreasing resilience, i.e., the first batch encompasses the six most resilient scenarios, the second batch encompasses the following six most resilient scenarios, and the third batch encompasses the six least resilient scenarios. Then, the rankings provided by MOPSUC1 and MOPSUC2 are compared looking for similarities in the batches of equal resilience and a degree of similarity is evaluated. For each one of the three couples of batches, the occurrences of the same scenario are enumerated. The occurrences of different items are added up to build a synthetic index which then is scaled in the interval [0%, 100%] through division by the maximum number of repeated occurrences: 100% similarity indicates that the same scenarios are included in the batches for various MOP; 0% simi- larity indicates that for different MOP the batches of equal resi- lience contain diverse components. The degree of similarity in the rankings provided by MOPSUC1 and MOPSUC2 for the three batches is [83%, 67%, 83%]. Therefore, MOPSUC1 and MOPSUC2 consistently rank

the most and the least resilient scenarios. In particular, the two rankings identify scenario MTTR ¼ 0.5 h & HEPA¼ 0.03 as the most resilient one.

The simulation results also demonstrate the consistency of the GR metric. Selecting two MOPs for the quantification of resilience, i.e. a topological and a functional MOP, leads to same conclusions about system resilience.

5.2. SCADA

Fig. 14 shows the performance of SCADA following the dis- ruptive event under various computer experiments related to strategy 3, in which the RBC varies from 5 to 30 min (HEPA is set to 0.3 and MTTR is set to 2 h).

As shown in Fig. 14, the transition from original steady phase to disruptive phase is significantly delayed due to the increase of the RBC value, affecting the system absorptive capability. The R value increases from 0.94 to 0.948 when RBC value is increased from 5 to 30 min, indicating the robustness of SCADA system is enhanced as a result of the increase of RTU battery capacity. The overall si- mulation results for strategy 3 are summarized in Table 6.

As shown in Table 6, strategy 3 enhances SCADA system robust- ness during the disruption phase, and, moreover, the system per- formance is less impacted during the disruptive phase. For instance, PLDP ¼0.017 when RBC is set to 5 min, while PLDP drops to 0.004 when RBC is set to 30 min. Compared to its strong impacts on the enhancement of absorptive capability, the impact on the adaptive and restorative capability of the system during recovery phase is unclear. According to the results shown in Tables 4 and 6, the SCADA is more resilient than SUC for the selected disruptive event, i.e. winter storm. The range of GRSCADA is [28.65, 38.83] and the range of GRSUC(MOPSUC1) value is [13.73, 20.82]. It should be noted that the two resilience metrics are calculated using topology-related MOP measures. The reason for this phenomenon might be that the se- lected disruptive event directly affects the performance of SUC, which in turn impacts on the performance of SCADA. The service disruption of SCADA is mainly caused by the second-order effects of the disconnection of the transmission lines of SUC, and not directly by the disruptive event, therefore, causing a less negative impact.

Fig. 15. Comparisons of effects of physical and informational dependency with respect to resilience (inset 1 and 4), performance loss in the disruptive phase (inset 2 and 5), and performance loss in the recovery phase (inset 3 and 6). The trendline in each inset represents linear correlation.

Table 7 Degree of similarity in the rankings provided by the two resilience and the two performability metrics under different simulation scenarios for strategy 1 and 2.

Degree of similarity Batch MOPSUC1–ASSAI MOPSUC1–ENS MOPSUC2–ASSAI MOPSUC2–ENS

6 67 67 83 83 6 50 50 83 67 6 83 83 100 83

C. Nan, G. Sansavini / Reliability Engineering and System Safety 157 (2017) 35–5350

5.3. Effects of interdependencies In this case study, the two technical systems (i.e. SCADA and

SUC) are dependent on each other for their operations. SCADA is

dependent upon SUC for the power supply (physical dependency), while SUC is dependent upon SCADA for the transmission of control actions and measurements (informational dependency). It should be noted that the interdependency between SCADA and SUC can be referred as internal interdependency since both two systems are included in one system, i.e. EPSS. This paper uses the general term interdependency, which includes both internal and external interdependency. Interdependencies between the two systems have direct impacts on resilience capabilities. These ef- fects are apparent when considering the resilience improving strategies. For example, strategy 1 and 2 target SUC and directly improve its resilience. On the other hand, the same strategies also

Fig. 16. The ASSAI under different simulation scenarios for strategy 1 and 2.

Fig. 17. The ENS under different simulation scenarios for strategy 1 and 2.

C. Nan, G. Sansavini / Reliability Engineering and System Safety 157 (2017) 35–53 51

affect and improve the resilience of SCADA system due to its physical dependencies with SUC. The effects of interdependency on resilience capabilities are estimated by cross-comparing the performances of the two interdependent systems, illustrated in Fig. 15.

In Fig. 15, the values of PL in the disruptive and recovery phases and the value of GR for SUC and for SCADA calculated using to- pology-related MOPs (MOPSUC1 and MOPSCADA1) are employed for comparing the impacts of interdependencies on the resilience capabilities. The effects of physical dependency between SUC and SCADA are shown in the left insets of Fig. 15 using the results related to strategy 1 and 2 (N¼329), and effects of informational dependency between SCADA and SUC are shown in the right insets of Fig. 15 using the results related to strategy 3 (N¼97). The impact of interdependency on the resilience capabilities and on overall system resilience can be quantified in terms of the degree of cor- relation measured by the correlation coefficient, and in terms of the coupling strength measured by the slope of the trendline. The implicit assumption is that the effects of interdependencies can be described to a first approximation by a linear relationship. The correlation coefficient indicates a measure of the linear correlation between two groups of variables, giving a value between 1 and �1 inclusive, where 1 and �1 is total positive and negative correla- tion respectively, 0 is no correlation. This measurement is used to indicate the extent to which resilience capabilities of the

interdependent systems are dependent on each other. Larger va- lues indicate that actions affecting a capability of the system, e.g., performance loss, will also affect the same capability of the other system that is dependent upon it.

According to Fig. 15, the degree of correlation among resilience capabilities for physical dependency is larger than for informa- tional dependency. For instance, in Fig. 15 inset 2 and 3, the cor- relation coefficient values for PLDP and PLRP are 0.996 and 0.918 respectively, indicating to what extent the performance of two systems during two phases are correlated due to physical de- pendency. Fig. 15 inset 4–6 also shows the relatively weak degree of correlation among resilience capabilities due to informational dependency except for the performance loss during restorative phase (correlation coefficient¼0.733), which demonstrates the importance of the SCADA system in terms of decreasing the per- formance loss of the SUC during the restorative phase. Both phy- sical and informational dependencies have non-negative effects on all the resilience capabilities based on the positive number related slope of trendlines shown in Fig. 15, i.e. the enhancement of these capabilities of one system has a positive impact on the same capabilities of the other one if the dependency exists between them.

Although the degree of correlation between the overall system resilience due to both types of dependencies is moderate accord- ing to the correlation coefficient (0.246 and 0.307 for physical and

C. Nan, G. Sansavini / Reliability Engineering and System Safety 157 (2017) 35–5352

informational dependency), the corresponding value of slope of trendline shows that the coupling due to physical dependency is much stronger than for informational dependency (0.551 vs 0.128). Therefore, the effects of enhancing the resilience of one system have a much more significant impact on an interdependent system in case physical dependencies are present. The correlation plot in Fig. 15 provides non-trivial insights regarding the ad- vantage of using one measure instead of another as well as re- garding their combined use. The quantification of the effects of interdependency is conditioned on the selection of the disruptive event, of the measure of performance, and of the different stra- tegies. Furthermore, the cost for the implementation of the mea- sures is not accounted for this case study, e.g., time for changes, required competence, and possible unknown risks added. The full characterization and quantification of the coupling strength be- tween interdependent systems require broad and focused in- vestigation which will be the object of further targeted studies.

5.4. Resilience metric vs performability metric Table 7 shows the degree of similarity in the rankings provided

by GRSUC(MOPSUC1), GRSUC(MOPSUC2) and the two performability metrics, i.e. ASSAI and ENS, evaluated as described in Section 5.1. Table 7 shows that performability and resilience are correlated, especially in case the functionality-related MOPSUC2 is adopted. GRSUC(MOPSUC2), ASSAI and ENS consistently rank the resilient scenarios. In particular, all three rankings identify the same three most resilient and five least resilient scenarios.

Figs. 16 and 17 illustrate the ASSAI and the ENS value for the SUC after the disruptive event with respect to strategy 1 and 2 using the 3D surface diagram.

Comparing Figs. 12 and 16, ASSAI shows a trend which is similar to the trend of the topology-related resilience metric GRSUC(MOPSUC1), indicating that both improvement strategies have non-negative effects on performability and resilience metrics. Figs. 13 and 17 consistently display the improvements due to the implementation of strategy 1 and 2 as quantified by ENS and by the functionality-related resilience metric GRSUC(MOPSUC2). ENS and GRSUC(MOPSUC2) are more sensitive to variations of MTTR and HEPA because they take into account the magnitude of the power not supplied to the customers.

6. Conclusion

A quantitative method aiming at the assessment of resilience in infrastructures is introduced in this paper. Within this method, three essential resilience capabilities are identified (i.e. absorptive, adaptive, and restorative capabilities) as well as different measures and phases related to these capabilities. The method includes two components, an integrated metric for resilience quantification, and a hybrid multi-layer modeling approach to capture and quantify the system performance across interdependencies. Three resi- lience improvement strategies are developed, each of which mainly focuses on the enhancement of a specific resilience cap- ability (i.e. improvement of the repair efficiency, improvement of human operator performance, and improvement of RTU battery capacity) of its target system.

The consistency of the proposed method is benchmarked with the application to an EPSS. The method is then applied to a case study considering system improvement strategies. The results provide non-trivial insights to decision makers, i.e. improving re- pair efficiency brings the largest enhancement in the system adaptive and restorative capability. On the other hand, improving human operator performance and RTU battery capacity has the largest impact in enhancing the system absorptive capability.

The hybrid multi-layer approach proves effective in capturing

the effects of interdependencies and in quantifying the coupling strength between interdependent infrastructures by modeling their functional behaviors. The simulation results also show how tightly the SUC and SCADA within the EPSS are interdependent with each other, and indicate how two types of dependencies impact resilience: the physical dependency has more significant impacts on system resilience than informational dependency. The developed resilience metric shows similarity to performability metrics, i.e. ASSAI and ENS, especially in case the functionality- related perspective on resilience is adopted. Nonetheless, it adds an additional perspective on system behavior by gaining insights into system performance in different phases and at the transition points among them. Therefore, resilience metrics can be a useful complement to performability metrics.

The results show the benefits of the proposed approach: dif- ferent layers of CI can be represented using appropriate methods which can capture their operations and physical and functional diversity. Within this approach, improvements in different sub- systems can be compared with respect to their effects on the performance of the entire CI. This approach is scalable, i.e. if the size of infrastructures and the complexity of research tasks in- crease, additional computational resources can make it manageable.

Future research work will apply the developed framework to assess the resilience of the EPSS to a group of multiple hazards besides the only event used in the paper. Furthermore, the concept of resilience costs, i.e. impact cost due to disruptive events and recovery costs, will be considered in order to improve the ap- plicability of the resilience metric. Trade-offs between the cost of loss of services (e.g., power loss) and the cost of resilience mea- sures need also to be considered.

References

[1] Kröger W, Zio E. Vulnerable systems. London: Springer-Verlag; 2011. [2] Kröger W. Critical infrastructure at risk: a need for a new conceptual approach

and extended analytical tools. Reliab Eng Syst Saf 2008;93:1781–7. [3] Eusgeld I, Kröger W, Sansavini G, Schläpfer M, Zio E. The role of network

theory and object-oriented modeling within a framework for the vulnerability analysis of critical infrastructures. Reliab Eng Syst Saf 2009;94:954–63.

[4] Rinaldi SM, Peerenboom JP, Kelly TK. Identifying, understanding, and analyz- ing critical infrastructure interdependencies. IEEE Control Syst 2001;21:11–25.

[5] US – Canada Power System Outage Task Force. Final Report on the August 14, 2003 Blackout in the United States and Canada: Causes and Recommenda- tions; 2004.

[6] Brummitt CD, D’Souza RM, Leicht EA. Suppressing cascades of load in inter- dependent networks. Proc Natl Acad Sci 2012;109:E680–9.

[7] Buldyrev SV, Parshani R, Paul G, Stanley HE, Havlin S. Catastrophic cascade of failures in interdependent networks. Nature 2010;464:1025–8.

[8] Vespignani A. Complex networks: the fragility of interdependency. Nature 2010;464:984–5.

[9] Leavitt WM, Kiefer JJ. Infrastructure interdependency and the creation of a normal disaster the case of hurricane Katrina and the city of New Orleans. Public Works Manag Policy 2006;10:306–14.

[10] Ouyang M, Dueñas-Osorio L, Min X. A three-stage resilience analysis frame- work for urban infrastructure systems. Struct Saf 2012;36–37:23–31.

[11] Fisher R, Bassett G, Buehring W, Collins M, Dickinson D, Eaton L, et al. Con- structing a Resilience Index for the Enhanced Critical Infrastructure Protection Program. Decision and Information Sciences Division, Argonne National La- boratory, Department of Energy, United States; 2010.

[12] Madni AM, Jackson S. Towards a conceptual framework for resilience en- gineering. IEEE Syst J 2009;3:181–91.

[13] Decò A, Bocchini P, Frangopol DM. A probabilistic approach for the prediction of seismic resilience of bridges. Earthq Eng Struct Dyn 2013;42:1469–87.

[14] Linkov I, Creutzig F, Decker J, Fox-Lent C, Kröger W, et al. Risking resilience: changing the paradigm. Nat Clim Change 2014;4:407–9.

[15] Holling CS. Resilience and stability of ecological systems. Annu Rev Ecol Syst 1973:1–23.

[16] Adger WN. Social and ecological resilience: are they related? Prog Hum Geogr 2000;24:347–64.

[17] Pant R, Barker K. Building dynamic resilience estimation metrics for inter- dependent infrastructures. ESREL 2012. Helsinki, Finland; 2012.

[18] Hollnagel E. Resilience engineering. Hampshire, England: Ashgate; 2006. [19] ⟨http://www.resalliance.org/⟩.

C. Nan, G. Sansavini / Reliability Engineering and System Safety 157 (2017) 35–53 53

[20] Haimes YY. On the definition of resilience in systems. Risk Anal 2009;29:498– 501.

[21] McCarthy JA, Pommering C, Perelman LJ, Scalingi PL, Garbin DA, Shortle JF, et al. Critical thinking: Moving from infrastructure protection to infrastructure resilience. George Mason University CIP Program Discussion Paper Series; 2007; p. 97–109.

[22] McDaniels T, Chang S, Cole D, Mikawoz J, Longstaff H. Fostering resilience to extreme events within infrastructure systems: characterizing decision con- texts for mitigation and adaptation. Glob Environ Change 2008;18:310–8.

[23] Fiksel J. Designing resilient, sustainable systems. Environ Sci Technol 2003;37:5330–9.

[24] Bruneau M, Chang SE, Eguchi RT, Lee GC, O’Rourke TD, Reinhorn AM, et al. A framework to quantitatively assess and enhance the seismic resilience of communities. Earthq Spectra 2003;19:733–52.

[25] Gao J, Buldyrev SV, Havlin S, Stanley HE. Robustness of a network of networks. Phys Rev Lett 2011;107:195701.

[26] Brand K-P, Lohmann V, Wimmer W. Substation automation handbook: Utility Automation Consulting Lohmann; 2003.

[27] Vugrin E, Warren D, Ehlen M, Camphouse RC. A Framework for assessing the resilience of infrastructure and economic systems. In: Gopalakrishnan K, Peeta S, editors. Sustainable and resilient critical infrastructure systems. Berlin, Heidelberg: Springer; 2010. p. 77–116.

[28] Bruneau M, Chang SE, Eguchi RT, Lee GC, O'Rourke TD, Reinhorn AM, et al. A framework to quantitatively assess and enhance the seismic resilience of communities. Earthq Spectra 2003;19:733–52.

[29] Chang SE, Shinozuka M. Measuring Improvements in the Disaster Resilience of Communities. Earthq Spectra 2004:20.

[30] Directive PP. Critical infrastructure security and resilience. PPD-21; 2013. [31] Bueno NP. Assessing the resilience of small socio-ecological systems based on

the dominant polarity of their feedback structure. Syst Dyn Rev 2012;28:351– 60.

[32] Filippini R, Silva A. A modeling framework for the resilience analysis of net- worked systems-of-systems based on functional dependencies. Reliab Eng Syst Saf 2013.

[33] Alessandri A, Filippini R. Evaluation of resilience of interconnected systems based on stability analysis. In: Hämmerli B, Kalstad Svendsen N, Lopez J, editors. Critical information infrastructures security. Berlin, Heidelberg: Springer; 2013. p. 180–90.

[34] Henry D, Emmanuel Ramirez-Marquez J. Generic metrics and quantitative approaches for system resilience as a function of time. Reliab Eng Syst Saf 2012;99:114–22.

[35] Nan C., Sansavini G., Kroeger W., Heinimann H.R. A quantitative method for assessing the resilience of infrastructure systems. In: Proceedings of the 12th Probabilistic safety assessment & management conference (PSAM12). Hono- lulu, USA; 2014.

[36] Ferreira C, Gama J, Miranda V, Botterud A. Probabilistic ramp detection and forecasting for wind power prediction. In: Billinton R, Karki R, Verma AK, editors. Reliability and risk evaluation of wind integrated power systems. In- dia: Springer; 2013. p. 9–44.

[37] Kamath C. Understanding wind ramp events through analysis of historical data. Transm Distrib Conf Expo 2010:1–6.

[38] Zheng H, Kusiak A. Prediction of wind farm power ramp rates: a data-mining approach. J Sol Energy Eng 2009;131:31011.

[39] Francis R, Bekera B. A metric and frameworks for resilience analysis of en- gineered and infrastructure systems. Reliab Eng Syst Saf 2014;121:90–103.

[40] Griot C. Modelling and simulation for critical infrastructure interdependency assessment: a meta-review for model characterisation. Int J Crit Infrastruct 2010;6:363–79.

[41] Pederson P, Dudenhoeffer D, Hartly S, Permann M. Critical infrastructure in- terdependency modeling: A survey of US and international research. Idaho National Laboratory; 2006.

[42] Zio E, Sansavini G. Modeling Interdependent Network Systems for Identifying Cascade-Safe Operating Margins. IEEE Trans Reliab 2011;60:94–101.

[43] Nan C, Eusgeld I, Kröger W. Analyzing vulnerabilities between SCADA system and SUC due to interdependencies. Reliab Eng Syst Saf 2013;113:76–93.

[44] Nan C, Eusgeld I, Adopting HLA. standard for interdependency study. Reliab Eng Syst Saf 2011;96:149–59.

[45] Zio E, Sansavini G. Vulnerability of smart grids with variable generation and consumption: a system of systems perspective. IEEE Trans Syst Man Cybern: Syst 2013;43:477–87.

[46] Bilis EI, Kroger W, Nan C. Performance of electric power systems under phy- sical malicious attacks. Syst J IEEE 2013;7:854–65.

[47] Chappin EJL, Dijkema GPJ. Agent-based modeling of energy infrastructure transitions. In: Proceedings of the first international conference on infra- structure systems and services: Building Networks for a Brighter Future (IN- FRA); 2008. p. 1–6.

[48] Tolk A, Uhrmacher AM. Agents: Agenthood, agent architectures, and agent taxonomies. Agent-Directed Simulation and Systems Engineering: Wiley-VCH Verlag GmbH & Co. KGaA; 2010. p. 73–109.

[49] Sharit J. Human error and human reliability analysis. Handbook of human factors and ergonomics, fourth edition; 2012. p. 734–800.

[50] Wood AJ, Wollenberg BF. Power generation, operation, and control. New York, USA: John Wiley & Sons; 2012.

[51] Schläpfer M, Kessler T, Kröger W. Reliability analysis of electric power systems using an object-oriented hybrid modeling approach. In: Proceedings of the 16th power systems computation conference. Glasgow; 2008.

[52] Nan C, Eusgeld I. Exploring impacts of single failure propagation between SCADA and SUC. In: Proceedings of the IEEE international conference on in- dustrial engineering and engineering management (IEEM); 2011. p. 1564–8.

[53] Hollnagel E. Cognitive reliability and error analysis method CREAM. Am- sterdam, The Netherlands: Elsevier; 1998.

[54] Nan C, Sansavini G. Developing an agent-based hierarchical modeling ap- proach to assess human performance of infrastructure systems. Int J Ind Ergon 2016;53:340–54.

[55] Nan C, Eusgeld I. Exploring impacts of single failure propagation between SCADA and SUC. In: Proceedings of the IEEE international conference on in- dustrial engineering and engineering management (IEEM); 2011. p. 1564–8.

[56] Bilis E, Raschke M, Kroeger W. Seismic response of the swiss transmission grid. Proc ESREL 2010:5–9.

[57] Raschke M, Bilis E, Kröger W. Vulnerability of the Swiss electric power transmission grid against natural hazards. Applications of Statistics and Probability in Civil Engineering: CRC Press; 2011. p. 1407–14.

[58] Cepin M. Relaibillity of power systems-catastrophic sleet and ice rain in Slo- venia 2014. ESRA (European Safety Reliability Association) Newsletter; March 2014.

[59] Samenow J. Positively caked in ice, crazy storm photos from Slovenia. The Washington Post; 2014 https://www.washingtonpost.com/news/capital- weather-gang/wp/2014/02/07/craziest-ice-storm-photos-youll-ever-seen- from-slovenia-photos/.

[60] Bründl M, Rickli C. The storm Lothar 1999 in Switzerland–an incident analysis. Snow Land Res 2002;77:2.

[61] Lachs WR. Transmission-line overloads: real-time control. Generation, Trans- mission and Distribution. IEE Proc C 1987;134:342–7.

[62] George CR, Norma Faris H, Douglas Carter, Montgomery EE. Engineering sta- tistics.New York, NY: Wiley; 2004.

  • A quantitative method for assessing resilience of interdependent infrastructures
    • Introduction
      • Critical infrastructures and resilience
      • Resilience capabilities
      • Frameworks for assessing resilience
    • Method component 1: development of an integrated resilience metric
      • Quantification of resilience capabilities
      • The integrated resilience metric
    • Method component 2: multi-layer hybrid modeling approach
      • Modeling the EPSS system using the hybrid multi-layer approach
        • Step 1: screening analysis
      • Step 2: individual model development
        • Step 3: model interaction
    • Case study for resilience assessment
      • Design of the experiment
      • Development of Resilience Improvement Strategies
    • Simulation results
      • SUC
      • SCADA
        • Effects of interdependencies
        • Resilience metric vs performability metric
    • Conclusion
    • References