Review on Energy Resilience
Reliability Engineering and System Safety 117 (2013) 89–97
Contents lists available at SciVerse ScienceDirect
Reliability Engineering and System Safety
0951-83 http://d
n Corr E-m
Jose.Ram
journal homepage: www.elsevier.com/locate/ress
Resilience-based network component importance measures
Kash Barker a, Jose Emmanuel Ramirez-Marquez b,n, Claudio M. Rocco c
a School of Industrial and Systems Engineering, University of Oklahoma, OK, USA b Engineering Management, School of Systems and Enterprises, Stevens Institute of Technology, NJ, USA c Facultad de Ingenieria, Universidad Central de Venezuela, Venezuela
a r t i c l e i n f o
Article history: Received 12 September 2012 Received in revised form 7 January 2013 Accepted 21 March 2013 Available online 8 April 2013
Keywords: Network resilience Component importance measure Vulnerability Recoverability
20/$ - see front matter & 2013 Elsevier Ltd. A x.doi.org/10.1016/j.ress.2013.03.012
esponding author. Tel.: +1 2012168003. ail addresses: [email protected], [email protected] (J.E. Ramirez-Mar
a b s t r a c t
Disruptive events, whether malevolent attacks, natural disasters, manmade accidents, or common failures, can have significant widespread impacts when they lead to the failure of network components and ultimately the larger network itself. An important consideration in the behavior of a network following disruptive events is its resilience, or the ability of the network to “bounce back” to a desired performance state. Building on the extensive reliability engineering literature on measuring component importance, or the extent to which individual network components contribute to network reliability, this paper provides two resilience-based component importance measures. The two measures quantify the (i) potential adverse impact on system resilience from a disruption affecting link i, and (ii) potential positive impact on system resilience when link i cannot be disrupted, respectively. The resilience-based component importance measures, and an algorithm to perform stochastic ordering of network components due to the uncertain nature of network disruptions, are illustrated with a 20 node, 30 link network example.
& 2013 Elsevier Ltd. All rights reserved.
1. Introduction and motivation
The ubiquitous nature of many infrastructure networks has led to their criticality in performing the functions of everyday life. Their criticality is due to their interconnectedness with other infrastructure networks, as well as industries and workforces that rely upon them. Among these critical infrastructure networks are electric power, telecommunications, and transportation, among others [11,29]. Disruptive events, whether malevolent attacks, natural disasters, manmade accidents, or common failures can have significant widespread impacts when they lead to the failure of network components. For example, the recent blackout that left over 600 million people without access to energy underscores the need for adequate protective measures against disruptive events and also for rapid and appropriate response when in the presence of such events.
While original work in planning for infrastructure network disruptions focused on prevention and protection, recent efforts have focused on “preparedness, timely response, and rapid recov- ery” [11] from such disruptive events. This emphasis on the inevitable occurrence of a disruptive event highlights the need for resilience in infrastructure networks, where resilience is often
ll rights reserved.
quez).
defined as an ability of a system to “bounce back” after a disrup- tion. Domain-specific discussion of resilience range from ecological systems [16,6], to economic systems [30,31,24], to organizational systems [17,18]. Specifically for infrastructure systems, the Infra- structure Security Partnership (2011) noted that a resilient infra- structure sector would “prepare for, prevent, protect against, respond or mitigate any anticipated or unexpected significant threat or event” and “rapidly recover and reconstitute critical assets, operations, and services with minimum damage and disruption.” In general, how- ever, no common definition or quantitative approach has been adopted [13].
The analysis of systems, regardless of domain, often includes determining which system components are most influential on the performance of the system. Some examples include the inter- connectedness of a node in a computer network may be of interest to a hacker whose objective is to maximize the damage of an attack [15], the interconnectedness of certain transportation assets may factor in investment priorities [20], and the influence of certain workforce-driven industry sectors on a regional economy in the event of an epidemic [4]. This is a well-studied topic in reliability, where component importance measures have been introduced to measure the influence of particular components on the overall reliability of the system. For all of the above domains, the study of influential components allows for focus to be given to those components that require more attention, perhaps through planning and investment to ensure a satisfactory performance of the system.
K. Barker et al. / Reliability Engineering and System Safety 117 (2013) 89–9790
The contributions of this paper are twofold: (i) to introduce two network component importance measures to identify the compo- nents that are most influential when considering the resilience of the entire network and given that resilience is stochastic in nature, and (ii) provide a discrimination algorithm to identify component importance.
The remainder of the paper is as follows: Section 2 provides a notional and quantitative discussion of resilience. Section 3 focuses on stochastic descriptors of component vulnerability and recoverability and how they fit into the larger quantification of network resilience, as well as two resilience-based component importance measures for describing how individual components contribute to network resilience. Section 4 illustrates these com- ponent importance measures, with concluding remarks provided in Section 5.
2. Methodological background
This section provides the definition of, as well as aspects of modeling, system (network) resilience originally discussed by Refs. [34,13]. Also, background on component importance mea- sures is provided.
Along the bottom of Fig. 1 is the state transition among three distinct states in which a system can operate
–
Fig fun
S0, Original, or as-planned (or baseline) state.
–Sd, Disrupted state, resulting from an event e
j disrupting the system.
–
Sf, Recovered state, that results from a recovery effort.
These states are related by the disruptive event ej as follows: the system operates in S0 until disruptive event e
j occurs at te, and up to time td when the system reaches its maximum disrupted state Sd. Recovery from the disruption commences at time ts, and state Sf is attained at time tf and is maintained thereafter.
System resilience can be defined as the time dependent ratio of recovery over loss, or Я(t)¼Recovery(t)/Loss(t). Note the notation for resilience, Я [34], as R is historically reserved for quantifying reliability. To quantify ЯðtÞ, system performance is overlaid with the system state transition in Fig. 1. System service function φ(t) describes the behavior or performance of the system at time t (e.g., φ(t) could describe traffic flow for a highway network, throughput for a manufacturing facility). Resilience in the system at time t is exhibited if there is an external disruptive event, ej from a set D of possible disruptive events, that affects the original system state, S0, where performance is measured as φ(t0). Disrupted at time te, a period of degradation of length (td−te) transitions the system to Sd with corresponding performance φ (td). After a period of time of
time
S0 Sd Sf
t0 td tftste
Disruptive Event ej
Resilience Action
Reliability Vulnerability-Survivability Recoverability
. 1. Description of state transitions over time with respect to the system service ction.
length (ts−td), the system restoration commences until the system reaches a stable recovered state, Sf, at time tf with corresponding performance φ(tf). Note that state Sf need not be the same as S0, as the new state may reach an alternative (φ(tf) may be lower, or perhaps higher) equilibrium level (e.g., the state of infrastructure following the 2010 earthquake in Haiti may be improved over pre- disruption levels). Based on this description, system resilience is provided in Eq. (1) [13]. Eq. (1), read resilience of the system at time tr given event ej, describes the point in time, tr, when resilience is quantified and provides the ratio of recovery over loss at such point in time. Also, the values of ts,tr and tf can change depending on the point in time when the recovery starts, the point in time when resilience needs to be quantified and the point time when recovery activities are finalized, respectively.
ЯFðtrjejÞ ¼ φðtrjejÞ−φðtdjejÞ φðt0Þ−φðtdjejÞ
; tr∈ðts; tf Þ ð1Þ
Fig. 1 also highlights four distinct dimensions of resilience: reliability, vulnerability, survivability, and recoverability. In the absence of an external disruptive event, the operation of the system during time period (te−t0) is governed by its reliability. Broadly, system availability can be described as the ratio of system up time to system uptime plus system downtime. Thus, availability describes the proportion of time that a system can be used. According to the description in Fig. 1, resilience describes at any point in time, after recovery actions have started, the proportion of service restored due to the loss associated to event ej. These two metrics share in that an approach to increase their values is to reduce system downtime/system restoration time.
Fig. 1 considers vulnerability as the study of the adverse effect on system performance caused by event ej [9,36,23,38], similar in concept to “robustness” in “resilience triangle” literature in civil infrastructure [5,39]. A mitigation approach to the vulnerability of a network is survivability, or the minimization of the original impact of disruptive events [33]. Finally, Ref. [32] explored an approach to quantify the effectiveness of a contingency system the reliability of a supply chain.
Finally, recoverability refers to “the speed at which an entity or system recovers from a severe shock to achieve a desired state” [30], similar in concept to “rapidity” in “resilience triangle” discussions [5,39].
This paper focuses on two dimensions of resilience over time, vulnerability and recoverability, for the development of a resilience-based component importance measure (CIM). CIMs identify system components that are more critical than others in terms of the reliability of the entire system [22,19]. Several reliability-based CIMs have been proposed [12,1,8,21,37,25], and they are generally calculated as some ratio of a measure of component contribution to system reliability and a measure of system reliability itself. The nature of system resilience in Eq. (1) lends itself to calculating the component contribution to system resilience when component i is isolated, discussed subsequently in Section 3.
3. Resilience-based component importance measure
Primary drivers in network resilience are vulnerability and recoverability. Means to measure these two dimensions are dis- cussed in this section, as well their role in measuring network resilience.
3.1. Defining the network
Let G ¼ ðN ; AÞ represent a network where N is the set of nodes, and A ¼ fij1≤i≤mg is the set of arcs. The state variable of link i at
K. Barker et al. / Reliability Engineering and System Safety 117 (2013) 89–97 91
time t is defined by xi(t). The state variable could evaluate, for example, the full traffic flow capacity across a bridge or the transfer of output from one member of a supply chain to another. The network state vector at time t, x(t)¼(x1(t), x2(t), …, xm(t)), denotes the state of all the arcs at time t. The service function, φ(x(t)), which can be analyzed for any possible realization of x(t), maps the performance of the collection of links into a measure of network performance at time t. Assume that disruptive event ej
leads to a degradation in the performance of the network.
3.2. Describing vulnerability in network components
Noted previously, vulnerability relates to the initial impact experienced by the network after a disruptive event. This corre- sponds to the reduced functionality in network components (e.g., reduced flow in network links after a disruption). Previous treat- ments of network resilience have assumed that the function of the links was binary [34], though a more general presentation is provided here. Let xiðt0Þ represent the as-planned state on the ith link prior to the onset of disruptive event ej. Assume that the effect of ej is a proportional reduction in the as-planned state of the ith link by VjiðejÞ ¼ V
j i.
For the case of flow, the effect of ej on the state variable associated to link i is provided in Eq. (2). Note that a complete reduction in the functionality of the link occurs when Vji ¼ 1. Preparedness activities (e.g., protection, system hardening, and false target) that reduce the initial impact on the system would reduce vulnerability. That is, investments in risk management can alter Vji. For example [27], tie vulnerability analysis and protection strategies into a new framework to guide the protection of critical infrastructures components.
xiðtdjVjiÞ ¼ ð1−V j iÞxiðt0Þ ð2Þ
Further, parameter Vji is likely stochastic due to the uncertainty associated with the nature of event j and the subsequent behavior of component i. Eq. (3) governs the behavior of Vji lying
S
A
B
C
E
D
T
(3,4)
(1,5)
(2,7)
(6,3) (10,9)
(12,6) (11,1)
(9,4)
(8,5)
(7,4)
(5,2)
(4,1)
Fig. 2. Illustrative network example, adapted from Hillier and Lieberman [14].
Fig. 3. Trajectory of resilience over time for (a) deterministic arc recovery activity time
in [a,b]∈[0,1].
PðaoVji ≤bÞ ¼ Z b a
f ðvjiÞdv j i ð3Þ
3.3. Describing recoverability in network components
As recoverability refers to the speed at which the network recovers, recoverability can manifest itself following a network disruption through the time required to recover the functionality of a network component. Recovery time for the ith link given its vulnerability, UjijV
j iðejÞ ¼ U
j iðV
j iÞ, is uncertain and thus, U
j iðV
j iÞ is a
stochastic term. The probability that link i recovers prior to time tr∈(ts,tf) is found in Eq. (4).
Pðts oUjiðV j iÞ≤trÞ ¼
Z tr ts
f ðujijV j iÞdv
j i ð4Þ
For this paper, it is assumed that xiðtrÞ ¼ xiðtdÞ until the recovery time is met, suggesting a step function to repair. This assumption could be relaxed with a known trajectory (e.g., linear, convex, and concave) relationship describing xiðtÞ for t∈½ts; tf �. It has also been assumed that Pðts oUjiðV
j iÞ≤trÞ ¼ Pðts oU
j iðV
k i Þ≤trÞ for every
ðVki Þ40. This assumption describes that recovery time is the same for any positive value of vulnerability.
Given a recoverability description for each component from Eq. (4), a metric describing the time until the system bounces back to its original state is Tφðxðt0ÞÞðejÞ, known as the time to full network service resilience.
This metric records the total time spent from the point when recovery activities start, at time ts, up to the time, tf n, when system service is completely restored to the initial service function value φ(x(t0)). That is, the random value tf n is given by such value ensuring that ЯFðtf njejÞ ¼ 1 and ЯFðtf−δjejÞo1 ∀δ40) given dis- ruptive event ej.
Since Tφðxðt0ÞÞðejÞ is stochastic, one can define the probability that network service resilience is reached before mission time tf as Pφðxðt0ÞÞðtfÞ ¼ PðTφðxðt0ÞÞðejÞ≤tfÞ.
3.3.1. Recovery time example To illustrate the Tφðxðt0ÞÞðejÞ metric, consider the network shown
in Fig. 2 (extended from Hillier and Lieberman [14] and Ref. [13]). The arc labels represent the arc index number and arc capacity, respectively. The service function for this network is the maximum network capacity between nodes S and T.
A single disruptive event, e1, is considered. This event renders arcs 1–5 completely inoperable (V1i ¼1, i≤5) with the remaining arcs unaffected (V1i ¼0, i45). Assume that recovery occurs in order of link label: link 1 is repaired first, then link 2, and so on until link 5. In the deterministic recovery time resilience analysis of Ref. [13], recovery activities had a fixed time duration of 10 time units. Fig. 3(a) illustrates the trajectory of resilience over time for
of 10 time units and (b) stochastic arc recovery activity time following UNI(8,12).
K. Barker et al. / Reliability Engineering and System Safety 117 (2013) 89–9792
this network, disruptive event, and recovery activity set when a deterministic time duration is assumed.
To illustrate Tφðxðt0ÞÞðe1Þ, consider a uniform distribution for arc recovery time, U1i ∼UNI(8,12), i¼1,…, 5. The trajectory of resilience under the stochastic recovery time assumption is provided in Fig. 3(b), where interval representations replace point estimates following a 1000-iteration discrete event simulation. Note that full network resilience is achieved after the first three recovery activities (the recovery of arcs 1, 2, and 3, leading to the full capability of the network), while full network restoration does not occur until all five disrupted arcs are recovered after all five recovery activities (as restoration implies that all arcs are fully functional). Fig. 3(b) highlights the difference in the average time to full network resilience and full network restoration.
Approximate probability distribution and cumulative distribu- tion function representations of Tφðxðt0ÞÞðe1Þ are provided in Fig. 4 for particular recovery order 1–2–3–4–5. Note that Tφðxðt0ÞÞðe1Þ is approximately bounded by the interval (24, 36).
3.4. Component importance from resilience
Mentioned in Section 1, the primary contribution of this paper is the introduction of two resilience-based component importance measures. The majority of CIMs have been developed to quantify importance of components to system reliability. In the reliability case [22,19], CIMs illustrate the effect on system reliability as a function of changes in the component state. Other work has extended traditional reliability-based component importance measures to various metrics of availability [7,3]. When considering vulnerability, CIMs illustrate the adverse effect on the system service function as a function of disruptions to the original state of the components.
Sections 3.2 and 3.3 described how individual components are initially impacted (via component vulnerability) and then recover (via component recoverability) following disruptive event ej. And Eq. (1) incorporates these dimensions for all components to quantify network resilience. When considering component impor- tance in a resilience setting, the interest is in understanding the effect that both disruption magnitude and recovery speed at the component level have on the time to full network service resilience, Tφðxðt0ÞÞðejÞ. Eq. (5) mathematically illustrates the first resilience-based CIM.
CIЯφ;iðtrjejÞ ¼ φðxðt0ÞÞ−φððxðt0Þ; xiðtdjVjiÞÞÞ
maxkfφðxðt0ÞÞ−φððxðt0Þ; xkðtdjVjkÞÞÞg T φðxðt0ÞjVjiÞ
ð5Þ
The numerator in the ratio of Eq. (5) describes the system service loss due to the disruption effect on link i, while the denominator describes the maximum loss among all the links. This ratio is the multiplied by the time required to restore the system service to its original state. As Vji and U
j i are stochastic
terms, CIЯφ;iðtrjejÞ has a distribution for tr∈½ts; tf�. CIЯφ;iðtrjejÞ is a
Fig. 4. Time to full network resilience results for 1000 simulations for stochastic a (pdf approximation), and (b) cdf approximation form.
measure of the weighted contribution of link i to the time until full network service is restored, with the weighting factor interpreted as the proportion of the maximum possible change in performance described by link i. This CIM can be regarded as comparable to risk reduction worth (RRW) [26], an index that quantifies the potential damage to a system caused by a particular component.
In reliability engineering, there is also interest in measuring the reliability achievement worth (RAW) of a component, or the maximum proportion increase in system reliability generated by that component. The second resilience-based CIM addresses this perspective as the “resilience worth” of link i, WЯφ;iðtrjejÞ, or an index that quantifies how the time total network service resilience is improved for scenario ej if link i is invulnerable. The mathema- tical representation of WЯφ;iðtrjejÞ is provided in Eq. (6).
WЯφ;iðtrjejÞ ¼ T φðxðt0ÞjVjiÞ
−T φðxðt0ÞjVji ¼ 0Þ
T φðxðt0ÞjVjiÞ
ð6Þ
3.5. Component ordering according to importance
Ordering links in terms of magnitude with respect to the resilience-based CIMs provides an account of greatest to least link impact at tr on (i) potential adverse impact on system resilience from a disruption to link i (with CIЯφ;iðtrjejÞ), and (ii) potential positive impact on system resilience when link i cannot be disrupted (with WЯφ;iðtrjejÞ). An algorithm for generating an order of component importance is described for CIЯφ;iðtrjejÞ below, though the same algorithm would apply for the WЯφ;iðtrjejÞ measure:
1.
rc re
For each link i, and based on probability distributions defined in Eqs. (3) and (4), generate a realization of VjiðejÞ and U
j iðV
j iðejÞÞ,
and calculate CIЯφ;iðtrjejÞ for tr∈½ts; tf �.
2.Repeat Step 1 for a chosen number, η, of iterations, producing a
distribution of CIЯφ;iðtrjejÞ at each time period tr∈½ts; tf �.
3.Given the distributions of CIЯφ;iðtrjejÞ for each link i, perform a
stochastic ranking of links according to ascending CIЯφ;iðtrjejÞ.
Note that when event ej affects more than one link, recoverability strategies (i.e., recovery orders) must be defined a priori [28].
3.5.1. Stochastic ranking To generate the stochastic order of links at Step 3, an approach
based on the Copeland Score method has been developed. The Copeland Score (CS) is a simple non-parametric ranking technique that does not require any information about decision maker preference and operates on a multi-indicator matrix, formed by m objects characterized by Ω attributes. The CS implemented here corresponds to a modification proposed by Al-Sharrah [2]. The CS is computed based on pair-wise comparisons between objects
covery activity time following U(8,12) provided in (a) frequency histogram
K. Barker et al. / Reliability Engineering and System Safety 117 (2013) 89–97 93
(each link) in a set (network) and is defined as the difference between the number of times object a is better (with respect to attribute qk) than the other objects and the number of times that object b is worse (with respect to the same attribute qk) than the other objects. The method then selects the object with the largest Copeland Score. Since comparisons are made for each attribute, no normalization is required. CS also assumes that each attribute has equal importance. If the decision maker has specific preferences as to how attributes should be weighted, then other parametric ranking techniques could be used (e.g., ordered weighted aver- aging [35])
The CS method is applied by first examining the cdf of CIЯφ;iðtrjejÞ for i¼1,…, m and then comparing the cdf of links a and b against a specific number of percentiles, where attribute qk refers to the kth percentile. The stochastic ordering problem is reduced to finding the rank of a set of objects (each link) with specific attributes (selected percentiles). The link ranked highest contributes most to the overall resilience of the system.
Ck(a,b) is a value based on a comparison between link a and link b for percentile qk, k¼1,…, Ω, performed according to the rule in Eq. (7). Before the first percentile, q1, C0(a,b) is initialized at zero, and Eq. (7) is iterated through all Ω percentiles.
Ckða; bÞ ¼ Ck−1ða; bÞ þ 1 qkðaÞoqkðbÞ Ck−1ða; bÞ−1 qkðaÞ4qkðbÞ Ck−1ða; bÞ qkðaÞ ¼ qkðbÞ
8>< >: ð7Þ
The method by Al-Sharrah [2] dictates that the CS of link a is obtained by summing Ck(a,b) over all b≠a, each representing the other objects, as shown in Eq. (8). The link with the largest CS value is assumed to stochastically dominate all other links with respect to the set of percentiles.
CSðaÞ ¼ ∑ b≠a
CΩða; bÞ ð8Þ
At the conclusion of Step 3 in the ranking algorithm, a stochastic order of component importance is produced for each tr∈½ts; tf�. This allows us to determine, for example, which links contribute to resilience early after the disruptive event, which contribute later in the recovery, and so on. Such would provide decision makers with an idea of where and when to place resilience building resources.
(1,11)
S
A
E
D
C
B
F
J
I
H
G
K
O
N
M
L
(2,8)
(3,11)
(4,13)
(5,6)
(6,8) (7,4) (8,9)
(9,13) (10,7) (11,6)
(12,5) (13,12) (14,7)
(16,15) (15,8)
(17,8)
Fig. 5. Illustrative network examp
4. Illustrative example
The resilience-based component importance measure approach is illustrated for a 20-node, 30-link network as depicted in Fig. 5 (adapted from Ref. [10]). Under baseline behavior, the network can handle a maximum flow of 44 units. The disruptive event, e1, causes component vulnerability in link i, as a uniform distribution in [0,1]. Recovery time is simplified to assume that regardless of component vulnerability, time to recover is uniformly distributed in [1,2] arbitrary time units.
Case 1. To illustrate the evaluation process considering vulner- ability and resilience, Fig. 6 illustrates the flow reduction of the network as a function of link vulnerability. To develop this figure, a disruption Vji∼UNI(0,1) is generated for each link i and the total loss is computed as the difference between baseline and disrupted maximum flow. Fig. 6 indicates that links 25 through 28 produce the largest vulnerability, or initial losses in functionality, with link 25, which connects nodes M and R, as the highest contributor. However, note that for component vulnerability≥0.4, the set of important links would include link 23.
Fig. 7 illustrates the cdf of CIЯφ;iðtrjejÞ for each link, using the procedure described in Section 3.5 (for η¼2000 samples). This figure illustrates the probability that CIЯφ;iðtrjejÞ will be less than or equal to a target value x. Considering a target value x¼0.10, as per Fig. 7, the CIЯφ;iðtrjejÞ value associated with link 4 will not have a value above 10 units; in terms of resilience, this link impacts system resilience the least. In contrast, link 25 will be below the same target value 15% of the time. Moreover, the curve for link 25 is always “dominated” by the remaining curves. Therefore, a disruption in link 25 has the most adverse effect on the resilience of the network. Finally, consider links 23 and 27. For a target value xo0.20, the curve for link 27 is below the curve of link 23 (i.e., link 27 is more important than link 23). However this behavior changes for x40.20 and link 23 becomes more important. This behavior underscores that excluding link 25, there is no clear second most important link. For this reason, CS is used to distinguish the importance of links. Fig. 8 shows the CS of all the links, ordered in descending order with link 25 as the most important, followed by links 28, 26, 27, and so on.
Case 2. To illustrate WЯφ;iðtrjejÞ, consider event ej as impacting links 9, 23, 24 and 25 with Vji ¼1 and V
j i ¼0 for every other link.
P
R
Q
T
(18,10) (19,4) (20,7)
(21,10) (22,11) (23,13)
(24,13)
(25,13) (26,4) (27,9)
(30,14)
(29,11)
(28,15)
le 2, adapted from Ref. [10].
Fig. 7. Cumulative probability distribution for the resilience-based component importance measure, CIЯF;iðtrjejÞ.
Fig. 8. Copeland score comparison for each link when comparing CIЯφ;iðtrjejÞ.
Fig. 6. Total network-wide vulnerability as a function of individual component vulnerability.
K. Barker et al. / Reliability Engineering and System Safety 117 (2013) 89–9794
Also, consider that recovery time for each of these links is assumed as uniformly distributed between [10,17]. There are 4!¼24 different recoverability combinations that can be imple- mented if the recovery of the four links is done in series (i.e., combinations {S1,…,S24}). The nine histograms in Fig. 9 illustrate the time to full network service resilience (TTSR) corresponding to a sample of nine out of the total 24 sequences (S1, S2, S3, S11, S12, S13, S21, S22, and S23). As evidenced by Fig. 9, the time to full service resilience associated to the event considered (i.e., failure of links 9, 23, 24 and 25), is at most 40 units roughly 50% of the time independent of the restoration sequence implemen- ted. Note that Fig. 9 depicts the denominator in Eq. (6). Fig. 10
illustrates the effect of T φðxðt0ÞjVji ¼ 0Þ
on links 9, 23, 24, and 25 one at a time and for the first sequence respectively; clearly time to total network service resilience has been reduced (note the maximum on the time axis in Fig. 10 relative to Fig. 9). For these cases, the time to full network service resilience is at most 30 units for roughly 50% of the time independent of the sequence implemented.
To identify the most important link, Fig. 11 illustrates the case when the worst-case (least resilient) restoration sequence is implemented. After applying the Copeland score approach, link 25 exhibits the highest stochastic order and has the most adverse effect on the resilience of the network in the worst case.
Fig. 9. Histograms of the time to total network service resilience for an event impacting links 9, 23, 24, and 25.
Fig. 10. Histograms for time to total network service resilience as a function of link invulnerability, for links 9, 23, 24, and 25.
K. Barker et al. / Reliability Engineering and System Safety 117 (2013) 89–97 95
Fig. 11. Cumulative probability distributions for TTSR for the first sequence.
K. Barker et al. / Reliability Engineering and System Safety 117 (2013) 89–9796
5. Concluding remarks
The purpose of this paper has been to highlight that considera- tions related to network resilience, as opposed to only network protection and disruption prevention, should become more pre- valent in risk-based analysis and planning efforts. The ability of a network to “bounce back” from seemingly inevitable disruptive events is a vital consideration. This work defines resilience as a function of four interacting paradigms: reliability, vulnerability, survivability, and recoverability. Modeling emphasis is devoted to vulnerability, or the initial impact experienced in a network following a disruptive event, and recoverability, or the ability of a network to recover functionality in a timely manner. As such the paper contributes two approaches to measure the importance of network components from the perspective of component contri- bution to network resilience as a function of stochastic vulner- ability and recoverability terms.
The first resilience-based component importance measure, CIЯφ;iðtrjejÞ in Eq. (5), quantifies the potential adverse impact on system resilience at time tr when disruption e
j affects link i. Analogous to the risk reduction worth CIM common in the reliability engineering literature, it measures the proportional contribution of link i to the time required to achieve full network service resilience. Due to the stochastic nature of the elements comprising CIЯφ;iðtrjejÞ, ordering the components according to this measure requires a stochastic ranking technique (e.g., the Copeland Score method). The comparison of the cdfs for the illustrative example in Fig. 7 demonstrate some non-obvious conclusions about the contributions of certain links to the resi- lience of the network in Fig. 5. The second resilience-based component importance measure, WЯφ;iðtrjejÞ in Eq. (6), quantifies the potential positive impact on network resilience when
vulnerability-strengthening measures are put into place such that link i cannot be disrupted.
As with all importance measures, the two proposed in this paper can serve as guides to prioritize resilience improvement activities. As illustrated by the results such activities can be in the form of vulnerability reduction policies (i.e., protecting or hard- ening components) or in the form of accelerating the speed of recovery activities. Future research should be focused on the optimal allocation of resources among these different activities, including the cost assessment of losses due to performance deterioration.
References
[1] Aggarwal A, Barlow R. A survey on network reliability and domination theory. Operations Research 1984;32(2):478–92.
[2] Al-Sharrah G. Ranking using the Copeland score: a comparison with the Hasse diagram. Journal of Chemical Information Models 2010;50(5):785–91.
[3] Barabady J, Kumar U. Availability allocation through importance measures. Inter- national Journal of Quality and Reliability Management 2007;24(6):643–57.
[4] Barker K, Santos JR. A risk-based approach for identifying key economic and infrastructure sectors. Risk Analysis 2010;30(6):962–74.
[5] Bruneau M, Chang SE, Eguchi RT, Lee GC, O'Rourke TD, Reinhorn AM, et al. A framework to quantitatively assess and enhance the seismic resilience of communities. Earthquake Spectra 2003;19(4):733–52.
[6] Carpenter S, Walker B, Anderies JM, Abel N. From metaphor to measurement: resilience of what to what? Ecosystems 2001;4(8):765–81.
[7] Cassady RC, Pohl EA, Song J. Managing availability improvement efforts with importance measures and optimization. Journal of Management Mathematics 2004;15(2):161–74.
[8] Cheok M, Parry G, Sherry R. Use of importance measures in risk informed applications. Reliability Engineering and System Safety 1998;60(3):213–26.
[9] Crucitti P, Latora V, Marchiori M. Locating critical lines in high-voltage electric power grids. Fluctuation and Noise Letters 2005;5(2):L201–8.
[10] Dai Y, Poh K-L. Solving the network interdiction problem with genetic algorithms. In: Proceedings of the Fourth Asia-Pacific Conference on Industrial Engineering and Management Systems. Taipei; 2002.
K. Barker et al. / Reliability Engineering and System Safety 117 (2013) 89–97 97
[11] Department of Homeland Security. National Infrastructure Protection Plan Washington, DC: Office of Secretary of Homeland Security; 2009.
[12] Fussell J. How to calculate system reliability and safety characteristics. IEEE Transactions on Reliability 1975;24(3):169–74.
[13] Henry D, Ramirez-Marquez JE. Generic metrics and quantitative approaches for system resilience as a function of time. Reliability Engineering and System Safety 2012;99(1):114–22.
[14] Hillier FS, Lieberman GJ. Introduction to operations research. 9th ed New- Jersey: McGraw-Hill; 2009.
[15] Holme P, Kim BJ, Yoon CN, Han SK. Attack vulnerability of complex networks. Physical Review E 2002;65:056101–14.
[16] Holling CS. Resilience and stability of ecological systems. Annual Review of Ecology and Systematics 1973;4(1):1–23.
[17] Jackson S. System resilience: capabilities, culture, and infrastructure. In: Proceedings of the INCOSE Annual International Symposium; 2007.
[18] Jackson S. Architecting resilient systems: accident avoidance and survival and recovery from disruptions. Hoboken, NJ: Wiley; 2010.
[19] Kuo W, Zuo MJ. Optimal reliability modeling: principles and applications. Hoboken, NJ: Wiley; 2003.
[20] Leung MF, Lambert JH, Mosenthal A. A risk -based approach to setting priorities in protecting bridges against terrorist attacks. Risk Analysis 2004;24(4):963–84.
[21] Levitin G, Podofillini L, Zio E. Generalized importance measures for multi-state elements based on performance level restrictions. Reliability Engineering and System Safety 2003;82(3):287–98.
[22] Modarres M, Kaminskiy M, Krivtsov V. Reliability engineering and risk analysis: a practical guide. 2nd ed. Boca Raton, FL: CRC Press; 2010.
[23] Nagurney A, Qiang Q. An efficiency measure for dynamic networks with application to the internet and vulnerability analysis. Netnomics 2008;9(1):1–20.
[24] Pant R, Barker K. Building dynamic resilience estimation metrics for inter- dependent infrastructures. In: Proceedings of the European Safety and Reliability Conference. Helsinki, Finland; 2012.
[25] Ramirez-Marquez JE, Coit DW. Composite Importance Measures for Multi- State Systems with Multi-State Components. IEEE Transactions on Reliability 2005;54(3):517–29.
[26] Ramirez-Marquez JE, Rocco CM, Gebre BA, Coit DW, Tortorella M. New insights on multi-state component criticality and importance. Reliability Engineering and System Safety 2006;91(8):894–904.
[27] Ramirez-Marquez JE, Rocco CM. Vulnerability based robust protection strategy selection in service networks. Computers and Industrial Engineer- ing 2012;63(1):235–42.
[28] Ramirez-Marquez, JE and Rocco CM. Towards a unified framework for network resilience. In: Proceedings of the Third International Engineering Systems Symposium CESUN 2012. Delft, Netherlands; 2012b.
[29] Reed DA, Kapur KC, Christie RD. Methodology for assessing the resilience of networked infrastructure. IEEE Systems Journal 2009;3(2):174–80.
[30] Rose A. Economic resilience to natural and man-made disasters: multi- disciplinary origins and contextual dimensions. Environmental Hazards 2007;7(4):383–98.
[31] Rose A. Economic resilience to disasters: Community and Regional Resilience Institute (CARRI). Oakridge, TN: CARRI Institute; 2009 [Research Report 8].
[32] Thomas M. Supply chain reliability for contingency operations. In: Proceedings of the 2002 Annual Reliability and Maintainability Symposium. Seattle, WA; 2002.
[33] West mark V. A definition for information system survivability. In: Proceeding of the 37th Hawaii International Conference on System Sciences. Honolulu, Hawaii; 2004.
[34] Whitson J, Ramirez-Marquez JE. Resiliency as a component importance measure in network reliability. Reliability Engineering and System Safety 2009;94(10):1685–93.
[35] Yager RR. On ordered weighted averaging aggregation operators in multi- criteria decision making. IEEE Transactions on Systems, Man, and Cybernetics 1988;18(1):183–90.
[36] Zhang C, Ramirez-Marquez JE, Rocco C. A new holistic method for reliability performance assessment and critical components detection in complex net- works. IIE Transactions 2011;43(9):661–75.
[37] Zio E, Podofillini L. Monte-Carlo simulation analysis of the effects on different system performance levels on the importance on multi-state components. Reliability Engineering and System Safety 2003;82(1):63–73.
[38] Zio E, Sansavini G, Maja R, Marchionni G. An analytical approach to the safety of road networks. International Journal of Reliability, Quality and Safety Engineering 2008;15(1):67–76.
[39] Zobel CW. Representing perceived tradeoffs in defining disaster resilience. Decision Support Systems 2011;50(2):394–403.
- Resilience-based network component importance measures
- Introduction and motivation
- Methodological background
- Resilience-based component importance measure
- Defining the network
- Describing vulnerability in network components
- Describing recoverability in network components
- Recovery time example
- Component importance from resilience
- Component ordering according to importance
- Stochastic ranking
- Illustrative example
- Concluding remarks
- References