Review on Energy Resilience

profileharsh55
Availability-based-engineering-resilience-metric-a_2018_Reliability-Engineer.pdf

Reliability Engineering and System Safety 172 (2018) 216–224

Contents lists available at ScienceDirect

Reliability Engineering and System Safety

journal homepage: www.elsevier.com/locate/ress

Availability-based engineering resilience metric and its corresponding evaluation methodology

Baoping Cai a , b , ∗ , Min Xie b , Yonghong Liu a , Yiliu Liu c , Qiang Feng d

a College of Mechanical and Electronic Engineering, China University of Petroleum, Qingdao, Shandong 266580, China b Department of Systems Engineering and Engineering Management, City University of Hong Kong, Kowloon, Hong Kong c Department of Mechanical and Industrial Engineering, Norwegian University of Science and Technology, N-7034 Trondheim, Norway d School of Reliability and Systems Engineering, Beihang University, Beijing 100191, China

a r t i c l e i n f o

Keywords:

Resilience

Availability

Metric

Engineering system

a b s t r a c t

Several resilience metrics have been proposed for engineering systems (e.g., mechanical engineering, civil engi-

neering, critical infrastructure, etc.); however, they are different from one another. Their difference is determined

by the performances of the objects of evaluation. This study proposes a new availability-based engineering re-

silience metric from the perspective of reliability engineering. Resilience is considered an intrinsic ability and an

inherent attribute of an engineering system. Engineering system structure and maintenance resources are prin-

cipal factors that affect resilience, which are integrated into the engineering resilience metric. A corresponding

dynamic-Bayesian-network-based evaluation methodology is developed on the basis of the proposed resilience

metric. The resilience value of an engineering system can be predicted using the proposed methodology, which

provides an implementation guidance for engineering planning, design, operation, construction, and manage-

ment. Some examples for common systems (i.e., series, parallel, and voting systems) and an actual application

example (i.e., a nine-bus power grid system) are used to demonstrate the application of the proposed resilience

metric and its corresponding evaluation methodology.

© 2017 Elsevier Ltd. All rights reserved.

1

n f o e s s t e c a i a e s r t H

s r o d

t [ v b s s i p o d fi d i a e

h

R

A

0

. Introduction

Resilience is the capability of an entity to recover from an exter- al disruptive event. To date, the concept of resilience has been spread rom ecology [1,2] to various fields, such as economics [3,4] , psychol- gy [5,6] , and sociology [7,8] . In comparison with the research in non- ngineering contexts, only a small proportion of resilience-related re- earch exists in the field of engineering [9–11] . For engineering systems, uch as mechanical engineering, civil engineering, critical infrastruc- ure, etc., different definitions are proposed depending on the objects of valuation. The National Infrastructure Advisory Council defines criti- al infrastructure resilience as the capability to reduce the magnitude nd/or duration of disruptive events. The effectiveness of a resilient nfrastructure or enterprise depends upon its capability to anticipate, bsorb, adapt to, and/or rapidly recover from a potentially disruptive vent [12] . The American Society of Mechanical Engineers defines re- ilience as the capability of a system to sustain external and internal dis- uptions without discontinuity of performing the system function or, if he function is disconnected, to fully recover the functions rapidly [13] . aimes [45] defined resilience as the capability of the system to with-

∗ Corresponding author.

E-mail address: [email protected] (B. Cai).

ttps://doi.org/10.1016/j.ress.2017.12.021

eceived 23 June 2017; Received in revised form 7 December 2017; Accepted 28 December 20

vailable online 28 December 2017

951-8320/© 2017 Elsevier Ltd. All rights reserved.

tand a major disruption within acceptable degradation parameters, and ecover within an acceptable time and composite costs and risks. Many ther researchers have also proposed their own engineering resilience efinitions from different perspectives [14–17] .

According to the definitions above, various resilience metrics and heir corresponding evaluation methodologies have been developed 18–23] . For example, Dessavre et al. [18] defined a new model and isual tools that improve the capabilities to characterize the resilience ehavior of complex systems by extending existing time-dependent re- ilience functions. Bruneau et al. [19] defined four dimensions of re- ilience, namely, robustness, rapidity, resourcefulness, and redundancy, n the well-known resilience triangle model in civil infrastructure and roposed a deterministic static metric for measuring the resilience loss f a community to an earthquake. Henry et al. [20] proposed a time- ependent quantifiable resilience metric corresponding to a specific gure-of-metric, which was evaluated at a certain time period under isruptive events. Francis et al. [21] proposed a resilience metric that ncorporates three resilience capabilities, including adaptive capacity, bsorptive capacity, recoverability, and the time to recovery. Hosseini t al. [22] used static Bayesian networks to model infrastructure re-

17

B. Cai et al. Reliability Engineering and System Safety 172 (2018) 216–224

Fig. 1. Resilience-related properties of engineering system.

s p

i b t e d r a s a b i r s t i i n r n d v i d

i n s p i c t i b m d e b m

Time

Availability

Shock

0 t1 t2 t3

A1 A3

A2

100%

Fig. 2. Availability of a system subject to degradation and shock.

2

2

p d t d s p o i A

c p t e t

i p o d u a t s t T i s

f d r e a o u l r A

r

𝜌

ilience and used a case of inland waterway port to demonstrate the roposed method.

Although various resilience metrics have been developed, quantify- ng the resilience for a specific engineering system remains a challenge ecause of internal and external factors involved in such metrics. From he above definitions and metrics, resilience overlaps with a number of xisting concepts, such as adaptability [24] , robustness [24,25] , redun- ancy [26] , flexibility [27] , survivability [27] , recoverability [28,29] , apidity [25,30] , and resourcefulness [30] . Here, we consider resilience s an intrinsic capability and an inherent attribute of an engineering ystem itself. It is composed of two properties, namely, performance- nd time-related properties. System structure determines performance- ased properties, such as robustness, adaptability, redundancy, flexibil- ty, and survivability, whereas maintenance resource determines time- elated properties, such as reparability, recoverability, rapidity, and re- ourcefulness. Similar with the reliability in reliability engineering, ex- ernal factors, such as disturbance, attack, and disaster events, are not ntrinsic properties of resilience in engineering system and are thus not nvolved in the resilience metric (see Fig. 1 ). Therefore, when an engi- eering system is designed and maintenance resource is allocated, the esilience of this system is determined. Hence, the structure and mainte- ance resource in the engineering system form a unified whole, thereby etermining the engineering resilience of the system. The promoted iewpoint may be different from the dominant ones [46,47] ; however, t can provide an implementation guidance for engineering planning, esign, operation, construction, and management.

In this study, we aim to develop a new availability-based engineer- ng resilience metric from the essence and property of resilience in engi- eering system. From the perspective of reliability engineering, steady- tate availability and steady-state time can be used to represent the erformance- and time-related properties. Each engineering system has ts own availability; thus, the resilience value can be obtained easily ac- ording to the steady-state availability and steady-state time. Therefore, he metric is suitable for every engineering system. The rest of this paper s organized as follows. Section 2 presents the proposed availability- ased engineering resilience metric and its corresponding evaluation ethodology. Section 3 adopts some examples for common systems to emonstrate the application of the proposed resilience metric and its valuation methodology. Section 4 adopts an actual example for a nine- us power grid system to demonstrate the application of the proposed ethod. Section 5 summarizes the contributions of this paper.

217

. Resilience metric and evaluation methodology

.1. Availability-based engineering resilience metric

Each engineering system has its own availability, where an item is ca- able to be in a state of performing a required function under given con- itions at a given time or time interval, assuming that the required ex- ernal resources are provided. The availability of an engineering system ecreases continuously to reach a steady-state availability A 1 at steady tate t 1 from the initial time with the initial availability of 100%. This rogress is caused by degradation of components and daily maintenance f the system. Suppose an external shock occurs at time t 2 , the availabil- ty instantaneously decreases to a post-shock transient-state availability 2 and then increases to a new equilibrium state A 3 . This progress is also aused by emergency repair after shock, as well as degradation of com- onents. The blue line in Fig. 2 represents the availability considering he degradation of components and daily system maintenance without xternal shocks, and the red line represents the availability considering he emergency repair after shock and degradation of components.

The steady-state availability A 1 , post-shock transient-state availabil- ty A 2 , post-shock steady-state availability A 3 , steady-state time t 1 , and ost-shock steady-state time ( t 3 − t 2 ) are determined by the structure f the engineering system and maintenance resource, such as redun- ant structure, failure rate, and repair rate. High redundancy, low fail- re rate, or high repair rate results in high steady-state availability nd short steady-state time before and after any shocks. This condi- ion accords with the essence and property of resilience. Therefore, teady-state availability and low steady-state time are used to represent he performance- and time-related properties of engineering resilience. hus, quantifying the resilience with an appropriate resilience metric

s no longer a challenge given that steady-state availability and steady- tate time are easy to obtain.

The proposed resilience metric aims to compare the resilience of dif- erent systems that achieve the same functions, thereby identifying the ifferent internal factors that contribute to it. In this study, we develop a esilience metric that incorporates performance- and time-related prop- rties using steady-state availability and steady-state time before and fter external shocks. The value of resilience increases with the increase f availability A and the decrease of recovery time t . Thus, A /ln( t ) is sed to describe the degree of resilience. The natural logarithm function n(x) is used to balance the level of effects between availability A and ecovery time t . The resilience metric is considered to be the product of /ln( t ) before and after external shocks. Therefore, the final developed esilience metric is given as follows:

= 𝐴 1

𝑛 ln ( 𝑡 1 )

𝑛 ∑ 𝑖 =1

𝐴 𝑖 2 𝐴 𝑖 3

ln ( 𝑡 𝑖 − 𝑡 𝑖 ) , (1)

3 2

B. Cai et al. Reliability Engineering and System Safety 172 (2018) 216–224

w

a T t s t s t n t t s f f

𝑝

w p u t s c

d m c i r

𝜇

w s a t

s e

𝐴

w d t

2

r s t c o t w v a i [ s c s s a d s

s o s

B

(

(

(

(

(

3

3

p a r t t s c r e s m r

o i w t o 2 w s o c p v

3

o r t S B t n w a i r b

B

here n is the number of shocks, and i ∈ [1, n ]. Given that the external factors are random and unpredictable, they

re not factors of resilience and not involved in the resilience metric. he external factors only trigger a “bounce back, ” which is similar to he spring system, where an external force F can extend or compress a pring by some distance X and the spring can bounce back to the ini- ial balance once the force F is removed. According to Hooke’s law, the tiffness of the spring can be expressed as k = F / X . It is a constant fac- or characteristic of spring, which is determined by the spring itself and ot by the force. The defined resilience 𝜌 is similar with k ; however, for he engineering system, the external factors and responses of this sys- em are not directly proportional. Therefore, we determine a series of hocks on the engineering system, which result in the common cause ailure of components. The prior probability of common cause failure or each component is defined as

𝑖 = 𝑖

𝑛 + 1 , (2)

here i ∈ [1, n ]. A shock can lead to a common cause failure of com- onents with any prior probability. The proposed resilience metric is sed to evaluate, optimize, compare, and design systems only if n is he same for each system. A larger value of n indicates more simulated hocks. When the number of shocks is more than 9, the resilience slightly hanges. Therefore, we select n = 9 to evaluate system resilience.

For different shocks, the repair rates of components are completely ifferent with fixed maintenance resources. When a shock is serious, the aintenance resources are dispersed, thereby causing low repair rates of

omponents. That is, a larger prior probability of common cause failure ndicates a smaller repair rate of components. To simplify, we define the epair rate of component under different shocks as

𝑖 = (1 − 𝑝 𝑖 ) 𝜇, (3)

here 𝜇 is the repair rate of each component under normal circum- tances. Notably, other relationships between repair rate and prior prob- bility of common cause failure, 𝜇i = f ( p i ), can also be modeled and used o calculate the availability and subsequent resilience.

Based on the prior probabilities of common cause failure and corre- ponding repair rates of components, the steady-state availability of the ngineering system can be obtained as follows:

(∞) = lim 𝑡 →∞

𝐴 ( 𝑡 ) , (4)

Notably, a real steady-state availability does not exist. In practice, e therefore define steady-state availability as the availability when the ifference within five continuous time point (hour) is equal to or less han 10 − 5 , and the time is termed as steady-state time.

.2. Engineering resilience evaluation methodology

Evaluating the availability of engineering systems is important in esilience evaluation. Several approaches can be used to evaluate the teady-state availability, such as reliability block diagram [31] , fault ree [32] , Monte Carlo simulation [33] , and Markov chain [34] . In this urrent work, a dynamic-Bayesian-network-based evaluation methodol- gy is proposed to calculate the steady-state availability, steady-state ime, and subsequent resilience of engineering systems. Bayesian net- ork is a probabilistic graphical model that represents a set of random ariables, including their conditional dependencies through directed cyclic graphs. It is considered to be one of the most useful models n the field of probabilistic knowledge representation and reasoning 35,43,44] . Dynamic Bayesian networks are a long-established exten- ion to ordinary Bayesian networks and allow the explicit modeling of hanges over time. In view of classical probabilistic temporal models, uch as Markov chains, dynamic Bayesian networks are stochastic tran- ition models factored over a number of random variables, over which set of conditional dependency assumption is defined [36] . We adopt ynamic Bayesian networks to predict the future state of variables con- idering the current observation of variables. That is, we can predict the

218

teady-state availability and steady-state time of an engineering system n the basis of the current state of components, such as when external hocks destroy some components.

Engineering resilience evaluation methodology with dynamic ayesian networks consists of the following five procedures:

1) Structural modeling of dynamic Bayesian networks is completed by using structural relationship methods, mapping algorithms, or struc- ture learning methods;

2) Expert elicitation with noisy models or parameter-learning methods are used to model the parameters of dynamic Bayesian networks;

3) Availability is evaluated by using exact or approximated inference algorithms;

4) Resilience is evaluated using the resilience metric shown in Eqs. (1) – (4) ; and

5) Sensitivity analysis is conducted to research the influences of failure and repair actions on the resilience of engineering systems.

. Examples for common systems

.1. Series, parallel, and voting systems

Many systems in practical engineering can be abstracted as series, arallel, or voting system. Taking subsea blowout preventer system as n example, the control stations, control pods, annular preventer, and am preventer are redundantly configured; thus, three control stations, wo control pods, several annular preventers, and several ram preven- ers are considered in parallel. The entire system can be considered a eries of control stations, triple modular redundancy controllers, subsea ontrol pods, annular preventers, lower marine riser package connector, am preventer, and wellhead connector because the complete failure of ach component category causes failure of the subsea blowout preventer ystem [34] . In the current work, we adopt some examples for the com- on systems to demonstrate the application of the availability-based

esilience metric and its evaluation methodology. Series, parallel, and voting systems are an abstract system composed

f three series components, three parallel components, and three vot- ng components, denoted by S3P3V3 (See Fig. 3 a). The series subsystem orks only when all of the three series components S1, S2, and S3 work;

he parallel subsystem works when any of the three components P1, P2, r P3 works; the 2-out-of-3 (2oo3) voting system works when at least components of V1, V2, and V3 work. The entire system works only hen all of the three subsystems work, which is equivalent to a series

ystem. Similarly, we use S2P3V3 to denote a series system composed f two series components, three parallel components, and three voting omponents (See Fig. 3 b), and S3P2V3 to denote a series system com- osed of three series components, two parallel components, and three oting components (See Fig. 3 c).

.2. Structural modeling of dynamic Bayesian networks

For these series, parallel, and voting systems, the structural models f dynamic Bayesian networks are established by using the structural elationship of each component. Taking S3P3V3 system as an example, he components and their states are denoted by root nodes, including 1, S2, S3, P1, P2, P3, V1, V2, and V3 at a specific time slice of dynamic ayesian networks (e.g., Slice1: t ; see Fig. 4 ). A is the final leaf node of he network, which represents the state of the entire system. Each root ode has two states (i.e., work and fail ). For node A, the probabilities of ork and fail indicate the transient availability and transient unavail- bility, respectively. We artificially added several intermediate nodes, ncluding S, P, and V, to simplify the conditional probability table of elated nodes. The causal relationship between the nodes are connected y intra slice arcs.

Dynamic Bayesian networks are essentially replications of static ayesian networks over n time slices between t and t + ( n − 1) Δt . A set

B. Cai et al. Reliability Engineering and System Safety 172 (2018) 216–224

2oo3

S1

S2

S3

P2

V2

P1

V1

P3

V3

(a)

2oo3

S1

S2

P2

V2

P1

V1

P3

V3

(b)

2oo3

S1

S2

S3

V2

P1

V1

P2

V3

(c)

Fig. 3. Simple systems composed of series, parallel, and voting subsystems: (a) S3P3V3, (b) S2P3V3, and (c) S3P2V3.

S1

S2

S3

P1

P2

P3

V1

V2

V3

S

P

V

A

S1

S2

S3

P1

P2

P3

V1

V2

V3

S

P

V

A

Slice1: t Slice2: t+Δt

S1

S2

S3

P1

P2

P3

V1

V2

V3

S

P

V

A

Slicen: t+(n-1)Δt

Fig. 4. Dynamic Bayesian networks of S3P3V3 system.

o s t A i r

3

o m t

Table 1

Failure and repair rates of components in series, parallel, and voting systems.

System Component Failure rate Repair rate

Series, parallel, and voting systems S 0.833e-3 0.500

P 2.083e-3 0.330

V 1.389e-3 0.670

a i t o

t n a f t t s

𝑝

𝑝

𝑝

𝑝

w r t (

3

t i f s a f w

𝑝

f inter arcs between adjacent time slices t and t + Δt connects the corre- ponding nodes of components, which represent the dynamic degrada- ion process and daily maintenance or emergency repair of components. ll the information required to predict a state at time t + Δt is contained

n the description at time t , and no information about earlier times is equired; thus, the process possesses the Markov property.

.3. Parameter modeling of dynamic Bayesian networks

The parameter model of dynamic Bayesian networks is composed f intra and inter slice parameter models. For the intra slice parameter odel, the marginal prior probabilities are assigned to them according

o the resilience metric in Eq. (2) , and the conditional probability tables

219

re determined using the series, parallel, and voting relationship. For the nter slice parameter model, we use Markov state transition relationship o determine the dynamic degradation process and daily maintenance r emergency repair of components.

In dynamic Bayesian networks, the inter slice parameter model is he probability of nodes between time slices t and t + Δt . For the compo- ents of S3V3P3 system, we suppose that the failure and repair follow n exponential distribution, that is, all of the transition rates, including ailure and repair rates, are constant. Given that the process possesses he Markov property, the probability is determined using a Markov-state ransition relationship. Hence, the transition relationships between con- ecutive nodes can be expressed as follows:

( 𝑋 𝑡 +Δ𝑡 = 𝑤𝑜𝑟𝑘 ||𝑋 𝑡 = 𝑤𝑜𝑟𝑘 ) = 𝑒 − 𝜆Δ𝑡 , (5)

( 𝑋 𝑡 +Δ𝑡 = 𝑓𝑎𝑖𝑙 ||𝑋 𝑡 = 𝑤𝑜𝑟𝑘 ) = 1 − 𝑒 − 𝜆Δ𝑡 , (6)

( 𝑋 𝑡 +Δ𝑡 = 𝑓𝑎𝑖𝑙 ||𝑋 𝑡 = 𝑓𝑎𝑖𝑙 ) = 𝑒 − 𝜇Δ𝑡 , (7)

( 𝑋 𝑡 +Δ𝑡 = 𝑤𝑜𝑟𝑘 ||𝑋 𝑡 = 𝑓𝑎𝑖𝑙 ) = 1 − 𝑒 − 𝜇Δ𝑡 , (8) here X is the root node, 𝜆 is the failure rate of a component, and 𝜇 is the

epair rate of a component. For the S3P3V3, S2P3V3, and S3P2V3 sys- ems, we provide the same failure and repair rates for each component see Table 1 ).

.4. Resilience evaluation

The goal of inference in a dynamic Bayesian network is to compute he marginal p ( X t + h | y 1: t ) when y 1: t is observation. h = 0, h < 0, and h > 0, ndicate filtering, smoothing, and prediction, respectively. We use the ollowing prediction of dynamic Bayesian networks to evaluate the re- ilience value of engineering systems. Junction tree algorithm for prop- gation analysis is conducted, where the joint probability for the model rom the conditional probability structure of the dynamic Bayesian net- orks is calculated in a computationally efficient manner.

( 𝑌 𝑡 + ℎ = ℎ |𝑦 1∶ 𝑡 ) =

∑ 𝑥 𝑝 ( 𝑌

𝑡 + ℎ = ℎ |||𝑋 𝑡 + ℎ = 𝑥 ) 𝑝 ( 𝑋 𝑡 + ℎ = 𝑥 |𝑦 1∶ 𝑡 ) (9)

B. Cai et al. Reliability Engineering and System Safety 172 (2018) 216–224

0.0

0.2

0.4

0.6

0.8

1.0

0 50 100 150 200 250

A va

ila bi

lit y

Time (h)

AY0 AY1 AY2 AY3 AY4 AY5 AY6 AY7 AY8 AY9

0.94

0.96

0.98

1.00

0 10 20 30 40 50

Fig. 5. Availability of S3P3V3 system subject to degradation and different shocks.

e t t i s a 9 t c s A u d r t a 2 s a fi t p

b c r

3

V s w p d p w u s f i

1.8951

1.8954

1.8957

1.8960

1.8963

1.8966

1.7500

1.8000

1.8500

1.9000

1.9500

2.0000

2.0500

0.4 1.4 2.4

R es

ili en

ce , P

a nd

V (%

)

R es

ili en

ce , S

( %

)

Times

S

P

V

Fig. 6. Sensitivity of failure rates of components of S3P3V3 system.

0.40

0.80

1.20

1.60

2.00

0.00

1.00

2.00

3.00

4.00

5.00

0 0.5 1 1.5 2 2.5 3

R es

ili en

ce , P

a nd

V (%

)

R es

ili en

ce , S

( %

)

Times

S P V

Fig. 7. Sensitivity of repair rates of components of S3P3V3 system.

s t

V s w p i i o i r r t i v t S

t t o r g

The trends of availability without external shocks or with differ- nt external shocks are different (See Fig. 5 ). The curve AY0 indicates he availability without any shocks, and the curves AY1–AY9 indicate he availability with different shocks, that is, different prior probabil- ties (p1–p9) of common cause failure for each component. When no hocks occur, that is, during normal running of the S3P3V3 system, the vailability decreases rapidly from 100% and reaches a stable level of 9.365% at the 15th hour. The steady-state availability and steady-state ime are therefore 99.365% and 15, respectively. During emergency cir- umstances, the availabilities decrease to a minimum valve the moment hocks occur and then increase to reach different stable levels. Taking Y1 as an example, the prior probability of 10% of common cause fail- re are assigned to each component at the original time. The availability ecreases to 70.790% immediately, increases rapidly with emergency epair, and reaches a stable level of 99.311% at the 27th hour. Hence, he post-shock transient-state availability, post-shock steady-state avail- bility, and post-shock steady-state time are 70.790%, 99.311%, and 7, respectively. With the increase in probability of shocks, the post- hock transient-state availability and the post-shock steady-state avail- bility decrease, and the post-shock steady-state time increases. This nding agrees with the fact. Using all the characteristic values, we ob- ain the resilience value of the S3P3V3 system of 1.90% using the pro- osed availability-based resilience metric.

Resilience is determined by the engineering system itself and not y external shocks; thus, the factors that affect system performance are ertain to affect the resilience value. System structures and failure and epair rates of components are main influencing factors.

.5. Sensitivity analysis

The sensitivity analysis of the failure rates of components S, P, and are conducted by changing the failure rates of each component of the ame category in multiples. The curves represent the resilience values ith the changes of the failure rates from 0.5 times to 2.5 times of series, arallel, and voting components. The resilience of the S3P3V3 system ecreases with the increase in time of failure rates (See Fig. 6 ). For com- onents S and P, the resilience values present a ladder-form decrease ith the increase of failure rates. For component V, the resilience val- es continuously decrease with the increase of failure rates. Under the ame time variation of failure rates of components, the resilience value or S decreases fastest, that for V decreases slowest, and that for P is n between. Therefore, the resilience of the S3P3V3 system is the most

220

ensitive to the failure rates of component S and is the least sensitive to he failure rates of component V.

The sensitivity analysis of the repair rates of components S, P, and are conducted by changing the repair rates of each component of the ame category in multiples. The curves represent the resilience values ith the changes of the repair rates from 0.1 times to 3.0 times of series, arallel, and voting components. The resilience of the S3P3V3 system ncreases with the increase in time of repair rates (See Fig. 7 ). With the ncrease of repair rates, the resilience value for component S continu- usly decreases when time is minimal, whereas it presents a ladder-form ncrease when the time is considerable. For components P and V, the esilience values increase rapidly when the times is minimal and then each stable levels with the increase of repair rates. Under the same ime variation of repair rates of components, the resilience value for S ncreases fastest, and those for P and V increase slowest; the resilience alue of V is slightly higher than that for P. Therefore, the resilience of he S3P3V3 system is the most sensitive to the repair rates of component and is the least sensitive to the failure rates of components P and V.

We suppose that S3P3V3, S2P3V3, and S3P2V3 systems can achieve he same functions. Comparison of S2P3V3 and S3P3V3 systems reveal hat less series components improve engineering resilience; comparison f S3P2V3 and S3P3V3 systems show that less parallel components can educe engineering resilience (See Fig. 8 ). Therefore, redundancy of en- ineering systems plays an important role in engineering resilience.

B. Cai et al. Reliability Engineering and System Safety 172 (2018) 216–224

0.00

0.50

1.00

1.50

2.00

2.50

3.00

S3P3V3 S2P3V3 S3P2V3

R es

ili en

ce (

% )

Fig. 8. Resilience values of S3P3V3, S2P3V3, and S3P2V3 systems.

Table 2

Failure and repair rates of components in the nine-bus power grid system.

System Component Failure rate Repair rate

Nine-bus power grid system Generator 1.631e-5 0.050

Transformer 1.903e-5 0.067

Line 2.854e-5 0.083

Bus 0.951e-5 0.250

4

4

i p r i p g l s C

s e t L s w

t r t A b t o c t a g p

4

4

t f i b r W t o p h o a t a s a i s t t f t o n

4

t t F t s s t t u c i p g

4

t t s u s r t ( r

4

b w v e t m b

. Actual application example for nine-bus power grid system

.1. Resilience of nine-bus power grid system

Aside from abstracting the series, parallel, and voting systems, us- ng a practical engineering system is necessary to demonstrate the ap- lication of the proposed availability-based resilience metric. Various esilience indexes or evaluation methods have been demonstrated us- ng power systems [37,38] . Here, we study the resilience of a nine-bus ower grid system (See Fig. 9 ). The system consists of nine buses, three enerators, three two-winding power transformers, six lines, and three oads. To simplify, we ignore the power energy volume and dynamic re- ponse and focus on the system structure. Hence, when loads A, B, and are powered, the entire nine-bus power grid system works.

The dynamic Bayesian network model of the nine-bus power grid ystem is established similar to the S3P3V3 system, and the resilience is valuated (See Fig. 10 and Table 2 ). Node B denotes the bus; G denotes he generator; T represents the transformer; L represents the line; and A, LB, and LC are loads A, B, and C, respectively. Node S represents the tate of the entire power grid system. Each root node has two states (i.e., ork and fail ). For node S, the probabilities of work and fail indicate the

ransient availability and transient unavailability of this power system, espectively. For the nine-bus power grid system, we review the litera- ure [39] and determine the failure and repair rates of each components. dditional general distributions, such as Weibull distribution, can also e modeled using dynamic Bayesian networks [40] . The results show hat the resilience value of the system is 0.54%. Although the resilience f the S3P3V3 system is larger than the nine-bus power grid system, omparing these two systems is illogical because they do not achieve he same functions. That is, the resilience of engineering systems that chieve the same functions is comparable. Therefore, for a specified en- ineering system, resilience provides an implementation guidance for lanning, design, operation, construction, and management.

221

.2. Discussion of the proposed resilience metric

.2.1. Resilience is an intrinsic capability and an inherent attribute

We consider resilience as an intrinsic capability and an inherent at- ribute of an engineering system itself. It is not influenced by external actors, such as disturbance, attack, and disaster events. This definition s different from other viewpoints [41,42] . The developed availability- ased resilience metric in this study involves the performance- and time- elated properties of the engineering system but not the external factors. e adopt steady-state availability and steady-state time before and af-

er several shocks to define the metric. Therefore, the resilience value f the system is determined when a system is developed and the re- air resources are assigned and fixed in this system. Resilience value is elpful in planning, design, operation, construction, and management f an engineering system. When an emergency incident occurs (e.g., n earthquake destroys some components of a power grid system), al- hough some emergency maintenance teams outside the system might be ssigned to this system and hence shorten the repair time, the primary ystem does not possess this intrinsic capability. Therefore, the repair ctions of the emergency maintenance teams from other systems cannot mprove the resilience of this system. Another example is two similar ystems in earthquake-prone area and non-earthquake area; these sys- ems should have the same resilience values when the same maintenance eams are involved because the systems have the same capability to suf- er from an earthquake and recover from the destruction. To improve he capability against earthquakes, one should improve the performance f the system or increase the maintenance teams of the system itself, but ot to obtain help from other systems.

.2.2. Comparison of resilience-based engineering systems

The proposed availability-based resilience metric aims to compare he resilience values of different systems that achieve the same func- ions, thereby identifying different internal factors that contribute to it. or systems with the same functions, a larger resilience value indicates hat the system is more resilient. Notably, the objects for comparison hould be systems that achieve the same functions. In this study, the eries, parallel, and voting systems S3P3V3, S2P3V3, and S3P2V3 aim o connect two terminations, and the nine-bus power grid system aims o supply power for loads A, B, and C. Comparing the resilience val- es of S3P3V3, S2P3V3, and S3P2V3 systems seems logical, whereas omparing the S3P3V3 system with the nine-bus power grid system is rrational. Similarly, if we develop a new power grid system to supply ower for loads A, B, and C, then comparing it with the nine-bus power rid system and selecting a more resilient system are important.

.2.3. Optimization of resilience-based engineering system

When an engineering system is planned and designed, identifying he weak components that affect the resilience significantly is impor- ant. Resilience is influenced by the internal factors of an engineering ystem; thus, the sensitivity of these factors (e.g., redundancy and fail- re and repair rates) to resilience can be quantified. We can change the ystem structure and failure or repair rates in multiples to evaluate the esilience values and analyze the sensitivity. We should pay more atten- ion to the structure or component that is most sensitive to resilience e.g., increase the redundancy of that structure to decrease the failure ate or increase the repair rate of that component).

.2.4. Design of resilience-based engineering system

From the perspective of reliability engineering, several reliability- ased engineering system design methods are available. For example, e can design an automation control system of subsea blowout pre- enters with the reliability of 99.9999%. Similarly, a resilience-based ngineering system design method might be useful because it involves he characteristics of recovery after shocks. In practical guidance docu- ents, the recommended resilience value of engineering systems should

e specified. For example, suppose that the resilience value for a power

B. Cai et al. Reliability Engineering and System Safety 172 (2018) 216–224

Fig. 9. Nine-bus power grid system.

G1 B1 T1 B4 L1 L2

M1

G2 B2 T2 B7 L3 L6

M2

G3 B3 T3 B9 L4 L5

M3

B5 B6 B8

M9 M7 M6 M4 M8 M5

LA LB LC

S

G1 B1 T1 B4 L1 L2

M1

G2 B2 T2 B7 L3 L6

M2

G3 B3 T3 B9 L4 L5

M3

B5 B6 B8

M9 M7 M6 M4 M8 M5

LA LB LC

S

Slice1: t

Slice2: t+Δt

Fig. 10. Dynamic Bayesian networks of the nine-bus power grid system.

222

B. Cai et al. Reliability Engineering and System Safety 172 (2018) 216–224

g s i m p v S b

5

f d b n t s c f e r g t t c u t e b

s p a i c f s a f u

A

t t f (

b K

R

[

[

[

[

[

[

[

[

[

[

[

[

[

[

[

[

[

[

[

[

[

[

[

[

[

[

[

[

[

[

[

[

[

rid system in Qingdao City is specified to be 0.60%; a new power grid ystem should be designed to meet this requirement. Notably, the spec- fied resilience values are determined by practical engineering environ- ent even for the same system. For example, the resilience value for a ower grid system in a large city might be 0.80%, but 0.50% in a small illage, 0.90% in an industrial park, and 0.60% in a residential area. uch values are determined by practical guidance documents produced y experts.

. Conclusion

A new availability-based engineering resilience metric is proposed rom the perspective of reliability engineering. The corresponding ynamic-Bayesian-network-based evaluation methodology is developed ased on the proposed metric. Series, parallel, and voting systems and a ine-bus power grid system are used to demonstrate the application of he metric and its corresponding evaluation methodology. The results how that the engineering resilience metric is reasonable and that the orresponding evaluation methodology is precise. System structures and ailure and repair rates of components are the main influencing factors of ngineering resilience. Redundancy of components plays an important ole in increasing resilience. The proposed metric can be used for en- ineering comparison, optimization, and design. Therefore, we can use he metric to compare the resilience of different systems that achieve he same functions, thereby identifying different internal factors that ontribute to it. Moreover, we can change the system structure and fail- re or repair rates in multiples to evaluate resilience values and analyze he sensitivity. Furthermore, we can conduct resilience-based design for ngineering systems based on the required resilience values determined y practical guidance documents.

Notably, only the series, parallel, voting, and nine-bus power grid ystems with binary state are used in this study to demonstrate the pro- osed availability-based resilience metric and its corresponding evalu- tion methodology. Bayesian networks are a powerful tool in model- ng any complex system, such as multistate system, linear consecutively onnected systems, and general networks with sources and sinks. There- ore, the proposed metric and methodology are general for any complex ystems. Future scopes of work can be directed toward resilience evalu- tion of a complex system (not limited to the binary state systems) and urther system comparison, system optimization, and system design by sing the proposed metric and methodology.

cknowledgments

This work was supported by the National Natural Science Founda- ion of China (No. 51779267 ), Fundamental Research Funds for the Cen- ral Universities (No. 17CX05022 and No. 14CX02197A ), and Program or Changjiang Scholars and Innovative Research Team in University IRT_ 14R58 ).

This work utilized the High-Performance Computer Cluster managed y the College of Science and Engineering of City University of Hong ong.

eferences

[1] Cole LE , Bhagwat SA , Willis KJ . Recovery and resilience of tropical forests after disturbance. Nat Commun 2014;5:3906 .

[2] Oliver TH , Isaac NJ , August TA , Woodcock BA , Roy DB , Bullock JM . Declining re- silience of ecosystem functions under biodiversity loss. Nat Commun 2015;6:10122 .

[3] Martin R . Regional economic resilience, hysteresis and recessionary shocks. J Econ Geogr 2011;12:1–32 .

[4] Rose A , Liao SY . Modeling regional economic resilience to disasters: a com- putable general equilibrium analysis of water service disruptions. J Region Sci 2005;45(1):75–112 .

[5] Southwick SM , Charney DS . The science of resilience: implications for the prevention and treatment of depression. Science 2012;338:79–82 .

[6] Pan JY , Chan CLW . Resilience: a new research area in positive psychology. Psycholo- gia 2007;50(3):164–76 .

223

[7] Olsson L , Jerneck A , Thoren H , Persson J , O’Byrne D . Why resilience is unappealing to social science: theoretical and empirical investigations of the scientific use of resilience. Sci Adv 2015;1(4):e1400217 .

[8] Keck M , Sakdapolrak P . What is social resilience? Lessons learned and ways forward. Erdkunde 2013;67(1):5–19 .

[9] Hosseini S , Barker K , Ramirez-Marquez JE . A review of definitions and measures of system resilience. Reliab Eng Syst Safety 2016;145:47–61 .

10] Gao J , Barzel B , Barabási AL . Universal resilience patterns in complex networks. Nature 2016;530:307–12 .

11] Fang YP , Pedroni N , Zio E . Resilience-based component importance measures for critical infrastructure network systems. IEEE Trans Reliab 2016;65(2):502–12 .

12] National Infrastructure Advisory Council (US) Critical infrastructure resilience: final report and recommendations.. Natl Infrastruct Advis Council 2009 .

13] American Society of Mechanical Engineers (US) All-Hazards risk and resilience: pri- oritizing critical infrastructures using the RAMCAP Plus SM approach. Am Soc Mech Eng 2009 .

14] Youn BD , Hu C , Wang P . Resilience-driven system design of complex engineered systems. J Mech Design 2011;133(10):101011 .

15] Ouyang M , Wang Z . Resilience assessment of interdependent infrastructure systems: with a focus on joint restoration modeling and analysis. Reliab Eng Syst Safety 2015;141:74–82 .

16] Yodo N , Wang P . Resilience modeling and quantification for engineered systems using Bayesian networks. J Mech Design 2016;138(3):031404 .

17] Roege PE , Collier ZA , Mancillas J , McDonagh JA , Linkov I . Metrics for energy re- silience. Energy Policy 2014;72:249–56 .

18] Dessavre DG , Ramirez-Marquez JE , Barker K . Multidimensional approach to complex system resilience analysis. Reliab Eng Syst Safety 2016;149:34–43 .

19] Bruneau M , Chang SE , Eguchi RT , Lee GC , O’Rourke TD , Reinhorn AM , von Winter- feldt D . A framework to quantitatively assess and enhance the seismic resilience of communities. Earthquake spectra 2003;19(4):733–52 .

20] Henry D , Ramirez-Marquez JE . Generic metrics and quantitative approaches for sys- tem resilience as a function of time. Reliab Eng Syst Safety 2012;99:114–22 .

21] Francis R , Bekera B . A metric and frameworks for resilience analysis of engineered and infrastructure systems. Reliab Eng Syst Safety 2014;121:90–103 .

22] Hosseini S , Barker K . Modeling infrastructure resilience using Bayesian networks: a case study of inland waterway ports. Comput Ind Eng 2016;93:252–66 .

23] Zhao S , Liu X , Zhuo Y . Hybrid hidden Markov models for resilience metrics in a dynamic infrastructure system. Reliab Eng Syst Safety 2017;164:84–97 .

24] Woods DD . Four concepts for resilience and the implications for the future of re- silience engineering. Reliab Engi Syst Safety 2015;141:5–9 .

25] McDaniels T , Chang S , Cole D , Mikawoz J , Longstaff H . Fostering resilience to ex- treme events within infrastructure systems: Characterizing decision contexts for mit- igation and adaptation. Global Environ Change 2008;18(2):310–18 .

26] Molyneaux L , Brown C , Wagner L , Foster J . Measuring resilience in energy sys- tems: Insights from a range of disciplines. Renewable Sustainable Energy Rev 2016;59:1068–79 .

27] Arghandeh R , von Meier A , Mehrmanesh L , Mili L . On the definition of cyber-physi- cal resilience in power systems. Renewable Sustainable Energy Rev 2016;58:1060–9 .

28] Filippini R , Silva A . A modeling framework for the resilience analysis of networked systems-of-systems based on functional dependencies. Reliability Eng Syst Safety 2014;125:82–91 .

29] Lundberg J , Johansson BJ . Systemic resilience model. Reliability Eng Syst Safety 2015;141:22–32 .

30] Zobel CW , Khansa L . Characterizing multi-event disaster resilience. Comput Oper Res 2014;42:83–94 .

31] Bourouni K . Availability assessment of a reverse osmosis plant: comparison be- tween reliability block diagram and fault tree analysis methods. Desalination 2013;313:66–76 .

32] Choi IH , Chang D . Reliability and availability assessment of seabed storage tanks using fault tree analysis. Ocean Eng 2016;120:1–14 .

33] Naseri M , Baraldi P , Compare M , Zio E . Availability assessment of oil and gas pro- cessing plants operating under dynamic arctic weather conditions. Reliability Eng Syst Safety 2016;152:66–82 .

34] Cai B , Liu Y , Liu Z , Tian X , Zhang Y , Liu J . Performance evaluation of subsea blowout preventer systems with common-cause failures. J Petroleum Sci Eng 2012;90:18–25 .

35] Cai B , Zhao Y , Liu H , Xie M . A data-driven fault diagnosis methodology in three-phase inverters for PMSM drive systems. IEEE Trans Power Electron 2017;32(7):5590–600 .

36] Cai B , Huang L , Xie M . Bayesian networks in fault diagnosis. IEEE Trans Ind Inform 2017;13(5):2227–40 .

37] Fang Y , Sansavini G . Optimizing power system investments and resilience against attacks. Reliability Eng Syst Safety 2017;159:161–73 .

38] Fotouhi H , Moryadee S , Miller-Hooks E . Quantifying the resilience of an urban traf- fic-electric power coupled system. Reliability Eng Syst Safety 2017;163:79–94 .

39] Yssaad B , Khiat M , Chaker A . Reliability centered maintenance optimization for power distribution systems. Int J Electric Power Energy Syst 2014;55:108–15 .

40] Zaidi A , Ould Bouamama B , Tagina M . Bayesian reliability models of Weibull sys- tems: state of the art. Int J Appl Math Comput Sci 2012;22(3):585–600 .

41] Panteli M, Mancarella P. Modeling and evaluating the resilience of critical electrical power infrastructure to extreme weather events. IEEE Syst J 2015. doi: 10.1109/JSYST.2015.2389272 .

42] Ji C , Wei Y , Mei H , Calzada J , Carey M , Church S , White J . Large-scale data analysis of power grid resilience across multiple US service regions. Nat Energy 2016;1:16052 .

B. Cai et al. Reliability Engineering and System Safety 172 (2018) 216–224

[

[

[

[

[

43] Cai B , Liu H , Xie M . A real-time fault diagnosis methodology of complex systems using object-oriented Bayesian networks. Mech Syst Signal Process 2016;80:31–44 .

44] Cai B , Liu Y , Fan Q , Zhang Y , Liu Z , Yu S , Ji R . Multi-source information fusion based fault diagnosis of ground-source heat pump using Bayesian network. Appl Energy 2014;114(2):1–9 .

224

45] Haimes YY . On the definition of resilience in systems. Risk Anal 2009;29(4):498–501 .

46] Hollnagel E , Woods D , Leveson N . Resilience engineering, concepts and precepts. Ashgate Publishing; 2006 .

47] Attoh-Okine NO . Resilience engineering, models and analysis. Cambridge University Press; 2016 .

  • Availability-based engineering resilience metric and its corresponding evaluation methodology
    • 1 Introduction
    • 2 Resilience metric and evaluation methodology
      • 2.1 Availability-based engineering resilience metric
      • 2.2 Engineering resilience evaluation methodology
    • 3 Examples for common systems
      • 3.1 Series, parallel, and voting systems
      • 3.2 Structural modeling of dynamic Bayesian networks
      • 3.3 Parameter modeling of dynamic Bayesian networks
      • 3.4 Resilience evaluation
      • 3.5 Sensitivity analysis
    • 4 Actual application example for nine-bus power grid system
      • 4.1 Resilience of nine-bus power grid system
      • 4.2 Discussion of the proposed resilience metric
        • 4.2.1 Resilience is an intrinsic capability and an inherent attribute
        • 4.2.2 Comparison of resilience-based engineering systems
        • 4.2.3 Optimization of resilience-based engineering system
        • 4.2.4 Design of resilience-based engineering system
    • 5 Conclusion
    • Acknowledgments
    • References