literature review

profilePrab
paper5.pdf

  1  

Assessing Systemic Risk to Cloud Computing Technology as Complex Interconnected Systems of Systems

Yacov Y. Haimes and Barry M. Horowitz

Zhenyu Guo, Eva Andrijcic, and Joshua Bogdanor Center for Risk Management of Engineering Systems,

University of Virginia

15 January 2014

ABSTRACT Modeling cloud computing technology (CCT), its users, and would-be malicious intruders as complex interdependent and interconnected systems of systems (S-o-S) stems from the premise that S-o-S are inherently composed of intrinsic shared states and subsystems. This paper posits and demonstrates that unless the cyber security tools for cloud computing infrastructures provide greater protection than those for non-CCT systems, users of CCT are at greater risk than users of non-CCT systems. Recognizing that the agility of CCT systems provide advantages relative to most non-CCT systems, this paper addresses the need for the CCT community to employ these advantages toward cyber security solutions. The risk comparison builds on the following theory and methodology: (i) CCT and its users constitute complex interconnected hardware and software subsystems that interact as S-o-S through shared states and shared subsystems, which are seen as connected in series (rather than mostly in parallel for non-CCT systems); (ii) this last characteristic of S-o-S is exploited through the use of fault-tree analysis to demonstrate the resulting unreliability of CCT S-o-S; and (iii) building on the published literature, Pareto-optimal frontiers compare the risks to security-conscious users of CCT (e.g., large corporations) vs. cost- conscious users (e.g., small or start-up companies).

  2  

Preface: For reasons related to economic advantage, cloud-computing technology (CCT) is gaining a rapidly increasing user base. Our national cyber infrastructure systems (CIS)— public and private—are subject to continuous intrusive attacks by adversaries and malicious intruders intent on exploiting critical databases and the theft of intellectual property (IP) and other resources. In this paper, we contend that CCT, together with its users and would-be malicious intruders form complex systems of systems (S-o-S). Given that service providers support multiple users, the security of CCT is just as challenging, if not more so, than that of conventional computing. Here we address and provide answers to the following set of epistemological questions: (i) Why the users of CIS-CCT are at risk, (ii) Why understanding the central role of the states of CIS-CCT is critical to answering risk and security questions, (iii) Why proper and effective hardware and software systems integration is critical to answering risk and security questions, (iv) Why studying and modeling CIS-CCT as a system of systems is critical to answering risk and security questions, (v) Why, and to what extent, users of CIS-CCT are at a higher risk than those who do not use CCT, and (vi) What are the associated tradeoffs between using CCT and non-CCT technology?

While we highlight the added risks associated with CCT, we do not analytically address certain advantages that CCT systems have over non-CCT systems regarding providing important, economically satisfying risk management opportunities [U. S. Defense Science Board 2013]. Some very important advantages include: (i) the opportunity to employ the agility of CCT S-o-S for moving target security solutions, (ii) the opportunity to use the diverse geographical locations of CCT S-o-S as part of moving target security solutions, and (iii) the opportunity to employ different CCT S-o- S to provide both diverse redundancy and moving target security solutions. In addition, the business mandate for CCT S-o-S to serve a diverse customer base and rapidly respond to new needs necessarily results in agile system structures and the employment of configuration control processes and tools that are optimized for managing change; more so than a single users to whom the system would ordinarily be designed. This results in the opportunity to respond to new cyber threats more rapidly and with more assurance. Given these advantages, and the sources of risk to CCT S-o-S presented in this paper, it follows that the developers of CCT S-o-S should introduce security solutions commensurate with their opportunity for providing economical solutions and their greater risk.

A. Introduction: Understanding and quantifying the risk function associated with cyber security in general, and with cloud computing technology in particular, requires an articulation of the essential building blocks or components of the risk function associated with CIS-CCT systems of systems (S-o-S): (i) Cyber infrastructure system (CIS) connotes a complex, large-scale interconnected and interdependent cyber network designed to achieve economies of scale, encompassing hardware, software, human involvement, protocols, Internet connections, culture, policies, and organizational procedures.

  3  

(ii) Emergent forced changes (EFC) connotes external or internal sources of risk to a system that may adversely affect specific states of that system and consequently affect the entire system. (iii) Cloud computing technology (CCT) (as defined by NIST [2011-12]) is a model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage, applications, and services) that can be rapidly provisioned and released with minimal management effort or service provider interaction. (iv) CIS-CCT connotes cyber infrastructure systems (CIS) serviced by cloud computing technology (CCT). This CIS-CCT (cloud) model promotes availability and is composed of five essential characteristics, three service models, and four deployment models: (A) Essential Characteristics: (i) On-demand self-service; (ii) Broad network access; (iii) Resource pooling; (iv) Rapid elasticity; and (v) Measured Service. (B) Service Models: (i) Cloud Software as a Service (SaaS); (ii) Cloud Platform as a Service (PaaS); and (iii) Cloud Infrastructure as a Service (IaaS). (C) Deployment Models: (i) Private cloud; (ii) Community cloud; (iii) Public cloud; and (iv) Hybrid cloud. In addition to the cloud S-o-S, we will build on the following systems-based building blocks of our model [Haimes 2012] (i) Given a system’s model, the states of the system are the smallest set of independent system variables whereby the values of the members of the set at time t0, along with known inputs, decisions, random and exogenous variables, completely determine the value of all system variables for all t > t0. (ii) Vulnerability is the manifestation of the inherent multidimensional states of the system (e.g., physical, technical, organizational, and cultural) that can be subjected to a natural hazard or be exploited to adversely affect (cause harm or damage to) that system, and it is a function of the specific threat to the system and of the time frame. (iii) The resilience of a system is also a manifestation of the inherent states of the system; it is a multidimensional vector that is time- and threat-dependent. More specifically, resilience represents the ability of the system to withstand a disruption within acceptable degradation parameters and to recover within acceptable loss and time parameters. (iv) Phantom System Models (PSM) are used for intrinsic meta-modeling of systems of systems using the basic assumption that some specific commonalities, interdependencies, interconnectedness, or other relationships exist through shared and unshared states, decisions, and inputs between any two systems within any system of systems [Haimes 2012].

B. Essential Interconnected and Interdependent Subsystems of the CIS-CCT (Cloud) Systems of Systems that Constitute Sources of Risk The first step in the risk assessment process [Kaplan and Garrick, 1981] is answering the

question: “What Can Go wrong?” by building on the Modified Theory of Scenario

Structuring. [Kaplan et. al., 2001] To do this, we use hierarchical holographic modeling

(HHM) to develop the following list of emergent forced changes associated with the

subsystems of the cloud S-o-S. [Haimes, 1981, 2009] (The term emergent forced changes

connotes trends in external or internal sources of risk to a system that may adversely

  4  

affect specific states of that system.) The list represents an initial examination of systemic

sources of risk that can affect the subsystems of the CCT systems of systems, and it

provides a baseline for what we know, or what we think we know:

(i) Shared states due to shared pools of configurable computing resources, e.g., networks, servers, storage, applications, client infrastructure, and server infrastructure.

(ii) Complexity of hardware and software system integration resulting from the large scale of the systems with millions of lines of code.

(iii) Complexity related to the buffer (virtual machines) between the memory (Hypervisor) and the operating system (OS). The buffer can be closed by an adversary and affect the OS. At the core, there is a central (global) entity that has complete control, and when penetrated, problems arise and pose risk to the system.

(iv) Excess uniformity (through standards) in protocols, e.g., all operating codes in Windows™ are the same; thus, if an intruder breaks one, the rest would follow.

(v) The tradeoffs between openness (exposure to attacks) and flexibility in CIS- CCT systems invariably have been driven by cutting costs resulting in exposure.

(vi) All electronic devices have digital fingerprints. Although we can’t see what the messages are, we can follow these fingerprints, and so can would-be intruders.

(vii) No known metrics are available for users to determine whether a CCT system is working appropriately. Users know the system’s states and this constitutes “security visibility.”

(viii) The interface between private and public CCTs--a Hybrid CCT--can provide an intrusion opportunity. A private cloud can be exploited by an adversary, and pass it to a public CCT.

(ix) Metrics available to users may address the following vulnerable features of CCT systems: Authentication; authorization; control over employees; state of patching; quality of failure (resilience: how quick they can recover from delay—based on actual events); outcomes from audit; use of audit; quality of audit and its frequency; what is the cloud owner’s trust model? How can the owner’s and users trust models be compared?

(x) Multiple layers of interacting parts designed for modularity and components acquired separately constitute sources of risk.

(xi) If we obtain access to the software we can modify the hardware and the data. There is a basic issue in access control to the data and the software, but also to the hardware. We can modify the hardware when we obtain access to the software, consequently, we can maliciously modify all of them, including the data. This constitutes sources of risk.

(xii) Data is commonly not insulated; thus, when it is in memory, a few algorithms can operate on encrypted data and potentially compromise the data.

(xiii) Lack of knowledge about where sources of data are coming from. (xiv) The ability to modify a chip. (xv) Careless logging in, transferring files, etc. by users.

  5  

(xvi) Difficulty in customizing the trust model in hybrid CCT systems where the software for each customer is separate.

(xvii) Unsafe methods for isolating concurrent programs. (xviii) Buffer overflow as a risk. There is a fixed size allocation for handling data; the

buffer may fail and someone can take control of the system by injecting a new code. This is a well-known problem, and there is no silver bullet against injecting mal-virus into the system.

(xix) Lack of understanding of how many security layers are needed to bring the risk down to an acceptable level.

(xx) Inability to safely update system configurations. The above sample of sources of risk illustrate the higher risks to users of CCT-based systems. We now explore the basic characteristics of the cloud as complex interconnected and interdependent systems of systems. C. Modeling Complex Interdependent and Interconnected Systems of Systems Complex systems of systems cannot be modeled through a single model, instead, a variety of models must be built to represent the diverse system of systems perspectives. We call these models sub-models. These sub-models are integrated (coordinated) into a meta-model by the process of intrinsic meta-model coordination through shared state variables. The objective of meta-model coordination and integration is to build on all relevant direct and indirect sources of information to gain insight into the interconnectedness and intra- and interdependencies among the sub-models and, on the basis of this insight, to develop representative models of the system of systems under consideration. The coordination and integration of the results of the multiple models are achieved in the meta-modeling phase within the Phantom System Models (PSM) methodology [Haimes, 2012], thereby yielding a better understanding of the system as a whole. The PSM is a modeling philosophy that legitimizes the exploration and experimentation of out-of-the-box and seemingly irrational ideas and potentially discovers insightful implications that otherwise would have been completely missed or dismissed. In this sense, it allows for non-consensus ideas or agree-to-disagree scenarios to be further explored and studied. Through logically organized and systemically executed models, the PSM provides a reasoned experimental modeling framework with which to explore, and thus understand, the intrinsic relationships that characterize the nature of multi-scale emergent systems. The PSM advocates the use of multiple models to uncover and represent different modeling perspectives, thus providing a holographic, multi- dimensional view of a problem. The PSM-based intrinsic meta-modeling of systems of systems stems from the basic assumption that some specific commonalities, interdependencies, interconnections, or other relationships must exist between any two systems within any system of systems. More specifically:

  6  

(i) A system of systems encompasses a specific group of subsystems. Conversely, the subsystems are members of a system of systems. A model of a subsystem will be denoted as a sub-model. (ii) A meta-model represents the overall coordinated and integrated sub-models of the system of systems. We define a meta-model as a family of coordinated/integrated sub- models, each representing specific aspects of the subsystem for the purpose of gaining knowledge and understanding of the multiple interdependencies among the sub-models, and thus allowing us to comprehend the system of systems as a whole. (iii) The essence of each subsystem can be represented by a finite number of essential state variables. (The term essence of a system connotes the quintessence of the system, the heart of the system; that is, all critical aspects of the system.) The states of a system, commonly a multidimensional vector, characterize the system as a whole and play a major role in estimating its future behavior for any given input. Thus, the behavior of the states of the system, as a function of time, enables modelers to determine, under certain conditions, its future behavior for any given input, or initiating event. Given that a system may have a large number of state variables, the term essential states of a system connotes the minimal number of state variables in a model needed to represent the system in a manner that permits the questions at hand to be answered effectively. Thus, these state variables are fundamental to an acceptable model representation. (iv) For a properly defined system of systems, any interconnected subsystem will have at least one (typically more) essential state variable(s) and objective(s) shared with at least one other subsystem. This requirement constitutes a necessary and sufficient condition for modeling interdependencies among the subsystems (and thus interdependencies across a system of systems.) This ensures an overlapping of state variables within the subsystems. Of course, the more we can identify and model joint (overlapping) state variables among the subsystems, the greater the representativeness of the sub-models and the meta-model of the system of systems. (v) Having multiple, albeit overlapping, databases used by multiple sub-models, each of which is built to answer specific questions, enhances the ability to identify shared variables. Each sub-model’s characterization, whether modeled separately or as part of a group, is likely to share common state variables—a fact that facilitates the ultimate coordination and integration of the modeled multiple sub-models at the meta-modeling level. Thus, a common database that supports a family of systems of systems must be available. (vi) Bringing together multiple sub-models via intrinsic meta-modeling coordination and integration enhances our understanding of the inherent behavior and interdependencies of existing and emergent complex systems. D: Higher Risk to Cloud Computing Technology and Its Users as Complex Systems of Systems This section builds on Part I and develops theoretical and test-bed results, positing that users of cloud computing technology are more at risk than non- users of the CCT.

  7  

D.1. Cloud Computing Technology Security Issues A variety of security issues and problems associated with cloud computing have been explored by researches, the majority of which are issues with cloud management consoles and with virtual machine (VM) images. For example, Bleikertz et al. [2010] discuss vulnerabilities of cloud management consoles from a highly theoretical perspective. Janc et al. [2010] and Weinberg et al. [2011] discuss browser history issues of OpenStack VNC (Virtual Network Computing) consoles. For security issues with VM images, Dhanjani et al. [2009] address images that can include a wide variety of security weaknesses--both intentional and accidental. Bindra et al. [2012] present a brief review of some image security mechanisms, but don’t provide any original analysis. Although Vaquero et al. [2010] do not focus on OpenStack specifically, they present an interesting perspective regarding VM image security. Wei et al. [2009] present valuable viewpoints on images and security. Other relevant papers include Slipetskyy [2011] and Cigoj [2012], both of which discuss general security in OpenStack. Slipetskyy [2011] presents a methodical search for flaws and bugs that might affect security but doesn’t provide much practical, useful information for experimentation. Cigoj [2012] also discusses general OpenStack security but, effectively does not include anything original. Khan [2011] investigates how to replace simple passwords for authentication by users of OpenStack with a system called OpenID. Rocha and Correia [2011] present a useful summary of insider attacks that would be possible without complete cloud control, for example, an attack would be possible for a technician with even limited authorization. Taheri, Monfared and Jaatun [2011] discuss how the initial consequences of a successful attack on OpenStack can be mitigated (i.e., fault tolerance.) Luo et al. [2011] discuss general classes of attack forms in virtualized environments. Nirmala [2012] presents methods of distributing data among clouds such that an insider at any one cloud cannot interpret the data meaningfully. The author also addresses how this affects processing speed. Bessani et al. [2011] discuss a similar use of multiple clouds rather than a single cloud for risk management of an insider attack. Roman et al. [2012] conduct a comparative security analysis of several cloud systems, including OpenStack. However, the authors focus only on storage such as Amazon’s S3 service rather than the more general case of combined processing and storage services. Kim et al. [2013] discuss some OpenStack security issues, but focus mostly on denial of service attacks, rather than attacks on confidentiality or integrity. The one exception is a SQL injection attack, but the authors give very few details. D.2. Premise: Users of cloud computing technology (CCT) are at a higher risk than users of non-CCT systems. Under the following principles and assumptions, and for a type of ubiquitous cyber attack targeting non-specific data/information, a public cloud system, being a system of systems, is more a risk to users than a non-cloud system. To validate the above premise, we must compare a cloud and a non-cloud system under similar conditions. In this research project, we explore the public CTT as a system of systems (CCT S-o-S.) In the following discussion we build on our earlier premises that the vulnerability and resilience of a system are manifestations of the states of that system and that they are functions of the specific emergent forced changes (threats); similarly,

  8  

the risk to a system is a function of all of the above. Furthermore, the multidimensional probabilistic consequences resulting from a malevolent intrusion into a CCT S-o-S, necessarily yield a multidimensional risk function whose modeling presents a considerable challenge. An example of this challenge is the selection of appropriate models to represent the essence of a CCT S-o-S including its multiple subsystems, functions, and capabilities. We build on the following assumptions: 1. The same type of intrusion/threat is applied to both systems. 2. Both CCT S-o-S and non-CCT systems operate under the same security standard. 3. The interconnected subsystems that constitute CCT S-o-S share functional components and states, and the multiple users are mostly unknown to one another. Some of the users could be intruders. 4. The CCT S-o-S is actually a complex interconnected and interdependent system of systems with shared states in the context of the Phantom System Model (PSM). 5. The CCT S-o-S reliability metric is used as a surrogate for intrusion risk.

a. The CCT as S-o-S share functional components, and thus, share states. Furthermore, two subsystems that share one or more states are more vulnerable to the same threat than a system that has no shared states. Thus, shared-state systems are more at risk because an intruder has more than one path to use in penetrating the CCT S-o-S. Therefore, the ability of an intruder to penetrate (compromise) a CCT S-o-S, which is characterized by shared states, implies that all subsystems of the CCT S-o-S with the shared states will be equally compromised.

b. To complement the above premises with a quantitative dimension, we build on fault-tree analysis and minimal-cut sets. The failure probabilities of subsystems of the CCT S-o-S are dependent on the failure of other interconnected subsystems with shared states. The failure probabilities can be evaluated using conditional failure probabilities because intrusion into any subsystem with shared states implies the ability to penetrate other subsystems with those shared states. Thus, in comparison with non-CCT systems, the shared states are synonymous with replacing independent subsystem failure probabilities with conditional subsystem failure probabilities (in the parlance of fault-tree analysis they become connected in series) and thus add vulnerability to the CCT S-o-S.

Furthermore, the reliability of interconnected subsystems with shared states (i.e., modeled by conditional failure probability, i.e., connected in series) is always lower than the reliability of subsystems that do not share states (i.e., modeled by independent failure probability.)

Let Ri be the independent reliability of subsystem i, Ri|j be the conditional reliability of subsystem i, given the failure of subsystem j, Rs-CCT be the reliability of the entire CCT system, and Rs-non-CCT be the reliability of the entire non-CCT system. The following inequality holds for subsystems connected in any form:

Ri|j < Ri and Rs-CCT < Rs-non-CCT. This is because for a system whose subsystems share states with other subsystems, the failure of any one subsystem would cause the failure of other interconnected subsystems.

  9  

Thus, from the perspective of reliability theory and fault-tree analysis, the CCT S-o-S will has a lower reliability, i.e., a higher risk of tangible and intangible risk than non-CCT systems.

In sum, the probability of adverse consequences to the CCT interconnected as a S-o-S, with the corresponding tangible and intangible risks, is higher than for the non-CCT systems.

D.3. Numerical Demonstration of the Premise with a Mini-CCT S-o-S This section provides a demonstration of the premise through a quantitative fault-tree analysis on a virtual mini-CCT S-o-S. The example virtual mini-CCT S-o-S is constructed according to literature reviews of different cloud system architectures. To focus on the proof of concept, only essential components and functionalities are included in the mini-CCT. Subsystems of the CCT S-o-S with their interconnectedness and interdependencies are identified. A fault tree model is developed based on identified subsystems and their potential failure modes. The reliability of the CCT S-o-S and the non-CCT system is calculated and compared assuming conditional failure probability and independent failure probability of each subsystem. The numerical analysis shows that the reliability of a CCT S-o-S is much lower than that of a non-CCT system, which is consistent with the theorem. Although this analysis is based on subjective assessment of failure probabilities of subsystems, it addresses what type of data needs to be collected and analyzed in future research. D.4. Mini-CCT S-o-S The National Institute of Standards and Technology (NIST) definition of cloud computing [NIST Special Publication 800-145] classifies CCT systems into different service models (i.e., Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (IaaS)) and deployment models (i.e., private, community, public, and hybrid.) Current CCT S-o-S providers in the marketplace employ specific technologies and cloud architectures to meet the different needs of a broad range of cloud users. Based on our literature review, there is no standard technology or system architecture for CCT S-o-S. To demonstrate the theorem with a real CCT S-o-S, a conceptual mini-IaaS public-CCT system is constructed in Figure 1. Each box in the figure represents an essential component or functionality of the CCT S-o-S.

  10  

Figure 1: Essential Components and Functionalities of a Simplified IaaS Public Cloud S-o-S

To demonstrate the theorem, the components and subsystems of the mini-CCT S-o-S must be identified. There are multiple perspectives through which a specific CCT S-o-S can be decomposed, such as physical, functional, and stakeholder perspectives. From a physical perspective, an IaaS public-CCT S-o-S can be decomposed into the following subsystems and components: (i) Hardware: (a) hard drives and memory; (b) CPU; and (c) switches and cables. (ii) Software: (a) system software (hypervisor, operating system, middleware); (b) application software (applications and APIs); (c) communication software; and (d) security software. (iii) Physical infrastructure: (a) data center; (b) physical security system; and (c) power supply. (iv) Human: (a) programmer/developers; (b) administrator; and (c) operation and maintenance staff. This decomposition perspective provides very few insights into the security aspects of the CCT S-o-S since the overall security functionality cannot be deduced from the functionality of the individual components. In most cases, hardware, software, and human perspectives need to be integrated and work together to provide meaningful functionality to the CCT S-o-S.

Cloud customer j

Cloud customer k

Security Service

Infrastructure (Storage, Computing, Network)

Cloud Provisioning and Management

Simplified Infrastructure as a Service Public Cloud System of Systems

Cloud Provider

Hypervisor

…Virtual Machine (VM) 1

VM 2

VM 3

VM n

  11  

From a functional perspective, an IaaS public-CCT S-o-S can be decomposed into the following subsystems and components: (i) Infrastructure: (a) storage resources; (b) computational resources; and (c) network resources. (ii) Platform: (a) provisioning tool; (b) monitoring and metering tool; (c) systems management (performance, capacity, availability); (d) service catalog; (e) database; and (d) runtime environment. (iii) Application: (a) user interface; (b) machine interface; and (c) service management. (iv) Security: (a) authentication; (b) authorization; (c) auditing and logging; (d) firewall; and (e) encryption and entitlement management. The relationship among these functionalities constitutes the architecture of the CCT S-o- S. A combination of these functions is used to achieve the “Essential Characteristics” of a CCT S-o-S defined by NIST (on-demand self-service, broad network access, resource pooling, rapid elasticity, and measured service.) In the above two decomposition schemes, each subsystem has a specific function and the failure of one subsystem leads to the failure of the CCT S-o-S. However, in this study we are not concerned with system failure, we are interested in the security failure of the CCT S-o-S due to internal or external emergent risks. A decomposition scheme based on the stakeholder’s perspective is proposed. From a stakeholder’s perspective, an IaaS public-CCT S-o-S can be decomposed into the following subsystems and components: (i) Normal user: (a) virtual machine; (b) physical machine; (c) computing tasks; and (d) data. (ii) Attacker; (b) virtual machine; (c) physical machine; and (d) malware. (iii) Operation agent within the provider: (a) physical machine; (b) Hhypervisor; and (c) provisioning and system management. (iv) Security agent within the provider: (a) firewall; (b) encryption; (c) authentication; and (d) authorization. Figure 2 shows the same mini-CCT S-o-S that has been decomposed into four subsystems based on the stakeholder’s perspective. Multiple subsystems in this decomposition scheme may share components (or states) of the CCT S-o-S. For example, the normal user (cloud customer j) may share hypervisor and infrastructure components with the intruder (cloud phantom-customer k); the intruder may share the virtual machine, hypervisor and infrastructure components with the operation agent; and the operation agent may share the hypervisor and infrastructure components with the security agent. These shared components constitute the interconnectedness and interdependency among these subsystems and a major source of emergent forced changes to the CCT S-o-S.

  12  

Figure 2: Identified Subsystems of a Simplified IaaS Public Cloud SoS

D.5. Building a Fault Tree for the Mini-CCT S-o-S To focus on the risks to CCT S-o-S from an intruder’s perspective, fault tree analysis is used to represent the functional relationship between subsystems and the failure event. A fault tree analysis can be simply described as an analytical technique, whereby an undesired state of the system is specified, and the system is then analyzed in the context of its environment and operation to find all credible ways in which the undesired event can occur. A fault tree is a complex of entities known as “gates” which serve to permit or inhibit the passage of fault logic up the tree. The gate symbol denotes the type of relationship of the input events required for the output event. A typical fault tree is composed of a number of symbols. There are two basic types of fault tree gates: the OR- gate and the AND-gate. The OR-gate is used to show that the output event occurs only if one or more of the input events occur. The AND-gate is used to show that the output fault occurs only if all the input faults occur [Fault Tree Handbook, USNRC 1981].

.

When subsystems are connected through an OR-gate (in series), the system fails when at least one of its components fails. The reliability of the system is represented as

Cloud customer j

Cloud customer k

Security Service

Infrastructure (Storage, Computing, Network)

Cloud Provisioning and Management

Simplified Infrastructure as a Service Public Cloud System of Systems

Cloud Provider

Hypervisor

…Virtual Machine (VM) 1

VM 2

VM 3

VM n

User j

User k

Security

Operation

Subsystems

  13  

, and . When subsystems are connected through an

AND-gate (in parallel), the system fails only when all of its components fail. The

reliability of the system is represented as , and

A minimal cut set is defined as the smallest combination of component failures, which if they all occur, will cause the top event (i.e., failure) to occur [U.S. Nuclear Regulatory Commission, 1981]. By definition, a minimal cut set is a combination of intersections of primary events in parallel sufficient for the top event to occur (if all parallel components fail.) This combination is the “smallest” combination in that all the failures in the minimal cut set need to occur for the top event (system failure) to occur. If any one component in the parallel combination does not occur, then the top event will not occur (by this combination.) A fault tree will consist of a finite number of minimal cut sets, all of which are in series, which are unique for the top event to occur. Since the combination of all minimal cut sets is in series, then the failure of any cut set will cause the failure of the entire system. That is, once the minimal cut sets are known, then any system can be written as a series arrangement of its cut sets with the components of each minimal cut set arranged in parallel. In sum, the one-component minimal cut set represents a single failure that will cause the top event to occur. The two-component minimal cut set represents double failures that together will cause the top event to occur. For an n- component minimal cut set, all n components in the cut set must fail in order for the top event to occur. The literature suggests that a cloud intruder constitutes a major source of risk for cloud users; thus, in this study the top event of the fault tree for the mini-CCT system is defined as “Attacker gains access to normal user’s confidential data.” Examples identified in the literature include Amazon’s cloud computing service Elastic Computer Cloud (EC2) being used by a researcher to fire 400,000 passwords a second at a secured Wi-Fi network. It took the researcher 20 minutes to hack into the system [The H Security 2011], Homeland Security News Wire, [2011]; Epsilon, a cloud-based marketing service, experienced a cyber attack in which data (e.g., email and bank account details) of its customers, including JP Morgan Chase, Citibank, Barclays Bank, Marriott and Hilton, was exposed to the hackers. [Bhadauria and Sanyal] Three potential events can individually cause the top event:

1. Event E1: Attacker accesses data by real-time intrusion (concurrent sharing) 2. Event E2: Attacker accesses data by leaving a Trojan horse (attacker-to-user

sharing) 3. Event E3: Attacker accesses data by reading residual traces (user-to-attacker

sharing) Each of these events alone will cause the top event to occur, thus they are connected through an OR-gate. The basic fault tree is shown in Figure 3.

∏ =

= n

i iS tRtR

1

)()( )}({min)( tRtR iiS <

∏ =

−−= n

i iS tRtR

1

)](1[1)(

)}({max)( tRtR i i

S >

  14  

Figure 3: Fault Tree Top Events and Three Potential Failure Modes

The necessary conditions (basic failure events) for each event are developed in detail in Figure 4.

  15  

  16  

Figure 4: Subsystem Failure for Each Sub-tree

Each basic failure event in a circle represents a specific failure mode of a subsystem. They must all occur to cause the top event (failure) to occur, thus they are connected in parallel, through an AND-gate. Quantitative Analysis To compare the reliability of a CCT system with a non-CCT system, a quantitative fault- tree analysis needs to be performed. As shown in the theorem, the failure probability of basic events (subsystems) in non-CCT systems is independent of each other while the failure probability of basic events (subsystems) in a CCT S-o-S is dependent on other interconnected subsystems through shared states. Two different sets of failure probabilities, independent probability for non-CCT systems and conditional probabilities for a CCT S-o-S are used for the same fault tree and the failure probabilities of the top event are compared for the two cases. In this paper, and at the current stage of this study, the failure probabilities of basic events are based on subjective assessment and they only represent an approximation of the order of magnitude of the events. A more objective estimation can be performed based on a Hadoop Cloud S-o-S. A summary of elicited failure probabilities for each basic event is shown here.

1. A1: Subsystem “User k” Failure – User k is a potential malicious attacker (P =

  17  

0.8) 2. B1: Subsystem “Operation” Failure – Hypervisor allocates user j and k to the

same infrastructure (P = 0.001) 3. B2: Subsystem “Operation” Failure – Malware is not deleted (P = 0.01) 4. B2’: Subsystem “Operation” Failure – Malware is not deleted, and user k is an

intruder (P = 0.1) 5. B3: Subsystem “Operation” Failure – User j’s data are not deleted (P = 0.001) 6. C1: Subsystem “Security” Failure – User k gains access to storage of other

cloud users (P = 0.001) 7. C1’: Subsystem “Security” Failure – User k gains access to storage of other

cloud users given user k is an intruder (P = 0.01) 8. C2: Subsystem “Security” Failure – Fails to detect malware (P = 0.1) 9. C2’: Subsystem “Security” Failure – Fails to detect malware, because user k is

an intruder (P = 0.5) 10. C3: Subsystem “Security” Failure – Establishment of unauthorized data

channel (P = 0.001) 11. C3’: Subsystem “Security” Failure – Establishment of unauthorized data

channel because user k is an intruder and malware is not deleted (P = 0.01) 12. D1: Subsystem “User j” Failure – Encryption key is compromised (P = 0.001) 13. D1’: Subsystem “User j” Failure – Encryption key is compromised because

malware is not deleted (P = 0.01) 14. D1’’: Subsystem “User j” Failure – Encryption key is compromised because

user k is an intruder (P = 0.01) Using Boolean algebra, the top event of the fault tree can be expressed as

, and three minimum cut sets can be identified: , , and . Based on above assumed failure probability, the failure probability of the non-CCT system (independent probability) is P(E) = 8.8E-09 and the failure probability of the CCT system (conditional probability) is P(E) = 4.08E-06. Summary of Fault-Tree Analyses The above results show that the reliability of CCT S-o-S is much lower than non-CCT systems due to subsystem interdependency, which is consistent with the theorem (the users of CCT S-o-S are at a higher risk than users of non-cloud technology. Although the magnitude in difference of reliabilities is currently based on a subjective assessment of the subsystem failure probabilities, the general conclusion above can be made as long as it is true that conditional failure probability is greater than the independent failure probability. Real CCT S-o-S are much more complex than the example virtual mini-CCT S-o-S built in this section, however, the concept of interdependent subsystems with shared states and the resulting interdependent failure probabilities can be used to analyze the reliability of such complex systems. The fault-tree model helps in developing risk management options for CCT S-o-S. The reliability of each minimum cut set mainly depends on the most reliable component in

1131132211111 DCBADCCBADCBAE ⋅⋅⋅+⋅⋅⋅⋅+⋅⋅⋅= 1111 DCBA ⋅⋅⋅ 13221 DCCBA ⋅⋅⋅⋅ 1131 DCBA ⋅⋅⋅

  18  

that minimum cut set. Limited risk management resources should be allocated to these components to improve the reliability of each minimum cut set. On the other hand, since all of the minimum cut sets are connected in series, the reliability of the system mainly depends on the most unreliable minimum cut set. Limited risk management resources should be allocated to that minimum cut set to improve the reliability of the whole system. The fault-tree model also points out the type of data to be observed, monitored, recorded, and analyzed in order to perform the above analysis. It is important for CCT S-o-S providers and users to use the correct performance metrics data to foresee emergent forced changes, and to make informed risk-assessment management decisions.

D.6. Economic Analysis of the Security of CCT S-o-S In this section we address the economics of cloud security. We introduce cloud- computing functionalities that can be used to create, or to enable, cloud cyber attacks, and we discuss the relatively low cost of carrying out cloud cyber attacks. We suggest that for certain types of ubiquitous, non-specific cyber attacks, companies that use the cloud are more at risk than companies that are not using the cloud. We then highlight the fact that, despite the likely increased risk of certain types of cyber attacks targeting confidential and sensitive data stored in the cloud, and despite the recognized fact that security may not be the highest priority for cloud providers, companies switch to the cloud to gain measurable economic benefits. The prior sections of this paper suggest that CCT is more prone to certain types of cyber attacks, which implies that companies that are considering switching to the cloud ought to be cognizant of the increased risk of loosing confidential and sensitive information when using cloud-based applications. However, we suggest in this section that a company’s decision to switch to the cloud is primarily governed by potential cost savings, and that if the savings are large enough, a company may switch to the cloud regardless of the increased risks. Our analysis points to the fact that the risks associated with cloud computing must include a consideration of the economic tradeoffs between switching and not switching to CCT. D.6.1. Cyber Attacks in the CCT S-o-S There are several cloud-computing functionalities that can be used to create or enable cloud cyber attacks originating from within the cloud. [CloudBus, 2011] These include rapid elasticity (a large amount of computational power can be quickly utilized to initiate and execute an internal cloud cyber attack); on demand service means that essentially anyone with a valid email and credit card can sign up for a cloud service and potentially launch attacks from the cloud; resource pooling suggests that malicious users launching an attack from within can utilize as many resources as they need because the service provider supports everyone’s applications; broad network access implies that cyber attacks can be launched from within a cloud from any digital device; and measured service combined with resource pooling implies that certain types of cyber attacks could be launched from within the cloud very cheaply. And the threat doesn’t stop there – once

  19  

the potential attacker is in the cloud, he or she has potential access to many the companies that store their information on cloud servers, thus, if an attacker is interested in ubiquitous non-specific attacks, i.e., is not looking for specific IP data belonging to a specific company, but is more interested in any type of information he/she can obtain, then, theoretically, with a single attack he or she can gain control of large amounts of confidential information. A theoretical example of such an attack is the hyperjacking attack in which a malicious user installs a rogue hypervisor that takes complete control of the servers. In such an attack regular security measures are ineffective because the OS will not be aware that the machine has been compromised. [Sarno and Rodriguez, 2011; McKay, 2011] But clouds might not be worse than traditional computing environments for all types of cyber attacks. For example, recent cyber-attacks have demonstrated that cloud computers can handle distributed denial-of-service attacks better than traditional servers can, precisely because they enable a quick instantiation of additional resources that can more efficiently handle increases in requests. [Johnson, 2010] Furthermore, non-cloud computing environments (e.g. traditional enterprise systems) may be more vulnerable to cyber attacks that target specific IP data, since such attacks in the cloud would be ineffective unless the exact location of the information on the server was known. D.6.2. The cost of cyber attacks internal to the cloud The Stratsec Winter School [Hayati, 2012] performed research to explore whether the cloud-computing environment provides certain benefits for cyber attackers and whether the cloud platform can be utilized to launch cyber attacks. Furthermore, they explored the ability and consistency of cloud providers to detect and report cyber security breaches in the cloud. They set up a botCloud (a group of cloud instances that are commanded and controlled by a malicious attacker), with which they attacked the victim hosts, which were also set up virtually in a controlled network environment. After performing four different experiments with different cases of malicious traffic and intrusion detection systems, the researchers concluded that, in general, cloud providers did not notify cloud users of the malicious activities occurring in the cloud, and they did not respond to the attacks by resetting or terminating connections. Aided by the inadequate response by cloud providers, these attacks were easy for malicious users to set up and use, as they required minimal knowledge of the internal architecture and system administration. Furthermore, the cloud environment made it effortless to create uncountable cloned opportunities from which to launch attacks, and they required attackers to have no significant physical infrastructure of their own. Finally, the availability and elasticity of resources made this attack very cheap.

The above example shows that the cloud-computing environment has opened doors to cyber attacks that can be significant in scope while requiring little physical infrastructure and costing little to design and execute. Another example of this is Kaspersky’s analysis of the Flame malware construction, in which, by utilizing Amazon’s EC2, Kaspersky was able to create a forged code-signing certificate at a cost of $200,000 [Cromwell]. Researchers estimate that the cost of creating this certificate without utilizing cloud-

  20  

computing resources would be over $1,000,000. Roth, who launched a dictionary attack on a secured Wi-Fi network through Amazon’s EC2, suggested that an optimized version of his software that could “crack Pre-Shared Keys (PSKs) in six minutes would enable an attack that would cost a total of $1.68.” [The H Security, 2011]

D.6.3. Security implications for cloud users Easy access to the cloud by anyone who has a valid email and credit card suggests that the cloud may be easier to infiltrate than a non-cloud computing environment, in the sense that the initial access to the cloud requires a low level of technical ability and a small investment. Furthermore, as was already discussed, the structure of the cloud and its functionalities facilitate an environment in which it is easy for somebody who is already in the cloud and has malicious intentions to create and propagate a cyber attack, again with a relatively low budget. Thus, the cloud is more vulnerable to ubiquitous, non- specific cyber attacks in which the attackers have no desire to access specific information, e.g., intellectual property of a specific company, but are more interested in accessing any confidential or sensitive information they can find. Several questions deserve clarification: What are cloud customers doing to protect themselves against such attacks? Are cloud customers aware that they might be exposed to a greater risk in the cloud? Are they requiring a higher level of security? How much are they willing to pay for security in the cloud? Surprisingly, studies show that cloud customers are not well acquainted with what security they are paying for, and what it protects against, and furthermore, many admit to having lowered their security requirements after switching to the cloud. Ponemon Institute’s study of cloud-computing service providers suggests that, of the U.S. providers surveyed, 73% believe that their cloud services do not secure and protect confidential information. [Ponemon Institute, 2011] The report suggests that most cloud providers do not consider security as their priority and believe that protecting the customer’s data is not their responsibility. The report suggests that “neither the company that provides the service nor the company that uses cloud computing seem willing to assume responsibility for security in the cloud. In addition, cloud computing users admit they are not vigilant in conducting audits or assessments of cloud computing providers before deployment.” [Ponemon Institute, 2010] The literature is sparse regarding data on the actual dollar value of security expenditures in the cloud, and most estimates are gathered at a higher level by considering total IT budgets of companies. Figure 5, obtained from PriceWaterhouseCoopers, shows the IT spending as a percent of revenues for major U.S. industries in 2005. [PriceWaterhouseCoopers, 2008]

  21  

Figure 5: IT expenditures as % of revenues for U.S. industries (obtained from Gartner IT Key Metrics Data 2006)

The question then becomes how we can utilize this data, along with some general economic performance data for relevant economic sectors, to determine how companies are implicitly trading increased profits for potentially decreased IT security. A majority of companies (73% according to the Ponemon Institute [2010]) switched to the cloud for performance reasons, e.g., savings in capital and operating costs that can be directed toward increasing net profit margins, potentially enabling the cloud-using companies to remain economically competitive or increase their market share. Many of the companies that switched to the cloud believed that in the process they accepted an overall lower level of security posture, yet many of them place sensitive and confidential information in the cloud. Even more importantly, many of the companies do not know what level of security service the cloud provides and they are doubtful of the cloud providers’ ability to protect their data. Even some cloud providers are doubtful of their ability to protect the data that is stored on their servers. This suggests that some companies value the additional profit they might receive by switching to the cloud more than they value data security and potential loss of reputation and business. In this analysis we examine what level of savings would justify companies of different sizes, different histories, and from different sectors accepting a potentially lower or ambiguous level of security that may be present in the cloud. Since we do not have data that would enable us to compare the levels of security of a specific company with respect to a specific threat to the cloud and non-cloud environments, we will instead compare the net profit margin increase that companies might be able to achieve by reducing their IT expenditures by moving to the cloud environment. In essence, through this analysis we are indirectly assessing how much security companies might be willing to give up in order to improve their net profit margin, i.e., to remain competitive. We assume the following:

1. Companies transfer sensitive and confidential information to the cloud. (According to a Ponemon Institute survey on encryption in the cloud [2012], 50%

  22  

of U.S. companies that participated in the survey transferred confidential and sensitive information to the cloud.)

2. Companies understand that their data might be more at risk in a cloud-computing environment. (According to a Ponemon Institute survey on encryption in the cloud [2012], 41% of U.S. companies that participated in the survey believe that their security posture was decreased by moving to the cloud.)

3. Companies in the cloud are not well aware of what the cloud provider is doing to protect their data. (Ponemon Institute’s survey on encryption in the cloud [2012], suggests that 54% or U.S. companies that responded to the survey were not aware of security features in the cloud.)

4. Companies in the cloud, as well as cloud providers themselves, are not confident in the security features provided by the cloud provider. (Ponemon’s Survey [2010] suggests that over 40% of cloud providers are not confident that they are meeting the customers’ security requirements.)

5. Companies transfer IT operations to the cloud primarily to save money, and security concerns are not of primary interest to most companies. (According to a Ponemon Institute Survey [2010] 73% of companies migrated to the cloud in order to reduce costs, and only 14% migrated to the cloud to improve security.)

6. On average, companies spend approximately 10% of their total IT budgets on IT security. [Kerk et al., 2006; Schwartz, 2011; Brenner, 2009]

7. On average, companies (not including governmental agencies) save 20-30% in costs by switching to the cloud. [Leung, 2010; Wright, 2012] We assume these savings come from reducing IT costs, thus we assume scenarios in which the IT costs are reduced 0 – 30% in increments of 5%.

8. We assume that companies are in business to make as much money as they can and any reduction in costs will be redirected to increasing net profits. Thus, we assume that the IT costs that are saved by switching to the cloud are directly transferred into net profit, i.e., for a company whose IT budget is 10% of revenues, and whose net profit margin is 2%, net profit margin will be increased by 2% if IT costs are reduced by 20%.

9. Companies of different sizes and of longer or shorter histories will have varied concerns over loss of reputation that might ensue from a cloud cyber-breach. For example, bigger, more established companies will, in general, care more about their reputations than smaller, start-up companies, and will thus require a larger net profit increase from an implied reduction in security. Thus, net profit increase vs. reduction in security cost curves will be different for companies in different sectors and of different sizes and histories.

10. The quantitative risk function of companies that are switching to the cloud is a multidimensional function composed of the following:

a. Risk of cyber incident (loss of reputation, clients, and money) – This may be of more concern for larger and more established companies that have created a high level of trust with their large client base, than with smaller, newer companies that are more driven to increase their profits at all costs, and do not have an established client base.

b. Risk of giving up existing IT staff, that is, not being able to return to the old configuration if the cloud environment proves inefficient. This may be

  23  

of more concern to larger and more established companies that have established legacy systems, and potentially large IT departments, than to smaller, newer companies that might not have extensive IT departments of their own.

c. Risk of not being able to switch from one cloud provider to another – contract breakage fees may be high, as well as the risk that some personal or confidential information might remain with the original cloud provider.

d. Risk of no cost guarantee - cloud costs are not guaranteed forever. Based on estimates of net profit margins of different economic sectors [Butler Consultants; Yahoo Finance, 2012] and estimates of IT expenses as a percentage of revenues for different economic sectors [PricewaterhouseCoopers, 2008], we consider the following five economic sectors: banking and financial services, professional services, construction and engineering, manufacturing, and health care, shown in Table 1.

Table 1: Net profit margins, IT expenditure as % revenues, and security expenditures as % of IT for five economic sectors in the U.S.

For these five sectors, we consider the following six scenarios:

1. Migration into cloud will reduce the IT expenses by 5% 2. Migration into cloud will reduce the IT expenses by 10% 3. Migration into cloud will reduce the IT expenses by 15% 4. Migration into cloud will reduce the IT expenses by 20% 5. Migration into cloud will reduce the IT expenses by 25% 6. Migration into cloud will reduce the IT expenses by 30%

For each of these sectors, we explore the increase in the net profit margin that results from the six levels of IT cost reduction. We assume that all IT cost savings are directly redistributed to net profit, hence, net profit increases with any reduction of IT costs. Results for the five economic sectors are shown in Figure 6.

  24  

Figure 6: Changes in net profit margin versus reductions in IT costs from switching to the cloud

Figure 6 indicates that under the aforementioned assumptions, a 30% reduction in IT costs for a company in the banking and finance sector could translate into a 1.6% increase in net profit. The same amount of IT cost reduction could translate into a 0.5% increase in net profit for the construction and engineering sectors. We can also consider how the potential increase in net profit translates into value for company shareholders. We do this by considering two financial metrics, namely the earnings per share and return on equity. The earnings per share (EPS) metric indicates the amount of earnings for each outstanding share of a company’s stock This metric is the single most important determinant of a company’s stock price and is computed as a ratio of net income to average common shares. Any increase in net income, while keeping average common shares constant, results an equal increase in the EPS. An increase in the EPS then translates into higher stock prices, which are desirable to a company’s shareholders, and are thus desirable to the company itself. In other words, companies do what they can to improve their EPS and stock prices, and this can be done efficiently by increasing their net incomes. The second metric, return on equity (ROE), indicates the rate of return of ownership interest, and like the EPS, a higher value is preferred. It is computed as the ratio of net income and average shareholder equity per period, thus any increase in net income, while keeping the average shareholder equity per period constant, results in an increase in the ROE. Figure 6 could be redone by switching the change in the net profit for a change in EPS or ROE. By relating the IT savings generated from switching to the cloud to EPS or ROE, companies may be persuaded to switch to the cloud in order to remain competitive in the marketplace. This will, of course, depend on the company’s existing net profit, size, and reputation concerns.

  25  

The generic curves in Figure 6 do not account for the size of a company, its economic sector, or for its relative concern with the four risk factors mentioned in Assumption 10 regarding qualitative risk functions. However, if we were to express these four risk factors in terms of multipliers, then we could apply those multipliers to these generic curves and get a better estimate of the actual relationship between net profit and reduction in IT costs for companies of different sizes and with different histories. For example, we could develop a reputation concern multiplier, which, as was already discussed, would generally be larger for larger companies having established histories and a large number of customers, as opposed to small start-ups that have not yet built credibility with their customers. Thus, a company more concerned with its reputation would require a larger increase in net-profit gained by switching to the cloud than a company that is less concerned with reputation. Similarly, we could develop a loss of in-house IT staff risk multiplier, which again, would generally be larger for larger companies. These multipliers could be broken down by industry type, company size, and history and they could then be applied to the generic curves shown in Figure 6, to show how valuable increases in net-profit are to different companies in different sectors, and how these potential increases might convince a company to switch to the cloud, thereby accepting a potentially lower level of security. While we currently do not have data to construct these multipliers, theoretically speaking such data could be obtained from case studies and surveys. With additional data about the cost of security provided in the cloud, and more accurate data about financial losses from breached confidential and sensitive data in the cloud, we could start to create Pareto-optimal curves like the one in Figure 7. This would enable us to more accurately say how much security various companies are willing to give up to reduce their costs.

  26  

Figure 7: Illustration of a Pareto optimal curve for different companies in the cloud*  

We show in this analysis that the risk to a company of not switching to the cloud is not trivial – it could mean loosing the competitive advantage by not being able to reduce expenditures, increase net profit, and increase the value of the company for shareholders. We argue that many companies take this short-term view of the high economic risk of not switching to the cloud without sufficiently considering, at the same time, the risks of the potential loss of confidential and sensitive information in the cloud. D.6.4. Conclusions When assessing risks of cloud computing, companies should not only consider the risks of possible loss of data and reputation, they should also consider the economic risks of not switching to the cloud due to market competitiveness. In reality, companies that are deciding whether to switch to the cloud must compare the risk of not switching to the risk of switching, and we provide a simple analysis to show that, in some cases, the potential economic losses of not switching to the cloud might be substantial. Our analysis is simplistic, but it is grounded on the basic business principle that companies are in the business of making money for shareholders. They do whatever they can to provide a higher return on investment for their shareholders, even if it means accepting a cloud- computing environment in which the level of security is not always clear to the company or to the provider. We do not have data to fully understand the risk of switching to the cloud, but by conducting this type of analysis we are “simulating” the type of thinking that occurs in companies when they are deciding whether to switch to the cloud. We are indirectly assessing how much security companies might be willing to give up in order to improve their net profit margin and remain competitive in the market. Hence, this type of

  27  

analysis could provide an insight into what companies implicitly believe the risks of cloud computing to be. E. Conclusions and Lessons Learned

Modeling cloud computing technology, its users and would-be malicious intruders as interconnected and interdependent complex systems of systems (S-o-S) enabled us to apply theoretical, analytical, and methodological foundations from three fields: (i) advances in systems engineering in modeling complex (S-o-S) by building shared states among subsystems and use of the Phantom Systems Model; (ii) principles and guidelines in risk assessment, management, and communication; and (iii) cyber security and cloud computing technology. Indeed, the builders, users, and would be malicious intruders all are learning as they go through the process of discovering the emergent benefits as well as the potentially undesirable risks associated with the significant tangible and intangible losses and other consequences of CCT. This is true for government, public, and private organizations.

We posited the following premise: “Cloud computing technology (CCT) as systems of

systems are at a higher risk than non-CCT systems.” Its validity was demonstrated under similar non-CCT conditions through quantitative fault-tree analysis on a virtual mini- CCT system. We further developed a fault tree of a virtual mini-CCT system of systems with the following components:

(i) an intruder constitutes a major risk to cloud users; (ii) the top event of the fault tree for the mini-CCT system is defined as “Attacker gains access to normal user’s confidential data;” (iii) three potential events were considered to cause the top event, namely, a successful malicious intrusion—(a) Event E1: Attacker accesses data by real- time intrusion (concurrent sharing); (b) Event E2: Attacker accesses data by leaving malware (attacker-to-user sharing); and (iii) Event E3: Attacker accesses data by reading residual traces (user-to-attacker sharing).

The fault-tree model provided a quantitative demonstration of the premise that users of CCT S-o-S are more at risk than users of non-CCT systems. The important findings from this demonstration are Event C3 (Establishment of unauthorized data channel) and D1 (Encryption key is compromised) in minimum cut set 2 means that cut set 2 should have top priority for improvement. The concept of interdependent subsystems with shared states and the resulting interdependent failure probabilities can be used to analyze the reliability of other complex S-o-S. Our results address what type of data is needed, and should be collected and analyzed, for future research.

On the economics of CCT S-o-S, we found that detailed information is available

regarding spending on security by cloud providers, or regarding what cloud users are willing to pay for a certain level of security in the cloud. Importantly, companies switch to the cloud by evaluating the potential level of profit increase that could result from a reduction in operating and capital costs, namely in IT.

Finally, while we have not performed a risk management analysis, we have highlighted that CCT system designs are well aligned with providing risk management opportunities that non-CCT systems cannot provide within the same economic profile or

  28  

with the same speed to implementation. It then becomes incumbent for those CCT providers who wish to address security that can compete with non-CCT systems to implement solutions that exploit the inherent advantages of CCT system designs. Acknowledgements The research presented in this paper was supported by the Institute for Information Infrastructure Protection (I3P) through a grant from the U.S. Department of Homeland Security (DHS), from the National Science Foundation (NSF), grant number XXXX, and from the University of Virginia funds in support of the Phantom Systems Models Laboratory (PSML) at the Center for Risk Management of Engineering systems. The authors acknowledge the valuable technical advice and support received from Dr. Thomas Longstaff, Dr. Matt Henry, and Dr. Dickie George; the help and support received from the I3P staff and leadership, the earlier contributions of our graduate student Alex Pape to the design of the mini-CCT that was essentially built in the PSML, and the help and invaluable advice and support that we received from Rick Jones and Ed Suhler of the Department of Systems and Information Engineering.

References Amazon Web Services “Amazon Elastic MapReduce (Amazon EMR)” Available at:

http://aws.amazon.com/elasticmapreduce/ A. Bessani, M. Correia, B. Quaresma, F. André and P. Sousa, "DepSky: Dependable and

secure storage in a cloud-of-clouds," in Proceedings of the Sixth Conference on Computer Systems, 2011, pp. 31-46.

R. Bhadauria and S. Sanyal, "Survey on Security Issues in Cloud Computing and Associated Mitigation Techniques," ArXiv Preprint arXiv:1204.0764, 2012.

G. S. Bindra, P. K. Singh, K. K. Kandwal and S. Khanna, "Cloud security: Analysis and risk management of VM images," in Information and Automation (ICIA), 2012 International Conference on, 2012, pp. 646-651.

S. Bleikertz, M. Schunter, C. W. Probst, D. Pendarakis and K. Eriksson, "Security audits of multi-tier virtual infrastructures in public infrastructure clouds," in Proceedings of the 2010 ACM Workshop on Cloud Computing Security Workshop, 2010, pp. 93-102.

B. Brenner, IT Security Spending Up For Some. Security Leadership. January 07, 2009. Available at: http://www.csoonline.com/article/474390/it-security-spending-up-for- some.

Butler Consultants, Free Industry Statistics – Sorted by Highest Gross Margin. Available at: http://research.financial-projections.com/IndustryStats-GrossMargin.

Y. Chen, V. Paxson and R. H. Katz, "What’s new about cloud computing security," University of California, Berkeley Report no.UCB/EECS-2010-5 January, vol. 20, pp. 2010-2015, 2010.

CloudBus, Cloud Computing – A Powerful tool for cyberattack! May, 2011. Available at: http://cloudbus.blogspot.com/2011/05/cloud-computing-powerful-tool-for.html

  29  

B. Cromwell, Cyber Attacks Have Been Monetized. Perspectives on Security & Information Assurance – Industry tips and trends in Cybersecurity. Available at: http://cybersecurity.learningtree.com/2012/08/28/cyber-attacks-have-been-monetized/.

Defense Science Board, Department of Defense: “Task Force Report, Cyber Security and Reliability in a Digital cloud.” January 2013.

N. Dhanjani, B. Rios and B. Hardin, Hacking: The Next Generation. O'Reilly Media, Inc., 2009.

E. Dumbill, “Big data in the cloud” O’Reilly Strata blog post, February 22, 2012. http://strata.oreilly.com/2012/02/big-data-in-the-cloud-microsoft-amazon-google.html

O. Eggen, A. Nganwa and A. Suka, "As Strong as The Weakest Link," 2010. D. Gottfrid, "Self-Service, Prorated Supercomputing Fun!" Open blog post, New York

Times, November 1, 2007 http://open.blogs.nytimes.com/2007/11/01/self-service- prorated-super-computing-fun/

B. Grobauer, T. Walloschek and E. Stocker, "Understanding cloud computing vulnerabilities," Security & Privacy, IEEE, vol. 9, pp. 50-57, 2011.

Y. Y. Haimes, "Hierarchical holographic modeling," Systems, Man and Cybernetics, IEEE Transactions on, vol. 11, pp. 606-617, 1981.

Y. Y. Haimes, Risk Modeling, Assessment, and Management. John Wiley & Sons, 2011. Y. Y. Haimes, "Modeling complex systems of systems with phantom system models,"

Systems Engineering, vol. 15, pp. 333-346, 2012. Y. Y. Haimes, "Systems-Based Guiding Principles for Risk Modeling, Planning,

Assessment, Management, and Communication," Risk Analysis, vol. 32, pp. 1451- 1467, 2012.

Y. Y. Haimes and C. C. Chittister, "Risk to cyberinfrastructure systems served by cloud computing technology as systems of systems," Systems Engineering, vol. 15, pp. 213- 224, 2012.

P. Hayati, botCloud – an emerging platform for cyber-attacks. Cyber Warfare Intelligence. November 2, 2012. Available at: http://0xicf.wordpress.com/2012/11/02/botcloud-an-emerging-platform-for-cyber- attacks/.

Homeland Security News Wire. Hackers using cloud networks to launch powerful attacks. June 3, 2011. Available at: http://www.homelandsecuritynewswire.com/hackers-using-cloud-networks-launch- powerful-attacks.

J. N. Hoover, “NSA's Big Data Platform Faces Enterprise Test,” Information Week Government, October 11, 2012 http://www.informationweek.com/government/enterprise-applications/nsas-big-data- platform-faces-enterprise/240008916

IBM, “InfoSphere BigInsights” http://www01.ibm.com/software/data/infosphere/biginsights/

A. Janc and L. Olejnik, "Web browser history detection as a real-world privacy threat," in Computer Security–ESORICS 2010Anonymous Springer, 2010, pp. 215-231.

R. C. Johnson, 3 Reasons Clouds Prevent Cyber-Attacks. Smarter Technology. December 13, 2010. Available at: http://www.smartertechnology.com/c/a/Smarter-Strategies/3- Reasons-Clouds-Prevent-CyberAttacks/.

  30  

S. Kaplan and B. J. Garrick, "On the quantitative definition of risk," Risk Analysis, vol. 1, pp. 11-27, 1981.

K. Kerk, J. Penn, C. Atwood and J. Albornoz. The State of Information Security Spending. Forrester Research Inc, 2006.

R. H. Khan, "Decentralized Authentication in OpenStack Nova: Integration of OpenID," 2011.

J. Kim, H. Jeong, I. Cho, S. M. Kang and J. H. Park, "A secure smart-work service model based OpenStack for Cloud computing," Cluster Computing, pp. 1-12, 2013.

L. Leung, Cloud Customers Report Capital Cost Savings. Data Center Knowledge. January 26, 2010. Available at: http://www.datacenterknowledge.com/archives/2010/01/26/cloud-customers-report- capital-cost-savings/

S. Luo, Z. Lin, X. Chen, Z. Yang and J. Chen, "Virtualization security for cloud computing service," in Cloud and Service Computing (CSC), 2011 International Conference on, 2011, pp. 174-179.

D. McKay, A Deep Dive Into Hyperjacking. Security Week. February 3, 2011. Available at: http://www.securityweek.com/deep-dive-hyperjacking.

P. Mell and T. Grance, "The NIST definition of cloud computing (draft)," NIST Special Publication, vol. 800, pp. 7, 2011.

National Institute of Standards and Technologies (NIST), The NIST Definition of Cloud Computing, Special Publication 800-145, 2011. Gaithersburg, MD: P. Mell and T. Grance. Available at: http://csrc.nist.gov/publications/nistpubs/800-145/SP800- 145.pdf Accessed 10-26-11.

S. J. Nirmala, S. M. S. Bhanu and A. A. Patel, "A Comparative study of the secret sharing algorithms for secure data in the cloud," A Comparative Study of the Secret Sharing Algorithms for Secure Data in the Cloud, 2 (4), 63-71., 2012.

OpenStack Open Source Cloud Computing Software. [Online]. Available: http://www.openstack.org/software/. [Accessed: 14-Apr-2013].

Oozie Homepage http://oozie.apache.org/ Ponemon Institute, Encryption in the Cloud. July 2012. http://www.thales-

esecurity.com/knowledge-base/gated-content?Article=%7Bec38ee6a-8130-407b- 8430-6fe45db372af%7D&cid=701D0000000aMMu

Ponemon Institute, Security of Cloud Computing Providers Study. April 2011. http://www.ca.com/~/media/Files/IndustryResearch/security-of-cloud-computing- providers-final-april-2011.pdf

Ponemon Institute, Security of Cloud Computing Users. May 2010. http://www.ca.com/~/media/Files/IndustryResearch/security-cloud-computing- users_235659.pdf

PriceWaterhouseCoopers, Why isn’t IT spending creating more value? How to start a new cycle of value creation. June 2008.

T. Ristenpart, E. Tromer, H. Shacham and S. Savage, "Hey, you, get off of my cloud: Exploring information leakage in third-party compute clouds," in Proceedings of the 16th ACM Conference on Computer and Communications Security, 2009, pp. 199- 212.

  31  

F. Rocha and M. Correia, "Lucy in the sky without diamonds: Stealing confidential data in the cloud," in Dependable Systems and Networks Workshops (DSN-W), 2011 IEEE/IFIP 41st International Conference on, 2011, pp. 129-134.

R. Roman, M. R. Felipe, P. E. Gene and J. Zhou, "Analysis of security components in cloud storage systems," in APMRC, 2012 Digest, 2012, pp. 1-4.

D. Sarno, and S. Rodriguez. Hacker attacks show vulnerability of cloud computing. Los Angeles Times. June 17, 2011. Available at: http://articles.latimes.com/2011/jun/17/business/la-fi-cloud-security-20110617.

M. Schwartz, Security Spending Grabs Greater Share of IT Budgets. InformationWeek Security. February 15, 2011.

The H Security, Cracking WPA keys in the cloud. January 12, 2011. Available at: http://www.h-online.com/security/news/item/Cracking-WPA-keys-in-the-cloud- 1168636.html

I. Seminar and P. Cigoj, "Security Aspects of OpenStack," . R. Slipetskyy, Security Issues in OpenStack, 2011. U. S. Defense Science Board (2013). Cyber Security and Reliability in a Digital

Cloud. Task Force Report Office of the Under Secretary of Defense for Acquisition, Technology, and Logistics Washington, D.C. http://www.acq.osd.mil/dsb/reports/CyberCloud.pdf

U.S. Nuclear Regulatory Commission, Fault tree handbook (NUREG-0492). Washington, D.C.: U.S. Government Printing Office, 1981

L. M. Vaquero, L. Rodero-Merino and D. Morán, "Locking the sky: a survey on IaaS cloud security," Computing, vol. 91, pp. 93-118, 2011.

VMware, “Apache Hadoop on vSphere” http://www.vmware.com/hadoop J. Wei, X. Zhang, G. Ammons, V. Bala and P. Ning, "Managing security of virtual

machine images in a cloud environment," in Proceedings of the 2009 ACM Workshop on Cloud Computing Security, 2009, pp. 91-96.

Z. Weinberg, E. Y. Chen, P. R. Jayaraman and C. Jackson, "I still know what you visited last summer: Leaking browsing history via user interaction and side channel attacks," in Security and Privacy (SP), 2011 IEEE Symposium on, 2011, pp. 147-161.

T. Wright, How much can I save with Cloud Computing? Hazaa. January 27, 2012. Available at: http://hazaa.com.au/sharepoint-consultant/blog/how-much-can-i-save- with-cloud-computing/.

Yahoo Finance. Industry Index. Available at: http://biz.yahoo.com/p/sum_peed.html