Case Scenario 3: The Security Plan

profilesecure2026
Chapter10Reading.docx

10

Management and Incidents

In this chapter:

· Security planning

· Incident response and business continuity planning

· Risk analysis

· Handling natural and human-caused disasters

In this chapter we introduce concepts of managing security. Many readers of this book are, or will be, practitioners or technologists: people who design, implement, and use security. Security devices, algorithms, architectures, protocols, and mechanisms are important for those readers to consider when making decisions about managing systems and services.

Some technologists think security involves just designing a stronger (faster, better, bigger) appliance or selecting the best cryptographic algorithm. That these are important considerations is true. But what if you build something that nobody adopts? Perhaps the user interface is inscrutable. Or people cannot figure out how to integrate your product into any existing system. Maybe it doesn’t really address the underlying security problem. Perhaps it is too restrictive, preventing users from getting real work done. And maybe it is or seems too expensive. Technology—even the best product—has to be used and usable.

In this chapter we consider two important topics: how security is managed and how it is used. These topics relate to human behavior and also to the business side of computing. We also address the physical side of security threats: natural disasters and those caused by humans.

10.1 Security Planning

Years ago, when most computing was done on mainframe computers, data processing centers protected both devices and data. Responsibility for security rested neither with the programmers nor the users but instead with the computing center staff itself. These centers developed expertise in security, and they implemented many forms of protection in the background, without users having to be conscious of protection needs and practices.

But beginning as far back as the 1980s, the introduction of personal computers and the general ubiquity of computing changed the way many of us work and interact with computers. In particular, significant responsibility for security has shifted to the user and away from the computing center. Alas, many users are unaware of (or choose to ignore) this responsibility, so they neither deal with the risks posed nor implement simple measures to prevent or mitigate problems.

You have probably seen common examples of this neglect in news stories: updates not applied, virus scans not performed, backups not created, for example. Moreover, neglect is exacerbated by the seemingly hidden nature of important data: Things we would protect if they were on paper, we ignore when they are stored electronically. For example, a person who carefully locks up paper copies of company confidential records overnight may leave running a personal computer on a desk. We access sensitive data from laptops, smartphones, and tablets, which we then leave on tables and chairs in restaurants, airports, bars, or coffee shops. In this situation, a curious or malicious person walking past can look at or even copy confidential memoranda and data. Similarly, the data on today’s devices are often more easily available than on older, more isolated systems. For instance, the large and cumbersome disk packs and tapes from years ago have been replaced by media such as flash drives that hold a huge volume of data but fit easily in a pocket or briefcase. Moreover, we all recognize that a single memory stick may contain many times more data than a printed report. But since the report is apparent and the stick is not, we leave the computer medium in plain view, easy to borrow or steal.

In all cases, whether the user initiates some computing action or simply interacts with an active application, every application has confidentiality, integrity, and availability requirements that relate to data, programs, and computing machinery. In these situations, users suffer from lack of sensitivity: They often do not appreciate the security risks associated with using computers.

And, as we have seen throughout this book, unprotected smart devices also offer opportunities to data thieves: purloined images from a doorbell or security camera, location history from an automobile’s route-mapping system, or stolen music files from an entertainment system.

For these reasons, every organization using computers to create and store valuable assets should perform thorough and effective security planning. A  security plan is a document that describes how an organization will address its security needs. The plan is subject to periodic review and revision as the organization’s security needs change. Even you and your family may want to create a security plan, to help you know how and when to protect your systems and data. In the following sections we focus primarily on organizational security, but most of the concepts and actions apply to personal computing too.

Both personal and enterprise security starts with a security plan that describes how to address security needs.

Organizations and Security Plans

Consider a simple example: You have several things you need to do in the next few days. You can keep them in your head, write them down on paper, or store them in an electronic device. In your head, it is easy to forget some or to focus on a less important activity. Writing matters down, whether on paper or with electrons, encourages you to think for a moment of other things you need to do. And a recorded list to which you can refer gives you a structure to remind you of the important items or help you choose something you can complete if you have a few free minutes. So it is best to record your security plan rather than try to remember its steps as you are performing other tasks.

A good security plan is an official record of current security practices, as well as a blueprint for orderly change to improve those practices. By following the plan, developers and users can measure the effect of proposed changes, leading eventually to further improvements. The impact of the security plan is important too. A carefully written plan, supported by management, notifies employees that security is important to everyone. Thus, the security plan must have appropriate content likely to produce desired effects.

In this section we study how to create and implement a security plan. We focus on three aspects of writing a security plan: what it should contain, who writes it, and how to obtain support for it. Then we address two specific cases of security plans: business continuity plans, to ensure that an organization continues to function despite a computer security incident, and incident response plans, to organize activities when a crisis occurs.

Contents of a Security Plan

A security plan identifies and organizes the actions needed to provide and protect security for a computing system and its users. The plan is both a description of the current situation and a map for improvement. Every security plan must address seven issues:

· policy, indicating the goals of a computer security effort and the willingness of the people involved to work to achieve those goals

· current state, describing the status of security at the time of the plan

· security requirements, describing security goals in terms of permissible and impermissible system actions

· recommended controls, mapping controls to the vulnerabilities identified in the policy, in the context of the security requirements

· accountability, documenting who is responsible for each security activity and outcome

· timetable, identifying when different security functions are to be done

· maintenance, specifying a structure for periodically updating the security plan

There are many approaches to creating and updating a security plan. Some organizations have a formal, defined security-planning process, much as they might have a defined and accepted development or software maintenance process. Others look to security professionals for guidance on how to perform security planning. But every security plan contains the same basic material, no matter the format. The following sections expand on the seven parts of a security plan.

Policy

A security policy documents an organization’s security needs and priorities.

A security policy is a high-level statement of purpose and intent that lays out the organization’s security needs and priorities. Initially, you might think that all policies would be the same: to prevent security breaches. But in fact the policy is one of the most difficult sections to write well because it reflects an analysis of what to do and what not to do in one particular setting.

Consider security needs for different types of organizations. What does an organization consider its most precious asset? A pharmaceutical company might value its scientific research on new drugs and its sales and marketing strategy as its most important assets. A hospital would likely find protecting the confidentiality of its patients’ records most crucial, perhaps followed closely by the integrity and availability of its systems that implement patient care (such as centralized monitors of vital signs). A television studio could decide its archive of previous broadcasts is most important. An online merchant might value highly its web presence and the associated back-end system for receiving and processing orders. A securities trading firm would be most concerned with the accuracy and completeness of its transaction records, including its log of executed trades. As you can see, organizations value different things, and following this analysis, the most significant threats will differ among organizations. In some cases confidentiality is paramount, but in others availability or integrity matters most.

In addition, an enterprise must decide for which security issues it will be directly responsible, and for which ones it will rely on insurance, third-party software, or security subcontractors. Sometimes it is cheaper and more effective to pay an insurance premium, especially when the security risk is low but the impact of a breach is significant. These are business decisions more than security decisions, but they must be considered in the security plan.

Later in this chapter we discuss these tradeoffs among the strength of the security, the cost, the impact on users, and more. And these aspects of security must also address timing: We must decide whether to implement very stringent—and possibly unpopular—controls to prevent security problems or simply mitigate the problem’s effects once they happen. For this reason, the policy statement must answer three essential questions about each type of security:

· Who should be allowed access?

· To what system and organizational  resources should access be allowed?

· What  types of access should each user be allowed for each resource ?

A security policy statement should specify the following:

· The organization’s  goals for security. For example, should the system protect data from leakage to outsiders, protect against loss of data due to physical disaster, protect the data’s integrity, or protect against loss of business when computing resources fail? What is the higher priority: serving customers or securing data?

· Where the  responsibility for security lies. For example, should the responsibility rest with a small computer security group, with each employee, or with relevant managers?

· The organization’s  commitment to security. For example, who provides security support for staff, and where does security fit into the organization’s structure?

Assessment of Current Security Status

To be able to plan for security, an organization must understand the vulnerabilities to which it may be exposed. The organization can determine the vulnerabilities by performing a  risk analysis: a systematic investigation of the system, its environment, and the things that might go wrong. The risk analysis forms the basis for describing the current status of security. This status can be expressed as a listing of organizational assets, the security threats to the assets, and the controls in place to protect the assets. We look at risk analysis in more detail later in this chapter.

The status portion of the plan also defines the limits of responsibility for security. It describes not only which assets are to be protected but also who is responsible for protecting them. The plan may note that some groups can be excluded from responsibility; for example, joint ventures with other organizations may designate one organization to provide security for all member organizations. The plan also defines the boundaries of responsibility, especially when networks are involved. For instance, the plan should clarify who provides the security for a network router, for a direct line to a remote site, or for data storage or processing in a cloud.

Even though the security plan should be thorough, there will necessarily be vulnerabilities that are not considered. These vulnerabilities are not always the result of ignorance or naïveté; rather, they can arise from the addition of new equipment or data as the system evolves. They can also result from new situations, such as when a system is used in ways not anticipated by its designers. The security plan should detail the process to be followed when someone identifies a new vulnerability. In particular, instructions should explain how to integrate controls for that vulnerability into the existing security procedures.

Security Requirements

The heart of the security plan is its set of  requirements: functional or performance demands placed on a system to ensure a desired level of security. The requirements are usually derived from organizational needs. Sometimes these needs include the need to conform to specific security mandates imposed from outside, such as by a government agency or a commercial standard.

Security requirements document organizational and external demands.

Shari Lawrence Pfleeger [ PFL91 ] points out that we must distinguish requirements from constraints and controls. A  constraint is an aspect of the security policy that constrains, circumscribes, or directs the implementation of the requirements. As defined in  Chapter 1 , a  control is an action, device, procedure, or technique that removes or reduces a vulnerability. To see the difference between requirements, constraints, and controls, consider the six “requirements” of the U.S. Department of Defense’s TCSEC, introduced in  Chapter 5 . These six items are listed in  Table 10-1 .

Table 10-1 The Six “Requirements” of the TCSEC

Security policy

There must be an explicit and well-defined security policy enforced by the system.

Identification

Every subject must be uniquely and convincingly identified. Identification is necessary so that subject/object access can be checked.

Marking

Every object must be associated with a label that indicates its security level. The association must be done so that the label is available for comparison each time an access to the object is requested.

Accountability

The system must maintain complete, secure records of actions that affect security. Such actions include introducing new users to the system, assigning or changing the security level of a subject or an object, and denying access attempts.

Assurance

The computing system must contain mechanisms that enforce security, and it must be possible to evaluate the effectiveness of these mechanisms.

Continuous protection

The mechanisms that implement security must be protected against unauthorized change.

Given our definitions of requirement, constraint, and control, you can see that the first “requirement” of the TCSEC is really a constraint: the security policy. The second and third “requirements” describe mechanisms for enforcing security, not descriptions of required behaviors. That is, the second and third “requirements” describe explicit implementations, not a general characteristic or property that the system must have. However, the fourth, fifth, and sixth TCSEC “requirements” are indeed true requirements. They state that the system must have certain characteristics, but they do not mandate a particular implementation.

These distinctions are important because the requirements explain  what should be accomplished, not  how. That is, the requirements should always leave the implementation approach to the designers, whenever possible. For example, rather than writing a requirement that certain data records should require passwords for access (an implementation decision), a security planner should state only that access to the data records should be restricted (and note to what categories of users the access should be restricted).

The requirement might also indicate strength, for example, preventing access by casual attempts (lightly restrictive) or protecting against concerted effort over weeks (highly protective). This more flexible requirement allows the designers to select among several controls (such as tokens or encryption) and to balance security requirements with other system requirements, such as performance and reliability.  Figure 10-1  illustrates how different aspects of system analysis support the security planning process.

An illustration shows the Inputs to the Security Plan.

FIGURE 10-1 Inputs to the Security Plan

As with the general software development process, security planning must allow customers or users to specify desired functions, independent of implementation. The requirements should address all aspects of security: confidentiality, integrity, and availability. They should also be reviewed to make sure they are of appropriate strength and quality. In particular, we should ensure that the requirements have these characteristics:

· Correctness. Are the requirements understandable? Are they stated without error?

· Consistency. Are there any conflicting or ambiguous requirements?

· Completeness. Are all possible situations addressed by the requirements?

· Realism. Is it possible to implement what the requirements mandate?

· Need. Are all requirements necessary; are they unnecessarily restrictive?

· Verifiability. Can tests be written to demonstrate conclusively and objectively that the requirements have been met? Can the system or its functionality be measured in some way that will assess the degree to which the requirements are met?

· Traceability. Can each requirement be traced to the functions and data related to it so that changes in a requirement can lead to easy reevaluation?

The requirements may then be constrained by budget, schedule, performance, policies, governmental regulations, and more. Given the requirements and constraints, developers then choose appropriate controls.

Recommended Controls

Security requirements lay out the system’s needs in terms of what should be protected. The security plan must also recommend what controls should be incorporated into the system to meet those requirements. Throughout this book you have seen many examples of controls, so we need not review them here. As we discuss later in this chapter, we can use risk analysis to create a map from vulnerabilities to controls. The mapping tells us how the system will meet the security requirements. That is, the recommended controls address implementation issues: how the system will be designed and developed to meet stated security requirements.

Accountability for Implementation

A security plan documents who is responsible for implementing security. No one responsible implies no action.

A section of the security plan will identify which people (usually listed as organizational titles, such as Head of Human Relations or the Network Security Administrator on duty) are responsible for implementing the security requirements. Note that describing responsibility by title not name is sensible: If Raesha is named, who becomes responsible after she leaves the organization or, worse, transfers to a position with entirely different duties? This documentation assists those who must coordinate their individual responsibilities with those of other developers. At the same time, the plan makes explicit who is accountable should some requirement not be met or some vulnerability not be addressed. That is, the plan notes who is responsible for implementing controls when a new vulnerability is discovered or a new kind of asset is introduced. (But see  Sidebar 10-1  on who is responsible.)

Sidebar 10-1 Who Is Responsible for Using Security?

We put a lot of responsibility on the user: Apply these patches, don’t download unknown code, keep sensitive material private, change your password frequently, don’t forget your umbrella. We are all fairly technology savvy, so we take in stride messages like “fatal error.” A neighbor once called in a panic, fearing that her entire machine and all its software and data were about to go up in a puff of electronic smoke because she had received a “fatal error” message; we explained calmly that the message was perhaps a bit melodramatic.

But that neighbor raises an important point: How can we expect people to use their computers securely when the terminology is confusing and the security steps so hard to do? Take, for example, the various steps necessary in securing a wireless access point (see  Chapter 6 ): Use WPA or WPA2, not WEP; set the access point to nonbroadcast mode, not open; choose a random 128-bit number for an initial value; and so forth. Whitten and Tygar [ WHI99 ] list four points critical to users’ security:

· Users must be aware of the security of tasks they need to perform.

· Users must be able to figure out how to perform those tasks successfully.

· Users must be prevented from making dangerous errors.

· Users must be sufficiently comfortable with the technology to continue using it.

PGP secure email (described in  Chapter 4 ) has a fairly good user interface. However, Whitten and Tygar conclude that the PGP product is not usable enough to provide effective security for most computer users. Furnell [ FUR05 ] reached a similar conclusion about the security features in Microsoft Word.

The field of human–computer interaction (HCI) is mature, guidance materials are available, and numerous good examples exist. Why, then, are security settings hidden on a sub-sub-tab and written in highly technical jargon? We cannot expect users to participate in security enforcement unless they can understand what they should do.

Ben Shneiderman counsels that the human–computer interface should be fun. Citing work others have done on computer game interfaces, Shneiderman notes that such interfaces satisfy needs for challenge, curiosity, and fantasy. He then argues that computer use must “(1) provide the right functions so that users can accomplish their goals, (2) offer usability plus reliability to prevent frustration from undermining the fun, and (3) engage users with fun-features” [ SHN04 ].

One can counter that security functionality is serious, unlike computer games or web browsers. Still, this does not relieve us from the need to make the interface consistent, informative, empowering, and error preventing. Ideally, users should want to make their systems more secure rather than be intimidated by security activities. Although we do not go so far as to expect the interface to be “fun” to use, it should be clear, straightforward, and succinct.

People building, using, and maintaining the system play many roles. Each role can take some responsibility for one or more aspects of security. Consider, for example, the groups listed below.

· Users of technologymay be responsible for the security of their own devices. Alternatively, the security plan may designate one person or group to be coordinator of personal computer security.

· Project leaders may be responsible for the security of data and computations.

· Managers may be responsible for seeing that the people they supervise implement security measures.

· Database administrators may be responsible for the access to and integrity of data in their databases.

· Information officers may be responsible for overseeing the creation and use of data; these officers may also be responsible for retention and proper disposal of data.

· Personnel staff members may be responsible for security involving employees, for example, screening potential employees for trustworthiness and arranging security training programs.

Timetable

A comprehensive security plan cannot be executed instantly. The security plan includes a timetable that shows how and when the elements of the plan will be performed. These dates also set milestones so that management can track the progress of implementation.

It may be desirable to implement the security practices over time rather than all at once. For example, if the controls are expensive, numerous, or complicated, they may be acquired and implemented gradually. Similarly, procedural controls may require staff training to ensure that everyone understands and accepts the reason for the control. The plan should specify the order in which the controls are to be implemented so that the most serious exposures are covered as soon as possible.

Furthermore, the plan must be extensible. Conditions will change: New equipment will be acquired, new degrees and modes of connectivity will be requested, and new threats will be identified. The plan must include a procedure for change and growth, so that the security aspects of changes are considered as a part of preparing for the change, not for adding security after the change has been made. The plan should also contain a schedule for periodic review. Even though there may have been no obvious, major growth, most organizations experience modest change every day. At some point the cumulative impact of the change is enough to require that the plan be modified.

Plan Maintenance

Security plans must be revisited periodically to adapt them to changing conditions.

Good intentions are not enough when it comes to security. We must not only take care in defining requirements and controls, but we must also find ways for evaluating a system’s security to be sure that the system is as secure as we intend it to be. Thus, the security plan must call for reviewing the security situation periodically. As users, data, uses, and equipment change, new exposures may develop. In addition, the current means of control may become obsolete or ineffective (such as when faster processor times enable attackers to break an encryption algorithm). The inventory of objects and the list of controls should periodically be scrutinized and updated, and risk analysis performed anew. The security plan should set times for these periodic reviews, based either on calendar time (such as, review the plan every nine months) or on the nature of system changes (such as, review the plan after every major system release).

Security Planning Team Members

Who performs the security analysis, recommends a security program, and writes the security plan? As with any such comprehensive task, these activities are likely to be performed by a committee that represents all the interests involved. The size of the committee depends on the size and complexity of the computing organization and the degree of its commitment to security. Organizational behavior studies suggest that the optimum size for a working committee is between five and nine members. Sometimes a larger committee may serve as an oversight body to review and comment on the products of a smaller working committee. Alternatively, a large committee might designate subcommittees to develop sections of the plan.

The membership of a computer security planning team must somehow relate to the different aspects of computer security described in this book. Security in operating systems and networks requires the cooperation of the systems administration staff. Program security measures can be understood and recommended by applications programmers. Physical security controls are implemented by those responsible for general physical security, both against human attacks and natural disasters. Because the plan will affect employees and their working conditions, human resources staff should be included. And to help in evaluating the costs of threats or countermeasures, someone from the finance group should participate. Finally, because controls affect system users, the plan should incorporate users’ views, especially with regard to usability and the general desirability of controls.

We have phrased this discussion in terms of a company, which might be a hospital, bank, or manufacturing plant. However, the security planning process is equally important for schools and universities, government agencies, nonprofit organizations, and other groups responsible for sensitive data and computing.

Thus, no matter how it is organized, a security planning team should represent each of the following groups:

· computer and network hardware group

· system administrators

· systems programmers

· applications programmers

· data entry personnel

· physical security personnel

· representative users

In some cases, a group can be adequately represented by someone who is consulted at appropriate times, rather than a committee member from each possible constituency being enlisted.

Assuring Commitment to a Security Plan

After the plan is written, it must be accepted and its recommendations enacted. Acceptance by the organization is key; a plan that has no organizational commitment is simply a plan that collects dust on the shelf. Commitment to the plan means that security functions will be implemented and security activities carried out. Three groups of people must contribute to making the plan a success.

· The planning team must be sensitive to the needs of each group affected by the plan.

· Those affected by the security recommendations must understand what the plan means for the way they will use the system and perform their business activities. In particular, they must see how their actions can affect other users and other systems.

· Management must be committed to using and enforcing the security aspects of the system.

Education and publicity can help people understand and accept a security plan. Acceptance involves not only the letter but also the spirit of the security controls. There is a story of an employee who went through 24 password changes at a time to get back to a favorite password, in a system that prevented use of any of the 23 most recently used passwords. Clearly, the employee either did not understand or did not agree with the reason for restrictions on password selection. If people understand the need for recommended controls and accept them as sensible, they will use the controls properly and effectively. If people think the controls are bothersome, capricious, or counterproductive, they will work to avoid or subvert them.

Management commitment is obtained through understanding. But this understanding is not just a function of what makes sense technologically; it also involves knowing the cause and the potential effects of lack of security. Managers must also weigh tradeoffs in terms of convenience and cost. The plan must present a picture of how cost effective the controls are, especially when compared to potential losses if security is breached without the controls. Thus, proper presentation of the plan is essential, in terms that relate to management as well as technical concerns.

A security plan positions technical issues in terms nontechnical people can appreciate.

Remember that some managers are not computing specialists. Instead, the system supports a manager who is an expert in some other business function, such as banking, medical technology, or sports. In such cases, the security plan must present security risks in language the managers understand. A useful security plan should avoid technical jargon and educate the readers about the nature of the perceived security risks in the context of the business the system supports. Sometimes outside experts can bridge the gap between the managers’ business and security.

Management is often reticent to allocate funds for controls before understanding the value of those controls. As we note later in this chapter, the results of a risk analysis can help communicate the financial tradeoffs and benefits of implementing controls. By describing vulnerabilities in financial terms and in the context of ordinary business activities (such as leaking data to a competitor or an outsider), security planners can help managers understand the need for controls.

The plans we have just discussed are part of normal business. They address how a business handles computer security needs. Similar plans might address how to increase sales or improve product quality, so these planning activities should be a natural part of management.

Next we turn to two particular kinds of business plans that address specific security problems: coping with and controlling activity during security incidents and ensuring that business activity continues in spite of an incident.

10.2 Business Continuity Planning

Small companies working with a low profit margin might be put out of business by even a moderate computer incident. By contrast, large, financially sound businesses can usually weather a modest (but still painful) incident that interrupts their use of computers for a while. But even rich companies do not want to spend money unnecessarily. The analysis supporting investment in security is sometimes as simple as:  No computers means no customers means no sales means no profit.

Government agencies, educational institutions, and nonprofit organizations also have limited budgets, which they want to use to address their mission. They may not have a direct profit motive, but being able to meet the needs of their customers—the public, students, and constituents—partially determines how well they will fare in the future. All kinds of organizations must plan for ways to cope with emergency situations.

business continuity plan 1  documents how a business will continue to function during or after a computer security incident. An ordinary security plan covers computer security during normal times and deals with protecting against a wide range of vulnerabilities from the usual sources. A business continuity plan deals with situations having two characteristics:

1.  The standard terminology is “business continuity plan,” even though such a plan is needed by and applies to a university’s “business” of educating students or a government’s “business” of serving the public.

· catastrophic situations, in which all or a major part of a computing capability is suddenly unavailable

· long duration, in which the outage is expected to last for so long that business will suffer

Business continuity planning guides response to a crisis that threatens a business’s existence.

A business continuity plan would be helpful in many situations. Here are some examples that typify what you might find in reading your daily newspaper:

· A fire destroys a company’s entire network.

· A seemingly permanent failure of a critical software component renders the computing system unusable.

· The abrupt failure of a supplier of electricity, telecommunications, network access, or other critical service limits or stops activity.

· A flood prevents the essential network support staff from getting to the operations center.

As you can see, the impact in each example is likely to continue for a long time, and each disables a vital function.

You may also have noticed how often “the computer” is blamed for an inability to provide a service or product. For instance, the clerk in a shop is unable to use the cash register because “the computer is down.” You may have an item in your hand, plus exactly the cash to pay for it. But the clerk will not take your money and send you on your way. Often, computer service is restored shortly. But sometimes it is not. Once we were delayed for over an hour in an airport because of an electrical storm that caused a power failure and disabled the airlines’ computers. Although our tickets showed clearly our reservations on a particular flight, the airline agents refused to let anyone board because they could not assign seats electronically. As the computer remained down, the agents were frantic 2  because the technology was delaying the flight and, more important, disrupting hundreds of connections.

2 . The obvious, at least to us, idea of telling passengers to “sit in any seat” seemed to be against airline policy. And this incident was long before the 9/11 terrorist attacks tightened airline security.

The key to coping with such disasters is advanced planning and preparation, identifying activities that will keep a business viable when the computing technology is disabled. The steps in business continuity planning are these:

· Assess the likely business impact of a crisis.

· Develop a strategy to control impact.

· Develop and implement a plan for the strategy

Assess Business Impact

To assess the impact of a failure on your business, you begin by asking two key questions:

· What are the  essential assets? What are the things that if lost will prevent the business from functioning? Answers are typically of the form “the network,” “the customer reservations database,” or “the system controlling traffic lights.”

· What could  disrupt use of these assets? The vulnerability is more important than the threat agent. For example, whether destroyed by a fire or zapped in an electrical storm, the network is nevertheless down. Answers might be “failure,” “corrupted,” or “loss of power.”

You probably will find only a handful of key assets when doing this analysis.

Do not overlook people and the things they need for support, such as documentation and communications equipment. Another way to think about your assets is to ask yourself, “What is the minimum set of things or activities needed to keep business operational, at least to some degree?” If a manual system could compensate for a failed computer system, albeit inefficiently, you may want to consider building such a manual system as a potential critical asset. Think of the airline unable to assign seats from a chart of the cabin.

Later in this chapter we study risk analysis, a comprehensive way to examine assets, vulnerabilities, and controls. For business continuity planning we do not need a full risk analysis. Instead, we focus on only those things that are critical to continued operation. We also look at larger classes of objects, such as “the network,” whose loss or compromise can have catastrophic effect.

Develop Strategy

The continuity strategy investigates how the key assets can be safeguarded. In some cases, a backup copy of data or redundant hardware or an alternative manual process is good enough. Sometimes, the most reasonable answer is reduced capacity. For example, a planner might conclude that if the call center in London fails, the business can divert all calls to Tokyo. Perhaps the staff in Tokyo cannot handle the full load of the London traffic; this situation may result in irritated or even lost customers, but at least some business can be transacted.

Ideally, you would like to continue business with no loss. But with catastrophic failures, usually only a portion of the business function can be preserved. In this case, you must develop a strategy appropriate for your business and customers. For instance, you can decide whether it is better to preserve half of function A and half of B, or most of A and none of B.

Business continuity planning forces a company to set base priorities.

You also must consider the time frame in which business is done. Some catastrophes last longer than others. For example, rebuilding after a fire is a long process and implies a long time in disaster mode. Your strategy may have several steps, each dependent on how long the business is disabled. Thus, you may take one action in response to a one-hour outage and another if the outage might last a day or longer.

Because you are planning in advance, you have the luxury of being able to think about possible circumstances and evaluate alternatives. For instance, you may realize that if the Tokyo site takes on work for the disabled London site, there will be a significant difference in time zones. It may be better to divert morning calls to Tokyo and afternoon ones to Dallas, to avoid asking Tokyo staff to work extra hours.

The result of a strategy analysis is a selection of the best actions, organized by circumstances. The strategy can then be used as the basis for your business continuity plan.

Develop the Plan

The business continuity plan specifies several important things:

· who is in charge when an incident occurs

· what to do

· who does it

The plan justifies making advance arrangements, such as acquiring redundant equipment, arranging for data backups, and stockpiling supplies, before the catastrophe. The plan also justifies advance training so that people know how they should react. A catastrophe will cause confusion; you do not want to add confused people to the already severe problem.

The person in charge declares the state of emergency and instructs people to follow the procedures documented in the plan, elaborating and improvising as necessary. The person in charge also declares when the emergency is over and conditions can revert to normal.

Seldom will the plan tell precise steps to take in a crisis because the nature of crises is too varied. Even in broad categories (such as, something causes the network to fail), the nature of “failure” and the prospects for recovery (one hour, one day, one week) are so imprecise that no plan can dictate what to do in each situation. Instead, the person in charge has latitude to take action that seems best at the time. The point is, one person is in charge and is authorized to organize action and spend money necessary to recover at least partially.

Thus, the business continuity planning addresses how to maintain some degree of critical business activity in spite of a catastrophe. Its focus is on keeping the business viable. It is based on the asset survey, which focuses on only a few critical assets and serious vulnerabilities that could threaten operation for a long or undetermined period of time.

A business continuity plan focuses on business needs.

The focus of the business continuity plan is to keep the business going while someone else addresses the crisis (for example, replacing faulty equipment). That is, the business continuity plan does not include calling the fire department or evacuating the building, important though those steps are. The focus of a business continuity plan is the  business and how to keep it functioning to the degree possible in the situation. Handling the emergency is someone else’s problem.

Now we turn to a different plan that deals specifically with computer crises.

10.3 Handling Incidents

The network grinds almost to a halt. A pop-up window advises you to patch an application immediately. A file disappears. An unusual name appears on the list of active processes. Are any of these situations normal? A concern? Something to report, and if yes, to whom? Any one of these situations could be a first sign of a security incident, or nothing at all. What should you do?

Individuals must take responsibility for their own environments. But students in a university or employees of a company or government agency sometimes assume it is someone else’s responsibility. Or they don’t want to bother a busy operations staff with something that may be nothing at all.

Organizations develop a capability to handle incidents from receiving the first report and investigating it. In this section we consider incident handling practices.

Incident Response Plans

An incident response plan details how to address security incidents of all types.

A (security)  incident response plan tells the staff how to deal with a security incident. In contrast to the business continuity plan, the goal of incident response is handling the current security incident, without direct regard for the business issues. The security incident may at the same time be a business catastrophe, as addressed by the business continuity plan. But as a specific security event, it might be less than catastrophic (that is, it may not severely interrupt business) but could be a serious breach of security, such as a hacker attack or a case of internal fraud. As we will see later in this chapter, an incident could be a single event, a series of events, or an ongoing problem.

An incident response plan should

· define what constitutes an  incident

· identify who is responsible for  taking charge of the situation

· describe the plan of  action

The plan usually has three phases: advance planning, triage, and running the incident. A fourth phase, review, is useful after the situation abates so that this incident’s discovery and resolution can lead to improvement in handling future incidents.

Advance Planning

As with all planning functions, advance planning works best because people can think logically, unhurried, and without pressure or emotion. What constitutes an incident may be vague. We cannot know the details of an incident in advance. Typical characteristics include harm or risk of harm to computer systems, data, processing, or people; initial uncertainty as to the extent of damage; and similar uncertainty as to the source or method of the incident. For example, you can see that the file is missing or the home page has been defaced, but you do not know how or by whom or what other damage there may be.

In organizations that have not done incident planning, chaos may develop at this point. Someone runs to the network manager. Someone sends email to the help desk. Someone calls the FBI, the CERT, the newspapers, or the fire department. Someone posts the situation to internal or external social media. People start to investigate on their own, without coordinating with the relevant staff in other departments, agencies, or businesses. And conversation, rumor, and misinformation ensue, often generating more noise than substance.

An incident response plan tells whom to contact in the event of an incident, which may be just an unconfirmed, unusual situation.

With an incident response plan in place, everybody is trained in advance to contact the designated leader. The plan establishes a list of people to alert, in order, in case the first person is unavailable. The leader decides what to do next, beginning by determining whether this is a real incident or a false alarm. Indeed, natural events sometimes look like incidents, and the situation’s facts should be established first. If the leader decides this is an incident of concern, he or she invokes the response team.

Responding

The  response team is the set of people charged with responding to the incident. The response team may include

· director: the person in charge of the incident, who decides what actions to take and when to terminate the response. The director is typically a management employee.

· technician(s): the people who perform the response’s technical activities. The lead technician decides where to focus attention, analyzes situation data, documents the incident and how it was handled, and calls for other technical people to assist with the analysis.

· advisor(s): the legal, human resources, or public relations staff members as appropriate, who will determine who else (inside and outside the organization, including customers) should be kept apprised of the incident and its resolution.

For a small incident, a single person may be able to handle more than one of these roles. Nevertheless, this designated leader directs the response work, acts as a single point of contact for “insiders” (employees, users) as well as a single official representative for informing the public.

To develop policy and identify a response team, you need to address several aspects of both incident and response.

· Legal issues. An incident may have legal ramifications. In some countries, computer intrusions are illegal, so law enforcement officials must be involved in the investigation. In other places, you have discretion in deciding whether to ask law enforcement to participate. In addition to criminal action, you may be able to bring a civil case against the perpetrator(s) for time or money lost. Both kinds of legal action have serious implications for the response. For example, evidence must be gathered and maintained in specific ways to be usable in court. Similarly, laws may limit what you can do against the alleged attacker: Cutting off a connection is probably acceptable, but launching a retaliatory denial-of-service attack may not be.

· Preserving evidence. The most common reaction in an incident is to assume the cause was internal or accidental. For instance, you may first assume that hardware has failed or software isn’t working correctly or even that you did something wrong. Someone who seems to be knowledgeable may direct people to change the configuration, reload the software, reboot the system, or similarly attempt to resolve the problem by adjusting the software. Unfortunately, each of these acts can irreparably distort or destroy evidence. When dealing with a possible incident, do as little as possible before securing the site and capturing possible evidence.

· Records. It may be difficult to remember what you have already done: Have you already reloaded a particular file? What steps led you to the prompt asking for the new DNS server’s address? If you call in an outside forensic investigator or the police, you will need to explain exactly what you have already done. A list of what was done can also help people who need to determine what happened, how to prevent it in the future, and how to restore data and computing capabilities. Photographs of a succession of computer screens can help to show what steps you took and how the system responded.

· Public relations. In handling an incident, your organization should speak with one voice, to avoid the possibility of confusing or conflicting messages. And those messages should be vetted by the legal team before being released. Otherwise, an unguarded comment may tip off the attacker or have a negative effect on the case. The spokesperson can simply say that an incident occurred, tell briefly and generally what it was, and state that the situation is now under control and when normal operation is expected to resume. All staff members should know that only one person will provide details to the public. Other staff members must not post to social media, give details to a journalist, or chat with friends and relatives.

“Is this really an incident?” is the most important question.

Incident responders first perform triage: They investigate what has happened. “The network is responding slowly” can have many causes, from heavy usage to electronic malfunction to terrorist attack. Based on first analysis, the team decides what steps to take to address the incident.

Some incidents resolve themselves (for example, the heavy usage ends), some stay the same (the malfunction does not heal itself), and some get worse (the fire spreads). Incident responders follow the case until they have identified the cause and done as much as possible to return the system to normal. Then the team finishes documenting its work and declares the incident over.

After the Incident Is Resolved

Eventually, the incident response team closes the case. At this point it will hold a review after the incident to consider two things:

· Is any security control action to be taken? Did an intruder compromise a system because security patches were not up to date? If so, should there be a procedure to ensure that patches are applied when they become available? Was access obtained because of a poorly chosen password? If so, should there be a campaign to encourage users to construct strong passwords? If there were control failures, what should be done to prevent similar attacks in the future?

· Did the incident response plan work? Did everyone know who to notify? Did the team have needed resources? Was the response fast enough? Were certain critical resources unnecessarily affected? What should be done differently next time?

The incident response plan ensures that incidents are handled promptly, efficiently, and with minimal harm.

Incident Response Teams

Many organizations name and maintain a team of people trained and authorized to handle a security incident. Such teams, called  computer security incident response teams ( CSIRTs) or  computer emergency response teams ( CERTs), are standard at large private and government organizations, as well as many smaller ones. A CSIRT can consist of one person, or it can be a flexible team of dozens of people on call for special skills they can contribute.

The September–October 2014 issue of  IEEE Security & Privacy magazine is devoted to CSIRTs. Papers include a case study of a national CSIRT and its coordination with other CSIRTs, how CSIRTs can (and must) automate the evaluation of millions of data items received hourly, and a study of CSIRT personnel from a psychological perspective to help teams be more effective.

Types of CSIRTs

When an incident occurs, responding to it may overtake a team member’s other ordinary responsibilities. For this reason, as an organization’s information technology operation becomes larger or more complex, the nature of its response capability often changes; a one-person incident response may grow to a larger, more dedicated team.

But an organizational incident response team does not operate in a vacuum. Although some incidents are confined to one organization, others often involve multiple targets, sometimes across organizational, political, and geographic boundaries. Here are some models for CSIRTs:

· Full organizational response team. This team covers all incidents and may include separate staff to deal with situations in different organizational units, such as plants in separate locations or distinct business units of a larger company.

· Coordination centers. These coordinate incident response activity across organizations. With this model, work is not duplicated unnecessarily, and efforts proceed toward the same goals.

· National CSIRTs. This model is responsible for coordinating within a country and communicating with the national CSIRTs of other countries.

· Sector CSIRTs. These assist with investigating and handling incidents specific to a particular business sector. For example, financial institutions or medical facilities may organize a sector CSIRT to address their common likely vulnerabilities and risks.

· Vendor CSIRTs. These teams address incidents or participate in responses involving one manufacturer’s products.

· Outsourced CSIRT teams. These groups are hired to perform incident response services on contract to other companies.

CSIRTs operate in organizations, nationally, internationally, by vendor, and by business sector.

An integral part of any response is the  security operations center ( SOC), which performs the day-to-day monitoring of a network and may be the first to detect and report an unusual situation. The SOCs work hand in hand with CSIRTs to identify problems, gather evidence, and participate in taking responsive action. At a higher level,  information sharing and analysis centers (ISACs) perform some CSIRT functions by sharing threat and incident data across CSIRTs. This collective action not only solves the immediate security problem but also heightens awareness and improves response options within a business sector or across countries.

CSIRT Activity

Responsibilities of a CSIRT include the following:

· reporting: receiving reports of suspected incidents and reporting as appropriate to senior management

· detection: investigating to determine whether an incident occurred

· triage: taking immediate action to address urgent needs

· response: coordinating effort to address all aspects in a manner appropriate to the incident’s severity and time demands

· postmortem: declaring the incident over, reviewing the incident and response, and making suggestions to improve future incident detection and response

· education: preventing harm by promulgating good security practices and disseminating lessons learned from past incidents

CSIRTs are effective not only in addressing incidents when they occur but also in studying trends, predicting future attacks, and taking action to prevent incidents before they happen. These proactive roles of a CSIRT in preventing attacks are increasing in importance, reports Robin Ruefle’s team [ RUE14 ]. CSIRTs’ predictions of future attack trends are useful in helping organizations determine where to invest their preventive resources. A CSIRT has a collection of individuals with broad and deep technical understanding of computer failures; during the time when the members are not occupied with an emergency they can use their expertise to develop greater understanding of how incidents happen, with the obvious goal of preventing future ones.

Team Membership

Not uncommonly, the incident response team of a large organization has 50 or more members. But not every team member needs to have the same skill set. At different times response teams need a variety of skills, including the ability to

· collect, analyze, and preserve digital forensic evidence

· analyze data to infer trends

· study the source, impact, and structure of malicious code

· help manage installations and networks by developing defenses such as signatures

· perform penetration testing and vulnerability analysis

· understand current technologies used in attacks

These skills are useful in many systems development capacities, not just for security. For instance, people with these skills can be part of teams evaluating system performance, designing new system capabilities, or testing new updates before they are fielded. So specialized skills can be brought into the response team as needed for specific incidents.

Even when the CSIRT is composed of members from other standing teams, naming the CSIRT members in advance lets an organization select people according to their personal and technical skills, try out different member groupings to determine whether the mix of people is effective, and let the CSIRT members develop camaraderie and trust before having to work together on an incident. Additionally, with advance notice, managers can plan for other people to take over the work of the person seconded to the incident response team for the duration of the incident.

Information Sharing

Incident response often requires sharing information—within an organization, with similarly affected ones, and with national officials.

As Robin Ruefle and colleagues [ RUE14 ] report, information sharing is a key responsibility of CSIRTs. An incident affecting one site may also affect another, and analysis from one place may help another. To date, however, there are no standards or even guidelines for automated information sharing between CSIRTs. Because of trust issues, much sharing now takes place informally, by word of mouth, in which one CSIRT member interacts with a known colleague at another. That model does not scale to larger-scale operation, nor does it support interchange with national and other coordinating CSIRTs. Information sharing is also stymied because of fears of competition, negative publicity, and regulations. However, the annual FIRST conference (Federation of Incident Response Security Teams) encourages informal interaction among teams and works toward developing more formal mechanisms for sharing information and building skills. (See  first.org  for more information.)

Determining Incident Scope

The scope of an incident is rarely obvious at the beginning. Heightened awareness may begin with someone’s noticing something irregular, no matter how inconsequential it may seem at first.  Sidebar 10-2  describes how a tiny irregularity exploded into a major incident. Although from some time ago, this example is instructive because it was heavily investigated and documented at the time.

Sidebar 10-2 Incorrect Account Balance Leads to Intruder

In 1986, Cliff Stoll was working as an astronomer at Lawrence Berkeley Laboratory when he noticed that the charges for computer accounts he managed did not add up properly. Although the mismatch was small—just a few cents—Stoll was unwilling to dismiss it as an inconsequential computer error. From the account listings, Stoll determined that someone had created an extra account that was being charged to Stoll’s projects. However, the monthly bill was not being delivered to Stoll or to anyone else because the account had no billing address. Coincidentally, Stoll received a report that someone from his site had been breaking into military computers, but he didn’t initially connect these two data points.

Stoll invalidated the unauthorized account but found that the attacker remained, having acquired system administrator privileges. Thinking the attacker was a student at a nearby university, Stoll and his colleagues planned a way to catch the attacker in the act. They soon found the flaw the attacker exploited but decided to keep the culprit engaged so they could investigate his actions, using an elaborate masquerade in which Stoll controlled everything the attacker could see and do [ STO88 STO89 ]. Stoll’s trap was one of the first examples of a honeypot (introduced in  Chapter 5 ).

After months of activity Stoll and authorities identified the attacker as a German agent named Markus Hess, recruited by the Soviet KGB. German authorities arrested Hess, who was convicted of espionage and sentenced to one to three years in prison.

Accounting records that did not balance—off by just US$0.75—led to investigation and conviction of an international spy. When you begin to investigate an incident, you may not know what its scope will be.

image1.jpeg