Clinical Application of Epidemiology

profilemauripereyra23
Chapter8clinicalEpidemiologyTheEssentials.docx

8

Prognosis

He, who would rightly distinguish those that will survive or die, as well as those that will be subject to disease a longer or shorter time, ought, from his knowledge and attention, to be able to form an estimate of all symptoms, and rationally to weigh their powers by comparison.

—Hippocrates 460-375 B.C.

Key Words

Prognosis

Prognostic factors

Clinical course

Natural history

Zero time

Inception cohort

Stage migration

Event

Survival analysis

Kaplan-Meier analysis

Time-to-event analysis

Censored

Hazard ratios

Case series

Case report

Clinical prediction rules

Training set

Test set

Validation

Prognostic stratification

Sampling bias

Migration bias

Dropouts

Measurement bias

Multiple imputation

Sensitivity analysis

Best-case/worst-case analysis

When people become sick, they have a great many questions about how their illness will affect them. Is it dangerous? Could I die of it? Will there be pain? How long will I be able to continue my present activities? Will it ever go away altogether? Most patients and their families want to know what to expect, even in situations where little can be done about their illness.

Prognosis is the prediction of the course of disease following its onset. This chapter reviews the ways in which the course of disease can be described. The intention is to give readers a better understanding of a difficult but indispensable task—predicting patients' futures as closely as possible. The objective is to avoid expressing prognoses with vagueness when unnecessary or with certainty when misleading.

Doctors and patients want to know the general course of the illness, but they want to go further and tailor this information to their particular situation as much as possible. For example, even though ovarian cancer is usually fatal in the long run, women with this cancer may live from a few months to many years, and they want to know where on this continuum their particular case is likely to fall.

Studies of prognosis are similar to cohort studies of risk factors. Patients are assembled who have a particular disease or illness in common, they are followed forward in time, and clinical outcomes are measured. Patient characteristics that are associated with an outcome of the disease, called  prognostic factors, are identified. Prognostic factors are analogous to risk factors, except that they represent a different part of the disease spectrum, from disease to outcomes. Case-control studies of people with the disease who do and do not have a bad outcome can also estimate the relative risk associated with various prognostic factors, but they are unable to provide information on outcome rates (see  Chapter 7 ).

DIFFERENCES IN RISK AND PROGNOSTIC FACTORS

Risk and prognostic factors differ from each other in several ways.

P.127

The Patients Are Different

Studies of risk factors usually deal with healthy people, whereas studies of prognostic factors are of sick people.

The Outcomes Are Different

For risk, the event being counted is usually the onset of disease. For prognosis, consequences of disease are counted, including death, complications, disability, and suffering.

The Rates Are Different

Risk factors are usually for low-probability events. Yearly rates for the onset of various diseases are on the order of 1/1,000 to 1/100,000 or less. As a result, relationships between exposure and disease are difficult to confirm in the course of day-today clinical experiences, even for astute clinicians. Prognosis, on the other hand, describes relatively frequent events. For example, several percent of patients with acute myocardial infarction die before leaving the hospital.

The Factors May be Different

Variables associated with an increased risk are not necessarily the same as those marking a worse prognosis. Often, they are considerably different for a given disease. For example, the number of well-established risk factors for cardiovascular disease (hypertension, smoking, dyslipidemia, diabetes, and family history of coronary heart disease) is inversely related to the risk of dying in the hospital after a first myocardial infarction  (1).

Clinicians can often form good estimates of short-term prognosis from their own personal experience. However, they may be less able to sort out, without the assistance of research, the various factors that are related to long-term prognosis or the complex ways in which prognostic factors are related to one another.

CLINICAL COURSE AND NATURAL HISTORY OF DISEASE

Prognosis can be described as either the clinical course or natural history of disease. The term  clinical course describes the evolution (prognosis) of a disease that has come under medical care and has been treated in a variety of ways that affect the subsequent course of events. Patients usually receive medical care at some time in the course of their illness when they have diseases that cause symptoms such as pain, failure to thrive, disfigurement, or unusual behavior. Examples include type 1 diabetes mellitus, carcinoma of the lung, and rabies. After such a disease is recognized, it is likely to be treated.

The prognosis of disease without medical intervention is termed the  natural history of disease. Natural history describes how patients fare if nothing is done about their disease. A great many health conditions do not come under medical care, even in countries with advanced health care systems. They remain unrecognized because they are asymptomatic (e.g., many cancers of the prostate are occult and slow growing) and are, therefore, unrecognized in life. For others, such as osteoarthritis, mild depression, or lowgrade anemia, people may consider their symptoms to be one of the ordinary discomforts of daily living, not a disease and, therefore, not seek medical care for them.

Example

Irritable bowel syndrome is a common condition that involves abdominal pain and disturbed bowel habits not caused by other diseases. How often do patients with this condition visit doctors? In a British cohort of 3,875 people without irritable bowel syndrome at baseline, 15% developed the syndrome over the next 10 years  (2). Of these, only 17% consulted their primary care physician with related symptoms at least once in 10 years, and 4% had consulted in the past year. In another study, characteristics of the abdominal complaints did not account for whether patients with irritable bowel syndrome sought health care for their symptoms  (3).

ELEMENTS OF PROGNOSTIC STUDIES

Figure 8.1 shows the basic design of a cohort study of prognosis. At best, studies of prognosis are of a defined clinical or geographic population, begin observation at a specified point in time in the course of disease, follow up all patients for an adequate period of time, and measure clinically important outcomes.

Patient Sample

The purpose of representative sampling from a defined population is to assure that study results have the greatest possible generalizability. It is sometimes possible to study prognosis in a complete sample of patients with new-onset disease in large regions. In some countries,

P.128

the existence of national medical records makes population-based studies of prognosis possible.

View Figure

Figure 8.1. Design of a cohort study of risk.

Example

Dutch investigators studied the risk of complications of pregnancy in women with type 1 diabetes mellitus  (4). The sample included all of the 323 women in the Netherlands with type 1 diabetes who had become pregnant during a 1-year period and had been under care in one of the nation's 118 hospitals. Most pregnancies were planned, and during pregnancy, most women took folic acid supplements and had good control of their blood sugar. Nevertheless, complication rates in newborns were much higher than in the general population. Neonatal morbidity (one or more complications) occurred in 80% of infants and rates of congenital malformations and unusually large newborns (macrosomia) were 3-fold to 12-fold higher than in the general population. This study suggests that good control of blood sugar alone was not sufficient to prevent complications of pregnancy in women with type 1 diabetes.

Even without national medical records, populationbased studies are possible. In the United States, the Network of Organ Sharing collects data on all patients with transplants, and the Surveillance, Epidemiology, and End Results (SEER) program collects incidence and survival data on all patients with new-onset cancers in several large areas of the country, comprising 28% of the US population. For primary care questions, in the United States and elsewhere, individual practices in communities have banded together into “primary care research networks” to collect research data on their patients' care.

Most studies of prognosis, especially for less common diseases, are of local patients. For these studies, it is especially important to provide the information that users can rely on to decide whether the results generalize to their own situation: patients' characteristics (e.g., age, severity of disease, comorbidity), the setting where they were found (e.g., primary care practices, community hospitals, referral centers), and how they were sampled (e.g., complete, random, or convenience sampling). Often, this information is sufficient to establish wide generalizability, for example, in studies of community-acquired pneumonia or thrombophlebitis in a local hospital.

Zero Time

Cohorts in prognostic studies should begin from a common point in time in the course of disease, called  zero time, such as at the time of the onset of symptoms, diagnosis, or the beginning of treatment. If observation begins at different points in the course of disease for the various patients in a cohort, the description of their prognosis will lack precision, and the timing of recovery, recurrence, death, and other outcome events will be difficult to interpret or will be misleading. The term  inception cohort is used to describe a group of patients that is assembled at the onset (inception) of their disease.

Prognosis of cancer is often described separately according to patients' clinical stage (extent of spread)

P.129

at the beginning of follow-up. If it is, a systematic change in how stage at zero time is established can result in a different prognosis for each stage even if the course of disease is unchanged for each patient in the cohort. This has been shown to happen during staging of cancer—assessing the extent of disease, with higher stages corresponding to more advanced cancer, which is done for the purposes of prognosis and choice of treatment.  Stage migration occurs when a newer technology is able to detect the spread of cancer better than an older staging method. Patients who used to be classified in a lower stage are, with the newer technology, classified as being in a higher (more advanced) stage. Removal of patients with more advanced disease from lower stages results in an apparent improvement in prognosis for each stage, regardless of whether treatment is more effective or prognosis for these patients as a whole is better. Stage migration has been called the “Will Rogers phenomenon” after the humorist who said of the geographic migration in the United States during the economic depression of the 1930s, “When the Okies left Oklahoma and moved to California, they raised the average intelligence in both states”  (5).

Example

Positron emission tomography (PET) scans, a sensitive test for metastases, are used to stage non-small cell lung cancers. Investigators compared cancer stages before and after PET scans were in general use and found a 5.4% decline in the number of patients with stage III disease (cancer spread within the chest) and a 8.4% increase in patients with stage IV disease (distant metastases)  (6). PET staging was associated with better survival in stage III and stage IV disease, but not earlier stages. The authors concluded that stage migration was responsible for at least some of the apparent improvement in survival in patients with stage III and stage IV lung cancer that occurred after PET scan staging was introduced.

Follow-Up

Patients must be followed for a long enough period of time for most of the clinically important outcome events to have occurred. Otherwise, the observed rate will understate the true one. The appropriate length of follow-up depends on the disease. For studies of surgical site infections, the follow-up period should last for a few weeks, and for studies of the onset of dementia and its complications in patients with mild cognitive impairment, the follow-up period should last at least several years.

Outcomes of Disease

Descriptions of prognosis should include the full range of manifestations of disease that would be considered important to patients. This means not only death and disease but also pain, anguish, and the inability to care for one's self or pursue usual activities. The 5 Ds—death, disease, discomfort, disability, and dissatisfaction—are a simple way to summarize important clinical outcomes (see  Table 1.2 ).

In their efforts to be “scientific,” physicians tend to value precise or technologically measured outcomes, sometimes at the expense of clinical relevance. As discussed in  Chapter 1 , clinical effects that cannot be directly perceived by patients, such as radiologic reduction in tumor size, normalization of blood chemistries, improvement in ejection fraction, or change in serology, are not clinically useful ends in themselves. It is appropriate to substitute these biologic phenomena for clinical outcomes only when the two are known to be related to each other. Thus, in patients with pneumonia, short-term persistence of abnormalities on chest radiographs may not be alarming if the patient's fever has subsided, energy has returned, and cough has diminished.

Ways to measure patient-centered outcomes are now used in clinical research.  Table 8.1 shows a simple measure of quality of life used in studies of cancer treatment. There are also research measures for performance status, health-related quality of life, pain, and other aspects of patient well-being.

DESCRIBING PROGNOSIS

It is convenient to summarize the course of disease as a single rate—the proportion of people experiencing an event during a fixed time period. Some rates used for this purpose are shown in  Table 8.2. These rates have in common the same basic components of incidence: events arising in a cohort of patients over time.

A Trade-Off: Simplicity Versus More Information

Summarizing prognosis by a single rate has the virtue of simplicity. Rates can be committed to memory and communicated succinctly. Their drawback is that relatively little information is conveyed. Large differences in prognosis can be hidden within similar summary rates.

P.130

TABLE 8.1 A Simple Measure of Quality of Life. The Eastern Collaborative Oncology Group's Performance Scale

Performance Status

Definition

0

Asymptomatic

1

Symptomatic, fully ambulatory

2

Symptomatic, in bed <50% of the day

3

Symptomatic, in bed >50% of the day

4

Bedridden

5

Dead

Adapted with permission from Oken MM, Creech RH, Tormey DC, et al. Toxicity and response criteria of the Eastern Cooperative Oncology Group.  Am J Clin Oncol 1982;5:649-655.

Figure 8.2  shows 5-year survival rates for patients with four conditions. For each condition, about 10% of the patients are alive at 5 years. However, the clinical courses are otherwise quite different in ways that are very important to patients. Early survival in patients with dissecting aneurysms is very poor, but if they survive the first few months, their risk of dying is much less affected by having had the aneurysm ( Fig. 8.2A ). Patients with locally invasive, non-small cell lung cancer experience a relatively constant mortality rate throughout the 5 years following diagnosis ( Fig. 8.2B ). The life of patients with amyotrophic lateral sclerosis (ALS, Lou Gehrig disease, a slowly progressive paralysis) and respiratory difficulties is not immediately threatened, but as neurologic function continues to decline over the years, the inability to breathe without assistance leads to death ( Fig. 8.2C ).  Figure 8.2D is a benchmark. Only at age 100 years do people in the general population have a 5-year survival rate comparable to that of patients with the three diseases.

TABLE 8.2 Rates Commonly Used to Describe Prognosis

Rate

Definition a

5-year survival

Percent of patients surviving 5 years from some point in the course of their disease

Case fatality

Percent of patients with a disease who die of it

Disease-specific mortality

Number of people per 10,000 (or 100,000) population dying of a specific disease

Response

Percent of patients showing some evidence of improvement following an intervention

Remission

Percent of patients entering a phase in which disease is no longer detectable

Recurrence

Percent of patients who have return of disease after a diseasefree interval

a Time under observation is either stated or assumed to be sufficiently long so that all events that will occur have been observed.

Survival Analysis

When interpreting prognosis, it is preferable to know the likelihood, on average, that patients with a given condition will experience an outcome at any point in time. Prognosis expressed as a summary rate does not contain this information. However, figures can show information about average time to event for any point in the course of disease. An  event refers to a dichotomous clinical outcome that can occur only once. The following discussion takes the common approach of describing outcomes in terms of “survival,” but the same methods apply to the reverse (time to death) and to any other outcome event such as cancer recurrence, cure of infection, freedom from symptoms, or arthritis becoming inactive.

Survival of a Cohort

The most straightforward way to learn about survival is to assemble a cohort of patients who have the condition of interest and are at the same point in the course of their illness (e.g., onset of symptoms, diagnosis, beginning of treatment), and then keep them all under observation until all experience the outcome or not. For a small cohort, one might then represent these patients' clinical course, as shown in  Figure 8.3A . The plot of survival against time displays steps corresponding to the death of each of the 10 patients in the cohort. If the number of patients were increased, the size of the steps would diminish. If a very large number of patients were studied, the figure would approximate a smooth curve ( Fig. 8.3B). This information could then be used to predict the year-by-year, or even week-by-week, prognosis of similar patients.

Unfortunately, obtaining the information in this way is impractical for several reasons. Some of the patients might drop out of the study before the end of the follow-up period, perhaps because of another illness, a move from the study area, or dissatisfaction with the study. These patients would have to be

P.131

P.132

excluded from the cohort even though considerable effort had been exerted to gather data on them until the point at which they dropped out. Also, it would be necessary to wait until all of the cohort's members had reached each point in follow-up before the probability of surviving to that point could be calculated. Because patients ordinarily become available for a study over a period of time, at any point in calendar time, there would be a relatively long follow-up for patients who had entered the study first, but only brief experience with those who had entered more recently. The last patient who entered the study would have to reach each year of follow-up before any information on survival to that year would be available.

View Figure

Figure 8.2. A limitation of 5-year survival rates: Four conditions with the same 5-year survival rate of 10%.

View Figure

Figure 8.3. Survival of two cohorts, small and large, when all members are observed for the full period of follow-up.

View Figure

Figure 8.4. Example of a survival curve, with detail for one part of the curve.

Survival Curves

To make efficient use of all available data from each patient in the cohort,  survival analysis has been developed to  estimate the survival of a cohort over time. The usual method is called  Kaplan-Meier analysis, after its originators. Survival analysis can be applied to any outcomes that are dichotomous and occur only once during follow-up (e.g., time to coronary event or to recurrence of cancer). A more general term, useful when an event other than survival is described, is  time-to-event analysis.

Figure 8.4  shows a simplified survival curve. On the vertical axis is the estimated probability of surviving, and on the horizontal axis is the period of time from the beginning of observation (zero time).

The probability of surviving to any point in time is estimated from the cumulative probability of surviving each of the time intervals that preceded it. Time intervals can be made as small as necessary; in Kaplan-Meier analyses, the intervals are between each new event, such as death, and the preceding one, however short or long that is. Most of the time, no one dies and the probability of surviving is 1. When a patient dies, the probability of surviving at that moment is calculated as the ratio of the number of patients surviving to the number at risk of dying at that point in time. Patients who have already died, dropped out, or have not yet been followed up to that point are not at risk of dying and are, therefore, not used to estimate survival for that time. The probability of surviving does not change during intervals in which no one dies, so it is recalculated only when there is a death. Although the probability at any given

P.133

interval is not very accurate, because either nothing has happened or there has been only one event in a large cohort, the overall probability of surviving up to each point in time (the product of all preceding probabilities) is remarkably accurate. When patients are lost from the study at any point in time, they are referred to as  censored and are no longer counted in the denominator from that point forward.

A part of the survival curve in  Figure 8.4  (from 3 to 5 years after zero time) is presented in detail to illustrate the data used to estimate survival: patients at risk, patients no longer at risk (censored), and patients experiencing outcome events at each point in time.

Variations on basic survival curves increase the amount of information they convey. Including the numbers of patients at risk at various points in time gives some idea of the contribution of chance to the observed rates, especially toward the end of follow-up. The vertical axis can show the proportion with, rather than without, the outcome event; the resulting curve will sweep upward and to the right. The precision of survival estimates, which declines with time because fewer and fewer patients are still under observation, can be shown by confidence intervals at various points in time (see  Chapter 11 ). Tics are sometimes added to the survival curves to indicate each time a patient is censored.

Interpreting Survival Curves

Several points must be kept in mind when interpreting survival curves. First, the vertical axis represents the  estimated probability of surviving for members of the cohort, not the cumulative incidence of surviving if all members of the cohort were followed up.

Second, points on a survival curve are the best estimate, for a given set of data, of the probability of survival for members of a cohort. However, the precision of these estimates depends on the number of patients on whom the estimate is based, as do all observations of samples. One can be more confident that the estimates on the left-hand side of the curve are sound, because more patients are at risk early in follow-up. But on the right-hand side, at the tail of the curve, the number of patients on whom estimates of survival are based may become relatively small because of deaths, dropouts, and late entrants to the study, so that fewer and fewer patients are followed for that length of time. As a result, estimates of survival toward the end of the follow-up period are imprecise and can be strongly affected by what happens to relatively few patients.

For example, in  Figure 8.4, only one patient was under observation at year 5. If that one remaining patient happened to die, the probability of surviving would fall from 8% to zero. Clearly, this would be a too literal reading of the data. Therefore, estimates of survival at the tails of survival curves must be interpreted with caution.

Finally, the shape of many survival curves gives the impression that outcome events occur more frequently early in follow-up than later on, when the slope approaches a plateau. But this impression is deceptive. As time passes, rates of survival are being applied to a diminishing number of patients, causing the slope of the curve to flatten even if the rate of outcome events did not change.

As with any estimate, Kaplan-Meier estimates of time to event depend on assumptions. It is assumed that being censored is not related to prognosis. To the extent that this is not true, a survival analysis may yield biased estimates of survival in cohorts. The Kaplan-Meier method may not be accurate enough if there are competing risks—more than one kind of outcome event—and the outcomes are not independent of each other such that one event changes the probability of experiencing the other. For example, patients with cancer who develop an infection related to aggressive chemotherapy and drop out for that reason may have had a different chance of dying of the cancer. There are other methods for estimating cumulative incidence in the presence of competing risks.

IDENTIFYING PROGNOSTIC FACTORS

Often, studies go beyond a simple description of prognosis in a homogeneous group of patients to compare prognosis in patients with different characteristics, that is, they identify prognostic factors. Multiple survival curves, one for patients with each of the characteristics, are represented on the same figure where they can be visually (and statistically) compared.

Example

Patients with gastric cancer, like many other cancers, have widely different chances of surviving over the several years after diagnosis. Prognosis varies according to characteristics of the cancer such as stage (the depth of the cancer from being limited to the surface at one extreme to invasion of adjacent organs at the other), location (proximity to the esophagus), and number of lymph nodes involved. A study combined these three characteristics into

P.134

seven prognostic groups for gastric cancer ( Fig. 8.5(7). At 5 years of follow-up, in the most favorable group, almost 90% of patients were alive whereas in the least favorable group, less than 10% were. This information might be especially useful in helping patients and doctors understand what lies ahead, and it is much more informative than simply saying that “overall, 30% of patients with gastric cancer survive 5 years.”

View Figure

Figure 8.5. Example of prognostic stratification. Survival from surgery in a patient with gastric cancer according to prognostic strata. (Redrawn with permission from Sano T, Coit DG, Kim HH, et al. Proposal of a new stage grouping of gastric cancer for TNM classification: International Gastric Cancer Association staging project.  Gastric Cancer 2017;20(2):217-225.)

The effects of one prognostic factor relative to the effects of another can be summarized from data in a time-to-even analysis by a  hazard ratio, which is analogous to a risk ratio (relative risk). Also, survival curves can be compared after taking into account other factors related to prognosis so that the independent effect of just one variable is examined.

CASE SERIES

case series is a description of the course of disease in a small number of cases, a few dozen at most. An even smaller report, with fewer than 10 patients, is called a  case report. Cases are typically found at a clinic or referral center and then followed forward in time to describe the course of disease and backward in time to describe what came earlier.

Such reports can make an important contribution to understanding of disease, primarily by describing experiences with newly defined syndromes or uncommon conditions. The reason for introducing case series into a chapter about prognosis is that they may masquerade as true cohort studies even though they do not have comparable strengths.

Example

Physicians in emergency departments see patients with bites from North American rattlesnakes. These bites are relatively uncommon at any one place, making it difficult to carry out large cohort studies of their clinical course, so physicians must rely mainly on case series. An example is a description of the clinical course of all 24 children managed at a children's hospital in California during a 10-year period  (8). Nineteen of the children were actually injected with venom, and they were managed with the aggressive use of antivenin. Three had surgical treatment to remove soft tissue debris or relieve tissue pressure. There were no serious reactions to antivenin, and all patients left the hospital without functional impairment.

Physicians caring for children with rattlesnake bites would be grateful for this and other case series if there were no better information to guide their care, but that is not to say that the case series provided a complete and fully reliable picture of snakebite care. Of all children bitten in that region, some may have been doing so well after a bite that they were not sent to the referral center. Others might have been doing so badly that they were rushed to the nearest hospital or even died before reaching a hospital at all. In other words, the case series does not describe the clinical course of all children from the time of snakebite (the inception) but rather a selected sample of children who happened to come under care at that particular hospital. In effect, case series describe the clinical course of prevalent, not necessarily a representative sample of incident cases, so they are “false” cohorts.

CLINICAL PREDICTION RULES

A combination of variables can provide a more precise prognosis than any of these variables taken one at a time. As discussed in  Chapter 5 clinical prediction rules estimate the probability of outcomes (either prognosis or diagnosis) according to a set of patient characteristics defined by history, physical examination, and simple laboratory tests. They are “rules” because they are often tied to recommendations about further diagnostic evaluation or treatment.

P.135

Many other names are used for the same concept, such as decision rules, prediction models, and risk scores. To make prediction rules workable in clinical settings, they depend on data that are available in the usual care of patients and scoring, the basis for the prediction, that has been simplified. Prediction models or rules typically use a fixed point in time in the future to determine outcome rates, and performance is often assessed with discrimination, calibration, and reclassification, as discussed in  Chapter 5 .

Example

Cirrhosis in advanced stages can be rapidly fatal. The Model for End Stage Liver Disease (MELD) is a scoring system that predicts mortality within 90 days for patients with cirrhosis, and has been used in the United States since 2002 to guide which patients on the waiting list should be prioritized for liver transplantation. The original score used three laboratory values that correlate with liver function and mortality: INR, creatinine, and bilirubin. A subsequent analysis showed prediction can be improved by including sodium levels  (9), with better discrimination and calibration. In particular, patients who died with low MELD scores in the old model had higher scores with the model that accounted for low sodium levels. The model was updated to include sodium in 2016. Mortality rates predicted by the score are shown in  Table 8.3. Mortality begins increasing steeply at scores of 15 or more; at this threshold, patients with cirrhosis are often considered transplant candidates, depending on organ availability. The impact of the updated score was assessed in subsequent cohorts. Mortality improved for those awaiting liver transplantation, particularly for patients with low sodium levels who would not have qualified for transplants in prior years  (10).

A clinical prediction rule should be developed in one setting and tested in others—with different patients, physicians, and usual care practices—to assure that predictions are good for a broad range of settings and not just for where it was developed, because it might have been the result of the particular characteristics of that setting. The data used to develop the prediction rule is called the  training set, and the data used to assess its validity is called the  test set, which is used for  validation of the prediction rule.

The process of separating patients into groups with different prognosis, as in the previous example, is called  prognostic stratification. In this case, cirrhosis is the disease and mortality is the outcome. The concept is analogous to risk stratification ( Chapter 5 ), where patients are divided into different strata of risk for developing disease.

TABLE 8.3 Prognosis Using the Model for End Stage Liver Disease (MELD) Score

□ Creatinine (mg/dL)

□ Bilirubin (mg/dL)

□ INR

□ Sodium

Calculators determine the score based on the following equation:

MELD(i) = 0.957 × Ln(creatinine) + 0. 378 × Ln(bilirubin) + 1.120 × Ln(INR) + 0.643

If initial MELD score greater than 11, the MELD score is then recalculated as follows:

MELD = MELD(i) + 1.32 × (137 - Na) - [0.033 - MELD(i) × (137 - Na)]

MELD

Predicted Mortality at 90 Days (%)

<15

1

15-20

4

21-22

7

23-26

13

27-31

27

32-40

81

Data from Kim WR, Biggins SW, Kremers WK, et al. Hyponatremia and mortality among patients on the liver-transplant waiting list.  N Engl J Med 2008;359(10):1018-1026; and Liver and Intestinal Organ Transplantation Committee. OPTN/UNOS Policy Notice Revisions to National Liver Review Board Policies (2019). Department of Health and Human Services, Health Resources and Services Administration, Healthcare Systems Bureau, Division of Transplantation, Rockville, MD; United Network for Organ Sharing, Richmond, VA.

BIAS IN COHORT STUDIES

In cohort studies of risk or prognosis, bias can alter the description of the course of disease. Bias can also create apparent differences between groups when differences do not actually exist in nature or obscure differences when they really do exist. These biases have their counterparts in case-control studies as well. They are a separate consideration from confounding and effect modification ( Chapter 6 ).

There are an almost infinite variety of systematic errors, many of which are given specific names, but some are more basic. They can be recognized more easily when one knows where they are most likely to

P.136

occur in the course of a study. With that in mind, this section describes some possibilities for bias in cohort studies and discusses them in relation to the following study of prognosis.

Example

Bell palsy is the sudden, one-sided, unexplained onset of weakness of the face in the area innervated by the facial nerve. Often, the cause is unknown, although some cases are associated with herpes simplex, diabetes, pregnancy, and a variety of other conditions. What is the clinical course? Investigators in Denmark followed up 1,701 patients with Bell palsy once a month until function returned or 1 year  (11). Recovery reached its maximum within 3 weeks in 85% of patients and in 3 to 5 months in the remaining 15%. By that time, 71% of patients had recovered fully; for the rest, the remaining loss of function was slight in 12%, mild in 13%, and severe in 4%. Recovery was related to younger age, less initial palsy, and earlier onset of recovery.

When thinking about the validity of this study, one should consider at least the following.

Sampling Bias

Sampling bias has occurred when the patients in a study are not like other patients with the condition. Were patients in this study like others with Bell palsy? The answer depends on the user's perspective. Patients were “from the Copenhagen area” and apparently under the care of an ear, nose, and throat specialist, so the results generalize to other referred patients (as long as we accept that Bell palsy in Denmark is similar to this condition in other parts of the world). However, mild cases might not have been included because they were managed by local clinicians and quickly recovered or not brought to medical attention at all, which limits the applicability in primary care settings.

Sampling bias can also be misleading when prognosis is compared across groups and sampling has produced groups that are systematically different with respect to prognosis, even before the factor of interest is considered. In the Bell palsy example, older patients might have had a worse prognosis because they are the ones who had an underlying herpes virus infection, not because of their age.

Is this not confounding? Strictly speaking, it is not because the study is for the purpose of prediction, not to identify independent “causes” of recovery. Also, it is not plausible to consider such phenomena as severity of palsy or onset of recovery as causes because they are probably part of the chain of events leading from disease to recovery. But to the extent that one wants to show that prognostic factors are independent predictors of outcome, the same approaches as are used for confounding ( Chapter 6 ) can be used to establish independence.

Migration Bias

Migration bias is present when some patients drop out of the study during follow-up and they are systematically different from those who remain. It is often the case that some members of the original cohort leave a study over time. (Patients are assured that this is their right as part of the ethical conduct of research on humans.) If  dropout occurs randomly, such that the characteristics of lost patients are on average similar to patients who remain, then there would be no bias. This is so regardless of whether the number of dropouts is large or similar in the cohorts being compared, but ordinarily the characteristics of lost patients are not the same as those who remain in a study. Dropping out tends to be related to prognosis. For example, patients who are doing especially well or badly with their Bell palsy may be more likely to leave the study, as would those who need care for other illnesses, for whom the extra visits related to the study would be burdensome. This would distort the main (descriptive) results of the study—rate and completeness of recovery. If the study also aims to identify prognostic factors (e.g., recovery in old vs. young patients), that also could be biased by patients dropping out, for the same reasons.

Migration bias might be seen as an example of selection bias because patients who were still in the study when outcomes were measured were selected from all those who began in the study. Migration bias might be considered an example of measurement bias because patients who migrate out of the study are no longer available when outcomes are measured.

Measurement Bias

Measurement bias is present when members of the cohort are not all assessed similarly for outcome. In the Bell palsy study, all members of the cohort were examined by a common protocol every month until they were no longer improving, ruling out this possibility. If it had been left to individual patients and physicians whether, when, and how they were examined, this would have diminished confidence in the description of time to and completeness of recovery.

Measurement bias also comes into play if prognostic groups are compared and patients in one

P.137

group have a systematically better chance of having outcomes detected than those in another. Some outcomes, such as death, cardiovascular catastrophes, and major cancers, are so obvious that they are unlikely to be missed. But for less clear-cut outcomes, including specific cause of death, subclinical disease, side effects, or disability, measurement bias can occur because of differences in the methods with which the outcome is sought or classified.

Measurement bias can be minimized in three general ways: (i) examine all members of the cohort equally for outcome events; (ii) if comparisons of prognostic groups are made, ensure that researchers are unaware of the group to which each patient belongs; and (iii) set up careful rules for deciding if an outcome event has occurred (and follow the rules). To help readers understand the extent of these kinds of biases in a given study, it is usual practice to include, with reports of the study, a flow diagram describing how the number of participants changed as the study progressed and why. It is also helpful to compare the characteristics of patients in and out of the study after sampling and follow-up.

Bias from “Non-differential” Misclassification

Until now, we have been discussing how the results of a study can be biased when there are systematic differences in how exposure or disease groups are classified. But bias can also result if misclassification is “non-differential,” that is, it occurs similarly in the groups being compared. In this case, the bias is toward finding no effect.

Example

When cigarette smoking is assessed by simply asking people whether they smoke, there is substantial misclassification relative to a gold standard such as the presence or absence of cigarette products in saliva. Yet in a cohort study of cigarette smoking and coronary heart disease (CHD), misclassification of smoking could not be different in people who did or did not develop CHD because the outcome was not known at the time exposure was assessed. Even so, when a real difference in CHD rates exists between smokers and nonsmokers, and smoking is incorrectly classified, then the measured difference is reduced, making a “null” effect more likely. At the extreme, if classifying smoking status were totally at random, there could be no association between smoking and CHD.

Bias from Missing Data

Missing data can lead to bias in studies. Migration bias can be considered a missing data problem, with baseline data such as patient characteristics perhaps available, but no outcome data on follow-up. Missing data can also lead to a form of measurement bias, where measurements are not just inconsistent, but completely missing in a systematic way that is related to the outcome.

Ideally, data are collected as thoroughly as possible during the course of the study. After completion of the study, some analytic techniques attempt to minimize the effects of bias due to missing data. A baseline analysis may use only patients who have complete data. However, if the reason data are missing is related to prognosis, then analysis using only complete cases gives biased results (as discussed for migration bias above). Another approach is imputation, making “educated guesses” to fill in the missing data.  Multiple imputation is one type of imputation, in which the investigator creates a statistical model, based on patients with complete data, to estimate what the missing values would be for patients with similar characteristics. This method assumes that all the factors that account for missing data are measured in the study, that these factors are appropriately modeled, and that otherwise the missing values occurred at random. Statistically modeling is discussed in  Chapter 11 .

BIAS, PERHAPS, BUT DOES IT MATTER?

Clinical epidemiology is not an error-finding game. Rather, it is meant to characterize the credibility of a study so that clinicians can decide how much to rely on its results when making high-stakes decisions about patients. It would be irresponsible to ignore results of studies that meet high standards, just as clinical decisions need not be bound by the results of weak studies.

With this in mind, it is not enough to recognize that bias might be present in a study. One must go on to determine if bias is actually present in the particular study. Beyond that, one must decide whether the consequences of bias are sufficiently large that they change the conclusions in a clinically important way. If damage to the study's conclusions is not very great, then the presence of bias is of little practical consequence and the study is still useful.

SENSITIVITY ANALYSIS

One way to decide how much bias might change the conclusions of a study is to do a  sensitivity analysis,

P.138

that is, to show how much larger or smaller the observed results might have been under various assumptions about the missing data or potentially biased measurements. The analysis could compare the results from using only patients with complete data to ones that use all patients, with missing data imputed under different assumptions. A  best-case/worst-case analysis tests the effects of the most extreme possible assumptions but is an unreasonably severe test for the effects of bias in most situations. More often, sensitivity analyses test the effects of somewhat unlikely values, as in the following example.