Module 5
Performance, Goals, and Alignment Assignment
A. Mission, Goals, and Objectives
Usually the most meaningful performance measures are derived from the mission,
goals, objectives, and, sometimes, service standards that have been established for a
particular program. This is because goals and objectives, and to a lesser extent mission
and service standards, define the desired results to be produced by an agency or program.
Clear goals and objectives are intended to improve organizational performance by
focusing employees’ energy and efforts on desired results (Locke & Latham, 1990).
Thus, there is usually a direct connection between goals and objectives on the one hand
and outcomes or effectiveness measures on the other. While it is often very useful to
develop logic models to fully understand all the performance dimensions of a public or
nonprofit program, depending on the purpose of the measurement system it is sometimes
sufficient to clarify goals and objectives and then define performance measures to track
their accomplishment. It should be understood that there are no universal distinctions
among these terms in the public management literature, and there is often considerable
overlap among them, but the definitions we use in this book are workable and not
severely incompatible with the distinctions others have made. Mission refers to the basic
purpose of an organization or a program, its reason for being, and the general means
through which it accomplishes that purpose. Goals are general statements about the
results to be produced by the program, and objectives are more specific milestones to be
achieved in order to accomplish the goals. Whereas goals are often formulated as very
general, often timeless, sometimes idealized outcomes, objectives should be specified in
more tangible terms.
The US Department of Health and Human Services (DHHS) is a good example of
a large federal department that has gone through the process of clarifying its mission,
goals, objectives, and performance measures in compliance with the Government
Performance and Results Act (GPRA) of 1993 and the more recent GPRA Modernization
Act of 2010. DHHS, the federal government’s principal agency for protecting the health
of Americans and providing essential human services, manages more than three hundred
programs through eleven operating agencies and an extended network of state, local, and
other grantees in a wide variety of areas, such as medical and social science research,
food and drug safety, financial assistance and health care for low-income individuals,
child support enforcement, maternal and infant health, substance abuse treatment and
prevention, health insurance, and services for older Americans. The department’s formal
mission statement is, “To enhance the health and well-being of Americans by providing
for effective health and human services and by fostering strong, sustained advances in the
sciences underlying medicine, public health, and social services.”
The Administration for Children and Families has lead responsibility for
performance in this area, and the data to operationalize this measure will be taken from
the adoption and foster care reporting system, while the target on this measure for FY
2013 is 92.5 percent or higher. Obviously each of these seven performance measures
represents one slice, or dimension, of this particular objective, one perspective on what
the results should look like. All seven indicators in this measure set are clearly aligned
with the objective of promoting the safety, well-being, resilience, and healthy
development of children and youth. Collectively this set of measures is intended to
provide a balanced perspective on whether and the extent to which progress is made in
accomplishing this objective over time.
While goals structures provide a different starting point as opposed to program
logic models for identifying the aspects of performance that should be captured in a
measurement system, the two are by no means incompatible. Indeed, in managing public
and nonprofit programs, program managers and others frequently establish goals and
objectives for their programs that are likely to focus on varying aspects of the program
and its underlying logic. In general, managers are concerned with ensuring that
programmatic activities are conducted efficiently and productively, that the quality of
these activities and the outputs they produce are of high quality, that outputs are produced
at the required levels, that the intended outcomes do in fact materialize, and that clients
are satisfied with both the services they receive and the outcomes they experience.
However, at any time, their goals and objectives are likely to focus in particular on those
program components and performance criteria where improvement is most needed, and
these focal points of interest are likely to change over time as conditions require.
Working back through the program logic, the data might indicate that the program
is not doing a good job of preparing clients for viable occupations, an initial outcome in
the logic model, and appropriate objectives might be set regarding better preparation of
clients for such occupations, which might be different from the kinds of occupations the
program has been focusing on. In turn, this might lead to a finding that the training
programs conducted, a chief output of this program, and new objectives might well focus
on strengthening those training programs. Alternatively, the program might be doing a
good job of providing training programs and preparing clients for viable occupations, but
the problem lies in the fact that staff has not been doing a good job of identifying good
prospective jobs. Thus, clients often must settle for jobs that are not particularly well
suited for them, and this is not working out well in the long run for either the clients or
the employers. This would likely lead to new objectives to increase both the volume and
quality of this particular output: suitable jobs identified. In addition, further investigation
may find that staff are not producing useful on-the-job evaluations of clients in the initial
jobs, allowing “misfits” to continue, and this might lead to clearer objectives regarding
the production of more discerning assessments with more helpful recommendations,
addressing another quality-of-output issue. The point here is that goals and objectives
might well pertain to any or all of the elements in a program logic, and while they might
shift over time as might a logic model itself, they often provide a good point of departure
regarding those aspects of a program’s performance that are important to measure.
B. SMART Objectives
It is often helpful for program objectives to specify milestones to be attained
within certain time periods, but in practice, objective statements are often overly general,
vague, and open-ended in terms of time. Poorly written objectives fail to convey any
management commitment to achieve particular results, and they provide little guidance
for defining meaningful measures to assess performance. However, specific goals tend to
help focus energy and attention on producing desired results in specific amounts rather
than being scattered across a range of necessary and unnecessary activities (Carroll &
Tosi, 1973). Truly useful program objectives can be developed using the SMART
convention, stating objectives that are Specific in terms of the results to be achieved:
Measurable, Ambitious but Realistic, and Time-bound (Broom, Harris, Jackson, &
Marshall, 1998). With respect to highway traffic safety programming, for example, the
objective of reducing the reported number of crashes on the nation’s highways to fewer
than 3 per 100 million vehicle-miles driven and the number of highway accident fatalities
down to no more than 10 per 100,000 US residents by the year 2020 would be a SMART
objective.
SMART objectives clearly indicate the kind of result or improvement to be
obtained within a specified time period. The measure to be used in determining whether
an objective has been achieved at the end of that period should also be identified along
with the SMART objective. For example, the measures to be used in conjunction with the
highway safety objectives stated above will be the number of reported crashes per 100
million vehicle-miles operated and the number of highway accident– related fatalities per
100,000 US residents, both to be drawn from the Fatality Analysis Reporting System
maintained by the National Highway and Transportation Administration. In addition,
SMART objectives establish targets—levels on performance measures that programs or
agencies have identified to be achieved within the specified time period. For the measure
of reported crashes per 100 million vehicle miles traveled, the target identified above is 3
or lower, while the target for the number of highway accident fatalities per 100,000
population is 10 or lower.
The idea underlying SMART objectives is that the targets should be set at levels
that are both ambitious and yet realistic. While some targets call for maintaining current
or minimally improving performance levels, in the context of results-oriented
management and the drive to improved performance, it is desirable to set targets that are
relatively ambitious “stretch objectives” designed to challenge the program or
organization to find ways to make meaningful improvement in performance. Very modest
targets are not likely to encourage people to work harder and smarter to strengthen
performance significantly. More ambitious targets, particularly when there is clearly
strong commitment to them from higher levels of authority, can have a galvanizing effect
on people and motivate program managers and employees to “stretch the envelope” in a
quest to really make a difference. Yet the targets established by SMART objectives must
also be realistic and achievable in order to be productive. Overly aggressive targets that
are out of reach or beyond the grasp of a program to achieve within the time period
specified can be counterproductive because by definition, they amount to programming
failure. Such targets are highly likely to backfire and create disincentives for working
toward improved performance in the long run. Thus, finding a happy medium in
establishing targets often requires careful assessment and sound judgment.
The term performance standard is often used interchangeably with targets,
particularly when they refer to outcomes and reflect performance expectations that are
fairly constant over time. Consider the child support enforcement program operated by a
state’s department of human resources or social services. The mission of this program is
to help families rise or remain out of poverty, and reduce their potential dependency on
public assistance, through the systematic enforcement of noncustodial parents’
responsibility to provide financial support for their children. Figure 4.1 presents the logic
model for this program, working through three basic components designed to obligate
support payments by absentee parents, collect payments that are obligated, and assist
absentee parents, if necessary, to secure employment so that they are financially able to
make support payments. The logic moves through locating absentee parents, establishing
paternity when necessary, and obtaining court orders to obligate support payments, as
well as helping absentee parents to earn wages, but the bottom line is collecting payments
and disbursing them to custodial parents to ensure that children receive adequate financial
support.
The most salient output measures are the number of absentee parent searches
conducted, the number of paternity investigations completed, and the number of court
orders sought. The productivity measures are the number of noncustodial searches
conducted per locator and the number of training programs conducted per training staff
member, as well as the number of active cases maintained per child support enforcement
agent, although the latter might well be considered to be more of a workload measure.
The efficiency measures represent unit costs of such outputs as paternity investigations
conducted, accounts established, and training programs completed. The one service
quality indicator shown is actually a customer service indicator: the percentage of
custodial parents who report being satisfied with the assistance they have received.
Although many performance measurement systems establish SMART objectives
with targets to be achieved on each indicator, other systems purposefully do not do so.
Whether to set targets depends on the purpose of the measurement system and the
management philosophy in the organization. For example, many state transportation
departments have been in the forefront of the performance management movement
(Transportation Research Board, 2001) and the majority of them incorporate targets in
the measures they track. In contrast, the New Mexico State Highway and Transportation
Department took a different approach with respect to the approximately eighty indicators
of performance in seventeen key result areas covered in its Compass system, which
became the driving force of management and decision making in the department in the
early part of the past decade (Poister, 2004). In keeping with the continuous improvement
philosophy underlying the department’s quality improvement program, from which the
Compass evolved, the department preferred not to establish targets on these measures.
This policy was based on the belief that targets can have ceiling effects and actually
inhibit improvement rather than provide incentives to strengthen performance. Thus, the
implicit objective is to continuously improve performance on these measures over time.
The most common approach is to look at current performance levels on the
indicators of interest, along with the past trends leading up to these current levels, and
then set targets that represent some reasonable degree of improvement over current
performance. Current performance levels often provide an appropriate point of departure,
but in a less-than-stellar agency, they may underrepresent the possibilities, so the
question to ask is, “To what degree should we be able to improve above where we are
now?” Extrapolating on past trends, an agency might develop forecasting models for
projecting future performance levels based on a continuation of past trends and
assumptions regarding future values of key driving forces incorporated in the model,
including both external, contextual factors and program delivery. The performance levels
predicted by the model for future years, assuming a constant cause system, can then be
used as a point of departure in setting targets representing incremental or perhaps more
dramatic improvement in performance levels that the agency aspires to achieve going
forward. This approach analyzes the service delivery process, assesses the production
possibilities, and determines what level of performance can reasonably be expected,
given constraints on the system. The analysis, which might be performed for subunits and
then rolled up to the agency or program as a whole, should take into account any changes
in resource levels, intervention strategies, treatments, program design, service delivery
arrangements, or operations that might be expected to have an impact on overall
productivity. This production function approach works particularly well for setting output
targets to be achieved by a production process. It may be less helpful in setting
appropriate targets for real outcomes when precise relationships between outputs and
outcomes are not clearly understood.
Finally, setting appropriate targets may be informed by comparative performance
data on other similar agencies or programs. Benchmarking performance against other
entities, as discussed in chapter 14, can help identify norms for public service industries
as well as star performers in the field, which can be helpful in setting targets for a
particular program or agency. A program that finds itself performing considerably lower
than other similar programs, for example, might first set targets for itself based on other
programs that are somewhat higher in the rankings but not necessarily at the top, while an
agency that is already in the top quartile might set targets that approximate the
performance of the leading performers in the field. A major challenge in using the
benchmarking approach is to find truly comparable programs or agencies in the first
place, or to make adjustments for differences in operating conditions in interpreting the
performance of other entities as the basis for setting targets for a particular program or
agency.
Whichever approach is used in developing SMART objectives, the intent should
be to set targets that are both ambitious and realistic. Thus, agencies might be well
advised to set moderately challenging targets that will motivate managers and employees
to find ways to improve performance but refrain from going over the top in setting targets
that are unrealistically high, in which case they would be preordaining failure. Perhaps
more important, as they set targets for moderate performance improvements and then
attain those target levels, they can continue to set incrementally higher targets and ratchet
up meaningful performance improvements over time. Again, target setting may be
partially an analytical exercise, but to do it well also requires sound judgment of the
possibilities and constraints involved.
C. Service Standards
Complicating the lexicon surrounding goals, objectives, and targets is the term
standards. The term standards is often used interchangeably with targets, but to some
people, standards refer to more routine performance expectations that are fairly constant
over time, whereas targets may be changed more frequently as actual and potential trends
change over time. Performance standards, then, tend to relate to programmatic or agency
outcomes, whereas service standards refer more often to service delivery processes.
Service standards are specific performance criteria intended to be attained on an ongoing
basis. They usually refer to characteristics of the service delivery process, service quality,
or productivity in producing outputs. In some cases, service standards are distinct from a
program’s objectives, but probably more often, service standards and objectives are
synonymous or closely related. In any case, if there is not a clear sense about what a
program’s mission, goals, objectives, and perhaps service standards are, it is important to
clarify them before attempting to identify meaningful measures of its performance.
These standards might be considered the objectives of the program, or there might
be other objectives, such as increasing the percentage of customers indicating on
response cards that they were satisfied with the service they received to 85 percent during
the next year. Alternatively, if the program has only been achieving a “fill rate” of 80
percent, a key objective may be to raise it to 90 percent during the next year and achieve
the standard of 95 percent by the following year. Understanding what a program is
supposed to accomplish through clarifying mission, goals, objectives, service or
performance standards, and targets, however they are configured, can help tremendously
in identifying critical performance measures.
Some service standards focus on service quality rather than timeliness. For
instance, a state highway maintenance program may set a standard for ride quality that
calls for maintaining all of its roads on the National Highway System at a roughness level
at or below 120 on the international roughness index, while a local public transit system
may work hard to adhere to a service standard calling for buses to arrive at all regular bus
stops within plus or minus three minutes of scheduled arrival times. Similarly, an after-
care program run by a state’s juvenile justice program may have a policy that
caseworkers or counselors should have weekly face-toface meetings with all juveniles
discharged from juvenile boot camps during the first six months after the date of
discharge, an output-oriented service standard.
Frequently public agencies establish targets for adherence to service standards,
especially when they are failing to do so successfully, but are motivated to improve
performance in those areas. For example, a state transportation department may have a
service standard calling for highway capacity on major interregional corridors to be
sufficient so that traffic can move at the posted speed limit, but its performance
monitoring indicates that it is meeting this standard on only 55 percent of the mileage on
those interregional roads. In an effort to improve its performance on this standard, it
establishes a target that calls for increasing the percentage of that road mileage on which
traffic does move at the posted speed limit up to 65 percent by the end of the following
fiscal year and up to 75 percent over the next three fiscal years. Consider a state juvenile
justice after-care program whose performance monitoring reveals that only 40 percent of
the juveniles discharged from boot camps within the past six months are being contacted
at least once per week by program staff, the service standard that has been established for
the program. This may lead the program to establish a target to the effect that at least 50
percent of juveniles who have been discharged from boot camp in the past six months
will in fact be contacted by their case manager or other appropriate program staff weekly
over the next fiscal year.
D. Programmatic versus Managerial Goals and Objectives
To be useful, performance measures should focus on whatever kinds of results
managers want to accomplish. From a purist program evaluation– based perspective,
appropriate measures are usually seen as focusing on programmatic goals and objectives,
the real outcomes produced by programs and organizations out in the field. However,
from a practical managerial perspective, performance measures focusing on
implementation goals and the production of outputs are often equally important. Thus,
public and nonprofit organizations often combine programmatic or outcome-based goals
and objectives along with more managerial or outputbased goals and objectives in the
same management systems. Both programmatic and managerial objectives should be
stated as SMART objectives and tracked with appropriate performance measures. For
example, the programmatic objectives of a community crime prevention program might
be to reduce personal crimes by 20 percent and property crimes by 25 percent in one
year, along with the goal of having at least 90 percent of all residents feeling safe and
secure in their own neighborhoods. These outcomes could be monitored with basic
reported crime statistics and an annual neighborhood-based survey. More managerial
objectives might include the initial implementation of community policing activities
within the first six months and the startup of at least twenty-five neighborhood watch
groups within the first year. These outputs could be tracked through internal reporting
systems.
Given current performance levels in this particular agency, all four of these
objectives are considered to be ambitious yet realistic, and they are all SMART
objectives in terms of specifying the nature and magnitude of expected results within a
particular time period. In addition, straightforward performance indicators can be readily
operationalized for each of these objectives, along with the service standards, and
collectively they will provide management with a clear picture of the overall performance
of this workers’ compensation program.
Clearly one useful framework for identifying appropriate goals and objectives is
the kind of program logic models and associated performance measures discussed in
chapter 3. Public and nonprofit organizations frequently set goals for increasing or
improving the quality of outputs produced by a particular program, and they also set
goals focusing on increasing the volume of outcomes or altering characteristics of
outcomes produced, as well as changing the mix of outcomes produced by a program.
Similarly, goals might be established for improving the quality of products or services
delivered by a program or increasing the efficiency and productivity of service delivery
processes, and other goals might be established for increasing customer satisfaction with
the services they receive from the program or the outcomes they experience as a result of
participating in a program or receiving services from a program. In addition, goals might
be defined in terms of the priority populations or target groups to be reached by the
program.
However, performance management systems often focus on an organization’s
performance rather than that of particular operating programs. Some public organizations
have full responsibility for a single program, while others, particularly larger department-
level agencies, are responsible for multiple programs and services, and agencies may also
share responsibility for some programs with other agencies. In any case, the goals that are
important to an agency almost always include some that are directly related to programs,
but the agencies are also likely to have other organizational or nonmission-oriented goals
concerning development, management capacity, technology, external support, and so
forth as well. In their book on the balanced scorecard as a framework for strategic
planning, as discussed in chapter 8, Kaplan and Norton (1996) proposed that private
firms should be establishing goals and attendant performance measures not only from the
perspective of financial performance or the bottom line, but also goals with respect to
customers, business processes, and learning and growth.
Many organizations in the public sector have developed their own balanced
scorecards, and most adopt the same four perspectives— financial, customer, internal
processes, and learning and growth— although they tend to identify the customer or
citizen perspective as the most important goals and establish goals in the other three
perspectives to support achievement of those customer- or citizen-oriented goals (Niven,
2003). In a similar vein, Boyne and Walker (2004) identified five “action areas” that
constitute strategy content in the public sector— markets, services, revenues, internal
organization, and external organization—and this model also provides a useful
framework for goal setting in the public sector.
Similar kinds of performance frameworks have been developed for the nonprofit
sector focusing on such perspectives as social mission achievement, program
effectiveness, and participant-centered outcomes in addition to organization and
management capacity and external support (Moore, 2003; Sowa, Selden, & Sandfort,
2004; Urban Institute, 2006). However, public and nonprofit organizations tend to differ
significantly with respect to emphasis on goals and performance measures focusing on
institutional support and revenues. While some public organizations such as toll roads,
public utilities, and regulatory agencies earning revenue through fees collected often
place substantial emphasis on goals and performance measures concerning revenues,
particularly those that operate in competitive markets such as public transit agencies, in
most public agencies financial resources are thought of as a given from dedicated revenue
sources or budget allocations, falling on the input side of the performance framework
rather than as results. Thus, as discussed in chapter 3, resource measures are typically
used in computing efficiency, productivity, and cost-effectiveness measures but are not
often considered as performance measures in their own right. However, as self-created
entities rather than government agencies with semiguaranteed financial revenues,
nonprofit organizations must of necessity secure their own revenue—from members or
regular contributors, charitable donors, and government grants or contracts in addition to
paying customers—in order to ensure their continued ability to pursue their social
missions. Thus, institutional support, and especially revenue and resources, tend to figure
much more prominently in the goals and performance measures set by nonprofit
organizations as compared with public agencies.
Obviously managers need to forge close linkages between goals and objectives,
on one hand, and performance measures, on the other. It is critical to monitor measures of
performance in terms of accomplishing outcome-oriented, programmatic objectives, but
often it is important to track measures focused on the achievement of more managerially
oriented objectives as well. In some instances, goals and objectives are stated in terms of
the general kinds of results intended to be produced by programmatic activity, and then
performance indicators must be developed to track their achievement. In other cases,
however, the objectives themselves are defined in terms of the measures that will be used
to track results. Sometimes performance standards or service standards are established
and tracked independently, while at other times, objectives or targets are set in terms of
improving performance on those standards. Although there is not one right way to do it,
the bottom line for results-oriented managers is to clearly define intended results through
some mix of goals, objectives, standards, and targets and then track performance
measures that are as closely aligned as possible with these results.
E. Operational Indicators
Once you have identified a program’s intended outcomes and other performance
criteria, how do you develop good measures of these things? What do useful performance
indicators look like? And what are the characteristics of effective sets of performance
measures? Where do you find the data to operationalize performance indicators? In order
for monitoring systems to convey meaningful information about program performance,
the measures used must be appropriate and meet the tests of sound measurement
principles. This chapter begins to focus on the how of performance measurement: how to
define measures of effectiveness, efficiency, productivity, quality, client satisfaction, and
so forth that are valid, reliable, and truly useful. In thinking about defining specific
performance measures, we first have to make a distinction between the measures
themselves and the operational indicators used to represent them. The measures of
performance that are identified through program logic models or goal structures, or
simply by decisions by managers or analysts or suggestions from other stakeholders,
provide a general sense of what the measures will focus on but not precisely how they
will be measured or observed or computed. Operational indicators redefine the measures
in terms of the data sources that will be accessed, the observations that will be made, and
criteria for what counts versus what does not count in operationalizing the measures so
that they can be monitored over time. The term metrics is often used in the field of
performance measurement to refer to the operational indicators that define the way a
measure is specified or observed and the categories or units on the sale used to
operationalize it.
Many, or at least most, performance measures—variable names in the language of
statistical analysis—such as recidivism in juvenile justice programs, cycle time in an
investigations process, or clients placed in employment situations by a job training
program can be operationalized in multiple ways, which creates a need for clear
definitions and rules regarding how a measure will be operationalized. Consider the
performance measure commonly referred to as the student-faculty ratio in the field of
higher education. What category of students should be included in the computation of this
measure: full-time students only or part-time students as well, undergraduate or graduate
students or both, students on campus or those taking courses online? What about students
who are matriculating in an academic degree program but are not enrolled in any courses
this particular semester? Similarly, what categories of faculty members should be
included in the ratio: teaching faculty only, research faculty, part-time versus full-time
faculty, visiting faculty, individuals with faculty status who are in administrative
positions, or faculty members on leave this semester?
As another example, how do we specify a measure of crime rate as an outcome of
a crime prevention program? Should it focus only on personal crimes or property crimes
—or both? Should we use reported crime statistics, or should we conduct victimization
surveys to compute crime rates? The alternative approaches to operationalizing either the
student-faculty ratio or crime rates are likely to generate differing results and different
impressions of what the performance of such programs actually looks like. Thus, it is
critical to define operational indicators carefully and precisely in order to ensure that the
various audiences for whom a measurement system is intended will have a clear
understanding with respect to what is actually being measured in a given instance. Before
we discuss the challenges of measurement issues, it may be helpful to picture the
numerical or statistical forms in which performance indicators can be specified. The most
common of these statistical formats— raw numbers, averages, percentages, ratios, rates,
and indexes—provide options for defining indicators that best represent the performance
dimensions to be measured.
Although some authorities on the subject might disagree, raw numbers often
provide the most straightforward portrayal of certain performance measures. For
example, program outputs and output targets are usually specified in raw numbers, such
as the miles of shoulders to be regraded by a county highway maintenance program, the
number of books circulated by a public library system, or the number of claims to be
cleared each month by a state government’s disability determination unit. Beyond
outputs, effectiveness measures often track program outcomes in the form of raw
numbers. For instance, a local economic development agency may track the number of
new jobs created in the county or the net gain or loss in jobs at the end of a year, and state
environmental protection agencies monitor the number of ozone action alert days in their
metropolitan areas.
Using raw numbers to measure outputs and outcomes has the advantage of
portraying the actual scale of operations and impacts, and this is often what line managers
are the most concerned with. In addition to programming and monitoring the number of
vehicle-miles and vehiclehours operated each month, for instance, a public transit system
in a small urban area might set as a key marketing objective the attainment of a total
ridership of more than 2 million passenger trips for the coming year. Although it might
also be useful to measure the number of passenger trips per vehicle-mile or per vehicle-
hour, the outcome measure of principal interest will be the raw number of passenger trips
carried for the year. In fact, the transit manager might also examine the seasonal patterns
over the past few years and then prorate the objective of 2 million passengers for the year
into numerical ridership targets for each month. He or she would then track the number of
passenger trips each month against those specific targets as the system’s most direct
outcome measure.
Sometimes statistical averages can be used to summarize performance data and
provide a clearer picture than raw numbers would. In an effort to improve customer
service, for example, a state department of motor vehicles might monitor the mean
average number of days required to process vehicle registration renewals by mail; a local
public school system may track the average staff development hours its teachers engage
in. Similarly, one measure of the effectiveness of an employment services program might
be the median weekly wages earned by former clients who have entered the workforce; a
public university system might track the effectiveness of its recruiting efforts by
monitoring the median verbal and mathematics SAT scores of each year’s freshman
class. Such averages are more readily interpreted because they express the measure on a
“typical case” basis rather than in the aggregate.
Even the particular type average being employed could make a significant
difference in the results generated, as well as the responses suggested by a performance
measurement system. Suppose, for example, that the top priority in a state’s
transportation department is to improve ride quality on its highways and that the metric
being used to track progress in this area is the median average score on an index of
pavement roughness of all these roads, which is heavily skewed to the high side (with a
low percentage of the roads showing very high roughness levels). If the district and area
maintenance managers across the state are strongly incentivized to improve performance
on this measure to the greatest extent possible, they might well focus on making fairly
modest improvements on a large number of those road segments with current roughness
scores that are slightly to somewhat higher than the median. If instead the operationalized
indicator is specified as the mean average roughness score and there are strong incentives
in place to improve performance on that measure, the maintenance managers would be
much more likely to focus attention on making dramatic improvements in pavement
smoothness on the relatively few roads that are in the high end of the skew of the
distribution. This same scenario could well apply to heavily skewed distributions to the
high side of a distribution of cycle time of investigations conducted by a regulatory
agency or, in the reverse, to a distribution of standardized test scores in a local school
district that is heavily skewed to the low side.
Percentages, rates, and ratios are relational statistics that can often express
performance measures in more meaningful context. Percentages can be especially useful
in conveying the number of instances with desired outcomes, or “successes,” as a share of
a total number of cases—for example, the percentage of teen mothers in a parenting
program who deliver healthy babies, the percentage of clients of a nonprofit agency
working with persons with mental disabilities who are placed in competitive
employment, and the percentage of youths discharged from juvenile justice programs
who don’t recidivate back into the criminal justice system within six months. Percentages
can often be more definitive performance measures than averages, particularly when
service standards or performance targets have been established. For instance, tracking the
average number of days required to process vehicle registration renewals by mail can be a
useful measure of service quality, but it doesn’t provide an indication of the number of
customers who do, or do not, receive satisfactory turnaround time. If a standard is set,
however, say to process vehicle renewals within three working days, then the percentage
of renewals actually processed within three working days is a much more informative
measure of performance.
Expressing performance measures as rates helps put performance in perspective
by relating it to some contextual measure representing exposure or potential. For
instance, a neighborhood watch program created to reduce crime in inner-city
neighborhoods might track the raw numbers of personal and property crimes reported
from one year to the next. However, to interpret crime trends in the context of population
size, the national Uniform Crime Reporting System tracks these statistics in terms of the
number of homicides, assaults, robberies, burglaries, automobile thefts, and so on
reported per 1,000 residents in a local jurisdiction. Similarly, the effectiveness of a birth
control program in an overpopulated country with an underdeveloped economy might be
monitored in terms of the number of births recorded per 1,000 women of childbearing
age. Accident rates are usually measured in terms of exposure factors, such as the number
of highway traffic accidents per 100 million vehicle-miles operated or the number of
commercial airliner collisions per 100 million passenger miles flown. In monitoring the
adequacy of health care resources in local communities, the Federal Health Care
Financing Administration looks at such measures as the number of physicians per 1,000
population, the number of hospitals per 100,000 population, and the number of hospital
beds per 1,000 population. Tracking such measures as rates helps interpret performance
in a meaningful context.
The use of ratios is prevalent in performance measurement systems because they
too express some performance dimension relative to some particular base. In particular,
ratios lend themselves to efficiency, productivity, and cost-effectiveness measures
because they are all defined in terms of input-output relationships. Operating efficiency is
usually measured in terms of unit costs—for example, the cost per vehicle-mile in a
transit system, the cost per detoxification procedure completed in a crisis stabilization
unit, the cost per course conducted by a teen parenting program, and the cost per
investigation completed by the US Environmental Protection Agency. Similarly,
productivity could be measured by such ratios as the tons of refuse collected per crew-
day, the number of cases cleared per disability adjudicator, flight segments handled per
air traffic controller, and the number of youths counseled per juvenile justice counselor.
Costeffectiveness measures are expressed in such ratios as the cost per client placed in
competitive employment or the parenting program cost per healthy infant delivered.
Percentages, rates, and ratios are often preferred because they express some dimension of
program performance within a relevant context. More important, however, they are useful
because as relational measures, they standardize the measure in terms of some basic
factor, which helps control for that factor in interpreting the results. As will be seen in
chapter 6, standardizing performance measures by expressing them as percentages, rates,
and ratios also helps afford valid comparisons of performance over time, across
subgroups, or between a particular agency and other similar agencies.
An index is a composite measure that is computed by combining multiple
measures or constituent variables into a single summary measure. For example, one way
the Federal Reserve Board monitors the effectiveness of its monetary policies in
preventing excessive inflation is the consumer price index (CPI), the calculated cost of
purchasing a standard set of household consumer items in various markets around the
country. Because indexes are derived by combining other indicators, scores, or repeated
measures into a new scale, some of them seem quite abstract, but ranges or categories are
often defined to help interpret the practical meaning of different scale values. Researchers
frequently develop and use indexes to measure multidimensional concepts such as
psychological well-being or quality of life. Thus, they can be particularly useful for
measuring outcomes in programmatic areas whose intended results are complex, as
illustrated in table 5.1. However, they can also apply to service quality and customer
satisfaction as well.
The Adaptive Behavior Scales (ABS) are standardized scales developed to assess
the level of functioning of individuals with mental disabilities in two areas: personal
independence and responsibility and social behaviors (Dixon, 2007). They are often used
by public and nonprofit agencies working with people with mental disabilities to assess
their needs and monitor the impact of various programs on clients’ level of functioning.
The overall scale consists of eighteen domains— for example, independent functioning,
physical development, language development, self-direction, self-abusive behavior, and
disturbing interpersonal behavior—which in turn have subdomains that are represented
by a series of items. For example, one subdomain of the independent functioning domain
concerns eating, and this is measured by four items focusing on the use of table utensils,
eating in public, drinking, and table manners. In human services programs, psychological
scales are often used to assess client outcomes. For example, the Duke Activity Status
Index is used to measure a client’s functional capacity. It is based on answers given on a
self-administered questionnaire containing twelve questions regarding an individual’s
ability to engage in or perform activities that are considered to be a normal part of daily
living, as shown in table 5.2. The items are weighted according to the difficulty or energy
required in each activity, and the responses are summed to an overall score that ranges
from 0 to 58.2.
Many performance monitoring systems include a mix of measures expressed in
the various forms we’ve discussed here. For example, table 5.3 illustrates a sample of
conventional measures used in tracking the performance of highway maintenance
programs. They include raw numbers of resource materials and outputs, ratios for
efficiency and productivity indicators, mean average quality assurance scores, median
pavement quality index scores, percentages of satisfactory roads and satisfied customers,
and accident rates related to road conditions. While managers, analysts, and consultants
involved in developing performance measurement systems often define indicators on
their own, working directly from logic models or goal structures, there are often
resources available that share information on measures used in a program area that might
well provide a starting point for developing more customized indictors in a given
instance. Such sources often provide information on the purpose and usefulness, and
strengths and weaknesses, of these measures.
F. Sources
The data used in performance measurement systems come from a wide variety of
sources, and this has implications regarding the cost and effort of data collection and
processing, as well as quality and appropriateness. In some cases, appropriate data exist
in files or systems that are used and maintained for other purposes. They can be extracted
or used for performance monitoring as well, whereas the data for other measures will
have to be collected specifically for the purpose of performance measurement. With
regard to the highway maintenance measures shown in table 5.3, information on the
gallons of patching material applied may be readily available from the highway
department’s inventory control system, and data on the number of lane-miles resurfaced,
the miles of shoulders graded, and the actual task and production hours taken to complete
these activities may be recorded in its maintenance management system. The cost of this
work is tracked in the department’s activity-based accounting system. The quality
assurance scores are generated by teams of inspectors who audit a sample of completed
maintenance jobs to assess compliance with prescribed procedures. The pavement quality
index and the percentage of roads in compliance with national American Association of
State Highway and Transportation Officials (AASHTO) standards require a combination
of mechanical measurements and physical inspection of highway condition and
deficiencies. The percentage of motorists rating the roads as satisfactory may require a
periodic mail-out survey of a sample of registered drivers. The accident rate data can
probably be extracted from a data file on recorded traffic accidents maintained by the
state police.
Sometimes existing databases that are maintained by agencies for other purposes
can meet selected performance measurement needs of particular programs. Many federal
agencies maintain compilations of data on demographics, housing, crime, transportation,
the economy, health, education, and the environment that may lend themselves to
tracking the performance of a particular program. Many state government agencies and
some nonprofit organizations maintain similar kinds of statistical databases, and a variety
of ongoing social surveys and citizen polls also produce data that might be useful as
performance measures.
By far the most common source of performance data consists of agency records.
Public and nonprofit agencies responsible for managing programs and delivering services
tend to store transactional data that record the flow of cases through a program, the
number of clients served, the number of projects completed, the number of services
provided, treatment modules completed, staff-client interactions documented, referrals
made, and so on. Much of this focuses on service delivery and outputs, but other
transactional data maintained in agency records relate further down the output chain
regarding the disposition of cases, results achieved, or numbers of complaints received,
for instance. In addition to residing in management information systems, these kinds of
data are also found in service requests, activity logs, case logs, production records,
records of permits issued and revoked, complaint files, incident reports, claims
processing systems, and treatment and follow-up records, among other sources.
Beyond working with transactional data relating specifically to particular
programs, you can also tap administrative data concerning personnel and expenditures,
for example, to operationalize performance data. In some cases, these administrative data
may also be housed in the same programmatic agencies that are responsible for service
delivery, but often they reside in central staff support units, such as personnel agencies,
training divisions, budget offices, finance departments, accounting divisions, and
planning and evaluation units. Sources of such administrative data might include time,
attendance, and salary reports, as well as budget and accounting systems and financial,
performance, and compliance audits.
In some program areas where the outcomes are expected to materialize outside the
agency and perhaps well after a program has been completed, it is necessary to make
follow-up contacts with clients to track effectiveness. Often this can be accomplished
through the context of follow-up services. For example, after juvenile offenders are
released from boot camp programs operated by a state’s department of juvenile justice,
the department may also provide after-care services in which counselors work with these
youths to help them readjust to their home or community settings; encourage them to
engage seriously in school, work, or other wholesome activities; and try to help them stay
away from further criminal activity. Through the follow-up contacts, the counselors keep
track of the juveniles’ status in terms of involvement in gainful activity versus recidivism.
Many times, measuring outcomes requires some type of direct observation, by
means of mechanical instruments or personal inspections, in contexts other than follow-
up client contacts. For example, state transportation departments use various kinds of
mechanical and electronic equipment to measure the condition and surface quality of the
highways they maintain, and environmental agencies use sophisticated measuring devices
to monitor air quality and water quality. In other cases, trained observers armed with
rating forms make direct physical inspections to obtain performance data. For instance,
local public works departments sometimes use trained observers to assess the condition
of city streets, sanitation departments may use trained observers to monitor the
cleanliness of streets and alleys, and transit authorities often use them to check the on-
time performance of the buses.
Some performance monitoring data come from a particular kind of direct
observation: clinical examinations. Physicians, psychiatrists, psychologists, occupational
therapists, speech therapists, and other professionals may all be involved in conducting
clinical examinations of program clients or other individuals on an ongoing basis,
generating streams of data that might feed into performance measurement systems. For
example, data from medical diagnoses or evaluations may be useful not only in tracking
the performance of health care programs but also in monitoring the effectiveness of crisis
stabilization units, teen parenting programs, vocational rehabilitation programs, disability
programs, and workers’ compensation return-to-work programs, among others. Similarly,
data from psychological evaluations might be useful as performance measures in
correctional facilities, drug and alcohol abuse programs, behavioralshaping programs for
persons with mental disabilities, and violence reduction programs in public schools.
Tests are instruments designed to measure individuals’ knowledge in a certain
area or their skill level in performing certain tasks. Obviously these are most relevant for
educational programs, as local public schools routinely use classroom tests to gauge
students’ learning or scholastic achievement. In addition, some states use uniform
“Regents”-type examinations, and a plethora of standardized exams that are on a
widespread basis, which facilitate tracking educational performance on a local, state, or
national level and allow individual schools or school districts to benchmark themselves
against others or national trends. Beyond education programs, testing is used to obtain
performance data in a wide variety of other kinds of training programs, generating
measures ranging from the job skills of persons working in sheltered workshops to the
flying skills of air force pilots and fitness ratings of police officers.
As will be discussed in chapter 13, public and nonprofit agencies also employ a
wide range of personal interview, telephone, mail-out, and other self-administered
surveys to generate performance data, most often focusing on feedback regarding service
quality, program effectiveness, and customer satisfaction. In addition to surveys of clients
and former clients are surveys of customers, service providers or contractors, other
stakeholders, citizens or the public at large, and even agency employees. However,
survey data are highly reactive, and great care is needed in the design and conduct of
surveys to ensure high-quality, objective feedback. One particular form of survey that is
becoming more prevalent as a source of performance data is the customer response card.
These are usually brief survey cards containing only a handful of straightforward
questions that are given to customers at the point of service delivery, or shortly after, to
monitor customers’ satisfaction with the service they received in that particular instance.
Such response cards might be given out, for example, to persons who just finished
renewing their driver’s license, individuals just about to be discharged from a crisis
stabilization unit, child support enforcement clients who have just made a visit to their
local office, or corporate representatives who have just attended a seminar about how
their firms can do business with state government. These response cards not only serve to
identify and, one hopes, resolve immediate service delivery problems but also generate
data that in the aggregate can be useful in monitoring service quality and customer
satisfaction with a program over time.
Although the vast majority of the measures used in performance monitoring
systems come from the conventional sources we have already discussed, in some cases it
is desirable or necessary to design special measurement instruments to gauge the
effectiveness of a particular program. For example, the national Keep America Beautiful
program and its state and local affiliates use the photometric index developed by the
American Public Works Association to monitor volumes of litter in local communities.
The photometric index is operationalized by taking color slides of a sample of ninety-six-
square-foot sites in areas that are representative of the community in terms of income and
land use. The specific kinds of sites include street curb fronts, sidewalks, vacant lots,
parking lots, dumpster sites, loading docks, commercial storage areas, and possibly rural
roads, beaches, and parks. There may be on the order of 120 such sites in the sample for
one community, and the same sites are photographed each year.
G. Validity and Reliability
As we have seen, for some performance measures good data may be readily at
hand, whereas other measures may require follow-up observation, surveys, or other
specially designed data collection procedures. Although available data sources can
obviously be advantageous in terms of time, effort, and cost, readily available data are
not always good data— but they aren’t always poor quality either. From a
methodological point of view, good data are those with a high degree of validity and
reliability—that is, they are unbiased indicators that are appropriate measures of
performance and provide a reasonable level of consistency, precision, and statistical
reliability. There are numerous good sources on the process of developing and testing
measures from a methodological or research perspective, such as those by Shulz and
Whitney (2005) and DeVellis.
Performance indicators are measures defined operationally in terms of how the
measure is taken or the data are collected. For example, the operational indicator for the
number of students entering a state’s university system each year might be the number
recorded as having enrolled in three or more classes for the first time during the
preceding academic year by the registrar’s office at each of the institutions in the system.
Similarly, the operational indicator for the number of passengers carried by an urban
transit system might be the number counted by automatic registering fare boxes; the
percentage of customers who are satisfied with the state patrol’s process for renewing
drivers’ licenses might be measured by the percentage who check off “satisfied” or “very
satisfied” on response cards that are handed out to people as they complete the process.
The reliability of such performance indicators is a matter of how objective,
precise, and dependable they are. Consistency over time is one aspect of reliability; if the
same measuring instrument—for example, a survey, test, or other observation—is used
on the same subject or subjects in the same way at different times, if the subject has really
not changed on the dimension of interest over that period of time, the measurement
should yield the same results in order to be considered reliable. To the extent that it
produces different results, the indicator lacks precision, producing a range of estimates of
the true value of the measure rather than a single value or only slight variation around it.
For instance, if repeated queries to a university registrar’s office asking how many
students are enrolled in classes during the current semester yield a different number every
time, the measure lacks consistency or dependability and thus is not very reliable. The
range of responses might provide an indication of roughly how many students are
enrolled in classes, but it certainly is not a precise indicator. With survey instruments and
other kinds of indicators observed on a sample of cases, clients in a program, for
example, test-retest reliability can be assessed by using the instrument at two points in
time on the same sample and running a correlation between the two sets of data. The
closer the correlation coefficient is to 1.0, the greater the reliability of the measure.
A lack of interrater reliability also presents problems in performance data. If a
number of trained observers rating the condition of city streets look at the same section of
street at the same time, using the same procedures, definitions, categories, and rating
forms, yet the rating they come up with varies substantially from observer to observer,
this measure of street condition clearly is not very reliable. Although the observers have
been trained to use this instrument the same way, the actual ratings that result appear to
be based more on the subjective impressions of the individual raters than on the objective
application of the standard criteria. Such problems with interrater reliability can occur
whenever the indicator is operationalized by different individuals observing cases and
making judgments, as might the case, for instance, when housing inspectors determine
the percentage of dwelling units that meet code requirements, when workers’
compensation examiners determine the percentage of employees injured on the job who
require longer-term medical benefits, or when staff psychologists rate the ability of
mildly and moderately mentally disabled clients of a nonprofit agency to function at a
higher level of independence. If they are not applying the measuring instrument, making
observations, or counting things the same way, the data will lack interrater reliability.
Reliability problems may also occur when the performance data are reported from
different sources or locations. For example, when the same indicators are reported
regularly by various district offices, subordinate work units, or project sites, there may be
differences in the way things are recorded that will cause reliability problems when the
data are aggregated and reported out at the department or program level. As will be seen
in chapter 14, the potential for this kind of reliability problem can be magnified
geometrically when data on the same measures are reported by separate agencies or
programs in a comparative measurement process to benchmark the performance of any
one agencies against the field at large. If these agencies are operationalizing the
indicators differently, even though they assume they are reporting on the same measures,
the data lack reliability and comparisons are likely to be meaningless or misleading.
From a measurement perspective, the perfect performance indicator may never
exist because there is always the possibility of some error in the measurement process. To
the extent that the error in a measure is random and unbiased in direction, this is a
reliability problem. Although quality assurance processes need to be built into data
processing procedures, there is always a chance of accidental errors in data reporting,
coding, and tabulating, and this creates reliability problems. For example, state child
support enforcement programs track the percentage of noncustodial parents who are
delinquent in making obligated payments, and computing this percentage would seem to
be a simple matter. At any given time, the parent is either up-to-date or delinquent in
making these payments. However, the information on the thousands of cases recorded in
the centralized database for this program pertaining to numbers of children in households,
establishment of paternity, obligation of payments, and current status comes from local
offices and a variety of other sources in piecemeal fashion. Although up-to-date accuracy
is critical in maintaining these records, errors are made and slippage in reporting does
occur, and the actual accounts are likely to be off the mark a little (or maybe a lot). Thus,
in a system that tracks this indicator monthly, the computed percentage of delinquent
parents may overstate the rate of delinquency some months and understate it other
months. Although there is no systematic tendency to overrepresent or underrepresent the
percentage of delinquent parents, this indicator will not be highly dependable or reliable.
Whereas reliability is a matter of objectivity and precision, the validity of a
performance measure concerns its appropriateness, that is, the extent to which an
indicator is directly related to and representative of the performance dimension of
interest. If a proposed indicator is largely irrelevant or only tangentially related to the
desired outcome of a particular program, then it will not provide a valid indication of that
program’s effectiveness. For example, scores on the verbal portion of the SATs have
sometimes been used as a surrogate indicator of the writing ability of twelfth graders in
public schools, but the focus of these tests is really on vocabulary and reading
comprehension, which are relevant but only partially indicative of writing capabilities. In
contrast, the more recently developed National Assessment of Educational Progress test
in writing provides a much more direct indicator of students’ ability to articulate points in
writing and to write effective, fully developed responses to questions designed
specifically to test their writing competence.
Consider the validity of unemployment statistics, for instance. The indicator of
the unemployment rate in the United States reported by the Bureau of Labor Statistics
each month is based on surveys of sixty thousand households concerning the employment
status of household members over sixteen years of age. Those who indicate that they do
not have jobs but have looked for work during the past four weeks are classified as being
unemployed; those who indicate that they are working in full-time or part-time jobs are
considered to be employed. The unemployment rate is computed as the number of those
classified as unemployed taken as a percentage of the total household members
considered to be in the labor force as represented by the number of employed plus
unemployed individuals. The validity of this standard measure is often questioned,
however, principally because it underestimates actual unemployment levels by not
including unemployed individuals who have looked for work sometime in the past year
but not in the past four weeks, or those who may be unemployed for much longer periods
of time but have given up looking for work.
As another example, the aim of a metropolitan transit authority’s welfare-to-work
initiative might be to facilitate moving employable individuals from dependence on
welfare to regular employment by providing access to work sites through additional
transportation services. As possible measures of effectiveness, however, the estimated
number of homeless individuals in the area would be largely irrelevant, and the total
number of employed persons and the average median income in the metropolitan area are
subject to whole hosts of factors and would be only marginally sensitive to the welfare-
to-work initiative. More relevant measures might focus on the number of individuals
reported by the welfare agency to have been moved off the welfare rolls, the number of
“third-shift” positions reported as filled by manufacturing plants and other employers, or
the number of passenger trips made on bus trips that have been instituted as part of the
welfare-to-work initiative. However, each of these measures still falls short as an
indicator of the number of individuals who were formerly without jobs and dependent on
welfare who now have jobs by virtue of being able to get to and from work on the transit
system.
Most proposed performance measures tend to be at least somewhat appropriate
and relevant to the program being monitored, but the issue of validity often boils down to
the extent to which they provide fair, unbiased indicators of the performance dimension
of interest. Whereas reliability problems result from random error in the measurement
process, validity problems arise when there is systematic bias in the measurement
process, producing a systematic tendency to overestimate or to underestimate program
performance. For instance, crime prevention programs may use officially reported crime
rates as the principal outcome measure, but as is well known, many crimes are not
reported to the police for a variety of reasons. Thus, these reported crime rates tend to
underestimate the number of crimes committed in a given area during a particular time
period. The percentage of total crimes reported as “solved” by a local police department
would systematically overstate the effectiveness of the police if it includes cases that
were initially recorded as crimes and subsequently determined not to constitute crimes
but were still carried on the books labeled as “solved” crimes. An alternative would be to
conduct victimization surveys to estimate the extent to which crimes occur, but for any
number of reasons, respondents may not supply candid or accurate information, which
could cause validity problems or reliability problems, or both.
For many human service programs, it is difficult to follow clients after they leave
the program, but that is often when the real outcomes occur. The crisis stabilization unit
observes consumers only while they are actually short-term residents of the facility, and
thus it cannot track whether they continue to take prescribed medications faithfully, begin
to use drugs or alcohol again, or continue participating in long-term care programs.
Appropriate measures of effectiveness are not difficult to define in this case, but
operationalizing them through systematic client follow-up would require significant
additional staff, time, and effort that is probably better invested in service delivery than in
performance measurement.
As another example, a teen mother parenting program can track clients’
participation in the training sessions, but it will have to stay in touch with all program
completers in order to determine the percentage who deliver healthy babies, babies of
normal birth weight, babies free from HIV, and so on. But what about the quality of
parental care given during the first year of infants’ lives? Consider the options for
tracking the extent to which the teen mothers provide the kind of care for their babies that
is imparted by the training program. Periodic telephone or mail-out surveys of the new
mothers could be conducted, but in at least some cases, their responses are likely to be
biased in terms of presenting a more favorable picture of reality. Alternatively, trained
professionals could make periodic follow-up visits to the clients’ homes, primarily to
help the mothers with any problems they are experiencing. By talking with the mothers
and observing the infants in their own households, they could also make assessments of
the adequacy of care given. This would be feasible if the program design includes follow-
up visits to provide further support, and it would probably provide a more satisfactory
indicator even though some of the mothers might be on their best behavior during these
short visits, possibly leading to more positive assessments that overstate the quality of
care given to the infants on a regular basis.
Researchers and others are continually working to improve the kinds of indicators
that are used in monitoring performance. For example, homelessness is a major concern
in many urban areas in the United States, as well as many other countries, but it is
difficult to compute with any assurance the percentage of homeless individuals in a local
area who are being served to some degree by homeless shelters run by public or nonprofit
organizations, for instance, or even determine whether the number of homeless
individuals has been increasing or decreasing over time. The various sources that might
shed light on the number of homeless people living in a community are likely to count
things differently; by the very nature of the problem, it may be difficult or impossible to
find or identify many individuals who are homeless; and when homeless people are in
fact interviewed, their memory of dates may be poor and fade as time from a period of
homelessness passes. One response to this issue has been the development of the
residential follow-back calendar to help subjects improve their recall of the number of
days they have been homeless over the past year by (1) taking more time to remember,
(2) decomposing a class of events into subclasses, (3) recalling events in reverse
chronology, and (4) listing boundaries or landmarks to assist accurate recall (Tsemberis,
McHog, Williams, Hanrahan, & Stefancic, 2007). This is a good example of the ongoing
efforts to refine indicators in order to strengthen the validity of performance measures in
many policy and program areas.
H. Common Measurement Problems
In working through the challenge of defining useful operational indicators, system
designers should always anticipate likely problems and try to avoid or circumvent them.
Common problems that can jeopardize reliability or validity, or both, include
noncomparable data, tenuously related proximate measures, tendencies to under- or
overreport data, poor instrument design, observer bias, instrument decay, reactive
measurement, nonresponse bias, and cheating. Whenever data are entered into the system
in a decentralized process, noncomparability of data is a possibility. Although uniform
data collection procedures are prescribed, there is no automatic guarantee that they will
be implemented exactly the same way from work station to work station or from site to
site. This can be a problem within a single agency or program, as people responsible for
data input from parallel offices, branches, or work units find their own ways to expedite
the process in the press of heavy workloads, and they may end up counting things
differently from one another. Thus, in a large agency with multiple data entry sites, care
must be taken to ensure uniform data entry.
In large agencies delivering programs through a decentralized structure, for
example, a state human services agency with 104 local offices, the central office may
wish to track certain measures in order to compare the performance of local offices, or it
may want to roll up the data to track performance on a statewide basis. Particularly if the
local offices operate with a fair degree of autonomy, there may be significant
inconsistencies in how the indicator is operationalized from one local office to the next.
This could jeopardize the validity of comparisons among the local offices as well as the
statewide data. The probability of noncomparable data is often greater in state and federal
grant programs, when the data input is done by the individual grantees—local
government agencies or nonprofit organizations—who, again, may set up somewhat
different processes for doing so.
When it is difficult to define direct indicators of program performance or is not
practical to operationalize them, it is often possible to use proximate measures instead.
Proximate measures are indicators that are thought to be approximately equivalent to
more direct measures of performance. In effect, proximate measures are less direct
indicators that are assumed to have some degree of correlational or predictive validity.
For example, records of customer complaints are often used as an indicator of customer
satisfaction with a particular program. Actually, customer complaints are an indicator of
dissatisfaction, whereas customer satisfaction is usually thought of as a much broader
concept. Nevertheless, in the absence of good customer feedback using surveys, response
cards, or focus groups, data on complaints often fill in as proximate measures for
customer satisfaction. Similarly, the commonly stated purposes of local public transit
systems are to meet the mobility needs of individuals who don’t have access to private
means of transportation and to reduce the use of private automobiles in cities by
providing a competitive alternative. Transit systems rarely track measures of these
intended outcomes directly. Instead, they monitor overall passenger trips as a proximate
measure that they believe to be correlated with these outcomes.
Sometimes when it is difficult to obtain real measures of program effectiveness,
monitoring systems rely on indicators of outputs or initial outcomes as proximate
measures of longer-term outcomes. For example, a state department of administrative
services may provide a number of support services, such as vehicle rentals, office
supplies, and printing services to the other operating agencies of state government. The
real impact of these services would be measured by the extent to which they enable these
operating departments, their customers, to perform their functions more effectively and
efficiently. However, the performance measures used by these other agencies are unlikely
to be at all sensitive to the marginal contribution of the support services. Thus, the
department of administrative services might well just monitor indicators of output and
service quality on the assumption that if the line agencies are using these support services
and are satisfied with them, then the services are in fact contributing to higher
performance levels on the part of these other agencies.
Although proximate measures can often be useful, validity problems emerge
when they are only tenuously related to the performance criteria of interest. Consider for
a moment a municipal government’s neighborhood revitalization program that is trying to
encourage the construction of infill residential and small business developments in order
to strengthen the economic viability of target areas within the city limits. The most direct
indicator of the success of the program might be the number of such housing and small
business units constructed in those areas, but a leading indicator that would be expected
to point in that direction would be the number of units for which building permits have
been issued. However, the building permits represent intentions rather than actions, and it
may be that many of the construction plans for which permits have been sought are never
realized. Thus, this proximate indicator would lead to biased counts that underrepresent
the number of units actually constructed.
Whereas some measures are simply sloppy and overrepresent some cases while
undercounting others, thereby eroding reliability, other performance indicators have a
tendency to underreport or overreport on a systematic basis, creating validity problems.
For instance, reported crime statistics tend to underestimate actual crimes committed
because for various reasons, many crimes go unreported to the police. Periodic
victimization surveys may provide more valid estimates of actual crimes committed, but
they require considerable time, effort, and resources. Thus, the official reported crime
statistics are often used as indicators of the effectiveness of crime prevention programs or
police crime-solving activities even though they are known to underestimate actual crime
rates. In part, this is workable because the reported crime statistics may be valid for
tracking trends over time, say on a monthly basis, as long as the tendency for crimes to be
reported or not reported is constant from month to month.
One critical concern of juvenile detention facilities is to eliminate, or at least
minimize, instances of physical or sexual abuse of children in their custody by other
detainees or by staff members. Thus, one performance measure that is important to them
is the number of child abuse incidents occurring per month. But what would be the
operationalized indicator for this measure? One possibility would be the number of such
incidents reported each month, but this really represents the number of allegations of
child abuse. Considering that some of these allegations may well be unfounded, this
indicator would systematically tend to overestimate the real number of such incidents. A
preferred indicator would probably be the number of child abuse incidents that are
recorded on the basis of full investigations when such allegations are made. However, as
is true of reported crime rates in general, this measure would underestimate the actual
number of child abuse incidents if some victims of child abuse in these facilities are
afraid to report them.
Sound design of measuring instruments is essential for effective performance
measurement. This is particularly important with surveys of customers or other
stakeholders; items that are unclear or that incorporate biases can lead to serious
measurement problems. Often such surveys include questions that are vague, double-
barreled, or ambiguous, and because respondents are likely to interpret them in different
ways, the resulting data include a considerable amount of “noise” and thus are not very
reliable. A more serious problem arises when surveys include biased items— leading
questions that, intentionally or not, prompt respondents to answer in a certain way. For
example, an agency’s ongoing customer satisfaction survey could include questions and
response choices that are worded in such a way as almost to force respondents to give
programs artificially high ratings. This would obviously overestimate customer
satisfaction with this program and invalidate the survey data. These kinds of problems
can also apply to other modes of performance measurement, such as trained observer
ratings and other specially designed measurement tools. The important point here is that
care should always be taken to design measurement instruments that are clear,
unambiguous, and unbiased.
Biased observers are another source of severe validity problems. Even with a
good survey instrument, for instance, an interviewer who has some definite bias, either in
favor of a program or opposed to it for some reason, can obviously bias the responses in
that direction by introducing the survey, setting the overall tone, and asking the questions
in a certain way. In the extreme, the performance data generated by the survey may
actually represent the interviewer’s biases more than they serve as a valid reflection of
the views of the respondents. Clearly the problem of observer bias is not limited to survey
data. Many performance measures are operationalized through observer ratings, including
inspection of physical conditions, observation of behavioral patterns, or quality assurance
audits. In addition, performance data from clinical evaluations by physicians,
psychologists, therapists, and other professionals can also be vulnerable to observer
biases. To control for this possibility, careful training of interviewers and observers and
emphasis on the need for fair, unbiased observation and assessment are essential.
In addition to sound instrument design, consistent application of the measure over
time is critical to performance monitoring systems, precisely because they are intended to
track key performance measures over time. If the instrument changes over time, it can be
difficult to assess the extent to which trends in the data reflect real trends in performance
versus changes in measurement procedures. For example, if a local police department
begins to classify as crimes certain kinds of reported incidents that previously were not
counted as crimes, then everything else being equal, reported crime rates will go up.
These performance data could easily be interpreted as indicating that crime is on the rise
in that area or that crime prevention programs are not working well there, when in reality
the upward trend simply reflects a change in recording procedures. Instrument decay
refers to the erosion of integrity of a measure as a valid and reliable performance
indicator over time. For instance, as part of a city sanitation department’s quality control
effort, it trains a few inspectors to conduct spot checks of neighborhoods where
residential refuse collection crews have recently passed through to observe the amount of
trash and litter that might have been left behind. At first the inspectors adhere closely to a
regular schedule of visiting randomly selected neighborhoods and are quite conscientious
about rating cleanliness according to prescribed guidelines, but after several months, they
begin to slack off, stopping through neighborhoods on a hit-or-miss basis and making
casual assessments that stray from the guidelines. Thus, the measure has decayed over
this period and lost much of its reliability, and the data therefore are no longer
meaningful.
Sometimes measurements can change because people involved in the process are
affected by the program in some way or react somehow to the fact that the data are being
monitored by someone else or used for some particular purpose. For instance, if a state
government introduces a new scholarship program that ties awards to the grades students
earn in high school, teachers might begin, consciously or unconsciously, to be more
lenient in their grading. In effect, their standards for grading, or how they actually rate
students’ academic performance, change in reaction to the new scholarship program, but
the resulting data would suggest that students are performing better in high school now
than before, which may not be true. Or consider an inner-city neighborhood that forms a
neighborhood watch program in cooperation with the local police department, aimed at
increasing personal safety and security, deterring crime, and helping the police solve
crimes. As this program becomes more and more established, residents’ attitudes toward
both crime and the police change, and their propensity to report crimes to the police
increases. This actually provides a more valid indicator of the actual crime level than
used to be the case, but the data are likely to show increases in reported crimes even
though this is simply an artifact of reactive measurement.
The quality of performance monitoring data is often called into question by virtue
of being incomplete. Even with routine record-keeping systems and transactional
databases, agencies often are unable for a variety of reasons to maintain up-to-date
information on all cases all the time. So at any one time when the observations are being
made or the data are being run, say, at the end of every month, there may be incomplete
data in the records. If the missing data are purely a random phenomenon, this weakens
the reliability of the data and can create problems of statistical instability. If, however,
there is some systematic pattern of missing data—if, for instance, the database tends to
have less complete information on the more problematic cases—this can inject a
systematic bias into the data and erode the validity of the performance measure. Although
the problem of nonresponse bias may technically be a sampling issue, its real impact is to
introduce bias or distortion into performance measures.
Thus, in considering alternative performance indicators, it is a good idea to
ascertain the basis on which the measure is drawn in order to assess whether missing data
might create problems of validity or reliability. For example, almost all colleges and
universities require applicants to submit SAT or ACT scores as part of the admissions
process. Although the primary purpose of these test scores is to help in the selection
process, average scores, or the midspread of these scores, can be used to compare the
quality of applicants to that of different institutions or to track the proficiency of a
particular university’s freshmen class over several years. However, SAT or ACT scores
are also sometimes used as a proximate measure of the academic achievement of the
students in individual high schools or school systems, and here there may be problems
due to missing cases. Not all high school students take these tests, and those who do tend
to be the better students; therefore, average SAT or ACT scores tend to overstate the
academic proficiency of the student body as a whole. In fact, teachers and administrators
can influence average SAT or ACT scores for their schools simply by encouraging some
students to take the test and discouraging others. Thus, as an indicator of academic
achievement for entire schools, SAT or ACT scores are much more questionable than, for
example, standard exams mandated by the state for all students. When missing cases pose
potential problems, it is important to interpret the data on the basis of the actual cases on
which the data are drawn. Thus, average SAT scores can be taken as an indicator of the
academic achievement of those students from a given high school who chose to take the
exam.
Nonresponse bias can be problematic when performance measures, especially
effectiveness measures, come from follow-up contacts with former clients, whether
through surveys, follow-up visits, or other direct contact. Especially with respect to
human service programs, it is often difficult to remain in contact with all the individuals
who have been served by a program or who completed treatment some time ago. And it
may be the case that certain kinds of clients are much less likely to remain in contact. As
would probably be the case with vocational rehabilitation programs, teen mother
parenting programs, and especially crisis stabilization units, for example, the former
clients who are the most difficult to track down are often those with the least positive
outcomes, that is, those for whom the program may have been least effective. They may
be the most likely ones to move from the area, drop out of sight, leave the system, or fall
through the cracks. Obviously the nonresponse bias of data based on follow-up contacts
that will necessarily exclude some of the most problematic clients could easily lead to
overstating program performance. This is not to say that such indicators should not be
included in performance monitoring systems—because often they are crucial indicators
of longterm program effectiveness—but rather that care must be taken to interpret them
within the confines of actual response rates.
In addition to all the methodological issues that can compromise the quality of
performance monitoring data, a common problem that can destroy validity and reliability
is cheating. If performance measurement systems are indeed used effectively as a
management tool, they carry consequences in terms of decisions regarding programs,
people, resources, and strategies. Thus, managers at all levels of the organization want to
“look good” in terms of the performance data. Suppose, for example, that air force bases
whose function is training pilots to fly combat missions are evaluated in part by how
close they come to hitting rather ambitious targets that have been set concerning the
number of sorties flown by these pilots. The sorties are the principal output of these
training operations, and there is a clear definition of what constitutes a completed sortie.
If the base commanders are under heavy pressure to achieve these targets, however, and
actual performance is lagging behind these objectives, they might begin to count all
sorties as full sorties even though some of these have to be cut short for various reasons
and are not completed. This tampering with the definition of a performance measure may
seem to be a rather subtle distinction—and an easy one for the commanders to rationalize
given the pressure to maintain high ratings—but it would represent willful misreporting
to make a program appear to be more effective than it really is.
Performance measurement systems provide incentives for organizations and
programs to perform at higher levels, and this is the core of the logic underlying the use
of monitoring systems as performance management tools. Human nature being what it is,
then, it is not surprising that people in public and nonprofit organizations are sometimes
tempted to cheat—selectively reporting data, purposefully falsifying data, or otherwise
“cooking the books” in order to present performance in a more favorable light. This kind
of cheating is a real problem, and it must be dealt with directly and firmly. One strategy
to ensure the quality of the data is to build sample audits into the overall design of the
system, aimed at ensuring the accurate reporting and keeping the system honest. In terms
of the performance measures themselves, sometimes it is possible to use complementary
measures that will help identify instances in which the data don’t seem to add up, thus
providing a check on cheating. In a state highway maintenance program, for instance,
foremen inputting production data from the field might be tempted to misrepresent the
level of output produced by their crews by overstating such indicators as the miles of
road resurfaced, the miles of shoulders graded, or the feet of guardrail replaced. If these
data are only marginally overstated, they will appear to be reasonable and will probably
not be caught as errors. If, however, a separate system is used to report on inventory
control and the use of resources, and these data are input by different individuals in a
different part of the organization, then it may be possible to track these different
indicators in tandem. Numbers that don’t seem to match up would trigger a data audit to
determine the reason for the apparent discrepancy. Such a safeguard might be an effective
deterrent against cheating.
I. Other Criteria for Performance Measures
Performance measures should be meaningful; that is, they should be directly
related to the mission, goals, and intended results of a program, and they should represent
performance dimensions that have been identified as part of the program logic. To be
meaningful, performance measures should be important to managers, policymakers,
employees, customers, or other stakeholders. Managers may be more concerned with
productivity and program impact, policymakers may care more about efficiency and cost-
effectiveness, and clients may be more directly concerned with service quality, but for a
performance indicator to be meaningful, it must be important to at least one of these
stakeholder groups. If no stakeholder is interested in a particular measure, then it cannot
be particularly useful as part of a performance measurement system. Performance
indicators must also be understandable to stakeholders. That is, the measures need to be
presented in such a way as to explain clearly what they consist of and how they represent
some aspect of performance.
Within the scope and purpose of a given monitoring system, a set of performance
measures should be balanced and comprehensive. A fully comprehensive measurement
system should incorporate all the performance dimensions and types of measures
discussed in chapter 3, including both outputs and outcomes and, if relevant, service
quality and customer satisfaction in addition to efficiency and productivity. Even with
systems that are more narrowly defined—focusing solely on strategic outcomes, for
example, or, at the other extreme, focusing solely on operations—the measurement
system should attempt to include indicators of every relevant aspect of performance.
Perhaps most important, the monitoring system for a program with multiple goals should
include a balanced set of effectiveness measures rather than emphasize some intended
outcomes while ignoring others that may be just as important.
In order for a performance indicator to be useful, there must be agreement on the
preferred direction of movement on the scale. If an indicator of customer satisfaction, for
example, is operationalized as the percentage of respondents to an annual survey who say
they were satisfied or very satisfied with the service they have received from a particular
agency, higher percentages are taken to represent stronger program performance, and
managers will want to see this percentage increase from year to year. Although it might
seem that this should go without saying, the preferred direction of movement is not
always so clear. For instance, such indicators as the student-faculty ratio or the average
class size are sometimes used as proximate measures of the quality of instructional
programs at public universities, on the theory that smaller classes offer greater
opportunity for participation in class discussions and increased attention to the needs of
individual students both in and out of class. On that score, then, the preferred direction of
movement would be to smaller class sizes. College deans, however, often like to see
classes filling up with more students in order to make more efficient use of faculty time
and cover a higher percentage of operating costs. Thus, from a budgetary standpoint,
larger class sizes might be preferred. Generally if agreement on targets and the preferred
direction of movement cannot be reached in such ambiguous situations, then the
proposed indicator should probably not be used.
To be useful, performance measures also should be timely and actionable. One of
managers’ most common complaints about performance measurement systems is that
they do not report the data in a timely manner. When performance measures are designed
to support a governmental unit’s budgeting process, for instance, the performance data
for the most recently completed fiscal year should be readily available when budget
requests or proposals are being developed. In practice, however, sometimes the only
available data pertain to two years earlier and are out-ofdate as a basis for making
decisions regarding the current allocation of resources. Performance data that are
intended to be helpful to managers with responsibility for ongoing operations—such as
highway maintenance work, central office supply, claims processing operations, or child
support enforcement customer service—should probably be monitored more frequently,
on a monthly or quarterly basis, in order to facilitate addressing operational problems
more immediately. Although reporting frequency is really an issue of overall system
design, as discussed in chapter 2, it also needs to be taken into account in the definition of
the measures themselves.
In contrast, some proposed performance indicators may be well beyond the
control of the program and thus not actionable. For instance, one way in which public
hospitals track their performance is through surveys of patients who have been recently
discharged, because they can provide useful feedback on the quality and responsiveness
of the services they received. Suppose, however, that one particular item on such a survey
refers to the availability of some specific service or treatment option. The responses to
this item are consistently and almost universally negative, but the reason for this is that
none of the insurance companies involved will cover this option, and thus it is beyond the
control of the hospital. Because there is little or no chance of improving performance in
this area, at least under the existing constraints, this measure cannot provide any new or
useful feedback to hospital administrators.