1 / 31100%
Module 7
Test Analysis and Grading
A. Interpreting Test Scores
After a test is scored, the teacher needs to interpret the results and use these
interpretations to make grading, selection, placement, or other decisions. To accurately
interpret test scores, the teacher needs to analyze the performance of the test as a whole
and of the individual test items, and to use these data to draw valid inferences about
student performance. This information also helps teachers prepare for posttest discussions
with students about the exam. This chapter discusses the process of performing test and
item analyses. It also suggests ways in which teachers can use posttest discussions to
contribute to student learning and seek student feedback that can lead to test-item
improvement.
As a measurement tool, a test results in a score—a number. A number, however,
has no intrinsic meaning and must be compared with something that has meaning to
interpret its significance. For a test score to be useful for making decisions about the test,
the teacher must interpret the score. Whether the interpretations are norm referenced or
criterion referenced, a basic knowledge of statistical concepts is necessary to assess the
quality of tests (whether teacher-made or published), understand standardized test scores,
summarize assessment results, and explain test scores to others. Some information about
how a test performed as a measurement instrument can be obtained from computer-
generated test- and item-analysis reports. In addition to providing item-analysis data such
as difficulty and discrimination indexes, such reports often summarize the characteristics
of the score distribution. If the teacher does not have access to electronic scoring and
computer software for test and item analysis, many of these analyses can be done by
hand, albeit more slowly.
To make it easier to see similar characteristics of scores, the teacher should
arrange them in rank order, from highest to lowest (Miller, Linn, & Gronlund, 2013), as
in Table 12.2. Ordering the scores in this way makes it obvious that they ranged from 42
to 60, and that one student’s score was much lower than those of the other students. But
the teacher still cannot visualize easily how a typical student performed on the test or the
general characteristics of the obtained scores. Removing student names, listing each score
once, and tallying how many times each score occurs results in a frequency distribution,
as in Table 12.3. By displaying scores in this way, it is easier for the teacher to identify
how well the group of students performed on the exam.
The characteristics of a score distribution can be described on the basis of its
symmetry, skewness, modality, and kurtosis. These characteristics are illustrated in
Figure 12.4. A symmetric distribution or curve is one in which there are two equal halves,
mirror images of each other. Nonsymmetric or asymmetric curves have a cluster of scores
or a peak at one end and a tail extending toward the other end. This type of curve is said
to be skewed; the direction in which the tail extends indicates whether the distribution is
positively or negatively skewed. The tail of a positively skewed curve extends toward the
right, in the direction of positive numbers on a scale, and the tail of a negatively skewed
curve extends toward the left, in the direction of negative numbers. A positively skewed
distribution, thus, has the largest cluster of scores at the low end of the distribution, which
seems counterintuitive. The distribution of test scores from Table 12.1 is nonsymmetric
and negatively skewed. Remember that the lowest possible score on this test was 0 and
the highest possible score was 65; the scores were clustered between 42 and 60.
Frequency polygons and histograms can differ in the number of peaks they
contain; this characteristic is called modality, referring to the mode or the most frequently
occurring score in the distribution. If a curve has one peak, it is unimodal; if it contains
two peaks, it is bimodal. A curve with many peaks is multimodal. The relative flatness or
peakedness of the curve is referred to as kurtosis. Flat curves are described as platykurtic,
moderate curves are said to be mesokurtic, and sharply peaked curves are referred to as
leptokurtic. The shape of a score distribution depends on the characteristics of the test as
well as the abilities of the students who were tested (Brookhart & Nitko, 2019). Some
teachers make grading decisions as if all test score distributions resemble a normal curve,
that is, they attempt to “curve” the grades. An understanding of the characteristics of a
normal curve would dispel this notion. A normal distribution is a bellshaped curve that is
symmetric, unimodal, and mesokurtic.
Many human characteristics, such as intelligence, weight, and height, are
normally distributed; the measurement of any of these attributes in a population would
result in more scores in the middle range than at either extreme. However, most score
distributions obtained from teacher-made tests do not approximate a normal distribution.
This is true for several reasons. The characteristics of a test greatly influence the resulting
score distribution; a very difficult test tends to yield a positively skewed curve. Likewise,
the abilities of the students influence the test score distribution. Regardless of the
distribution of the attribute of intelligence among the human population, this
characteristic is not likely to be distributed normally among a class of nursing students or
a group of newly hired RNs. Because admission and hiring decisions tend to select those
individuals who are most likely to succeed in the nursing program or job, a distribution of
IQ scores from a class of 16 nursing students or 16 newly hired RNs would tend to be
negatively skewed. Likewise, knowledge of nursing content is not likely to be normally
distributed because those who have been admitted to a nursing education program or
hired as staff nurses are not representative of the population in general. Therefore,
grading procedures that attempt to apply the characteristics of the normal curve to a test
score distribution are likely to result in unwise and unfair decisions.
One of the questions to be answered when interpreting test scores is, “What score
is most characteristic or typical of this distribution?” A typical score is likely to be in the
middle of a distribution with the other scores clustered around it; measures of central
tendency provide a value around which the test scores cluster. Three measures of central
tendency commonly used to interpret test scores are the mode, median, and mean. The
mode, sometimes abbreviated Mo, is the most frequently occurring score in the
distribution; it must be a score actually obtained by a student. It can be identified easily
from a frequency distribution or graphic display without mathematical calculation. As
such, it provides a rough indication of central tendency. The mode, however, is the least
stable measure of central tendency because it tends to fluctuate considerably from one
sample to another drawn from the same population (Miller et al., 2013). That is, if the
same 65-item test that yielded the scores in Table 12.1 were administered to a different
group of 16 nursing students in the same program who had taken the same course, the
mode might differ considerably. In addition, as in the distribution depicted in Figure 12.1,
the mode has two or more values in some distributions, making it difficult to specify one
typical score. A uniform distribution of scores has no mode; such distributions are likely
to be obtained when the number of students is small, the range of scores is large, and
each score is obtained by only one student.
It is possible for two score distributions to have similar measures of central
tendency and yet be very different. The scores in one distribution may be tightly clustered
around the mean, and in the other distribution, the scores may be widely dispersed over a
range of values. Measures of variability are used to determine how similar or different the
students are with respect to their scores on a test. The simplest measure of variability is
the range, the difference between the highest and lowest scores in the distribution. For the
test score distribution in Table 12.3, the range is 18 (60 − 42 = 18). The range is
sometimes expressed as the highest and lowest scores, rather than a difference score.
Because the range is based on only two values, it can be highly unstable. The range also
tends to increase with sample size; that is, test scores from a large group of students are
likely to be scattered over a wide range because of the likelihood that an extreme score
will be obtained.
In addition to interpreting the test score distribution and measures of central
tendency and variability, teachers should examine test items in the aggregate for evidence
of bias. For example, although there may be no obvious gender bias in any single test
item, such a bias may be apparent when all items are reviewed as a group. Similar cases
of ethnic, racial, religious, and cultural bias may be found when items are grouped and
examined together. The effect of bias on testing and evaluation is discussed in detail in
Chapter 16, Social, Ethical, and Legal Issues. The ability to interpret the characteristics of
a distribution of scores will assist the teacher to make norm-referenced interpretations of
the meaning of any individual score in that distribution. However, most nurse educators
need to make criterion-referenced interpretations of individual test scores. A student’s
score on the test is compared with a preset standard or criterion, and the scores of the
other students are not considered. The percentage-correct score is a derived score that is
often used to report the results of tests that are intended for criterion-referenced
interpretation. The percentage correct is a comparison of a student’s score with the
maximum possible score; it is calculated by dividing the raw score by the total number of
items on the test (Miller et al., 2013). Although many teachers believe that percentage-
correct scores are an objective indication of how much students really know about a
subject, in fact they can change significantly with the difficulty of the test items. Because
percentage-correct scores are often used as a basis for assigning letter grades according to
a predetermined grading system, it is important to recognize that they are determined
more by test difficulty than by true quality of performance. For tests that are more
difficult than they were expected to be, the teacher may want to adjust the raw scores
before calculating the percentage correct on that test.
The percentage-correct score should not be confused with percentile rank, often
used to report the results of standardized tests. The percentile rank describes the student’s
relative standing within a group and therefore is a norm-referenced interpretation. The
percentile rank of a given raw score is the percentage of scores in the distribution that
occur at or below that score. A percentile rank of 83, therefore, means that the student’s
score is equal to or higher than the scores made by 83% of the students in that group; one
cannot assume, however, that the student answered 83% of the test items correctly.
Because there are 99 points that divide a distribution into 100 groups of equal size, the
highest percentile rank that can be obtained is the 99th. The median is at the 50th
percentile. Differences between percentile ranks mean more at the highest and lowest
extremes than they do near the median.
The results of standardized tests usually are intended to be used to make
normreferenced interpretations. Before making such interpretations, the teacher should
keep in mind that standardized tests are more relevant to general rather than specific
instructional goals. In addition, the results of standardized tests are more appropriate for
evaluations of groups rather than individuals. Consequently, standardized test scores
should not be used to determine grades for a specific course or to make a decision to hire,
promote, or terminate an employee. Like most educational measures, standardized tests
provide gross, not precise, data about achievement. Actual differences in performance
and achievement are reflected in large score differences.
Standardized test results usually are reported in derived scores such as percentile
ranks, standard scores, and norm group scores. Because all of these derived scores should
be interpreted in a norm-referenced way, it is important to specify an appropriate norm
group for comparison. The user’s manual for any standardized test typically presents
norm tables in which each raw score is matched with an equivalent derived score.
Standardized test manuals may contain a number of norm tables; the norm group on
which each table is based should be fully described. The teacher should take care to select
the norm group that most closely matches the group whose scores will be compared to it
(Miller et al., 2013). For example, when interpreting the results of standardized tests in
nursing, the performance of a group of baccalaureate nursing students should be
compared with a norm group of baccalaureate nursing students. Norm tables sometimes
permit finer distinctions such as size of program, geographical region, and public versus
private affiliation.
B. Item Analysis
In addition to test statistics, teachers also should examine indicators of
performance quality for each item on the exam. When used together, multiple data points
— difficulty index, discrimination index, and point biserial correlation coefficient—
provide a rich source of information about the performance quality of test items (Ermie,
n.d.). However, teachers should not depend solely on these statistical data to judge the
quality of exam items. Decisions about individual test items should be made in the
context of the content and structure of the item, the teacher’s expectations about how the
items would perform, and an accurate interpretation of the item statistics. Computer
software for item analysis is widely available for use with electronic answer sheet
scanning equipment. Commercially available computer testing applications usually also
provide services that produce user reports of item-analysis statistics. For teachers who do
not have access to such equipment and software, procedures for analyzing student
responses to test items by hand are described in detail later in this section. Regardless of
the method used for analysis, teachers should be familiar enough with the meaning of
each item-analysis statistic to correctly interpret the results.
It is important to keep in mind that item-discriminating power does not indicate
item validity. To gather evidence of item validity, the teacher would have to compare
each test item to an independent measure of achievement, which is seldom possible for
teacher-constructed tests. Standardized tests in the same content area usually measure the
achievement of more general objectives, so they are not appropriate as independent
criteria. The best measure of the domain of interest usually is the total score on the test if
the test has been constructed to correspond to specific instructional objectives and
content. Thus, comparing each item’s discriminating power to the performance of the
entire test determines how effectively each item measures what the entire test measures.
However, retaining very easy or very difficult items despite low discriminating power
may be desirable so as to measure a representative sample of learning objectives and
content. This statistic indicates the correlation between a student’s response to an item
and his or her overall performance on the exam. This statistic frequently is calculated and
provided as part of an item-analysis report from a commercial test development, scoring,
and analysis application. The statistic is interpreted in the same manner as the
discrimination index previously described. Values range from −1.00 to +1.00, with higher
positive values indicating that students who performed well on the exam tended to
answer the item correctly and lower positive values indicating that students whose overall
test performance was poor tended to answer the item incorrectly. A negative point
biserial correlation coefficient indicates a negative correlation between the total score and
performance on that item; students with low scores tended to answer the item correctly
whereas high-scoring students tended to answer incorrectly. As previously discussed,
negative correlations may indicate items that are flawed and need to be revised.
Teachers should not make decisions about retaining a test item in its present form,
revising it, or eliminating it from future use on the basis of the item statistics alone. Item
difficulty and discrimination indexes are not fixed, unchanging characteristics. Item-
analysis data for a given test item will vary from one administration to another because of
factors such as students’ ability levels, quality of instruction, and the size of the group
tested. With very small groups of students, if a few students would have changed their
responses to the test item, the difficulty and discrimination indexes could change
considerably (Miller et al., 2013). Thus, when using these indexes to identify
questionable items, the teacher should carefully examine each test item for evidence of
poorly functioning distractors, ambiguous alternatives, and miskeying.
Ideally, every distractor should be selected by at least one student in the lower
group, and more lower group students than higher group students should select it. A
distractor that is not selected by any student in the lower group may contain a technical
flaw or may be so implausible as to be obvious even to students who lack knowledge of
the correct answer. A distractor may be ambiguous if upper group students tend to choose
it with about the same frequency as the keyed, or correct, response. This result usually
indicates that there is no single clearly correct or best answer. Poorly functioning and
ambiguous distractors may be revised to make them more plausible or to eliminate the
ambiguity. If a large number of higher scoring students select a particular incorrect
response, the teacher should check to see whether the answer key is correct. However, as
previously mentioned, the content of the item along with the statistics should guide the
teacher’s decision-making.
Multiple factors other than the content and structure of the exam item also may
affect students’ answers to a test item. Supplemental readings may contradict what was
discussed in class, students may misinterpret poorly worded distractors, or an item may
be miskeyed. Psychometric data cannot identify these factors, hence the need for item
analysis accompanied by review of the actual item (Ermie, n.d.). The following examples
from a commercial item-analysis report illustrate how to use item statistics and distractor
analysis in the context of actual item content to determine whether the items performed as
expected and desired. Assume that these sample items were included on a unit exam on
nursing of patients with endocrine disorders. The exam contained 85 items, each worth 1
point, and 70 students took the exam.
Item 48 (Exhibit 12.5) is a multiple-response (select all that apply) item with five
response options. Only 2 of 70 students chose combination E, and a closer look at this
combination reveals that including mutually exclusive alternatives might have been an
unintentional clue that it was incorrect, apparent even to students who were in the low-
scoring group. This combination should be revised before future use of this item. There is
no obvious reason for the slightly negative discrimination index, but a large number of
students selected the incorrect combination A, and some of them likely were high-scoring
students. Comparing combination A with the correct answer reveals that option A
excludes alternative 2, a statement about food intake. Students may have eliminated this
alternative because the first sentence of the item stem focuses on exercise. This item also
should be reviewed with students during the posttest discussion to determine why they
chose this response instead of the correct one. In this case, however, the teacher should
not accept combination A as a second correct response because it is not complete.
Giving students feedback about test results can be an opportunity to reinforce
learning, to correct misinformation, and to solicit their input for improvement of test
items. But a feedback session also can be an invitation to engage in battle, with students
attacking to gain extra points and the teacher defending the honor of the test and, it often
seems, the very right to give tests. Discussions with students about the test should be
rational rather than opportunities for the teacher to assert power and authority. Posttest
discussions can be beneficial to both teachers and students if they are planned in advance
and not emotionally charged. The teacher should prepare for a posttest discussion by
completing a test analysis and an item analysis and reviewing the items that were most
difficult for the majority of students.
Teachers may use this information about item effectiveness as an aid to posttest
discussion. The items with the lowest difficulty index (the ones answered incorrectly by
the largest number of students) can be discussed at greater length, and the teacher can ask
students why they selected the correct or wrong answer for such items. A discussion of
the rationale for their choices may reveal common errors and misconceptions that may be
corrected at that time, serve as a basis for remedial study, or contribute to a revision of
those items. Ierardi (2014) described a student-centered approach to posttest exam review
in which a student representative volunteers to moderate the discussion of test items, with
a faculty member present to clarify. Students who answered test items correctly provide
insight into how they approached the items and chose the correct responses. The faculty
member benefits from hearing the student discussion because it often reveals important
information about items that can be used to revise them for future use.
Discussing why students chose correct or incorrect answers can reveal students’
thought processes that can contribute to better teaching, improved test construction skills,
and better test-taking skills. To use time efficiently, the teacher should read the correct
answers aloud quickly. If the test is hand-scored, correct answers also may be indicated
by the teacher on the students’ answer sheets or test booklets. If machine-scoring is used,
the answer key may be projected as a scanned document from a computer or via a
document camera or overhead projector. Many electronic scoring applications allow an
option for marking the correct or incorrect answers directly on each student’s answer
sheet.
During test administration, some teachers allow students to record their answers
on the test booklets, where the students also record their names, as well as on a separate
answer sheet. At the completion of the exam, students submit the answer sheets and their
test booklets to the teacher. When all students have finished the exam, they return to the
room to check their answers using only their test booklets. The teacher might project the
answers onto a screen as described previously. At the conclusion of this session, the
teacher collects the test booklets again. It is important not to review and discuss
individual items at this time because the test has not yet been scored and analyzed.
However, the teacher may ask students to indicate problematic items and give a rationale
for their answers. As discussed earlier, the teacher can use this item in conjunction with
the item-analysis results to evaluate the effectiveness of test items. One disadvantage to
this method of giving posttest feedback is that because the test has not yet been scored
and analyzed, the teacher would not have an opportunity to thoroughly prepare for the
session; feedback consists only of the correct answers, and no discussion takes place.
With item effectiveness information, the teacher can identify and point out defective test
items and discuss how they will be treated in scoring, rather than feel the need to defend
the fairness of the items without data to support it.
Teachers often debate the merits of adjusting test scores by eliminating items or
adding points to compensate for real or perceived deficiencies in test construction or
performance. For example, during a posttest discussion, students may argue that if they
all answered an item incorrectly, the item should be omitted or all students should be
awarded an extra point to compensate for the “bad item.” It is interesting to note that
students seldom propose subtracting a point from their scores if they all answer an item
correctly. In any case, how should the teacher respond to such requests? In this
discussion, a distinction is made between test items that are technically flawed and those
that do not function as intended. If test items are properly constructed, critiqued, and
proofread, it is unlikely that serious flaws will appear on the test. However, errors that do
appear may have varying effects on students’ scores. For example, if the correct answer
to a multiple-choice item is inadvertently omitted from the test, no student will be able to
answer the item correctly. In this case, the item simply should not be scored. That is, if
the error is discovered during or after test administration and before the test is scored, the
item is omitted from the answer key; a test that was intended to be worth 73 points then is
worth 72 points. If the error is discovered after the tests are scored, they can be rescored.
Students often worry about the effect of this change on their scores and may argue that
they should be awarded an extra point in this case.
If the technical flaw consists of a misspelled word in a true–false item that does
not change the meaning of the statement, no adjustment should be made. The teacher
should avoid lengthy debate about item semantics if it is clear that such errors are
unlikely to have affected the students’ scores. Feedback from students can be used to
revise items for later use and sometimes make changes in teaching that concept or skill.
As previously discussed, teachers should resist the temptation to eliminate items from the
test solely on the basis of low difficulty and discrimination indices. Omission of items
may affect the validity of the scores from the test, particularly if several items related to
one content area or objective are eliminated, resulting in inadequate sampling of that
content (Miller et al., 2013). This is particularly true for quizzes that contain a small
number of items. Because identified flaws in test construction do contribute to
measurement error, the teacher should consider taking them into account when using the
test scores to make grading decisions and set cutoff scores. That is, the teacher should not
fix cutoff scores for assigning grades until after all tests have been given and analyzed.
The proposed grading scale can then be adjusted if necessary to compensate for
deficiencies in test construction. It should be made clear to students that any changes in
the grading scale because of flaws in test construction would not adversely affect their
grades.
C. Developing a Test-Item Bank
Because considerable effort goes into developing, administering, and analyzing
test items, teachers should develop a system for maintaining and expanding a pool or
bank of items from which to select items for future tests. Teachers can maintain databases
of test items on their computers with backups on storage devices. When teachers store
test-item databases electronically, the files must be passwordprotected and test security
maintained. When developing test banks, the teacher can record the following data with
each test item: (a) the correct response for objective-type items and a brief scoring key
for completion or essay items; (b) the course, unit, content area, or objective for which it
was designed; and (c) the item-analysis results for a specified period of time. Exhibit 12.6
offers one such example. Commercially produced software applications can be used in a
similar way to develop and store a database of test items. Each test item is a record in the
database. The test items can then be sorted according to the fields in which the data are
entered; for example, the teacher could retrieve all items that are classified as Objective
3, with a moderate difficulty index.
Many publishers also offer test-item banks that relate to the content contained in
their textbooks. However, faculty members need to be cautious about using these items
for their own examinations. The purpose of the test, relevant characteristics of the
students to be tested, and the balance and emphasis of content as reflected in the teacher’s
test blueprint are the most important criteria for selecting test items. Although some
teachers would consider these item banks to be a shortcut to test development, items
should be evaluated carefully before they are used. There is no guarantee that the quality
of test items in a published item bank is superior to that of test items that a skilled teacher
can construct. Many of the items may be of questionable quality. Often, a teacher can
improve the quality of commercial test-bank items that are congruent with his or her test
blueprint by modifying an item stem, substituting better answer options, eliminating
technical flaws, or changing the item format. In addition, published test-item banks
seldom contain item-analysis information such as difficulty and discrimination indices.
However, the teacher can calculate this information for each item used or modified from
a published item bank, and can develop and maintain an item file.
To accurately interpret test scores, the teacher needs to analyze the performance
of the test as a whole as well as the individual test items. Information about how the test
performed helps teachers to give feedback to students about test results and to improve
test items for future use. Scoring a test results in a collection of numbers known as raw
scores. To make raw scores understandable, they can be arranged in frequency
distributions or displayed graphically as histograms or frequency polygons. Score
distribution characteristics such as symmetry, skewness, modality, and kurtosis can assist
the teacher in understanding how the test performed as a measurement tool as well as to
interpret any one score in the distribution. Measures of central tendency and variability
also aid in interpreting individual scores. Measures of central tendency include the mode,
median, and mean; each measure has advantages and disadvantages for use. In a normal
distribution, these three measures will coincide. Most score distributions from teacher-
made tests do not meet the assumptions of a normal curve. The shape of the distribution
can determine the most appropriate index of central tendency to use. Variability in a
distribution can be described roughly as the range of scores or, more precisely, as the
standard deviation.
A percentage-correct score is calculated by dividing the raw score by the total
possible score; thus, it compares the student’s score to a preset standard or criterion. A
percentage-correct score is not an objective indication of how much a student really
knows about a subject because it is affected by the difficulty of the test items. The
percentage-correct score should not be confused with percentile rank, which describes the
student’s relative standing within a group and therefore is a normreferenced
interpretation. The percentile rank of a given raw score is the percentage of scores in the
distribution that occurs at or below that score. The results of standardized tests usually
are reported as percentile ranks or other norm-referenced scores. Teachers should be
cautious when interpreting standardized test results so that comparisons with the
appropriate norm group are made. Standardized test scores should not be used to
determine grades, and results should be interpreted with the understanding that only large
differences in scores indicate real differences in achievement levels.
Item analysis typically is performed by the use of a computer program, either as
part of a test scoring application or computer testing software. The difficulty index (P),
ranging from 0 to 1.00, indicates the percentage of students who answered the item
correctly. Items with P-values of .20 and below are considered to be difficult, and those
with P-values of .80 and above are considered to be easy. However, interpretation of the
difficulty index should take into account the quality of the instruction and the abilities of
the students in the group. The discrimination index (D), ranging from −1.00 to +1.00, is
an indication of the extent to which high-scoring students answered the item correctly
more often than low-scoring students did. In general, the higher the positive value, the
better the test item; desirable discrimination indexes should be at least +.20. An item’s
power to discriminate is highly related to its difficulty index. An item that is answered
correctly by all students has a difficulty index of 1.00; the discrimination index for this
item is 0, because there is no difference in performance on that item between high scorers
and low scorers. Flaws in test construction may have varying effects on students’ scores
and therefore should be handled differently. If the correct answer to a multiple-choice
item is inadvertently omitted from the test, no student will be able to answer the item
correctly. In this case, the item simply should not be scored. If a flaw consists of a
misspelled word that does not change the meaning of the item, no adjustment should be
made.
D. Purposes and Criticisms of Grades
In earlier chapters, there was extensive discussion about formative and summative
evaluation. Through formative evaluation, the teacher provides feedback to the learner on
a continuous basis. In contrast, summative evaluation is conducted periodically to
indicate the student’s achievement at the end of the course or at a point during the course.
Summative evaluation provides the basis for arriving at grades in the course. Grading, or
marking, is defined as the use of symbols, for instance, the letters A to F, for reporting
student achievement. Grading is used for summative purposes, indicating through the use
of symbols how well the student performed in individual assignments, clinical practice,
laboratories (skills, simulation, others), and the course as a whole. To reflect valid
judgments about student achievement, grades need to be based on careful evaluation
practices, reliable test results, and multiple assessment methods. No grade should be
determined by one method or one assignment completed by the students; grades reflect
instead a combination of various tests and other assessment methods.
Absolutely, incorporating assignments that are not graded can be highly
beneficial, especially when the focus is on formative evaluation and fostering a
supportive learning environment. Ungraded assignments allow educators to assess
student learning progress without the pressure of assigning a formal grade. This approach
emphasizes feedback and improvement over evaluation, helping students identify areas
for growth and development. Students may be more willing to engage in exploratory or
creative tasks when they know their efforts will not be graded. This promotes a risk-free
environment where students can experiment, make mistakes, and learn from their
experiences without fear of academic consequences. Assignments that are not graded can
focus on developing specific skills or competencies relevant to the course objectives.
Students can practice critical thinking, problem-solving, communication, or collaboration
skills in a supportive context that encourages experimentation and learning.
Ungraded assignments can include activities such as reflective journals, self-
assessments, or peer reviews. These promote metacognitive awareness as students reflect
on their learning process, strengths, and areas needing improvement. Providing
opportunities for ungraded assignments allows educators to cater to diverse learning
styles and preferences. Some students may excel in tasks that require creativity,
teamwork, or independent exploration, which may not always align with traditional
graded assessments. By focusing on learning rather than grades, students are encouraged
to delve deeper into course content and develop a more comprehensive understanding of
the material. This promotes intrinsic motivation and a genuine interest in learning.
Educators can provide timely and constructive feedback on ungraded
assignments, guiding students toward academic success and mastery of course concepts.
This feedback-oriented approach supports ongoing improvement and helps students
achieve their learning goals. Ungraded assignments contribute to a less stressful learning
environment, where students can approach tasks with a growth mindset and focus on
learning rather than solely on achieving a certain grade. In summary, integrating
ungraded assignments into the learning process can enhance student engagement,
promote deeper learning, and support formative assessment practices. By emphasizing
feedback, skill development, and reflective learning, educators create a supportive
atmosphere conducive to academic growth and achievement.
Grades in a course serve multiple purposes beyond just evaluating and assessing
students' performance. Grades provide feedback to students on their progress and
understanding of course material. They highlight strengths and weaknesses, guiding
students on where to focus their efforts for improvement. Grades can motivate students to
engage actively in learning activities, complete assignments on time, and strive for
academic excellence. Assigning grades encourages students to take responsibility for
their learning outcomes and academic performance. Grades serve as a formal record of
students' academic achievements and progress throughout the course. Grades are essential
for transcripts and official academic records, documenting students' academic
performance for future educational pursuits, employment opportunities, and professional
licensure. Grades may be used to determine whether students have met the academic
requirements for graduation or progression within their program of study.
Grades inform academic advisors and counselors about students' academic
standing, helping them provide informed guidance on course selection, career paths, and
academic planning. Grades can signal when students may need additional support or
intervention, such as tutoring, counseling services, or academic mentoring. Grades
provide a basis for setting academic goals and developing strategies for academic
success, helping students track their progress toward achieving these goals. Overall,
while not all student activities need to be graded, grades play a vital role in the
educational process by providing feedback, motivating students, documenting academic
achievements, and guiding academic and career development. Balancing these purposes
ensures that grades serve as a meaningful tool for both students and educators in fostering
learning and academic success.
Grades for instructional purposes indicate the achievement of students in the
course. They provide a measure of what students have learned and their competencies at
the end of the course or at a certain point within it. A “pass” grade in the clinical
practicum and a grade of “B” in the nursing course are examples of using grades for
instructional purposes. The third use of grades is for guidance and counseling. Grades can
be used to make decisions about courses to select, including more advanced courses to
take or remedial courses that might be helpful. Grades also suggest academic resources
that students might benefit from such as reading, study, and test-taking workshops and
support. In some situations, grades assist students in making career choices, including a
change in the direction of their careers.
E. Types of Grading Systems
There are different types of grading systems or methods of reporting grades. Most
nursing education programs use a letter system for grading (A, B, C, D, E or A, B, C, D,
F), which may be combined with “+” and “−.” The integers 5, 4, 3, 2, and 1 (or 9–1) also
may be used. These two systems of grading are convenient to use, yield grades that are
able to be averaged within a course and across courses, and present the grade concisely.
Grades also may be indicated by percentages (100, 99, 98, …). Most programs use
percentages as a basis for assigning letter grades—90% to 100% represents an A, 80% to
89% represents a B, and so forth. In some nursing programs, the percentages for each
letter grade are higher, for example, 93% to 100% for an A, 85% to 92% for a B, 76% to
84% for a C, 67% to 75% for a D, and 67% and below for an E or F. It is not uncommon
in nursing education programs to specify that students need to achieve at least a C in each
nursing course at the undergraduate level and a B or better at the graduate level.
Requirements such as these are indicated in the school policies and course syllabi.
Another type of grading system is two-dimensional: pass–fail, satisfactory–
unsatisfactory, credit–no credit, and met–not met. For determining clinical grades, some
programs add a third honors category, creating three levels: honors– pass–fail. One
advantage of a two-dimensional grading system is that the grade is not calculated in the
GPA. This allows students to take new courses and explore different areas of learning
without concern about the grades in these courses affecting their overall GPA. This also
may be viewed as a disadvantage, however, in that clinical performance in a nursing
course graded on a pass–fail basis is not calculated as part of the overall course grade. A
pass indicates that students met the outcomes or demonstrated satisfactory performance
of the clinical competencies.
Grading systems for clinical practice are a critical aspect of nursing education,
ensuring that students' clinical skills, competencies, and professional behaviors are
effectively evaluated in real-world healthcare settings. Clinical grading systems aim to
assess students' ability to apply theoretical knowledge in practical, patient care situations.
They evaluate clinical skills, decision-making abilities, and adherence to professional
standards. These systems provide structured feedback to students, identifying strengths
and areas for improvement in their clinical performance. This feedback supports ongoing
professional development and enhances learning outcomes. Assessment of technical
skills such as medication administration, wound care, physical assessments, and
procedural competencies. Evaluation of students' ability to analyze patient data, prioritize
nursing interventions, and make evidence-based decisions.
Assessment of students' communication with patients, families, and healthcare
teams, including therapeutic communication, empathy, and teamwork. Evaluation of
students' professionalism, including ethical behavior, respect for patient rights, adherence
to healthcare policies, and accountability. In this system, students are assessed against
predetermined criteria or standards of competence. It focuses on whether students have
achieved specific learning outcomes and proficiency levels. This approach measures
students' mastery of competencies essential for safe and effective nursing practice. It
emphasizes demonstrated proficiency in skills and knowledge rather than time spent in
clinical settings. Rubrics provide a structured framework for evaluating different aspects
of clinical performance, including observable behaviors, quality of interactions, and
achievement of learning objectives.
Ensuring consistency and fairness in grading across different clinical settings and
instructors. Managing the inherent subjectivity in clinical assessment due to varying
perceptions and interpretations of clinical performance. Providing constructive feedback
that supports student learning and development without discouraging or demoralizing
learners. Aligning clinical grading with overall curriculum goals, learning outcomes, and
accreditation standards. Effective clinical grading systems contribute to the professional
growth of students, preparing them for entry into the nursing profession with confidence
and competence.
By ensuring that students meet established standards of practice, these systems
promote safe and high-quality patient care delivery. Preparing students to meet clinical
competency requirements for licensure examinations such as the NCLEX-RN, ensuring
they are well-prepared for professional practice. In summary, discussions on grading
systems for clinical practice in nursing education emphasize their role in evaluating
competencies, providing feedback, and preparing students for professional nursing roles.
These systems are designed to foster continuous improvement, uphold professional
standards, and support students in becoming proficient and compassionate healthcare
providers.
F. Assigning Letter Grades
Because most nursing education programs use the letter system for grading
nursing courses, this framework will be used for discussing how to assign grades. These
principles, however, are applicable to the other grading systems as well. There are two
major considerations in assigning letter grades: deciding what to include in the grade and
selecting a grading framework. Grades in nursing courses should reflect the student’s
achievement and not be biased by the teacher’s own values, beliefs, and attitudes. If the
student did not attend class or appeared to be inattentive during lectures, this behavior
should not be incorporated into the course grade unless criteria were established at the
outset for class attendance and participation.
The student’s grade is based on the tests and assessment methods developed for
the course. Multiple assessment methods should be used to determine course grades. The
weight given to each of these in the overall grade should reflect the emphasis of the
objectives and the content measured by them. Tests and other assessment methods
associated with important content, for which more time was probably spent in the
instruction, should receive greater weight in the course grade. For example, a midterm
examination in a community health nursing course should be given more weight in the
course grade than a short paper that students completed about community resources for a
family under their care. How much weight should be given in the course grade to each
test and other types of assessment methods used in the course? The teacher begins by
listing the tests, quizzes, papers, presentations, and other assessment methods in the
course that should be included in the course grade.
In nursing education, once students have undergone clinical assessments based on
various components such as skill proficiency, critical thinking, communication, and
professionalism, the teacher or clinical instructor plays a crucial role in determining the
importance of each of these components within the overall course grade. The teacher
establishes the weighting or importance of each component in the overall course grade.
For example, skill proficiency might carry a higher weight if the course focuses heavily
on technical competencies, while professionalism and communication skills may also be
significant but weighted differently based on the course objectives. This weighting
reflects the course's learning outcomes and ensures that assessment aligns with the
educational priorities and professional standards relevant to nursing practice.
Teachers often use grading rubrics that outline specific criteria and performance
expectations for each component being assessed. These rubrics provide transparency to
students regarding how their clinical performance will be evaluated and graded. Criteria
may include specific behaviors, skills, or competencies that students are expected to
demonstrate during clinical practice, allowing for consistent evaluation across different
clinical settings and instructors. While grading systems provide structure, teachers also
exercise professional judgment in evaluating students' clinical performance. This involves
assessing the quality of students' interactions, their clinical reasoning processes, and their
overall readiness for professional nursing practice. Teachers strive for consistency in
grading by adhering to established criteria, using rubrics consistently, and engaging in
calibration exercises with other faculty or clinical instructors to ensure fairness and
reliability in assessment.
Alongside assigning grades, teachers provide constructive feedback to students.
This feedback focuses on strengths, areas for improvement, and actionable
recommendations to support students' ongoing learning and development. Feedback plays
a critical role in helping students understand their clinical strengths and areas needing
improvement, fostering a culture of continuous learning and improvement in nursing
education The weighting of each component in the overall grade reflects the educational
goals of the nursing program and the specific course objectives. For instance, a course
aiming to develop advanced clinical judgment may place greater emphasis on critical
thinking skills in the grading structure. Alignment with accreditation standards and
licensure requirements ensures that students are adequately prepared for professional
practice and success on licensure examinations such as the NCLEX-RN. In summary, the
teacher's role in determining the importance of each component in the overall course
grade involves careful consideration of educational objectives, alignment with
professional standards, and consistent application of assessment criteria. This approach
ensures that grading practices in nursing education support the development of
competent, skilled, and ethical nurses ready to meet the demands of contemporary
healthcare practice.
G. Criterion, Norm, and Self Referenced Grading
In criterion-referenced grading, grades are based on students’ achievement of the
outcomes of the course, the extent of content learned in the course, or how well they
performed in clinical practice. Students who achieve more of the objectives, acquire more
knowledge, and can perform more competencies or with greater proficiency receive
higher grades. The meaning assigned to grades, then, is based on these absolute standards
without regard to the achievement of other students. Using this frame of reference for
grading means that it is possible for all students to achieve an A or a B in a course, if they
meet the standards, or a D or F if they do not. This framework is appropriate for most
nursing courses because it focuses on outcomes and competencies to be achieved in the
course. Criterion-referenced grading indicates how students are progressing toward
meeting those outcomes (formative evaluation) and whether they have achieved them at
the end of the course (summative evaluation). Norm-referenced grading, in contrast, is
not appropriate for use in nursing courses that are based on standards or learning
outcomes because it focuses on comparing students with one another, not on how they
are progressing or on their achievement. For example, formative evaluation in a norm-
referenced framework would indicate how each student ranks among the group rather
than provide feedback on student progress in meeting the outcomes of the course and
strategies for further learning.
There are several ways of assigning grades using a criterion-referenced system.
One is called the fixed-percentage method. This method uses fixed ranges of
percentcorrect scores as the basis for assigning grades (Miller, Linn, & Gronlund, 2013).
A common grading scale is 93% to 100% for an A, 85% to 92% for a B, 76% to 84% for
a C, 67% to 75% for a D, and below 67% for an E or F. Each component of the course
grade—written tests, quizzes, papers, case presentations, and other assignments—is given
a percentage-correct score or percentage of the total points possible. For example, the
student might have a score of 21 out of 25 on a quiz, or 84%. The component grades are
then weighted, and the percentages are averaged to get the final grade, which is converted
to a letter grade at the end of the course. With all grading systems, the students need to be
informed as to how the grade will be assigned. If the fixed-percentage method is used, the
students should know the scale for converting percentages to letter grades; this should be
in the course syllabus with a clear explanation of how the course grade will be
determined.
In determining the composite score for the course, the student’s percentage for
each of the components of the grade is multiplied by the weight and summed; the sum is
then divided by the sum of the weights. This procedure is shown in Table 17.3.
Generally, test and other component scores should not be converted to grades for the
purpose of later computing a final average grade. Instead, the teacher should record
actual test scores and then combine them into a composite score that can be converted to
a final grade. The second method of assigning grades in a criterion-referenced system is
the total-points method. In this method, each component of the grade is assigned a
specific number of points, for example, a paper may be worth 100 points and midterm
examination 75 points. The number of points assigned reflects the weights given to each
component within the course, that is, what each one is “worth.”
One problem with this method is that the decision about the points to allot to each
test and evaluation method in the course is made before the teacher has developed them
(Brookhart & Nitko, 2019). For example, to end with 500 points for the course, the
teacher may need 75 points for the midterm exam. However, in preparing that exam, the
teacher finds that 73 items adequately cover the content and reflect the emphasis given to
the content in the instruction. If this were known during the course planning, the teacher
could assign 2 fewer points to the midterm exam and add 2 points to the final exam or
one of the assignments, or merely alter the total number of points for the course grade.
However, when the course is already underway, changes such as these cannot be made in
the grading scheme, and the teacher needs to develop a 75-point midterm exam even if
fewer items would have adequately sampled the content. The next time the course is
offered, the teacher can modify the points allotted for the midterm exam in the course
grade.
In a norm-referenced grading system using relative standards, grades are assigned
by comparing a student’s performance with that of others in the class. Students who
perform better than their peers receive higher grades. When using a normreferenced
system, the teacher decides on the reference group against which to compare a student’s
performance. Should students be compared with others in the course? Should they be
compared with students only in their section of the course? Or with students who
completed the course the prior semester or previous year? One issue with norm-
referenced grading is that high performance in a particular group may not be indicative of
mastery of the content or what students have learned; it reflects instead a student’s
standing in that group.
Two methods of assigning grades using a norm-referenced system are (a)
“grading on the curve” and (b) using standard deviations. Grading on the curve refers to
the score distribution curve. In this method, students’ scores are rank-ordered from
highest to lowest, and grades are assigned according to the rank order. After the quotas
are set, grades are assigned without considering actual achievement. For example, the top
20% of students will receive an A even if their scores are close to the next group that gets
a B. The students assigned lower grades may in fact have acquired sufficient knowledge
in the course but unfortunately had lower scores than the other students. In these two
examples, the decisions on the percentages of As, Bs, Cs, and lower grades are made
arbitrarily by the teacher. The teacher determines the proportion of grades at each level;
this approach is not based on a normal curve. For “grading on the curve” to work
correctly, student scores need to be distributed based on the normal curve. However, the
abilities of nursing students tend not to be heterogeneous, especially late in the nursing
education program, and therefore their scores on tests and other evaluation products are
not normally distributed. They are carefully selected for admission into the program, and
they need to achieve certain grades in courses and earn minimum GPAs to progress in the
program. With grading on the curve, even if most students achieved high grades on a test
and mastered the content, some would still be assigned lower grades.
The second method is based on standard deviations. With this method, the teacher
determines the cutoff points for each grade. The grades are based on how far they are
from the mean of raw scores for the class. To use the standard deviation method, the
teacher first prepares a frequency distribution of the final scores and then calculates the
mean score. With this method, the teacher has a reference point (mean) and the average
distance of scores from the mean (Miller et al., 2013). The grade boundaries are then
based on the standard deviation. For example, the cutoff points for a C grade might range
from one half the standard deviation below the mean to one half above the mean. To
identify the A–B cutoff scores, the teacher adds one standard deviation to the upper
cutoff number of the C range. Subtracting one standard deviation from the lower C cutoff
provides the range for the D–F grades.
Self-referenced grading is based on standards of growth and change in the
student. With this method, grades are assigned by comparing the student’s performance
with the teacher’s perceptions of the student’s capabilities or the student’s own progress
over the course (Brookhart & Nitko, 2019). Did the student achieve at a higher level than
deemed capable regardless of the knowledge and competencies acquired? Did the student
improve performance throughout the course? Table 17.2 compares self-referencing with
criterion- and norm-referenced grading. One major problem with this method is the
unreliability of the teacher’s perceptions of student capability and growth, and the
student’s own assessment of performance. A second issue occurs with students who enter
the course or clinical practice with a high level of achievement and proficiency in many
of the clinical competencies. These students may have the least amount of growth and
change but nevertheless exit the course with the highest achievement and clinical
competency. Ultimately, judgments about the quality of a nursing student’s performance
are more important than judgments about the degree of improvement. It is difficult to
make valid predictions about future performance on licensure or certification exams, or in
clinical practice, based on self-referenced grades. For these reasons, self-referenced
grades are not widely used in nursing education programs.
H. Grading Clinical Practice
Arriving at grades for clinical practice is difficult because of the nature of clinical
practice and the need for judgments about performance. Issues in evaluating clinical
practice and rating performance were discussed in Chapter 13, Clinical Evaluation
Process and Chapter 14, Clinical Evaluation Methods. Many teachers constantly revise
their rating forms for clinical evaluation and seek new ways of grading clinical practice.
Although these changes may create a fairer grading system, they will not eliminate the
problems inherent in judging clinical performance. The different types of grading systems
described earlier may be used for grading clinical practice. In general, these include
systems using letter grades, A to F; integers, 5 to 1; and percentages. Grading systems for
clinical practice also may use pass– fail, satisfactory–unsatisfactory, and met–did not
meet the clinical objectives. Some programs add a third category, honors, to acknowledge
performance that exceeds the level required. Pass–fail is used most frequently in nursing
programs (Oermann, Yarbrough, Ard, Saewert, & Charasika, 2009). With any of the
grading systems, it is not always easy to summarize the multiple types of evaluation data
collected on the student’s performance in a symbol representing a grade. This is true even
in a pass–fail system; it may be difficult to arrive at a judgment as to pass or fail based on
the evaluation data and the circumstances associated with the student’s clinical,
simulated, and laboratory practice.
Regardless of the grading system for clinical practice, there are two criteria to be
met: (a) the evaluation methods for collecting data about student performance should
reflect the outcomes and clinical competencies for which a grade will be assigned, and
(b) students must understand how their clinical practice will be evaluated and graded.
Decisions about assigning letter grades for clinical practice are the same as grading any
course: identifying what to include in the clinical grade and selecting a grading
framework. The first consideration relates to the evaluation methods used in the course to
provide data for determining the clinical grade. Some of these evaluation methods are for
summative evaluation, thereby providing a source of information for inclusion in the
clinical grade. Other strategies, though, are used in clinical practice for feedback only and
are not incorporated into the grade.
Categories for grading clinical practice, such as pass–fail, satisfactory–
unsatisfactory, and met–not met, have some advantages over a system with multiple
levels, although there are disadvantages as well. Pass–fail places greater emphasis on
giving feedback to the learner because only two categories of performance need to be
determined. With a pass–fail grading system, teachers may be more inclined to provide
continual feedback to learners because ultimately they do not have to differentiate
performance according to four or five levels of proficiency such as with a letter system.
Performance that exceeds the requirements and expectations, however, is not reflected in
the grade for clinical practice unless a third category is included: honors–pass–fail. A
pass–fail system requires only two types of judgment about clinical performance. Do the
evaluation data indicate that the student has met the outcomes or has demonstrated
satisfactory performance of the clinical competencies to indicate a pass? Or do the data
suggest that the performance of those competencies is not at a satisfactory level? Arriving
at a judgment as to pass or fail is often easier for the teacher than using the same
evaluation information for deciding on multiple levels of performance. Use of a letter
system for grading clinical practice, however, acknowledges the different levels of
clinical proficiency students may have demonstrated in their clinical practice.
A disadvantage of pass–fail for grading clinical practice is the difficulty of
including a clinical grade in the course grade. One strategy is to separate nursing courses
into two components for grading, one for theory and another for clinical practice
(designated as pass–fail), even though the course may be considered as a whole.
Typically, guidelines for the course indicate that the students must pass the clinical
component to pass the course. An alternative mechanism is to offer two separate courses
with the clinical course graded on a pass–fail basis or by using a letter system. Once the
grading system is determined, there are various ways of using it to arrive at the clinical
grade. In one method, the grade is assigned based on the outcomes or competencies
achieved by the student. To use this method, the teacher should consider designating
some of the outcomes or competencies as critical for achievement.
Teachers will be faced with determining when students have not met the
outcomes of clinical practice, that is, have not demonstrated sufficient competence to
pass the clinical course. There are principles that should be followed in evaluating and
grading clinical practice, which are critical if a student fails a clinical course or has the
potential to fail it. The evaluation methods used in a clinical course, the manner in which
each will be graded if at all, and how the clinical grade will be assigned should be
documented in writing and communicated to the students. The practices of the teacher in
evaluating and grading clinical performance must reflect this written information. In
courses with preceptors, it is critical that preceptors and others involved in teaching and
assessing student performance understand the outcomes of the course, the evaluation
methods, how to observe and rate performance, and the preceptor’s responsibilities when
students are not performing adequately. Preceptors are reluctant to assign failing grades
to students whose competence is questionable (Anthony & Wickman, 2015). There is a
need for faculty development, especially for new and part-time teachers. As part of this
education teachers should explore their beliefs and values about grading clinical
performance and their expectations of students in the clinical setting.
Students should sign any written clinical evaluation documents—notes about the
student’s performance in clinical practice, rating forms (of clinical practice, clinical
examinations, and performance in simulations), narrative comments about the student’s
performance, and summaries of conferences in which performance was discussed. Their
signatures do not mean they agree with the ratings or comments, only that they have read
them. Students should have an opportunity to write in their own comments. These
materials are important because they document the student’s performance and indicate
that the teacher provided feedback and shared concerns about that performance. This is
critical in situations in which students may be failing the clinical course because of
performance problems. Students need continuous feedback on their clinical performance.
Observations made by the teacher, the preceptor, and others, as well as evaluation data
from other sources, should be shared with the student. Performance data should be
discussed together. Students may have different perceptions of their performance and in
some cases may provide new information that influences the teacher’s judgment about
clinical competencies. When the teacher or preceptor identifies performance problems
and clinical deficiencies that may affect passing the course, conferences should be held
with the student to discuss these areas of concern and develop a plan for remediation. It is
critical that these conferences focus on problems in performance combined with specific
learning activities meant to address them. The conferences should not consist of the
teacher telling the student everything that is wrong with his or her clinical performance;
the student needs an opportunity to respond to the teacher’s concerns and identify how to
address them.
One of the goals of the conference is to develop a plan with learning activities for
the student to correct deficiencies and develop competencies further. This plan serves as a
learning contract, an agreement between the teacher and student. A learning contract
specifies the clinical competencies to be developed by the student, learning activities
planned collaboratively by the teacher and student to guide learning and improve
performance, and expected outcomes with “due dates.” Exhibit 17.1 is an example of a
format that can be used to develop a learning contract for any level of learner. If the
student is failing clinical practice, the contract should indicate that (a) completing the
remedial learning activities does not guarantee that the student will pass the course, (b)
one satisfactory performance of the competencies will not constitute a pass clinical grade
(the improvement must be sustained), and (c) the student must demonstrate satisfactory
performance of the competencies by the end of the course.
Any discussions with students at risk of failing clinical practice should focus on
the student’s inability to achieve the outcomes of the clinical course and perform the
specified competencies, not on the teacher’s perceptions of the student’s intelligence,
overall ability, or perceived motivation or effort. In addition, opinions about the student’s
ability in general should not be discussed with others. Conferences should be held in
private, and a summary of the discussion should be prepared. The summary should
include the date and time of the conference, who participated, areas of concern about
clinical performance, and the learning plan with a time frame for completion (Oermann,
Shellenbarger, & Gaberson, 2018). The summary should be signed by the teacher, the
student, and any other participants. Faculty members should review related policies of the
nursing education program because they might specify other requirements.
As the clinical course progresses, the teacher should give feedback to the student
about performance and continue to guide learning. It is important to document the
observations made, other types of evaluation data collected, and the learning activities
completed by the student. The documentation should be shared routinely with students,
discussions about performance should be summarized, and students should sign these
summaries to confirm that they read them. The teacher cannot observe and document the
performance only of the student at risk for failing the course. There should be a minimum
number of observations and documentation of other students in the clinical group, or the
student failing the course might believe that he or she was treated differently than others
in the group. One strategy is to plan an approximate number of observations of
performance to be made for each student in the clinical group to avoid focusing only on
the student with performance problems. However, teachers may observe students who are
believed to be at risk for failure more closely, and document their observations and
conferences with those students more thoroughly and frequently than is necessary for the
majority of students. When observations result in feedback to students that can be used to
improve performance, at-risk students usually do not object to this extra attention.
There should be a policy in the nursing program about actions to be taken if a
student’s work in clinical practice is unsafe. Students who are not meeting the outcomes
of the course or have problems performing some of the competencies can continue in the
clinical course as long as they demonstrate safe care. This is because the outcomes and
clinical competencies are identified for achievement at the end of the course, not during
it. If the student demonstrates performance that is potentially unsafe, however, the
teacher can remove the student from the clinical setting when following the policy and
procedures of the nursing education program. Specific learning activities outside of the
clinical setting need to be offered to help students develop the knowledge and skills they
lack; simulation and practice in the skills laboratory are valuable in these situations. A
learning plan should be prepared and implemented as described earlier.
In all instances, the teacher must follow the policies of the nursing program. If the
student fails the clinical course, the student must be notified of the failure and its
consequences as indicated in these policies. In some nursing education programs,
students are allowed to repeat only one clinical course, and there may be other
requirements to be met. If the student will be dismissed from the program because of the
failure, the student must be informed of this in writing. Generally, there is a specific time
frame outlined for each step in the process, which must be adhered to by the faculty,
administrators, and students. It is critical that all teachers know the policies and
procedures to be implemented when students have performance problems or are at risk
for failing the clinical course. These policies and procedures must be followed for all
students.
A number of the procedures used to determine grades are time-consuming to use,
particularly if the class of students is large. Although a calculator may be used, student
grades can be calculated easily with a spreadsheet application such as Microsoft Excel or
in an online learning management system. With a spreadsheet application, teachers can
enter individual scores, include the weights of each component of the grade, and compute
final grades. Many statistical functions can be performed with a spreadsheet application.
Learning management systems provide grade books for teachers to manage all aspects of
student grades. The grades can be weighted and a final grade calculated. One advantage
to using a learning management system grade book is that students usually have online
access to their own scores and grades as soon as the teacher has entered them. There are
also a number of grading software programs on the market that include a premade
spreadsheet for grading purposes; these have different grading frameworks that may be
used to calculate the grade and enable the teacher to carry out the tasks needed for
grading. Not all grading software programs are of high quality, however, and should be
reviewed prior to purchase.
Students also viewed