Wk 3, IOP 490: Creating a Change Strategy

profileMesamada
TheEffectsofTemporalPlacementofFeedbackonPerformance.pdf

ORIGINAL ARTICLE

The Effects of the Temporal Placement of Feedback on Performance

Nathan T. Bechtel & Heather M. McGee & Bradley E. Huitema & Alyce M. Dickinson

Published online: 19 March 2015 # Association for Behavior Analysis International 2015

Abstract The purpose of this pilot study was to compare the effects of the temporal placement of feedback on task perfor- mance and skill acquisition. Two temporal placements were examined: feedback immediately after and feedback immedi- ately prior to performance. A two-factor mixed design was used. Participants were randomly assigned to one of three groups, which differed in the order of condition implementa- tion. Participants performed a computerized data entry task. The primary dependent variable was the number of correctly completed patient records per session. During feedback con- ditions, participants were provided with individual, graphic feedback and no feedback was provided during baseline. The results of this study indicate no significant differences in performance or the speed of skill acquisition associated with the experimental conditions. Participants indicated a strong preference for any type of feedback over no feedback, as well as a strong preference for feedback prior to performance over feedback after performance.

Keywords Feedback delivery . Skill acquisition .

Performance improvement . Feedback preference

The use of performance feedback is seen in a variety settings, including organizations, schools, sports teams, and clinics, to improve performance (Brobst and Ward 2002; Codding et al. 2005; Isaacs et al. 1982). Reviews of the associated literature within the field of organizational behavior management (OBM), industrial organizational (I/O) psychology, and ap- plied behavior analysis (ABA), show the extent to which feed- back is used (Balcazar et al. 1989; Nolan et al. 1999). Balcazar

et al., in a review of the first ten years (1974-1984) of the Journal of Organizational Behavior Management (JOBM), determined that performance feedback was used in 50 % of the articles. Furthermore, Nolan et al. found that the use of performance feedback increased to 71 % over the following 14 years (1985-1998).

The importance of performance feedback is also visible in the feedback literature reviews that traverse four journals (i.e., Journal of Organizational Behavior Management, Journal of Applied Psychology, Journal of Applied Behavior Analysis, and Academy of Management Journal), and 25 years of appli- cation between them (Alvero et al. 2001; Balcazar et al. 1985- 1986). Balcazar et al. found that these four journals yielded 126 applications of feedback over 11 years (1974-1984) and Alvero et al. found 68 applications within the same journals over 14 years (1985-1998). Feedback is, by a widemargin, the most prevalent intervention in the field of OBM, and a highly prevalent intervention in the field of ABA, due primarily to its ease of use, low cost, flexibility, and programmatic simplicity (Prue and Fairbank 1981); however, little is known about the behavioral principles affecting its use.

While feedback is used extensively to improve perfor- mance, it must be noted that feedback can vary along numer- ous dimensions, including (a) the recipients of feedback (e.g., group, individual), (b) feedback mechanisms (e.g., written, graphic, verbal), (c) content of feedback (e.g., comparison of performance with past performance, comparison of perfor- mance with group performance), (d) the source of feedback (e.g., supervisor, researcher), (e) temporal characteristics of feedback (e.g., duration, latency) (Prue and Fairbank 1981). These factors, and their effects on the efficacy of feedback, are examined within the two major feedback literature reviews discussed previously (Alvero et al. 2001; Balcazar et al. 1985-1986).

Absent from both reviews is the temporal placement of the feedback relative to the performance. There are four logical

N. T. Bechtel (*) :H.M.McGee :B. E. Huitema :A.M.Dickinson Department of Psychology, Western Michigan University, 3700 Wood Hall, Kalamazoo, MI 49008-5439, USA e-mail: [email protected]

Psychol Rec (2015) 65:425–434 DOI 10.1007/s40732-015-0117-4

reasons for this omission: (a) the research articles examined did not specify the temporal placement of feedback, (b) the majority of feedback applications provided feedback at the same time relative to the performance, (c) the temporal place- ment of feedback had no effect on feedback efficacy, and/or (d) the importance of the temporal placement of feedback was not recognized by the authors. Two studies that specifically address the efficacy of feedback delivered at different times (Alavosius and Sulzer-Azaroff 1990; So et al. 2013) were not included in the relevant literature reviews. Alvero et al. (2001) discuss adding feedback characteristic categories that were not presented in the original literature review conducted by Balcazar et al. (1985-1986). Temporal placement and feed- back timing were not included among these characteristics, which implies that there were no studies that warranted such a category. This suggests that the research articles examined did not specify the timing of feedback implementation in most cases. Whatever the reason(s) for omitting this factor, the fact remains that the temporal placement of feedback is a compo- nent which needs to be examined.

Determining the most appropriate temporal placement of feedback may help in determining the behavioral principles affecting feedback. Knowledge of feedback efficacy as it re- lates to temporal placement may also enable researchers and practitioners to maximize the efficacy of feedback and in- crease the celerity of acquisition of new skills. Previous re- search in this area has not used the term Btemporal placement^ to describe the timing of feedback, but rather Bimmediacy^ or Bfeedback timing^ (Alavosius and Sulzer-Azaroff 1990; Brewer 1989; Roberts 1997; So et al. 2013). For the purposes of this paper, temporal placement of feedback is defined as the point in time when feedback is delivered relative to the per- formance on which the feedback is based or the performance feedback is meant to effect. For example, feedback may be provided immediately after the performance of some task, in which case the temporal placement would be described in relation to that performance of the task (immediately after). However, if the feedback was provided immediately prior to the next performance of the same task, the temporal placement would be described in relation to the second occurrence of the task (immediately prior).

One factor related to the temporal placement of feedback has been examined in numerous research studies: immediacy of feedback. Immediacy of feedback refers to how quickly the feedback is provided in relation to the performance it follows. Daniels and Daniels (2004) state B…the general rule on feed- back is this: The sooner the better^ (p. 177). This is a common notion about feedback; however, research involving the im- mediacy of feedback has a tendency to providemore feedback (or more intensive feedback) to those in the immediate feed- back group, as opposed to the latent feedback group (Alavosius and Sulzer-Azaroff 1990; Goomas et al. 2011). Therefore, the heightened improvements for those receiving

immediate feedbackmay be caused by their repeated exposure to the feedback, rather than its temporal placement.

A study by Brewer (1989) attempted to determine the im- plications of the bi-functional theory on formative feedback. Brewer found that B…providing formative feedback immedi- ately prior to the occurrence of the targeted behavior will be more effective in cueing changes in the targeted behavior than if formative feedback is provided at other times^ (p. 47). This finding is dissimilar to many other studies pertaining to im- mediacy of feedback. This is most likely because (a) many of the other studies were actually focused on feedback frequency and/or (b) many of the other studies were aimed at increasing an existing response, rather than a new behavior.

Annett (1969) argued that because feedback is provided between two behaviors, the fact that it comes before one re- sponse may be as important as the fact that it comes after another response. In other words, feedback efficacy may be equally tied to both pre- and post-performance of the behavior in question. For example, a person receiving feedback on their teaching performance may require immediate positive feed- back regarding appropriate behaviors exhibited as well as feedback regarding behaviors they should avoid prior to the next occurrence of the behavior. The present account is a pilot study designed to assess the effects of the temporal placement of feedback on responding. Our hope is that this pilot research will motivate more robust future research examining the effi- cacy of pre- and post-performance feedback.

Method

Participants and Setting

Participants were 45 undergraduate students from a mid- western university, recruited using a combination of flyers and announcements in undergraduate-level psychology clas- ses. Prior to recruitment, Western Michigan University’s Hu- man Subjects Institutional Review Board approved the study and participants were required to read and sign an informed consent form.

Participants were screened based upon four exclusionary criteria. First, the participants had to report availability for 15 – 24 30-min sessions. Second, participants had to report using at least one of the available computer activities a minimum of 2 h per month. Alternative activities were included in order to replicate the work environment and provide realistic non- work activities. The available alternative activities were An- gry Birds, Solitaire, Spider Solitaire, Mahjong, and Bejew- eled, all of which were available on the computer used for the work task. To ensure that participants were not performing exemplary due to a lack of alternative activities, the partici- pants needed to report a certain level of interest in the alterna- tives. Interest in these activities was determined by asking

426 Psychol Rec (2015) 65:425–434

participants how many hours per week they spent engaging in the activities, with a threshold of 2 h being determined as an adequate level of interest. Third, participants could not have participated in any previous psychology department perfor- mance management research or held any positions involving data entry. Because this study intended to determine the ef- fects of feedback on skill acquisition, prior knowledge of the experimental task could easily confound any data related to learning the program. Lastly, participants were required to demonstrate a comprehensive understanding of the graphs being used to display the individual feedback in order to en- sure that performance changes were the result of the feedback being provided. Potential participants were provided with a brief quiz to determine their understanding of the graphs being used. To qualify for inclusion, participants had to obtain a score of 100 % on the quiz.

Apparatus

Sessions were conducted in a small university laboratory containing three work areas, separated by dividers. The work areas comprised a desktop computer, mouse, gel wrist-rest, keyboard, and adjustable chair. Participants were not allowed to bring in outside devices (e.g., phones, iPods) because we were unable to monitor off-task behav- iors that were not occurring on the computer itself. Off- task behaviors were only measured by the program if they occurred for at least 30 s, at which point a timer would track the amount of time before the next on-task behavior. If a participant were to respond to a text or check an email on their phone, these behaviors would most likely take less than 30 s and subsequently not be recorded as off- task. Multiple occurrences of such behaviors could have eventually skewed the on-task time. By limiting partici- pants’ access to such devices, we hoped to limit extrane- ous variables such as these from confounding our measurements.

The experimental task was a computer-based data entry task, intended to replicate the job of a medical data entry professional. The program instructed participants to perform two tasks per patient record: (a) type a patient’s identification number into a text box; and (b) determine whether the pa- tient’s heart rate (HR) is within a specified range. Typing the patient’s ID number required the participant merely to retype a number located on the side of the screen. Determining wheth- er the patient’s HR was within range required two separate steps. First, the participant ascertained the gender of the pa- tient by inspecting the box labeled Bgender^ on the side of the page. Second, the participant compared the patient’s HR to the HR range for his or her gender. The participant then clicked on a Bwithin range^ or Bout of range^ button, depending on the HR. This process was then repeated for the duration of the session.

Dependent Variables

The first dependent variable was the number of slides a par- ticipant correctly completed in each session, as recorded by the data entry program. The second dependent variable, used to analyze skill acquisition, was the rate of completion. Rate of completion was calculated as the number of correctly com- pleted slides per on-task minute in a session. On-task time was recorded as the amount of time that the participant actively used the program with less than 60 s of inactivity. Inactivity was measured as not clicking or typing in any interactive fields of the program. The experimental task was computer- ized and the program recorded all of the necessary variables automatically. To address specifically the effects of feedback timing, variables often used in conjunction with feedback, such as monetary incentives, goals, praise, and other social contingencies, were not used in this study.

Participants were also required to complete a brief ques- tionnaire at the conclusion of the study. The questionnaire was a seven-question survey designed to measure participants’ sat- isfaction and preference for the different feedback conditions. All questions were rated on a 5-point Likert scale. The data obtained from this survey were analyzed separately from the other dependent variables, and are considered a secondary dependent variable.

Design and Procedures

This experiment attempted to answer two experimental ques- tions: (a) Does the temporal placement of feedback affect overall participant performance on a data entry task?; and (b) Does the temporal placement of feedback affect the speed of skill acquisition on a data entry task?

The experimental design was a two-factor mixed design. Two factors were used to answer both experimental questions simultaneously with one study. The first factor was a between- groups factor. The groups within this factor differed only in the order of treatment introduction. Three groups were used, with each group employing a different condition order. Partic- ipants were randomly assigned to one of the three experimen- tal groups. The second factor was the experimental condition, each of which was introduced to all participants. There were three experimental conditions: (A) baseline, during which time no graphic feedback was provided; (B) feedback imme- diately after, during which graphic feedback was provided immediately after the performance on which it was based; and (C) feedback immediately prior, during which graphic feedback was provided immediately prior to the next perfor- mance of the task on which it was based. Participants were not informed of changes in experimental conditions. It is likely the participants were aware of the changes, however, because the timing of the feedback delivery was noticeably altered. One clear alteration in the sessions was an increased time

Psychol Rec (2015) 65:425–434 427

requirement for participants in the Bfeedback immediately after^ condition. This condition required roughly one to two extra post-session minutes during which time the participant’s graph was printed and delivered. The condition order of the three groups was (1) ABC, (2) BCA, and (3) CAB. Each group was exposed to all three conditions, but because the order of the conditions differed for each group, we use Bphase^ to describe the participants’ exposure to the condi- tions over time. In other words, phase one was condition A for the first group, condition B for the second group, and condi- tion C for the final group; phase two was condition B for the first group, condition C for the second group, and condition A for the final group; and phase three was condition C for the first group, condition A for the second group, and condition B for the final group. Because each experimental condition was the first condition experienced by at least one group, effects on skill acquisition could be determined. Although there were six possible grouping orders (ABC, ACB, BAC, BCA, CAB, and CBA) and using all of these orders would have been ideal, three groups were used for two reasons. First, the incomplete Latin square design, which uses three groups, was the most appropriate design to answer both experimental questions si- multaneously. Because the major concern regarding condition order was ensuring that each condition correlated with the beginning, middle, and end of a group’s exposure to the task, three groups were sufficient. Second, the resources required to run 90 participants as opposed to 45 participants were not readily available.

Graphic feedback consisted of a time-series graph that depicted the total number of slides correctly completed in each session by a participant, up to the point when the feedbackwas delivered. The graphs were printed by either the first author or one of the research assistants, and given to the participants at the appropriate time, depending on the experimental condition to which they were being exposed. No written or verbal feed- back supplemented the graphs. Scripted responses to partici- pant questions regarding feedback were available; however, no participants asked any questions that required the use of scripts; therefore, these scripts were not used in the study.

During the introductory session, the investigator provided an explanation of the study and obtained informed consent. If informed consent was obtained, potential participants were asked to complete the inclusion questionnaire and graph com- prehension quiz. If potential participants met the four inclu- sion criteria, they were scheduled for their first session. All participants (regardless of inclusion) were assigned a six-digit participant number to ensure anonymity. Participants who agreed to participate in the study and met the inclusion criteria were randomly assigned to one of three experimental groups. Participants were scheduled for two to five sessions per week. During the participant’s final session of each week, the next week’s sessions were scheduled. Participants remained under the same experimental condition until a stability criteria of +/-

10 % over three sessions occurred. If stability did not occur within eight sessions, the participant proceeded to the next experimental condition. This resulted in a slightly different number of total sessions for each participant and each condi- tion. The average number of sessions for each condition was approximately the same at 6.67, 6.4, and 6.27 for conditions A, B, and C, respectively. Similarly, the average number of sessions for each phase was approximately the same at 6.67, 6.33, and 6.29 for phases one, two, and three, respectively.

Participants were paid the equivalent of $4.50 per 30-min session in a single payment at the end of the study. Payment was provided at the end of the study to avoid confounding the Bfeedback immediately after performance^ phase with pay- ment. Payment given in combination with this form of feed- back, but not the Bfeedback immediately prior^ form, could have altered the effect the feedback had on performance.

Participants were free to take a break at any point during the sessions to prevent fatigue. A minimum performance cri- terion of 70 slides per session was also included to receive payment for that session. This criterion was implemented to replicate further a real work environment as well as to ensure that participants actually performed the task. If a person per- forms below some minimum level in any job, his or her em- ployment will most likely be terminated.

Upon completion of the last experimental session, partici- pants were scheduled for one, 5-10 min debriefing session. During this session, participants were provided with a com- pensation receipt and the money they had earned up to that point. The debriefing session included: (a) a description of the purpose of the study; (b) an explanation of the experimental phases; (c) an explanation of the use of feedback in the study; (d) a brief satisfaction survey used to determine participant preference for the different feedbacks; and (e) an opportunity for participants to ask any questions they may have had re- garding their participation. Following this, participants were debriefed on their participation in the study. After being debriefed, the participant’s obligation to the study was com- plete and they were free to leave.

Results

Figure 1 depicts the average number of correctly completed slides for each participant in each condition. A visual analysis of this figure indicates there is no distinct difference in perfor- mance between conditions. The average number of correctly completed slides in each phase can be seen in Fig. 2. Similarly, few robust statements can be made from a visual analysis of this figure. There is a noticeable increase in performance from Phase II to Phase III, but nothing extraordinarily noteworthy. A statistical analysis was completed on these data to analyze the results further.

428 Psychol Rec (2015) 65:425–434

A form of Latin square analysis of variance was used for the s ta t i s t i ca l ana lys i s . This des ign a l lows for counterbalancing while keeping the experiment size manage- able. In a complete three-condition counterbalanced design, six experimental groups would be required. The Latin square counterbalanced measures design allows us to reduce the number of groups to three, while still controlling for potential carry over effects. The number of correctly completed slides was the primary variable used in the Latin square ANOVA. The data met the assumptions of parametric tests (homogene- ity of variance, normality, and independence); therefore, non- parametric tests were not necessary.

Feedback prior to performance yielded the highest total average of correctly completed slides per session (214.9), followed by no feedback (206.8), and finally feedback after performance (202.6). However, statistical analysis indicated that the differences between the three treatment conditions (feedback prior, feedback after, or baseline) were not signifi- cant. For the conditions in the analysis, the Latin square ANOVA resulted in F(2, 42)=1.924, p=0.152. This finding indicates that participants performed similarly regardless of the feedback condition.

Differences between phases were analyzed next. Phase I had the lowest average number of slides completed, while Phases II and III had significantly higher averages. These data indicate that participants performed better after Phase I,

regardless of the experimental condition to which they were exposed. These results are supported by the statistical analysis of the data. The resulting statistics from the Latin square ANOVAwere F(2, 42)=14.94, p<.001. A Fisher-Hayter test was then conducted to determine between which pair(s) of phases the differences occurred. There was a statistically sig- nificant difference between Phases I and II (M=188.01 &M= 217.13, respectively), and Phases I and III (M=188.01 &M= 219.22, respectively). There was no statistically significant difference between Phases II and III (M=217.13 & M= 219.22, respectively). Thus, participants performed signifi- cantly better in the second phase and third phase of the exper- iment, regardless of the feedback they received in that phase, and there was no significant difference in performance be- tween the second phase and third phase, regardless of the feedback received in that phase.

The condition averages by group or treatment order were also analyzed. These data indicate that Group 2 (i.e., treatment order BCA) performed significantly better than Group 1 (i.e., ABC) and Group 3 (CAB), on average. A statistically signif- icant difference was found between experimental groups. The resulting statistics for the Latin square ANOVA were F(2, 42)=4.03, p=0.025. A Fisher-Hayter test was then conducted to determine between which pair(s) of groups the difference occurred. The only significant, pairwise difference was found between Group 2 and Group 3 (M=237.84 & M=208.12, respectively). There was no statistically significant difference between Group 1 and Group 2 (M=198.79 & M=237.84, respectively) or Group 1 and Group 3 (M=198.79 & M= 208.12, respectively). Thus, Group 2 performed significantly better than Group 3, and there was no significant difference in performance between Group 1 and Group 3 or Group 1 and Group 2.

To assess whether these group differences existed prior to any experimental manipulation, or were caused by a sequenc- ing effect, the average first session score for each group was calculated and analyzed. Average scores in the first session for each group were as follows: Group 1,M=5.18 slides per min- ute; Group 2, M=5.57 slides per minute; and Group 3, M= 4.83 slides per minute. A statistical analysis of these data indicates that Group 2 was more adept at the data entry task prior to the receipt of any feedback.

The secondary dependent variables (accuracy and time on- task) were also analyzed; however, these variables yielded no new information regarding the effects of the temporal place- ment of feedback. All groups averaged over 95 % accuracy in all conditions. This indicates that the task was simple enough that participants did not require feedback to maintain a high rate of accuracy. Rate of completion data was used to analyze the skill acquisition question (see BSkill Acquisition Analysis^ section). Average time on-task was also calculated for all participants in the study. The average time spent on-task was approximately M=26.2 min per session for all

C o n d it io

n A

C o n d it io

n B

C o n d it io

n C

0

100

200

300

400 A

v e

r a

g e

N u

m b

e r o

f

C o

r r e

c t ly

C o

m p

le t e

d S

li d

e s

Group 1

Group 2

Group 3

Fig. 1 Average number of correctly completed slides for each participant, separated by group and condition

P h a s e I

P h a s e I I

P h a s e I II

0

100

200

300

400

A v

e r a

g e

N u

m b

e r o

f

C o

r r e

c t ly

C o

m p

le t e

d S

li d

e s

Group 1

Group 2

Group 3

Fig. 2 Average number of correctly completed slides for each participant, separated by group and phase

Psychol Rec (2015) 65:425–434 429

participants, indicating that participants did not take exceed- ingly long breaks from the task (approximately 87 % of time spent on-task).

Skill Acquisition Analysis

Rate of completion was the primary variable used to an- alyze skill acquisition. The rate of completion for the first five sessions of the first condition in each group was used to calculate the simple linear regression lines for each experimental condition. These simple linear regression lines reflect the rate of skill acquisition for each condition. The individual regression lines for all participants were then compared using an ANOVA to determine the homo- geneity of slopes. An F(2, 42)=0.36, p=0.701 was found for the groups in the analysis, indicating that there was no statistically significant difference between conditions re- garding the factor of skill acquisition. In other words, no specific feedback condition produced faster learning than any other feedback condition.

Questionnaire Analysis

The complete results of the Feedback Satisfaction Survey can be seen in Table 1. A cross-question analysis (specifically an ANOVA) was conducted on Questions 5-7 to determine whether there was a statistically significant difference in feed- back preference.

Statistically significant differences were found between all three questions. An F(2, 132)=48.89, p=<0.001 was found when comparing these preference questions. A Fisher-Hayter test was then conducted to compare each pair of questions. All of the Fisher-Hayter results indicat- ed significant differences between the questions. Feed- back prior was strongly preferred to feedback after (M= 3.778 & M=3.0, respectively) and no feedback (M= 1.689), and feedback after was strongly preferred to no feedback (M=3.0 & M=1.689, respectively).

Discussion

Overall Performance

The primary purpose of this study was to determine whether the temporal placement of feedback altered the efficacy of feedback on performance. There were statistically significant differences in performance, but these differences appear to be due to performance improvements throughout the study (phase to phase increases) and intrinsic group differences, as opposed to the feedback type or the order in which the condi- tions were presented.

There are several possible reasons for this outcome. First, it is possible that the forms of feedback we evaluated in this study have no inherent differences and, therefore, produce similar effects on performance. The results indicate that feed- back provided prior and after performance were equally inef- fective because performance increases were comparable to those of baseline. However, it is possible that these results were because feedback was provided without any supplemen- tary reinforcement (i.e., no praise, evaluation, or incentives). These results complement those of Johnson et al. (2008), who found that objective feedback (i.e., feedback absent of any evaluative statements) did not improve performance on a check-processing task.

Another possible reason for the lack of differences between experimental conditions is the experimental task we used. The experimental task was relatively simple, and once participants became familiar with it, they could complete slides with little effort. This is supported by the finding that all groups demon- strated significant increases in performance after the first phase with smaller improvements for two of the groups be- tween the second and third phases, indicating an increase in fluency throughout the study. The simplicity of the task might have resulted in reinforcement for the behavior of engaging in the task, as very little response effort was required and the participants may have received natural feedback regarding the speed of typing.

The experimental task may have also provided other natural consequences. Upon typing in the patient’s ID

Table 1 Group averages of feedback satisfaction survey

Group Feedback Satisfaction Survey Questions

I tried to alter my performance based on my feedback graph

I worked harder after seeing my feedback graph

I enjoyed seeing my feedback graph each session

The sessions were not long enough for me to need a break

I preferred to receive my feedback at the beginning of the session

I preferred to receive my feedback at the end of the session

I preferred receiving no feedback at all

Group 1 4.067 3.333 3.933 3.333 4.067 2.800 1.867

Group 2 4.133 4.000 4.333 3.800 3.800 3.200 1.533

Group 3 4.000 4.133 4.200 3.333 3.467 3.000 1.689

Total 4.067 3.822 4.156 3.489 3.778 3.000 1.689

430 Psychol Rec (2015) 65:425–434

number, participants were able to see whether or not it matched the display and adjust it accordingly. Although the program did not tell participants the number of slides they had completed correctly, it is feasible that they were able to count the number of slides they completed. Be- cause participants were only required to complete 70 slides to earn their payment for each session, it is reason- able to suspect that some participants counted until they completed this number and then engaged in alternative activities for the remainder of the session. Two partici- pants whose data indicated such behavior were questioned by the experimenter during the debriefing session and acknowledged counting. While most participants’ data did not indicate such behavior, it is still a factor that bears consideration.

Lastly, the length of the sessions may have had an effect on participant performance levels. Sessions were only a half hour long, and it is likely that participants did not require a break during sessions. This assumption is supported by the time on-task data which shows that participants were on task an average of 26 min per ses- sion. This implies that participants were working at full effort throughout the majority of each session, regardless of the availability of alternative activities.

Skill Acquisition

The secondary purpose of this study was to determine if the temporal placement of feedback had an effect on skill acqui- sition. A statistical analysis comparing the feedback condi- tions determined that the differences in performance were not statistically significant, thus the temporal placement of feedback appeared to have no effect on the speed of skill acquisition for this task.

There are several possible reasons for this outcome. First, it is possible that the temporal placement of feedback has no effect on the speed of skill acquisition. While both forms of feedback increased the speed of skill acquisition slightly over the baseline condition, there was almost no difference between the effects of the two forms. Similar to the effects on overall performance, it is possible that these results were because feedback was provided without any supplementary reinforce- ment or evaluative statements.

The experimental task may have also contributed to the lack of differences in skill acquisition between experimental conditions. The experimental task was very easy to learn, and participants were capable of seeing correct typing perfor- mance as they completed the task. Therefore, the participants might not have necessarily required feedback to determine how well they were performing and subsequently increase their completion speed. The task also lent itself to counting as discussed above.

Questionnaire

According to the results of the questionnaire, participants pre- ferred to receive feedback prior to performance as opposed to feedback after performance or no feedback at all. Because the two feedback types did not differ in their effects on perfor- mance or skill acquisition, it is important to note these social validity results. Although performance was not significantly affected, participants’ feedback preference may have affected their overall satisfaction with the task or the experiment as a whole.

It is possible that participants preferred receiving feedback prior to performance due to the increased session time associ- ated with receiving feedback after performance. When partic- ipants were subject to feedback after performance condition, they were required to sit in the meeting room after each ses- sion for one to two extra minutes while the experimenter printed a feedback graph. The graphs for the feedback prior to performance condition were printed out before the arrival of participants, resulting in a time requirement that was essen- tially the same as baseline. Extrapolated across sessions, this extra time added to approximately six to 13 additional minutes in the feedback after performance condition (average of 6.4 sessions in this condition). This extra time may have been aversive and repeated pairings with the feedback after perfor- mance condition may have subsequently decreased participant preference for this condition. Additionally, participants were not compensated for this additional time, which may have further reduced preference. However, the baseline condition was not associated with this post-session waiting period and resulted in a significantly lower participant preference. This may indicate that the post-session waiting period was aversive enough to differentiate the two feedback conditions, but not so aversive as to increase preference for no feedback. Another possibility is that the lower overall rate of pay (due to the additional time across sessions) affected preference for the feedback conditions.

It is also important to note the real world implications of these results. Even if the post-session waiting time acted as a punisher, the same waiting period would be associated with feedback in a real world work environment. Presenting feed- back prior to the next performance of some task (e.g., the next day) would not only negate any post-session waiting period, but would also allow those providing the feedback more time to prepare any graphs or supplementary reinforcement. Thus, when implementing a feedback contingency in the workplace, it appears that feedback prior to performance is the more so- cially valid method of delivery.

The questionnaire results also indicated that participants enjoyed seeing the feedback graphs. This corroborates the social validity of providing feedback to performers, indicating that they enjoy seeing how well they are performing on a task. Participants also reported that they tried to alter their

Psychol Rec (2015) 65:425–434 431

performance based on their feedback graphs, and that they worked harder after seeing their feedback graphs.

Limitations

There were several limitations to this pilot study which should be noted. The first limitation involves the possible simplicity of the experimental task. The task may have been too simple to illuminate performance improvements adequately. The task required a minimal amount of comparison (patient heart rate was compared to the acceptable range). It may be appropriate to add one or more additional variables for comparison to increase the difficulty of the task. Participants also received some natural feedback from their typing behavior due to the fact that they could see what they were typing and compare it to the data presented by the program. A good counter-measure for this limitation would be to refine the program so that as- terisks appear on the screen when typing (similar to a pass- word bar on a website); however, this option has the limitation of reducing ecological validity, because data entry tasks like this do not normally have such a counter-measure in place. Another option would be to use a writing quality or critical thinking task. These alterations would reduce the amount of natural feedback participants receive, making them more reli- ant on experimenter-provided feedback graphs.

Related to this limitation is the argument that this study did not sufficiently replicate the real work environment, with the myriad of factors such as distractions and social interactions. Similarly, the length of the experimental sessions did not closely match that which would be found in a real work envi- ronment. Participants averaged 26 min on-task per session (87 % of time spent on-task). Locke (1986) provides numerous plausible arguments for the generalizability of laboratory stud- ies to the real-world work environment, however. Locke ar- gues that Bfield studies of the kind typically done by industrial-organizational psychologists are no more represen- tative of the real world than are laboratory studies^ (p. 4). Additional arguments for laboratory studies are provided by Mook (1983), who argues that, although laboratory studies are not always generalizable, they should not be justified based solely on this factor. Mook argues that lab settings are an excellent environment to manipulate phenomena under favor- able or unfavorable conditions that do not readily equate to the real-world environment.

Another limitation is that participants experienced different exposures to each of the experimental conditions depending upon the stability of their performance. Participants with less stable performance were likely to experience each condition for eight sessions, as opposed to those with highly stable per- formance who were likely to experience only five sessions for each condition. This may have had an effect on participant performance during each condition, or across multiple conditions.

Another possible limitation was the potential for carryover effects. Despite the counterbalancing of groups, it is impossi- ble to control for the permanent effects of feedback once it is provided. Although the groups may have been subject to a baseline condition after feedback was implemented, this does not guarantee that performance levels would decrease to those expected in a baseline condition. Without such a change in performance, it may not be possible to attribute performance changes to the feedback conditions.

Lastly, the design which was used may have affected the results. The skill acquisition question required three groups, each beginning with a different experimental condition. By using this design, we were able to determine the effects of the temporal placement of feedback on skill acquisition; how- ever, the primary question of effects on performancemay have been more easily answered with a multiple-baseline design.

Future Research

This pilot study lends itself to many possible future research and replication options. We recommend that replications and extrapolations of this study alter only one variable at a time to isolate the variables of interest. For instance, if future research were to alter the dependent variable used in this study, and similar results were found, it could be inferred that the depen- dent variable was not at issue. If numerous variables were altered and the same results achieved, it may be possible to conclude that temporal placement of feedback simply does not affect performance to a great extent. Some possible variables of interest and alterations are discussed below.

The lack of significant differences between feedback prior to performance and feedback after performance may be attrib- utable to many factors of the study. Future research on the effects of the temporal placement of feedback on performance may utilize a more difficult experimental task, or alter the current experimental task. One way to alter the current exper- imental task would be to program a cumulative recorder into the task. This would allow for detailed examination of within- session changes in performance. For example, it is possible that Bfeedback prior^ effected early-session performance more than late-session performance.

A task which provides little to no natural feedback, allows for longer experimental sessions, and reduces the possibility of counting behaviors as much as possible would be ideal. This type of task would allow the feedback conditions to exert greater control over performance by limiting other contingen- cies. Increased task difficulty is also essential to future re- search in this area. One possibility is a task which requires participants to complete multiple tasks concurrently. A task of this type was adopted by Bucklin et al. (2003). The task (SYNWORK) used four tasks simultaneously: memory, arith- metic, visual, and auditory monitoring. An altered task would also allow for a different method of feedback delivery.Making

432 Psychol Rec (2015) 65:425–434

feedback automated would help to ensure the precision of feedback delivery timing and ameliorate any human error in delivery. These program alterations would undoubtedly re- quire alterations to the feedback being presented, a factor that should be considered for any extension of this research.

Although the repetitious nature of the experimental task was appropriate for our research question, tasks that require true skill acquisition, beyond simple typing and clicking, may offer more opportunities to explore behavioral variations. It is possible that the simplicity of the task minimized differences in skill acquisition, whereas a more difficult task may have produced greater performance differences between conditions during initial learning. Another option for altering the exper- imental task would be gamification of everyday life routines, outside of the workplace. The efficacy of feedback before and after certain exercise, diet, and other routine behaviors could be examined in a simulated game-like environment. This op- tion could help increase ecological validity and decrease is- sues with the simplicity of a task.

Another option for future research would be to alter the experimental design. A combination between- and within- groups design was used to allow us to measure skill acquisi- tion. Excluding this particular dependent variable would allow for more options in experimental design which would in turn allow for easier analysis of the performance question. A multiple-baseline or altering treatments design could be im- plemented to analyze the effects of both types of feedback on performance. Another possibility is to design a study that fo- cuses on measuring skill acquisition instead of performance (i.e., correctly completed slides). This option would still re- quire a between-groups design because skill acquisition ef- fects are permanent. Three groups (baseline, feedback prior, and feedback after) would be required to make this viable. This would allow for more participants because only five ses- sions would be required per participant, as opposed to 15 – 24. As many as 135 participants could be recruited for the same cost as the 45 participants who were recruited for the current study, resulting in increased power in the statistical analysis.

Another area of concentration for future research would involve adding supplementary reinforcement to the feedback contingency. While this would admittedly reduce the amount of control feedback exerts over behavior, it is likely to increase the efficacy of conditions other than baseline. Feedback in the workplace is often related to some form of praise or incentive, so it is reasonable to study these contingencies. A pay for performance contingency would be an acceptable addition to the experiment, as long as participants still received their com- pensation at the end of the study as opposed to receiving it after each session. The feedback graphs could be altered to represent the amount of money earned, as well as the number of correctly completed slides. This method would hopefully reduce any direct effects of compensation on performance and cause feedback to result in improved performance.

Future research may also consider equalizing or yoking the time requirements across conditions. The additional one to two post-session minutes associated with the feedback after performance condition could be extended to the other two groups. This may result in altered preference for the different feedback conditions compared to those found with the present setup.

Lastly, future research could be conducted in an applied setting as opposed to a laboratory setting. Medicine, aviation, and industrial factory work offer many jobs wherein workers are provided with regular feedback, which lends itself well to this type of research. While this option would limit experi- mental control, it has numerous benefits that may outweigh the potential limitations. First, an applied setting would pro- vide a more realistic (and likely more difficult) task for partic- ipants to complete. While experimental tasks are generally developed to replicate real-world work environments, perfect replication is often impossible. Second, an applied setting would help determine the logistics of each type of feedback. While feedback delivered immediately after performance is often the preferred method, it is difficult to provide immedi- ately any graphic or comparative feedback after performance. Were feedback to be delivered prior to the next performance of some task, it would give the person delivering the feedback more time to create a visually stimulating graphic. This factor in itself could make a feedback prior condition more appealing to managers and supervisors. Lastly, the social validity ques- tion could be revisited in an applied setting. The results of the current study indicate that participants preferred receiving feedback prior to performance, and an applied replication could help determine if these results hold true across tasks and populations.

Acknowledgments This research was supported in part by a grant from the Graduate College at Western Michigan University.

References

Alavosius, M. P., & Sulzer-Azaroff, B. (1990). Acquisition and mainte- nance of health-care routines as a function of feedback density. Journal of Applied Behavior Analysis, 23, 151–162. doi:10.1901/ jaba. 1990.23-151.

Alvero, A. M., Bucklin, B. R., & Austin, J. (2001). An objective review of the effectiveness and essential characteristics of performance feedback in organizational settings (1985-1998). Journal of Organizational Behavior Management, 21(1), 3–29. doi:10.1300/ J075v21n01_02.

Annett, J. (1969). Feedback and human behaviour: The effects of knowl- edge of results, incentives and reinforcement on learning and performance. Oxford: Penguin Books.

Balcazar, F., Hopkins, B. L., & Suarez, Y. (1985–1986). A critical, ob- jective review of performance feedback. Journal of Organizational Behavior Management, 7(3), 65–89. doi:10.1300/J075v07n03_05.

Balcazar, F. E., Shupert, M. K., Daniels, A. C., Mawhinney, T. C., & Hopkins, B. L. (1989). An objective review and analysis of ten years

Psychol Rec (2015) 65:425–434 433

of publication in the Journal of Organizational Behavior Management. Journal of Organizational Behavior Management, 10(1), 7–37. doi:10.1300/J075v10n01_02.

Brewer, A. (1989). Feedback in training: Optimizing the effects of feed- back timing (Unpublished master’s thesis). The University of the Pacific, Stockton, CA.

Brobst, B., &Ward, P. (2002). Effects of public posting, goal setting, and oral feedback on the skills of female soccer players. Journal of Applied Behavior Analysis, 35(3), 247–257. doi:10.1901/jaba. 2002.35-247.

Bucklin, B. R., McGee, H.M., & Dickinson, A. M. (2003). The effects of individual monetary incentives with and without feedback. Journal of Organizational Behavior Management, 23(2), 65–94. doi:10. 1300/J075v23n02_05.

Codding, R. S., Feinburg, A. B., Dunn, E. K., & Pace, G. M. (2005). Effects of immediate performance feedback on implementation of behavior support plans. Journal of Applied Behavior Analysis, 38(2), 205–219. doi:10.1901/jaba. 2005.98-04.

Daniels, A. C., & Daniels, J. E. (2004). R+ performance management: Changing behavior that drives organizational effectiveness. Atlanta: Performance Management Publications.

Goomas, D. T., Smith, S. M., & Ludwig, T. D. (2011). Business activity monitoring: real-time group goals and feedback using an overhead scoreboard in a distribution center. Journal of Organizational Behavior Management, 31(3), 196–209. doi:10.1080/01608061. 2011.589715.

Isaacs, C. D., Embry, L. H., & Baer, D. M. (1982). Training family therapists: an experimental analysis. Journal of Applied Behavior Analysis, 15(4), 505–520. doi:10.1901/jaba. 1982.15-505.

Johnson, D. A., Dickinson, A. M., & Huitema, B. E. (2008). The effects of objective feedback on performance when individuals receive fixed and individual incentive pay. Performance Improvement Quarterly, 20(3/4), 53–74. doi:10.1002/piq.20003.

Locke, E. A. (Ed.) (1986). Generalizing from laboratory to field settings. Lexington, MA: Lexington Books.

Mook, D. G. (1983). In defense of external invalidity. American Psychologist, 38(4), 379–387. doi:10.1037/0003-066X.38.4.379.

Nolan, T. V., Jarema, K. A., & Austin, J. (1999). An objective review of the Journal of Organizational Behavior Management: 1987-1997. Journal of Organizational Behavior Management, 19(3), 83–114. doi:10.1300/J075v19n03_09.

Prue, D. M., & Fairbank, J. A. (1981). Performance feedback in organizational behavior management: a review. Journal of Organizational Behavior Management, 3(1), 1–16. doi:10. 1300/J075v03n01_01.

Roberts, P. (1997). Should corrective feedback come before or after responding to establish a Bnew^ behavior? (Unpublished master’s thesis). University of North Texas, Denton, TX.

So, Y., Lee, K., & Oah, S. (2013). Relative effects of daily feedback and weekly feedback on customer service behavior at a gas station. Journal of Organizational Behavior Management, 33(2), 137–151. doi:10.1080/01608061.2013.785898.

434 Psychol Rec (2015) 65:425–434

Copyright of Psychological Record is the property of Springer Science & Business Media B.V. and its content may not be copied or emailed to multiple sites or posted to a listserv without the copyright holder's express written permission. However, users may print, download, or email articles for individual use.

  • The Effects of the Temporal Placement of Feedback �on Performance
    • Abstract
    • Method
      • Participants and Setting
      • Apparatus
      • Dependent Variables
      • Design and Procedures
    • Results
      • Skill Acquisition Analysis
      • Questionnaire Analysis
    • Discussion
      • Overall Performance
      • Skill Acquisition
      • Questionnaire
      • Limitations
      • Future Research
    • References