Business
Relationships Among Team Ability Composition, Team Mental Models, and Team Performance
Bryan D. Edwards Tulane University
Eric Anthony Day University of Oklahoma
Winfred Arthur Jr. and Suzanne T. Bell Texas A&M University
This study examined the relationship between the similarity and accuracy of team mental models and compared the extent to which each predicted team performance. The relationship between team ability composition and team mental models was also investigated. Eighty-three dyadic teams worked on a complex skill task in a 2-week training protocol. Results indicated that although similarity and accuracy of team mental models were significantly related, accuracy was a stronger predictor of team performance. In addition, team ability was more strongly related to the accuracy than to the similarity of team mental models and accuracy partially mediated the relationship between team ability and team performance, but similarity did not.
Keywords: team mental models, knowledge structures, team ability, team training, team performance
Many current theories emphasize the importance of team mental models1 to team training and performance (e.g., Mohammed & Dumville, 2001). Mental models are based on the premise that people organize information into patterns that reflect existing relationships between concepts and the features that define them (Johnson-Laird, 1983). As such, the measurement of mental mod- els goes beyond the amount of declarative knowledge acquired but instead refers to an organized understanding or mental represen- tation of that knowledge (Cannon-Bowers, Salas, & Converse, 1993; Klimoski & Mohammed, 1994). It has been hypothesized that team mental models enable team members to form common expectations, coordinate actions, adapt their behaviors to task
demands, facilitate information processing, provide support, and diagnose deficiencies. As such, team mental models are an emer- gent characteristic of teams that influences both team processes (e.g., communication, conflict) and team outputs (e.g., perfor- mance; Klimoski & Mohammed, 1994; Kraiger & Wenzel, 1997; Marks, Mathieu, & Zaccaro, 2001).
Much discussion has been devoted to the number of different ways team mental models can be operationalized (e.g., Cooke, Salas, Cannon-Bowers, & Stout, 2000; Mohammed, Klimoski, & Rentsch, 2000). For instance, the similarity or sharedness of team mental models generally describes the degree to which members’ mental models are similar or overlapping. Another characteristic of team mental models is accuracy, which refers to the degree to which members’ mental models adequately represent a given knowledge or skill domain. Researchers (e.g., Stout, Salas, & Kraiger, 1997) have acknowledged that there is sparse research examining measures of team mental model accuracy. Indeed, we were able to locate only two studies (Marks, Zaccaro, & Mathieu, 2000; Webber, Chen, Payne, Marsh, & Zaccaro, 2000) that com- pared the contributions of similarity and accuracy to team perfor- mance. Therefore, our objective in the present study was to exam- ine the relationship between the similarity and accuracy of team mental models and also to compare the unique contribution of each in the prediction of team performance using a longitudinal design. Additionally, consonant with calls for research on effective team composition strategies for the development of team mental models (Mathieu, Heffner, Goodwin, Salas, & Cannon-Bowers, 2000), another objective was to investigate the relationships among team
1 In addition to the label mental models, several other labels have been used to describe the construct of knowledge organization such as knowl- edge structures, schemas, cognitive maps, and conceptual frameworks (Klimoski & Mohammed, 1994).
Bryan D. Edwards, Department of Psychology, Tulane University; Eric Anthony Day, Department of Psychology, University of Oklahoma; Win- fred Arthur Jr. and Suzanne T. Bell, Department of Psychology, Texas A&M University.
Suzanne T. Bell is now in the Department of Psychology at DePaul University.
An earlier version of this article was presented at the 15th Annual Conference of the Society for Industrial/Organizational Psychology, New Orleans, Louisiana, April 2000.
This research was sponsored under contract to Winfred Arthur Jr. from the U. S. Air Force Research Laboratory, Warfighter Training Research Division, Williams Air Force Base, Arizona. The views expressed herein are our own and do not necessarily reflect the official position or opinion of our respective organizations.
Correspondence concerning this article should be addressed to Bryan D. Edwards, Department of Psychology, 2007 Percival Stern Hall, Tulane University, New Orleans, LA 70118 or to Winfred Arthur Jr., Department of Psychology, Texas A&M University, 4235 TAMU, College Station, TX 77843-4235. E–mail: [email protected] or [email protected]. After August 1, 2006, correspondence should be addressed to Bryan D. Edwards, Department of Psychology, 226 Thach, Auburn University, Auburn, AL 36849. E-mail: [email protected]
Journal of Applied Psychology Copyright 2006 by the American Psychological Association 2006, Vol. 91, No. 3, 727–736 0021-9010/06/$12.00 DOI: 10.1037/0021-9010.91.3.727
727
ability, team mental model similarity and accuracy, and team performance. Finally, we tested the proposition that team mental models would mediate the relationship between team ability and team performance.
Similarity and Accuracy of Team Mental Models
In spite of several calls for the measurement of both the simi- larity and accuracy of team mental models (e.g., Cannon-Bowers & Salas, 2001; Cooke et al., 2000) surprisingly, the preponderance of the team mental models research (e.g., Converse, Cannon- Bowers, & Salas, 1991; Mathieu et al., 2000; Rentsch & Hall, 1994) has focused only on similarity. We identified three probable explanations for this apparent emphasis on similarity and limited attention to accuracy. First, the emphasis on similarity is consistent with the previous studies’ focus on team process variables because the formation of similar mental models serves as a critical process for achieving effective communication and coordination, which subsequently results in overall improved team performance (Mathieu et al., 2000).
Second, the assessment of team mental model accuracy requires a known “true state of the world” against which a team’s model is compared. However, it would seem that most of the tasks used in the extant literature (e.g., decision making and other tasks for which there is no clearly specified single correct way to perform the task) do not readily lend themselves to generating “true” scores. For instance, instead of using an expert referent model, Marks et al. (2000) had raters judge the accuracy of teams’ concept maps. They argued that there may be multiple correct mental models of the given task, which precludes using a single expert referent model (i.e., “true” score) to represent a single best way to perform the task.
Third, and related to the preceding, measuring accuracy can be a methodological challenge. For instance, Webber et al. (2000) assessed accuracy by comparing team member mental models to an average expert rating that served as the “true” score. However, the measurement and operationalization of mental models by Web- ber et al. was somewhat different from typical conceptualizations of mental models (e.g., Klimoski & Mohammed, 1994). Specifi- cally, they obtained ratings of the appropriateness of 17 actions or strategies that indicated if they were effective, ineffective, or neutral strategies, but there was no measurement of how these ratings represented relationships among the actions and/or strategies.
Given the issues identified above, we consider the present study to be a constructive replication of previous research because it compared the similarity and accuracy of only taskwork mental models. This is in contrast to Marks et al. (2000) who combined both teamwork (e.g., reporting what other team members are doing) and taskwork (e.g., shooting pillbox, hiding in forest) components in the measurement of team mental models and did not separate them. We also extended the literature by measuring team mental models at two points in time in a longitudinal design, training teams for a longer period of time (i.e., 10 hr over a 2-week interval vs. 1–3 hr), and assessing the effects of team ability composition. We also used a different operationalization of accu- racy by comparing trainees’ mental models to an expert referent model that served as the “true” score (see Day, Arthur, & Gettman, 2001, for an example at the individual level).
Contrary to Marks et al. (2000) and Webber et al. (2000), who found a predictive advantage for similarity over accuracy, we posited that there are many tasks for which accuracy would be expected to be a stronger predictor of performance than would similarity. Specifically, accuracy is particularly important in situ- ations or tasks in which there is one best way or a limited set of effective strategies or ways to successfully perform the task—a feature that characterized the task used in the present study. In their analysis of the task used in the present study, Frederiksen and White (1989) identified a single set of internally consistent strat- egies that distinguish expert performance. The importance of ac- curacy is based on the rationale that although team members may have shared knowledge, it is possible for this shared knowledge to be inaccurate. As argued by Acton, Johnson, and Goldsmith (1994), when individuals develop expertise, their mental models approximate an expert model and therefore increase in accuracy. Furthermore, as team members develop expertise and their repre- sentations converge on the “true” mental model, similarity is expected to increase as accuracy also increases. Thus, team mem- bers can have similar but not accurate mental models, although teams with accurate mental models will by definition have similar mental models. So, in some instances, it is plausible that similarity and accuracy could have unique relationships with performance.
A characteristic of past research is that teams have typically been trained for limited amounts of time (i.e., 1–3 hr). Thus, these studies have assessed mental models after very short training sessions. So, as previously noted, it is conceivable that similarity and accuracy will display different patterns of relationships over longer time frames of training (Acton et al., 1994). In summary, on the basis of the preceding conceptual arguments, we tested the following hypotheses.
Hypothesis 1a: The similarity of team mental models will increase over time.
Hypothesis 1b: The accuracy of team mental models will increase over time.
Because there is a clearly defined set of optimal performance strategies for the task used in this study (Frederiksen & White, 1989), we also tested the following hypothesis.
Hypothesis 2: The relationship between team mental model accuracy and team performance will be stronger than the relationship between team mental model similarity and team performance
Team Ability Composition and Team Mental Models
Researchers have stressed the need for more research that de- scribes effective team composition strategies for the development of team mental models (e.g., Mathieu et al., 2000). Kraiger and Wenzel (1997) argued that of all the theoretical antecedents of team mental models, “individual differences are the most proximal variables to shared mental models and thus may have an important impact in their development” (p. 77). Day et al. (2001) showed that ability was related to mental model accuracy for individuals, but no research has examined the relationship between team ability and team mental models.
728 RESEARCH REPORTS
Given that general mental ability is related to performance through knowledge acquisition (Ree, Carretta, & Teachout, 1995; Schmidt & Hunter, 1992), it is reasonable to posit that team ability should be significantly related to team performance through the development of accurate mental models. That is, teams consisting of members with high mental ability should develop more accurate mental models and consequently have higher team performance than teams with low-ability members. As team members’ mental models converge on the “true” score (i.e., became more accurate), they should also become more similar, such that higher ability teams should also develop similar mental models and have higher team performance than those of lower ability teams. In the present study we manipulated the ability composition of teams by creating teams of uniformly high (HH), mixed (HL), and uniformly low (LL) ability members on the basis of general mental ability. We expected HH teams to have more accurate and similar mental models than HL and LL teams. In addition, we expected that team mental model similarity and accuracy would mediate the relation- ship between team ability and team performance. Accordingly, we tested the following hypotheses.
Hypothesis 3a: The mental models of HH teams will be more similar than the models of HL teams, which in turn, will be more similar than the models of LL teams.
Hypothesis 3b: The mental models of HH teams will be more accurate than the models of HL teams, which in turn will be more accurate than the models of LL teams.
Hypothesis 4a: The similarity of team mental models will mediate the relationship between team ability composition and team performance.
Hypothesis 4b: The accuracy of team mental models will mediate the relationship between team ability composition and team performance.
Method
Participants
An initial pool of 1,266 male volunteers from Texas A&M University and its community were recruited via advertisements on campus and in local newspapers. Because of hardware constraints, participation was lim- ited to right-handed volunteers. Taking the Raven’s Advanced Progressive Matrices (APM; Raven, Raven, & Court, 1998) as a measure of general mental ability, we used ability scores to create HH, HL, and LL teams. Individuals were retained if they scored 21 or lower or 27 or higher on the APM. These cut scores represented one standard error of measurement above and below the mean APM score on the basis of a standardization sample (N � 496) of college men. This approach ensured that the low- and high-ability participants would indeed be categorically different. At the conclusion of the screening process, 194 individuals were selected and randomly assigned, within ability level, to high-, mixed-, and low-ability teams. Trainees were assigned the same partner throughout training. Twenty-eight of these participants did not complete the study, resulting in an attrition rate of 14%. Chi-square tests indicated no differences in attrition across the three levels of team ability. Thus, the final sample size was 166, which translated into 83 dyadic teams (i.e., 30 HH, 31 HL, and 22 LL dyads). The mean age of the final sample was 19.62 (SD � 2.30).
Participants were paid $75 to participate in a total of 10 days of training. Participants trained for one hr per day on Monday–Friday for 2 consecutive
weeks. Participants also had the opportunity to receive a bonus of $50, $30, or $20, which was awarded to each member of the teams with the three highest team scores on the performance task.
Measures
Raven’s APM. The APM is a measure of general mental ability con- sisting of 36 design problems arranged in an ascending order of difficulty. Because its stimuli are nonverbal and do not require a specific knowledge base to be understood, the APM offers scores that are posited to be uninfluenced by a respondent’s acquired knowledge or reading ability (Saccuzzo & Johnson, 1995). This has led experts to conclude that it is one of the better measures of general mental ability (e.g., Carpenter, Just, & Shell, 1990). We used an administration time of 40 min and obtained a Spearman–Brown odd-even split-half reliability of .84.
Space Fortress. The performance task was the video game Space Fortress (Donchin, 1989; Mané & Donchin, 1989). Space Fortress is “an experimental game which was designed to simulate a complex and dy- namic aviation environment” (Gopher, 1993, p. 299). Space Fortress rep- resents important information-processing demands that are present in avi- ation and other complex tasks (Gopher, Weil, & Bareket, 1994; Hart & Battiste, 1992). These processing demands include short- and long-term memory load, high workload, dynamic attention allocation, decision mak- ing, prioritization, resource management, discrete motor responses, and difficult manual control elements (Gopher, Weil, & Siegel, 1989). Per- forming Space Fortress involves coordinating mouse and joystick functions to control a spaceship’s flight path and shoot missiles at a fortress. See Arthur et al.’s (1995) article for a more detailed description of Space Fortress.
Pathfinder (Schvaneveldt, 1990; Schvaneveldt, Durso, & Dearholt, 1989). Pathfinder, a computerized structural assessment technique that generates concept similarity maps, was used for the elicitation and analysis of mental models. Pathfinder is a network-scaling procedure (Schvan- eveldt, 1990) used to summarize and graphically display relatedness rat- ings. Pathfinder renders network structures that capture local relationships among concepts. The resulting networks are rich representations that can be quantified and compared (Goldsmith & Davenport, 1990). Two param- eters, r and q, determine how network distance is calculated and affect the density of the network. For the present study, the networks were derived with the parameters set to r � infinity and q � the number of concepts (or nodes) minus one.
To generate mental models, we used a set of 14 Space Fortress concepts from Frederiksen and White’s (1989) cognitive task analysis of Space Fortress. Respectively, Table 1 and the Appendix present the concepts and instructions used in the administration of Pathfinder. Trainees made relat- edness ratings on all possible pairs of the 14 Space Fortress concepts, which were presented sequentially and randomly to the trainees, resulting in a total of 91 ratings, n(n – 1)/2 � 91, where n � the number of concepts. For each pair of concepts, trainees were asked to indicate the extent to which they were related by using a 9-point Likert scale (1 � not at all related; 9 � highly related).
We used Pathfinder to generate the team mental model indices of similarity and accuracy. Similarity and accuracy were represented by two derivations of the Pathfinder metric of closeness (C), which represents the ratio of common links between two networks divided by the total number of links in both networks. We used C because it is considered to be superior to other Pathfinder metrics such as correlation and number of links (Gold- smith, Johnson, & Acton, 1991; Johnson, Goldsmith, & Teague, 1994) and is also the most commonly used Pathfinder index in the literature. We generated similarity by calculating C between team members’ individual network structures. Team mental model accuracy was defined by the degree of overlap between trainee mental models and an expert referent model. Specifically, we generated accuracy by taking the mean C between each team member’s structure and the expert referent structure. The values
729RESEARCH REPORTS
of C for similarity and accuracy range from 0 to 1, with 1 representing perfect similarity and accuracy.
To obtain the expert referent model, we asked three subject-matter experts to independently complete Pathfinder. These experts had previ- ously worked in research labs that used Space Fortress. In addition, they consistently achieved high scores (i.e., above 4,750) each time they per- formed Space Fortress. Subsequently, we averaged the three models within the Pathfinder program to yield one referent model, which had an average C of .49. The Cs for each comparison among the three expert models were .46, .58, and .44. Research has indicated that referent models derived from an average of expert judgments yield stronger correlations with perfor- mance than those based on a single expert’s judgment (Acton et al., 1994; Day et al., 2001).
Design and Procedure
Participation involved 10 days of training over a 2-week period. On the Monday of the first week, trainees began with 20 min of instructions that explained the rules and optimal strategies of Space Fortress. Trainees then performed two 3-min baseline Space Fortress games followed by a 5-min review of the instructions. Trainees then underwent 11 more Space Fortress training sessions, which took place over the next 13 days. Training sessions were not scheduled on weekends, and there were 2 training sessions on the Monday and Friday of the second week. Trainees completed Pathfinder at the end of Session 1 (i.e., Time 1, which corresponded to the 2nd day of the 10-day training protocol) and Session 3 (i.e., Time 2, which corresponded to the 4th day of training). Although mental model measurement at the end of training would have been optimal, we experienced administrative and computer problems with mental model data collection at Time 3 (i.e., the 9th day of training) that resulted in the loss of data for most teams. Consequently, we have only Time 1 and Time 2 data for all teams. Nevertheless, a 2-day interval between measurements in a 10-day training protocol is still a marked improvement over intervals in other longitudinal designs in which team mental model measurements were taken 20 –30 min
apart in a 1–3 hr training protocol (e.g., Marks et al., 2000; Mathieu et al., 2000).
During a standard training session, the teams performed six practice and two test games. All games lasted 3 min. For each practice and test game 1 trainee, using his left hand, controlled all functions related to the mouse (managing mines and missiles), and the other trainee, using his right hand, controlled all functions related to the joystick and trigger (piloting and firing the gun). Trainees alternated roles, which called for physically switching places, at the end of each game. Communication between train- ees was encouraged. A typical training and testing session lasted approx- imately 1 hr, and trainees were scheduled to train at the same 1-hr slot for their 2 weeks of participation. For each session, performance was opera- tionalized as the average of the total scores from the two test games.
Results
Relationship Between Similarity and Accuracy
Table 2 presents the descriptive statistics and intercorrelations among all study variables. The correlation between the similarity and accuracy indices was large and statistically significant for both Time 1 (r � .61, p � .01, 95% confidence intervals [CIs] � .45–.73) and Time 2 (r � .67, p � .01, 95% CI � .53 to .77) measurements. Although team mental model similarity and accu- racy increased from Time 1 to Time 2, the increases were not significant: similarity, t(82) � 1.70, p � .09, d � 0.17; accuracy, t(82) � 1.56, p � .12, d � 0.12. So Hypotheses 1a and 1b were not supported.
Comparative Criterion-Related Validity of Similarity and Accuracy
Table 3 reports the means and standard deviations for Space Fortress team performance for the baseline and all 11 training
Table 1 Space Fortress Concepts With Descriptions
Concept Description
1. Control ship speed and distance Maintenance of proper speed and distance from the fortress 2. Change trajectory of ship Direction of the ship in frictionless space 3. Correct press of IFF (mouse) Accurate identification and resultant response to the appearance
of a mine 4. Incorrect press of IFF (mouse) Inaccurate identification and resultant response to the
appearance of a mine 5. Select points bonus (mouse) Choice of this bonus increases points score 6. Select missiles bonus (mouse) Choice of this bonus increases missile supply 7. Friend or foe identifier
(instrument panel) Preestablished mine identifiers, shown on the instrument panel
8. INTRVL (instrument panel) Interval in milliseconds between mouse button presses indicating when a foe mine is vulnerable, shown on the instrument panel
9. Shots counter less than 50 (instrument panel)
Important information regarding selection of bonus shown on the instrument panel
10. Shots counter more than 50 (instrument panel)
Important information regarding selection of bonus shown on the instrument panel
11. Recognize second $ Important to the acquisition of bonuses 12. Scoring or losing points Objective indicator of task performance 13. Destroy or avoid mines Handling of friend and foe mines 14. VLNER (instrument panel) Number of hits the fortress has suffered that primes the player
to apply a “double shot” to destroy the fortress, as shown on the display panel
Note. IFF � identify friend or foe; INTRVL � interval; VLNER � vulnerability.
730 RESEARCH REPORTS
sessions along with the correlations between the team mental model indices and team performance. These correlations, which are also plotted in Figure 1, indicate that for both Time 1 and Time 2 mental model assessments, the relationships between accuracy and performance were consistently stronger than the relationships between similarity and performance. Indeed the pattern of results indicates that contrary to the relationships between accuracy and performance, the magnitude of the relationships between similarity and performance decreased with time. Thus, these results provided initial support for Hypothesis 2.
To permit a more succinct and parsimonious presentation of additional tests of Hypothesis 2, we computed the average of the Space Fortress performance scores across Sessions 4 –11. We included only Sessions 4 –11 in this team performance index because mental models were measured after training Sessions 1
(Time 1) and 3 (Time 2) and we were interested in the predictive criterion-related validity of the mental model indices. In addition, limiting the average team performance to Sessions 4 –11 allowed us to hold the time frames of the criterion variable (i.e., team performance) constant to permit interpretable comparisons across the predictors (i.e., team ability, mental model indices). Finally, although team performance substantially improved from Session 1 (M � 349.66, SD � 1,136.44) to Session 11 (M � 3,607.35, SD � 1,733.32), t(82) � 20.61, p � .01, d � 2.22, the decision to average across sessions was deemed appropriate given the magni- tude of the intercorrelations among the Space Fortress perfor- mance scores (mean r � .89).
Consistent with the pattern of results in Figure 1, analyses of the correlations presented in Table 2 indicated that the relations be- tween accuracy and performance (Time 1 r � .34, p � .01, 95%
Table 2 Descriptive Statistics and Intercorrelations for Study Variables
Variable M SD 1 2 3 4 5 6 7
1. Team ability — 2. Similarity Time 1 .33 .11 .15 — 3. Accuracy Time 1 .34 .08 .42 .61 — 4. Similarity Time 2 .35 .13 .36 .56 .50 — 5. Accuracy Time 2 .35 .08 .46 .53 .57 .67 — 6. Average team performance (Sessions 4–11)a 2,999.62 1,592.46 .47 .26 .34 .27 .46 — 7. Average team performance (All sessions)b 2,139.08 1,374.85 .47 .29 .37 .30 .47 .99 —
Note. N � 83 dyadic teams. LL � low-ability team; HL � mixed ability team; HH � high-ability team. Team ability was coded as LL � 1, HL � 2, HH � 3. Similarity � team mental model similarity at the given time (Time 1 or Time 2); Accuracy � team mental model accuracy at the given time (Time 1 or Time 2). If r � .21 to .28, then p � .05; if r � .29, then p � .01. a Average team performance was the average of Sessions 4 –11. b Average team performance was the average of the baseline and Sessions 1–11; this information is presented for the sake of completeness.
Table 3 Descriptive Statistics for Space Fortress Team Performance and Correlations Between Team Mental Model Indices and Team Performance
SF session M SD
Team mental model indices
Similarity Time 1
Accuracy Time 1
Similarity Time 2
Accuracy Time 2
Baseline –1,854.29 713.49 .26 .29 .27 .32 1 349.66 1,136.44 .29 .35 .28 .36 2 1,228.93 1,352.44 .32 .38 .35 .40 3 1,947.71 1,528.90 .30 .38 .27 .45 4 2,100.14 1,487.38 .30 .34 .33 .45 5 2,452.04 1,540.87 .30 .36 .24 .43 6 2,717.27 1,692.89 .23 .29 .26 .44 7 3,027.35 1,695.94 .23 .33 .27 .44 8 3,148.32 1,928.13 .25 .31 .22 .44 9 3,432.21 1,657.50 .24 .31 .26 .42
10 3,512.24 1,636.10 .23 .30 .26 .45 11 3,607.35 1,733.32 .25 .37 .25 .44 Average team performance
(Sessions 4–11)a 2,999.62 1,592.46 .26 .34 .27 .46 Average team performance
(All sessions)b 2,139.08 1,374.85 .28 .37 .30 .47
Note. All correlations were significant at p � .05. SF � Space Fortress. a Average team performance was the average of Sessions 4 –11. b Average team performance was the average of the baseline and Sessions 1–11; this information is presented for the sake of completeness.
731RESEARCH REPORTS
CI � .13 to .52; Time 2 r � .46, p � .01, 95% CI � .27 to .61) were stronger than those between similarity and performance (Time 1 r � .26, p � .05, 95% CI � .05 to .45; Time 2 r � .27, p � .01, 95% CI � .06 to .46). A test for differences between dependent rs revealed that although the difference between the similarity and accuracy correlations with performance was not significant at Time 1, t(80) � 0.86, ns, it was significant at Time 2, t(80) � 2.36, p � .05, lending additional support for Hypothe- sis 2.
To investigate the incremental variance explained in team performance by similarity and accuracy, we computed four hierarchical regression equations (see Table 4). In all models, we entered team ability composition in the first step of the hierarchical regression. In the first model, we entered similarity in the second step, followed by accuracy in the third step to examine the incremental validity of accuracy beyond team ability and similarity measured at Time 1. In the second model, we reversed the entry order of similarity and accuracy to examine the unique contribution of similarity beyond team ability and accuracy measured at Time 1. Models 3 and 4 were replications of the first two models using similarity and accu- racy measured at Time 2. The results revealed that team ability composition accounted for 22% of the variance in team perfor- mance. The results of Models 1 and 2 indicated that neither accuracy (Model 1 �R2 � .00, ns) nor similarity (Model 2 �R2 � .02, ns) explained unique variance in team performance beyond team ability and the other mental model index. How- ever, for Time 2 data, accuracy explained unique variance in team performance beyond team ability and similarity (Model 3 �R2 � .07, p � .01), but similarity did not explain unique variance in team performance (Model 4 �R2 � .01, ns).
Figure 1. Correlations between team similarity and accuracy measured at Time 1 and Time 2, and team performance measured at the baseline and Sessions 1–11; Sim T1 � Similarity measured at Time 1. Acc T1 � Accuracy measured at Time 1. Sim T2 � Similarity measured at Time 2; Acc T2 � Accuracy measured at Time 2. B � Baseline performance session.
Table 4 Summary of Hierarchical Regression Results Using Team Ability and Similarity and Accuracy of Team Mental Models to Predict Team Performance
Variable
�
R2 �R2Step 1 Step 2 Step 3
Time 1
Model 1 Team Ability .47** .44** .41** .22** Mental model similarity .20* .16 .26** .04* Mental model accuracy .07 .26** .00
Model 2 Team ability .47** .39** .41** .22** Mental model accuracy .17 .07 .24** .02 Mental model similarity .16 .26** .02
Time 2
Model 3 Team ability .47** .42** .33** .22** Mental model similarity .12 �.10 .23** .01 Mental model accuracy .37** .30** .07**
Model 4 Team ability .47** .32** .33** .22** Mental model accuracy .31** .37** .29** .07** Mental model similarity �.10 .30** .01
Note. N � 83 dyadic teams. LL � low-ability team; HL � mixed-ability team; HH � high ability team. Team ability was coded as LL � 1, HL � 2, HH � 3. The dependent variable was average team performance opera- tionalized as the average of Sessions 4 –11. * p � .05. ** p � .01.
732 RESEARCH REPORTS
Team Ability, Mental Models, and Team Performance
Hypothesis 3a predicted that the mental models of HH teams would be more similar than mental models of HL teams, which in turn would be more similar than the models of LL teams. Hypoth- esis 3b made the same predictions for accuracy. The results in Table 5 show that the similarity and accuracy of team mental models differed across the three ability compositions in the pre- dicted direction. A 3 (team ability) � 2 (mental models adminis- tration) analysis of variance showed a significant main effect for ability, F(2, 80) � 3.98, p � .05, �2 � .09, no significant main effect for mental models administration, F(1, 80) � 1.79, ns, �2 � .00, and a significant interaction, F(2, 80) � 3.22, p � .05, �2 � .02. Paired comparisons showed that there were no differences between the ability compositions for similarity measured at Time 1. However, for Time 2 there were significant differences for the HH–LL comparisons, t(50) � 3.70, p � .01, d � 0.96 and the HL–LL comparisons, t(51) � 2.34, p � .05, d � 0.58, although the HH–HL difference was not statistically significant, t(59) � 1.28, ns, d � 0.35. These results indicated that mental model similarity did not differ by team ability level when measured at Time 1, but after two more sessions of training, the HH and HL teams had developed more similar mental models than had the LL teams.
Regarding Hypothesis 3b, a 3 � 2 analysis of variance for accuracy showed a significant main effect for ability, F(2, 80) � 13.09, p � .01, �2 � .32, but the main effect for mental models administration, F(1, 80) � 2.37, ns, �2 � .00, and the interaction, F(2, 80) � 0.44, ns, �2 � .00, were not significant. Paired comparisons showed that for Time 1 there were significant differ- ences in mental model accuracy among all three ability composi- tions: HH–LL, t(50) � 3.94, p � .01, d � 1.08, HL–LL, t(51) � 2.54, p � .01, d � 0.73, HH–HL, t(59) � 1.91, p � .05, d � 0.46. Comparisons for the accuracy index measured at Time 2 also revealed statistically significant differences among all three team
ability compositions: HH–HL t(59) � 2.62, p � .01, d � 0.67; HL–LL t(51) � 2.31, p � .05, d � 0.54; HH–LL t(50) � 4.36, p � .01, d � 1.13. In general, the pattern of results demonstrated support for Hypotheses 3a and 3b in that the similarity and accu- racy of team mental models differed among all three team ability compositions. The only exceptions to this pattern of results were the nonsignificant effects for similarity measured at Time 1 and the difference between the HH and HL teams for similarity measured at Time 2.
Hypotheses 4a and 4b stated that mental models (similarity and accuracy, respectively) would mediate the relationship between team ability and team performance and were tested in accordance with standards outlined by Baron and Kenny (1986). We con- ducted the tests of mediation for similarity and accuracy sepa- rately, using mental models data collected at Time 2 because we considered these data to be more robust as they encompassed more training. The results of the similarity mediation test, which are presented as Model 3 in Table 4, indicate a small decrease in the effect of team ability after controlling for mental model similarity (� � .47 vs. � � .42). However, contrary to Hypothesis 4a, Sobel’s (1982) test showed that the indirect effect of team ability and team performance through mental model similarity was not significantly different from 0, t(81) � 1.09, ns; indirect effect � 88.83, 95% CI � �70.67 to 248.34.
The results of the accuracy mediation test, which are presented as Model 4 in Table 4, indicate a decrease in the effect of team ability when controlling for mental model accuracy (� � .47 vs. � � .32). An examination of the 95% CI for the indirect effect of team ability and team performance through mental model accuracy indicates that the effect was significantly different from 0, indirect effect � 284.40, 95% CI � 54.75 to 514.04; t(81)� 2.43, p � .05. Therefore, in support of Hypothesis 4b, mental model accuracy partially mediated the relationship between team ability and team performance.
Table 5 Descriptive Statistics for Mental Models and Team Performance of the Three Team Ability Compositions
Variable
Team ability composition
HH HL LL
F(2, 80) �2M SD M SD M SD
Team abilitya 29.98 1.85 23.95 1.99 18.45 1.88 234.34** .85 Similarity Time 1 .35 .12 .33 .11 .30 .10 0.90 .00 Accuracy Time 1 .37 .07 .34 .06 .29 .08 8.77** .16 Similarity Time 2 .39 .11 .35 .12 .28 .12 6.45** .12 Accuracy Time 2 .39 .08 .34 .07 .30 .08 10.60** .19 Average team performance
(Sessions 4–11)b 3,924.82 1,303.20 2,761.11 1,591.72 2,074.05 1,316.54 11.45** .21 Average team performance
(All sessions)c 2,968.46 1,158.79 1,895.65 1,372.67 1,351.10 1,055.51 12.16** .21
Note. HH � two high-ability members (n � 30 dyadic teams); HL � one high- and one low-ability member (n � 31 dyadic teams); LL � two low-ability members (n � 22 dyadic teams). The F values and eta squared were generated from five one-way analyses of variance using team ability composition as the independent variable and each index and average team performance as the dependent variables. a Team ability was the average Advanced Progressive Matrices scores of members within each ability compo- sition. b Average team performance was the average of Sessions 4 –11. c Average team performance was the average of the baseline and Sessions 1–11; this information is presented for the sake of completeness. ** p � .01.
733RESEARCH REPORTS
Discussion
Our findings support a growing body of research that indicates team mental models play an important role in the development of complex skills and subsequent team performance (Marks et al., 2000; Mathieu et al., 2000). Previous research has primarily fo- cused on the measurement of team mental model similarity at the exclusion of accuracy (e.g., Levesque, Wilson, & Wholey, 2001; Mathieu et al., 2000; Peterson, Mitchell, Thompson, & Burr, 2000; Rentsch & Klimoski, 2001). Therefore, one of our objectives in the present study was to compare the similarity and accuracy of team mental models in a longitudinal research design.
Although our results did not show a significant increase in similarity and accuracy over time—most likely because of the relatively short time interval between the two mental model ad- ministrations—the similarity and accuracy of team mental models were strongly related. We also showed that although the similarity and accuracy of mental model indices taken early in team training (i.e., at Time 1) were equally predictive of team performance, after 4 days of training (i.e., at Time 2), the accuracy of team mental models was a stronger predictor of subsequent team performance. Prior research on team mental models tends to favor similarity as the stronger predictor of team performance, but these data have been typically collected in designs with 1–3 hr training and per- formance sessions (e.g., Marks et al., 2000; Mathieu et al., 2000). Thus, the differences in our findings and those previously reported for similarity might be due to our use of a longer team-training time frame. Our longitudinal data suggest that we would have obtained results similar to those reported in the extant literature if we had terminated our data collection at the end of the first mental model assessment (i.e., after 2 hr of training and performance). However, our results (e.g., Figure 1) clearly show that as teams acquired more skill and converged on the “true” mental model with increased training (Acton et al., 1994), the comparative va- lidity of similarity and accuracy changed, with accuracy becoming a stronger predictor than similarity.
In contrast to previous research (e.g., Marks et al., 2000; Web- ber et al., 2000), the present study used a task in which there was a limited number of effective strategies and focused on only taskwork mental models. These features allowed us to obtain an expert referent mental model that served as the “true” score and subsequently operationalize team mental model accuracy as the degree of overlap between trainee mental models and an expert referent model. The differences in tasks and our use of longer training and performance time frames may be plausible explana- tions for the differences in our results and those of Marks et al. (2000), who used a decision-making task and 3 hr of training and performance and found stronger effects for mental model similarity.
Another contribution of the present study is its investigation of the relationships among team ability, team mental model similarity and accuracy, and team performance. Specifically, our results indicate that the similarity and accuracy of team mental models are related to team general mental ability. However, team ability is more strongly related to the accuracy than to the similarity of team mental models. Thus, we demonstrated that team ability is an important predictor of the accuracy and, to a lesser extent, the similarity of team mental models.
However, because we focused exclusively on taskwork mental models, our data do not speak to the role of team ability in the development of other forms of knowledge organization, such as teamwork models. It is conceivable that the comparative validity of similarity and accuracy may be a function of whether the focus is on taskwork versus teamwork. Nevertheless, we demonstrated that when the focus is on taskwork mental models only, accuracy of mental models is a better predictor of team performance than is their similarity. Now that we have established this boundary con- dition for taskwork mental models, future research could focus on measuring taskwork and teamwork mental models separately in a single study.
We also demonstrated that the accuracy of team mental models partially mediates the relationship between team ability and team performance. General mental ability is related to performance through knowledge acquisition (Schmidt & Hunter, 1992). Given that mental models are representations of knowledge in a given domain, it is not surprising that HH teams developed more accu- rate mental models and subsequently higher team performance.
Although a strength of the present study is that teams partici- pated in a much longer training protocol (2 weeks) than in most team training laboratory studies (typically 1–3 hr), the external validity of our findings may still be somewhat restricted because of the limited life span of our teams. In spite of this, the results of the present study are most likely to generalize to teams that perform tasks for which there is a demonstrable best or limited set of effective strategies. In contrast, for tasks with multiple correct ways or effective strategies, there are likely to be multiple, accu- rate team mental models. Consistent with the concept of equifi- nality, it would be difficult, if not impossible, to determine the definitive accurate mental model. Consequently, under these con- ditions, similarity (and teamwork mental models) may be more important than accuracy.
Also, although we used dyadic teams and larger teams have more complex dynamics than dyads, it is not unreasonable to posit that our findings may generalize to larger teams responsible for tasks similar to the one used in the present study; that is, tasks for which there is a limited number of optimal strategies. However, further research is needed to test this proposition. Finally, our results have additional implications for research and the training and development of teams in the field. First, they would suggest that wherever possible, one could train for accuracy with the expectation that similarity would follow. Second, where it is possible to generate expert referent mental models, one could investigate their efficacy as interventions to facilitate the develop- ment of accurate and shared team mental models. Third, the mental models of trainees could be assessed during training, and trainers could use the expert model as a means of providing corrective feedback. Fourth, our results suggest that when feasible one can influence team mental models, and subsequently team perfor- mance by manipulating team composition in terms of team general mental ability. Fifth, whereas the measurement of mental models in the field may be an administrative challenge, there is no reason why the processes would be any different from measurements used in the laboratory as long as the task concepts can be explicitly defined (e.g., Smith-Jentsch, Campbell, Milanovich, & Reynolds, 2001).
734 RESEARCH REPORTS
Conclusion
We presented evidence that for a task with a defined set of optimal strategies, team mental model accuracy is a stronger predictor of team performance than team mental model similarity. However, unlike previous research that tends to favor similarity, this pattern of results did not emerge until later in training. In response to calls for the exploration of the determinants of team mental models, the present study also provides evidence that team members’ ability is related to the development of similar and accurate mental models and that the accuracy of mental models partially mediates the relationship between team ability and team performance.
References
Acton, W. H., Johnson, P. J., & Goldsmith, T. E. (1994). Structural knowledge assessment: Comparison of referent structures. Journal of Educational Psychology, 86, 303–311.
Arthur, W., Jr., Strong, M. H., Jordan, J. A., Williamson, J. E., Shebilske, W. L., & Regian, J. W. (1995). Visual attention: Individual differences in training and predicting complex task performance. Acta Psychologica, 88, 3–23.
Baron, R. M., & Kenny, D. A. (1986). The moderator–mediator variable distinction in social psychological research: Conceptual, strategic, and statistical considerations. Journal of Personality and Social Psychology, 51, 1173–1182.
Cannon-Bowers, J. A., & Salas, E. (2001). Reflections on shared cognition. Journal of Organizational Behavior, 22, 195–202.
Cannon-Bowers, J. A., Salas, E., & Converse, S. A. (1993). Shared mental models in expert team decision making. In N. J. Castellan Jr. (Ed.), Individual and group decision making: Current issues (pp. 221–246). Hillsdale, NJ: Erlbaum.
Carpenter, P. A., Just, M. A., & Shell, P. (1990). What one intelligence test measures: A theoretical account of the processing in the Raven Progres- sive Matrices test. Psychological Review, 97, 404 – 431.
Converse, S. A., Cannon-Bowers, J. A., & Salas, E. (1991). Shared mental models: A theory and some methodological issues. Proceedings of the Human Factors Society 35th Annual Meeting (pp. 1417–1421). Santa Monica, CA: Human Factors Society.
Cooke, N. J., Salas, E., Cannon-Bowers, J. A., & Stout, R. (2000). Mea- suring team knowledge. Human Factors, 42, 151–173.
Day, E. A., Arthur, W., Jr., & Gettman, D. (2001). Knowledge structures and the acquisition of a complex skill. Journal of Applied Psychology, 86, 1022–1033.
Donchin, E. (1989). The learning strategies project: Introductory remarks. Acta Psychologica, 71, 1–15.
Frederiksen, J. R., & White, B. Y. (1989). An approach to training based on principled task decomposition. Acta Psychologica, 71, 89 –146.
Goldsmith, T. E., & Davenport, D. M. (1990). Assessing structural simi- larity of graphs. In R. W. Schvaneveldt (Ed.), Pathfinder associative networks: Studies in knowledge organization (pp. 75– 87). Westport, CT: Ablex Publishing.
Goldsmith, T. E., Johnson, P. J., & Acton, W. H. (1991). Assessing structural knowledge. Journal of Educational Psychology, 83, 88 –96.
Gopher, D. (1993). The skill of attention control: Acquisition and execu- tion of attention strategies. In D. E. Meyer & S. Kornblum (Eds.), Attention and performance XIV: Synergies in experimental psychology, artificial intelligence, and cognitive neuroscience (pp. 299 –322). Cam- bridge, MA: MIT Press.
Gopher, D., Weil, M., & Bareket, T. (1994). The transfer of skill from a computer game trainer to actual flight. Human Factors, 36, 387– 405.
Gopher, D., Weil, M., & Siegel, D. (1989). Practice under changing priorities: An approach to the training of complex skills. Acta Psycho- logica, 71, 147–177.
Hart, S. G., & Battiste, V. (1992). Field test of a video game trainer. Proceedings of the Human Factors Society 36th Annual Meeting (pp. 1291–1295). Santa Monica, CA: Human Factors Society.
Johnson, P. J., Goldsmith, T. E., & Teague, K. W. (1994). Locus of the predictive advantage in Pathfinder-based representations of classroom knowledge. Journal of Educational Psychology, 86, 617– 626.
Johnson-Laird, P. N. (1983). Mental models: Towards a cognitive science of language, inference, and consciousness. Cambridge, MA: Harvard University Press.
Klimoski, R., & Mohammed, S. (1994). Team mental model: Construct or metaphor? Journal of Management, 20, 403– 437.
Kraiger, K., & Wenzel, L. H. (1997). Conceptual development and empir- ical evaluation of measures of shared mental models as indicators of team effectiveness. In M. T. Brannick, E. Salas, & C. Prince (Eds.), Team performance assessment and measurement: Theory, methods, and applications (pp. 63– 84). Mahwah, NJ: Erlbaum.
Levesque, L. L., Wilson, J. M., & Wholey, D. R. (2001). Cognitive divergence and shared mental models in software development project teams. Journal of Organizational Behavior, 22, 135–144.
Mané, A. M., & Donchin, E. (1989). The Space Fortress game. Acta Psychologica, 71, 17–22.
Marks, M. A., Mathieu, J. E., & Zaccaro, S. J. (2001). A temporally based framework and taxonomy of team processes. Academy of Management Review, 26, 356 –376.
Marks, M. A., Zaccaro, S. J., & Mathieu, J. E. (2000). Performance implications of leader briefings and team–interaction training for team adaptation to novel environments. Journal of Applied Psychology, 85, 971–986.
Mathieu, J. E., Heffner, T. S., Goodwin, G. F., Salas, E., & Cannon- Bowers, J. A. (2000). The influence of shared mental models on team process and performance. Journal of Applied Psychology, 85, 273–283.
Mohammed, S., & Dumville, B. C. (2001). Team mental models in a team knowledge framework: Expanding theory and measurement across dis- ciplinary boundaries. Journal of Organizational Behavior, 22, 89 –106.
Mohammed, S., Klimoski, R., & Rentsch, J. R. (2000). The measurement of team mental models: We have no shared schema. Organizational Research Methods, 3, 123–165.
Peterson, E., Mitchell, T. R., Thompson, L., & Burr, R. (2000). Collective efficacy and aspects of shared mental models as predictors of perfor- mance over time in work groups. Group Processes and Intergroup Relations, 3, 296 –316.
Raven, J., Raven, J. C., & Court, J. H. (1998). Manual for Raven’s Progressive Matrices and Vocabulary Scales. Oxford, England: Oxford Psychologists Press.
Ree, M. J., Carretta, T. R., & Teachout, M. S. (1995). Role of ability and prior job knowledge in complex training performance. Journal of Ap- plied Psychology, 80, 721–730.
Rentsch, J. R., & Hall, R. J. (1994). Members of great teams think alike: A model of team effectiveness and schema similarity among team members. In M. M. Beyerlein & D. A. Johnson (Eds.), Advances in interdisciplinary studies of work teams: Vol. 1 Theories of self-managing work teams (pp. 223–261). Stamford, CT: JAI Press.
Rentsch, J. R., & Klimoski, R. J. (2001). Why do “great minds” think alike?: Antecedents of team member schema agreement. Journal of Organizational Behavior, 22, 107–120.
Saccuzzo, D. P., & Johnson, N. E. (1995). Traditional psychometric tests and proportionate representation: An intervention and program evalua- tion study. Psychological Assessment, 7, 183–194.
Schmidt, F. L., & Hunter, J. E. (1992). Development of a causal model of processes determining job performance. Current Directions in Psycho- logical Science, 1, 89 –92.
735RESEARCH REPORTS
Schvaneveldt, R. W. (1990). Pathfinder associative networks: Studies in knowledge organization. Westport, CT: Ablex Publishing.
Schvaneveldt, R. W., Durso, F. T., & Dearholdt, D. W. (1989). Network structures in proximity data. In G. H. Bower (Ed.), The psychology of learning and motivation: Advances in research and theory (pp. 249 – 284). New York: Academic Press.
Smith-Jentsch, K. A., Campbell, G. E., Milanovich, D. M., & Reynolds, A. M. (2001). Measuring teamwork mental models to support training needs assessment, development, and evaluation: Two empirical studies. Journal of Organizational Behavior, 22, 179 –194.
Sobel, M. E. (1982). Asymptotic confidence intervals for indirect effects in structural equations models. In S. Leinhart (Ed.), Sociological method- ology 1982 (pp. 290 –312). San Francisco: Jossey-Bass.
Stout, R. J., Salas, E., & Kraiger, K. (1997). Role of trainee mental models in aviation team environments. International Journal of Aviation Psy- chology, 7, 235–250.
Webber, S. S., Chen, G., Payne, S. C., Marsh, S. M., & Zaccaro, S. J. (2000). Enhancing team mental model measurement with perfor- mance appraisal practices. Organizational Research Methods, 3, 307–322.
Appendix Instructions for Making Relatedness Ratings in Pathfinder
Your task on the computer is to make judgments about the “relatedness” of pairs of terms that have to do with playing the Space Fortress game. There are several ways one might think about the terms being judged. For instance, two terms might be related because they share common features or because they frequently occur together. For this task think about the terms as they relate to playing Space Fortress.
YOU SHOULD BASE YOUR RELATEDNESS JUDGMENTS ON HOW THE TERMS WORK TOGETHER TO HELP YOU PLAY THE GAME WELL.
The major goal of Space Fortress is to maximize your game points. This is accomplished by: (1) Destroying the Fortress as many times as you can, (2) Hitting as many mines as possible, and (3) Protecting your own ship from being hit or damaged. Thus, when making your relatedness judgments you should think about the Space Fortress in relation to these 3 sub-goals.
Each pair of terms will be presented on the screen along with a “relatedness” scale. You can think of the points along the scale as representing degrees of relatedness ranging from “1”, not at all related to “9”, highly related.
You are to indicate your judgement of relatedness for each pair of terms by pressing the number key that represents your rating. Upon responding, a bar marker will move directly above the number you pressed. If you wish to change your rating, simply press another number. When you are satisfied with the rating you have given, press the �SPACE BAR� to enter your response. Following this, the next pair of terms will be displayed.
A complete list of terms will be presented prior to beginning the task. This will give you a general idea of the scope of the Space Fortress terms you will be rating.
WHEN MAKING YOUR RATINGS, REMEMBER BACK TO PLAYING THE SPACE FORTRESS GAME AND THINK ABOUT HOW THE TERMS ARE RELATED IN ACCOMPLISHING THE GOAL OF MAXIMIZING YOUR SCORE.
Received April 30, 2004 Revision received February 24, 2005
Accepted March 10, 2005 �
736 RESEARCH REPORTS