wk4dis/talk
Business Failure Prediction using Decision Trees
ADRIAN GEPP,1 KULDEEP KUMAR1* AND SUKANTO BHATTACHARYA2 1 School of Business, Bond University, Gold Coast, Queensland, Australia 2 Deakin Business School, Deakin University, Burwood, Victoria, Australia
ABSTRACT Accurate business failure prediction models would be extremely valuable to many industry sectors, particularly fi nancial investment and lending. The potential value of such models is emphasised by the extremely costly failure of high-profi le companies in the recent past. Consequently, a signifi cant interest has been generated in business failure prediction within academia as well as in the fi nance industry. Statistical business failure prediction models attempt to predict the failure or success of a business. Discriminant and logit analyses have traditionally been the most popular approaches, but there are also a range of promising non-parametric techniques that can alternatively be applied. In this paper, the relatively new technique of decision trees is applied to business failure prediction. The numerical results suggest that decision trees could be superior predictors of business failure as compared to discriminant analysis. Copyright © 2009 John Wiley & Sons, Ltd.
key words business failure; bankruptcy; credit risk; insolvency; decision trees
INTRODUCTION
In light of the recent global fi nancial crisis, it has become imperative to question the long-term, real-world applicability of some of the hitherto popular parametric credit risk evaluation and fi nancial insolvency prediction models. While banks and lending institutions even in this time of worldwide economic gloom are not desperately trying to avoid creating credit, they are obviously desperate in their search for alternative, more transparent and functionally reliable tools to appropriately evaluate the risk of default.
The appropriate pricing of a complex fi nancial instrument like a complex derivative critically depends on the appropriate evaluation of the risk associated with the instrument. Methodological validity of pricing derivatives (including credit derivatives) using a Black–Scholes type closed-form
Journal of Forecasting J. Forecast. 29, 536–555 (2010) Published online 25 November 2009 in Wiley InterScience (www.interscience.wiley.com) DOI: 10.1002/for.1153
* Correspondence to: Kuldeep Kumar, Department of Economics and Statistics, Faculty of Business, Technology and Sustainable Development, Bond University, Gold Coast, Queensland 4229, Australia. E-mail: [email protected]
Copyright © 2009 John Wiley & Sons, Ltd.
Business Failure Prediction using Decision Trees 537
Copyright © 2009 John Wiley & Sons, Ltd. J. Forecast. 29, 536–555 (2010) DOI: 10.1002/for
model under the assumptions of continuous diffusion, log-normality and constant variance have been questioned by academicians as well as practitioners ever since the publication in 1973 of Black and Scholes’ landmark paper (Black and Scholes, 1973). However, rather than looking elsewhere, the general response by fi nancial mathematicians thus far to the drawbacks of the classical Black–Scholes model has been to venture deeper and deeper into the dense domain of stochastic calculus in search of a functionally reliable risk pricing tool.
While taking nothing away from the applicational utility of a closed-form risk evaluation model, one must always remember that the veracity of such a model wholly depends on the underlying parameter estimation process. Estimating the underlying parameters for a fi nancial risk evaluation model is always fraught with danger, given that much of fi nancial risk emanates from the market mood at any given point of time, which basically is a collection of mass human sentiments. All parametric risk evaluation models necessarily depend on the underlying assumption of rationality of human economic decision making in so far as they depend heavily on the presupposed omnipotence of the governing distributions and are independent of the ultimately human origins of risk.
In the specifi c context of pricing credit instruments, the key issue boils down to an appropriate evaluation of the primary default risk associated with the instrument. A popular parametric model constructs state transition probability matrices of credit ratings by rating agencies like Moody’s or Standard & Poor’s by applying Markov chain rule. The state transition probability matrices that are assumed to be stochastic and latent factor-driven may be predicted via a factor probit model (Feng et al., 2008). The basic presupposition is that the current credit rating of a fi rm by an agency completely determines the distribution of its future rating by the same agency (or one that is very similar). This is obviously a very strong presupposition and Eberlein et al. (2007) point out that a ‘tradeoff between tractability and realism is typical for the application of mathematical models in fi nance in general’. This, however, gives all the more reason to question the observed fascination of fi nancial mathematicians with parametric models.
In the backdrop of parametric credit risk pricing models being fraught with obvious methodological drawbacks and failing to deliver in this hour of global crisis, it is certainly worthwhile at this time to take stock of the other available options. One obvious area to explore is the plethora of non- parametric computational techniques that are available in this day and age of almost unbridled computing power at our fi ngertips. As non-parametric risk prediction models are distribution-free, they come closer to truly refl ecting the human element behind the surge of data generated by the fi nancial markets.
Non-parametric techniques have been successfully developed and applied in recent times in the pricing of fi nancial derivatives (Ait-Sahalia and Duarte, 2003) and also in an actuarial context (Serfl ing, 2000), so there is evidently much promise for successful applications of non-parametric techniques in credit risk estimation. Although not specifi cally dealing with credit default risk prediction, this paper applies a non-parametric technique to the closely related problem area of forecasting business failure.
The fi eld of business failure prediction has many aliases, such as bankruptcy prediction, fi rm failure prediction and fi nancial (di)stress prediction. Hereafter it will be referred to as business failure prediction (BFP). As the name suggests, BFP involves developing models that attempt to predict the fi nancial failure of a business before it actually happens. Statistical BFP models attempt to predict the failure or success of a business based on publicly available information about that business, such as the fi nancial ratios obtained from published fi nancial statements. In addition, some studies also include indicators of industry- and economy-wide performance to aid in the business failure predictions. Academic and industry interest in BFP models has increased as their value has been
538 A. Gepp, K. Kumar and S. Bhattacharya
Copyright © 2009 John Wiley & Sons, Ltd. J. Forecast. 29, 536–555 (2010) DOI: 10.1002/for
highlighted by extremely costly failures of high-profi le businesses in recent past (failures of HIH and OneTel are two examples in the South Pacifi c region).
Amongst other benefi ts, with accurate business failure prediction:
• commercial and investment banks, credit unions, and other fi nancial institutions would be able to better assess the primary default risk when creating credit; and
• investors would be able to better manage their investment portfolios by reallocating the funds locked in the stocks of companies that are about to fail.
In brief, accurate business failure prediction models would result in increased and stable economic growth for the benefi t of all involved. Discriminant and logit analyses have traditionally been the popular tools for BFP but suffer from the obvious drawbacks of parametric, distribution-dependent approaches—so search for a new approach is justifi ed.
The most important characteristics of a BFP model are its two forms of accuracy, namely classifi cation and prediction. A model’s classifi cation accuracy is obtained by assessing its accuracy on the dataset from which it was developed. Following that, the more important prediction accuracy of the model is assessed by its application to a brand new set of data, which refl ects how well the model will perform on future predictions. Nevertheless, when measuring either classifi cation or prediction accuracy there is a pragmatic important consideration that should be noted. It is a more critical error to classify a failing business as successful (Type I error) than to classify a successful business as failing (Type II error). The reason for this is that a Type II error only creates a lost opportunity cost from not dealing with a successful business: for example, missed potential investment gains. In contrast, a Type I error results in a realized fi nancial loss due to involvement with a business that will fail: for example, losing all money invested in an impending bankrupt business. Thus, misclassifi cation costs are not equal in the real world. Therefore, when analysing BFP models a higher-weighted penalty should be imposed for a misclassifi cation of a truly failing business (Type I error). A quantifi able difference in misclassifi cation costs has not been agreed upon in the literature, as it seems to vary for different circumstances and usually involves subjective decision making.
The remainder of this paper contains a brief review of BFP models and its relation to other comparable computational techniques that have appeared in the forecasting literature, followed by an in-depth analysis and review of decision trees for BFP. The data and methodology used for the numerical study are then explained, prior to the results and conclusions sections. Gepp’s (2005) thesis (which is available for download from http://www.it.bond.edu.au/publications/Theses.htm) contains more detail on some of the conceptual issues encountered in this paper; specifi cally, Chapter 2 contains a detailed review of BFP models and Chapter 3 details the research on BFP using DT techniques.
BRIEF REVIEW OF BUSINESS FAILURE PREDICTION MODELS
Many different techniques have been applied to BFP since its beginnings in the 1960s. The fi eld arguably started earlier, but the fi rst statistical and mathematical models for BFP were published in the 1960s. Beaver (1966) presented a univariate model, then Altman (1968) pioneered the use of multiple discriminant analysis (MDA), which was further developed by Deakin (1972), Edminster (1972) and others. Ohlson (1980) in his pioneer study, to avoid some signifi cant problems associated with MDA, employed conditional logit analysis (LA) for predicting the survival of businesses. LA does not require normality or equal covariances, which are prerequisites for MDA. Kumar and
Business Failure Prediction using Decision Trees 539
Copyright © 2009 John Wiley & Sons, Ltd. J. Forecast. 29, 536–555 (2010) DOI: 10.1002/for
Ganesalingam (2001) have since focused on predicting the fi nancial distress of a selection of major Australian companies. This research used principal component analysis, factor analysis, discriminant analysis and cluster analysis.
Theodossiou (1993) introduced sequential cumulative sum (CUSUM) procedures for BFP with excellent empirical results. The soft computing methods known as artifi cial neural networks (ANNs) have also been used in BFP with success (see Tan, 2001, for a summary). There are also vast numbers of other techniques that have been applied to BFP, including human information processing (HIP) and survival analysis.
DECISION TREES
Decision tree (DT) techniques generate a set of tree-based classifi cation rules used to construct a DT (also known as a classifi cation tree). DTs assign data to predefi ned classifi cation groups: in the case of BFP, a DT usually assigns each business to a successful or failing group. In general, DTs are binary trees, which consist of a root node, non-leaf nodes and leaf nodes connected by branches, whereby each non-leaf node has two branches leading to two distinct nodes, as shown in Figure 1.
When applied to classifi cation problems such as BFP, leaf nodes represent classifi cation groups (fail or success) and the non-leaf nodes each contain a splitting (or decision) rule. Thus the tree is built by a recursive process of splitting the data when moving from a higher to a lower level of the tree. The splitting rules comprise an expression (usually containing one fi nancial ratio) that is evaluated for each case (business) and compared to a cut-off value. An example splitting rule might be to classify a business into:
• the left sub-tree if current ratio <2.5; or • the right sub-tree if current ratio ≥2.5.
Figure 1. The basic structure of a binary tree
540 A. Gepp, K. Kumar and S. Bhattacharya
Copyright © 2009 John Wiley & Sons, Ltd. J. Forecast. 29, 536–555 (2010) DOI: 10.1002/for
Splitting rules are usually univariate as shown above, but the same variable can be used in zero, one or many splitting rules.
Similar to supervised learning with ANNs, DT building algorithms are used to manage the creation of DTs. The two main tasks a DT building algorithm performs are:
1. choosing the best splitting rule at each non-leaf node that discriminates between successful and failing businesses; and
2. managing the complexity (number of leaf nodes) of the DT. Most DT building algorithms fi rst create a very complex DT, then ‘prune’ the DT to the desired complexity, which involves replacing multiple node sub-trees with single leaf nodes. Pre-pruning is also possible at the initial DT creation phase, which would create a simpler tree at the expense of not exploring a more complex tree. Pre-pruning increases the effi ciency of the tree generation process, but with the risk of a possible loss in classifi cation and prediction accuracy.
Different DT building algorithms can be used to generate different decision trees. Although similar tree structures are created by all building algorithms, the choice of which algorithm to use often has a large infl uence on the accuracy of the fi nal DT. The important DT building algorithms from a BFP viewpoint are the recursive partitioning algorithm, and entropy algorithms such as classifi cation and regression trees (CART) and See5.
Theoretical analysis The major advantage of DTs is that they are non-parametric. Unlike the parametric discriminant and logit analyses, DTs have no distribution assumptions to violate and thus there is no need to consider transforming variables. The only assumptions that DTs possess are that the successful and failing groups are discrete, non-overlapping and distinctly identifi able, which are common assumptions in BFP.
DTs can handle missing values and qualitative data, as well as being easily represented in a user- friendly graphical format (Joos et al., 1998). Another advantage of DTs is that they can take different misclassifi cation costs for Type I and Type II error as inputs, which can then be incorporated into the DT building process at all stages. This is preferable to adjusting the cut-off values after model generation, as is the case for the more popular techniques of MDA and LA. However, a disadvantage of many DT building processes is that they require prior probabilities of successful and failed businesses as inputs. These prior probabilities are usually arbitrarily estimated, which adds an element of imprecision to the DT.
The interpretation of DTs is simple as splitting rules are usually univariate. This allows for easy identifi cation of signifi cant variables, where the root node contains the most signifi cant variable. The relative signifi cance of other variables can be found by comparing their proximity to the root node, whereby closer nodes contain more signifi cant variables.1 Thus DTs only identify the relative signifi cance of variables, unlike MDA and LA, which provide quantifi ed statistical fi gures that represent each variable’s signifi cance. DTs still have the predictive power of a multivariate approach as there are sequences of nodes (univariate rules) that lead to each classifi cation leaf node. Furthermore, these sequences of splitting rules (in nodes) can naturally model interactions between variables without including an interaction term in the model, as has to be done with MDA and LA. Although linear combinations of variables could also be used in splitting rules, the potentially increased
1 If a variable appears in more than one splitting rule, then its signifi cance is measured by the smallest distance to the root node.
Business Failure Prediction using Decision Trees 541
Copyright © 2009 John Wiley & Sons, Ltd. J. Forecast. 29, 536–555 (2010) DOI: 10.1002/for
predictive power is not thought to outweigh the reduction in simplicity, ease of interpretation and identifi cation of signifi cant variables.
The DT approach has a discrete scoring system whereby the probability of group membership (failure or success) is not produced and thus there is no way to distinguish businesses in the same classifi cation. These are the main disadvantages of DTs compared with the more popular MDA and LA models that have a continuous scoring system. However, the discrete scoring system removes the need for cut-off values and the arbitrary assignment of them, which has been a major criticism of both MDA and LA. Thus the discrete scoring system of DTs has both advantages and disadvantages.
DTs have been criticised for their building algorithms not reviewing previous rules when determining future rules possibly containing a repeat variable (Zopounidis and Dimitras, 1998). This is a theoretical weakness of DTs, but all forward stepwise approaches, such as those commonly used in MDA and LA, have a similar weakness of not reviewing previously included variables. Furthermore, there is no evidence to suggest that this weakness will signifi cantly reduce the classifi cation or predictive ability of DTs.
Relation of DTs to other comparable computational techniques in forecasting literature There have been several applications of DTs in conjunction with other comparable computational techniques, especially domain-specifi c expert systems, which have appeared to date in the forecasting literature—more in the fi eld of medical diagnostics than in any other fi eld. For example, Zelič et al. (1997) applied DTs in conjunction with Bayesian rule-based classifi cation in the diagnosis of sports injuries. McSherry (1999) proposed an algorithm for DT induction where attribute selection depended on the sequential evidence-gathering strategies used by doctors in clinical situations. Going beyond domain-specifi c expert systems, Kumar et al. (2009) very recently proposed a domain- independent decision support for the intensive care unit of a hospital using a hybrid of rule-based and case-based reasoning which incorporates Quinlan’s Iterative Dichotomiser 3 (ID3) DT algorithm for effective knowledge mining (Quinlan, 1986). The proliferation of hybrid systems in the research literature on decision/prediction systems for medical diagnosis is primarily attributable to the fact that medical cases are often scenario-specifi c and hence do not adequately yield to computational algorithms designed to extract general decision rules. As such, purely rule-based expert systems are easily outperformed by hybrid systems that combine rule-based and case-based reasoning. Interestingly, in such hybrid diagnostic decision support systems that combine rule-based and case- based reasoning, case libraries to retain the new and original cases are often constructed using DTs (Nilsson and Sollenborn, 2004).
Rule-based forecasting (RBF) systems have also found applications in business and fi nancial decision-making, though their prevalence is much less as compared to the time series forecasting methods. Collopy and Armstrong (1992) examined the feasibility of RBF by combining forecasting expertise and domain knowledge to produce forecasts according to time series features of the data. After reviewing various methods to determine a logical approach to analyse the results of sequential decision making in developing an expert system for RBF in business, Chrysler (2005) observed that DTs can be a very effi cient methodological route in this direction and proposed such a DT-based method that a knowledge engineer can actually use to develop a rule-based expert system.
DTs have also been used in conjunction with artifi cial neural networks (ANNs) to extract decision rules from feed-forward neural networks with continuous outputs without the need for any assumptions about the network’s internal structure or specifi c data features (Schmitz et al., 1999). Univariate DTs have been used in conjunction with a backbone ANN predictor model to improve its expressiveness
542 A. Gepp, K. Kumar and S. Bhattacharya
Copyright © 2009 John Wiley & Sons, Ltd. J. Forecast. 29, 536–555 (2010) DOI: 10.1002/for
and alleviate the black-box nature of ANNs (Abbass et al., 1999). In the context of a business- relevant forecasting application, a DT model was recently seen to easily outperform both ANN and support vector machine (SVM) models in predicting auditor choice for a sample of companies (Kirkos et al., 2008).
The predictive power of DTs can be further ‘boosted’ by applying the predictive function iteratively in a series and recombining the output with nodal weighting in order to minimise the forecast error. Commercial application softwares such as DTREG (http://www.dtreg.com/) are available that can perform boosting of DTs. Although not yet tested with business or fi nancial data vis-à-vis other computational methods like ANNs or SVMs in terms of prediction performance, boosted DTs have actually already been tested and found to be a considerably more effi cient alternative to back- propagation-type ANNs for elementary particle identifi cation in nuclear physics (Roe et al., 2005).
Apart from expert systems and ANNs, there is a substantial literature on the successful application of DTs in relation to other computational methods as well. For example, Bala et al. (1995) combined genetic algorithms and the classical ID3 algorithm of DTs to evolve optimal subsets of distinctly identifi able features for robust pattern detection. Braaten (1996) used a cluster analysis-derived tree to cross-validate predictions obtained using the ID3 algorithm of DTs, thus highlighting their mutual comparability.
Review of DTs in BFP Following the successful application of DTs (using the recursive partitioning algorithm) to medical decision making by Goldman et al. (1982, 1988), DTs were fi rst applied to BFP in the seminal work of Frydman et al. (1985). Since then, the successful applications of DT algorithms by subsequent researchers as mentioned in the previous subsection have inspired further applications of DTs to BFP with varying success.
Frydman et al.’s RPA Frydman et al. (1985) were the fi rst to apply a non-parametric technique to BFP. They used the recursive partitioning algorithm to build DTs, a method that has become known as recursive partitioning analysis (RPA). RPA’s inputs are the prior probabilities of failing and successful businesses, misclassifi cation costs, and a set of training data. RPA constructs each DT in a way to minimise the expected misclassifi cation cost (Jones, 1987), referred to as resubstitution risk. Resubstitution risk is calculated using probability theory and Bayesian reasoning (refer to Frydman et al., 1985, for more details).
To assess the suitability of DTs (specifi cally RPA) to BFP, Frydman et al. compared two RPA- built DTs with two MDA models. The difference between the two DT models and the two MDA models was the model size, in relation to the number of explanatory variables incorporated. The larger RPA-DT was created using the approach outlined in the paragraph above, while the smaller tree was chosen as the sub-tree with the best cross-validation performance. V-fold cross-validation models are calculated by dividing the data into V approximately equal groups and generating a DT for each (V-1) group (Frydman et al. used V = 5). Each DT was then used to classify the group left out, and the best DT was chosen as the tree with the lowest average resubstitution risk. Similarly, the smaller and larger MDA models were constructed using the four and 10 most signifi cant fi nancial variables, respectively, according to a forward stepwise method.
The prior probabilities were set at 2% for failing and 98% for successful businesses, and the techniques were compared over eight different misclassifi cation costs. The dataset used is the same as used in this paper, which is discussed in the next section. The RPA-DT models were found to be
Business Failure Prediction using Decision Trees 543
Copyright © 2009 John Wiley & Sons, Ltd. J. Forecast. 29, 536–555 (2010) DOI: 10.1002/for
superior classifi ers of business failure compared with the MDA models. In addition, as may be expected, the more parsimonious (smaller) models outperformed their more complex counterparts on the cross-validation analysis. More importantly, the more parsimonious RPA-DT model was found to be a superior predictor to the MDA models for almost all misclassifi cation costs on the cross-validation analysis. However, the more complex RPA-DT was found to be the worst predictor, which highlights the potential risk of over-training. Overall, after considering their theoretical analysis and empirical results, Frydman et al. concluded that decision trees using RPA were useful tools for predicting business failure.
Other relevant studies Very few papers have focused on DTs since Frydman et al. (1985), despite the existence of many different DT-building algorithms, including ID3, See4.5, See5 and CART.2 All of these DT-building algorithms are entropy algorithms, which select splitting rules that minimise entropy (or noise), rather than minimising expected misclassifi cation costs as occurs with RPA. Entropy-building algorithms manage complexity by choosing a DT that only retains splitting rules based on general patterns in the data. There is also another related technique that has been applied to BFP, known as multivariate adaptive regression splines (MARS). MARS is an extension of the CART technique that incorporates the DT concept of splitting the data into subregions using stepwise regression. Thus MARS can be thought of as a procedure that searches for the optimal piecewise regression model (which may include higher order terms).
Joos et al. (1998) used LA and See5 to predict credit classifi cations for one of Belgium’s largest banks. Nine models were created: three models from each technique based on three different datasets comprising a full set of fi nancial variables, a reduced smaller set of fi nancial variables and a set of qualitative variables. As expected, Joos et al. concluded that, regardless of the technique used, the classifi cation ability of these models was ranked in the order mentioned in the previous sentence. Additionally, LA was found to outperform See5 on the full dataset, but See5 outperformed LA on both the reduced and qualitative datasets. When different misclassifi cation costs were introduced See5 was superior for fi ve (out of eight) of these trials, but LA was better when the Type I error was the highest. This study showed no obvious overall superiority between See5 and LA.
Huarng et al. (2005) also used the See5 package, as well as being the fi rst to apply CART to BFP. CART was found to be empirically superior to See5. However, the datasets comprised fewer than 12 businesses and fi ve variables, which is too small to obtain reliable results. It should also be noted that Shirata (1998) used CART prior to 2005, but only for selecting variables to use in their MDA model of Japanese bankruptcies. CART was used in this case as it calculates a variable importance score for each variable.
Despite the lack of papers focused on DTs, various studies have used them as a comparison technique. Tam (1991) and Tam and Kiang (1992) found that ANNs, MDA and LA all outperformed ID3 decision trees. Fernández and Olmeda (1995) also found ANNs to be slightly superior to See4.5 and MARS. Additionally, despite the lower classifi cation accuracy of See4.5 and MARS (except for MARS compared with MDA), they were found to have superior predictive ability over both LA and MDA on the hold-out dataset. Fernández and Olmeda (1995) also empirically tested combining some of these techniques, with promising results. In contrast to these studies, Martinelli et al. (1999) found that See4.5 outperformed ANNs. Furthermore, Laitinen and Kankaanpää (1999) showed that RPA
2 See4.5, an earlier version of See5, was presented as an improved version of ID3. See5 is considered to be superior to both See4.5 and ID3, but an overall ‘better than’ comparison cannot be made with CART.
544 A. Gepp, K. Kumar and S. Bhattacharya
Copyright © 2009 John Wiley & Sons, Ltd. J. Forecast. 29, 536–555 (2010) DOI: 10.1002/for
produced similar results to MDA and LA. However, none of these studies have considered varying misclassifi cation costs.
DATA AND METHODOLOGY
This research is designed to further the assessment of DTs for use in BFP. The comparison of RPA- DTs and MDA by Frydman et al. (1985) will be extended to include See5 and CART decision trees, such that models of both high and low complexity will be compared over various misclassifi cation costs.
Data The dataset used for this research is an unaltered copy of the data used by Frydman et al. (1985)3 to ensure the validity of comparisons with Frydman et al.’s results. The main properties of this cross-sectional dataset are presented in Table I. In this dataset, Frydman et al. considered a business to be failed if it had legally fi led for bankruptcy according to the Wall Street Journal Index.
Methodology As done in Frydman et al. (1985), each technique’s classifi cation accuracy was fi rst tested on the original training dataset. The prediction accuracy of each model was then estimated by using the V- fold cross-validation procedure where V = 5. The V-fold cross-validation procedure was used in preference to the hold-out sample procedure as it more accurately estimates prediction accuracy when the sample size is small (Frydman et al., 1985): the sample size used in this research is 200 businesses, which is considered small.
A simple and a complex model of both techniques have been generated to keep in line with Frydman et al. (1985). These models were then tested over the same range of Type I to Type II error ratios as Frydman et al. (1985). CART has the same input formats as RPA, so the exact inputs used by Frydman et al. (1985) were used to produce the CART models: 2% failed and 98% successful prior probabilities for misclassifi cation costs of 1, 10, 20 . . . 60, 70. However, See5 does not take prior probabilities as inputs or assess models based on resubstitution risk (discussed in depth in Frydman et al., 1985), so the misclassifi cation cost inputs needed to be adjusted by a factor of approximately 0.05 calculated as
3 The dataset is available from Frydman et al. upon request.
Table I. Description of data used
Property Value
Businesses (type) US manufacturing and retail (COMPUSTAT database) Selection procedure Random selection Year of study Randomly chosen from 1971–1981 (fi nancial years) Businesses (number) 200: 142 successful and 58 failed Explanatory variables 20 fi nancial variables: mostly fi nancial ratios. A list of the
variables is available in the appendix of Frydman et al.’s paper.
Business Failure Prediction using Decision Trees 545
Copyright © 2009 John Wiley & Sons, Ltd. J. Forecast. 29, 536–555 (2010) DOI: 10.1002/for
priorProb NoFirms PriorProbabilityfailure failed successful×( ) ××( ) = ×( ) ×( ) ≈
NoFirmssuccessful 0 02 58 0 98 142 0 05. . .
This formula is mathematically derived from the ratio of misclassifi cations (misclassifi ed as failed/ misclassifi ed as successful) in the resubstitution risk formula from Frydman et al. (1985). This results in the values shown in Table II. Note that these fi gures are being called error cost ratios, denoted ECRs hereafter.
CART models Version 5.0 of the CART software package, produced by Salford Systems in San Diego, was used. This software claims to be the only DT software based on the original (proprietary) code of the creators of classifi cation and regression trees (CART) (Breiman et al., 1984). The ‘entropy’ method for selecting splitting variables was used as it had been successfully used in diverse applications. The two models generated were:
• the simple CART model (CART_s), chosen as the model with the smallest resubstitution risk from the fi vefold cross-validation procedure; and
• the complex CART model (CART_c), chosen as CART_s prior to its last level of pruning as defi ned by the CART software.
See5 models See5 Release 2.02 was used to produce all See5 models.4 The ‘minimum cases per leaf node’ option was set to 2 to prevent pre-pruning. The ‘pruning CF’ setting controls tree complexity whereby larger values result in less pruning. ‘Pruning CF’ is expressed as a percentage similar to a signifi cance level in hypothesis testing, such that each sub-tree is not pruned if it is signifi cant at the ‘pruning CF’ level. The default of 25% produces relatively large trees, so:
• the simple See5 model (See5_s) was generated using 1% ‘pruning CF’; and • the complex See5 Model (See5_c) was generated using 10% ‘pruning CF’.
RESULTS
The results are presented as follows. The infl uence of model complexity will fi rst be analysed. The four techniques will then be compared for both classifi cation and then prediction accuracy. Finally, the relative signifi cance and importance of the explanatory variables will be analysed. These results are presented as an extension to the comparison of RPA decision trees and MDA undertaken by Frydman et al. (1985); therefore, the resubstitution risk has been used to compare techniques, such
Table II. Error cost ratio (ECR) values
Frydman et al. (1985) measure 1 10 20 30 40 50 60 70
ECR 0.05 0.5 1 1.5 2 2.5 3 3.5
4 Quinlan (1993) discusses the underlying algorithm in See4.5, which is still present in See5.
546 A. Gepp, K. Kumar and S. Bhattacharya
Copyright © 2009 John Wiley & Sons, Ltd. J. Forecast. 29, 536–555 (2010) DOI: 10.1002/for
that the lower the resubstitution risk the better the model. For a more detailed comparison of RPA and MDA, refer to Frydman et al. (1985). Tables of the resubstitution risk and error breakdown data for the four techniques are presented in Appendix A and Appendix B, respectively.
Effect of model complexity The in-sample classifi cation accuracy of the four complex models is superior to that of their simple model counterparts, except for a few situations with the See5 package where the models have equal classifi cation accuracy. These results are expected, as the simple models are fully contained within their complex counterparts. The cross-validation prediction accuracy of complex models is guaranteed to be better than the simple models. Moreover, the results in Table III show that in most cases the simple models produce more accurate predictions. In the case of MDA and See5 the results are not conclusive, but the simpler models were superior predictors in the majority of cases. Moreover, generating simpler models has been conclusively shown to increase predictive power when using the RPA and CART approaches. Thus, overall, the simpler models had better predictive power, even though they are outperformed by more complex models when classifying the original dataset. This empirical fi nding is consistent with the principle of parsimony, whereby most statisticians prefer less complex models given the choice.
In-sample classifi cations There was a large variation between the in-sample classifi cation ability of RPA, MDA, See5 and CART. Overall, the simple MDA model was the worst, while See5 was the best-performing simple model except for the two lowest ECRs. As the complex models are superior classifi ers they are the focus of this analysis, and their classifi cation accuracy is presented in Figure 2. The See5 and MDA models still remain the best and worst classifi ers, respectively. RPA and CART also have superior classifi cation ability over MDA for more than 50% of the ECRs, including CART being the choice model for low ECR values. For greater ECRs, the ranking order is consistently See5, RPA, CART and then MDA. Overall, it is clear that all the DT techniques had superior classifi cation ability compared with MDA, and See5 was the best of these DT techniques.
Cross-validated predictions Compared with the large variability in classifi cation ability, the four techniques had more comparable prediction abilities. The fi rst observation noted from both the complex and simple models is that
Table III. The most accurate predictive model (complex, simple or equal) for each technique and misclassifi cation cost, determined by the lowest resubstitution risk
ECR RPA MDA See5 CART
0.05 Simple Simple Equal Simple 0.5 Equal Simple Simple Simple 1 Simple Simple Complex Simple 1.5 Simple Complex Complex Simple 2 Simple Complex Simple Simple 2.5 Simple Complex Simple Simple 3 Simple Simple Simple Simple 3.5 Simple Simple Complex Simple
Business Failure Prediction using Decision Trees 547
Copyright © 2009 John Wiley & Sons, Ltd. J. Forecast. 29, 536–555 (2010) DOI: 10.1002/for
there appears to be zero, or negative, correlation between the best models for classifi cation and the best models for prediction. That is, MDA_c is on balance the best predictor amongst the complex models, but it was the worst classifi er. CART_c is the best of the DT complex models at prediction. Further analysis of the prediction accuracy of the complex models will not be included here as the main purpose is to analyse the best predictive model and, as mentioned above, the simple models are better predictors.
The cross-validated prediction accuracy of the simple models is presented in Figure 3. This graph reveals that the four simple models have similar prediction ability. Owing to this similarity, each model’s results will be analysed separately.
The reduction in complexity has not increased the predictive power of the MDA model as signifi cantly as the DT models. Consequently, the MDA model has the largest for the majority of ECRs, with a notable exception being for the highest ECR value. In fact, the MDA model’s resubstitution risk actually decreases when the ECR is greater than 2.5, which causes it to have the lowest resubstitution risk and be the best predictor for the highest ratio of Type I to Type II error.
The See5 model has the highest resubstitution risk of all DT models, with the exception of ECR values of 2.5 and 3. This suggests that See5 decision trees might be better for large Type I error costs, but the counter-example for this is for an ECR of 3.5 where See5 was the worst predictor by a signifi cant margin. See5 is arguably a better overall predictor than MDA, but it is the worst predictor amongst the DT techniques.
The CART model has the most consistent predictive ability over varying misclassifi cation costs: that is, in Figure 3 CART has the straightest line without any unexpected rises or drops. CART was superior to MDA and See5 for ECR values up to 2.5 (except for See5’s slightly smaller resubstitution risk when the ECR is 2.5). However, CART did not adjust well to the highest two ECRs, as shown
Figure 2. Classifi cation accuracy of the four complex models
548 A. Gepp, K. Kumar and S. Bhattacharya
Copyright © 2009 John Wiley & Sons, Ltd. J. Forecast. 29, 536–555 (2010) DOI: 10.1002/for
by its poor performance for ECRs of 3 and 3.5. Overall, CART had superior predictive power compared with MDA and See5.
The reduction in complexity from the complex to the simple model had the largest corresponding increase in predictive power for the RPA-DTs. Consequently, the RPA_s model had the best overall performance in terms of predictive ability. However, CART was almost identical for ECR values of 2.5 and lower.
RPA and CART models had similar predictive power and superiority over the similar predictive power of See5 and MDA. The qualifi er to this is when ECR values are extremely high; then MDA becomes superior to RPA, CART and See5 (in order). The technique with the best overall predictive ability is RPA-built DTs. In general, DTs have slightly better predictive power than MDA, but caution must be taken to limit the complexity of DT models.
Model complexity revisited It is important to discuss the remarkable difference between the classifi cation and prediction ability of the See5 technique. See5 was clearly the best classifi cation technique, but its predictive ability was poor in comparison. In addition, while large ECR values do not affect the classifi cation ability of See5, they signifi cantly and adversely affect its prediction ability. A partial reason for this is that, as Figure 4 shows, See5 normally generated more complex trees compared with CART.5 The performance gap with See5 can then be described with the earlier fi nding that increases improve classifi cation, but worsen prediction, accuracy. Additionally, as ECR increases CART models reduce
Figure 3. Prediction accuracy of the four simple models
5 RPA-DT complexity data were not published by Frydman et al. (1985), but from partial information published it seems that RPA-DTs were more comparable to the complexity of CART-DTs.
Business Failure Prediction using Decision Trees 549
Copyright © 2009 John Wiley & Sons, Ltd. J. Forecast. 29, 536–555 (2010) DOI: 10.1002/for
in complexity, while See5 models increase in complexity. This difference between CART and See5 may hold the key to their performance difference, but more research would have to be done to make a more defi nitive proposition.
Importance of analysis of variables The variables important in CART and See5 models are in most cases different from both MDA and RPA. Unlike RPA, CART generally used almost identical decision trees, despite varying mis- classifi cation costs. A typical CART model is shown in Figure 5 and depicts the common most important variables in the CART models:
• the activity ratio CA/SLS; • the business size variable Ln(TA) (also found to be important by MDA); and • the coverage ratio TL/TA and liquidity ratio WC/SLS.
Contrary to the general case that DT techniques do not provide a numeric measure of variable importance, CART provides a unique statistical ranking of variable importance based on how accurate the model would have been produced if it did include each variable. The details of this are not within the scope of this paper, but it is important to note that this added feature of CART is considered the best way for determining variable importance by some researchers. Applied in this case, this procedure still identifi ed CA/SLS as the most important variable, but surprisingly WC/SLS was listed as second, which highlights the uniqueness of this procedure.
CA/SLS was also at the root of every See5 DT, except for an ECR value of 0.5 where the WC/TS was at the root. In contrast to CART, however, the remainder of the See5 DTs vary considerably over different misclassifi cation costs. It was also common for See5 to use variables in more
Figure 4. Comparison of the complexity of CART and See5 decision trees
550 A. Gepp, K. Kumar and S. Bhattacharya
Copyright © 2009 John Wiley & Sons, Ltd. J. Forecast. 29, 536–555 (2010) DOI: 10.1002/for
than one splitting rule. Overall, the majority of the important variables were profi tability and liquidity ratios.
See5 and CART are the only techniques with the most important variable (CA/SLS) in common, while the most important variables from MDA and RPA (CA/CL and CF/TL, respectively) were not found to be very important in See5 or CART. Overall, the only common conclusion that can be made is that MDA, RPA, CART and See5 models tend to include a range of variables types rather than just variables from one classifi cation group.
CONCLUSION
By extending the work of Frydman et al. (1985) to include See5 and CART we have provided further empirical evidence to support the claim that less complex, more parsimonious models are better predictors than more complex models. More importantly, this research has provided further evidence to suggest that DT techniques are superior classifi ers and predictors of business failure. See5 had the best in-sample classifi cation ability, although it had the worst predictive power amongst the DT techniques. The CART and RPA decision tree techniques produced very similar results and were the best overall predictors. In particular, RPA was preferred as it had slightly better predictive ability with high Type I error (relative to Type II error) costs. However, we would ideally want to conclude by providing some insight to the readers as to certain elements of model choice in a forecasting context so that they are able to better generalise methodological ideas rather than opt for a method based simply on empirical results.
There is a popular belief that DTs perform well with discrete/categorical datasets but not so well with continuous data. While this remains to be exhaustively tested, Koh (2004) showed that DTs
Figure 5. CART_s model for ECR of 0.5. This graphical tree representation was automatically created by the CART software package, a feature not available with See5
Business Failure Prediction using Decision Trees 551
Copyright © 2009 John Wiley & Sons, Ltd. J. Forecast. 29, 536–555 (2010) DOI: 10.1002/for
outperformed ANNs and LA in predicting a fi rm’s ‘going concern’ qualifi cations using an input vector that contained six fi nancial ratios and hence was continuous data. We too have used a set of fi nancial ratios to perform the classifi cation between failing and non-failing businesses and our results show that all of the different DT techniques we have applied have comfortably outperformed MDA. However, we are of the opinion that much more exhaustive performance testing is required before a conclusive statement can be made about the capability of DTs to handle continuous data. So if the data is continuous, we would have to advise the reader to use his/her own discretion in whether to go by the empirical results that we have obtained and select DTs as the methodology of choice, or rather go by the popular belief and select ANNs or SVMs, which are said to perform better than DTs with continuous data.
It has also been argued that ‘top-down’ binary DTs can ‘force’ orthogonal partitions on to data where a non-orthogonal partition is better suited and that a ‘bottom-up’ method like hierarchical cluster analysis can better discern the true shape as compared to DTs (Wishart, 1998). This might be true in the case of certain shape-dependent, knowledge discovery problems in the presence of correlated variables, and we advise the reader to try out alternative methodologies if his/her primary research problem is such.
To conclude, we would like to emphasize that the best forecasting technique for any prediction problem depends on the specifi cation of the misclassifi cation cost, which in turn depends on the problem context and will unavoidably involve some degree of subjective judgement irrespective of the sophistication of the ultimately chosen method.
APPENDIX A: RESUBSTITUTION RISKS
The table below is based on, and extends, Table IVa in Fydman et al. (1985). It shows the resubstitution risk of each model for all error cost ratios (ECRs), which were calculated using the number of Type I and Type II errors produced by each model (shown in Appendix B).
ECR RPA_c RPA_s MDA_c MDA_s See5_c See5_s CART_c CART_s
In-sample 0.05 0.006 0.020 0.017 0.026 0.020 0.020 0.008 0.020 0.5 0.069 0.086 0.121 0.155 0.059 0.059 0.017 0.038 1 0.076 0.110 0.193 0.207 0.097 0.097 0.076 0.097 1.5 0.138 0.224 0.231 0.235 0.062 0.128 0.134 0.159 2 0.159 0.255 0.255 0.297 0.069 0.103 0.193 0.255 2.5 0.190 0.286 0.272 0.324 0.066 0.114 0.228 0.286 3 0.186 0.248 0.290 0.324 0.062 0.117 0.255 0.317 3.5 0.173 0.269 0.321 0.352 0.062 0.128 0.269 0.348
Cross-validated 0.05 0.091 0.020 0.056 0.035 0.020 0.020 0.068 0.020 0.5 0.124 0.124 0.179 0.155 0.152 0.141 0.186 0.131 1 0.235 0.221 0.296 0.249 0.221 0.228 0.241 0.214 1.5 0.376 0.269 0.272 0.307 0.321 0.335 0.307 0.276 2 0.442 0.324 0.329 0.365 0.414 0.359 0.352 0.331 2.5 0.497 0.386 0.386 0.402 0.441 0.373 0.407 0.379 3 0.497 0.386 0.427 0.397 0.497 0.400 0.469 0.428 3.5 0.531 0.455 0.518 0.369 0.486 0.528 0.517 0.476
552 A. Gepp, K. Kumar and S. Bhattacharya
Copyright © 2009 John Wiley & Sons, Ltd. J. Forecast. 29, 536–555 (2010) DOI: 10.1002/for
A P
P E
N D
IX B
: T
Y P
E B
R E
A K
D O
W N
O F
E R
R O
R S
T he
t ab
le b
el ow
s ho
w s
th e
er ro
rs p
ro du
ce d
by e
ac h
m od
el f
or a
ll e
rr or
c os
t ra
ti os
( E
C R
s) . T
he a
bb re
vi at
io ns
I , I
I an
d T
i nd
ic at
e T
yp e
I er
ro r,
T yp
e II
e rr
or ,
an d
to ta
l er
ro r,
r es
pe ct
iv el
y. A
ll R
P A
a nd
M D
A d
at a
ha ve
b ee
n ob
ta in
ed f
ro m
F ry
dm an
e t
a l.
( 19
85 ),
w hi
ch
di d
no t
re ve
al t
he n
um be
r of
e rr
or s
pr od
uc ed
b y
th e
R P
A a
nd M
D A
m od
el s
in t
he c
ro ss
-v al
id at
io n
tr ia
ls .
E C
R :
0. 05
0. 5
1 1.
5 2
2. 5
3 3.
5
E rr
or t
yp e:
I II
T I
II T
I II
T I
II T
I II
T I
II T
I II
T I
II T
In -s
a m
p le
R P
A _c
18 0
18 14
3 17
9 2
11 6
11 17
6 11
17 5
15 20
5 12
17 2
18 20
R P
A _s
58 0
58 19
3 22
14 2
16 9
19 28
9 19
28 9
19 28
9 9
18 4
25 29
M D
A _c
48 0
48 25
5 30
17 11
28 13
14 27
11 15
26 9
17 26
8 18
26 7
22 29
M D
A _s
54 1
55 33
6 39
18 12
30 12
16 28
11 21
32 8
27 35
6 29
35 6
30 36
S ee
5_ c
58 0
58 17
0 17
11 3
14 6
0 6
5 0
5 3
2 5
0 9
9 0
9 9
S ee
5_ s
58 0
58 17
0 17
11 3
14 9
5 14
5 5
10 3
9 12
3 8
11 3
8 11
C A
R T
_c 23
0 23
5 0
5 11
0 11
11 3
14 10
8 18
10 8
18 5
22 27
4 25
29 C
A R
T _s
58 0
58 11
0 11
11 3
14 10
8 18
9 19
28 9
19 28
9 19
28 9
19 28
C ro
ss -v
a li
d a te
d S
ee 5_
c 58
0 58
22 11
33 18
14 32
19 18
37 18
24 42
16 24
40 16
24 40
13 25
38 S
ee 5_
s 58
0 58
27 7
34 17
16 33
23 14
37 15
22 37
10 29
39 12
22 34
13 31
44 C
A R
T _c
38 8
46 16
19 35
22 13
35 17
19 36
15 21
36 14
24 38
14 26
40 14
26 40
C A
R T
_s 58
0 58
16 11
27 21
10 31
16 16
32 14
20 34
14 20
34 14
20 34
14 20
34
Business Failure Prediction using Decision Trees 553
Copyright © 2009 John Wiley & Sons, Ltd. J. Forecast. 29, 536–555 (2010) DOI: 10.1002/for
ACKNOWLEDGEMENT
The authors are grateful to the referee for comments on the earlier version of this paper.
REFERENCES
Abbass HA, Towsey M, Finn G. 1999. C-Net: generating multivariate decision trees from artifi cial neural networks using C5. University of Queensland technical report: FIT-TR-99-04. http://portal.acm.org/citation. cfm?id=869850 [29 April 2009].
Ait-Sahalia Y, Duarte J. 2003. Nonparametric option pricing under shape restrictions. Journal of Econometrics 116: 9–47.
Altman EI. 1968. Financial ratios, discriminant analysis and the prediction of corporate bankruptcy. Journal of Finance 23(4): 589–609.
Bala J, Huang J, Vafaie H, DeJong K, Wechsler H. 1995. Hybrid learning using genetic algorithms and decision trees for pattern classifi cation. In Proceedings of the International Joint Conference on Artifi cial Intelligence (IJCAI), Montréal, Canada, 19–25 August.
Beaver WH. 1966. Financial ratios as predictors of failure. Journal of Accounting Research (Supplement) 4(3): 71–111.
Black F, Scholes M. 1973. The pricing of options and corporate liabilities. Journal of Political Economy 81(3): 637–654.
Braaten Ø. 1996. Artifi cial intelligence in pediatrics: important clinical signs in newborn syndromes. Computers and Biomedical Research 29: 153–161.
Breiman L, Friedman JH, Olshen R, Stone CJ. 1984. Classifi cation and Regression Trees. Wadsworth & Brooks: Monterey, CA.
Chrysler E. 2005. Using decision tree analysis to develop an expert system. Proceedings of ISECON 22nd Annual Information Systems Education Conference, Ohio, 6–9 October.
Collopy F, Armstrong JS. 1992. Rule-based forecasting: development and validation of an expert systems approach to combining time series extrapolations. Management Science 38(10): 1394–1414.
Deakin E. 1972. A discriminant analysis of predictors of business failure. Journal of Accounting Research Spring: 167–179.
Eberlein E, Frey R, Kalkbrener M, Overbeck L. 2007. Mathematics in Financial risk management. http://www. defaultrisk.com/pp_other147.htm [29 April 2009].
Edminster R. 1972. An empirical test of fi nancial ratio analysis for small business failure prediction. Journal of Financial and Quantitative Analysis 2(7): 1477–1493.
Feng D, Gourieroux C, Jasiak J. 2008. The ordered qualitative model for credit rating transitions. Journal of Empirical Finance 15: 111–130.
Fernández E, Olmeda I. 1995. Bankruptcy prediction with artifi cial neural networks. In Natural to Artifi cial Neural Computation: Proceedings of the International Workshop on Artifi cial Neural Networks (IWANN ’95). Lecture Notes in Computer Science. Springer: Berlin; 1142–1146.
Frydman H, Altman EI, Kao DL. 1985. Introduction recursive partitioning for fi nancial classifi cation: the case of fi nancial distress. Journal of Finance 40(1): 269–291.
Gepp A. 2005. An evaluation of decision tree and survival analysis techniques for business failure prediction. Masters thesis, Bond University, Australia.
Goldman L, Cook EF, Brand DA, Lee TH, Rouan GW, Weisberg MC, Acampora D, Stasiulewicz C, Walshon J, Terranova G, Gottlieb L, Kobernick M, Goldstein-Wayne B, Copen D, Daley K, Brandt AA, Jones D, Mellors J, Jakubowski R. 1982. A computer protocol to predict myocardial information in emergency department patients with chest pain. New England Journal of Medicine 307: 588–597.
Goldman L, Weinberg M, Olshen R, Cook EF, Sargent RK, Lamas GA, Dennis C, Wilson C, Deckelbaum L, Fineberg H, Stiratelli R. 1988. A computer-derived protocol to aid in the diagnosis of emergency room patients with acute chest pain. New England Journal of Medicine 318: 797–803.
Huarng K, Yu HK, Chen CJ. 2005. The application of decision trees to forecast fi nancial distressed companies. In Proceedings of International Conference on Intelligent Technologies and Applied Statistics, Taipei, Taiwan.
554 A. Gepp, K. Kumar and S. Bhattacharya
Copyright © 2009 John Wiley & Sons, Ltd. J. Forecast. 29, 536–555 (2010) DOI: 10.1002/for
Jones FL. 1987. Current techniques in bankruptcy prediction. Journal of Accounting Literature 6: 131–164. Joos P, Vanhoof K, Ooghe H, Sierens N. 1998. Credit classifi cation: a comparison of logit models and decision
trees. In Proceedings Notes of the Workshop on Application of Machine Learning and Data Mining in Finance: 10th European Conference on Machine Learning, Chemnitz, Germany; 59–72.
Kirkos E, Spathis C, Manolopoulos Y. 2008. Support vector machines, decision trees and neural networks for auditor selection. Journal of Computational Methods in Sciences and Engineering 8(3): 213–224.
Koh HC. 2004. Going concern predictions using data mining techniques. Managerial Auditing Journal 19: 462–476.
Kumar K, Ganesalingam S. 2001. Detection of fi nancial distress via multivariate statistical analysis. Detection and Prediction of Financial Distress 27(4): 45–55.
Kumar KA, Singh Y, Sanyal S. 2009. Hybrid approach using case-based reasoning and rule-based reasoning for domain independent clinical decision support in ICU. Expert Systems with Applications 36(1): 65–71.
Laitinen T, Kankaanpää M. 1999. Comparative analysis of failure prediction methods: the Finnish case. European Accounting Review 8(1): 67–92.
Martinelli E, de Carvalho A, Rezende S, Matias A. 1999. Rules extractions from banks’ bankrupt data using connectionist and symbolic learning algorithms. In Proceedings of the Computational Finance Conference, New York, USA.
McSherry D. 1999. Strategic induction of decision trees. Knowledge-Based Systems 12(5–6): 269–275. Nilsson M, Sollenborn M. 2004. Advancements and trends in medical case-based reasoning: an overview of
systems and system development. In Proceedings of the 17th International FLAIRS Conference, Special Track on Case-Based Reasoning, AAAI, Miami, FL.
Ohlson JA. 1980. Financial ratios and the probabilistic prediction of bankruptcy. Journal of Accounting Research 18(1): 109–131.
Quinlan JR. 1986. Induction of decision trees. Machine Learning 1: 81–106. Quinlan JR. 1993. C4.5: Programs for Machine Learning. Morgan Kaufmann: San Francisco, CA. Roe BP, Yang H-J, Zhu J, Liu Y, Stancu I, McGregor G. 2005. Boosted decision trees as an alternative to artifi cial
neural networks for particle identifi cation. Nuclear Instruments and Methods in Physics Research A 543: 577–584.
Schmitz GPJ, Aldrich C, Gouws FS. 1999. ANN-DT: an algorithm for extraction of decision trees from artifi cial neural networks. IEEE Transactions on Neural Networks 10(6): 1392–1401.
Serfl ing R. 2000. Robust and nonparametric estimation via generalized L-statistics: theory, applications, and perspectives. In Advances in Methodological and Applied Aspects of Probability and Statistics, Balakrishnan N (ed.). Gordon & Breach; 197–217.
Shirata CY. 1998. Financial ratios as predictors of bankruptcy in Japan: an empirical research. In Proceedings of The Second Asian Pacifi c Interdisciplinary Research in Accounting Conference, Osaka, Japan; 437–445.
Tam KY. 1991. Neural networks and the prediction of bank bankruptcy. Omega 19: 429–445. Tam KY, Kiang M. 1992. Managerial applications of neural networks: the case of bank failure predictions.
Management Science 38(7): 926–947. Tan CNW. 2001. Artifi cial Neural Networks: Applications in Financial Distress Prediction and Foreign Exchange
Trading. Wilberto Publishing: Gold Coast, Australia. Theodossiou PT. 1993. Predicting shifts in the mean of a multivariate time series process: an application in
predicting business failure. Journal of the American Statistical Association 88(422): 441–449. Wishart D. 1998. Effi cient hierarchical cluster analysis for data mining and knowledge discovery. Computing
Science and Statistics 30: 257–263. Zelič I, Kononenko I, Lavrač N, Vuga V. 1997. Induction of decision trees and Bayesian classifi cation applied to
diagnosis of sport injuries. Journal of Medical Systems 21(6): 429–444. Zopounidis C, Dimitras AI. 1998. Multicriteria Decision Aid Methods for the Prediction of Business Failure.
Kluwer: Boston, MA.
Authors’ biographies: Adrian Gepp ([email protected]) currently teaches economics and fi nance at undergraduate and post- graduate levels at Bond University. He has published papers in the area of bankruptcy prediction, quantum algo- rithms and river system management. He graduated as a Master of Information Technology (Honours) in 2006 following undergraduate degrees in both commerce and information technology. He has also led a commercial software development project.
Business Failure Prediction using Decision Trees 555
Copyright © 2009 John Wiley & Sons, Ltd. J. Forecast. 29, 536–555 (2010) DOI: 10.1002/for
Kuldeep Kumar is Professor and Head, Department of Economics and Statistics at Bond University, Gold Coast, Australia. His research interests are in the areas of time series analysis and forecasting, bankruptcy prediction, fraud detection and breast cancer detection. Winner of several awards including Commonwealth Scholarship, Young Statistician Award, Bond-Oxford Fellowship, VC Quality award for research supervision, Dr Kumar has published more than 90 papers in various journals including Journal of Time Series Analysis and International Journal of Forecasting.
Sukanto Bhattacharya’s primary research interests are in the areas of credit rating forecast, bankruptcy prediction and fi nancial fraud detection. He has also published several papers in the areas of portfolio insurance and fi nancial engineering. He has a PhD degree in Information Technology specializing in Computational Finance from Bond University, Australia. Sukanto is presently a Senior Lecturer in the Deakin Business School, Deakin University, Australia and may be contacted by e-mail at [email protected]
Authors’ addresses: Adrian Gepp and Kuldeep Kumar, Department of Economics and Statistics, Faculty of Business, Technology and Sustainable Development, Bond University, Gold Coast, Queensland 4229, Australia.
Sukanto Bhattacharya, Deakin Business School, Deakin University, Burwood, Victoria 3125, Australia.
Copyright of Journal of Forecasting is the property of John Wiley & Sons, Inc. and its content may not be
copied or emailed to multiple sites or posted to a listserv without the copyright holder's express written
permission. However, users may print, download, or email articles for individual use.