Please answer the following questions below: Answers are giving in the manuals. I just need to you to shorten and edit them where it is not word from word. You can put each answer in excel sheet Chapter 8: Question(s) 8-1, 8-8, 8-12, 8-15, 8-22, 8-26, 8-2
SOLUTIONS CHAPTER 9
Transforming Data into Evidence (Part 2)
COVERAGE OF LEARNING OBJECTIVES
|
LEARNING OBJECTIVE |
QUESTIONS |
WORKPLACE APPLICATIONS |
CHAPTER PROBLEMS and CASES |
|
LO1. Explain the application of descriptive statistics in forensic accounting engagements. |
1, 2, 3, 4, 5, 6, 9, 10, 39–48 |
93 |
|
|
LO2. Identify and describe various methods for displaying data. |
11, 19, 20, 21, 49–58, 59–68 |
89, 93
|
|
|
LO3. Explain the purpose and application of data mining in forensic accounting engagements. |
22, 23, 24, 25, 26, 27, 69–78 |
90, 91, 92, 93 |
|
|
LO4. Identify examples of data analysis software, and explain the advantages and disadvantages of each. |
28, 29, 30, 31, 32 |
90, 91, 92 |
97 |
|
LO5. Explain Benford’s Law and describe specific digital analysis tests. |
33, 34, 35, 36, 37, 79–88 |
90, 91, 92 |
94, 95, 96 |
Questions
9-1. The purpose of the use of statistics in an analytical procedure is to summarize data, analyze them, and draw meaningful inferences that lead to improved decisions.
9-2. No. In statistical analysis, there is some element of imprecision. Thus, a key function of statistics is to define imprecision built into a statistical test. Defining imprecision is an important aspect of expert testimony, because such testimony must be based on sufficient relevant data and stated within a reasonable degree of professional certainty. Neither of these concepts requires absolute precision, but the forensic accountant must be aware of the level of potential imprecision represented by terms such as the error rate, significance level, or confidence level. Understanding the measure of imprecision allows the forensic accountant to determine if the data meets the requirements of sufficient relevant data and is stated within a reasonable degree of professional certainty.
9-3. The purpose of descriptive statistics is to describe data using various numerical measures and graphical depictions.
9-4. A population is the entire group of observations in which a forensic accountant is interested. A sample is a subset of the observations comprising a population.
9-5. The purpose of inferential statistics is to draw conclusions (or inferences) about a population based on information obtained from a sample.
9-6. Forensic accountants are less likely to use inferential statistics to analyze data in an engagement because the use of such statistics requires the employment of a random sample, and forensic engagement samples are normally drawn for a specific purpose and are not random.
9-7. The number of observations impacts the application of analytic methods employed in an analysis because it establishes what methods can be applied to the data, what technological resources will be applied, and how long the analysis will take.
9-8. Understanding the time and magnitude of data impacts an analysis as it indicates the nature of a distribution of data under examination. Data can be compiled at a single point in time (cross-sectional) or over some period of time (time series). The magnitude can be positive, negative, or zero.
9-9. Three measures of central tendency are the mean, median, and mode. The mean is the average of all observations in a data set and considers the magnitude of all the observations. The median is the center point of the data set and provides information as to whether observations are in the upper or lower half of a distribution. Because of this, the median is not impacted by extreme values (outliers). The mode is the most frequently occurring number in the data set.
9-10. Variability describes how observations are dispersed around the mean. Three measures of variability employed in the data analysis process are the range, the variance, and the standard deviation. The range is the difference between the values of the largest and smallest observations. The variance is the average of the squared deviations of the observation from the mean. The standard deviation is the variance squared.
9-11. A histogram is simply a bar chart.
9-12. The definition of intervals is important as the selection of intervals used to construct a histogram impacts its shape and explanatory power. In other words, the smaller the size of each interval, the larger the number of intervals used. Using a larger number of intervals provides a more detailed picture of the data distribution, but it may be misleading if the observations are heavily weighted in only a few intervals. In most cases, between ten and fifteen intervals is sufficient.
9-13. Absolute frequency is the count of observations within an interval. Relative frequency is the percentage of the total number of observations that fall within an interval.
9-14. Relative frequencies provide the analyst with standardized data (described relative to a standard quantity, such as the total number of observations). This is useful, as relative frequencies can be interpreted as probabilities—or in other words, the percentage of observations that fall within a selected interval range.
9-15. Skewness is the measure of the degree of asymmetry of a data distribution around its mean. Understanding these concepts impacts data analysis as it provides insight into where the majority of observations are located. If positively skewed data, the distribution is more heavily weighted toward smaller numbers. If negatively skewed, the distribution is more heavily weighted toward larger numbers.
9-16. In a discrete distribution, the observations are countable, and there are discreet gaps between successive values. In a continuous distribution, the observations can be measured to an infinitesimally small degree with no gaps between successive values. Examples are time, weight, and distance.
9-17. A normal distribution is symmetrically distributed around its mean (bell shaped), and the mean, median, and mode are equal. It is completely described by its mean and standard deviation.
9-18. The timing of transactions might impact the evaluation of the distribution of a data set as comparisons for different time periods might not be appropriate, such as sales for a calendar year compared to sales for a quarter or for the year to date. In addition, outside influences, such as a recession, can temporarily or permanently alter the shape of an expected distribution.
9-19. A pie chart is a graphical representation of a data set divided into logical components (or slices). Each component represents some percentage of the total data set, and the sum of all components must equal the total of the data set—whether the data are presented as percentages or absolute numbers. This is useful to viewers, as they do not need to know the specific values used to construct the chart but can tell from the sizes of the slices what the largest components are.
9-20. A bar chart presents observations within the distribution in terms of either absolute or relative frequency. The use of a bar chart allows for more than one category of data to be presented on the same graph. A line chart presents observations as points along a line rather than as the heights of bars. It can include more than one category of data. A line graph often provides a better representation of the relationship between the two data series.
9-21. The most common way in which bias might be incorporated into a graph is through the manipulation of the scale. Another way in which graphs can be manipulated is the inclusion (or lack thereof) of labels, including the title of the graph, labels of the horizontal and vertical axes, and labels of individual data points. A graph should always have a title, as this is the first and primary means by which the graph communicates information. Descriptive labels for the two axes may or may not be used, depending on their contribution to the overall clarity of the graph.
9-22. Data mining is a method that drills down into a data set to reduce a large number of observations to a smaller number that can be examined more closely. This is important, as time and resources normally do not allow for the examination of each and every item in a large data set.
9-23. Applications to which data mining might be employed are (students will select three):
1. Marketing research—predicting consumer demand and sales
2. Drug research—predicting the effectiveness of drugs and the likelihood of side effects
3. Credit scoring—predicting the likelihood of default or bankruptcy
4. Operations management—predicting input usage and productive efficiency
5. Investment analysis—predicting future changes in asset prices
6. Actuarial analysis—predicting life expectancies and probabilities of other insurable events
7. Fraud detection—predicting the likelihood that irregular transactions reflect unlawful practices
9-24. Data can be sorted by any of the individual data fields included in a database. Benefits that can be realized through the analysis of sorted data include the following items set forth in the chapter:
1. Identify duplicate entries.
2. Identify transactions with round numbers.
3. Identify gaps in the data sequence (such as dates, check numbers, or invoice numbers).
4. Identify matches in data fields (such as employees and vendors with the same contact information).
5. Compute category totals (such as total payments made to a specific vendor or employee or total payments for a specific expense category).
6. Highlight blanks (or lack of data) in a particular data field (such as employees without a Social Security number).
7. Identify inconsistencies among data fields (such as incompatible telephone numbers and addresses or back-dated checks).
9-25. Ratio analysis involves the calculation of ratios for key observations in a data set. This form of data mining often involves the use of the largest and second largest values as well as the smallest value. Four ratios that might be calculated once data have been sorted are (as presented in the chapter):
1. Ratio of the largest value to the smallest value, which is similar in concept to the range. A larger ratio indicates greater variation in the data set.
2. Ratio of the highest value to the second highest value, known as the relative size factor (RSF). A large RSF indicates an outlier in the data set. The value of the RSF is that it provides a numeric measure that can be compared to a benchmark or tracked over time.
3. Ratio of the smallest value to the second smallest value, Which, like the RSF, identifies outliers, but on the opposite side of the distribution.
4. Ratio of the largest (or smallest) value to mean, another means of identifying an outlier, using a different reference point.
9-26. A Type I error (or false positive) is the identification of an item in a data set that should NOT have been selected. A Type II error (or false negative) is not identifying an item from a data set that should have been selected.
9-27. A forensic accountant is primarily concerned with minimizing the risk of identifying false positives because after specific observations have been identified, substantial effort is required to individually examine each of them. In contrast to the data mining process, in which technological tools can be leveraged, examination of individual observations requires the close involvement of the forensic accountant. Thus, false positives unnecessarily consume the time and resources of a forensic accountant.
9-28. The advantages of using spreadsheet programs, such as Excel and Access, in the data analysis process include:
1. Affordability
2. Ease of use
3. Flexibility of application
4. The variety of functions available for use
9-29. Excel and Access have similar capabilities. The attributes of each, as they were presented in the chapter, are set forth below:
Excel is more popular among practitioners, probably due to their past experience with the program for general accounting tasks. A key difference between the two programs is that Access forces some structure on the data analysis project, while Excel allows more flexibility. For example, in Excel, either a number or a formula can be entered into a cell. If a number appears in a cell, it is difficult to determine whether it is a data point or a calculation (that is, the result of a formula) until the cell is clicked. Similarly, column headings can be duplicated, and there is no requirement that consistency be maintained between column headings and the type of data contained in the column.
Access has a well-defined structure that includes tables, queries, and reports. Data are stored in tables, which have a layout similar to an Excel spreadsheet. Each row in the table contains data for a single record (such as a transaction), the discrete components of which are stored in fields (columns). Each field can store only one type of data for all the records. All data-related functions are applied through the creation of queries, the results of which are displayed in reports. Thus, with Access, it is always clear whether you are looking at data (input) or results (output)
9-30. Two add-ins that can be acquired to extend the capabilities of spreadsheet products are Analysis Toolpak, which is included in the Excel software, and ActiveData for Excel, which must be purchased separately and is available in professional and business (fewer features) versions.
9-31. Disadvantages related to the use of spreadsheets in the data analysis process include:
1. Spreadsheets allow data to be altered, either intentionally or unintentionally, without a record of such alteration.
2. Errors can be easily introduced into spreadsheets, commonly through formulas, the copy/paste function, incorrect cell references, or improperly defined cell ranges.
3. Spreadsheet programs cannot accommodate data in certain formats, in which case the data must be converted. This conversion process introduces yet another opportunity for the integrity of the data to be compromised.
9-32. Two common specialized software packages that are useful for auditing and fraud investigations are Audit Command Language (ACL) and Interactive Data Extraction Analysis (IDEA). The primary advantage of these programs is their user-interface attributes, which are customized for specific tasks. In contrast to spreadsheet programs, which can facilitate data analysis for a variety of purposes (but less efficiently), GAS programs are designed to analyze financial data in an auditing environment. Other advantages are:
1. They can process data in a wide variety of formats, which eliminates the need for data conversion.
2. They can record the analytics that have been performed, thus creating an audit trail or log.
9-33. Digital analysis is founded on the counterintuitive observation that individual digits of multidigit numbers are not random but follow a pattern known as Benford’s Law.
9-34. Benford’s Law posits that the distribution of first digits is positively skewed, or more heavily weighted toward smaller numbers. In other words, the first digit (the left-most digit) of numbers is more often low than high—digit 1 appears more frequently than digit 2, which appears more frequently than digit 3 and so on.
9-35. In order for a data set to conform to Benford’s Law, its number series must approximately follow a geometric sequence, in which each successive number is calculated as a fixed percentage increase over the previous number. Almost all natural numbers display such a geometric tendency, including city populations, the sizes of geologic objects, and accounting numbers (such as stock prices, company revenues, and trading volume).
9-36. Four studies that have used Benford’s Law to successfully analyze data are (as set forth in the chapter):
1. Net income, based on analyses of income data for New Zealand and U.S. companies.
2. Earnings per share, in a study of U.S. corporate reports displaying unusually high frequencies of 5- and 10-cent multiples.
3. Income tax, as evident in the tendency among individuals to claim additional deductions for the purpose of decreasing taxable income to fall within the next-lowest tax bracket.
4. Fraud detection, first employed in a 1994 digital analysis of records in long-term payroll fraud; the comparative analysis identified fraudulent payroll data and showed that deviations increased over time.
9-37. A first-digit test compares the first-digit profile of a data set to Benford’s first-digit profile. It is useful in data analysis because variations in the profiles of numbers from 1 to 9 of the data set are easily identified when compared to Benford’s profile.
9-38. Size is a consideration in the determination to employ a Benford test because close conformity to Benford’s Law requires a large data set (often defined as at least 1,000 observations) with numbers having at least four digits. The size of the data set is important because it is more difficult to identify significant deviations from Benford’s profile in small data sets.
Multiple-Choice Questions
Select the best response to the following questions related to descriptive statistics:
9-39. A
9-40. B
9-41. D
9-42. B
9-43. A
9-44. D
9-45. A
9-46. C
9-47. B
9-48. D
Select the best response to the following questions related to the shape of data distributions:
9-49. C
9-50. A
9-51. A
9-52. D
9-53. A
9-54. B
9-55. A
9-56. B
9-57. B
9-58. C
Select the best response to the following questions related to methods of displaying data:
9-59. B
9-60. D
9-61. A
9-62. A
9-63. A
9-64. B
9-65. C
9-66. D
9-67. A
9-68. A
Select the best response to the following questions related to data mining:
9-69. B
9-70. A
9-71. D
9-72. C
9-73. D
9-74. B
9-75. A
9-76. A
9-77. D
9-78. B
Select the best response to the following questions related to Benford’s Law:
9-79. B
9-80. D
9-81. B
9-82. A
9-83. A
9-84. D
9-85. C
9-86. C
9-87. C
9-88. B
Workplace Applications
9-89. This questions relates to presentation of data in charts.
1. The two options are a line and a bar chart. For comparison of the three earnings streams, the line graph presents the better visual representation. Samples include:
Line Graph
Bar Chart
2. The line graph clearly shows a relationship among the three earnings measures. The strength of the earnings relationship can be quantified through the use of correlation analysis. As the following table shows, the Pearson correlation coefficient is strong for each earnings measure, but only significant for the relationship between net income and the net cash flow from operations earnings streams.
|
Correlations |
||||
|
|
IncomeOPER |
NetIncome |
NCFOPER |
|
|
IncomeOPER |
Pearson Correlation |
1 |
.876 |
.720 |
|
|
Sig. (2-tailed) |
|
.052 |
.170 |
|
|
N |
5 |
5 |
5 |
|
NetIncome |
Pearson Correlation |
.876 |
1 |
.895* |
|
|
Sig. (2-tailed) |
.052 |
|
.040 |
|
|
N |
5 |
5 |
5 |
|
NCFOPER |
Pearson Correlation |
.720 |
.895* |
1 |
|
|
Sig. (2-tailed) |
.170 |
.040 |
|
|
|
N |
5 |
5 |
5 |
|
*Correlation is significant at the 0.05 level (2-tailed).
|
3. The line graph can be altered to make the changes look more dramatic by narrowing the time columns (interval) and starting the graph at $4,000 rather than at zero.
9-90. This question relates to Benford’s Law and an Excel spreadsheet model called Fraud Buster that can be used to develop statistics and graphs that show the data’s distribution compared to that of a Benford distribution. The Excel spreadsheet needed to complete this problem is obtainable from the AICPA web site via a link at www.pearsonhighered.com/rufus.
1. The Excel data file for 1978 taxable income contains 157,519 observations. Obtained from http://www.nigrini.com/benfordslaw.htm, or via a link at www.pearsonhighered.com/rufus.
2. The first 6,000 taxable incomes (observations) when analyzed using Fraud Buster finds 5,780 observations with digits between 1 and 9. The table output for the first-digit test follows:
Benford’s Law First-Digit Test
|
|
|
|
|
|
|
|
Sample |
Benford |
Sample Data |
|
|
Digit |
Frequency |
Rate |
Rate |
Difference |
|
1 |
1587 |
0.30103 |
0.274567474 |
-0.02646252 |
|
2 |
1202 |
0.1760913 |
0.207958478 |
0.03186722 |
|
3 |
727 |
0.1249387 |
0.125778547 |
0.00083981 |
|
4 |
747 |
0.09691 |
0.129238754 |
0.03232874 |
|
5 |
579 |
0.0791812 |
0.10017301 |
0.02099176 |
|
6 |
294 |
0.0669468 |
0.050865052 |
-0.01608174 |
|
7 |
216 |
0.0579919 |
0.037370242 |
-0.0206217 |
|
8 |
213 |
0.0511525 |
0.036851211 |
-0.01430131 |
|
9 |
215 |
0.0457575 |
0.037197232 |
-0.00856026 |
The results show that the taxable income sample follows a Benford distribution. This can be seen visually in the following graph.
Benford’s Law Second-Digit Test
The results of the second-digit test show consistency between the 1978 taxable income distribution and Benford’s distribution.
|
|
Sample |
Benford |
Sample Data |
|
|
Digit |
Frequency |
Rate |
Rate |
Difference |
|
0 |
725 |
0.11968 |
0.125432526 |
0.005752526 |
|
1 |
619 |
0.11389 |
0.107093426 |
-0.00679657 |
|
2 |
620 |
0.10822 |
0.107266436 |
-0.00095356 |
|
3 |
608 |
0.10433 |
0.105190311 |
0.000860311 |
|
4 |
556 |
0.10031 |
0.096193772 |
-0.00411623 |
|
5 |
559 |
0.09668 |
0.096712803 |
3.28028E-05 |
|
6 |
537 |
0.09337 |
0.092906574 |
-0.00046343 |
|
7 |
497 |
0.09035 |
0.085986159 |
-0.00436384 |
|
8 |
530 |
0.08757 |
0.091695502 |
0.004125502 |
|
9 |
529 |
0.085 |
0.091522491 |
0.006522491 |
9-91. Using the Fraud Buster spreadsheet and the information on 1978 taxable incomes obtained from Dr. Nigrini’s web site, students will run a Fraud Buster test on only the first 375 observations (373 observations meet the 1–9 digit requirement) appearing in the data set.
1. The results of the first-digit test compared to Benford’s predicted distribution are as follows. There is clearly a significant difference at the first- and second-digit level. It would be appropriate to investigate the cause of this disparity in results.
The likely cause is the small sample. As discussed in the chapter, Benford’s law is more applicable to large data sets that approximate a normal distribution. While we do not know for sure, the taxable incomes for the first 375 observations in the data set might have been drawn from a less affluent segment of the population exhibiting low levels of taxable income—for example, a disproportionate number of incomes in the $10,000 to $30,000 range.
This problem illustrates the need of the forensic accountant to understand his/her data when using digital analysis.
|
|
Sample |
Benford |
Sample Data |
|
|
Digit |
Frequency |
Rate |
Rate |
Difference |
|
1 |
201 |
0.30103 |
0.538873995 |
0.237844 |
|
2 |
117 |
0.1760913 |
0.313672922 |
0.13758166 |
|
3 |
8 |
0.1249387 |
0.021447721 |
-0.10349102 |
|
4 |
4 |
0.09691 |
0.010723861 |
-0.08618615 |
|
5 |
9 |
0.0791812 |
0.024128686 |
-0.05505256 |
|
6 |
9 |
0.0669468 |
0.024128686 |
-0.0428181 |
|
7 |
8 |
0.0579919 |
0.021447721 |
-0.03654423 |
|
8 |
9 |
0.0511525 |
0.024128686 |
-0.02702384 |
|
9 |
8 |
0.0457575 |
0.021447721 |
-0.02430977 |
9-92. The Fraud Buster spreadsheet described in problems 9-90 and 9-91 must be downloaded in order to complete this problem. In the data entry column, students should use the Excel Rand function to generate random numbers. The formula is =RAND()*(1000-1)+1. This should provide a random number between $1 and $999. Once students generate the first random number in cell A1, they must copy the formula so that they have a data set of 1,000 randomly generated observations. They then run Fraud Buster and analyze the first-digit results.
While each randomly generated data set will contain different numbers, the distribution should be 11.11% for each of the digits in the random sample. For example, see the following output from this test for 1,000 observations. The sample data rate varies slightly from the expected 11.11% of observations expected from a randomly drawn sample. This is attributable to the small sample size.
First-Digit Test 1,000 Random Numbers
|
|
Sample |
Benford |
Sample Data |
Benford |
Difference |
|
Digit |
Frequency |
Rate |
Rate |
Difference |
from 11.11 |
|
1 |
125 |
0.30103 |
0.125 |
-0.17603 |
0.0139 |
|
2 |
108 |
0.1760913 |
0.108 |
-0.06809126 |
-0.0031 |
|
3 |
124 |
0.1249387 |
0.124 |
-0.00093874 |
0.0129 |
|
4 |
109 |
0.09691 |
0.109 |
0.01208999 |
-0.0021 |
|
5 |
119 |
0.0791812 |
0.119 |
0.03981875 |
0.0079 |
|
6 |
92 |
0.0669468 |
0.092 |
0.02505321 |
-0.0191 |
|
7 |
99 |
0.0579919 |
0.099 |
0.04100805 |
-0.0121 |
|
8 |
113 |
0.0511525 |
0.113 |
0.06184748 |
0.0019 |
|
9 |
111 |
0.0457575 |
0.111 |
0.06524251 |
-0.0001 |
When the sample size is increased to 25,000 observations, the variance of each digit’s proportion from the expected 11.11% is significantly reduced as shown in the table below.
First-Digit Test 25,000 Random Numbers
|
|
Sample |
Benford |
Sample Data |
Benford |
Difference |
|
Digit |
Frequency |
Rate |
Rate |
Difference |
from 11.11 |
|
1 |
2797 |
0.30103 |
0.11188 |
-0.18915 |
0.00078 |
|
2 |
2818 |
0.1760913 |
0.11272 |
-0.06337126 |
0.00162 |
|
3 |
2776 |
0.1249387 |
0.11104 |
-0.01389874 |
-6E-05 |
|
4 |
2808 |
0.09691 |
0.11232 |
0.01540999 |
0.00122 |
|
5 |
2720 |
0.0791812 |
0.1088 |
0.02961875 |
-0.0023 |
|
6 |
2777 |
0.0669468 |
0.11108 |
0.04413321 |
-2E-05 |
|
7 |
2805 |
0.0579919 |
0.1122 |
0.05420805 |
0.0011 |
|
8 |
2805 |
0.0511525 |
0.1122 |
0.06104748 |
0.0011 |
|
9 |
2694 |
0.0457575 |
0.10776 |
0.06200251 |
-0.00334 |
9-93. This data mining exercise has six parts.
1. Sorting the file by graduation rates by categories of 25% or less, 26% to 50%, 51% to 75%, 76% to 100% yields the following results.
The following solution can be obtained using the COUNTIF function. This results in the following:
25% of less42
26% to 50%324
51% to 75%571
76% to 100%267
Total1204
Count IF Test on Grad Rate
2. One version of a Pie Chart developed from the step 1 solution is:
The graph shows that 74% of all institutions have a graduation rate of more than 50% and only 4% of institutions have a graduation rate of 25% or lower.
3. Two types of data were not usable. First, 98 colleges had no graduation rate available in the data set. Second, one college showed a 118% graduation rate. Both categories of unusable data were omitted from further analysis.
4. It might be useful to compare the graduation rates for the bottom 5% and the upper 5% to see if there are important differences.
It might also be useful to analyze colleges with graduation rates of 95% to 100%, as the student-to-faculty ratio drops to 11:1 for this subset of colleges.
Coding the data by state and comparing average graduation rates might provide additional insight.
5. Data that might cause the analyses to be misleading or suggest to an analyst that further investigation is warranted are:
a. There are missing data. Most columns have missing data, such as SAT and ACT scores, and % new students from the top 10% and 25% of their high school classes, to name a few.
b. The 75% to 100% graduation rate category is highly populated by schools with quality brands. Separating this group by perceived quality might be revealing.
c. Professional skepticism suggests that not even the best schools graduate 100% of enrolled students. Further investigation into this area is warranted.
d. Not all higher education institutions are listed in the data set. For example, the Excel search feature was not able to find a listing for large for-profit institutions such as Kaplan University, the University of Phoenix, and Strayer University. Other small not-for-profit colleges are not included as well (such as Southeastern University in Lakeland, Florida). Thus, understanding how colleges were selected for inclusion in the data set would be important to interpreting the results of the analysis.
6. Statistical tests that might reveal relationships between graduation rates and other data presented in the data set would include correlation and regression analysis.
This problem does not expect students to perform regression and correlation analyses since they are not reviewed in the chapter. However, a correlation analysis is provided for the benefit of instructors who might require students to conduct these tests as part of their analysis.
A summary of several separate correlation analyses is presented in the following table.
Graduation
Rate
Categories
Public (1)/
Private (2)
Math SATVerbal
SAT
Total
SAT
ACT# appli.
rec'd
# appl.
accept
ed
# new
stud.
enrolled
% new
stud.
from top
10%
% new
stud.
from top
25%
# FT
undergra
d
# PT
undergra
d
in-state
tuition
out-of-
state
tuition
roomboardadd. feesestim.
book
costs
estim.
persona
l $
% fac.
w/PHD
stud./f
ac.
ratio
Gradu
ation
rate
25% of less1.4448412461201808131274911373250210739396100224517122965541642631620
26% to 50%1.4468423891212371176689916414563189943416552210017403605571629651741
51% to 75%1.751046397322287620768002553387782582529598260021074165391362691562
76% or more1.956151110722536551975698406729024621291713301289824184035631150771386
Average1.6497452849222678178278623493648132373628888246119943695531446691552
Correlation to Graduation Rate
25% or less-0.0440.198-0.1420.181-0.003-0.092-0.113-0.177-0.205-0.2720.0940.1260.1390.2310.3930.0040.157-0.140-0.0730.082-0.1341.000
26% to 50%0.1250.3100.3280.0880.2860.1130.0950.0460.1360.1910.0560.0050.1650.2710.1310.1010.0050.0600.0600.076-0.1111.000
51% to 75%0.2530.2780.2870.2930.3170.0680.0680.0110.2680.273-0.016-0.0880.3550.3800.0900.2220.025-0.006-0.1120.192-0.2121.000
76% or more0.1150.3460.3780.3670.2490.141-0.0340.0520.3460.3490.003-0.1200.2470.2640.1020.314-0.0450.180-0.0030.130-0.2651.000
Average0.1120.2830.2130.2320.2120.0580.004-0.0170.1360.1350.034-0.0190.2260.2860.1790.1600.0350.023-0.0320.120-0.1811.000
95% to 100%1.960254411472650861848769537831474061409214407315127493756231190771197
Correlation0.077-0.131-0.076-0.109-0.314-0.137-0.184-0.132-0.024-0.105-0.1170.061-0.185-0.242-0.071-0.226-0.113-0.2200.175-0.0560.2541.000
To develop the above summary, the data set was subdivided by category and correlations were calculated using the Excel Data Analysis tool. For each of the categories, the correlations were run for all data columns, but only the correlations for each data column as it relates to graduation rate is shown in the lower portion of the table. In the top portion of the table, averages for all institutions in each category are shown for presentation purposes.
Did any of the correlations provide insight as to the differences in graduation rates? Yes, but most correlation coefficients are not revealing in terms of graduation rate percent. This observation is very clear for the 25% or less category.
As graduation rates go up, SAT and ACT scores have moderate correlation coefficients, as do the percent of new students in the top 10% and 25% of their high school classes. Correlation coefficients are higher for the 51% to 75% and the 75% to 100% categories.
Likewise, in- and out-of-state tuition rates and percent of faculty with doctoral degrees are more highly correlated for the top two graduation rate categories.
Finally, class size matters as shown by the negative correlation coefficients between class size and higher graduation rates. Thus, as class size goes down, graduation rates increase. This becomes more apparent when the data is parsed and a 95% to 100% graduation category is examined. For this group of colleges, the class size drops to 11.1 students.
Based on the correlation analysis, it appears that recruiting higher achieving students, keeping the percent of faculty with Ph.D.s at a high level, and maintaining low student/faculty ratios are key determinants of graduation rates.
Chapter Problems
9-94. The Durtschi, Hillison, and Pacini article entitled “The Effective Use of Benford’s Law to Assist in Detecting Fraud in Accounting Data,” published in the Journal of Forensic Accounting, presents a nice summary of the issues related to Benford’s Law as it applies to accounting. The emphasis of the paper is to provide auditors with guidance that highlights how digital analysis might be useful in detecting fraud. Guidance on how to interpret the results of a Benford test is provided as well.
Student papers will include sections on the following aspects of the article:
1. The history of Benford’s Law. This section relates how Simon Newcomb discovered the concept in 1881 by observing wear and tear on pages in a book of logarithms. The math underlying this concept is presented, and a table with Benford’s percentages for digits 1, 2, 3, and 4 is provided.
2. The application of Benford’s Law to auditing and accounting. This section credits Nigrini as the first researcher to apply Benford’s Law to accounting numbers in an effort to detect fraud.
3. When to use digital analysis. This section begins with a caution that Benford’s Law is less effective as the number of contaminated observations increases. It also points out that fraud is rare and many data sets that show nonconforming distributions do not contain fraud. Statistical considerations are covered, and cases where digital analysis is not valuable are also presented (e.g., duplicate address, bank account numbers, and other numbers that have numerical constraints or are arbitrarily ordered for some administrative purpose).
The use of digital analysis as a screening device is discussed, and a sample problem is presented to illustrate the use of digital analysis as a screening tool.
9-95. The article written by Rose and Rose entitled “Turn Excel into a Financial Sleuth,” published in the Journal of Accountancy, is available via a link at www.pearsonhighered.com/rufus.
1. The article presents a step-by-step outline, including relevant screen shots, about how to perform a digital analysis using an Excel-based model called Fraud Buster. The Fraud Buster Excel spreadsheet can be downloaded from the AICPA web site using the link provided by the authors, available at www.pearsonhighered.com/rufus.
2. The Fraud Buster spreadsheet includes tabs containing the data set, as well as tabs for the first-digit, second-digit, and first-two digit tests.
The analysis tables are presented in both table and chart formats. For the first-digit test, the output appears as follows: The output shows significant difference in the actual distribution compared to the Benford distribution at the 1st, 2nd, 5th, 6th, and 7th digits. Additional tests would be appropriate to determine the cause of the variation.
First-Digit Test of Data Set Provided by the Authors
SampleBenfordSample Data
DigitFrequencyRateRateDifference
150.301030.064935065-0.23609493
230.17609130.038961039-0.13713022
390.12493870.116883117-0.00805562
4100.096910.129870130.03296012
5220.07918120.2857142860.20653304
6110.06694680.1428571430.07591035
7110.05799190.1428571430.0848652
820.05115250.025974026-0.0251785
940.04575750.0519480520.00619056
3. Memos from students explaining what they discovered about the distribution of the data set included in the Fraud Buster spreadsheet and what they learned about digital testing will vary based on individual experience with the concept. The memos should, at a minimum, include the above observations.
9-96. A Google search using the search term “Using Digital Analysis to Enhance Data Integrity” will locate this article by Ettredge and Srivastava. The first nine pages of this document present a lesson on digital analysis, and the remaining twenty pages include teaching notes on how to use digital analysis in a classroom setting. Student papers should discuss each of the following concepts covered in the first nine pages of the paper.
1. The introduction uses an example as a way to define digital analysis and suggests that digital analysis can be used to help identify deviant number patterns in a variety of settings.
2. Examples of how digital analysis has been successfully employed include a finding that in certain circumstances taxpayers bias their taxable income downward. The use of this technique in fraud detection and auditing was also mentioned.
The authors point out that fraud is not a common occurrence and that digital analysis is useful as a way to find errors and to monitor compliance with established policies.
3. The final section sets forth the mathematics underlying a Benford distribution as well as the probability of the number 1 being the first digit of a set of observations occurring 30.1% of the time. A full table of probabilities is not provided.
Finally, the article points out that even if a set of data is not constructed in a manner that accommodates a Benford test, an organization can use its unique data to develop a distribution of probabilities of occurrence that can be expected and then use the expected distribution to test future transactions for variation from what is expected/normal for that particular function.
9-97. Via a link at www.pearsonhighered.com/rufus, students may obtain the areas of the 254 counties in Texas.
1. Once they copy and paste this data into an Excel spreadsheet, they extract the first digit from the data series using the LEFT function. This data is not shown as it is simply a number 1 through 9. Using this data, they complete step 2.
2. Complete the Actual Frequency and Relative Frequency columns in the following table. (Note: Use the COUNTIF function for Absolute Frequency).
|
Digit |
Absolute Frequency |
Relative Frequency |
Benford Frequency |
Difference |
|
1 |
63 |
24.8% |
0.301 |
-5.3% |
|
2 |
11 |
4.3% |
0.176 |
-13.3% |
|
3 |
7 |
2.8% |
0.125 |
-9.7% |
|
4 |
6 |
2.4% |
0.097 |
-7.3% |
|
5 |
10 |
3.9% |
0.079 |
-4.0% |
|
6 |
15 |
5.9% |
0.067 |
-0.8% |
|
7 |
19 |
7.5% |
0.058 |
1.7% |
|
8 |
47 |
18.5% |
0.051 |
13.4% |
|
9 |
76 |
29.9% |
0.046 |
25.3% |
|
Totals |
254 |
100.0% |
100.0% |
0.0% |
3. A line graph of the relative frequency and Benford frequency is as follows. Also provided is a bar graph generated using the Fraud Buster spreadsheet used in prior problems.
Line Graph
Fraud Buster Bar Chart
4. The data set does not conform to Benford’s Law. Digits 2, 3, 4 and 7, 8, 9 vary from the expected probabilities of occurrence.
5. It is not reasonable to expect the areas of the counties in Texas to follow Benford’s Law. Benford’s Law assumes that a number series must approximately follow a geometric sequence, in which each successive number is calculated as a fixed percentage increase over the previous number. The land areas are determined using factors that are based on some prescribed criteria and would not meet the natural progression of numbers requirement.
Most likely, the intent was to keep the counties roughly similar in size, which appears to be between 800 and 1,200. The exception is ten westernmost counties, which are significantly larger than those in the remainder of the state. A Google search using “Texas Counties” will provide maps that can be used to visualize the county sizes. Thus, more digits would fall in the 8, 9, and 1 places in the county’s area profile.
PAGE
180
Copyright © 2015 Pearson Education, Inc.