Quantitative Methods in Finance
CFA Institute is the premier association for investment professionals around the world, with over 130,000 members in 151 countries and territories. Since 1963 the organization has developed and administered the renowned Chartered Financial analyst® Program. With a rich history of leading the investment profession, CFA Institute has set the highest standards in ethics, education, and professional excellence within the global investment community and is the foremost authority on investment profession conduct and practice. Each book in the CFA Institute investment Series is geared toward industry practitioners along with graduate-level finance students and covers the most important topics in the industry. The authors of these cutting-edge books are themselves industry professionals and academics and bring their wealth of knowledge and expertise to this series.
QUANTITATIVE INVESTMENT ANALYSIS Third Edition
Richard A. DeFusco, CFA
Dennis W. McLeavey, CFA
Jerald E. Pinto, CFA
David E. Runkle, CFA
Cover image: © r.nagy/Shutterstock Cover design: Wiley Copyright © 2004, 2007, 2015 by CFA insti tute. All rights reserved.
Published by John Wiley & Sons, Inc., Hoboken, New Jersey. Published simultaneously in Canada.
No part of this publication may be reproduced, stored in a retrieval system, or transmitted in any form or by any means, electronic, mechanical , photocopying, recording, scanning, or otherwise, except as permitted under Section 107 or 108 of the 1976 United States Copyright act, without ei ther the prior wri tten permission of the Publisher, or authorization through payment of the appropriate per-copy fee to the Copyright Clearance Center, Inc., 222 Rosewood Drive, Danvers, MA 01923, (978) 750-8400, fax (978) 646-8600, or on the Web at www.copyright.com. Requests to the Publisher for permission should be addressed to the Permissions Department, John Wiley & Sons, Inc., 111 River Street, Hoboken, NJ 07030, (201) 748-6011, fax (201) 748-6008, or onl ine at http://www.wiley.com/go/permissions.
Limit of l iabi l i ty/Disclaimer of Warranty: While the publisher and author have used their best efforts in preparing this book , they make no representations or warranties with respect to the accuracy or completeness of the contents of this book and specifical ly disclaim any implied warranties of merchantabi l i ty or fi tness for a particular purpose. No warranty may be created or extended by sales representatives or wri tten sales materials. The advice and strategies contained herein may not be sui table for your si tuation. You should consult with a professional where appropriate. Neither the publisher nor author shal l be l iable for any loss of profi t or any other commercial damages, including but not l imited to special , incidental , consequential , or other damages.
For general information on our other products and services or for technical support, please contact our Customer Care Department within the uni ted States at (800) 762-2974, outside the uni ted States at (317) 572-3993, or fax (317) 572-4002.
Wiley publishes in a variety of print and electronic formats and by print-on-demand. Some material included with standard print versions of this book may not be included in e-books or in print-on-demand. If this book refers to media such as a CD or DVD that is not included in the version you purchased, you may download this material at http://booksupport.wiley.com. For more information about Wiley products, visi t www.wiley.com.
ISBN 978-1-119-10422-3 (Hardcover) ISBN 978-1-119-10459-9 (ePDF)
ISBN 978-1-119-10460-5 (ePub)
CONTENTS 1. Foreword 2. Preface 3. Acknowledgment 4. About the CFA Institute Investment Series 5. CHAPTER 1: The Time Value of Money
1. 1. Introduction 2. 2. Interest Rates: Interpretation 3. 3. The Future Value of a Single Cash Flow 4. 4. The Future Value of a Series of Cash Flows 5. 5. The Present Value of a Single Cash Flow 6. 6. The Present Value of a Series of Cash Flows 7. 7. Solving for Rates, Number of Periods, or Size of Annuity Payments 8. 8. Summary 9. Problems 10. Notes
6. CHAPTER 2: Discounted Cash Flow Applications 1. 1. Introduction 2. 2. Net Present Value and Internal Rate of Return 3. 3. Portfolio Return Measurement 4. 4. Money Market Yields 5. 5. Summary 6. References 7. Problems 8. Notes
7. CHAPTER 3: Statistical Concepts and Market Returns 1. 1. Introduction 2. 2. Some Fundamental Concepts 3. 3. Summarizing Data Using Frequency Distributions 4. 4. The Graphic Presentation of Data
5. 5. Measures of Central Tendency 6. 6. Other Measures of Location: Quantiles 7. 7. Measures of Dispersion 8. 8. Symmetry and Skewness in Return Distributions 9. 9. Kurtosis in Return Distributions 10. 10. Using Geometric and Arithmetic Means 11. 11. Summary 12. References 13. Problems 14. Notes
8. CHAPTER 4: Probability Concepts 1. 1. Introduction 2. 2. Probability, Expected Value, and Variance 3. 3. Portfolio Expected Return and Variance of Return 4. 4. Topics in Probability 5. 5. Summary 6. References 7. Problems 8. Notes
9. CHAPTER 5: Common Probability Distributions 1. 1. Introduction to Common Probability Distributions 2. 2. Discrete Random Variables 3. 3. Continuous Random Variables 4. 4. Monte Carlo Simulation 5. 5. Summary 6. References 7. Problems 8. Notes
10. CHAPTER 6: Sampling and Estimation 1. 1. Introduction 2. 2. Sampling
3. 3. Distribution of the Sample Mean 4. 4. Point and Interval Estimates of the Population Mean 5. 5. More on Sampling 6. 6. Summary 7. References 8. Problems 9. Notes
11. CHAPTER 7: Hypothesis Testing 1. 1. Introduction 2. 2. Hypothesis Testing 3. 3. Hypothesis Tests Concerning the Mean 4. 4. Hypothesis Tests Concerning Variance 5. 5. Other Issues: Nonparametric Inference 6. 6. Summary 7. References 8. Problems 9. Notes
12. CHAPTER 8: Correlation and Regression 1. 1. Introduction 2. 2. Correlation Analysis 3. 3. Linear Regression 4. 4. Summary 5. Problems 6. Notes
13. CHAPTER 9: Multiple Regression and Issues in Regression Analysis 1. 1. Introduction 2. 2. Multiple Linear Regression 3. 3. Using Dummy Variables in Regressions 4. 4. Violations of Regression Assumptions 5. 5. Model Specification and Errors in Specification 6. 6. Models with Qualitative Dependent Variables
7. 7. Summary 8. References 9. Problems 10. Notes
14. CHAPTER 10 : Time-Series Analysis 1. 1. Introduction to Time-Series Analysis 2. 2. Challenges of Working with Time Series 3. 3. Trend Models 4. 4. Autoregressive (AR) Time-Series Models 5. 5. Random Walks and Unit Roots 6. 6. Moving-Average Time-Series Models 7. 7. Seasonality in Time-Series Models 8. 8. Autoregressive Moving-Average Models 9. 9. Autoregressive Conditional Heteroskedasticity Models 10. 10. Regressions with More than One Time Series 11. 11. Other Issues in Time Series 12. 12. Suggested Steps in Time-Series Forecasting 13. 13. Summary 14. Problems 15. Notes
15. CHAPTER 11: An Introduction to Multifactor Models 1. 1. Introduction 2. 2. Multifactor Models and Modern Portfolio Theory 3. 3. Arbitrage Pricing Theory 4. 4. Multifactor Models: Types 5. 5. Multifactor Models: Selected Applications 6. 6. Summary 7. References 8. Problems
16. Appendices 17. Glossary
18. About the Editors and Authors 19. About the CFA Program 20. Index 21. Advert 22. EULA
List of Tables
1. Chapter 1
1. TABLE 1
2. TABLE 2
3. TABLE 3
4. TABLE 4
5. TABLE 5
6. TABLE 6
2. Chapter 2
1. TABLE 1
2. TABLE 2
3. TABLE 3
4. TABLE 4
5. TABLE 5
6. TABLE 6
7. TABLE 7
8. TABLE 8
3. Chapter 3
1. TABLE 1
2. TABLE 2
3. TABLE 3
4. TABLE 4
5. TABLE 5
6. TABLE 6
7. TABLE 7
8. TABLE 8
9. TABLE 9
10. TABLE 10
11. TABLE 11
12. TABLE 12
13. TABLE 13
14. TABLE 14
15. TABLE 15
16. TABLE 16
17. TABLE 17
18. TABLE 18
19. TABLE 19
20. TABLE 20
21. TABLE 21
22. TABLE 22
23. TABLE 23
24. TABLE 24
25. TABLE 25
26. TABLE 26
27. TABLE 27
28. TABLE 28
29. TABLE 29
30. TABLE 30
4. Chapter 4
1. TABLE 1
2. TABLE 2
3. TABLE 3
4. TABLE 4
5. TABLE 5
6. TABLE 6
7. TABLE 7
8. TABLE 8
9. TABLE 9
10. TABLE 10
11. TABLE 11
12. TABLE 12
13. TABLE 13
5. Chapter 5
1. TABLE 1
2. TABLE 2
3. TABLE 3
4. TABLE 4
5. TABLE 5
6. TABLE 6
7. TABLE 7
8. TABLE 8
9. TABLE 9
6. Chapter 6
1. TABLE 1
2. TABLE 2
3. TABLE 3
4. TABLE 4
7. Chapter 7
1. TABLE 1
2. TABLE 2
3. TABLE 3
4. TABLE 4
5. TABLE 5
6. TABLE 6
7. TABLE 7
8. TABLE 8
9. TABLE 9
10. TABLE 10
11. TABLE 11
8. Chapter 8
1. TABLE 1
2. TABLE 2
3. TABLE 3
4. TABLE 4
5. TABLE 5
6. TABLE 6
7. TABLE 2
8. TABLE 7
9. TABLE 8
10. TABLE 9
11. TABLE 10
12. TABLE 11
13. TABLE 12
9. Chapter 9
1. TABLE 1
2. TABLE 2
3. TABLE 3
4. TABLE 2
5. TABLE 1
6. TABLE 4
7. TABLE 5
8. TABLE 6
9. TABLE 7
10. TABLE 8
11. TABLE 9
12. TABLE 10
13. TABLE 11
14. TABLE 12
15. TABLE 13
16. TABLE 14
17. TABLE 15
18. TABLE 16
19. TABLE 17
20. TABLE 18
10. Chapter 10
1. TABLE 1
2. TABLE 2
3. TABLE 3
4. TABLE 4
5. TABLE 5
6. TABLE 6
7. TABLE 7
8. TABLE 8
9. TABLE 9
10. TABLE 10
11. TABLE 11
12. TABLE 12
13. TABLE 13
14. TABLE 14
15. TABLE 15
16. TABLE 16
17. TABLE 17
18. TABLE 18
19. TABLE 1
20. TABLE 2
21. TABLE 3
22. TABLE 4
23. TABLE 5
24. TABLE 6
25. TABLE 7
26. TABLE 8
27. TABLE 9
28. TABLE 10
List of Illustrations
1. Chapter 1
1. Figure 1 The Relationship between an Initial Investment, PV, and Its Future Value, FV
2. Figure 2 The Future Value of a Lump Sum, Initial Investment Not at t = 0
3. Figure 3 The Future Value of a Five-Year Ordinary Annuity
4. Figure 4 The Present Value of a Lump Sum to Be Received at Time t = 6
5. Figure 5 The Relationship between Present Value and Future Value
6. Figure 6 An Annuity Due of $100 per Period
7. Figure 7 The Present Value of an Ordinary Annuity with First Payment at Time t = 10 (in Millions)
8. Figure 8 Solving for Missing Annuity Payments (in Thousands)
9. Figure 9 The Additivity of Two Series of Cash Flows
2. Chapter 3
1. Figure 1 Histogram of S&P 500 Annual Total Returns: 1926 to 2012
2. Figure 2 Histogram of S&P 500 Monthly Total Returns: January 1926 to December 2012
3. Figure 3 Frequency Polygon of S&P 500 Monthly Total Returns: January 1926 to December 2012
4. Figure 4 Cumulative Absolute Frequency Distribution of S&P 500 Monthly Total Returns: January 1926 to December 2012
5. Figure 5 Center of Gravity Analogy for the Arithmetic Mean
6. Figure 6 Properties of a Normal Distribution (EV 5 Expected Value)
7. Figure 7 Properties of a Skewed Distribution
8. Figure 8 Leptokurtic: Fat Tailed
3. Chapter 4
1. Figure 1 Addition Rule for Probabilities
2. Figure 2 BankCorp’s Forecasted EPS
3. Figure 3 BankCorp’s Forecasted Operating Costs
4. Chapter 5
1. Figure 1 One-Period Stock Price as a Bernoulli Random Variable
2. Figure 2 A Binomial Model of Stock Price Movement
3. Figure 3 Continuous Uniform Distribution
4. Figure 4 Continuous Uniform Cumulative Distribution
5. Figure 5 Two Normal Distributions
6. Figure 6 Units of Standard Deviation
7. Figure 7 Two Lognormal Distributions
5. Chapter 6
1. Figure 1 Student’s t-Distribution versus the Standard Normal Distribution
6. Chapter 7
1. Figure 1 Rejection Points (Critical Values), 0.05 Significance Level, Two- Sided Test of the Population Mean Using a z-Test
2. Figure 2 Rejection Point (Critical Value), 0.05 Significance Level, One- Sided Test of the Population Mean Using a z-Test
7. Chapter 8
1. Figure 1 Scatter Plot of Annual Money Supply Growth Rate and Inflation Rate by Country, 1980–2012
2. Figure 2 Variables with a Correlation of 1
3. Figure 3 Variables with a Correlation of –1
4. Figure 4 Variables with a Correlation of 0
5. Figure 5 Variables with a Strong Nonlinear Association
6. Figure 6 US Inflation and Stock Returns: 1990–2013
7. Figure 7 Actual Change in Euro Area HICP versus Predicted Change
8. Figure 8 Fitted Regression Line Explaining the Inflation Rate Using Growth in the Money Supply by Country, 1980–2012
9. Figure 9 Actual Change in Euro Area HICP versus Predicted Change
10. Figure 10 Fitted Regression Line Explaining Stock Returns by Inflation
during 1990–2013
11. Figure 11 Fitted Regression Line Explaining Enterprise Value/Invested Capital Using ROIC–WACC Spread for the Food Industry
8. Chapter 9
1. Figure 1 Regression with Homoskedasticity
2. Figure 2 Regression with Heteroskedasticity
3. Figure 3 Value of the Durbin–Watson Statistic
4. Figure 4 Linear Regression When Two Variables Have a Linear Relation
5. Figure 5 Linear Regression When Two Variables Have a Nonlinear Relation
6. Figure 6 Plot of Two Series with Changing Means
9. Chapter 10
1. Figure 1 Swiss Franc/US Dollar Exchange Rate, Monthly Average of Daily Data
2. Figure 2Monthly US Retail Sales
3. Figure 3 Monthly CPI Inflation, Not Seasonally Adjusted
4. Figure 4Monthly CPI Inflation with Trend
5. Figure 5Starbucks Quarterly Sales by Fiscal Year
6. Figure 6 Starbucks Quarterly Sales with Trend
7. Figure 7 Residual from Predicting Starbucks Sales with a Trend
8. Figure 8Natural Log of Starbucks Quarterly Sales
9. Figure 9 Monthly CPI Inflation
10. Figure 10 Log of AstraZeneca’s Quarterly Sales
11. Figure 11 Log Difference, AstraZeneca’s Quarterly Sales
12. Figure 12 Monthly US Real Retail Sales and 12-Month Moving Average of Retail Sales
13. Figure 13 Monthly Europe Brent Crude Oil Price and 12-Month Moving Average of Prices
14. Figure 1 Predicted and Actual Civilian Unemployment Rates
15. Figure 2 Lightweight Vehicle Sales
16. Figure 3 Change in Civilian Unemployment Rate
17. Figure 4 Lightweight Vehicle Sales
18. Figure 5 Change in Natural Log of Lightweight Vehicle Sales
19. Figure 6 Quarterly Sales at Cisco
20. Figure 7 Quarterly Sales at Avon
FOREWORD “Central limits,” “probability distributions,” “hypothesis test”— investors have a bit of trouble generating enthusiasm for such terms. Yet, they should be enthusiastic because every investor needs these tools to analyze, compete, and succeed in today's economic environment. The financial markets and the participants in them become increasingly sophisticated every year. So, at times, it seems like you need a PhD in mathematics just to keep up with the markets.
Fortunately, a PhD is not necessary to succeed. In fact, the financial market battlefield is littered with the credentials of highly educated individuals who have failed spectacularly despite their intense education. Nonetheless, the better equipped you are with the basic tools of financial calculus, the better your chance of success.
Quantitative Investment Analysis provides the necessary utensils for success. In this volume, you will find all the statistical gadgets you need to be a confident and knowledgeable investor. Math need not be a four letter word. It can make your wealth analysis sharper, your investment theme more precise, your portfolio construction more successful.
Furthermore, this book is chock full of examples, practice problems (with answers!), charts, tables, and graphs that bring home in clear detail the concepts and tools of financial calculus. Whether you are a novice investor or an experienced practitioner, this book has something for you. In fact, as I read the book in preparation for writing this foreword, I kept getting unconsciously pulled into the examples; unwittingly, I became engaged in the book before I knew it. But that effect is part of the beauty of this book: It is an easy-to-read and easy-to-use handbook. I wanted to know more with each example I read. I know that you, too, will find that this book stimulates your curiosity while having the same ease of use that I found. Enjoy!
Mark J. P. Anson, PhD, CFA, CAIA, CPA President & Chief Investment Officer
Acadia Capital Bass Family Office
PREFACE We are pleased to bring you Quantitative Investment Analysis, Third Edition, which focuses on key tools that are needed for today's professional investor. In addition to classic time value of money, discounted cash flow applications, and probability material, the text covers advanced concepts such as correlation and regression that ultimately figure into the formation of hypotheses for purposes of testing. The text teaches critical skills that challenge many professionals, including the ability to distinguish useful information from the overwhelming quantity of available data.
The content was developed in partnership by a team of distinguished academics and practitioners, chosen for their acknowledged expertise in the field, and guided by CFA Institute. It is written specifically with the investment practitioner in mind and is replete with examples and practice problems that reinforce the learning outcomes and demonstrate real-world applicability.
The CFA Program Curriculum, from which the content of this book was drawn, is subjected to a rigorous review process to assure that it is:
Faithful to the findings of our ongoing industry practice analysis
Valuable to members, employers, and investors
Globally relevant
Generalist (as opposed to specialist) in nature
Replete with sufficient examples and practice opportunities
Pedagogically sound
The accompanying workbook is a useful reference that provides Learning Outcome Statements, which describe exactly what readers will learn and be able to demonstrate after mastering the accompanying material. Additionally, the workbook has summary overviews and practice problems for each chapter.
We hope you will find this and other books in the CFA Institute Investment Series helpful in your efforts to grow your investment knowledge, whether you are a relatively new entrant or an experienced veteran striving to keep up to date in the ever-changing market environment. CFA Institute, as a long-term committed participant in the investment profession and a not-for-profit global membership association, is pleased to provide you with this opportunity.
ACKNOWLEDGMENT We would like to thank Eugene L. Podkaminer, CFA, for his contribution to revising coverage of multifactor models. We are indebted to Professor Sanjiv Sabherwal for his painstaking work in updating self-test examples in all chapters except the chapter on multifactor models. His contribution was essential to the fresh look of this third edition.
We are indebted to Wendy L. Pirie, CFA, Gregory Siegel, CFA, and Stephen E. Wilcox, CFA, for their help in verifying the accuracy of the text. Margaret Hill, Wanda Lauziere, and Julia MacKesson and the production team at CFA Institute provided essential support through the various stages of production. We thank Robert E. Lamy, CFA, and Christopher B. Wiese, CFA, for their encouragement and oversight of the production of a third edition.
ABOUT THE CFA INSTITUTE INVESTMENT SERIES CFA Institute is pleased to provide you with the CFA Institute Investment Series, which covers major areas in the field of investments. We provide this best-in-class series for the same reason we have been chartering investment professionals for more than 45 years: to lead the investment profession globally by setting the highest standards of ethics, education, and professional excellence.
The books in the CFA Institute Investment Series contain practical, globally relevant material. They are intended both for those contemplating entry into the extremely competitive field of investment management as well as for those seeking a means of keeping their knowledge fresh and up to date. This series was designed to be user friendly and highly relevant.
We hope you find this series helpful in your efforts to grow your investment knowledge, whether you are a relatively new entrant or an experienced veteran ethically bound to keep up to date in the ever-changing market environment. As a long-term, committed participant in the investment profession and a not-for-profit global membership association, CFA Institute is pleased to provide you with this opportunity.
THE TEXTS Corporate Finance: A Practical Approach is a solid foundation for those looking to achieve lasting business growth. In today's competitive business environment, companies must find innovative ways to enable rapid and sustainable growth. This text equips readers with the foundational knowledge and tools for making smart business decisions and formulating strategies to maximize company value. It covers everything from managing relationships between stakeholders to evaluating merger and acquisition bids, as well as the companies behind them. Through extensive use of real-world examples, readers will gain critical perspective into interpreting corporate financial data, evaluating projects, and allocating funds in ways that increase corporate value. Readers will gain insights into the tools and strategies used in modern corporate financial management.
Equity Asset Valuation is a particularly cogent and important resource for anyone involved in estimating the value of securities and understanding security pricing. A well-informed professional knows that the common forms of equity valuation— dividend discount modeling, free cash flow modeling, price/earnings modeling, and residual income modeling—can all be reconciled with one another under certain assumptions. With a deep understanding of the underlying assumptions, the professional investor can better understand what other investors assume when calculating their valuation estimates. This text has a global orientation, including emerging markets.
All books in the CFA Institute Investment Series are available through all major booksellers. And, all titles are available on the Wiley Custom Select platform at http://customselect .wiley.com/ where individual chapters for all the books may be mixed and matched to create custom textbooks for the classroom.
Fixed Income Analysis has been at the forefront of new concepts in recent years, and this particular text offers some of the most recent material for the seasoned professional who is not a fixed-income specialist. The application of option and derivative technology to the once staid province of fixed income has helped contribute to an explosion of thought in this area. Professionals have been challenged to stay up to speed with credit derivatives, swaptions, collateralized mortgage securities, mortgage-backed securities, and other vehicles, and this explosion of products has strained the world's financial markets and tested central banks to provide sufficient oversight. Armed with a thorough grasp of the new exposures, the professional investor is much better able to anticipate and understand the challenges our central bankers and markets face.
International Financial Statement Analysis is designed to address the ever-
increasing need for investment professionals and students to think about financial statement analysis from a global perspective. The text is a practically oriented introduction to financial statement analysis that is distinguished by its combination of a true international orientation, a structured presentation style, and abundant illustrations and tools covering concepts as they are introduced in the text. The authors cover this discipline comprehensively and with an eye to ensuring the reader's success at all levels in the complex world of financial statement analysis.
Investments: Principles of Portfolio and Equity Analysis provides an accessible yet rigorous introduction to portfolio and equity analysis. Portfolio planning and portfolio management are presented within a context of up-to-date, global coverage of security markets, trading, and market-related concepts and products. The essentials of equity analysis and valuation are explained in detail and profusely illustrated. The book includes coverage of practitionerimportant but often neglected topics, such as industry analysis. Throughout, the focus is on the practical application of key concepts with examples drawn from both emerging and developed markets. Each chapter affords the reader many opportunities to self- check his or her understanding of topics.
One of the most prominent texts over the years in the investment management industry has been Maginn and Tuttle's Managing Investment Portfolios: A Dynamic Process. The third edition updates key concepts from the 1990 second edition. Some of the more experienced members of our community own the prior two editions and will add the third edition to their libraries. Not only does this seminal work take the concepts from the other readings and put them in a portfolio context, but it also updates the concepts of alternative investments, performance presentation standards, portfolio execution, and, very importantly, individual investor portfolio management. Focusing attention away from institutional portfolios and toward the individual investor makes this edition an important and timely work.
The New Wealth Management: The Financial Advisor's Guide to Managing and Investing Client Assets is an updated version of Harold Evensky's mainstay reference guide for wealth managers. Harold Evensky, Stephen Horan, and Thomas Robinson have updated the core text of the 1997 first edition and added an abundance of new material to fully reflect today's investment challenges. The text provides authoritative coverage across the full spectrum of wealth management and serves as a comprehensive guide for financial advisors. The book expertly blends investment theory and real-world applications and is written in the same thorough but highly accessible style as the first edition.
All books in the CFA Institute Investment Series are available through all major booksellers. And, all titles are available on the Wiley Custom Select platform at http://customselect.wiley.com/ where individual chapters for all the books may be
mixed and matched to create custom textbooks for the classroom.
CHAPTER 1 THE TIME VALUE OF MONEY Richard A. DeFusco, CFA
Dennis W. McLeavey, CFA
Jerald E. Pinto, PhD, CFA
David E. Runkle, PhD, CFA
LEARNING OUTCOMES
After completing this chapter, you will be able to do the following:
interpret interest rates as required rates of return, discount rates, or opportunity costs;
explain an interest rate as the sum of a real risk-free rate and premiums that compensate investors for bearing distinct types of risk;
calculate and interpret the effective annual rate, given the stated annual interest rate and the frequency of compounding;
solve time value of money problems for different frequencies of compounding;
calculate and interpret the future value (FV) and present value (PV) of a single sum of money, an ordinary annuity, an annuity due, a perpetuity (PV only), and a series of unequal cash flows;
demonstrate the use of a time line in modeling and solving time value of money problems.
1. Introduction As individuals, we often face decisions that involve saving money for a future use, or borrowing money for current consumption. We then need to determine the amount we need to invest, if we are saving, or the cost of borrowing, if we are shopping for a loan. As investment analysts, much of our work also involves evaluating transactions with present and future cash flows. When we place a value on any security, for example, we are attempting to determine the worth of a stream of future cash flows. To carry out all the above tasks accurately, we must understand the mathematics of time value of money problems. Money has time value in that individuals value a given amount of money more highly the earlier it is received. Therefore, a smaller amount of money now may be equivalent in value to a larger amount received at a future date. The time value of money as a topic in investment mathematics deals with equivalence relationships between cash flows with different dates. Mastery of time value of money concepts and techniques is essential for investment analysts.
The reading1 is organized as follows: Section 2 introduces some terminology used throughout the reading and supplies some economic intuition for the variables we will discuss. Section 3 tackles the problem of determining the worth at a future point in time of an amount invested today. Section 4 addresses the future worth of a series of cash flows. These two sections provide the tools for calculating the equivalent value at a future date of a single cash flow or series of cash flows. Sections 5 and 6 discuss the equivalent value today of a single future cash flow and a series of future cash flows, respectively. In Section 7, we explore how to determine other quantities of interest in time value of money problems.
2. Interest Rates: Interpretation In this reading, we will continually refer to interest rates. In some cases, we assume a particular value for the interest rate; in other cases, the interest rate will be the unknown quantity we seek to determine. Before turning to the mechanics of time value of money problems, we must illustrate the underlying economic concepts. In this section, we briefly explain the meaning and interpretation of interest rates.
Time value of money concerns equivalence relationships between cash flows occurring on different dates. The idea of equivalence relationships is relatively simple. Consider the following exchange: You pay $10,000 today and in return receive $9,500 today. Would you accept this arrangement? Not likely. But what if you received the $9,500 today and paid the $10,000 one year from now? Can these amounts be considered equivalent? Possibly, because a payment of $10,000 a year from now would probably be worth less to you than a payment of $10,000 today. It would be fair, therefore, to discount the $10,000 received in one year; that is, to cut its value based on how much time passes before the money is paid. An interest rate, denoted r, is a rate of return that reflects the relationship between differently dated cash flows. If $9,500 today and $10,000 in one year are equivalent in value, then $10,000 − $9,500 = $500 is the required compensation for receiving $10,000 in one year rather than now. The interest rate—the required compensation stated as a rate of return—is $500/$9,500 = 0.0526 or 5.26 percent.
Interest rates can be thought of in three ways. First, they can be considered required rates of return—that is, the minimum rate of return an investor must receive in order to accept the investment. Second, interest rates can be considered discount rates. In the example above, 5.26 percent is that rate at which we discounted the $10,000 future amount to find its value today. Thus, we use the terms “interest rate” and “discount rate” almost interchangeably. Third, interest rates can be considered opportunity costs. An opportunity cost is the value that investors forgo by choosing a particular course of action. In the example, if the party who supplied $9,500 had instead decided to spend it today, he would have forgone earning 5.26 percent on the money. So we can view 5.26 percent as the opportunity cost of current consumption.
Economics tells us that interest rates are set in the marketplace by the forces of supply and demand, where investors are suppliers of funds and borrowers are demanders of funds. Taking the perspective of investors in analyzing market- determined interest rates, we can view an interest rate r as being composed of a real risk-free interest rate plus a set of four premiums that are required returns or compensation for bearing distinct types of risk:
The real risk-free interest rate is the single-period interest rate for a completely risk-free security if no inflation were expected. In economic theory, the real risk-free rate reflects the time preferences of individuals for current versus future real consumption.
The inflation premium compensates investors for expected inflation and reflects the average inflation rate expected over the maturity of the debt. Inflation reduces the purchasing power of a unit of currency—the amount of goods and services one can buy with it. The sum of the real risk-free interest rate and the inflation premium is the nominal risk-free interest rate.2 Many countries have governmental short-term debt whose interest rate can be considered to represent the nominal risk-free interest rate in that country. The interest rate on a 90-day US Treasury bill (T-bill), for example, represents the nominal risk-free interest rate over that time horizon.3 US T-bills can be bought and sold in large quantities with minimal transaction costs and are backed by the full faith and credit of the US government.
The default risk premium compensates investors for the possibility that the borrower will fail to make a promised payment at the contracted time and in the contracted amount.
The liquidity premium compensates investors for the risk of loss relative to an investment’s fair value if the investment needs to be converted to cash quickly. US T-bills, for example, do not bear a liquidity premium because large amounts can be bought and sold without affecting their market price. Many bonds of small issuers, by contrast, trade infrequently after they are issued; the interest rate on such bonds includes a liquidity premium reflecting the relatively high costs (including the impact on price) of selling a position.
The maturity premium compensates investors for the increased sensitivity of the market value of debt to a change in market interest rates as maturity is extended, in general (holding all else equal). The difference between the interest rate on longer-maturity, liquid Treasury debt and that on short-term Treasury debt reflects a positive maturity premium for the longer-term debt (and possibly different inflation premiums as well).
Using this insight into the economic meaning of interest rates, we now turn to a discussion of solving time value of money problems, starting with the future value of a single cash flow.
3. The Future Value of a Single Cash Flow In this section, we introduce time value associated with a single cash flow or lump- sum investment. We describe the relationship between an initial investment or present value (PV), which earns a rate of return (the interest rate per period) denoted as r, and its future value (FV), which will be received N years or periods from today.
The following example illustrates this concept. Suppose you invest $100 (PV = $100) in an interest-bearing bank account paying 5 percent annually. At the end of the first year, you will have the $100 plus the interest earned, 0.05 × $100 = $5, for a total of $105. To formalize this one-period example, we define the following terms:
For N = 1, the expression for the future value of amount PV is
(1)
For this example, we calculate the future value one year from today as FV1 = $100(1.05) = $105.
Now suppose you decide to invest the initial $100 for two years with interest earned and credited to your account annually (annual compounding). At the end of the first year (the beginning of the second year), your account will have $105, which you will leave in the bank for another year. Thus, with a beginning amount of $105 (PV = $105), the amount at the end of the second year will be $105(1.05) = $110.25. Note that the $5.25 interest earned during the second year is 5 percent of the amount invested at the beginning of Year 2.
Another way to understand this example is to note that the amount invested at the beginning of Year 2 is composed of the original $100 that you invested plus the $5 interest earned during the first year. During the second year, the original principal again earns interest, as does the interest that was earned during Year 1. You can see how the original investment grows:
Original investment $100.00 Interest for the first year ($100 × 0.05) 5.00
Interest for the second year based on original investment ($100 × 0.05) 5.00
Interest for the second year based on interest earned in the first year (0.05 × $5.00 interest on interest)
0.25
Total $110.25
The $5 interest that you earned each period on the $100 original investment is known as simple interest (the interest rate times the principal). Principal is the amount of funds originally invested. During the two-year period, you earn $10 of simple interest. The extra $0.25 that you have at the end of Year 2 is the interest you earned on the Year 1 interest of $5 that you reinvested.
The interest earned on interest provides the first glimpse of the phenomenon known as compounding. Although the interest earned on the initial investment is important, for a given interest rate it is fixed in size from period to period. The compounded interest earned on reinvested interest is a far more powerful force because, for a given interest rate, it grows in size each period. The importance of compounding increases with the magnitude of the interest rate. For example, $100 invested today would be worth about $13,150 after 100 years if compounded annually at 5 percent, but worth more than $20 million if compounded annually over the same time period at a rate of 13 percent.
To verify the $20 million figure, we need a general formula to handle compounding for any number of periods. The following general formula relates the present value of an initial investment to its future value after N periods:
(2)
where r is the stated interest rate per period and N is the number of compounding periods. In the bank example, FV2 = $100(1 + 0.05)2 = $110.25. In the 13 percent investment example, FV100 = $100(1.13)100 = $20,316,287.42.
The most important point to remember about using the future value equation is that the stated interest rate, r, and the number of compounding periods, N, must be compatible. Both variables must be defined in the same time units. For example, if N is stated in months, then r should be the one-month interest rate, unannualized.
A time line helps us to keep track of the compatibility of time units and the interest rate per time period. In the time line, we use the time index t to represent a point in time a stated number of periods from today. Thus the present value is the amount available for investment today, indexed as t = 0. We can now refer to a time N periods from today as t = N. The time line in Figure 1 shows this relationship.
In Figure 1, we have positioned the initial investment, PV, at t = 0. Using Equation 2, we move the present value, PV, forward to t = N by the factor (1 + r)N. This factor is
called a future value factor. We denote the future value on the time line as FV and position it at t = N. Suppose the future value is to be received exactly 10 periods from today’s date (N = 10). The present value, PV, and the future value, FV, are separated in time through the factor (1 + r)10.
The fact that the present value and the future value are separated in time has important consequences:
We can add amounts of money only if they are indexed at the same point in time.
For a given interest rate, the future value increases with the number of periods.
For a given number of periods, the future value increases with the interest rate.
Figure 1 The Relationship between an Initial Investment, PV, and Its Future Value, FV
To better understand these concepts, consider three examples that illustrate how to apply the future value formula.
EXAMPLE 1 The Future Value of a Lump Sum with Interim Cash Reinvested at the Same Rate
You are the lucky winner of your state’s lottery of $5 million after taxes. You invest your winnings in a five-year certificate of deposit (CD) at a local financial institution. The CD promises to pay 7 percent per year compounded annually. This institution also lets you reinvest the interest at that rate for the duration of the CD. How much will you have at the end of five years if your money remains invested at 7 percent for five years with no withdrawals?
Solution: To solve this problem, compute the future value of the $5 million investment using the following values in Equation 2:
At the end of five years, you will have $7,012,758.65 if your money remains invested at 7 percent with no withdrawals.
In this and most examples in this reading, note that the factors are reported at six decimal places but the calculations may actually reflect greater precision. For example, the reported 1.402552 has been rounded up from 1.40255173 (the calculation is actually carried out with more than eight decimal places of precision by the calculator or spreadsheet). Our final result reflects the higher number of decimal places carried by the calculator or spreadsheet.4
EXAMPLE 2 The Future Value of a Lump Sum with No Interim Cash
An institution offers you the following terms for a contract: For an investment of ¥2,500,000, the institution promises to pay you a lump sum six years from now at an 8 percent annual interest rate. What future amount can you expect?
Solution: Use the following data in Equation 2 to find the future value:
You can expect to receive ¥3,967,186 six years from now.
Our third example is a more complicated future value problem that illustrates the importance of keeping track of actual calendar time.
EXAMPLE 3 The Future Value of a Lump Sum
A pension fund manager estimates that his corporate sponsor will make a $10 million contribution five years from now. The rate of return on plan assets has been estimated at 9 percent per year. The pension fund manager wants to calculate the future value of this contribution 15 years from now, which is the date at which the funds will be distributed to retirees. What is that future value?
Solution: By positioning the initial investment, PV, at t = 5, we can calculate the future value of the contribution using the following data in Equation 2:
Figure 2 The Future Value of a Lump Sum, Initial Investment Not at t = 0
This problem looks much like the previous two, but it differs in one important respect: its timing. From the standpoint of today (t = 0), the future amount of $23,673,636.75 is 15 years into the future. Although the future value is 10 years from its present value, the present value of $10 million will not be received for another five years.
As Figure 2 shows, we have followed the convention of indexing today as t = 0 and indexing subsequent times by adding 1 for each period. The additional contribution of $10 million is to be received in five years, so it is indexed as t = 5 and appears as such in the figure. The future value of the investment in 10 years is then indexed at t = 15; that is, 10 years following the receipt of the $10 million contribution at t = 5. Time lines like this one can be extremely useful when dealing with more complicated problems, especially those involving
more than one cash flow.
In a later section of this reading, we will discuss how to calculate the value today of the $10 million to be received five years from now. For the moment, we can use Equation 2. Suppose the pension fund manager in Example 3 above were to receive $6,499,313.86 today from the corporate sponsor. How much will that sum be worth at the end of five years? How much will it be worth at the end of 15 years?
and
These results show that today’s present value of about $6.5 million becomes $10 million after five years and $23.67 million after 15 years.
3.1. The Frequency of Compounding In this section, we examine investments paying interest more than once a year. For instance, many banks offer a monthly interest rate that compounds 12 times a year. In such an arrangement, they pay interest on interest every month. Rather than quote the periodic monthly interest rate, financial institutions often quote an annual interest rate that we refer to as the stated annual interest rate or quoted interest rate. We denote the stated annual interest rate by rs. For instance, your bank might state that a particular CD pays 8 percent compounded monthly. The stated annual interest rate equals the monthly interest rate multiplied by 12. In this example, the monthly interest rate is 0.08/12 = 0.0067 or 0.67 percent.5 This rate is strictly a quoting convention because (1 + 0.0067)12 = 1.083, not 1.08; the term (1 + rs) is not
meant to be a future value factor when compounding is more frequent than annual.
With more than one compounding period per year, the future value formula can be expressed as
(3)
where rs is the stated annual interest rate, m is the number of compounding periods per year, and N now stands for the number of years. Note the compatibility here between the interest rate used, rs/m, and the number of compounding periods, mN. The periodic rate, rs/m, is the stated annual interest rate divided by the number of compounding periods per year. The number of compounding periods, mN, is the number of compounding periods in one year multiplied by the number of years. The periodic rate, rs/m, and the number of compounding periods, mN, must be compatible.
EXAMPLE 4 The Future Value of a Lump Sum with Quarterly Compounding
Continuing with the CD example, suppose your bank offers you a CD with a two-year maturity, a stated annual interest rate of 8 percent compounded quarterly, and a feature allowing reinvestment of the interest at the same interest rate. You decide to invest $10,000. What will the CD be worth at maturity?
Solution: Compute the future value with Equation 3 as follows:
At maturity, the CD will be worth $11,716.59.
The future value formula in Equation 3 does not differ from the one in Equation 2. Simply keep in mind that the interest rate to use is the rate per period and the exponent is the number of interest, or compounding, periods.
EXAMPLE 5 The Future Value of a Lump Sum with Monthly Compounding
An Australian bank offers to pay you 6 percent compounded monthly. You decide to invest A$1 million for one year. What is the future value of your investment if interest payments are reinvested at 6 percent?
Solution: Use Equation 3 to find the future value of the one-year investment as follows:
If you had been paid 6 percent with annual compounding, the future amount would be only A$1,000,000(1.06) = A$1,060,000 instead of A$1,061,677.81 with monthly compounding.
3.2. Continuous Compounding The preceding discussion on compounding periods illustrates discrete compounding, which credits interest after a discrete amount of time has elapsed. If the number of compounding periods per year becomes infinite, then interest is said to compound continuously. If we want to use the future value formula with continuous compounding, we need to find the limiting value of the future value factor for m → ∞ (infinitely many compounding periods per year) in Equation 3. The expression for the future value of a sum in N years with continuous compounding is
(4)
The term is the transcendental number e ≈ 2.7182818 raised to the power rsN. Most financial calculators have the function ex.
EXAMPLE 6 The Future Value of a Lump Sum with Continuous Compounding
Suppose a $10,000 investment will earn 8 percent compounded continuously for two years. We can compute the future value with Equation 4 as follows:
With the same interest rate but using continuous compounding, the $10,000 investment will grow to $11,735.11 in two years, compared with $11,716.59 using quarterly compounding as shown in Example 4.
Table 1 shows how a stated annual interest rate of 8 percent generates different ending dollar amounts with annual, semiannual, quarterly, monthly, daily, and continuous compounding for an initial investment of $1 (carried out to six decimal places).
TABLE 1 The Effect of Compounding Frequency on Future Value
Frequency rs/m mN Future Value of $1 Annual 8%/1 = 8% 1 × 1 = 1 $1.00(1.08) = $1.08
Semiannual 8%/2 = 4% 2 × 1 = 2 $1.00(1.04)2 = $1.081600 Quarterly 8%/4 = 2% 4 × 1 = 4 $1.00(1.02)4 = $1.082432 Monthly 8%/12 = 0.6667% 12 × 1 = 12 $1.00(1.006667)12 = $1.083000 Daily 8%/365 = 0.0219% 365 × 1 = 365 $1.00(1.000219)365 = $1.083278
Continuous $1.00e0.08(1) = $1.083287
As Table 1 shows, all six cases have the same stated annual interest rate of 8 percent; they have different ending dollar amounts, however, because of differences in the frequency of compounding. With annual compounding, the ending amount is $1.08. More frequent compounding results in larger ending amounts. The ending dollar amount with continuous compounding is the maximum amount that can be earned with a stated annual rate of 8 percent.
Table 1 also shows that a $1 investment earning 8.16 percent compounded annually grows to the same future value at the end of one year as a $1 investment earning 8 percent compounded semiannually. This result leads us to a distinction between the stated annual interest rate and the effective annual rate (EAR).6 For an 8 percent stated annual interest rate with semiannual compounding, the EAR is 8.16 percent.
3.3. Stated and Effective Rates The stated annual interest rate does not give a future value directly, so we need a formula for the EAR. With an annual interest rate of 8 percent compounded semiannually, we receive a periodic rate of 4 percent. During the course of a year, an investment of $1 would grow to $1(1.04)2 = $1.0816, as illustrated in Table 1. The interest earned on the $1 investment is $0.0816 and represents an effective annual rate of interest of 8.16 percent. The effective annual rate is calculated as follows:
(5)
The periodic interest rate is the stated annual interest rate divided by m, where m is the number of compounding periods in one year. Using our previous example, we can solve for EAR as follows: (1.04)2 − 1 = 8.16 percent.
The concept of EAR extends to continuous compounding. Suppose we have a rate of 8 percent compounded continuously. We can find the EAR in the same way as above by finding the appropriate future value factor. In this case, a $1 investment would grow to $1e0.08(1.0) = $1.0833. The interest earned for one year represents an effective annual rate of 8.33 percent and is larger than the 8.16 percent EAR with semiannual compounding because interest is compounded more frequently. With continuous compounding, we can solve for the effective annual rate as follows:
(6)
We can reverse the formulas for EAR with discrete and continuous compounding to find a periodic rate that corresponds to a particular effective annual rate. Suppose we want to find the appropriate periodic rate for a given effective annual rate of
8.16 percent with semiannual compounding. We can use Equation 5 to find the periodic rate:
To calculate the continuously compounded rate (the stated annual interest rate with continuous compounding) corresponding to an effective annual rate of 8.33 percent, we find the interest rate that satisfies Equation 6:
To solve this equation, we take the natural logarithm of both sides. (Recall that the natural log of is ln .) Therefore, ln 1.0833 = rs, resulting in rs = 8 percent. We see that a stated annual rate of 8 percent with continuous compounding is equivalent to an EAR of 8.33 percent.
4. The Future Value of a Series of Cash Flows In this section, we consider series of cash flows, both even and uneven. We begin with a list of terms commonly used when valuing cash flows that are distributed over many time periods.
An annuity is a finite set of level sequential cash flows.
An ordinary annuity has a first cash flow that occurs one period from now (indexed at t = 1).
An annuity due has a first cash flow that occurs immediately (indexed at t = 0).
A perpetuity is a perpetual annuity, or a set of level never-ending sequential cash flows, with the first cash flow occurring one period from now.
4.1. Equal Cash Flows—Ordinary Annuity Consider an ordinary annuity paying 5 percent annually. Suppose we have five separate deposits of $1,000 occurring at equally spaced intervals of one year, with the first payment occurring at t = 1. Our goal is to find the future value of this ordinary annuity after the last deposit at t = 5. The increment in the time counter is one year, so the last payment occurs five years from now. As the time line in Figure 3 shows, we find the future value of each $1,000 deposit as of t = 5 with Equation 2, FVN = PV(1 + r)N. The arrows in Figure 3 extend from the payment date to t = 5. For instance, the first $1,000 deposit made at t = 1 will compound over four periods. Using Equation 2, we find that the future value of the first deposit at t = 5 is $1,000(1.05)4 = $1,215.51. We calculate the future value of all other payments in a similar fashion. (Note that we are finding the future value at t = 5, so the last payment does not earn any interest.) With all values now at t = 5, we can add the future values to arrive at the future value of the annuity. This amount is $5,525.63.
We can arrive at a general annuity formula if we define the annuity amount as A, the number of time periods as N, and the interest rate per period as r. We can then define the future value as
which simplifies to
(7)
The term in brackets is the future value annuity factor. This factor gives the future value of an ordinary annuity of $1 per period. Multiplying the future value annuity factor by the annuity amount gives the future value of an ordinary annuity. For the ordinary annuity in Figure 3, we find the future value annuity factor from Equation 7 as
With an annuity amount A = $1,000, the future value of the annuity is $1,000(5.525631) = $5,525.63, an amount that agrees with our earlier work.
Figure 3 The Future Value of a Five-Year Ordinary Annuity
The next example illustrates how to find the future value of an ordinary annuity using the formula in Equation 7.
EXAMPLE 7 The Future Value of an Annuity
Suppose your company’s defined contribution retirement plan allows you to invest up to €20,000 per year. You plan to invest €20,000 per year in a stock index fund for the next 30 years. Historically, this fund has earned 9 percent per year on average. Assuming that you actually earn 9 percent a year, how much money will you have available for retirement after making the last payment?
Solution: Use Equation 7 to find the future amount:
Assuming the fund continues to earn an average of 9 percent per year, you will have €2,726,150.77 available at retirement.
4.2. Unequal Cash Flows In many cases, cash flow streams are unequal, precluding the simple use of the future value annuity factor. For instance, an individual investor might have a savings plan that involves unequal cash payments depending on the month of the year or lower savings during a planned vacation. One can always find the future value of a series of unequal cash flows by compounding the cash flows one at a time. Suppose you have the five cash flows described in Table 2, indexed relative to the present (t = 0).
TABLE 2 A Series of Unequal Cash Flows and Their Future Values at 5 Percent
Time Cash Flow ($) Future Value at Year 5 t = 1 1,000 $1,000(1.05)4 = $1,215.51 t = 2 2,000 $2,000(1.05)3 = 2,315.25 t = 3 4,000 $4,000(1.05)2 = $4,410.00 t = 4 5,000 $5,000(1.05)1 = $5,250.00 t = 5 6,000 $6,000(1.05)0 = $6,000.00
Sum = $19,190.76
All of the payments shown in Table 2 are different. Therefore, the most direct approach to finding the future value at t = 5 is to compute the future value of each payment as of t = 5 and then sum the individual future values. The total future value at Year 5 equals $19,190.76, as shown in the third column. Later in this reading, you will learn shortcuts to take when the cash flows are close to even; these shortcuts will allow you to combine annuity and single-period calculations.
5. The Present Value of a Single Cash Flow
5.1. Finding the Present Value of a Single Cash Flow Just as the future value factor links today’s present value with tomorrow’s future value, the present value factor allows us to discount future value to present value. For example, with a 5 percent interest rate generating a future payoff of $105 in one year, what current amount invested at 5 percent for one year will grow to $105? The answer is $100; therefore, $100 is the present value of $105 to be received in one year at a discount rate of 5 percent.
Given a future cash flow that is to be received in N periods and an interest rate per period of r, we can use the formula for future value to solve directly for the present value as follows:
(8)
We see from Equation 8 that the present value factor, (1 + r)−N, is the reciprocal of the future value factor, (1 + r)N.
EXAMPLE 8 The Present Value of a Lump Sum
An insurance company has issued a Guaranteed Investment Contract (GIC) that promises to pay $100,000 in six years with an 8 percent return rate. What amount of money must the insurer invest today at 8 percent for six years to make the promised payment?
Solution: We can use Equation 8 to find the present value using the following data:
Figure 4 The Present Value of a Lump Sum to Be Received at Time t = 6
We can say that $63,016.96 today, with an interest rate of 8 percent, is equivalent to $100,000 to be received in six years. Discounting the $100,000 makes a future $100,000 equivalent to $63,016.96 when allowance is made for the time value of money. As the time line in Figure 4 shows, the $100,000 has been discounted six full periods.
EXAMPLE 9 The Projected Present Value of a More Distant Future Lump Sum
Suppose you own a liquid financial asset that will pay you $100,000 in 10 years from today. Your daughter plans to attend college four years from today, and you want to know what the asset’s present value will be at that time. Given an 8 percent discount rate, what will the asset be worth four years from today?
Solution: The value of the asset is the present value of the asset’s promised payment. At t = 4, the cash payment will be received six years later. With this information, you can solve for the value four years from today using Equation 8:
FIGURE 5 The Relationship between Present Value and Future Value
The time line in Figure 5 shows the future payment of $100,000 that is to be received at t = 10. The time line also shows the values at t = 4 and at t = 0. Relative to the payment at t = 10, the amount at t = 4 is a projected present value, while the amount at t = 0 is the present value (as of today).
Present value problems require an evaluation of the present value factor, (1 + r)−N. Present values relate to the discount rate and the number of periods in the following ways:
For a given discount rate, the further in the future the amount to be received, the smaller that amount’s present value.
Holding time constant, the larger the discount rate, the smaller the present value of a future amount.
5.2. The Frequency of Compounding Recall that interest may be paid semiannually, quarterly, monthly, or even daily. To handle interest payments made more than once a year, we can modify the present value formula (Equation 8) as follows. Recall that rs is the quoted interest rate and equals the periodic interest rate multiplied by the number of compounding periods in each year. In general, with more than one compounding period in a year, we can express the formula for present value as
(9)
where
m = number of compounding periods per year
rs = quoted annual interest rate
N = number of years
The formula in Equation 9 is quite similar to that in Equation 8. As we have already noted, present value and future value factors are reciprocals. Changing the frequency of compounding does not alter this result. The only difference is the use of the periodic interest rate and the corresponding number of compounding periods.
The following example illustrates Equation 9.
EXAMPLE 10 The Present Value of a Lump Sum with Monthly Compounding
The manager of a Canadian pension fund knows that the fund must make a lump-sum payment of C$5 million 10 years from now. She wants to invest an amount today in a GIC so that it will grow to the required amount. The current interest rate on GICs is 6 percent a year, compounded monthly. How much should she invest today in the GIC?
Solution: Use Equation 9 to find the required present value:
In applying Equation 9, we use the periodic rate (in this case, the monthly rate) and the appropriate number of periods with monthly compounding (in this case, 10 years of monthly compounding, or 120 periods).
6. The Present Value of a Series of Cash Flows Many applications in investment management involve assets that offer a series of cash flows over time. The cash flows may be highly uneven, relatively even, or equal. They may occur over relatively short periods of time, longer periods of time, or even stretch on indefinitely. In this section, we discuss how to find the present value of a series of cash flows.
6.1. The Present Value of a Series of Equal Cash Flows We begin with an ordinary annuity. Recall that an ordinary annuity has equal annuity payments, with the first payment starting one period into the future. In total, the annuity makes N payments, with the first payment at t = 1 and the last at t = N. We can express the present value of an ordinary annuity as the sum of the present values of each individual annuity payment, as follows:
(10)
where
A = the annuity amount
r = the interest rate per period corresponding to the frequency of annuity payments (for example, annual, quarterly, or monthly)
N = the number of annuity payments
Because the annuity payment (A) is a constant in this equation, it can be factored out as a common term. Thus the sum of the interest factors has a shortcut expression:
(11)
In much the same way that we computed the future value of an ordinary annuity, we find the present value by multiplying the annuity amount by a present value annuity factor (the term in brackets in Equation 11).
EXAMPLE 11 The Present Value of an Ordinary Annuity
Suppose you are considering purchasing a financial asset that promises to pay €1,000 per year for five years, with the first payment one year from now. The required rate of return is 12 percent per year. How much should you pay for this asset?
Solution: To find the value of the financial asset, use the formula for the present value of an ordinary annuity given in Equation 11 with the following data:
The series of cash flows of €1,000 per year for five years is currently worth €3,604.78 when discounted at 12 percent.
Figure 6 An Annuity Due of $100 per Period
Keeping track of the actual calendar time brings us to a specific type of annuity with level payments: the annuity due. An annuity due has its first payment occurring today (t = 0). In total, the annuity due will make N payments. Figure 6 presents the time line for an annuity due that makes four payments of $100.
As Figure 6 shows, we can view the four-period annuity due as the sum of two parts: a $100 lump sum today and an ordinary annuity of $100 per period for three periods. At a 12 percent discount rate, the four $100 cash flows in this annuity due example will be worth $340.18.7
Expressing the value of the future series of cash flows in today’s dollars gives us a convenient way of comparing annuities. The next example illustrates this approach.
EXAMPLE 12 An Annuity Due as the Present Value of an Immediate Cash Flow Plus an Ordinary Annuity
You are retiring today and must choose to take your retirement benefits either as a lump sum or as an annuity. Your company’s benefits officer presents you with two alternatives: an immediate lump sum of $2 million or an annuity with 20 payments of $200,000 a year with the first payment starting today. The interest rate at your bank is 7 percent per year compounded annually. Which option has the greater present value? (Ignore any tax differences between the two options.)
Solution: To compare the two options, find the present value of each at time t = 0 and choose the one with the larger value. The first option’s present value is $2 million, already expressed in today’s dollars. The second option is an annuity due. Because the first payment occurs at t = 0, you can separate the annuity benefits into two pieces: an immediate $200,000 to be paid today (t = 0) and an ordinary annuity of $200,000 per year for 19 years. To value this option, you need to find the present value of the ordinary annuity using Equation 11 and then add $200,000 to it.
The 19 payments of $200,000 have a present value of $2,067,119.05. Adding the initial payment of $200,000 to $2,067,119.05, we find that the total value of the annuity option is $2,267,119.05. The present value of the annuity is greater than the lump sum alternative of $2 million.
We now look at another example reiterating the equivalence of present and future values.
EXAMPLE 13 The Projected Present Value of an Ordinary Annuity
A German pension fund manager anticipates that benefits of €1 million per year must be paid to retirees. Retirements will not occur until 10 years from now at time t = 10. Once benefits begin to be paid, they will extend until t = 39 for a total of 30 payments. What is the present value of the pension liability if the appropriate annual discount rate for plan liabilities is 5 percent compounded annually?
Solution: This problem involves an annuity with the first payment at t = 10. From the perspective of t = 9, we have an ordinary annuity with 30 payments. We can compute the present value of this annuity with Equation 11 and then look at it on a time line.
On the time line, we have shown the pension payments of €1 million extending from t = 10 to t = 39. The bracket and arrow indicate the process of finding the present value of the annuity, discounted back to t = 9. The present value of the pension benefits as of t = 9 is €15,372,451.03. The problem is to find the present value today (at t = 0).
Now we can rely on the equivalence of present value and future value. As Figure 7 shows, we can view the amount at t = 9 as a future value from the
vantage point of t = 0. We compute the present value of the amount at t = 9 as follows:
The present value of the pension liability is €9,909,219.00.
Figure 7 The Present Value of an Ordinary Annuity with First Payment at Time t = 10 (in Millions)
Example 13 illustrates three procedures emphasized in this reading:
finding the present or future value of any cash flow series;
recognizing the equivalence of present value and appropriately discounted future value; and
keeping track of the actual calendar time in a problem involving the time value of money.
6.2. The Present Value of an Infinite Series of Equal Cash Flows— Perpetuity Consider the case of an ordinary annuity that extends indefinitely. Such an ordinary annuity is called a perpetuity (a perpetual annuity). To derive a formula for the present value of a perpetuity, we can modify Equation 10 to account for an infinite series of cash flows:
(12)
As long as interest rates are positive, the sum of present value factors converges and
(13)
To see this, look back at Equation 11, the expression for the present value of an ordinary annuity. As N (the number of periods in the annuity) goes to infinity, the term 1/(1 + r)N approaches 0 and Equation 11 simplifies to Equation 13. This equation will reappear when we value dividends from stocks because stocks have no predefined life span. (A stock paying constant dividends is similar to a perpetuity.) With the first payment a year from now, a perpetuity of $10 per year with a 20 percent required rate of return has a present value of $10/0.2 = $50.
Equation 13 is valid only for a perpetuity with level payments. In our development above, the first payment occurred at t = 1; therefore, we compute the present value as of t = 0.
Other assets also come close to satisfying the assumptions of a perpetuity. Certain government bonds and preferred stocks are typical examples of financial assets that make level payments for an indefinite period of time.
EXAMPLE 14 The Present Value of a Perpetuity
The British government once issued a type of security called a consol bond, which promised to pay a level cash flow indefinitely. If a consol bond paid £100 per year in perpetuity, what would it be worth today if the required rate of return were 5 percent?
Solution: To answer this question, we can use Equation 13 with the following data:
The bond would be worth £2,000.
6.3. Present Values Indexed at Times Other than t = 0 In practice with investments, analysts frequently need to find present values indexed at times other than t = 0. Subscripting the present value and evaluating a perpetuity beginning with $100 payments in Year 2, we find PV1 = $100/0.05 = $2,000 at a 5
percent discount rate. Further, we can calculate today’s PV as PV0 = $2,000/1.05 = $1,904.76.
Consider a similar situation in which cash flows of $6 per year begin at the end of the fourth year and continue at the end of each year thereafter, with the last cash flow at the end of the 10th year. From the perspective of the end of the third year, we are facing a typical seven-year ordinary annuity. We can find the present value of the annuity from the perspective of the end of the third year and then discount that present value back to the present. At an interest rate of 5 percent, the cash flows of $6 per year starting at the end of the fourth year will be worth $34.72 at the end of the third year (t = 3) and $29.99 today (t = 0).
The next example illustrates the important concept that an annuity or perpetuity beginning sometime in the future can be expressed in present value terms one period prior to the first payment. That present value can then be discounted back to today’s present value.
EXAMPLE 15 The Present Value of a Projected Perpetuity
Consider a level perpetuity of £100 per year with its first payment beginning at t = 5. What is its present value today (at t = 0), given a 5 percent discount rate?
Solution: First, we find the present value of the perpetuity at t = 4 and then discount that amount back to t = 0. (Recall that a perpetuity or an ordinary annuity has its first payment one period away, explaining the t = 4 index for our present value calculation.)
i. Find the present value of the perpetuity at t = 4:
ii. Find the present value of the future amount at t = 4. From the perspective of t = 0, the present value of £2,000 can be considered a future value. Now we need to find the present value of a lump sum:
Today’s present value of the perpetuity is £1,645.40.
As discussed earlier, an annuity is a series of payments of a fixed amount for a specified number of periods. Suppose we own a perpetuity. At the same time, we issue a perpetuity obligating us to make payments; these payments are the same size as those of the perpetuity we own. However, the first payment of the perpetuity we issue is at t = 5; payments then continue on forever. The payments on this second perpetuity exactly offset the payments received from the perpetuity we own at t = 5 and all subsequent dates. We are left with level nonzero net cash flows at t = 1, 2, 3, and 4. This outcome exactly fits the definition of an annuity with four payments. Thus we can construct an annuity as the difference between two perpetuities with equal, level payments but differing starting dates. The next example illustrates this result.
EXAMPLE 16 The Present Value of an Ordinary Annuity as the Present Value of a Current Minus Projected Perpetuity
Given a 5 percent discount rate, find the present value of a four-year ordinary annuity of £100 per year starting in Year 1 as the difference between the following two level perpetuities:
Solution: If we subtract Perpetuity 2 from Perpetuity 1, we are left with an ordinary annuity of £100 per period for four years (payments at t = 1, 2, 3, 4). Subtracting the present value of Perpetuity 2 from that of Perpetuity 1, we arrive at the present value of the four-year ordinary annuity:
The four-year ordinary annuity’s present value is equal to £2,000 – £1,645.40 = £354.60.
6.4. The Present Value of a Series of Unequal Cash Flows When we have unequal cash flows, we must first find the present value of each individual cash flow and then sum the respective present values. For a series with many cash flows, we usually use a spreadsheet. Table 3 lists a series of cash flows with the time periods in the first column, cash flows in the second column, and each cash flow’s present value in the third column. The last row of Table 3 shows the sum of the five present values.
TABLE 3 A Series of Unequal Cash Flows and Their Present Values at 5 Percent
Time Period Cash Flow ($) Present Value at Year 0 1 1,000 $1,000(1.05)−1 = $952.38 2 2,000 $2,000(1.05)−2 = $1,814.06 3 4,000 $4,000(1.05)−3 = $3,455.35 4 5,000 $5,000(1.05)−4 = $4,113.51 5 6,000 $6,000(1.05)−5 = $4,701.16
Sum = $15,036.46
We could calculate the future value of these cash flows by computing them one at a time using the single-payment future value formula. We already know the present value of this series, however, so we can easily apply time-value equivalence. The future value of the series of cash flows from Table 2, $19,190.76, is equal to the single $15,036.46 amount compounded forward to t = 5:
7. Solving for Rates, Number of Periods, or Size of Annuity Payments In the previous examples, certain pieces of information have been made available. For instance, all problems have given the rate of interest, r, the number of time periods, N, the annuity amount, A, and either the present value, PV, or future value, FV. In real-world applications, however, although the present and future values may be given, you may have to solve for either the interest rate, the number of periods, or the annuity amount. In the subsections that follow, we show these types of problems.
7.1. Solving for Interest Rates and Growth Rates Suppose a bank deposit of €100 is known to generate a payoff of €111 in one year. With this information, we can infer the interest rate that separates the present value of €100 from the future value of €111 by using Equation 2, FVN = PV(1 + r)N, with N = 1. With PV, FV, and N known, we can solve for r directly:
The interest rate that equates €100 at t = 0 to €111 at t = 1 is 11 percent. Thus we can state that €100 grows to €111 with a growth rate of 11 percent.
As this example shows, an interest rate can also be considered a growth rate. The particular application will usually dictate whether we use the term “interest rate” or “growth rate.” Solving Equation 2 for r and replacing the interest rate r with the growth rate g produces the following expression for determining growth rates:
(14)
Below are two examples that use the concept of a growth rate.
EXAMPLE 17 Calculating a Growth Rate (1)
Hyundai Steel, the first Korean steelmaker, was established in 1953. Hyundai Steel’s sales increased from 10,503.0 billion in 2008 to 14,146.4 billion in 2012. However, its net profit declined from 822.5 billion in 2008 to 796.4 billion in 2012. Calculate the following growth rates for Hyundai Steel for the four-year period from the end of 2008 to the end of 2012:
1. Sales growth rate.
2. Net profit growth rate.
Solution to 1: To solve this problem, we can use Equation 14, g = (FVN/PV)1/N
– 1. We denote sales in 2008 as PV and sales in 2012 as FV4. We can then solve for the growth rate as follows:
The calculated growth rate of about 7.7 percent a year shows that Hyundai Steel’s sales grew substantially during the 2008–2012 period.
Solution to 2: In this case, we can speak of a positive compound rate of decrease or a negative compound growth rate. Using Equation 14, we find
In contrast to the positive sales growth, the rate of growth in net profit was approximately –0.80 percent during the 2008–2012 period.
EXAMPLE 18 Calculating a Growth Rate (2)
Toyota Motor Corporation, one of the largest automakers in the world, had consolidated vehicle sales of 7.35 million units in 2012. This is substantially less than consolidated vehicle sales of 8.52 million units five years earlier in 2007. What was the growth rate in number of vehicles sold by Toyota from 2007 to 2012?
Solution: Using Equation 14, we find
The rate of growth in vehicles sold was approximately −2.9 percent during the 2007–2012 period. Note that we can also refer to −2.9 percent as the compound annual growth rate because it is the single number that compounds the number of vehicles sold in 2007 forward to the number of vehicles sold in 2012. Table 4 lists the number of vehicles sold by Toyota from 2007 to 2012.
Table 4 also shows 1 plus the one-year growth rate in number of vehicles sold. We can compute the 1 plus five-year cumulative growth in number of vehicles sold from 2007 to 2012 as the product of quantities (1 + one-year growth rate). We arrive at the same result as when we divide the ending number of vehicles sold, 7.35 million, by the beginning number of vehicles sold, 8.52 million:
TABLE 4 Number of Vehicles Sold, 2007–2012
Year Number of Vehicles Sold (Millions) (1 + g)t t 2007 8.52 0 2008 8.91 8.91/8.52 = 1.045775 1 2009 7.57 7.57/8.91 = 0.849607 2 2010 7.24 7.24/7.57 = 0.956407 3 2011 7.31 7.31/7.24 = 1.009669 4 2012 7.35 7.35/7.31 = 1.005472 5
Source: www.toyota.com.
The right-hand side of the equation is the product of 1 plus the one-year growth rate in number of vehicles sold for each year. Recall that, using Equation 14, we took the fifth root of 7.35/8.52 = 0.862676. In effect, we were solving for the single value of g which, when compounded over five periods, gives the correct product of 1 plus the one-year growth rates.8
In conclusion, we do not need to compute intermediate growth rates as in Table
4 to solve for a compound growth rate g. Sometimes, however, the intermediate growth rates are interesting or informative. For example, at first (from 2007 to 2008), Toyota Motors increased its number of vehicles sold. We can also analyze the variability in growth rates when we conduct an analysis as in Table 4. Most of the decline in Toyota Motor ’s sales occurred in 2009. Elsewhere in Toyota Motor ’s disclosures, the company noted that the substantial decline in vehicle sales in 2009 was due to the steep downturn in the global economy. Sales declined further in 2010 as the market conditions remained difficult. Each of the next two years saw a slight increase in sales.
The compound growth rate is an excellent summary measure of growth over multiple time periods. In our Toyota Motors example, the compound growth rate of −2.9 percent is the single growth rate that, when added to 1, compounded over five years, and multiplied by the 2007 number of vehicles sold, yields the 2012 number of vehicles sold.
7.2. Solving for the Number of Periods In this section, we demonstrate how to solve for the number of periods given present value, future value, and interest or growth rates.
EXAMPLE 19 The Number of Annual Compounding Periods Needed for an Investment to Reach a Specific Value
You are interested in determining how long it will take an investment of €10,000,000 to double in value. The current interest rate is 7 percent compounded annually. How many years will it take €10,000,000 to double to €20,000,000?
Solution: Use Equation 2, FVN = PV(1 + r)N, to solve for the number of periods, N, as follows:
With an interest rate of 7 percent, it will take approximately 10 years for the initial €10,000,000 investment to grow to €20,000,000. Solving for N in the expression (1.07)N = 2.0 requires taking the natural logarithm of both sides and using the rule that ln(xN) = N ln(x). Generally, we find that N = [ln(FV/PV)]/ln(1
+ r). Here, N = ln(€20,000,000/ €10,000,000)/ln(1.07) = ln(2)/ln(1.07) = 10.24.9
7.3. Solving for the Size of Annuity Payments In this section, we discuss how to solve for annuity payments. Mortgages, auto loans, and retirement savings plans are classic examples of applications of annuity formulas.
EXAMPLE 20 Calculating the Size of Payments on a Fixed-Rate Mortgage
You are planning to purchase a $120,000 house by making a down payment of $20,000 and borrowing the remainder with a 30-year fixed-rate mortgage with monthly payments. The first payment is due at t = 1. Current mortgage interest rates are quoted at 8 percent with monthly compounding. What will your monthly mortgage payments be?
Solution: The bank will determine the mortgage payments such that at the stated periodic interest rate, the present value of the payments will be equal to the amount borrowed (in this case, $100,000). With this fact in mind, we can use
Equation 11, , to solve for the annuity amount, A, as the present value divided
by the present value annuity factor:
The amount borrowed, $100,000, is equivalent to 360 monthly payments of $733.76 with a stated interest rate of 8 percent. The mortgage problem is a relatively straightforward application of finding a level annuity payment.
Next, we turn to a retirement-planning problem. This problem illustrates the complexity of the situation in which an individual wants to retire with a specified retirement income. Over the course of a life cycle, the individual may be able to save only a small amount during the early years but then may have the financial resources to save more during later years. Savings plans often involve uneven cash flows, a topic we will examine in the last part of this reading. When dealing with uneven cash flows, we take maximum advantage of the principle that dollar amounts indexed at the same point in time are additive—the cash flow additivity principle.
EXAMPLE 21 The Projected Annuity Amount Needed to Fund a Future-Annuity Inflow
Jill Grant is 22 years old (at t = 0) and is planning for her retirement at age 63 (at t = 41). She plans to save $2,000 per year for the next 15 years (t = 1 to t = 15). She wants to have retirement income of $100,000 per year for 20 years, with the first retirement payment starting at t = 41. How much must Grant save each year from t = 16 to t = 40 in order to achieve her retirement goal? Assume she plans to invest in a diversified stock-and-bond mutual fund that will earn 8 percent per year on average.
Solution: To help solve this problem, we set up the information on a time line. As Figure 8 shows, Grant will save $2,000 (an outflow) each year for Years 1
to 15. Starting in Year 41, Grant will start to draw retirement income of $100,000 per year for 20 years. In the time line, the annual savings is recorded in parentheses ($2) to show that it is an outflow. The problem is to find the savings, recorded as X, from Year 16 to Year 40.
Figure 8 Solving for Missing Annuity Payments (in Thousands)
Solving this problem involves satisfying the following relationship: the present value of savings (outflows) equals the present value of retirement income (inflows). We could bring all the dollar amounts to t = 40 or to t = 15 and solve for X.
Let us evaluate all dollar amounts at t = 15 (we encourage the reader to repeat the problem by bringing all cash flows to t = 40). As of t = 15, the first payment of X will be one period away (at t = 16). Thus we can value the stream of Xs using the formula for the present value of an ordinary annuity.
This problem involves three series of level cash flows. The basic idea is that the present value of the retirement income must equal the present value of Grant’s savings. Our strategy requires the following steps:
1. Find the future value of the savings of $2,000 per year and index it at t = 15.
This value tells us how much Grant will have saved.
2. Find the present value of the retirement income at t = 15. This value tells us how much Grant needs to meet her retirement goals (as of t = 15). Two substeps are necessary. First, calculate the present value of the annuity of $100,000 per year at t = 40. Use the formula for the present value of an annuity. (Note that the present value is indexed at t = 40 because the first payment is at t = 41.) Next, discount the present value back to t = 15 (a total of 25 periods).
3. Now compute the difference between the amount Grant has saved (Step 1) and the amount she needs to meet her retirement goals (Step 2). Her savings from t = 16 to t = 40 must have a present value equal to the difference between the future value of her savings and the present value of her retirement income.
Our goal is to determine the amount Grant should save in each of the 25 years from t = 16 to t = 40. We start by bringing the $2,000 savings to t = 15, as follows:
At t = 15, Grant’s initial savings will have grown to $54,304.23.
Now we need to know the value of Grant’s retirement income at t = 15. As stated earlier, computing the retirement present value requires two substeps. First, find the present value at t = 40 with the formula in Equation 11; second, discount this present value back to t = 15. Now we can find the retirement income present value at t = 40:
The present value amount is as of t = 40, so we must now discount it back as a lump sum to t = 15:
Now recall that Grant will have saved $54,304.23 by t = 15. Therefore, in
present value terms, the annuity from t = 16 to t = 40 must equal the difference between the amount already saved ($54,304.23) and the amount required for retirement ($143,362.53). This amount is equal to $143,362.53 − $54,304.23 = $89,058.30. Therefore, we must now find the annuity payment, A, from t = 16 to t = 40 that has a present value of $89,058.30. We find the annuity payment as follows:
Grant will need to increase her savings to $8,342.87 per year from t = 16 to t = 40 to meet her retirement goal of having a fund equal to $981,814.74 after making her last payment at t = 40.
7.4. Review of Present and Future Value Equivalence As we have demonstrated, finding present and future values involves moving amounts of money to different points on a time line. These operations are possible because present value and future value are equivalent measures separated in time. Table 5 illustrates this equivalence; it lists the timing of five cash flows, their present values at t = 0, and their future values at t = 5.
To interpret Table 5, start with the third column, which shows the present values. Note that each $1,000 cash payment is discounted back the appropriate number of periods to find the present value at t = 0. The present value of $4,329.48 is exactly equivalent to the series of cash flows. This information illustrates an important point: A lump sum can actually generate an annuity. If we place a lump sum in an account that earns the stated interest rate for all periods, we can generate an annuity that is equivalent to the lump sum. Amortized loans, such as mortgages and car
loans, are examples of this principle.
TABLE 5 The Equivalence of Present and Future Values
Time Cash Flow ($) Present Value at t = 0 Future Value at t = 5 1 1,000 $1,000(1.05)−1= $952.38 $1,000(1.05)4= $1,215.51 2 1,000 $1,000(1.05)−2= $907.03 $1,000(1.05)3= $1,157.63 3 1,000 $1,000(1.05)−3= $863.84 $1,000(1.05)2= $1,102.50 4 1,000 $1,000(1.05)−4= $822.70 $1,000(1.05)1= $1,050.00 5 1,000 $1,000(1.05)−5= $783.53 $1,000(1.05)0= $1,000.00
Sum: $4,329.48 Sum: $5,525.64
To see how a lump sum can fund an annuity, assume that we place $4,329.48 in the bank today at 5 percent interest. We can calculate the size of the annuity payments by using Equation 11. Solving for A, we find
Table 6 shows how the initial investment of $4,329.48 can actually generate five $1,000 withdrawals over the next five years.
To interpret Table 6, start with an initial present value of $4,329.48 at t = 0. From t = 0 to t = 1, the initial investment earns 5 percent interest, generating a future value of $4,329.48(1.05) = $4,545.95. We then withdraw $1,000 from our account, leaving $4,545.95 − $1,000 = $3,545.95 (the figure reported in the last column for time period 1). In the next period, we earn one year ’s worth of interest and then make a $1,000 withdrawal. After the fourth withdrawal, we have $952.38, which earns 5 percent. This amount then grows to $1,000 during the year, just enough for us to make the last withdrawal. Thus the initial present value, when invested at 5 percent for five years, generates the $1,000 five-year ordinary annuity. The present value of the initial investment is exactly equivalent to the annuity.
Now we can look at how future value relates to annuities. In Table 5, we reported that the future value of the annuity was $5,525.64. We arrived at this figure by compounding the first $1,000 payment forward four periods, the second $1,000
forward three periods, and so on. We then added the five future amounts at t = 5. The annuity is equivalent to $5,525.64 at t = 5 and $4,329.48 at t = 0. These two dollar measures are thus equivalent. We can verify the equivalence by finding the present value of $5,525.64, which is $5,525.64 × (1.05)−5 = $4,329.48. We found this result above when we showed that a lump sum can generate an annuity.
To summarize what we have learned so far: A lump sum can be seen as equivalent to an annuity, and an annuity can be seen as equivalent to its future value. Thus present values, future values, and a series of cash flows can all be considered equivalent as long as they are indexed at the same point in time.
TABLE 6 How an Initial Present Value Funds an Annuity
Time Period
Amount Availableat the Beginning ofthe Time
Period ($)
Ending Amount before
Withdrawal
Withdrawal ($)
Amount Available after
Withdrawal ($)
1 4,329.48 $4,329.48(1.05) =$4,545.95 1,000 3,545.95
2 3,545.95 $3,545.95(1.05) =$3,723.25 1,000 2,723.25
3 2,723.25 $2,723.25(1.05) =$2,859.41 1,000 1,859.41
4 1,859.41 $1,859.41(1.05) =$1,952.38 1,000 952.38
5 952.38 $952.38(1.05) =$1,000 1,000 0
7.5. The Cash Flow Additivity Principle The cash flow additivity principle—the idea that amounts of money indexed at the same point in time are additive—is one of the most important concepts in time value of money mathematics. We have already mentioned and used this principle; this section provides a reference example for it.
Consider the two series of cash flows shown on the time line in Figure 9. The series are denoted A and B. If we assume that the annual interest rate is 2 percent, we can find the future value of each series of cash flows as follows. Series A’s future value is $100(1.02) + $100 = $202. Series B’s future value is $200(1.02) + $200 = $404. The future value of (A + B) is $202 + $404 = $606 by the method we have used up to
this point. The alternative way to find the future value is to add the cash flows of each series, A and B (call it A + B), and then find the future value of the combined cash flow, as shown in Figure 9.
The third time line in Figure 9 shows the combined series of cash flows. Series A has a cash flow of $100 at t = 1, and Series B has a cash flow of $200 at t = 1. The combined series thus has a cash flow of $300 at t = 1. We can similarly calculate the cash flow of the combined series at t = 2. The future value of the combined series (A + B) is $300(1.02) + $300 = $606—the same result we found when we added the future values of each series.
The additivity and equivalence principles also appear in another common situation. Suppose cash flows are $4 at the end of the first year and $24 (actually separate payments of $4 and $20) at the end of the second year. Rather than finding present values of the first year ’s $4 and the second year ’s $24, we can treat this situation as a $4 annuity for two years and a second-year $20 lump sum. If the discount rate were 6 percent, the $4 annuity would have a present value of $7.33 and the $20 lump sum a present value of $17.80, for a total of $25.13.
Figure 9 The Additivity of Two Series of Cash Flows
8. Summary In this reading, we have explored a foundation topic in investment mathematics, the time value of money. We have developed and reviewed the following concepts for use in financial applications:
The interest rate, r, is the required rate of return; r is also called the discount rate or opportunity cost.
An interest rate can be viewed as the sum of the real risk-free interest rate and a set of premiums that compensate lenders for risk: an inflation premium, a default risk premium, a liquidity premium, and a maturity premium.
The future value, FV, is the present value, PV, times the future value factor, (1 + r)N.
The interest rate, r, makes current and future currency amounts equivalent based on their time value.
The stated annual interest rate is a quoted interest rate that does not account for compounding within the year.
The periodic rate is the quoted interest rate per period; it equals the stated annual interest rate divided by the number of compounding periods per year.
The effective annual rate is the amount by which a unit of currency will grow in a year with interest on interest included.
An annuity is a finite set of level sequential cash flows.
There are two types of annuities, the annuity due and the ordinary annuity. The annuity due has a first cash flow that occurs immediately; the ordinary annuity has a first cash flow that occurs one period from the present (indexed at t = 1).
On a time line, we can index the present as 0 and then display equally spaced hash marks to represent a number of periods into the future. This representation allows us to index how many periods away each cash flow will be paid.
Annuities may be handled in a similar fashion as single payments if we use annuity factors instead of single-payment factors.
The present value, PV, is the future value, FV, times the present value factor, (1 + r)−N.
The present value of a perpetuity is A/r, where A is the periodic payment to be
received forever.
It is possible to calculate an unknown variable, given the other relevant variables in time value of money problems.
The cash flow additivity principle can be used to solve problems with uneven cash flows by combining single payments and annuities.
Problems Practice Problems and Solutions: 1–20 taken from Quantitative Methods for Investment Analysis, Second Edition, by Richard A. DeFusco, CFA, Dennis W. McLeavey, CFA, Jerald E. Pinto, CFA, and David E. Runkle, CFA. Copyright © 2004 by CFA Institute. All other problems and solutions copyright © CFA Institute. 1. The table below gives current information on the interest rates for two two-
year and two eight-year maturity investments. The table also gives the maturity, liquidity, and default risk characteristics of a new investment possibility (Investment 3). All investments promise only a single payment (a payment at maturity). Assume that premiums relating to inflation, liquidity, and default risk are constant across all time horizons.
Investment Maturity (in Years) Liquidity Default Risk Interest Rate (%) 1 2 High Low 2.0 2 2 Low Low 2.5 3 7 Low Low r3 4 8 High Low 4.0 5 8 Low High 6.5
Based on the information in the above table, address the following: 1. Explain the difference between the interest rates on Investment 1 and
Investment 2.
2. Estimate the default risk premium.
3. Calculate upper and lower limits for the interest rate on Investment 3, r3.
2. A client has a $5 million portfolio and invests 5 percent of it in a money market fund projected to earn 3 percent annually. Estimate the value of this portion of his portfolio after seven years.
3. A client invests $500,000 in a bond fund projected to earn 7 percent annually. Estimate the value of her investment after 10 years.
4. For liquidity purposes, a client keeps $100,000 in a bank account. The bank quotes a stated annual interest rate of 7 percent. The bank’s service
representative explains that the stated rate is the rate one would earn if one were to cash out rather than invest the interest payments. How much will your client have in his account at the end of one year, assuming no additions or withdrawals, using the following types of compounding?
1. Quarterly.
2. Monthly.
3. Continuous.
5. A bank quotes a rate of 5.89 percent with an effective annual rate of 6.05 percent. Does the bank use annual, quarterly, or monthly compounding?
6. A bank pays a stated annual interest rate of 8 percent. What is the effective annual rate using the following types of compounding?
1. Quarterly.
2. Monthly.
3. Continuous.
7. A couple plans to set aside $20,000 per year in a conservative portfolio projected to earn 7 percent a year. If they make their first savings contribution one year from now, how much will they have at the end of 20 years?
8. Two years from now, a client will receive the first of three annual payments of $20,000 from a small business project. If she can earn 9 percent annually on her investments and plans to retire in six years, how much will the three business project payments be worth at the time of her retirement?
9. To cover the first year ’s total college tuition payments for his two children, a father will make a $75,000 payment five years from now. How much will he need to invest today to meet his first tuition goal if the investment earns 6 percent annually?
10. A client has agreed to invest €100,000 one year from now in a business planning to expand, and she has decided to set aside the funds today in a bank account that pays 7 percent compounded quarterly. How much does she need to set aside?
11. A client can choose between receiving 10 annual $100,000 retirement payments, starting one year from today, or receiving a lump sum today. Knowing that he can invest at a rate of 5 percent annually, he has decided to take the lump sum. What lump sum today will be equivalent to the future annual payments?
12. A perpetual preferred stock position pays quarterly dividends of $1,000 indefinitely (forever). If an investor has a required rate of return of 12 percent per year compounded quarterly on this type of investment, how much should he be willing to pay for this dividend stream?
13. At retirement, a client has two payment options: a 20-year annuity at €50,000 per year starting after one year or a lump sum of €500,000 today. If the client’s required rate of return on retirement fund investments is 6 percent per year, which plan has the higher present value and by how much?
14. You are considering investing in two different instruments. The first instrument will pay nothing for three years, but then it will pay $20,000 per year for four years. The second instrument will pay $20,000 for three years and $30,000 in the fourth year. All payments are made at year-end. If your required rate of return on these investments is 8 percent annually, what should you be willing to pay for:
1. The first instrument?
2. The second instrument (use the formula for a four-year annuity)?
15. Suppose you plan to send your daughter to college in three years. You expect her to earn two-thirds of her tuition payment in scholarship money, so you estimate that your payments will be $10,000 a year for four years. To estimate whether you have set aside enough money, you ignore possible inflation in tuition payments and assume that you can earn 8 percent annually on your investments. How much should you set aside now to cover these payments?
16. A client is confused about two terms on some certificate-of-deposit rates quoted at his bank in the United States. You explain that the stated annual interest rate is an annual rate that does not take into account compounding within a year. The rate his bank calls APY (annual percentage yield) is the effective annual rate taking into account compounding. The bank’s customer service representative mentioned monthly compounding, with $1,000 becoming $1,061.68 at the end of a year. To prepare to explain the terms to your client, calculate the stated annual interest rate that the bank must be quoting.
17. A client seeking liquidity sets aside €35,000 in a bank account today. The account pays 5 percent compounded monthly. Because the client is concerned about the fact that deposit insurance covers the account for only up to €100,000, calculate how many months it will take to reach that amount.
18. A client plans to send a child to college for four years starting 18 years from now. Having set aside money for tuition, she decides to plan for room and
board also. She estimates these costs at $20,000 per year, payable at the beginning of each year, by the time her child goes to college. If she starts next year and makes 17 payments into a savings account paying 5 percent annually, what annual payments must she make?
19. A couple plans to pay their child’s college tuition for 4 years starting 18 years from now. The current annual cost of college is C$7,000, and they expect this cost to rise at an annual rate of 5 percent. In their planning, they assume that they can earn 6 percent annually. How much must they put aside each year, starting next year, if they plan to make 17 equal payments?
20. You are analyzing the last five years of earnings per share data for a company. The figures are $4.00, $4.50, $5.00, $6.00, and $7.00. At what compound annual rate did EPS grow during these years?
21. An analyst expects that a company’s net sales will double and the company’s net income will triple over the next five-year period starting now. Based on the analyst’s expectations, which of the following best describes the expected compound annual growth?
1. Net sales will grow 15% annually and net income will grow 25% annually.
2. Net sales will grow 20% annually and net income will grow 40% annually.
3. Net sales will grow 25% annually and net income will grow 50% annually.
Notes 1 Examples in this reading and other readings in quantitative methods at Level I
were updated in 2013 by Professor Sanjiv Sabherwal of the University of Texas, Arlington.
2 Technically, 1 plus the nominal rate equals the product of 1 plus the real rate and 1 plus the inflation rate. As a quick approximation, however, the nominal rate is equal to the real rate plus an inflation premium. In this discussion we focus on approximate additive relationships to highlight the underlying concepts.
3 Other developed countries issue securities similar to US Treasury bills. The French government issues BTFs or negotiable fixed-rate discount Treasury bills (Bons du Tésor àtaux fixe et à intérêts pécomptés) with maturities of up to one year. The Japanese government issues a short-term Treasury bill with maturities of 6 and 12 months. The German government issues at discount both Treasury financing paper (Finanzierungsschätzedes Bundes or, for short, Schätze) and Treasury discount paper (Bubills) with maturities up to 24 months. In the United Kingdom, the British government issues gilt-edged Treasury bills with maturities ranging from 1 to 364 days. The Canadian government bond market is closely related to the US market; Canadian Treasury bills have maturities of 3, 6, and 12 months.
4 We could also solve time value of money problems using tables of interest rate factors. Solutions using tabled values of interest rate factors are generally less accurate than solutions obtained using calculators or spreadsheets, so practitioners prefer calculators or spreadsheets.
5 To avoid rounding errors when using a financial calculator, divide 8 by 12 and then press the %i key, rather than simply entering 0.67 for %i, so we have (1 + 0.08/12)12 = 1.083000.
6 Among the terms used for the effective annual return on interest-bearing bank deposits are annual percentage yield (APY) in the United States and equivalent annual rate (EAR) in the United Kingdom. By contrast, the annual percentage rate (APR) measures the cost of borrowing expressed as a yearly rate. In the United States, the APR is calculated as a periodic rate times the number of payment periods per year and, as a result, some writers use APR as a general synonym for the stated annual interest rate. Nevertheless, APR is a term with legal connotations; its calculation follows regulatory standards that vary
internationally. Therefore, “stated annual interest rate” is the preferred general term for an annual interest rate that does not account for compounding within the year.
7 There is an alternative way to calculate the present value of an annuity due. Compared to an ordinary annuity, the payments in an annuity due are each discounted one less period. Therefore, we can modify Equation 11 to handle annuities due by multiplying the right-hand side of the equation by (1 + r):
8 The compound growth rate that we calculate here is an example of a geometric mean, specifically the geometric mean of the growth rates. We define the geometric mean in the reading on statistical concepts.
9 To quickly approximate the number of periods, practitioners sometimes use an ad hoc rule called the Rule of 72: Divide 72 by the stated interest rate to get the approximate number of years it would take to double an investment at the interest rate. Here, the approximation gives 72/7 = 10.3 years. The Rule of 72 is loosely based on the observation that it takes 12 years to double an amount at a 6 percent interest rate, giving 6 × 12 = 72. At a 3 percent rate, one would guess it would take twice as many years, 3 × 24 = 72.
CHAPTER 2 DISCOUNTED CASH FLOW APPLICATIONS Richard A. DeFusco, CFA
Dennis W. McLeavey, CFA
Jerald E. Pinto, PhD, CFA
David E. Runkle, PhD, CFA
LEARNING OUTCOMES
After completing this chapter, you will be able to do the following:
calculate and interpret the net present value (NPV) and the internal rate of return (IRR) of an investment;
contrast the NPV rule to the IRR rule, and identify problems associated with the IRR rule;
calculate and interpret a holding period return (total return);
calculate and compare the money-weighted and time-weighted rates of return of a portfolio and evaluate the performance of portfolios based on these measures;
calculate and interpret the bank discount yield, holding period yield, effective annual yield, and money market yield for US Treasury bills and other money market instruments;
convert among holding period yields, money market yields, effective annual yields, and bond equivalent yields.
1. Introduction As investment analysts, much of our work includes evaluating transactions involving present and future cash flows. In the reading on the time value of money (TVM), we presented the mathematics needed to solve those problems and illustrated the techniques for the major problem types. In this reading we turn to applications. Analysts must master the numerous applications of TVM or discounted cash flow analysis in equity, fixed income, and derivatives analysis as they study each of those topics individually. In this reading, we present a selection of important TVM applications: net present value and internal rate of return as tools for evaluating cash flow streams, portfolio return measurement, and the calculation of money market yields. Important in themselves, these applications also introduce concepts that reappear in many other investment contexts.
The reading is organized as follows. Section 2 introduces two key TVM concepts, net present value and internal rate of return. Building on these concepts, Section 3 discusses a key topic in investment management, portfolio return measurement. Investment managers often face the task of investing funds for the short term; to understand the choices available, they need to understand the calculation of money market yields. The reading thus concludes with a discussion of that topic in Section 4.
2. Net Present Value and Internal Rate of Return In applying discounted cash flow analysis in all fields of finance, we repeatedly encounter two concepts, net present value and internal rate of return. In the following sections we present these keystone concepts.
We could explore the concepts of net present value and internal rate of return in many contexts, because their scope of application covers all areas of finance. Capital budgeting, however, can serve as a representative starting point. Capital budgeting is important not only in corporate finance but also in security analysis, because both equity and fixed income analysts must be able to assess how well managers are investing the assets of their companies. There are three chief areas of financial decision-making in most businesses. Capital budgeting is the allocation of funds to relatively long-range projects or investments. From the perspective of capital budgeting, a company is a portfolio of projects and investments. Capital structure is the choice of long-term financing for the investments the company wants to make. Working capital management is the management of the company’s short-term assets (such as inventory) and short-term liabilities (such as money owed to suppliers).
2.1. Net Present Value and the Net Present Value Rule Net present value (NPV) describes a way to characterize the value of an investment, and the net present value rule is a method for choosing among alternative investments. The net present value of an investment is the present value of its cash inflows minus the present value of its cash outflows. The word “net” in net present value refers to subtracting the present value of the investment’s outflows (costs) from the present value of its inflows (benefits) to arrive at the net benefit.
The steps in computing NPV and applying the NPV rule are as follows: 1. Identify all cash flows associated with the investment—all inflows and
outflows.1
2. Determine the appropriate discount rate or opportunity cost, r, for the investment project.2
3. Using that discount rate, find the present value of each cash flow. (Inflows have a positive sign and increase NPV; outflows have a negative sign and decrease NPV.)
4. Sum all present values. The sum of the present values of all cash flows (inflows
and outflows) is the investment’s net present value.
5. Apply the NPV rule: If the investment’s NPV is positive, an investor should undertake it; if the NPV is negative, the investor should not undertake it. If an investor has two candidates for investment but can only invest in one (i.e., mutually exclusive projects), the investor should choose the candidate with the higher positive NPV.
What is the meaning of the NPV rule? In calculating the NPV of an investment proposal, we use an estimate of the opportunity cost of capital as the discount rate. The opportunity cost of capital is the alternative return that investors forgo in undertaking the investment. When NPV is positive, the investment adds value because it more than covers the opportunity cost of the capital needed to undertake it. So a company undertaking a positive NPV investment increases shareholders’ wealth. An individual investor making a positive NPV investment increases personal wealth, but a negative NPV investment decreases wealth.
When working problems using the NPV rule, it will be helpful to refer to the following formula:
(1)
where
CFt = the expected net cash flow at time t
N = the investment’s projected life
r = the discount rate or opportunity cost of capital
As always, we state the inputs on a compatible time basis: If cash flows are annual, N is the project’s life in years and r is an annual rate. For instance, suppose you are reviewing a proposal that requires an initial outlay of $2 million (CF0 = −$2 million). You expect that the proposed investment will generate net positive cash flows of CF1 = $0.50 million at the end of Year 1, CF2 = $0.75 million at the end of Year 2, and CF3 = $1.35 million at the end of Year 3. Using 10 percent as a discount rate, you calculate the NPV as follows:
Because the NPV of $88,655 is positive, you accept the proposal under the NPV rule.
Consider an example in which a research and development program is evaluated using the NPV rule.
EXAMPLE 1 Evaluating a Research and Development Program Using the NPV Rule
As an analyst covering the RAD Corporation, you are evaluating its research and development (R&D) program for the current year. Management has announced that it intends to invest $1 million in R&D. Incremental net cash flows are forecasted to be $150,000 per year in perpetuity. RAD Corporation’s opportunity cost of capital is 10 percent.
1. State whether RAD’s R&D program will benefit shareholders, as judged by the NPV rule.
2. Evaluate whether your answer to Part 1 changes if RAD Corporation’s opportunity cost of capital is 15 percent rather than 10 percent.
Solution to 1: The constant net cash flows of $150,000, which we can denote as , form a perpetuity. The present value of the perpetuity is , so we calculate
the project’s NPV as
With an opportunity cost of 10 percent, the present value of the program’s cash inflows is $1.5 million. The program’s cost is an immediate outflow of $1 million; therefore, its net present value is $500,000. As NPV is positive, you conclude that RAD Corporation’s R&D program will benefit shareholders.
Solution to 2: With an opportunity cost of capital of 15 percent, you compute the NPV as you did above, but this time you use a 15 percent discount rate:
With a higher opportunity cost of capital, the present value of the inflows is smaller and the program’s NPV is smaller: At 15 percent, the NPV exactly equals $0. At NPV = 0, the program generates just enough cash flow to compensate shareholders for the opportunity cost of making the investment. When a company undertakes a zero-NPV project, the company becomes larger but shareholders’ wealth does not increase.
2.2. The Internal Rate of Return and the Internal Rate of Return Rule Financial managers often want a single number that represents the rate of return generated by an investment. The rate of return computation most often used in investment applications (including capital budgeting) is the internal rate of return (IRR). The internal rate of return rule is a second method for choosing among investment proposals. The internal rate of return is the discount rate that makes net present value equal to zero. It equates the present value of the investment’s costs (outflows) to the present value of the investment’s benefits (inflows). The rate is “internal” because it depends only on the cash flows of the investment; no external data are needed. As a result, we can apply the IRR concept to any investment that can be represented as a series of cash flows. In the study of bonds, we encounter IRR under the name of yield to maturity. Later in this reading, we will explore IRR as the money-weighted rate of return for portfolios.
Before we continue, however, we must add a note of caution about interpreting IRR: Even if our cash flow projections are correct, we will realize a compound rate of return that is equal to IRR over the life of the investment only if we can reinvest all interim cash flows at exactly the IRR. Suppose the IRR for a project is 15 percent but we consistently reinvest the cash generated by the project at a lower rate. In this case, we will realize a return that is less than 15 percent. (This principle can work in our favor if we can reinvest at rates above 15 percent.)
To return to the definition of IRR, in mathematical terms we said the following:
(2)
Again, the IRR in Equation 2 must be compatible with the timing of the cash flows. If the cash flows are quarterly, we have a quarterly IRR in Equation 2. We can then state the IRR on an annual basis. For some simple projects, the cash flow at t = 0, CF0, captures the single capital outlay or initial investment; cash flows after t = 0 are the positive returns to the investment. In such cases, we can say CF0 = – Investment (the negative sign indicates an outflow). Thus we can rearrange Equation 2 in a form that is helpful in those cases:
For most real-life problems, financial analysts use software, spreadsheets, or financial calculators to solve this equation for IRR, so you should familiarize
yourself with such tools.3
The investment decision rule using IRR, the IRR rule, states the following: “Accept projects or investments for which the IRR is greater than the opportunity cost of capital.” The IRR rule uses the opportunity cost of capital as a hurdle rate, or rate that a project’s IRR must exceed for the project to be accepted. Note that if the opportunity cost of capital is equal to the IRR, then the NPV is equal to 0. If the project’s opportunity cost is less than the IRR, the NPV is greater than 0 (using a discount rate less than the IRR will make the NPV positive). With these comments in mind, we work through two examples that involve the internal rate of return.
EXAMPLE 2 Evaluating a Research and Development Program Using the IRR Rule
In the previous RAD Corporation example, the initial outlay is $1 million and the program’s cash flows are $150,000 in perpetuity. Now you are interested in determining the program’s internal rate of return. Address the following:
1. Write the equation for determining the internal rate of return of this R&D program.
2. Calculate the IRR.
Solution to 1: Finding the IRR is equivalent to finding the discount rate that makes the NPV equal to 0. Because the program’s cash flows are a perpetuity, you can set up the NPV equation as
or as
Solution to 2: We can solve for IRR as IRR = $150,000/$1,000,000 = 0.15 or 15 percent. The solution of 15 percent accords with the definition of IRR. In Example 1, you found that a discount rate of 15 percent made the program’s NPV equal to 0. By definition, therefore, the program’s IRR must be 15 percent. If the opportunity cost of capital is also 15 percent, the R&D program just covers its opportunity costs and neither increases nor decreases shareholder wealth. If it is less than 15 percent, the IRR rule indicates that management should invest in the program because it more than covers its opportunity cost.
If the opportunity cost is greater than 15 percent, the IRR rule tells management to reject the R&D program. For a given opportunity cost, the IRR rule and the NPV rule lead to the same decision in this example.
EXAMPLE 3 The IRR and NPV Rules Side by Side
The Japanese company Kageyama Ltd. is considering whether or not to open a new factory to manufacture capacitors used in cell phones. The factory will require an investment of ¥1,000 million. The factory is expected to generate level cash flows of ¥294.8 million per year in each of the next five years. According to information in its financial reports, Kageyama’s opportunity cost of capital for this type of project is 11 percent.
1. Determine whether the project will benefit Kageyama’s shareholders using the NPV rule.
2. Determine whether the project will benefit Kageyama’s shareholders using the IRR rule.
Solution to 1: The cash flows can be grouped into an initial outflow of ¥1,000 million and an ordinary annuity of five inflows of ¥294.8 million. The expression for the present value of an annuity is A[1 − (1 + r)−N]/r, where A is the level annuity payment. Therefore, with amounts shown in millions of Japanese yen,
Because the project’s NPV is positive ¥89.55 million, it should benefit Kageyama’s shareholders.
Solution to 2: The IRR of the project is the solution to
This project’s positive NPV tells us that the internal rate of return must be greater than 11 percent. Using a calculator, we find that IRR is 0.145012 or 14.50 percent. Table 1 gives the keystrokes on most financial calculators.
Because the IRR of 14.50 percent is greater than the opportunity cost of the project, the project should benefit Kageyama’s shareholders. Whether it uses
the IRR rule or the NPV rule, Kageyama makes the same decision: Build the factory.
TABLE 1 Computing IRR
Notation Used on Most Calculators Numerical Value for This Problem N 5
%i compute X PV −1,000 PMT 294.8 FV n/a (= 0)
In the previous example, value creation is evident: For a single ¥1,000 million payment, Kageyama creates a project worth ¥1,089.55 million, a value increase of ¥89.55 million. Another perspective on value creation comes from converting the initial investment into a capital charge against the annual operating cash flows that the project generates. Recall that the project generates an annual operating cash flow of ¥294,800,000. If we subtract a capital charge of ¥270,570,310 (the amount of a five-year annuity having a present value of ¥1,000 million at 11 percent), we find ¥294,800,000 − ¥270,570,310 = ¥24,229,690. The amount of ¥24,229,690 represents the profit in each of the next five years after taking into account opportunity costs. The present value of a five-year annuity of ¥24,229,690 at an 11 percent cost of capital is exactly what we calculated as the project’s NPV: ¥89.55 million. Therefore, we can also calculate NPV by converting the initial investment to an annual capital charge against cash flow.
2.3. Problems with the IRR Rule The IRR and NPV rules give the same accept or reject decision when projects are independent—that is, when the decision to invest in one project does not affect the decision to undertake another. When a company cannot finance all the projects it would like to undertake—that is, when projects are mutually exclusive—it must rank the projects from most profitable to least. However, rankings according to IRR and NPV may not be the same. The IRR and NPV rules rank projects differently when
TABLE 2 IRR and NPV for Mutually Exclusive Projects of Different Size
Project Investment at t = 0 Cash Flow at t = 1 IRR (%) NPV at 8% A −€10,000 €15,000 50 €3,888.89
B −€30,000 €42,000 40 €8,888.89
the size or scale of the projects differs (measuring size by the investment needed to undertake the project), or
the timing of the projects’ cash flows differs.
When the IRR and NPV rules conflict in ranking projects, we should take directions from the NPV rule. Why that preference? The NPV of an investment represents the expected addition to shareholder wealth from an investment, and we take the maximization of shareholder wealth to be a basic financial objective of a company. To illustrate the preference for the NPV rule, consider first the case of projects that differ in size. Suppose that a company has only €30,000 available to invest.4 The company has available two one-period investment projects described as A and B in Table 2.
Project A requires an immediate investment of €10,000. This project will make a single cash payment of $15,000 at t = 1. Because the IRR is the discount rate that equates the present value of the future cash flow with the cost of the investment, the IRR equals 50 percent. If we assume that the opportunity cost of capital is 8 percent, then the NPV of Project A is €3,888.89. We compute the IRR and NPV of Project B as 40 percent and €8,888.89, respectively. The IRR and NPV rules indicate that we should undertake both projects, but to do so we would need €40,000—more money than is available. So we need to rank the projects. How do the projects rank according to IRR and NPV?
The IRR rule ranks Project A, with the higher IRR, first. The NPV rule, however, ranks Project B, with the higher NPV, first—a conflict with the IRR rule’s ranking. Choosing Project A because it has the higher IRR would not lead to the largest increase in shareholders’ wealth. Investing in Project A effectively leaves €20,000 (€30,000 minus A’s cost) uninvested. Project A increases wealth by almost €4,000, but Project B increases wealth by almost €9,000. The difference between the two projects’ scale creates the inconsistency in the ranking between the two rules.
IRR and NPV can also rank projects of the same scale differently when the timing of cash flows differs. We can illustrate this principle with Projects A and D, presented in Table 3.
TABLE 3 IRR and NPV for Mutually Exclusive Projects with Different Timing of Cash Flows
Project CF0 (€) CF1 (€) CF2 (€) CF3 (€) IRR (%) NPV at 8%
A −10,000 15,000 0 0 50.0 €3,888.89 D −10,000 0 0 21,220 28.5 €6,845.12
The terms CF0, CF1, CF2, and CF3 represent the cash flows at time periods 0, 1, 2, and 3. The IRR for Project A is the same as it was in the previous example. The IRR for Project D is found as follows:
The IRR for Project D is 28.5 percent, compared with 50 percent for Project A. IRRs and IRR rankings are not affected by any external interest rate or discount rate because a project’s cash flows alone determine the internal rate of return. The IRR calculation furthermore assumes reinvestment at the IRR, so we generally cannot interpret them as achievable rates of return. For Project D, for example, to achieve a 28.5 percent return we would need to earn 28.5 percent on €10,000 for the first year, earn 28.5 percent on €10,000(1.285) = €12,850 the second year, and earn 28.5 percent on €10,000(1.285)2 = €16,512.25 the third year.5 A reinvestment rate such as 50 percent or 28.5 percent may be quite unrealistic. By contrast, the calculation of NPV uses an external market-determined discount rate, and reinvestment is assumed to take place at that discount rate. NPV rankings can depend on the external discount rate chosen. Here, Project D has a larger but more distant cash inflow (€21,220 versus €15,000). As a result, Project D has a higher NPV than Project A at lower discount rates.6 The NPV rule’s assumption about reinvestment rates is more realistic and more economically relevant because it incorporates the market- determined opportunity cost of capital as a discount rate. As a consequence, the NPV is the expected addition to shareholder wealth from an investment.
In summary, when dealing with mutually exclusive projects, choose among them using the NPV rule when the IRR rule and NPV rule conflict.7
3. Portfolio Return Measurement Suppose you are an investor and you want to assess the success of your investments. You face two related but distinct tasks. The first is performance measurement, which involves calculating returns in a logical and consistent manner. Accurate performance measurement provides the basis for your second task, performance appraisal.8 Performance measurement is thus of great importance for all investors and investment managers because it is the foundation for all further analysis.
In our discussion of portfolio return measurement, we will use the fundamental concept of holding period return (HPR), the return that an investor earns over a specified holding period. For an investment that makes one cash payment at the end of the holding period, HPR = (P1 − P0 + D1)/P0, where P0 is the initial investment, P1 is the price received at the end of the holding period, and D1 is the cash paid by the investment at the end of the holding period.
Particularly when we measure performance over many periods, or when the portfolio is subject to additions and withdrawals, portfolio performance measurement is a challenging task. Two of the measurement tools available are the money-weighted rate of return measure and the time-weighted rate of return measure. The first measure we discuss, the money-weighted rate of return, implements a concept we have already covered in the context of capital budgeting: internal rate of return.
3.1. Money-Weighted Rate of Return The first performance measurement concept that we will discuss is an internal rate of return calculation. In investment management applications, the internal rate of return is called the money-weighted rate of return because it accounts for the timing and amount of all cash flows into and out of the portfolio.9
To illustrate the money-weighted return, consider an investment that covers a two- year horizon. At time t = 0, an investor buys one share at $200. At time t = 1, he purchases an additional share at $225. At the end of Year 2, t = 2, he sells both shares for $235 each. During both years, the stock pays a per-share dividend of $5. The t = 1 dividend is not reinvested. Table 4 shows the total cash inflows and outflows.
The money-weighted return on this portfolio is its internal rate of return for the two-year period. The portfolio’s internal rate of return is the rate, r, for which the present value of the cash inflows minus the present value of the cash outflows equals 0, or
The left-hand side of this equation details the outflows: $200 at time t = 0 and $225 at time t = 1. The $225 outflow is discounted back one period because it occurs at t = 1. The right-hand side of the equation shows the present value of the inflows: $5 at time t = 1 (discounted back one period) and $480 (the $10 dividend plus the $470 sale proceeds) at time t = 2 (discounted back two periods).
TABLE 4 Cash Flows
Time Outlay 0 $200 to purchase the first share 1 $225 to purchase the second share
Time Proceeds 1 $5 dividend received from first share (and not reinvested) 2 $10 dividend ($5 per share × 2 shares) received 2 $470 received from selling two shares at $235 per share
To solve for the money-weighted return, we use either a financial calculator that allows us to enter cash flows or a spreadsheet with an IRR function.10 The first step is to group net cash flows by time. For this example, we have −$200 for the t = 0 net cash flow, −$220 = −$225 + $5 for the t = 1 net cash flow, and $480 for the t = 2 net cash flow. After entering these cash flows, we use the spreadsheet’s or calculator ’s IRR function to find that the money-weighted rate of return is 9.39 percent.11
Now we take a closer look at what has happened to the portfolio during each of the two years. In the first year, the portfolio generated a one-period holding period return of ($5 + $225 − $200)/$200 = 15 percent. At the beginning of the second year, the amount invested is $450, calculated as $225 (per share price of stock) × 2 shares, because the $5 dividend was spent rather than reinvested. At the end of the second year, the proceeds from the liquidation of the portfolio are $470 (as detailed in Table 4) plus $10 in dividends (as also detailed in Table 4). So in the second year the portfolio produced a holding period return of ($10 + $470 − $450)/$450 = 6.67 percent. The mean holding period return was (15% + 6.67%)/2 = 10.84 percent. The money-weighted rate of return, which we calculated as 9.39 percent, puts a greater weight on the second year ’s relatively poor performance (6.67 percent) than the first year ’s relatively good performance (15 percent), as more money was invested in the second year than in the first. That is the sense in which returns in this method of
calculating performance are “money weighted.”
As a tool for evaluating investment managers, the money-weighted rate of return has a serious drawback. Generally, the investment manager ’s clients determine when money is given to the investment manager and how much money is given. As we have seen, those decisions may significantly influence the investment manager ’s money-weighted rate of return. A general principle of evaluation, however, is that a person or entity should be judged only on the basis of their own actions, or actions under their control. An evaluation tool should isolate the effects of the investment manager ’s actions. The next section presents a tool that is effective in that respect.
3.2. Time-Weighted Rate of Return An investment measure that is not sensitive to the additions and withdrawals of funds is the time-weighted rate of return. In the investment management industry, the time- weighted rate of return is the preferred performance measure. The time-weighted rate of return measures the compound rate of growth of $1 initially invested in the portfolio over a stated measurement period. In contrast to the money-weighted rate of return, the time-weighted rate of return is not affected by cash withdrawals or additions to the portfolio. The term “time-weighted” refers to the fact that returns are averaged over time. To compute an exact time-weighted rate of return on a portfolio, take the following three steps: 1. Price the portfolio immediately prior to any significant addition or withdrawal
of funds. Break the overall evaluation period into subperiods based on the dates of cash inflows and outflows.
2. Calculate the holding period return on the portfolio for each subperiod.
3. Link or compound holding period returns to obtain an annual rate of return for the year (the time-weighted rate of return for the year). If the investment is for more than one year, take the geometric mean of the annual returns to obtain the time-weighted rate of return over that measurement period.
Let us return to our money-weighted example and calculate the time-weighted rate of return for that investor ’s portfolio. In that example, we computed the holding period returns on the portfolio, Step 2 in the procedure for finding the time- weighted rate of return. Given that the portfolio earned returns of 15 percent during the first year and 6.67 percent during the second year, what is the portfolio’s time- weighted rate of return over an evaluation period of two years?
We find this time-weighted return by taking the geometric mean of the two holding period returns, Step 3 in the procedure above. The calculation of the geometric
mean exactly mirrors the calculation of a compound growth rate. Here, we take the product of 1 plus the holding period return for each period to find the terminal value at t = 2 of $1 invested at t = 0. We then take the square root of this product and subtract 1 to get the geometric mean. We interpret the result as the annual compound growth rate of $1 invested in the portfolio at t = 0. Thus we have
The time-weighted return on the portfolio was 10.76 percent, compared with the money-weighted return of 9.39 percent, which gave larger weight to the second year ’s return. We can see why investment managers find time-weighted returns more meaningful. If a client gives an investment manager more funds to invest at an unfavorable time, the manager ’s money-weighted rate of return will tend to be depressed. If a client adds funds at a favorable time, the money-weighted return will tend to be elevated. The time-weighted rate of return removes these effects.
In defining the steps to calculate an exact time-weighted rate of return, we said that the portfolio should be valued immediately prior to any significant addition or withdrawal of funds. With the amount of cash flow activity in many portfolios, this task can be costly. We can often obtain a reasonable approximation of the time- weighted rate of return by valuing the portfolio at frequent, regular intervals, particularly if additions and withdrawals are unrelated to market movements. The more frequent the valuation, the more accurate the approximation. Daily valuation is commonplace. Suppose that a portfolio is valued daily over the course of a year. To compute the time-weighted return for the year, we first compute each day’s holding period return:
where MVBt equals the market value at the beginning of day t and MVEt equals the market value at the end of day t. We compute 365 such daily returns, denoted r1, r2, …, r365. We obtain the annual return for the year by linking the daily holding period returns in the following way: (1 + r1) × (1 + r2) × … × (1 + r365) − 1. If withdrawals and additions to the portfolio happen only at day’s end, this annual return is a precise time-weighted rate of return for the year. Otherwise, it is an approximate time-weighted return for the year.
If we have a number of years of data, we can calculate a time-weighted return for each year individually, as above. If ri is the time-weighted return for year i, we
calculate an annualized time-weighted return as the geometric mean of N annual returns, as follows:
Example 4 illustrates the calculation of the time-weighted rate of return.
EXAMPLE 4 Time-Weighted Rate of Return
Strubeck Corporation sponsors a pension plan for its employees. It manages part of the equity portfolio in-house and delegates management of the balance to Super Trust Company. As chief investment officer of Strubeck, you want to review the performance of the in-house and Super Trust portfolios over the last four quarters. You have arranged for outflows and inflows to the portfolios to be made at the very beginning of the quarter. Table 5 summarizes the inflows and outflows as well as the two portfolios’ valuations. In the table, the ending value is the portfolio’s value just prior to the cash inflow or outflow at the beginning of the quarter. The amount invested is the amount each portfolio manager is responsible for investing.
Based on the information given, address the following.
1. Calculate the time-weighted rate of return for the in-house account.
2. Calculate the time-weighted rate of return for the Super Trust account.
Solution to 1: To calculate the time-weighted rate of return for the in-house account, we compute the quarterly holding period returns for the account and link them into an annual return. The in-house account’s time-weighted rate of return is 27 percent, calculated as follows:
TABLE 5 Cash Flows for the In-House Strubeck Account and the Super Trust Account
Quarter 1 ($) 2 ($) 3 ($) 4 ($)
In-House Account Beginning value 4,000,000 6,000,000 5,775,000 6,720,000
Beginning of period inflow (outflow) 1,000,000 (500,000) 225,000 (600,000)
Amount invested 5,000,000 5,500,000 6,000,000 6,120,000
Ending value 6,000,000 5,775,000 6,720,000 5,508,000 Super Trust Account Beginning value 10,000,000 13,200,000 12,240,000 5,659,200
Beginning of period inflow (outflow) 2,000,000 (1,200,000) (7,000,000) (400,000)
Amount invested 12,000,000 12,000,000 5,240,000 5,259,200 Ending value 13,200,000 12,240,000 5,659,200 5,469,568
Solution to 2: The account managed by Super Trust has a time-weighted rate of return of 26 percent, calculated as follows:
The in-house portfolio’s time-weighted rate of return is higher than the Super Trust portfolio’s by 100 basis points.
Having worked through this exercise, we are ready to look at a more detailed case.
EXAMPLE 5 Time-Weighted and Money-Weighted Rates of Return Side by Side
Your task is to compute the investment performance of the Walbright Fund during 2014. The facts are as follows:
On 1 January 2014, the Walbright Fund had a market value of $100 million.
During the period 1 January 2014 to 30 April 2014, the stocks in the fund showed a capital gain of $10 million.
On 1 May 2014, the stocks in the fund paid a total dividend of $2 million. All dividends were reinvested in additional shares.
Because the fund’s performance had been exceptional, institutions invested an additional $20 million in Walbright on 1 May 2014, raising assets under management to $132 million ($100 + $10 + $2 + $20).
On 31 December 2014, Walbright received total dividends of $2.64 million. The fund’s market value on 31 December 2014, not including the $2.64 million in dividends, was $140 million.
The fund made no other interim cash payments during 2014.
Based on the information given, address the following.
1. Compute the Walbright Fund’s time-weighted rate of return.
2. Compute the Walbright Fund’s money-weighted rate of return.
3. Interpret the differences between the time-weighted and money-weighted rates of return.
Solution to 1: Because interim cash flows were made on 1 May 2014, we must compute two interim total returns and then link them to obtain an annual return. Table 6 lists the relevant market values on 1 January, 1 May, and 31 December as well as the associated interim four-month (1 January to 1 May) and eight- month (1 May to 31 December) holding period returns.
Now we must geometrically link the four-and eight-month returns to compute an annual return. We compute the time-weighted return as follows:
In this instance, we compute a time-weighted rate of return of 21.03 percent for one year. The four-month and eight-month intervals combine to equal one year. (Taking the square root of the product 1.12 × 1.0806 would be appropriate only if 1.12 and 1.0806 each applied to one full year.)
TABLE 6 Cash Flows for the Walbright Fund
1 January 2014 Beginning portfolio value = $100 million 1 May 2014 Dividends received before additional investment = $2 million
Ending portfolio value = $110 million
New investment = $20 million Beginning market value for last 2/3 of year = $132 million
31 December 2014 Dividends received = $2.64 million Ending portfolio value = $140 million
Solution to 2: To calculate the money-weighted return, we find the discount rate that sets the present value of the outflows (purchases) equal to the present value of the inflows (dividends and future payoff). The initial market value of the fund and all additions to it are treated as cash outflows. (Think of them as expenditures.) Withdrawals, receipts, and the ending market value of the fund are counted as inflows. (The ending market value is the amount investors receive on liquidating the fund.) Because interim cash flows have occurred at four-month intervals, we must solve for the four-month internal rate of return. Table 6 details the cash flows and their timing.
The present value equation (in millions) is as follows:
The left-hand side of the equation shows the investments in the fund or outflows: a $100 million initial investment followed by the $2 million dividend reinvested and an additional $20 million of new investment (both occurring at the end of the first four-month interval, which makes the exponent in the denominator 1). The right-hand side of the equation shows the payoffs or inflows: the $2 million dividend at the first four-month interval followed by the $2.64 million dividend and the terminal market value of $140 million (both occurring at the end of the third four-month interval, which makes the exponent in the denominator 3). The second four-month interval has no cash flow. We can bring all the terms to the right of the equal sign, arranging them in order of time. After simplification,
Using a spreadsheet or IRR-enabled calculator, we use −100, −20, 0, and
$142.64 for the t = 0, t = 1, t = 2, and t = 3 net cash flows, respectively.12 Using either tool, we get a four-month IRR of 6.28 percent. The quick way to annualize this is to multiply by 3. A more accurate way is (1.0628)3 − 1 = 0.20 or 20 percent.
Solution to 3: In this example, the time-weighted return (21.03 percent) is greater than the money-weighted return (20 percent). The Walbright Fund’s performance was relatively poorer during the eight-month period, when the fund owned more shares, than it was overall. This fact is reflected in a lower money-weighted rate of return compared with time-weighted rate of return, as the money-weighted return is sensitive to the timing and amount of withdrawals and additions to the portfolio.
The accurate measurement of portfolio returns is important to the process of evaluating portfolio managers. In addition to considering returns, however, analysts must also weigh risk. When we worked through Example 4, we stopped short of suggesting that in-house management was superior to Super Trust because it earned a higher time-weighted rate of return. With risk in focus, we can talk of risk- adjusted performance and make comparisons—but only cautiously. In other readings, we will discuss the Sharpe ratio, an important risk-adjusted performance measure that we might apply to an investment manager ’s time-weighted rate of return. For now, we have illustrated the major tools for measuring the return on a portfolio.
4. Money Market Yields In our discussion of internal rate of return and net present value, we referred to the opportunity cost of capital as a market-determined rate. In this section, we begin a discussion of discounted cash flow analysis in actual markets by considering short- term debt markets.
To understand the various ways returns are presented in debt markets, we must discuss some of the conventions for quoting yields on money market instruments. The money market is the market for short-term debt instruments (one-year maturity or less). Some instruments require the issuer to repay the lender the amount borrowed plus interest. Others are pure discount instruments that pay interest as the difference between the amount borrowed and the amount paid back.
In the US money market, the classic example of a pure discount instrument is the US Treasury bill (T-bill) issued by the federal government. The face value of a T-bill is the amount the US government promises to pay back to a T-bill investor. In buying a T-bill, investors pay the face amount less the discount, and receive the face amount at maturity. The discount is the reduction from the face amount that gives the price for the T-bill. This discount becomes the interest that accumulates, because the investor receives the face amount at maturity. Thus, investors earn a dollar return equal to the discount if they hold the instrument to maturity. T-bills are by far the most important class of money market instruments in the United States. Other types of money market instruments include commercial paper and bankers’ acceptances, which are discount instruments, and negotiable certificates of deposit, which are interest-bearing instruments. The market for each of these instruments has its own convention for quoting prices or yields. The remainder of this section examines the quoting conventions for T-bills and other money market instruments. In most instances, the quoted yields must be adjusted for use in other present value problems.
Pure discount instruments such as T-bills are quoted differently from US government bonds. T-bills are quoted on a bank discount basis, rather than on a price basis. The bank discount basis is a quoting convention that annualizes, based on a 360-day year, the discount as a percentage of face value. Yield on a bank discount basis is computed as follows:
(3)
where
The bank discount yield (often called simply the discount yield) takes the dollar discount from par, D, and expresses it as a fraction of the face value (not the price) of the T-bill. This fraction is then multiplied by the number of periods of length t in one year (that is, 360/t), where the year is assumed to have 360 days. Annualizing in this fashion assumes simple interest (no compounding). Consider the following example.
EXAMPLE 6 The Bank Discount Yield
Suppose a T-bill with a face value (or par value) of $100,000 and 150 days until maturity is selling for $98,000. What is its bank discount yield?
Solution: For this example, the dollar discount, D, is $2,000. The yield on a bank discount basis is 4.8 percent, as computed with Equation 3:
The bank discount formula takes the T-bill’s dollar discount from face or par as a fraction of face value, 2 percent, and then annualizes by the factor 360/150 = 2.4. The price of discount instruments such as T-bills is quoted using discount yields, so we typically translate discount yield into price.
Suppose we know the bank discount yield of 4.8 percent but do not know the price. We solve for the dollar discount, D, as follows:
With rBD = 4.8 percent, the dollar discount is D = 0.048 × $100,000 × 150/360 = $2,000. Once we have computed the dollar discount, the purchase price for the T-bill is its face value minus the dollar discount, F − D = $100,000 − $2,000 = $98,000.
Yield on a bank discount basis is not a meaningful measure of investors’ return, for three reasons. First, the yield is based on the face value of the bond, not on its purchase price. Returns from investments should be evaluated relative to the amount
that is invested. Second, the yield is annualized based on a 360-day year rather than a 365-day year. Third, the bank discount yield annualizes with simple interest, which ignores the opportunity to earn interest on interest (compound interest).
We can extend Example 6 to discuss three often-used alternative yield measures. The first is the holding period return over the remaining life of the instrument (150 days in the case of the T-bill in Example 6). It determines the return that an investor will earn by holding the instrument to maturity; as used here, this measure refers to an unannualized rate of return (or periodic rate of return). In fixed income markets, this holding period return is also called a holding period yield (HPY).13 For an instrument that makes one cash payment during its life, HPY is
(4)
where
When we use this expression to calculate the holding period yield for an interest- bearing instrument (for example, coupon-bearing bonds), we need to observe an important detail: The purchase and sale prices must include any accrued interest added to the trade price because the bond was traded between interest payment dates. Accrued interest is the coupon interest that the seller earns from the last coupon date but does not receive as a coupon, because the next coupon date occurs after the date of sale.14
For pure discount securities, all of the return is derived by redeeming the bill for more than its purchase price. Because the T-bill is a pure discount instrument, it makes no interest payment and thus D1 = 0. Therefore, the holding period yield is the dollar discount divided by the purchase price, HPY = D/P0, where D = P1 − P0. The holding period yield is the amount that is annualized in the other measures. For the T-bill in Example 6, the investment of $98,000 will pay $100,000 in 150 days. The holding period yield on this investment using Equation 4 is ($100,000 − $98,000)/$98,000 = $2,000/$98,000 = 2.0408 percent. For this example, the periodic return of 2.0408 percent is associated with a 150-day period. If we were to use the T- bill rate of return as the opportunity cost of investing, we would use a discount rate of 2.0408 percent for the 150-day T-bill to find the present value of any other cash flow to be received in 150 days. As long as the other cash flow has risk characteristics similar to those of the T-bill, this approach is appropriate. If the other
cash flow were riskier than the T-bill, then we could use the T-bill’s yield as a base rate, to which we would add a risk premium. The formula for the holding period yield is the same regardless of the currency of denomination.
The second measure of yield is the effective annual yield (EAY). The EAY takes the quantity 1 plus the holding period yield and compounds it forward to one year, then subtracts 1 to recover an annualized return that accounts for the effect of interest- on-interest.15
(5)
TABLE 7 Three Commonly Used Yield Measures
Holding Period Yield (HPY)
Effective Annual Yield (EAY)
Money Market Yield (CD Equivalent Yield)
In our example, we can solve for EAY as follows:
This example illustrates a general rule: The bank discount yield is less than the effective annual yield.
The third alternative measure of yield is the money market yield (also known as the CD equivalent yield). This convention makes the quoted yield on a T-bill comparable to yield quotations on interest-bearing money market instruments that pay interest on a 360-day basis. In general, the money market yield is equal to the annualized holding period yield; assuming a 360-day year, rMM = (HPY)(360/t). Compared to the bank discount yield, the money market yield is computed on the purchase price, so rMM = (rBD)(F/P0). This equation shows that the money market yield is larger than the bank discount yield. In practice, the following expression is more useful because it does not require knowing the T-bill price:
(6)
For the T-bill example, the money market yield is rMM = (360)(0.048)/[360 − (150)
(0.048)] = 4.898 percent.16
Table 7 summarizes the three yield measures we have discussed.
The next example will help you consolidate your knowledge of these yield measures.
EXAMPLE 7 Using the Appropriate Discount Rate
You need to find the present value of a cash flow of $1,000 that is to be received in 150 days. You decide to look at a T-bill maturing in 150 days to determine the relevant interest rate for calculating the present value. You have found a variety of yields for the 150-day bill. Table 8 presents this information.
TABLE 8 Short-Term Money Market Yields
Holding period yield 2.0408% Bank discount yield 4.8% Money market yield 4.898% Effective annual yield 5.0388%
Which yield or yields are appropriate for finding the present value of the $1,000 to be received in 150 days?
Solution: The holding period yield is appropriate, and we can also use the money market yield and effective annual yield after converting them to a holding period yield.
Holding period yield (2.0408 percent). This yield is exactly what we want. Because it applies to a 150-day period, we can use it in a straightforward fashion to find the present value of the $1,000 to be received in 150 days. (Recall the principle that discount rates must be compatible with the time period.) The present value is
Now we can see why the other yield measures are inappropriate or not as easily applied.
Bank discount yield (4.8 percent). We should not use this yield measure to determine the present value of the cash flow. As mentioned earlier, the bank discount yield is based on the face value of the bill and not on its price.
Money market yield (4.898 percent). To use the money market yield, we need to convert it to the 150-day holding period yield by dividing it by (360/150). After obtaining the holding period yield 0.04898/(360/150) = 0.020408, we use it to discount the $1,000 as above.
Effective annual yield (5.0388 percent). This yield has also been annualized, so it must be adjusted to be compatible with the timing of the cash flow. We can obtain the holding period yield from the EAY as follows:
Recall that when we found the effective annual yield, the exponent was 365/150, or the number of 150-day periods in a 365-day year. To shrink the effective annual yield to a 150-day yield, we use the reciprocal of the exponent that we used to annualize.
In Example 7, we converted two short-term measures of annual yield to a holding period yield for a 150-day period. That is one type of conversion. We frequently also need to convert periodic rates to annual rates. The issue can arise both in money markets and in longer-term debt markets. As an example, many bonds (long- term debt instruments) pay interest semiannually. Bond investors compute IRRs for bonds, known as yields to maturity (YTM). If the semiannual yield to maturity is 4 percent, how do we annualize it? An exact approach, taking account of compounding, would be to compute (1.04)2 − 1 = 0.0816 or 8.16 percent. This is what we have been calling an effective annual yield. An approach used in US bond markets, however, is to double the semiannual YTM: 4% × 2 = 8%. The yield to maturity calculated this way, ignoring compounding, has been called a bond equivalent yield. Annualizing a semiannual yield by doubling is putting the yield on a bond-equivalent basis. In practice the result, 8 percent, would be referred to simply as the bond’s yield to maturity. In money markets, if we annualized a six-month period yield by doubling it, in order to make the result comparable to bonds’ YTMs we would also say that the result was a bond equivalent yield.
5. Summary In this reading, we applied the concepts of present value, net present value, and internal rate of return to the fundamental problem of valuing investments. We applied these concepts first to corporate investment, the well-known capital budgeting problem. We then examined the fundamental problem of calculating the return on a portfolio subject to cash inflows and outflows. Finally we discussed money market yields and basic bond market terminology. The following summarizes the reading’s key concepts:
The net present value (NPV) of a project is the present value of its cash inflows minus the present value of its cash outflows. The internal rate of return (IRR) is the discount rate that makes NPV equal to 0. We can interpret IRR as an expected compound return only when all interim cash flows can be reinvested at the internal rate of return and the investment is maintained to maturity.
The NPV rule for decision making is to accept all projects with positive NPV or, if projects are mutually exclusive, to accept the project with the higher positive NPV. With mutually exclusive projects, we rely on the NPV rule. The IRR rule is to accept all projects with an internal rate of return exceeding the required rate of return. The IRR rule can be affected by problems of scale and timing of cash flows.
Money-weighted rate of return and time-weighted rate of return are two alternative methods for calculating portfolio returns in a multiperiod setting when the portfolio is subject to additions and withdrawals. Time-weighted rate of return is the standard in the investment management industry. Money- weighted rate of return can be appropriate if the investor exercises control over additions and withdrawals to the portfolio.
The money-weighted rate of return is the internal rate of return on a portfolio, taking account of all cash flows.
The time-weighted rate of return removes the effects of timing and amount of withdrawals and additions to the portfolio and reflects the compound rate of growth of one unit of currency invested over a stated measurement period.
The bank discount yield for US Treasury bills (and other money market instruments sold on a discount basis) is given by rBD = (F − P0)/F × 360/t = D/F × 360/t, where F is the face amount to be received at maturity, P0 is the price of the Treasury bill, t is the number of days to maturity, and D is the dollar discount.
For a stated holding period or horizon, holding period yield (HPY) = (Ending price − Beginning price + Cash distributions)/(Beginning price). For a US Treasury bill, HPY = D/P0.
The effective annual yield (EAY) is (1 + HPY)365/t − 1.
The money market yield is given by rMM = HPY × 360/t, where t is the number of days to maturity.
For a Treasury bill, money market yield can be obtained from the bank discount yield using rMM = (360 × rBD)/(360 − t × rBD).
We can convert back and forth between holding period yields, money market yields, and effective annual yields by using the holding period yield, which is common to all the calculations.
The bond equivalent yield of a yield stated on a semiannual basis is that yield multiplied by 2.
References
Brealey, Richard A., Stewart C. Myers, and Franklin Allen. 2014. Principles of Corporate Finance. 11th ed. New York: McGraw-Hill.
Fabozzi, Frank J. 2007. Fixed Income Analysis. 2nd ed. Hoboken, NJ: Wiley.
Problems Practice Problems and Solutions: Quantitative Methods for Investment Analysis, Second Edition, by Richard A. DeFusco, CFA, Dennis W. McLeavey, CFA, Jerald E. Pinto, CFA, and David E. Runkle, CFA. Copyright © 2004 by CFA Institute. 1. Waldrup Industries is considering a proposal for a joint venture that will
require an investment of C$13 million. At the end of the fifth year, Waldrup’s joint venture partner will buy out Waldrup’s interest for C$10 million. Waldrup’s chief financial officer has estimated that the appropriate discount rate for this proposal is 12 percent. The expected cash flows are given below. Year Cash Flow (C$) 0 −13,000,000 1 3,000,000 2 3,000,000 3 3,000,000 4 3,000,000 5 10,000,000
1. Calculate this proposal’s NPV.
2. Make a recommendation to the CFO (chief financial officer) concerning whether Waldrup should enter into this joint venture.
2. Waldrup Industries has committed to investing C$5,500,000 in a project with expected cash flows of C$1,000,000 at the end of Year 1, C$1,500,000 at the end of Year 4, and C$7,000,000 at the end of Year 5.
1. Demonstrate that the internal rate of return of the investment is 13.51 percent.
2. State how the internal rate of return of the investment would change if Waldrup’s opportunity cost of capital were to increase by 5 percentage points.
3. Bestfoods, Inc. is planning to spend $10 million on advertising. The company expects this expenditure to result in annual incremental cash flows of $1.6 million in perpetuity. The corporate opportunity cost of capital for this type of project is 12.5 percent.
1. Calculate the NPV for the planned advertising.
2. Calculate the internal rate of return.
3. Should the company go forward with the planned advertising? Explain.
4. Trilever is planning to establish a new factory overseas. The project requires an initial investment of $15 million. Management intends to run this factory for six years and then sell it to a local entity. Trilever ’s finance department has estimated the following yearly cash flows: Year Cash Flow ($) 0 −15,000,000 1 4,000,000 2 4,000,000 3 4,000,000 4 4,000,000 5 4,000,000 6 7,000,000
Trilever ’s CFO decides that the company’s cost of capital of 19 percent is an appropriate hurdle rate for this project.
1. Calculate the internal rate of return of this project.
2. Make a recommendation to the CFO concerning whether to undertake this project.
5. Westcott–Smith is a privately held investment management company. Two other investment counseling companies, which want to be acquired, have contacted Westcott–Smith about purchasing their business. Company A’s price is £2 million. Company B’s price is £3 million. After analysis, Westcott–Smith estimates that Company A’s profitability is consistent with a perpetuity of £300,000 a year. Company B’s prospects are consistent with a perpetuity of £435,000 a year. Westcott–Smith has a budget that limits acquisitions to a maximum purchase cost of £4 million. Its opportunity cost of capital relative to undertaking either project is 12 percent.
1. Determine which company or companies (if any) Westcott–Smith should purchase according to the NPV rule.
2. Determine which company or companies (if any) Westcott–Smith should purchase according to the IRR rule.
3. State which company or companies (if any) Westcott–Smith should purchase. Justify your answer.
6. John Wilson buys 150 shares of ABM on 1 January 2012 at a price of $156.30 per share. A dividend of $10 per share is paid on 1 January 2013. Assume that this dividend is not reinvested. Also on 1 January 2013, Wilson sells 100 shares at a price of $165 per share. On 1 January 2014, he collects a dividend of $15 per share (on 50 shares) and sells his remaining 50 shares at $170 per share.
1. Write the formula to calculate the money-weighted rate of return on Wilson’s portfolio.
2. Using any method, compute the money-weighted rate of return.
3. Calculate the time-weighted rate of return on Wilson’s portfolio.
4. Describe a set of circumstances for which the money-weighted rate of return is an appropriate return measure for Wilson’s portfolio.
5. Describe a set of circumstances for which the time-weighted rate of return is an appropriate return measure for Wilson’s portfolio.
7. Mario Luongo and Bob Weaver both purchase the same stock for €100. One year later, the stock price is €110 and it pays a dividend of €5 per share. Weaver decides to buy another share at €110 (he does not reinvest the €5 dividend, however). Luongo also spends the €5 per share dividend but does not transact in the stock. At the end of the second year, the stock pays a dividend of €5 per share but its price has fallen back to €100. Luongo and Weaver then decide to sell their entire holdings of this stock. The performance for Luongo and Weaver ’s investments are as follows:
Briefly explain any similarities and differences between the performance of Luongo’s and Weaver ’s investments.
8. A Treasury bill with a face value of $100,000 and 120 days until maturity is selling for $98,500.
1. What is the T-bill’s bank discount yield?
2. What is the T-bill’s money market yield?
3. What is the T-bill’s effective annual yield?
9. Jane Cavell has just purchased a 90-day US Treasury bill. She is familiar with yield quotes on German Treasury discount paper but confused about the bank discount quoting convention for the US T-bill she just purchased.
1. Discuss three reasons why bank discount yield is not a meaningful measure of return.
2. Discuss the advantage of money market yield compared with bank discount yield as a measure of return.
3. Explain how the bank discount yield can be converted to an estimate of the holding period return Cavell can expect if she holds the T-bill to maturity.
Notes 1 In developing cash flow estimates, we observe two principles. First, we include
only the incremental cash flows resulting from undertaking the project; we do not include sunk costs (costs that have been committed prior to the project). Second, we account for tax effects by using after-tax cash flows. For a full discussion of these and other issues in capital budgeting, see Brealey, Myers, and Allen (2014).
2 The weighted average cost of capital (WACC) is often used to discount cash flows. This value is a weighted average of the after-tax required rates of return on the company’s common stock, preferred stock, and long-term debt, where the weights are the fraction of each source of financing in the company’s target capital structure. For a full discussion of the issues surrounding the cost of capital, see Brealey, Myers, and Allen (2014).
3 In some real-world capital budgeting problems, the initial investment (which has a minus sign) may be followed by subsequent cash inflows (which have plus signs) and outflows (which have minus signs). In these instances, the project can have more than one IRR. The possibility of multiple solutions is a theoretical limitation of IRR.
4 Or suppose the two projects require the same physical or other resources, so that only one can be undertaken.
5 The ending amount €10,000(1.285)3 = €21,218 differs from the €21,220 amount listed in Table 3 because we rounded IRR.
6 There is a crossover discount rate above which Project A has a higher NPV than Project D. This crossover rate is 18.94 percent.
7 Technically, different reinvestment rate assumptions account for this conflict between the IRR and NPV rules. The IRR rule assumes that the company can earn the IRR on all reinvested cash flows, but the NPV rule assumes that cash flows are reinvested at the company’s opportunity cost of capital. The NPV assumption is far more realistic. For further details on this and other topics in capital budgeting, see Brealey, Myers, and Allen (2014).
8 The term “performance evaluation” has been used as a synonym for performance appraisal.
9 In the United States, the money-weighted return is frequently called the dollar- weighted return. We follow a standard presentation of the money-weighted return as an IRR concept.
10 In this particular case we could solve for r by solving the quadratic equation 480x2 − 220x − 200 = 0 with x = 1/(1 + r), using standard results from algebra. In general, however, we rely on a calculator or spreadsheet software to compute a money-weighted rate of return.
11 Note that the calculator or spreadsheet will give the IRR as a periodic rate. If the periods are not annual, we annualize the periodic rate.
12 By convention, we denote outflow with a negative sign, and we need 0 as a placeholder for the t = 2.
13 Bond-market participants often use the term “yield” when referring to total returns (returns incorporating both price change and income), as in yield to maturity. In other cases, yield refers to returns from income alone (as in current yield, which is annual interest divided by price). As used in this volume and by many writers, holding period yield is a bond market synonym for holding period return, total return, and horizon return.
14 The price with accrued interest is called the full price. Trade prices are quoted “clean” (without accrued interest), but accrued interest, if any, is added to the purchase price. For more on accrued interest, see Fabozzi (2007).
15 Effective annual yield was called the effective annual rate (Equation 5) in the reading on the time value of money.
16 Some national markets use the money market yield formula, rather than the bank discount yield formula, to quote the yields on discount instruments such as T- bills. In Canada, the convention is to quote Treasury bill yields using the money market formula assuming a 365-day year. Yields for German Treasury discount paper with a maturity less than one year and French BTFs (T-bills) are computed with the money market formula assuming a 360-day year.
CHAPTER 3 STATISTICAL CONCEPTS AND MARKET RETURNS Richard A. DeFusco, CFA
Dennis W. McLeavey, CFA
Jerald E. Pinto, PhD, CFA
David E. Runkle, PhD, CFA
LEARNING OUTCOMES
After completing this chapter, you will be able to do the following:
distinguish between descriptive statistics and inferential statistics, between a population and a sample, and among the types of measurement scales;
define a parameter, a sample statistic, and a frequency distribution;
calculate and interpret relative frequencies and cumulative relative frequencies, given a frequency distribution;
describe the properties of a data set presented as a histogram or a frequency polygon;
calculate and interpret measures of central tendency, including the population mean, sample mean, arithmetic mean, weighted average or mean, geometric mean, harmonic mean, median, and mode;
calculate and interpret quartiles, quintiles, deciles, and percentiles;
calculate and interpret 1) a range and a mean absolute deviation and 2) the variance and standard deviation of a population and of a sample;
calculate and interpret the proportion of observations falling within a specified number of standard deviations of the mean using Chebyshev’s inequality;
calculate and interpret the coefficient of variation and the Sharpe ratio;
explain skewness and the meaning of a positively or negatively skewed return distribution;
describe the relative locations of the mean, median, and mode for a unimodal, nonsymmetrical distribution;
explain measures of sample skewness and kurtosis;
compare the use of arithmetic and geometric means when analyzing investment returns.
1. INTRODUCTION Statistical methods provide a powerful set of tools for analyzing data and drawing conclusions from them. Whether we are analyzing asset returns, earnings growth rates, commodity prices, or any other financial data, statistical tools help us quantify and communicate the data’s important features. This reading presents the basics of describing and analyzing data, the branch of statistics known as descriptive statistics. The reading supplies a set of useful concepts and tools, illustrated in a variety of investment contexts. One theme of our presentation, reflected in the reading’s title, is the demonstration of the statistical methods that allow us to summarize return distributions.1 We explore four properties of return distributions:
where the returns are centered (central tendency);
how far returns are dispersed from their center (dispersion);
whether the distribution of returns is symmetrically shaped or lopsided (skewness); and
whether extreme outcomes are likely (kurtosis).
These same concepts are generally applicable to the distributions of other types of data, too.
The reading is organized as follows. After defining some basic concepts in Section 2, in Sections 3 and 4 we discuss the presentation of data: Section 3 describes the organization of data in a table format, and Section 4 describes the graphic presentation of data. We then turn to the quantitative description of how data are distributed: Section 5 focuses on measures that quantify where data are centered, or measures of central tendency. Section 6 presents other measures that describe the location of data. Section 7 presents measures that quantify the degree to which data are dispersed. Sections 8 and 9 describe additional measures that provide a more accurate picture of data. Section 10 provides investment applications of concepts introduced in Section 5.
2. SOME FUNDAMENTAL CONCEPTS Before starting the study of statistics with this reading, it may be helpful to examine a picture of the overall field. In the following, we briefly describe the scope of statistics and its branches of study. We explain the concepts of population and sample. Data come in a variety of types, affecting the ways they can be measured and the appropriate statistical methods for analyzing them. We conclude by discussing the basic types of data measurement.
2.1. The Nature of Statistics The term statistics can have two broad meanings, one referring to data and the other to method. A company’s average earnings per share (EPS) for the last 20 quarters, or its average returns for the past 10 years, are statistics. We may also analyze historical EPS to forecast future EPS, or use the company’s past returns to infer its risk. The totality of methods we employ to collect and analyze data is also called statistics.
Statistical methods include descriptive statistics and statistical inference (inferential statistics). Descriptive statistics is the study of how data can be summarized effectively to describe the important aspects of large data sets. By consolidating a mass of numerical details, descriptive statistics turns data into information. Statistical inference involves making forecasts, estimates, or judgments about a larger group from the smaller group actually observed. The foundation for statistical inference is probability theory, and both statistical inference and probability theory will be discussed in later readings. Our focus in this reading is solely on descriptive statistics.
2.2. Populations and Samples Throughout the study of statistics we make a critical distinction between a population and a sample. In this section, we explain these two terms as well as the related terms “parameter” and “sample statistic.”2
Definition of Population. A population is defined as all members of a specified group.
Any descriptive measure of a population characteristic is called a parameter. Although a population can have many parameters, investment analysts are usually concerned with only a few, such as the mean value, the range of investment returns, and the variance.
Even if it is possible to observe all the members of a population, it is often too expensive in terms of time or money to attempt to do so. For example, if the population is all telecommunications customers worldwide and an analyst is interested in their purchasing plans, she will find it too costly to observe the entire population. The analyst can address this situation by taking a sample of the population.
Definition of Sample. A sample is a subset of a population.
In taking a sample, the analyst hopes it is characteristic of the population. The field of statistics known as sampling deals with taking samples in appropriate ways to achieve the objective of representing the population well. A later reading addresses the details of sampling.
Earlier, we mentioned statistics in the sense of referring to data. Just as a parameter is a descriptive measure of a population characteristic, a sample statistic (statistic, for short) is a descriptive measure of a sample characteristic.
Definition of Sample Statistic. A sample statistic (or statistic) is a quantity computed from or used to describe a sample.
We devote much of this reading to explaining and illustrating the use of statistics in this sense. The concept is critical also in statistical inference, which addresses such problems as estimating an unknown population parameter using a sample statistic.
2.3. Measurement Scales To choose the appropriate statistical methods for summarizing and analyzing data, we need to distinguish among different measurement scales or levels of measurement. All data measurements are taken on one of four major scales: nominal, ordinal, interval, or ratio.
Nominal scales represent the weakest level of measurement: They categorize data but do not rank them. If we assigned integers to mutual funds that follow different investment strategies, the number 1 might refer to a small-cap value fund, the number 2 to a large-cap value fund, and so on for each possible style. This nominal scale categorizes the funds according to their style but does not rank them.
Ordinal scales reflect a stronger level of measurement. Ordinal scales sort data into categories that are ordered with respect to some characteristic. For example, the Morningstar and Standard & Poor ’s star ratings for mutual funds represent an ordinal scale in which one star represents a group of funds judged to have had
relatively the worst performance, with two, three, four, and five stars representing groups with increasingly better performance, as evaluated by those services.
An ordinal scale may also involve numbers to identify categories. For example, in ranking balanced mutual funds based on their five-year cumulative return, we might assign the number 1 to the top 10 percent of funds, and so on, so that the number 10 represents the bottom 10 percent of funds. The ordinal scale is stronger than the nominal scale because it reveals that a fund ranked 1 performed better than a fund ranked 2. The scale tells us nothing, however, about the difference in performance between funds ranked 1 and 2 compared with the difference in performance between funds ranked 3 and 4, or 9 and 10.
Interval scales provide not only ranking but also assurance that the differences between scale values are equal. As a result, scale values can be added and subtracted meaningfully. The Celsius and Fahrenheit scales are interval measurement scales. The difference in temperature between 10°C and 11°C is the same amount as the difference between 40°C and 41°C. We can state accurately that 12°C = 9°C + 3°C, for example. Nevertheless, the zero point of an interval scale does not reflect complete absence of what is being measured; it is not a true zero point or natural zero. Zero degrees Celsius corresponds to the freezing point of water, not the absence of temperature. As a consequence of the absence of a true zero point, we cannot meaningfully form ratios on interval scales.
As an example, 50°C, although five times as large a number as 10°C, does not represent five times as much temperature. Also, questionnaire scales are often treated as interval scales. If an investor is asked to rank his risk aversion on a scale from 1 (extremely risk-averse) to 7 (extremely risk-loving), the difference between a response of 1 and a response of 2 is sometimes assumed to represent the same difference in risk aversion as the difference between a response of 6 and a response of 7. When that assumption can be justified, the data are measured on an interval scale.
Ratio scales represent the strongest level of measurement. They have all the characteristics of interval measurement scales as well as a true zero point as the origin. With ratio scales, we can meaningfully compute ratios as well as meaningfully add and subtract amounts within the scale. As a result, we can apply the widest range of statistical tools to data measured on a ratio scale. Rates of return are measured on a ratio scale, as is money. If we have twice as much money, then we have twice the purchasing power. Note that the scale has a natural zero—zero means no money.
Now that we have addressed the important preliminaries, we can discuss summarizing and describing data.
EXAMPLE 1 Identifying Scales of Measurement
State the scale of measurement for each of the following:
1. Credit ratings for bond issues.3
2. Cash dividends per share.
3. Hedge fund classification types.4
4. Bond maturity in years.
Solution to 1: Credit ratings are measured on an ordinal scale. A rating places a bond issue in a category, and the categories are ordered with respect to the expected probability of default. But the difference in the expected probability of default between AA− and A+, for example, is not necessarily equal to that between BB− and B+. In other words, letter credit ratings are not measured on an interval scale.
Solution to 2: Cash dividends per share are measured on a ratio scale. For this variable, 0 represents the complete absence of dividends; it is a true zero point.
Solution to 3: Hedge fund classification types are measured on a nominal scale. Each type groups together hedge funds with similar investment strategies. In contrast to credit ratings for bonds, however, hedge fund classification schemes do not involve a ranking. Thus such classification schemes are not measured on an ordinal scale.
Solution to 4: Bond maturity is measured on a ratio scale.
3. SUMMARIZING DATA USING FREQUENCY DISTRIBUTIONS In this section, we discuss one of the simplest ways to summarize data—the frequency distribution.
Definition of Frequency Distribution. A frequency distribution is a tabular display of data summarized into a relatively small number of intervals.
Frequency distributions help in the analysis of large amounts of statistical data, and they work with all types of measurement scales.
Rates of return are the fundamental units that analysts and portfolio managers use for making investment decisions and we can use frequency distributions to summarize rates of return. When we analyze rates of return, our starting point is the holding period return (also called the total return).
Holding Period Return Formula. The holding period return for time period t, Rt, is
(1)
where
Thus the holding period return for time period t is the capital gain (or loss) plus distributions divided by the beginning-period price. (For common stocks, the distribution is a dividend; for bonds, the distribution is a coupon payment.) Equation 1 can be used to define the holding period return on any asset for a day, week, month, or year simply by changing the interpretation of the time interval between successive values of the time index, t.
The holding period return, as defined in Equation 1, has two important characteristics. First, it has an element of time attached to it. For example, if a monthly time interval is used between successive observations for price, then the
rate of return is a monthly figure. Second, rate of return has no currency unit attached to it. For instance, suppose that prices are denominated in euros. The numerator and denominator of Equation 1 would be expressed in euros, and the resulting ratio would not have any units because the units in the numerator and denominator would cancel one another. This result holds regardless of the currency in which prices are denominated.5
With these concerns noted, we now turn to the frequency distribution of the holding period returns on the S&P 500 Index.6 First, we examine annual rates of return; then we look at monthly rates of return. The annual rates of return on the S&P 500 calculated with Equation 1 span the period January 1926 to December 2012, for a total of 87 annual observations. Monthly return data cover the period January 1926 to December 2012, for a total of 1,044 monthly observations.
We can state a basic procedure for constructing a frequency distribution as follows.
Construction of a Frequency Distribution.
1. Sort the data in ascending order.
2. Calculate the range of the data, defined as Range = Maximum value − Minimum value.
3. Decide on the number of intervals in the frequency distribution, k.
4. Determine interval width as Range/k.
5. Determine the intervals by successively adding the interval width to the minimum value, to determine the ending points of intervals, stopping after reaching an interval that includes the maximum value.
6. Count the number of observations falling in each interval.
7. Construct a table of the intervals listed from smallest to largest that shows the number of observations falling in each interval.
TABLE 1 Endpoints of Intervals
−4.57 + 4.00 = −0.57 −0.57 + 4.00 = 3.43 3.43 + 4.00 = 7.43
7.4 + 4.00 = 11.43
In Step 4, when rounding the interval width, round up rather than down, to ensure that the final interval includes the maximum value of the data.
As the above procedure makes clear, a frequency distribution groups data into a set of intervals.7 An interval is a set of values within which an observation falls. Each observation falls into only one interval, and the total number of intervals covers all the values represented in the data. The actual number of observations in a given interval is called the absolute frequency, or simply the frequency. The frequency distribution is the list of intervals together with the corresponding measures of frequency.
To illustrate the basic procedure, suppose we have 12 observations sorted in ascending order: −4.57, −4.04, −1.64, 0.28, 1.34, 2.35, 2.38, 4.28, 4.42, 4.68, 7.16, and 11.43. The minimum observation is −4.57 and the maximum observation is +11.43, so the range is +11.43 − (−4.57) = 16. If we set k = 4, the interval width is 16/4 = 4. Table 1 shows the repeated addition of the interval width of 4 to determine the endpoints for the intervals (Step 5).
Thus the intervals are [−4.57 to −0.57), [−0.57 to 3.43), [3.43 to 7.43), and [7.43 to 11.43].8 Table 2 summarizes Steps 5 through 7.
Note that the intervals do not overlap, so each observation can be placed uniquely into one interval.
In practice, we may want to refine the above basic procedure. For example, we may want the intervals to begin and end with whole numbers for ease of interpretation. We also need to explain the choice of the number of intervals, k. We turn to these issues in discussing the construction of frequency distributions for the S&P 500.
TABLE 2 Frequency Distribution
Interval Absolute Frequency A −4.57 ≤ observation < −0.57 3 B −0.57 ≤ observation < 3.43 4 C 3.43 ≤ observation < 7.43 4 D 7.43 ≤ observation ≤ 11.43 1
We first consider the case of constructing a frequency distribution for the annual returns on the S&P 500 over the period 1926 to 2012. During that period, the return on the S&P 500 had a minimum value of −43.34 percent (in 1931) and a maximum
value of +53.99 percent (in 1933). Thus the range of the data was +54% − (−43%) = 97%, approximately. The question now is the number k of intervals into which we should group observations. Although some guidelines for setting k have been suggested in statistical literature, the setting of a useful value for k often involves inspecting the data and exercising judgment. How much detail should we include? If we use too few intervals, we will summarize too much and lose pertinent characteristics. If we use too many intervals, we may not summarize enough.
We can establish an appropriate value for k by evaluating the usefulness of the resulting interval width. A large number of empty intervals may indicate that we are trying to organize the data to present too much detail. Starting with a relatively small interval width, we can see whether or not the intervals are mostly empty and whether or not the value of k associated with that interval width is too large. If intervals are mostly empty or k is very large, we can consider increasingly larger intervals (smaller values of k) until we have a frequency distribution that effectively summarizes the distribution. For the annual S&P 500 series, return intervals of 1 percent width would result in 97 intervals and many of them would be empty because we have only 87 annual observations. We need to keep in mind that the purpose of a frequency distribution is to summarize the data. Suppose that for ease of interpretation we want to use an interval width stated in whole rather than fractional percents. A 2 percent interval width would have many fewer empty intervals than a 1 percent interval width and effectively summarize the data. A 2 percent interval width would be associated with 97/2 = 48.5 intervals, which we can round up to 49 intervals. That number of intervals will cover 2% × 49 = 98%. We can confirm that if we start the smallest 2 percent interval at the whole number −44.0 percent, the final interval ends at −44.0% + 98% = 54% and includes the maximum return in the sample, 53.99 percent. In so constructing the frequency distribution, we will also have intervals that end and begin at a value of 0 percent, allowing us to count the negative and positive returns in the data. Without too much work, we have found an effective way to summarize the data. We will use return intervals of 2 percent, beginning with −44% ≤ Rt < −42% (given as “−44% to −42%” in the table) and ending with 52% ≤ Rt ≤ 54%. Table 3 shows the frequency distribution for the annual total returns on the S&P 500.
Table 3 includes three other useful ways to present data, which we can compute once we have established the frequency distribution: the relative frequency, the cumulative frequency (also called the cumulative absolute frequency), and the cumulative relative frequency.
Definition of Relative Frequency. The relative frequency is the absolute frequency of each interval divided by the total number of observations.
The cumulative relative frequency cumulates (adds up) the relative frequencies as we move from the first to the last interval. It tells us the fraction of observations that are less than the upper limit of each interval. Examining the frequency distribution given in Table 3, we see that the first return interval, −44 percent to −42 percent, has one observation; its relative frequency is 1/87 or 1.15 percent. The cumulative frequency for this interval is 1 because only one observation is less than −42 percent. The cumulative relative frequency is thus 1/87 or 1.15 percent. The next return interval has zero observations; therefore, its cumulative frequency is 0 plus 1 and its cumulative relative frequency is 1.15 percent (the cumulative relative frequency from the previous interval). We can find the other cumulative frequencies by adding the (absolute) frequency to the previous cumulative frequency. The cumulative frequency, then, tells us the number of observations that are less than the upper limit of each return interval.
TABLE 3 Frequency Distribution for the Annual Total Return on the S&P 500, 1926–2012
Return Interval (%)
Frequency Relative Frequency
(%)
Cumulative Frequency
Cumulative Relative Frequency
(%)
Return Interval (%)
Frequency Relative Frequency
(%)
−44.0 to −42.0 1 1.15 1 1.15
4.0 to 6.0 6 6.90
−42.0 to −40.0 0 0.00 1 1.15
6.0 to 8.0 4 4.60
−40.0 to −38.0 0 0.00 1 1.15
8.0 to 10.0 1 1.15
−38.0 to −36.0 1 1.15 2 2.30
10.0 to 12.0 4 4.60
−36.0 to −34.0 1 1.15 3 3.45
12.0 to 14.0 1 1.15
−34.0 to −32.0 0 0.00 3 3.45
14.0 to 16.0 4 4.60
−32.0 to −30.0 0 0.00 3 3.45
16.0 to 18.0 2 2.30
−30.0 to −28.0 0 0.00 3 3.45
18.0 to 20.0 6 6.90
−28.0 to −26.0
1 1.15 4 4.60 20.0 to 22.0
3 3.45
−26.0 to −24.0 1 1.15 5 5.75
22.0 to 24.0 5 5.75
−24.0 to −22.0 1 1.15 6 6.90
24.0 to 26.0 2 2.30
−22.0 to −20.0 0 0.00 6 6.90
26.0 to 28.0 2 2.30
−20.0 to −18.0 0 0.00 6 6.90
28.0 to 30.0 2 2.30
−18.0 to −16.0 0 0.00 6 6.90
30.0 to 32.0 5 5.75
−16.0 to −14.0 1 1.15 7 8.05
32.0 to 34.0 4 4.60
−14.0 to −12.0 0 0.00 7 8.05
34.0 to 36.0 0 0.00
−12.0 to −10.0 4 4.60 11 12.64
36.0 to 38.0 4 4.60
−10.0 to −8.0 7 8.05 18 20.69
38.0 to 40.0 0 0.00
−8.0 to −6.0 1 1.15 19 21.84
40.0 to 42.0 0 0.00
−6.0 to −4.0 1 1.15 20 22.99
42.0 to 44.0 2 2.30
−4.0 to −2.0 1 1.15 21 24.14
44.0 to 46.0 0 0.00
−2.0 to 0.0 3 3.45 24 27.59
46.0 to 48.0 1 1.15
0.0 to 2.0 2 2.30 26 29.89
48.0 to 50.0 0 0.00
2.0 to 4.0 1 1.15 27 31.03
50.0 to 52.0 0 0.00
52.0 to 54.0 2 2.30
Note: The lower class limit is the weak inequality (≤) and the upper class limit is the strong inequality (<). Cumulative relative frequency totals reflect calculations using
full precision, with results rounded to two decimal places.
Source: Ibbotson Associates.
As Table 3 shows, return intervals have frequencies from 0 to 7 in this sample. The interval encompassing returns between −10 percent and −8 percent (−10% ≤ Rt < −8%) has the most observations, seven. Next most frequent are returns between 4 percent and 6 percent (4% ≤ Rt < 6%) and between18 percent and 20 percent (18% ≤ Rt < 20%), with six observations in each interval. From the cumulative frequency column, we see that the number of negative returns is 24. The number of positive returns must then be equal to 87 − 24, or 63. We can express the number of positive and negative outcomes as a percentage of the total to get a sense of the risk inherent in investing in the stock market. During the 87-year period, the S&P 500 had negative annual returns 27.6 percent of the time (that is, 24/87). This result appears in the fifth column of Table 3, which reports the cumulative relative frequency.
The frequency distribution gives us a sense of not only where most of the observations lie but also whether the distribution is evenly distributed, lopsided, or peaked. In the case of the S&P 500, we can see that more than half of the outcomes are positive and most of those annual returns are larger than 10 percent. (Only 14 of the 63 positive annual returns—about 22 percent—were between 0 and 10 percent.)
Table 3 permits us to make an important further point about the choice of the number of intervals related to equity returns in particular. From the frequency distribution in Table 3, we can see that only six outcomes fall between −44 percent to −16 percent and only five outcomes fall between 38 percent to 54 percent. Stock return data are frequently characterized by a few very large or small outcomes. We could have collapsed the return intervals in the tails of the frequency distribution by choosing a smaller value of k, but then we would have lost the information about how extremely poorly or well the stock market had performed. A risk manager may need to know the worst possible outcomes and thus may want to have detailed information on the tails (the extreme values). A frequency distribution with a relatively large value of k is useful for that. A portfolio manager or analyst may be equally interested in detailed information on the tails; however, if the manager or analyst wants a picture only of where most of the observations lie, he might prefer to use an interval width of 4 percent (25 intervals beginning at −44 percent), for example.
The frequency distribution for monthly returns on the S&P 500 looks quite different from that for annual returns. The monthly return series from January 1926 to December 2012 has 1,044 observations. Returns range from a minimum of approximately −30 percent to a maximum of approximately +43 percent. With such a large quantity of monthly data we must summarize to get a sense of the
distribution, and so we group the data into 37 equally spaced return intervals of 2 percent. The gains from summarizing in this way are substantial. Table 4 presents the resulting frequency distribution. The absolute frequencies appear in the second column, followed by the relative frequencies. The relative frequencies are rounded to two decimal places. The cumulative absolute and cumulative relative frequencies appear in the fourth and fifth columns, respectively.
TABLE 4 Frequency Distribution for the Monthly Total Return on the S&P 500, January 1926 to December 2012
Return Interval (%)
Absolute Frequency
Relative Frequency
(%)
Cumulative Absolute Frequency
Cumulative Relative Frequency (%)
−30.0 to −28.0 1 0.10 1 0.10
−28.0 to −26.0 0 0.00 1 0.10
−26.0 to −24.0 1 0.10 2 0.19
−24.0 to −22.0 1 0.10 3 0.29
−22.0 to −20.0 2 0.19 5 0.48
−20.0 to −18.0 2 0.19 7 0.67
−18.0 to −16.0 3 0.29 10 0.96
−16.0 to −14.0 2 0.19 12 1.15
−14.0 to −12.0 6 0.57 18 1.72
−12.0 to −10.0 7 0.67 25 2.39
−10.0 to −8.0 23 2.20 48 4.60
−8.0 to −6.0 34 3.26 82 7.85 −6.0 to −4.0 59 5.65 141 13.51
−4.0 to −2.0 98 9.39 239 22.89 −2.0 to 0.0 157 15.04 396 37.93 0.0 to 2.0 220 21.07 616 59.00 2.0 to 4.0 173 16.57 789 75.57 4.0 to 6.0 137 13.12 926 88.70 6.0 to 8.0 63 6.03 989 94.73 8.0 to 10.0 25 2.39 1,014 97.13 10.0 to 12.0 15 1.44 1,029 98.56 12.0 to 14.0 6 0.57 1,035 99.14 14.0 to 16.0 2 0.19 1,037 99.33 16.0 to 18.0 3 0.29 1,040 99.62 18.0 to 20.0 0 0.00 1,040 99.62 20.0 to 22.0 0 0.00 1,040 99.62 22.0 to 24.0 0 0.00 1,040 99.62 24.0 to 26.0 1 0.10 1,041 99.71 26.0 to 28.0 0 0.00 1,041 99.71 28.0 to 30.0 0 0.00 1,041 99.71 30.0 to 32.0 0 0.00 1,041 99.71 32.0 to 34.0 0 0.00 1,041 99.71 34.0 to 36.0 0 0.00 1,041 99.71
36.0 to 38.0 0 0.00 1,041 99.71 38.0 to 40.0 2 0.19 1,043 99.90 40.0 to 42.0 0 0.00 1,043 99.90 42.0 to 44.0 1 0.10 1,044 100.00
Note: The lower class limit is the weak inequality (≤) and the upper class limit is the strong inequality (<). The relative frequency is the absolute frequency or cumulative
frequency divided by the total number of observations. Cumulative relative frequency totals reflect calculations using full precision, with results rounded to two
decimal places.
Source: Ibbotson Associates.
The advantage of a frequency distribution is evident in Table 4, which tells us that the vast majority of observations (687/1,044 = 66 percent) lie in the four intervals
spanning −2 percent to +6 percent. Altogether, we have 396 negative returns and 648 positive returns. Almost 62 percent of the monthly outcomes are positive. Looking at the cumulative relative frequency in the last column, we see that the interval −2 percent to 0 percent shows a cumulative frequency of 37.93 percent, for an upper return limit of 0 percent. This means that 37.93 percent of the observations lie below the level of 0 percent. We can also see that not many observations are greater than +12 percent or less than −12 percent. Note that the frequency distributions of annual and monthly returns are not directly comparable. On average, we should expect the returns measured at shorter intervals (for example, months) to be smaller than returns measured over longer periods (for example, years).
Next, we construct a frequency distribution of average inflation-adjusted returns over 1900–2010 for 19 major equity markets.
EXAMPLE 2 Constructing a Frequency Distribution
How have equities rewarded investors in different countries in the long run? To answer this question, we could examine the average annual returns directly.9 The worth of a nominal level of return depends on changes in the purchasing power of money, however, and internationally there have been a variety of experiences with price inflation. It is preferable, therefore, to compare the average real or inflation-adjusted returns earned by investors in different countries. Dimson, Marsh, and Staunton (2011) presented authoritative evidence on asset returns in 19 countries for the 111 years 1900–2010. Table 5 excerpts their findings for average inflation-adjusted returns.
TABLE 5 Real (Inflation-Adjusted) Equity Returns: Nineteen Major Equity Markets, 1900–2010
Country Arithmetic Mean (%) Australia 9.1 Belgium 5.1 Canada 7.3 Denmark 6.9 Finland 9.3 France 5.7 Germany 8.1 Ireland 6.4
Italy 6.1 Japan 8.5
Netherlands 7.1 New Zealand 7.6 Norway 7.2
South Africa 9.5 Spain 5.8 Sweden 8.7
Switzerland 6.1 United Kingdom 7.2 United States 8.3
Source: Dimson, Marsh, and Staunton (2011), Table 1.
Table 6 summarizes the data in Table 5 into five intervals spanning 5 percent to 10 percent. With nineteen markets, the relative frequency for the 5.0 to 6.0 percent return interval is calculated as 3/19 = 15.79 percent, for example.
As Table 6 shows, there is substantial variation internationally of average real equity returns. More than a quarter of the observations fall in the 7.0 to 8.0 percent interval, which has a relative frequency of 26.32 percent. Either three or four observations fall in each of the other four intervals.
TABLE 6 Frequency Distribution of Average Real Equity Returns
Return Interval (%)
Absolute Frequency
Relative Frequency
(%)
Cumulative Absolute Frequency
Cumulative Relative Frequency (%)
5.0 to 6.0 3 15.79 3 15.79 6.0 to 7.0 4 21.05 7 36.84 7.0 to 8.0 5 26.32 12 63.16 8.0 to 9.0 4 21.05 16 84.21 9.0 to 10 3 15.79 19 100.00
4. THE GRAPHIC PRESENTATION OF DATA A graphical display of data allows us to visualize important characteristics quickly. For example, we may see that the distribution is symmetrically shaped, and this finding may influence which probability distribution we use to describe the data. In this section, we discuss the histogram, the frequency polygon, and the cumulative frequency distribution as methods for displaying data graphically. We construct all of these graphic presentations with the information contained in the frequency distribution of the S&P 500 shown in either Table 3 or Table 4.
4.1. The Histogram A histogram is the graphical equivalent of a frequency distribution.
Definition of Histogram. A histogram is a bar chart of data that have been grouped into a frequency distribution.
The advantage of the visual display is that we can see quickly where most of the observations lie. To see how a histogram is constructed, look at the return interval 18% ≤ Rt < 20% in Table 3. This interval has an absolute frequency of 6. Therefore, we erect a bar or rectangle with a height of 6 over that return interval on the horizontal axis. Continuing with this process for all other return intervals yields a histogram. Figure 1 presents the histogram of the annual total return series on the S&P 500 from 1926 to 2012.
FIGURE 1 Histogram of S&P 500 Annual Total Returns: 1926 to 2012
Note: Because of space limitations, only every other return interval is labeled below the horizontal axis.
Source: Ibbotson Associates
FIGURE 2 Histogram of S&P 500 Monthly Total Returns: January 1926 to December 2012
Source: Ibbotson Associates
In the histogram in Figure 1, the height of each bar represents the absolute frequency for each return interval. The return interval −10% ≤ Rt < −8% has a frequency of 7 and is represented by the tallest bar in the histogram. Because there are no gaps between the interval limits, there are no gaps between the bars of the histogram. Many of the return intervals have zero frequency; therefore, they have no height in the histogram.
Figure 2 presents the histogram for the distribution of monthly returns on the S&P 500. Somewhat more symmetrically shaped than the histogram of annual returns shown in Figure 1, this histogram also appears more bell-shaped than the distribution of annual returns.
4.2. The Frequency Polygon and the Cumulative Frequency Distribution Two other graphical tools for displaying data are the frequency polygon and the cumulative frequency distribution. To construct a frequency polygon, we plot the
midpoint of each interval on the x-axis and the absolute frequency for that interval on the y-axis; we then connect neighboring points with a straight line. Figure 3 shows the frequency polygon for the 1,044 monthly returns for the S&P 500 from January 1926 to December 2012.
FIGURE 3 Frequency Polygon of S&P 500 Monthly Total Returns: January 1926 to December 2012
Source: Ibbotson Associates
In Figure 3, we have replaced the bars in the histogram with points connected with straight lines. For example, the return interval 0 percent to 2 percent has an absolute frequency of 220. In the frequency polygon, we plot the return-interval midpoint of 1 percent and a frequency of 220. We plot all other points in a similar way.10 This form of visual display adds a degree of continuity to the representation of the distribution.
Another form of line graph is the cumulative frequency distribution. Such a graph can plot either the cumulative absolute or cumulative relative frequency against the upper interval limit. The cumulative frequency distribution allows us to see how many or what percent of the observations lie below a certain value. To construct the cumulative frequency distribution, we graph the returns in the fourth or fifth column of Table 4 against the upper limit of each return interval. Figure 4 presents a graph of the cumulative absolute distribution for the monthly returns on the S&P 500.
Notice that the cumulative distribution tends to flatten out when returns are extremely negative or extremely positive. The steep slope in the middle of Figure 4 reflects the fact that most of the observations lie in the neighborhood of −2 percent to 6 percent.
FIGURE 4 Cumulative Absolute Frequency Distribution of S&P 500 Monthly Total Returns: January 1926 to December 2012
Source: Ibbotson Associates
We can further examine the relationship between the relative frequency and the cumulative relative frequency by looking at the two return intervals reproduced in Table 7. The first return interval (0 percent to 2 percent) has a cumulative relative frequency of 59 percent. The next return interval (2 percent to 4 percent) has a cumulative relative frequency of 75.57 percent. The change in the cumulative relative frequency as we move from one interval to the next is the next interval’s relative frequency. For instance, as we go from the first return interval (0 percent to 2 percent) to the next return interval (2 percent to 4 percent), the change in the cumulative relative frequency is 75.57% − 59.00% = 16.57%. (Values in the table have been rounded to two decimal places.) The fact that the slope is steep indicates that these frequencies are large. As you can see in the graph of the cumulative distribution, the slope of the curve changes as we move from the first return interval to the last. A fairly small slope for the cumulative distribution for the first few return intervals tells us that these return intervals do not contain many observations.
You can go back to the frequency distribution in Table 4 and verify that the cumulative absolute frequency is only 25 observations (the cumulative relative frequency is 2.39 percent) up to the 10th return interval (−12 percent to −10 percent). In essence, the slope of the cumulative absolute distribution at any particular interval is proportional to the number of observations in that interval.
TABLE 7 Selected Class Frequencies for the S&P 500 Monthly Returns
Return Interval (%)
Absolute Frequency
Relative Frequency
(%)
Cumulative Absolute Frequency
Cumulative Relative Frequency (%)
0.0 to 2.0 220 21.07 616 59.00 2.0 to 4.0 173 16.57 789 75.57
5. MEASURES OF CENTRAL TENDENCY So far, we have discussed methods we can use to organize and present data so that they are more understandable. The frequency distribution of an asset class’s return series, for example, reveals the nature of the risks that investors may encounter in a particular asset class. As an illustration, the histogram for the annual returns on the S&P 500 clearly shows that large positive and negative annual returns are common. Although frequency distributions and histograms provide a convenient way to summarize a series of observations, these methods are just a first step toward describing the data. In this section we discuss the use of quantitative measures that explain characteristics of data. Our focus is on measures of central tendency and other measures of location or location parameters. A measure of central tendency specifies where the data are centered. Measures of central tendency are probably more widely used than any other statistical measure because they can be computed and applied easily. Measures of location include not only measures of central tendency but other measures that illustrate the location or distribution of data.
In the following subsections we explain the common measures of central tendency —the arithmetic mean, the median, the mode, the weighted mean, and the geometric mean. We also explain other useful measures of location, including quartiles, quintiles, deciles, and percentiles.
5.1. The Arithmetic Mean Analysts and portfolio managers often want one number that describes a representative possible outcome of an investment decision. The arithmetic mean is by far the most frequently used measure of the middle or center of data.
Definition of Arithmetic Mean. The arithmetic mean is the sum of the observations divided by the number of observations.
We can compute the arithmetic mean for both populations and samples, known as the population mean and the sample mean, respectively.
5.1.1. The Population Mean
The population mean is the arithmetic mean computed for a population. If we can define a population adequately, then we can calculate the population mean as the arithmetic mean of all the observations or values in the population. For example, analysts examining the fiscal 2013 year-over-year growth in same-store sales of major US wholesale clubs might define the population of interest to include only
three companies: BJ’s Wholesale Club (a private company since 2011), Costco Wholesale Corporation (NASDAQ: COST), and Sam’s Club, part of Wal-Mart Stores (NYSE: WMT).11 As another example, if a portfolio manager ’s investment universe (the set of securities he or she must choose from) is the Nikkei 225 Index, the relevant population is the 225 shares on the First Section of the Tokyo Stock Exchange that compose the Nikkei.
Population Mean Formula. The population mean, μ, is the arithmetic mean value of a population. For a finite population, the population mean is
(2)
where N is the number of observations in the entire population and Xi is the ith observation.
The population mean is an example of a parameter. The population mean is unique; that is, a given population has only one mean. To illustrate the calculation, we can take the case of the population mean of profit as a percentage of revenue of US companies running major wholesale clubs for 2012. During the year, profit as a percentage of revenue for BJ’s Wholesale club, COST, and WMT was 0.9 percent, 1.6 percent, and 3.5 percent, respectively, according to the Fortune 500 list for 2012. Thus the population mean profit as a percentage of revenue was μ = (0.9 + 1.6 + 3.5)/3 = 6/3 = 2 percent.
5.1.2. The Sample Mean
The sample mean is the arithmetic mean computed for a sample. Many times we cannot observe every member of a set; instead, we observe a subset or sample of the population. The concept of the mean can be applied to the observations in a sample with a slight change in notation.
Sample Mean Formula. The sample mean or average, (read “X-bar”), is the arithmetic mean value of a sample:
(3)
where n is the number of observations in the sample.
Equation 3 tells us to sum the values of the observations (Xi) and divide the sum by the number of observations. For example, if a sample of price-to-earnings (P/E) multiples for six publicly traded companies contains the values 35, 30, 22, 18, 15, and 12, the sample mean P/E is 132/6 = 22. The sample mean is also called the arithmetic average.12 As we discussed earlier, the sample mean is a statistic (that is, a descriptive measure of a sample).
Means can be computed for individual units or over time. For instance, the sample might be the 2013 return on equity (ROE) for the 100 companies in the FTSE Eurotop 100, an index of Europe’s 100 largest companies. In this case, we calculate mean ROE in 2013 as an average across 100 individual units. When we examine the characteristics of some units at a specific point in time (such as ROE for the FTSE Eurotop 100), we are examining cross-sectional data. The mean of these observations is called a cross-sectional mean. On the other hand, if our sample consists of the historical monthly returns on the FTSE Eurotop 100 for the past five years, then we have time-series data. The mean of these observations is called a time-series mean. We will examine specialized statistical methods related to the behavior of time series in the reading on times-series analysis.
Next, we show an example of finding the sample mean return for 16 European equity markets for 2012. In this case, the mean is cross-sectional because we are averaging individual country returns.
EXAMPLE 3 Calculating a Cross-Sectional Mean
The MSCI EAFE (Europe, Australasia, and Far East) Index is a free float- adjusted market capitalization index designed to measure developed-market equity performance excluding the United States and Canada.13 As of September 2013, the EAFE consisted of 22 developed market country indexes, including indexes for 16 European markets, 2 Australasian markets (Australia and New Zealand), 3 Far Eastern markets (Hong Kong, Japan, and Singapore), and Israel.
Suppose we are interested in the local currency performance of the 16 European markets in the EAFE in 2012. We want to find the sample mean total return for 2012 across these 16 markets. The return series reported in Table 8 are in local currency (that is, returns are for investors living in the country). Because this return is not stated in any single investor ’s home currency, it is not a return any single investor would earn. Rather, it is an average of returns in local currencies of the 16 countries.
TABLE 8 Total Returns for European Equity Markets, 2012
Market Total Return in Local Currency (%) Austria 20.72 Belgium 33.99 Denmark 28.09 Finland 8.27 France 15.90 Germany 25.24 Greece −2.35 Ireland 2.24 Italy 6.93
Netherlands 15.36 Norway 6.05 Portugal −2.22 Spain −4.76 Sweden 12.66
Switzerland 14.83 United Kingdom 5.93
Source: www.msci.com.
Using the data in Table 8, calculate the sample mean return for the 16 equity markets in 2012.
Solution: The calculation applies Equation 3 to the returns in Table 8: (20.72 + 33.99 + 28.09 + 8.27 + 15.90 + 25.24 − 2.35 + 2.24 + 6.93 + 15.36 + 6.05 − 2.22 − 4.76 + 12.66 + 14.83 + 5.93)/16 = 186.88/16 = 11.68 percent.
In Example 3, we can verify that eight markets had returns less than the mean and eight had returns that were greater. We should not expect any of the actual observations to equal the mean, because sample means provide only a summary of the data being analyzed. Also, although in this example the number of values below the mean is equal to the number of values above the mean, that need not be the case. As an analyst, you will often need to find a few numbers that describe the characteristics of the distribution. The mean is generally the statistic that you will use as a measure of the typical outcome for a distribution. You can then use the mean to compare the performance of two different markets. For example, you
might be interested in comparing the stock market performance of investments in Pacific Rim countries with investments in European countries. You can use the mean returns in these markets to compare investment results.
5.1.3. Properties of the Arithmetic Mean
The arithmetic mean can be likened to the center of gravity of an object. Figure 5 expresses this analogy graphically by plotting nine hypothetical observations on a bar. The nine observations are 2, 4, 4, 6, 10, 10, 12, 12, and 12; the arithmetic mean is 72/9 = 8. The observations are plotted on the bar with various heights based on their frequency (that is, 2 is one unit high, 4 is two units high, and so on). When the bar is placed on a fulcrum, it balances only when the fulcrum is located at the point on the scale that corresponds to the arithmetic mean.
FIGURE 5 Center of Gravity Analogy for the Arithmetic Mean
As analysts, we often use the mean return as a measure of the typical outcome for an asset. As in the example above, however, some outcomes are above the mean and some are below it. We can calculate the distance between the mean and each outcome and call it a deviation. Mathematically, it is always true that the sum of the deviations around the mean equals 0. We can see this by using the definition of the arithmetic
mean shown in Equation 3, multiplying both sides of the equation by n: . The sum of the deviations from the mean can thus be calculated as follows:
Deviations from the arithmetic mean are important information because they indicate risk. The concept of deviations around the mean forms the foundation for the more complex concepts of variance, skewness, and kurtosis, which we will discuss later in this reading.
An advantage of the arithmetic mean over two other measures of central tendency, the median and mode, is that the mean uses all the information about the size and
magnitude of the observations. The mean is also easy to work with mathematically.
A property and potential drawback of the arithmetic mean is its sensitivity to extreme values. Because all observations are used to compute the mean, the arithmetic mean can be pulled sharply upward or downward by extremely large or small observations, respectively. For example, suppose we compute the arithmetic mean of the following seven numbers: 1, 2, 3, 4, 5, 6, and 1,000. The mean is 1,021/7 = 145.86 or approximately 146. Because the magnitude of the mean, 146, is so much larger than that of the bulk of the observations (the first six), we might question how well it represents the location of the data. In practice, although an extreme value or outlier in a financial dataset may only represent a rare value in the population, it may also reflect an error in recording the value of an observation, or an observation generated from a different population from that producing the other observations in the sample. In the latter two cases in particular, the arithmetic mean could be misleading. Perhaps the most common approach in such cases is to report the median in place of or in addition to the mean.14 We discuss the median next.
5.2. The Median A second important measure of central tendency is the median.
Definition of Median. The median is the value of the middle item of a set of items that has been sorted into ascending or descending order. In an odd- numbered sample of n items, the median occupies the (n + 1)/2 position. In an even-numbered sample, we define the median as the mean of the values of items occupying the n/2 and (n + 2)/2 positions (the two middle items).15
Earlier we gave the profit as a percentage of revenue of three wholesale clubs as 0.9, 1.6, and 3.5. With an odd number of observations (n = 3), the median occupies the (n + 1)/2 = 4/2 = 2nd position. The median was 1.6 percent. The value of 1.6 percent is the “middlemost” observation: One lies above it, and one lies below it. Whether we use the calculation for an even-or odd-numbered sample, an equal number of observations lie above and below the median. A distribution has only one median.
A potential advantage of the median is that, unlike the mean, extreme values do not affect it. The median, however, does not use all the information about the size and magnitude of the observations; it focuses only on the relative position of the ranked observations. Calculating the median is also more complex; to do so, we need to order the observations from smallest to largest, determine whether the sample size is even or odd and, on that basis, apply one of two calculations. Mathematicians express this disadvantage by saying that the median is less mathematically tractable
than the mean.
To demonstrate finding the median, we use the data from Example 3, reproduced in Table 9 in ascending order of the 2012 total return for European equities. Because this sample has 16 observations, the median is the mean of the values in the sorted array that occupy the 16/2 = 8th and 18/2 = 9th positions. Finland’s return occupies the eighth position with a return of 8.27 percent, and Sweden’s return occupies the ninth position with a return of 12.66 percent. The median, as the mean of these two returns, is (8.27 + 12.66)/2 = 10.465 percent. Note that the median is not influenced by extremely large or small outcomes. Had Spain’s total return been a much lower value or Belgium’s total return a much larger value, the median would not have changed. Using a context that arises often in practice, Example 4 shows how to use the mean and median in a sample with extreme values.
TABLE 9 Total Returns for European Equity Markets, 2012 (in Ascending Order)
No. Market Total Returnin Local Currency (%) 1 Spain −4.76 2 Greece −2.35 3 Portugal −2.22 4 Ireland 2.24 5 United Kingdom 5.93 6 Norway 6.05 7 Italy 6.93 8 Finland 8.27 9 Sweden 12.66 10 Switzerland 14.83 11 Netherlands 15.36 12 France 15.90 13 Austria 20.72 14 Germany 25.24 15 Denmark 28.09 16 Belgium 33.99
Source: www.msci.com.
EXAMPLE 4 Median and Arithmetic Mean: The Case of the
Price–Earnings Ratio
Suppose a client asks you for a valuation analysis on the seven-stock US common stock portfolio given in Table 10. The stocks are equally weighted in the portfolio. One valuation measure that you use is P/E, the ratio of share price to earnings per share (EPS). Many variations exist for the denominator in the P/E, but you are examining P/E defined as current price divided by the current mean of all analysts’ EPS estimates for the company for the fiscal year 2013 (“Consensus Current EPS” in the table).16 The values in Table 10 are as of 9 September 2013. For comparison purposes, the average current P/E on the companies in the S&P 500 index was 18.80 at that time.
Using the data in Table 10, address the following: 1. Calculate the arithmetic mean P/E.
2. Calculate the median P/E.
3. Evaluate the mean and median P/Es as measures of central tendency for the above portfolio.
TABLE 10 P/Es for a Client Portfolio
Stock ConsensusCurrent EPS Consensus Current P/E
Caterpillar, Inc. (NYSE: CAT) 6.34 13.15 Ford Motor Company (NYSE: F) 1.55 10.97 General Dynamics (NYSE: GD) 6.96 12.15 Green Mountain Coffee Roasters
(NASDAQ: GMCR) 3.25 25.27
McDonald’s Corporation (NYSE: MCD) 5.61 17.16 Qlik Technologies (NASDAQ: QLIK) 0.17 204.82 Questcor Pharmaceuticals (NASDAQ:
QCOR) 4.79 13.94
Note: Consensus current P/E was calculated as price as of 9 September 2013 divided by consensus EPS as of the same date.
Source: www.nasdaq.com.
Solution to 1: The mean P/E is (13.15 + 10.97 + 12.15 + 25.27 + 17.16 + 204.82 + 13.94)/7 = 297.46/7 = 42.49.
Solution to 2: The P/Es listed in ascending order are:
10.97 12.15 13.15 13.94 17.16 25.27 204.82
The sample has an odd number of observations with n = 7, so the median occupies the (n + 1)/2 = 8/2 = 4th position in the sorted list. Therefore, the median P/E is 13.94.
Solution to 3: Qlik Technologies’ P/E of approximately 205 tremendously influences the value of the portfolio’s arithmetic mean P/E. The mean P/E of about 42 is much larger than the P/E of six of the seven stocks in the portfolio. The mean P/E also misleadingly suggests an orientation to stocks with high P/Es. The mean P/E of the stocks excluding Qlik Technologies, or excluding the largest-and smallest-P/E stocks (Qlik Technologies and Ford Motor Company), is below the average P/E of 18.80 for the companies in the S&P 500 Index. The median P/E of 13.94 appears to better represent the central tendency of the P/Es.
It frequently happens that when a company’s EPS is quite low—at a low point in the business cycle, for example—its P/E is extremely high. The high P/E in those circumstances reflects an anticipated future recovery of earnings. Extreme P/E values need to be investigated and handled with care. For reasons related to this example, analysts often use the median of price multiples to characterize the valuation of industry groups.
5.3. The Mode The third important measure of central tendency is the mode.
Definition of Mode. The mode is the most frequently occurring value in a distribution.17
A distribution can have more than one mode, or even no mode. When a distribution has one most frequently occurring value, the distribution is said to be unimodal. If a distribution has two most frequently occurring values, then it has two modes and we say it is bimodal. If the distribution has three most frequently occurring values, then it is trimodal. When all the values in a data set are different, the distribution has no mode because no value occurs more frequently than any other value.
Stock return data and other data from continuous distributions may not have a modal outcome. When such data are grouped into intervals, however, we often find an interval (possibly more than one) with the highest frequency: the modal interval (or intervals). For example, the frequency distribution for the monthly returns on the S&P 500 has a modal interval of 0 percent to 2 percent, as shown in Figure 2; this return interval has 220 observations out of a total of 1,044. The modal interval always has the highest bar in the histogram.
The mode is the only measure of central tendency that can be used with nominal data. When we categorize mutual funds into different styles and assign a number to each style, the mode of these categorized data is the most frequent mutual fund style.
5.4. Other Concepts of Mean Earlier we explained the arithmetic mean, which is a fundamental concept for describing the central tendency of data. Other concepts of mean are very important in investments, however. In the following, we discuss such concepts.
EXAMPLE 5 Calculating a Mode
Table 11 gives the credit ratings on senior unsecured debt as of September 2002 of nine US department stores rated by Moody’s Investors Service. In descending order of credit quality (increasing expected probability of default), Moody’s ratings are Aaa, Aa1, Aa2, Aa3, A1, A2, A3, Baa1, Baa2, Baa3, Ba1, Ba2, Ba3, B1, B2, B3, Caa1, Caa2, Caa3, Ca, and C.18
Using the data in Table 11, address the following concerning the senior unsecured debt of US department stores:
1. State the modal credit rating.
2. State the median credit rating.
TABLE 11 Senior Unsecured Debt Ratings: US Department Stores, September 2013
Company Credit Rating Bon-Ton Stores Inc. B3
Dillards, Inc. Ba2 Kohl’s Corporation Baa1
Macy’s, Inc. Baa3 Neiman Marcus Group, Inc. B2
Nordstrom, Inc. Baa1 Penney, JC, Company, Inc. Caa1
Saks Incorporated Ba2 Sears, Roebuck and Co. B3
Source: www.moodys.com.
TABLE 12 Senior Unsecured Debt Ratings: US Department Stores, Distribution of Credit Ratings
Credit Rating Frequency Baa1 2 Baa3 1 Ba2 2 B2 1 B3 2 Caa1 1
Solution to 1: The group of companies represents six distinct credit ratings, ranging from Baa1 to Caa1. To make our task easy, we first organize the ratings into a frequency distribution.
Credit ratings Baa1, Ba2, and B3 have a frequency of 2, and the other three ratings have a frequency of 1. Therefore, the credit rating of US department stores in September 2013 was trimodal, with Baa1, Ba2, and B3 being the three modes. Moody’s considers bonds rated Baa to be of moderate credit risk, Ba to be of substantial credit risk, and B to be of high credit risk.
Solution to 2: For the group n = 9, an odd number. The group’s median occupies the (n + 1)/2 = 10/2 = 5th position. We see from Table 12 that Ba2 occupies the fifth position. Therefore the median credit rating at September 2013 was Ba2.
5.4.1. The Weighted Mean
The concept of weighted mean arises repeatedly in portfolio analysis. In the arithmetic mean, all observations are equally weighted by the factor 1/n (or 1/N). In working with portfolios, we need the more general concept of weighted mean to
allow different weights on different observations.
To illustrate the weighted mean concept, an investment manager with $100 million to invest might allocate $70 million to equities and $30 million to bonds. The portfolio has a weight of 0.70 on stocks and 0.30 on bonds. How do we calculate the return on this portfolio? The portfolio’s return clearly involves an averaging of the returns on the stock and bond investments. The mean that we compute, however, must reflect the fact that stocks have a 70 percent weight in the portfolio and bonds have a 30 percent weight. The way to reflect this weighting is to multiply the return on the stock investment by 0.70 and the return on the bond investment by 0.30, then sum the two results. This sum is an example of a weighted mean. It would be incorrect to take an arithmetic mean of the return on the stock and bond investments, equally weighting the returns on the two asset classes.
Consider a portfolio invested in Canadian stocks and bonds. The stock component of the portfolio includes the RBC Canadian Index Fund, which tracks the performance of the S&P/TSX Composite Total Return Index. The bond component of the portfolio includes the RBC Bond Fund, which invests in high-quality fixed- income securities issued by Canadian governments and corporations. The portfolio manager allocates 60 percent of the portfolio to the Canadian stock fund and 40 percent to the Canadian bond fund. Table 13 presents total returns for these funds from 2008 to 2012.
TABLE 13 Returns for Canadian Equity and Bond Funds], 2008–2012
Year Equity Fund (%) Bond Fund (%) 2008 −33.1 −0.1 2009 34.1 11.0 2010 16.8 6.4 2011 −9.2 8.4 2012 6.4 3.8
Source: funds.rbcgam.com.
Weighted Mean Formula. The weighted mean (read “X-bar sub-w”), for a set of observations X1, X2, …, Xn with corresponding weights of w1, w2, …, wn is computed as
(4)
where the sum of the weights equals 1; that is, .
In the context of portfolios, a positive weight represents an asset held long and a negative weight represents an asset held short.19
The return on the portfolio under consideration is the weighted average of the return on the Canadian stock fund and the Canadian bond fund (the weight of the stock fund is 0.60; that of the bond fund is 0.40). We find, using Equation 4, that
It should be clear that the correct mean to compute in this example is the weighted mean and not the arithmetic mean. If we had computed the arithmetic mean for 2008, we would have calculated a return equal to ½(−33.1%) + ½(−0.1%) = (−33.1% − 0.1%)/2 = −16.6%. Given that the portfolio manager invested 60 percent in stocks and 40 percent in bonds, the arithmetic mean would underweight the investment in stocks and overweight the investment in bonds, resulting in a number for portfolio return that is too high by 3.3 percentage points (−16.6% − (−19.9%) = −16.6% + 19.9%).
Now suppose that the portfolio manager maintains constant weights of 60 percent in stocks and 40 percent in bonds for all five years. This method is called a constant- proportions strategy. Because value is price multiplied by quantity, price fluctuation causes portfolio weights to change. As a result, the constant-proportions strategy requires rebalancing to restore the weights in stocks and bonds to their target levels. Assuming that the portfolio manager is able to accomplish the necessary rebalancing, we can compute the portfolio returns in 2009, 2010, 2011, and 2012 with Equation 4 as follows:
We can now find the time-series mean of the returns for 2008 through 2012 using Equation 3 for the arithmetic mean. The time-series mean total return for the portfolio is (−19.9 + 24.9 + 12.6 − 2.2 + 5.4)/5 = 20.8/5 = 4.2 percent.
Instead of calculating the portfolio time-series mean return from portfolio annual returns, we can calculate the arithmetic mean stock and bond fund returns for the five years and then apply the portfolio weights of 0.60 and 0.40, respectively, to
those values. The mean stock fund return is (−33.1 + 34.1 + 16.8 − 9.2 + 6.4)/5 = 15.0/5 = 3.0 percent. The mean bond fund return is (−0.1 + 11.0 + 6.4 + 8.4 + 3.8)/ 5 = 29.5/5 = 5.9 percent. Therefore, the mean total return for the portfolio is 0.60(3.0) + 0.40(5.9) = 4.2 percent, which agrees with our previous calculation.
EXAMPLE 6 Portfolio Return as a Weighted Mean
Table 14 gives information on the asset allocation of the pension plan of the Canadian Broadcasting Corporation in 2012 as well as the returns on these asset classes in 2012.20
TABLE 14 Asset Allocation for the Pension Plan of the Canadian Broadcasting Corporation in 2012
Asset Class Asset Allocation(Weight) Asset Class Return
(%) Cash and short-term
investments 3.8 1.3
Nominal bonds 33.7 6.6 Real return bonds 14.8 2.9 Canadian equities 10.4 8.8 Global equities 21.4 13.3
Strategic investments 15.8 9.5 Bond overlay 0.1 0.8
Source: Canadian Broadcasting Corporation Pension Plan, 2012 Annual Report
Using the information in Table 14, calculate the mean return earned by the pension plan in 2012.
Solution: Converting the percent asset allocation to decimal form, we find the mean return as a weighted average of the asset class returns. We have
The previous examples illustrate the general principle that a portfolio return is a weighted sum. Specifically, a portfolio’s return is the weighted average of the returns on the assets in the portfolio; the weight applied to each asset’s return is the fraction of the portfolio invested in that asset.
Market indexes are computed as weighted averages. For market-capitalization indexes such as the CAC-40 in France or the TOPIX in Japan or the S&P 500 in the United States, each included stock receives a weight corresponding to its outstanding market value divided by the total market value of all stocks in the index.
Our illustrations of weighted mean use past data, but they might just as well use forward-looking data. When we take a weighted average of forward-looking data, the weighted mean is called expected value. Suppose we make one forecast for the year-end level of the S&P 500 assuming economic expansion and another forecast for the year-end level of the S&P 500 assuming economic contraction. If we multiply the first forecast by the probability of expansion and the second forecast by the probability of contraction and then add these weighted forecasts, we are calculating the expected value of the S&P 500 at year-end. If we take a weighted average of possible future returns on the S&P 500, we are computing the S&P 500’s expected return. The probabilities must sum to 1, satisfying the condition on the weights in the expression for weighted mean, Equation 4.
5.4.2. The Geometric Mean
The geometric mean is most frequently used to average rates of change over time or to compute the growth rate of a variable. In investments, we frequently use the geometric mean to average a time series of rates of return on an asset or a portfolio, or to compute the growth rate of a financial variable such as earnings or sales. In the reading on the time value of money, for instance, we computed a sales growth rate (Example 17). That growth rate was a geometric mean. Because of the subject’s importance, in a later section we will return to the use of the geometric mean and offer practical perspectives on its use. The geometric mean is defined by the following formula.
Geometric Mean Formula. The geometric mean, G, of a set of observations X1, X2, …, Xn is
(5)
Equation 5 has a solution, and the geometric mean exists, only if the product under the radical sign is non-negative. We impose the restriction that all the observations
Xi in Equation 5 are greater than or equal to zero. We can solve for the geometric mean using Equation 5 directly with any calculator that has an exponentiation key (on most calculators, yx). We can also solve for the geometric mean using natural logarithms. Equation 5 can also be stated as
or as
When we have computed ln G, then G = elnG (on most calculators, the key for this step is ex).
Risky assets can have negative returns up to −100 percent (if their price falls to zero), so we must take some care in defining the relevant variables to average in computing a geometric mean. We cannot just use the product of the returns for the sample and then take the nth root because the returns for any period could be negative. We must redefine the returns to make them positive. We do this by adding 1.0 to the returns expressed as decimals. The term (1 + Rt) represents the year- ending value relative to an initial unit of investment at the beginning of the year. As long as we use (1 + Rt), the observations will never be negative because the biggest negative return is −100 percent. The result is the geometric mean of 1 + Rt; by then subtracting 1.0 from this result, we obtain the geometric mean of the individual returns Rt. For example, the returns on RBC Canadian Index Fund during the 2008– 2012 period were given in Table 13 as −0.331, 0.341, 0.168, −0.092, and 0.064, putting the returns into decimal form. Adding 1.0 to those returns produces 0.669, 1.341, 1.168, 0.908, and 1.064. Using Equation 5 we have = = 1.002455.
This number is 1 plus the geometric mean rate of return. Subtracting 1.0 from this result, we have 1.002455 − 1.0 = 0.002455 or approximately 0.25 percent. The geometric mean return of RBC Canadian Index Fund during the 2008–2012 period was 0.25 percent.
An equation that summarizes the calculation of the geometric mean return, RG, is a slightly modified version of Equation 5 in which the Xi represent “1 + return in decimal form.” Because geometric mean returns use time series, we use a subscript t indexing time as well.
which leads to the following formula.
Geometric Mean Return Formula. Given a time series of holding period returns Rt, t = 1, 2, …, T, the geometric mean return over the time period spanned by the returns R1 through RT is
(6)
We can use Equation 6 to solve for the geometric mean return for any return data series. Geometric mean returns are also referred to as compound returns. If the returns being averaged in Equation 6 have a monthly frequency, for example, we may call the geometric mean monthly return the compound monthly return. The next example illustrates the computation of the geometric mean while contrasting the geometric and arithmetic means.
EXAMPLE 7 Geometric and Arithmetic Mean Returns (1)
As a mutual fund analyst, you are examining, as of early 2013, the most recent five years of total returns for two US large-cap value equity mutual funds.
Based on the data in Table 15, address the following: 1. Calculate the geometric mean return of SLASX.
2. Calculate the arithmetic mean return of SLASX and contrast it to the fund’s geometric mean return.
3. Calculate the geometric mean return of PRFDX.
4. Calculate the arithmetic mean return of PRFDX and contrast it to the fund’s geometric mean return.
Solution to 1: Converting the returns on SLASX to decimal form and adding 1.0 to each return produces 0.6056, 1.3164, 1.1253, 0.9565, and 1.1282. We use Equation 6 to find SLASX’s geometric mean return:
Solution to 2: For SLASX, = (−39.44 + 31.64 + 12.53 − 4.35 +12.82)/ 5 = 13.20/5 = 2.64%. The arithmetic mean return for SLASX exceeds the geometric mean return by 2.64 − (−0.65) = 3.29% or 329 basis points.
TABLE 15 Total Returns for Two Mutual Funds, 2008–2012 (Repeated)
Year Selected AmericanShares(SLASX) T. Rowe Price Equity Income(PRFDX)
2008 −39.44% −35.75% 2009 31.64 25.62 2010 12.53 15.15 2011 −4.35 −0.72 2012 12.82 17.25
Source: performance.morningstar.com.
Solution to 3: Converting the returns on PRFDX to decimal form and adding 1.0 to each return produces 0.6425, 1.2562, 1.1515, 0.9928, and 1.1725. We use Equation 6 to find PRFDX’s geometric mean return:
Solution to 4: PRFDX, = (−35.75 + 25.62 + 15.15 − 0.72 + 17.25)/5 = 21.55/5 = 4.31%. The arithmetic mean for PRFDX exceeds the geometric mean return by 4.31 − 1.59 = 2.72% or 272 basis points. The table below summarizes the findings.
In Example 7, for both mutual funds, the geometric mean return was less than the arithmetic mean return. In fact, the geometric mean is always less than or equal to the arithmetic mean.21 The only time that the two means will be equal is when there is no variability in the observations—that is, when all the observations in the series are the same.22 In Example 7, there was variability in the funds’ returns; thus for both funds, the geometric mean was strictly less than the arithmetic mean. In general, the difference between the arithmetic and geometric means increases with
the variability in the period-by-period observations.23 This relationship is also illustrated by Example 7. Casual inspection suggests that the returns of SLASX are somewhat more variable than those of PRFDX, and consequently, the spread between the arithmetic and geometric mean returns is larger for SLASX (329 basis points) than for PRFDX (272 basis points).24 Arithmetic and geometric returns need not always rank funds similarly, however, in this example, PRFDX has both higher arithmetic and geometric mean returns than SLASX. However, the difference between the geometric mean returns of the two funds (2.24%) is greater than the difference between the arithmetic mean returns of the two funds (1.67%). How should the analyst interpret these results?
TABLE 16 Mutual Fund Arithmetic and Geometric Mean Returns: Summary of Findings
Fund Arithmetic Mean (%) Geometric Mean (%) SLASX 2.64 −0.65 PRFDX 4.31 1.59
The geometric mean return represents the growth rate or compound rate of return on an investment. One dollar invested in SLASX at the beginning of 2008 would have grown (or, in this case, decreased) to (0.6056)(1.3164)(1.1253)(0.9565)(1.1282) = $0.9681, which is equal to 1 plus the geometric mean return compounded over five periods: [1 + (−0.006466)]5 = (0.993534)5 = $0.9681, confirming that the geometric mean is the compound rate of return. For PRFDX, one dollar would have grown to a larger amount, (0.6425)(1.2562)(1.1515)(0.9928)(1.1725) = $1.0819, equal to (1.015861)5. With its focus on the profitability of an investment over a multiperiod horizon, the geometric mean is of key interest to investors. The arithmetic mean return, focusing on average single-period performance, is also of interest. Both arithmetic and geometric means have a role to play in investment management, and both are often reported for return series. Example 8 highlights these points in a simple context.
EXAMPLE 8 Geometric and Arithmetic Mean Returns (2)
A hypothetical investment in a single stock initially costs €100. One year later, the stock is trading at €200. At the end of the second year, the stock price falls back to the original purchase price of €100. No dividends are paid during the two-year period. Calculate the arithmetic and geometric mean annual returns.
Solution: First, we need to find the Year 1 and Year 2 annual returns with
Equation 1.
The arithmetic mean of the annual returns is (100% − 50%)/2 = 25%.
Before we find the geometric mean, we must convert the percentage rates of return to (1 + Rt). After this adjustment, the geometric mean from Equation 6 is
–1 = 0 percent.
The geometric mean return of 0 percent accurately reflects that the ending value of the investment in Year 2 equals the starting value in Year 1. The compound rate of return on the investment is 0 percent. The arithmetic mean return reflects the average of the one-year returns.
5.4.3. The Harmonic Mean
The arithmetic mean, the weighted mean, and the geometric mean are the most frequently used concepts of mean in investments. A fourth concept, the harmonic mean, , is appropriate in a limited number of applications.25
Harmonic Mean Formula. The harmonic mean of a set of observations X1, X2, …, Xn is
(7)
The harmonic mean is the value obtained by summing the reciprocals of the observations—terms of the form 1/Xi—then averaging that sum by dividing it by the number of observations n, and, finally, taking the reciprocal of the average.
The harmonic mean may be viewed as a special type of weighted mean in which an observation’s weight is inversely proportional to its magnitude. The harmonic mean is a relatively specialized concept of the mean that is appropriate when averaging ratios (“amount per unit”) when the ratios are repeatedly applied to a fixed quantity to yield a variable number of units. The concept is best explained through an illustration. A well-known application arises in the investment strategy known as cost averaging, which involves the periodic investment of a fixed amount of money. In this application, the ratios we are averaging are prices per share at purchases dates, and we are applying those prices to a constant amount of money to yield a variable number of shares.
Suppose an investor purchases €1,000 of a security each month for n = 2 months. The share prices are €10 and €15 at the two purchase dates. What is the average price paid for the security?
In this example, in the first month we purchase €1,000/€10 = 100 shares and in the second month we purchase €1,000/€15 = 66.67, or 166.67 shares in total. Dividing the total euro amount invested, €2,000, by the total number of shares purchased, 166.67, gives an average price paid of €2,000/166.67 = €12. The average price paid is in fact the harmonic mean of the asset’s prices at the purchase dates. Using Equation 7, the harmonic mean price is 2/[(1/10) + (1/15)] = €12. The value €12 is less than the arithmetic mean purchase price (€10 + €15)/2 = €12.5. However, we could find the correct value of €12 using the weighted mean formula, where the weights on the purchase prices equal the shares purchased at a given price as a proportion of the total shares purchased. In our example, the calculation would be (100/166.67)€10.00 + (66.67/166.67)€15.00 = €12. If we had invested varying amounts of money at each date, we could not use the harmonic mean formula. We could, however, still use the weighted mean formula in a manner similar to that just described.
A mathematical fact concerning the harmonic, geometric, and arithmetic means is that unless all the observations in a data set have the same value, the harmonic mean is less than the geometric mean, which in turn is less than the arithmetic mean. In the illustration given, the harmonic mean price was indeed less than the arithmetic mean price.
6. OTHER MEASURES OF LOCATION: QUANTILES Having discussed measures of central tendency, we now examine an approach to describing the location of data that involves identifying values at or below which specified proportions of the data lie. For example, establishing that 25, 50, and 75 percent of the annual returns on a portfolio are at or below the values −0.05, 0.16, and 0.25, respectively, provides concise information about the distribution of portfolio returns. Statisticians use the word quantile (or fractile) as the most general term for a value at or below which a stated fraction of the data lies. In the following, we describe the most commonly used quantiles—quartiles, quintiles, deciles, and percentiles—and their application in investments.
6.1. Quartiles, Quintiles, Deciles, and Percentiles We know that the median divides a distribution in half. We can define other dividing lines that split the distribution into smaller sizes. Quartiles divide the distribution into quarters, quintiles into fifths, deciles into tenths, and percentiles into hundredths. Given a set of observations, the yth percentile is the value at or below which y percent of observations lie. Percentiles are used frequently, and the other measures can be defined with respect to them. For example, the first quartile (Q1) divides a distribution such that 25 percent of the observations lie at or below it; therefore, the first quartile is also the 25th percentile. The second quartile (Q2) represents the 50th percentile, and the third quartile (Q3) represents the 75th percentile because 75 percent of the observations lie at or below it.
When dealing with actual data, we often find that we need to approximate the value of a percentile. For example, if we are interested in the value of the 75th percentile, we may find that no observation divides the sample such that exactly 75 percent of the observations lie at or below that value. The following procedure, however, can help us determine or estimate a percentile. The procedure involves first locating the position of the percentile within the set of observations and then determining (or estimating) the value associated with that position.
Let Py be the value at or below which y percent of the distribution lies, or the yth percentile. (For example, P18 is the point at or below which 18 percent of the observations lie; 100 −18 = 82 percent are greater than P18.) The formula for the position of a percentile in an array with n entries sorted in ascending order is
(8)
where y is the percentage point at which we are dividing the distribution and Ly is the location (L) of the percentile (Py) in the array sorted in ascending order. The value of Ly may or may not be a whole number. In general, as the sample size increases, the percentile location calculation becomes more accurate; in small samples it may be quite approximate.
As an example of the case in which Ly is not a whole number, suppose that we want to determine the third quartile of returns for 2012 (Q3 or P75) for the 16 European equity markets given in Table 8. According to Equation 8, the position of the third quartile is L75 = (16 + 1)(75/100) = 12.75, or between the 12th and 13th items in Table 9, which ordered the returns into ascending order. The 12th item in Table 9 is the return to equities in France in 2012, 15.90 percent. The 13th item is the return to equities in Austria in 2012, 20.72 percent. Reflecting the “0.75” in “12.75,” we would conclude that P75 lies 75 percent of the distance between 15.90 percent and 20.72 percent.
To summarize:
When the location, Ly, is a whole number, the location corresponds to an actual observation. For example, if Denmark had not been included in the sample, then n + 1 would have been 16 and, with L75 = 12, the third quartile would be P75 = X12, where Xi is defined as the value of the observation in the ith (i = L75) position of the data sorted in ascending order (i.e., P75 =15.90).
When Ly is not a whole number or integer, Ly lies between the two closest integer numbers (one above and one below), and we use linear interpolation between those two places to determine Py. Interpolation means estimating an unknown value on the basis of two known values that surround it (lie above and below it); the term “linear” refers to a straight-line estimate. Returning to the calculation of P75 for the equity returns, we found that Ly = 12.75; the next lower whole number is 12 and the next higher whole number is 13. Using linear interpolation, P75 ≈ X12 + (12.75 − 12) (X13 − X12). As above, in the 12th position is the return to equities in France, so X12 = 15.90 percent; X13 = 20.72 percent, the return to equities in Austria. Thus our estimate is P75 ≈ X12 + (12.75 − 12)(X13 − X12) = 15.90 + 0.75 [20.72 − 15.90] = 15.90 + 0.75(4.82) = 15.90 + 3.62 = 19.52 percent. In words, 15.90 and 20.72 bracket P75 from below and above, respectively. Because 12.75 − 12 = 0.75, using linear interpolation we move 75 percent of the distance from 15.90 to 20.72 as our estimate of P75.
We follow this pattern whenever Ly is a non-integer: The nearest whole numbers below and above Ly establish the positions of observations that bracket Py and then interpolate between the values of those two observations.
Example 9 illustrates the calculation of various quantiles for the dividend yield on the components of a major European equity index.
EXAMPLE 9 Calculating Percentiles, Quartiles, and Quintiles
The EURO STOXX 50 is an index of 50 publicly traded companies, which provides a blue-chip representation of supersector leaders in the Eurozone. Table 17 shows the market capitalization on the 50 component stocks in the index, as provided by STOXX Ltd. in September 2013. The market capitalizations are ranked in ascending order.
TABLE 17 Market Capitalizations of the Components of the EURO STOXX 50
No. Company Market Cap (Euro Billion) 1 Arcelor-Mittal 8.83 2 CRH 10.99 3 RWE 11.92 4 Carrefour 12.13 5 Repsol 12.84 6 Saint-Gobain 13.60 7 France Telecom 14.09 8 Unibail-Rodamco 15.96 9 Enel 16.33 10 Essilor International 16.85 11 Intesa Sanpaolo 17.00 12 Assicurazioni Generali 17.76 13 Vivendi 17.84 14 VINCI 18.64 15 Philips 19.04 16 EADS 19.37 17 Inditex 19.66
18 UniCredit 19.69 19 Iberdrola 20.29 20 BMW 20.69 21 ASML 20.71 22 Société Générale 20.92 23 GDF Suez 21.10 24 Volkswagen 21.57 25 Munich RE 22.25 26 E.ON 24.83 27 Deutsche Telekom 25.60 28 ING 25.93 29 Air Liquide 28.98 30 L’Oreal 29.20 31 Schneider Electric 29.75 32 AXA 30.13 33 Deutsche Bank 30.92 34 LVMH Moët Hennessy 32.36 35 Danone 33.36 36 BBVA 34.56 37 Telefonica 39.00 38 ENI 41.42 39 Daimler 42.42 40 BNP Paribas 43.09 41 Unilever 46.04 42 Allianz 47.72 43 Anheuser-Busch InBev 49.40 44 SAP 50.93 45 BCO Santander 51.17 46 BASF 63.88 47 Siemens 64.27 48 Bayer 65.83 49 Total 81.06
50 Sanofi 93.29
Source: www.stoxx.com accessed 27 September 2013.
Using the data in Table 17, address the following: 1. Calculate the 10th and 90th percentiles.
2. Calculate the first, second, and third quartiles.
3. State the value of the median.
4. How many quintiles are there, and to what percentiles do the quintiles correspond?
5. Calculate the value of the first quintile.
Solution to 1: In this example, n = 50. Using Equation 8, Ly = (n + 1)y/100 for position of the yth percentile, so for the 10th percentile we have
L10 is between the fifth and sixth observations with values X5 = 12.84 and X6 = 13.60. The estimate of the 10th percentile (first decile) for dividend yield is
For the 90th percentile,
L90 is between the 45th and 46th observations with values X45 = 51.17 and X46 = 63.88, respectively. The estimate of the 90th percentile (ninth decile) is
Solution to 2: The first, second, and third quartiles correspond to P25, P50, and P75, respectively.
L25 = (51)(25/100) = 12.75 L25 is between the 12th and 13th entrieswith values
X12 = 17.76 and X13 = 17.84.
P25
= Q1 ≈ X12 + (12.75 – 12) (X13 – X12)
= 17.76 + 0.75(17.84 – 17.76)
= 17.76 + 0.75(0.08) => 17.82
L50 = (51)(50/100) = 25.5 L50 is between the 25th and 26th entrieswith values,
X25 = 22.25 and X26 = 24.83.
P50
= Q2 ≈ X25 + (25.50 – 25) (X26 – X25)
= 22.25 + 0.50(24.83 – 22.25)
= 22.25 + 0.50(2.58) = 23.54
L75 = (51)(75/100) = 38.25 L75 is between the 38th and 39th entrieswith values
X38 = 41.42 and X39= 42.42.
P75
= Q3 ≈ X38 + (42.42 – 41.42)(X39 – X38)
= 41.42 + 0.25(42.42 – 41.42)
= 41.42 + 0.25(1.00) = 41.67
Solution to 3: The median is the 50th percentile, 23.54. This is the same value that we would obtain by taking the mean of the n/2 = 50/2 = 25th item and (n + 2)/2 = 52/2 = 26th items, consistent with the procedure given earlier for the median of an even-numbered sample.
Solution to 4: There are five quintiles, and they are specified by P20, P40, P60, and P80.
Solution to 5: The first quintile is P20.
L20 = (50 + 1) (20/100) = 10.2
L20is between the 10th and 11th observations with values X10 = 16.85 and X11 = 17.00.
The estimate of the first quintile is
6.2. Quantiles in Investment Practice In this section, we discuss the use of quantiles in investments. Quantiles are used in portfolio performance evaluation as well as in investment strategy development and research.
Investment analysts use quantiles every day to rank performance—for example, the performance of portfolios. The performance of investment managers is often characterized in terms of the quartile in which they fall relative to the performance of their peer group of managers. The Morningstar mutual fund star rankings, for example, associates the number of stars with percentiles of performance relative to similar-style mutual funds.
Another key use of quantiles is in investment research. Analysts refer to a group defined by a particular quantile as that quantile. For example, analysts often refer to the set of companies with returns falling below the 10th percentile cutoff point as the bottom return decile. Dividing data into quantiles based on some characteristic allows analysts to evaluate the impact of that characteristic on a quantity of interest. For instance, empirical finance studies commonly rank companies based on the market value of their equity and then sort them into deciles. The 1st decile contains the portfolio of those companies with the smallest market values, and the 10th decile contains those companies with the largest market value. Ranking companies by decile allows analysts to compare the performance of small companies with large ones.
We can illustrate the use of quantiles, in particular quartiles, in investment research using the example of Ibbotson et al. (2013). That study proposed an investment style based on liquidity—buying stocks of less liquid stocks and selling stocks of more liquid stocks. It compared the performance of this style with three already popular investment styles, which include (1) firm size (buying stocks of small firms and selling stocks of large firms), (2) value/growth (buying stocks of value firms, defined as firms for which the stock price is relatively low in relation to earnings per share, book value per share, or dividends per share, and selling stocks of growth firms, defined as firms for which the stock price is relatively high in relation to those same measures), and (3) momentum (buying stocks of firms with a high momentum in returns, or winners, and selling stocks of firms with a low momentum, or losers.)
Ibbotson et al. examined the top 3,500 US stocks by market capitalization for the period of 1971–2011. For each stock, they computed yearly measures of liquidity as the annual share turnover (the sum of the 12 monthly volumes divided by each month’s shares outstanding), size as the year-end market capitalization, value as the trailing earnings-to-price ratio as of the year end, and momentum as the annual return. They assigned one-fourth of the total sample with the lowest liquidity in a year to Quartile 1 and the one-fourth with the highest liquidity in that year to Quartile 4. The stocks with the second-highest liquidity formed Quartile 3 and the stocks with the second-lowest liquidity, Quartile 2. Treating each quartile group as a portfolio composed of equally weighted stocks, they measured the returns on each liquidity quartile in the following year (so that the quartiles are constructed “before the fact”). The authors repeated this process for each of the other three investment styles (size, value, and momentum.) The results from Table 1 of their study are included in Table 18. We have added a column with the spreads in returns from Quartile 1 to Quartile 4.
TABLE 18 Cross-Sectional Investment Style Returns (%) and Standard Deviations of Returns (%), 1972–2011
Investment Style Q1 Q2 Q3 Q4 Spreadin Return,Q1 to Q4 Size
(Q1 = micro; Q4 = large) Geometric mean 13.04 11.93 11.95 10.98 +2.06 Arithmetic mean 16.42 14.69 14.14 12.61 +3.81 Standard deviation 27.29 24.60 21.82 18.35
Value
(Q1 = value; Q4 = growth) Geometric mean 16.13 13.60 10.10 7.62 +8.51 Arithmetic mean 18.59 15.42 12.29 11.56 +7.03 Standard deviation 23.31 20.17 21.46 29.42
Momentum
(Q1 = winners; Q4 = losers) Geometric mean 12.85 14.25 13.26 7.18 +5.67 Arithmetic mean 15.37 16.03 15.29 11.16 +4.21 Standard deviation 23.46 19.79 21.21 29.49
Liquidity
(Q1 = low; Q4 = high)
Geometric mean 14.50 13.97 11.91 7.24 +7.26 Arithmetic mean 16.38 16.05 14.39 11.04 +5.34 Standard deviation 20.41 21.50 23.20 28.48
Note: Each investment style portfolio contains as average of 742 stocks a year.
Source: Ibbotson et al.
Table 18 reports each investment style’s geometric and arithmetic mean returns and standard deviation of returns for each quartile grouping. In each style, moving from Quartile 1 to Quartile 4, mean returns decrease. For example, the geometric mean return for the least liquid stocks is 14.50% and for the most liquid stocks is 7.24%. Only for the case of size does standard deviation decrease at each step moving from Quartile 1 to Quartile 4. Thus, the table provides evidence that the investment styles generally having incremental value in explaining returns in relation to standard deviation. The authors conclude that liquidity appears to differentiate the returns approximately as well as the other styles.
To address the concern that liquidity may simply be a proxy for firm size, with investing in less liquid firms being equivalent to investing in small firms, the authors examined how less liquid stocks performed relative to more liquid stocks while controlling for firm size. This step involved constructing equally weighted double-sorted portfolios in firm size and liquidity quartiles. That is, they constructed 16 different liquidity and size portfolios (4 × 4 = 16) and investigated the interaction between these two styles. The results from Table 2 of their article are included in Table 19. We have added a column with the spreads in returns from Quartile 1 to Quartile 4 for each size category.
TABLE 19 Mean Annual Returns (%) and Standard Deviations of Returns (%) of Size and Liquidity Quartile Portfolios, 1972–2011
Quartile Q1(Lowliquidity) Q2 Q3 Q4(High liquidity)
Spreadin Return,Q1 to Q4
Micro cap Geometric mean 15.36 16.21 9.94 1.32 +14.04
Arithmetic mean 17.92 20.00 15.40 6.78 +11.14
Standard 23.77 29.41 35.34 34.20
deviation Small cap Geometric mean 15.30 14.09 11.80 5.48 +9.82
Arithmetic mean 17.07 16.82 15.38 9.89 +7.18
Standard deviation 20.15 24.63 28.22 31.21
Mid cap Geometric mean 13.61 13.57 12.24 7.85 +5.76
Arithmetic mean 15.01 15.34 14.51 11.66 +3.35
Standard deviation 17.91 20.10 22.41 28.71
Large cap Geometric mean 11.53 11.66 11.19 8.37 +3.16
Arithmetic mean 12.83 12.86 12.81 11.58 +1.25
Standard deviation 16.68 15.99 18.34 25.75
Source: Ibbotson et al.
The table shows that within the quartile with the smallest firms, the low-liquidity portfolio earned an annual geometric mean return of 15.36 percent, in contrast to the high-liquidity portfolio return of 1.32 percent, producing a liquidity effect of 14.04 percentage points (1,404 basis points). While the liquidity effect is strongest for the smallest firms, it does persist in the other three size quartiles also. These results indicate that size does not capture liquidity (i.e., the liquidity effect holds regardless of size group).
7. MEASURES OF DISPERSION As the well-known researcher Fischer Black has written, “[t]he key issue in investments is estimating expected return.”26 Few would disagree with the importance of expected return or mean return in investments: The mean return tells us where returns, and investment results, are centered. To completely understand an investment, however, we also need to know how returns are dispersed around the mean. Dispersion is the variability around the central tendency. If mean return addresses reward, dispersion addresses risk.
In this section, we examine the most common measures of dispersion: range, mean absolute deviation, variance, and standard deviation. These are all measures of absolute dispersion. Absolute dispersion is the amount of variability present without comparison to any reference point or benchmark.
These measures are used throughout investment practice. The variance or standard deviation of return is often used as a measure of risk pioneered by Nobel laureate Harry Markowitz. William Sharpe, another winner of the Nobel Prize in economics, developed the Sharpe ratio, a measure of risk-adjusted performance. That measure makes use of standard deviation of return. Other measures of dispersion, mean absolute deviation and range, are also useful in analyzing data.
7.1. The Range We encountered range earlier when we discussed the construction of frequency distribution. The simplest of all the measures of dispersion, range can be computed with interval or ratio data.
Definition of Range. The range is the difference between the maximum and minimum values in a data set:
(9)
As an illustration of range, the largest monthly return for the S&P 500 in the period from January 1926 to December 2012 is 42.56 percent (in April 1933) and the smallest is −29.73 percent (in September 1931). The range of returns is thus 72.29 percent [42.56 percent − (−29.73 percent)]. An alternative definition of range reports the maximum and minimum values. This alternative definition provides more information than does the range as defined in Equation 9.
One advantage of the range is ease of computation. A disadvantage is that the range uses only two pieces of information from the distribution. It cannot tell us how the
data are distributed (that is, the shape of the distribution). Because the range is the difference between the maximum and minimum returns, it can reflect extremely large or small outcomes that may not be representative of the distribution.27
7.2. The Mean Absolute Deviation Measures of dispersion can be computed using all the observations in the distribution rather than just the highest and lowest. The question is, how should we measure dispersion? Our previous discussion on properties of the arithmetic mean introduced the notion of distance or deviation from the mean as a fundamental piece of information used in statistics. We could compute measures of dispersion as the arithmetic average of the deviations around the mean, but we would encounter a problem: The deviations around the mean always sum to 0. If we computed the mean of the deviations, the result would also equal 0. Therefore, we need to find a way to address the problem of negative deviations canceling out positive deviations.
One solution is to examine the absolute deviations around the mean as in the mean absolute deviation.
Mean Absolute Deviation Formula. The mean absolute deviation (MAD) for a sample is
(10)
where is the sample mean and n is the number of observations in the sample.
In calculating MAD, we ignore the signs of the deviations around the mean. For example, if Xi = −11.0 and = 4.5, the absolute value of the difference is |−11.0 − 4.5| = |−15.5| = 15.5. The mean absolute deviation uses all of the observations in the sample and is thus superior to the range as a measure of dispersion. One technical drawback of MAD is that it is difficult to manipulate mathematically compared with the next measure we will introduce, variance.28 Example 10 illustrates the use of the range and the mean absolute deviation in evaluating risk.
EXAMPLE 10 The Range and the Mean Absolute Deviation
Having calculated mean returns for the two mutual funds in Example 7, the analyst is now concerned with evaluating risk.
TABLE 15 Total Returns for Two Mutual Funds, 2008–2012 (Repeated)
Year Selected AmericanShares(SLASX) T. Rowe Price Equity Income(PRFDX)
2008 −39.44% −35.75% 2009 31.64 25.62 2010 12.53 15.15 2011 −4.35 −0.72 2012 12.82 17.25
Source: performance.morningstar.com.
Based on the data in Table 15 on the previous page, answer the following: 1. Calculate the range of annual returns for (A) SLASX and (B) PRFDX, and state
which mutual fund appears to be riskier based on these ranges.
2. Calculate the mean absolute deviation of returns on (A) SLASX and (B) PRFDX, and state which mutual fund appears to be riskier based on MAD.
Solution to 1:
A.For SLASX, the largest return was 31.64 percent and the smallest was −39.44 percent. The range is thus 31.64 − (−39.44) = 71.08%.
B.For PFRDX, the range is 25.62 − (−35.75) = 61.37%. With a larger range of returns than PRFDX, SLASX appeared to be the riskier fund during the 2008– 2012 period.
Solution to 2: 1. The arithmetic mean return for SLASX as calculated in Example 7 is 2.64
percent. The MAD of SLASX returns is
2. The arithmetic mean return for PRFDX as calculated in Example 7 is 4.31 percent. The MAD of PRFDX returns is
SLASX, with a MAD of 19.63 percent, appears to be slightly riskier than PRFDX, with a MAD of 18.04 percent.
7.3. Population Variance and Population Standard Deviation The mean absolute deviation addressed the issue that the sum of deviations from the mean equals zero by taking the absolute value of the deviations. A second approach to the treatment of deviations is to square them. The variance and standard deviation, which are based on squared deviations, are the two most widely used measures of dispersion. Variance is defined as the average of the squared deviations around the mean. Standard deviation is the positive square root of the variance. The following discussion addresses the calculation and use of variance and standard deviation.
7.3.1. Population Variance
If we know every member of a population, we can compute the population variance. Denoted by the symbol σ2, the population variance is the arithmetic average of the squared deviations around the mean.
Population Variance Formula. The population variance is
(11)
where μ is the population mean and N is the size of the population.
Given knowledge of the population mean, μ, we can use Equation 11 to calculate the sum of the squared differences from the mean, taking account of all N items in the population, and then to find the mean squared difference by dividing the sum by N. Whether a difference from the mean is positive or negative, squaring that difference results in a positive number. Thus variance takes care of the problem of negative
deviations from the mean canceling out positive deviations by the operation of squaring those deviations. The profit as a percentage of revenue for BJ’s Wholesale Club, Costco, and Walmart was given earlier as 0.9, 1.6, and 3.5, respectively. We calculated the mean profit as a percentage of revenue as 2.0. Therefore, the population variance of the profit as a percentage of revenue is (1/3)[(0.9 − 2.0)2 + (1.6 − 2.0)2 + (3.5 − 2.0)2] = (1/3)(−1.12 + −0.42 + 1.52) = (1/3)(1.21 + 0.16 + 2.25) = (1/3)(3.62) = 1.21.
7.3.2. Population Standard Deviation
Because the variance is measured in squared units, we need a way to return to the original units. We can solve this problem by using standard deviation, the square root of the variance. Standard deviation is more easily interpreted than the variance because standard deviation is expressed in the same unit of measurement as the observations.
Population Standard Deviation Formula. The population standard deviation, defined as the positive square root of the population variance, is
(12)
where μ is the population mean and N is the size of the population.
Using the example of the profit as a percentage of revenue for BJ’s Wholesale Club, Costco, and Walmart, according to Equation 12 we would calculate the variance, 1.21, then take the square root: = 1.10.
Both the population variance and standard deviation are examples of parameters of a distribution. In later readings, we will introduce the notion of variance and standard deviation as risk measures.
In investments, we often do not know the mean of a population of interest, usually because we cannot practically identify or take measurements from each member of the population. We then estimate the population mean with the mean from a sample drawn from the population, and we calculate a sample variance or standard deviation using formulas different from Equations 11 and 12. We shall discuss these calculations in subsequent sections. However, in investments we sometimes have a defined group that we can consider to be a population. With well-defined populations, we use Equations 11 and 12, as in the following example.
EXAMPLE 11 Calculating the Population Standard Deviation
Table 20 gives the yearly portfolio turnover for the 12 US equity funds that composed the 2013 Forbes Magazine Honor Roll.29 Portfolio turnover, a measure of trading activity, is the lesser of the value of sales or purchases over a year divided by average net assets during the year. The number and identity of the funds on the Forbes Honor Roll changes from year to year.
TABLE 20 Portfolio Turnover: 2013 Forbes Honor Roll Mutual Funds
Fund Yearly Portfolio Turnover(%) Bruce Fund (BRUFX) 10
CGM Focus Fund (CGMFX) 360 Hotchkis and Wiley Small Cap Value A Fund
(HWSAX) 37
Aegis Value Fund (AVALX) 20 Delafield Fund (DEFIX) 49
Homestead Small Company Stock Fund (HSCSX) 1 Robeco Boston Partners Small Cap Value II Fund
(BPSCX) 32
Hotchkis and Wiley Mid Cap Value A Fund (HWMAX) 72
T Rowe Price Small Cap Value Fund (PRSVX) 9 Guggenheim Mid Cap Value Fund Class A (SEVAX) 19
Wells Fargo Advantage Small Cap Value Fund (SSMVX) 16
Stratton Small Cap Value Fund (STSCX) 11
Source: Forbes (2013).
Based on the data in Table 20, address the following: 1. Calculate the population mean portfolio turnover for the period used by Forbes
for the 12 Honor Roll funds.
2. Calculate the population variance and population standard deviation of portfolio turnover.
3. Explain the use of the population formulas in this example.
Solution to 1:
Solution to 2: Having established that μ = 53, we can calculate by first calculating the numerator in the expression and then dividing by N = 12. The numerator (the sum of the squared differences from the mean) is
To calculate standard deviation, = 94.51 percent. (The unit of variance is percent squared so the unit of standard deviation is percent.)
Solution to 3: If the population is clearly defined to be the Forbes Honor Roll funds in one specific year (2013), and if portfolio turnover is understood to refer to the specific one-year period reported upon by Forbes, the application of the population formulas to variance and standard deviation is appropriate. The results of 8,932.50 and 94.51 are, respectively, the cross-sectional variance and standard deviation in yearly portfolio turnover for the 2013 Forbes Honor Roll Funds.30
7.4. Sample Variance and Sample Standard Deviation In the following discussion, note the switch to Roman letters to symbolize sample quantities.
7.4.1. Sample Variance
In many instances in investment management, a subset or sample of the population is all that we can observe. When we deal with samples, the summary measures are called statistics. The statistic that measures the dispersion in a sample is called the sample variance.
Sample Variance Formula. The sample variance is
(13)
where is the sample mean and n is the number of observations in the sample.
Equation 13 tells us to take the following steps to compute the sample variance: 1. Calculate the sample mean, .
2. Calculate each observation’s squared deviation from the sample mean, .
3. Sum the squared deviations from the mean: .
4. Divide the sum of squared deviations from the mean by n – 1: .
We will illustrate the calculation of the sample variance and the sample standard deviation in Example 12.
We use the notation s2 for the sample variance to distinguish it from population variance, σ2. The formula for sample variance is nearly the same as that for population variance except for the use of the sample mean, , in place of the population mean, μ, and a different divisor. In the case of the population variance, we divide by the size of the population, N. For the sample variance, however, we divide by the sample size minus 1, or n − 1. By using n − 1 (rather than n) as the divisor, we improve the statistical properties of the sample variance. In statistical terms, the sample variance defined in Equation 13 is an unbiased estimator of the population variance.31 The quantity n − 1 is also known as the number of degrees of freedom in estimating the population variance. To estimate the population variance with s2, we must first calculate the mean. Once we have computed the sample mean, there are only n − 1 independent deviations from it.
7.4.2. Sample Standard Deviation
Just as we computed a population standard deviation, we can compute a sample standard deviation by taking the positive square root of the sample variance.
Sample Standard Deviation Formula. The sample standard deviation, s, is
(14)
where is the sample mean and n is the number of observations in the sample.
To calculate the sample standard deviation, we first compute the sample variance using the steps given. We then take the square root of the sample variance. Example 12 illustrates the calculation of the sample variance and standard deviation for the two mutual funds introduced earlier.
EXAMPLE 12 Calculating Sample Variance and Sample Standard Deviation
After calculating the geometric and arithmetic mean returns of two mutual funds in Example 7, we calculated two measures of dispersions for those funds, the range and mean absolute deviation of returns, in Example 10. We now calculate the sample variance and sample standard deviation of returns for those same two funds.
Based on the data in Table 15 repeated below, answer the following: 1. Calculate the sample variance of return for (A) SLASX and (B) PRFDX.
2. Calculate the sample standard deviation of return for (A) SLASX and (B) PRFDX.
3. Contrast the dispersion of returns as measured by standard deviation of return and mean absolute deviation of return for each of the two funds.
Solution to 1: To calculate the sample variance, we use Equation 13. (Deviation answers are all given in percent squared.)
A. SLASX
The sample mean is
The squared deviations from the mean are
TABLE 15 Total Returns for Two Mutual Funds, 2008–2012
Year Selected AmericanShares(SLASX) T. Rowe Price Equity Income(PRFDX)
2008 −39.44% −35.75% 2009 31.64 25.62 2010 12.53 15.15 2011 −4.35 −0.72 2012 12.82 17.25
Source: performance.morningstar.com.
The sum of the squared deviations from the mean is 1,770.73 + 841.00 + 97.81 + 48.86 + 103.63 = 2,862.03.
Divide the sum of the squared deviations from the mean by n − 1: 2,862.03/(5 − 1) = 2,862.03/4 = 715.51.
B.PRFDX
The sample mean is
The squared deviations from the mean are
The sum of the squared deviations from the mean is 1,604.80 + 454.12 + 117.51 + 25.30 + 167.44 = 2,369.17.
Divide the sum of the squared deviations from the mean by n − 1: 2,369.17/4 = 592.29.
Solution to 2: To find the standard deviation, we take the positive square root of variance.
A.For SLASX, = 26.7%.
B.For PRFDX, = 24.3%.
Solution to 3: Table 21 summarizes the results from Part 2 for standard deviation and incorporates the results for MAD from Example 10.
Note that the mean absolute deviation is less than the standard deviation. The mean absolute deviation will always be less than or equal to the standard deviation because the standard deviation gives more weight to large deviations than to small ones (remember, the deviations are squared).
TABLE 21 Two Mutual Funds: Comparison of Standard Deviation and Mean Absolute Deviation
Fund Standard Deviation (%) Mean Absolute Deviation (%) SLASX 26.7 19.6 PRFDX 24.3 18.0
Because the standard deviation is a measure of dispersion about the arithmetic mean, we usually present the arithmetic mean and standard deviation together when summarizing data. When we are dealing with data that represent a time series of percent changes, presenting the geometric mean—representing the compound rate of growth—is also very helpful. Table 22 presents the historical geometric and arithmetic mean returns, along with the historical standard deviation of returns, for the S&P 500 annual and monthly return series. We present these statistics for nominal (rather than inflation-adjusted) returns so we can observe the original magnitudes of the returns.
TABLE 22 Equity Market Returns: Means and Standard Deviations
Return Series GeometricMean(%) ArithmeticMean
(%) StandardDeviation
Ibbotson Associates Series: 1926–2012 S&P 500 (Annual) 9.84 11.82 20.18
S&P 500 (Monthly) 0.79 0.94 5.50
Source: Ibbotson.
7.5. Semivariance, Semideviation, and Related Concepts An asset’s variance or standard deviation of returns is often interpreted as a measure of the asset’s risk. Variance and standard deviation of returns take account of returns above and below the mean, but investors are concerned only with downside risk, for example, returns below the mean. As a result, analysts have developed semivariance, semideviation, and related dispersion measures that focus on downside risk. Semivariance is defined as the average squared deviation below the mean. Semideviation (sometimes called semistandard deviation) is the positive square root of semivariance.32 To compute the sample semivariance, for example, we take the following steps:
Calculate the sample mean.
Identify the observations that are smaller than or equal to the mean (discarding observations greater than the mean).
Compute the sum of the squared negative deviations from the mean (using the observations that are smaller than or equal to the mean).
Divide the sum of the squared negative deviations from Step iii by the total sample size minus 1: n − 1. A formula for semivariance approximating the unbiased estimator is
To take the case of Selected American Shares with returns (in percent) of −39.44, 31.64, 12.53, −4.35, and 12.82, we earlier calculated a mean return of 2.64 percent. Two returns, −39.44 and −4.35, are smaller than 2.64. We compute the sum of the squared negative deviations from the mean as (−39.44 − 2.64)2 + (−4.35 − 2.64)2 = −42.082 + −6.992 = 1,770.73 + 48.86 = 1,819.59. With n − 1 = 4, we conclude that semivariance is 1,819.59/4 = 454.9 and that semideviation is = 21.3 percent, approximately. The semideviation of 21.3 percent is less than the standard deviation of 26.7 percent. From this downside risk perspective, therefore, standard deviation overstates risk.
In practice, we may be concerned with values of return (or another variable) below
some level other than the mean. For example, if our return objective is 12.75 percent annually, we may be concerned particularly with returns below 12.75 percent a year. We can call 12.75 percent the target. The name target semivariance has been given to average squared deviation below a stated target, and target semideviation is its positive square root. To calculate a sample target semivariance, we specify the target as a first step. After identifying observations below the target, we find the sum of the squared negative deviations from the target and divide that sum by the number of observations minus 1. A formula for target semivariance is
where B is the target and n is the number of observations. With a target return of 12.75 percent, we find in the case of Selected American Shares that three returns (−39.44, 12.53, and −4.35) were below the target. The target semivariance is [(−39.44 − 12.75)2 + (12.53 − 12.75)2 + (−4.35 − 12.75)2]/(5 – 1) = 754.06, and the target semideviation is = 27.5 percent, approximately.
When return distributions are symmetric, semivariance and variance are effectively equivalent. For asymmetric distributions, variance and semivariance rank prospects’ risk differently.33 Semivariance (or semideviation) and target semivariance (or target semideviation) have intuitive appeal, but they are harder to work with mathematically than variance.34 Variance or standard deviation enters into the definition of many of the most commonly used finance risk concepts, such as the Sharpe ratio and beta. Perhaps because of these reasons, variance (or standard deviation) is much more frequently used in investment practice.
7.6. Chebyshev’s Inequality The Russian mathematician Pafnuty Chebyshev developed an inequality using standard deviation as a measure of dispersion. The inequality gives the proportion of values within k standard deviations of the mean.
Definition of Chebyshev’s Inequality. According to Chebyshev’s inequality, for any distribution with finite variance, the proportion of the observations within k standard deviations of the arithmetic mean is at least 1 − 1/k2 for all k > 1.
Table 23 illustrates the proportion of the observations that must lie within a certain number of standard deviations around the sample mean.
TABLE 23 Proportions from Chebyshev’s Inequality
k Interval around the Sample Mean Proportion (%) 1.25 36 1.50 56 2.00 75 2.50 84 3.00 89 4.00 94
Note: Standard deviation is denoted as s.
When k = 1.25, for example, the inequality states that the minimum proportion of the observations that lie within ± 1.25s is 1 − 1/(1.25)2 = 1 − 0.64 = 0.36 or 36 percent.
The most frequently cited facts that result from Chebyshev’s inequality are that a two-standard-deviation interval around the mean must contain at least 75 percent of the observations, and a three-standard-deviation interval around the mean must contain at least 89 percent of the observations, no matter how the data are distributed.
The importance of Chebyshev’s inequality stems from its generality. The inequality holds for samples and populations and for discrete and continuous data regardless of the shape of the distribution. As we shall see in the reading on sampling, we can make much more precise interval statements if we can assume that the sample is drawn from a population that follows a specific distribution called the normal distribution. Frequently, however, we cannot confidently assume that distribution.
The next example illustrates the use of Chebyshev’s inequality.
EXAMPLE 13 Applying Chebyshev’s Inequality
According to Table 22, the arithmetic mean monthly return and standard deviation of monthly returns on the S&P 500 were 0.94 percent and 5.50 percent, respectively, during the 1926–2012 period, totaling 1,044 monthly observations. Using this information, address the following:
1. Calculate the endpoints of the interval that must contain at least 75 percent of
monthly returns according to Chebyshev’s inequality.
2. What are the minimum and maximum number of observations that must lie in the interval computed in Part 1, according to Chebyshev’s inequality?
Solution to 1: According to Chebyshev’s inequality, at least 75 percent of the observations must lie within two standard deviations of the mean, ± 2s. For the monthly S&P 500 return series, we have 0.94% ± 2(5.50%) = 0.94% ± 11.00%. Thus the lower endpoint of the interval that must contain at least 75 percent of the observations is 0.94% − 11.00% = −10.06%, and the upper endpoint is 0.94% + 11.00% = 11.94%.
Solution to 2: For a sample size of 1,044, at least 0.75(1,044) = 783 observations must lie in the interval from −10.06% to 11.94% that we computed in Part 1. Chebyshev’s inequality gives the minimum percentage of observations that must fall within a given interval around the mean, but it does not give the maximum percentage. Table 4, which gave the frequency distribution of monthly returns on the S&P 500, is excerpted below. The data in the excerpted table are consistent with the prediction of Chebyshev’s inequality. The set of intervals running from −10.0% to 12.0% is about equal in width to the two-standard-deviation interval −10.06% to 11.94%. A total of 1,004 observations (approximately 96 percent of observations) fall in the range from −10.0% to 12.0%.
TABLE 4 Frequency Distribution for the Monthly Total Return on the S&P 500, January 1926 to December 2012 (Excerpt)
Return Interval (%) Absolute Frequency −10.0 to −8.0 23 −8.0 to −6.0 34 −6.0 to −4.0 59 −4.0 to −2.0 98 −2.0 to 0.0 157 0.0 to 2.0 220 2.0 to 4.0 173 4.0 to 6.0 137 6.0 to 8.0 63 8.0 to 10.0 25 10.0 to 12.0 15
1,004
7.7. Coefficient of Variation We noted earlier that standard deviation is more easily interpreted than variance because standard deviation uses the same units of measurement as the observations. We may sometimes find it difficult to interpret what standard deviation means in terms of the relative degree of variability of different sets of data, however, either because the data sets have markedly different means or because the data sets have different units of measurement. In this section we explain a measure of relative dispersion, the coefficient of variation that can be useful in such situations. Relative dispersion is the amount of dispersion relative to a reference value or benchmark.
We can illustrate the problem of interpreting the standard deviation of data sets with markedly different means using two hypothetical samples of companies. The first sample, composed of small companies, includes companies with 2003 sales of €50 million, €75 million, €65 million, and €90 million. The second sample, composed of large companies, includes companies with 2003 sales of €800 million, €825 million, €815 million, and €840 million.
We can verify using Equation 14 that the standard deviation of sales in both samples is €16.8 million.35 In the first sample, the largest observation, €90 million, is 80 percent larger than the smallest observation, €50 million. In the second sample, the largest observation is only 5 percent larger than the smallest observation. Informally, a standard deviation of €16.8 million represents a high degree of variability relative to the first sample, which reflects mean 2003 sales of €70 million, but a small degree of variability relative to the second sample, which reflects mean 2003 sales of €820 million.
The coefficient of variation is helpful in situations such as that just described.
Coefficient of Variation Formula. The coefficient of variation, CV, is the ratio of the standard deviation of a set of observations to their mean value:36
(15)
where s is the sample standard deviation and is the sample mean.
When the observations are returns, for example, the coefficient of variation measures the amount of risk (standard deviation) per unit of mean return. Expressing the magnitude of variation among observations relative to their average
size, the coefficient of variation permits direct comparisons of dispersion across different data sets. Reflecting the correction for scale, the coefficient of variation is a scale-free measure (that is, it has no units of measurement).
We can illustrate the application of the coefficient of variation using our earlier example of two samples of companies. The coefficient of variation for the first sample is (€16.8 million)/(€70 million) = 0.24; the coefficient of variation for the second sample is (€16.8 million)/(€820 million) = 0.02. This confirms our intuition that the first sample had much greater variability in sales than the second sample. Note that 0.24 and 0.02 are pure numbers in the sense that they are free of units of measurement (because we divided the standard deviation by the mean, which is measured in the same units as the standard deviation). If we need to compare the dispersion among data sets stated in different units of measurement, the coefficient of variation can be useful because it is free from units of measurement. Example 14 illustrates the calculation of the coefficient of variation.
EXAMPLE 14 Calculating the Coefficient of Variation
Table 24 summarizes annual mean returns and standard deviations computed using monthly return data for major stock market indexes of four Asia-Pacific markets. The indexes are: S&P/ASX 200 Index (Australia), Hang Seng Index (Hong Kong), Straits Times Index (Singapore), and KOSPI Composite Index (South Korea).
TABLE 24 Arithmetic Mean Annual Return and Standard Deviation of Returns, Asia-Pacific Stock Markets, 2003–2012
Market Arithmetic MeanReturn (%) Standard Deviationof Return (%) Australia 5.0 13.6 Hong Kong 9.4 22.4 Singapore 9.3 19.2 South Korea 12.0 21.5
Source: finance.yahoo.com.
Using the information in Table 24, address the following: 1. Calculate the coefficient of variation for each market given.
2. Rank the markets from most risky to least risky using CV as a measure of relative dispersion.
3. Determine whether there is more difference between the absolute or the relative riskiness of the Hong Kong and Singapore markets. Use the standard deviation as a measure of absolute risk and CV as a measure of relative risk.
Solution to 1:
Solution to 2: Based on CV, the ranking for the 2003–2012 period examined is Australia (most risky), Hong Kong, Singapore, and South Korea (least risky).
Solution to 3: As measured both by standard deviation and CV, the Hong Kong market was riskier than the Singapore market. The standard deviation of Hong Kong returns was (22.4 − 19.2)/19.2 = 0.167 or about 17 percent larger than Singapore returns, compared with a difference in the CV of (2.383 − 2.065)/2.065 = 0.154 or about 15 percent. Thus, the CVs reveal slightly less difference between Hong Kong and Singapore return variability than that suggested by the standard deviations alone.
7.8. The Sharpe Ratio Although CV was designed as a measure of relative dispersion, its inverse reveals something about return per unit of risk because the standard deviation of returns is commonly used as a measure of investment risk. For example, a portfolio with a mean monthly return of 1.19 percent and a standard deviation of 4.42 percent has an inverse CV of 1.19%/4.42% = 0.27. This result indicates that each unit of standard deviation represents a 0.27 percent return.
A more precise return–risk measure recognizes the existence of a risk-free return, a return for virtually zero standard deviation. With a risk-free asset, an investor can choose a risky portfolio, p, and then combine that portfolio with the risk-free asset to achieve any desired level of absolute risk as measured by standard deviation of return, sp. Consider a graph with mean return on the vertical axis and standard deviation of return on the horizontal axis. Any combination of portfolio p and the risk-free asset lies on a ray (line) with slope equal to the quantity (Mean return − Risk-free return) divided by sp. The ray giving investors choices offering the most reward (return in excess of the risk-free rate) per unit of risk is the one with the
highest slope. The ratio of excess return to standard deviation of return for a portfolio p—the slope of the ray passing through p—is a single-number measure of a portfolio’s performance known as the Sharpe ratio, after its developer, William F. Sharpe.
Sharpe Ratio Formula. The Sharpe ratio for a portfolio p, based on historical returns, is defined as
(16)
where is the mean return to the portfolio, is the mean return to a risk-free asset, and sp is the standard deviation of return on the portfolio.37
The numerator of the Sharpe measure is the portfolio’s mean return minus the mean return on the risk-free asset over the sample period. The term measures the extra reward that investors receive for the added risk taken. We call this difference the mean excess return on portfolio p. Thus the Sharpe ratio measures the reward, in terms of mean excess return, per unit of risk, as measured by standard deviation of return. Those risk-averse investors who make decisions only in terms of mean return and standard deviation of return prefer portfolios with larger Sharpe ratios to those with smaller Sharpe ratios.
To illustrate the calculation of the Sharpe ratio, consider the performance of two exchange traded funds. SPDR S&P 500 (NYSE: SPY) seeks to track the investment results of the S&P 500 Index (large capitalization US stocks) and iShares Russell 2000 Index (NYSE: IWM) seeks to track the investment results of the Russell 2000 Index (small capitalization US stocks). Table 25 presents the historical arithmetic mean return, along with the historical standard deviation of returns, for annual returns series of these two funds and the US 30 day T-bill during the 2003–2012 period.
TABLE 25 Exchange Traded Fund and US 30 Day T-Bill Mean Return and Standard Deviation of Return, 2003–2012
Fund/T-Bill Arithmetic Mean (%) Standard Deviation of Return (%) IWM 9.26 22.36 SPY 6.77 19.99
30 day T-bill 1.58 1.78
Sources: finance.yahoo.com and www.federalreserve.gov.
Using the mean US 30 day T-bill return to represent the risk-free rate, we find the following Sharpe ratios
Although US small stocks (IWM) had a higher standard deviation, they performed better than the US large stocks (SPY), as measured by the Sharpe ratio.
The Sharpe ratio is a mainstay of performance evaluation. We must issue two cautions concerning its use, one related to interpreting negative Sharpe ratios and the other to conceptual limitations.
Finance theory tells us that in the long run, investors should be compensated with additional mean return above the risk-free rate for bearing additional risk, at least if the risky portfolio is well diversified. If investors are so compensated, the numerator of the Sharpe ratio will be positive. Nevertheless, we often find that portfolios exhibit negative Sharpe ratios when the ratio is calculated over periods in which bear markets for equities dominate. This raises a caution when dealing with negative Sharpe ratios. With positive Sharpe ratios, a portfolio’s Sharpe ratio decreases if we increase risk, all else equal. That result is intuitive for a risk- adjusted performance measure. With negative Sharpe ratios, however, increasing risk results in a numerically larger Sharpe ratio (for example, doubling risk may increase the Sharpe ratio from −1 to −0.5). Therefore, in a comparison of portfolios with negative Sharpe ratios, we cannot generally interpret the larger Sharpe ratio (the one closer to zero) to mean better risk-adjusted performance.38 Practically, to make an interpretable comparison in such cases using the Sharpe ratio, we may need to increase the evaluation period such that one or more of the Sharpe ratios becomes positive; we might also consider using a different performance evaluation metric.
The conceptual limitation of the Sharpe ratio is that it considers only one aspect of risk, standard deviation of return. Standard deviation is most appropriate as a risk measure for portfolio strategies with approximately symmetric return distributions. Strategies with option elements have asymmetric returns. Relatedly, an investment strategy may produce frequent small gains but have the potential for infrequent but extremely large losses.39 Such a strategy is sometimes described as picking up coins in front of a bulldozer; for example, some hedge fund strategies tend to produce that return pattern. Calculated over a period in which the strategy is
working (a large loss has not occurred), this type of strategy would have a high Sharpe ratio. In this case, the Sharpe ratio would give an overly optimistic picture of risk-adjusted performance because standard deviation would incompletely measure the risk assumed.40 Therefore, before applying the Sharpe ratio to evaluate a manager, we should judge whether standard deviation adequately describes the risk of the manager ’s investment strategy.
Example 15 illustrates the calculation of the Sharpe ratio in a portfolio performance evaluation context.
EXAMPLE 15 Calculating the Sharpe Ratio
In earlier examples, we computed the various statistics for two mutual funds, Selected American Shares (SLASX) and T. Rowe Price Equity Income (PRFDX), for a five-year period ending in December 2012. Table 26 summarizes selected statistics for these two mutual funds for a longer period, the 10-year period ending in 2012.
The US 30-day T-bill rate is frequently used as a proxy for the risk-free rate. Earlier, Table 25 gave the average annual return on T-bills for the 2003–2012 period as 1.58 percent.
Using the information in Table 25 and the average annual return on 30-day T- bills, address the following:
1. Calculate the Sharpe ratios for SLASX and PRFDX during the 2003–2012
period.
2. State which fund had superior risk-adjusted performance during this period, as measured by the Sharpe ratio.
Solution to 1: We already have in hand the means of the portfolio return and standard deviations of returns and the mean annual risk-free rate of return from 2003 to 2012.
Solution to 2: PRFDX had a higher positive Sharpe ratio than SLASX during the period. As measured by the Sharpe ratio, PRFDX’s performance was superior. This is not surprising as PRFDX had a higher return and a lower
standard deviation than SLASX.
TABLE 26 Mutual Fund Mean Return and Standard Deviation of Return, 2003– 2012
Fund Arithmetic Mean (%) Standard Deviation of Return (%) SLASX 8.60 20.02 PRFDX 8.91 18.12
Source: performance.morningstar.com.
8. SYMMETRY AND SKEWNESS IN RETURN DISTRIBUTIONS Mean and variance may not adequately describe an investment’s distribution of returns. In calculations of variance, for example, the deviations around the mean are squared, so we do not know whether large deviations are likely to be positive or negative. We need to go beyond measures of central tendency and dispersion to reveal other important characteristics of the distribution. One important characteristic of interest to analysts is the degree of symmetry in return distributions.
If a return distribution is symmetrical about its mean, then each side of the distribution is a mirror image of the other. Thus equal loss and gain intervals exhibit the same frequencies. Losses from −5 percent to −3 percent, for example, occur with about the same frequency as gains from 3 percent to 5 percent.
One of the most important distributions is the normal distribution, depicted in Figure 6. This symmetrical, bell-shaped distribution plays a central role in the mean–variance model of portfolio selection; it is also used extensively in financial risk management. The normal distribution has the following characteristics:
Its mean and median are equal.
It is completely described by two parameters—its mean and variance.
Roughly 68 percent of its observations lie between plus and minus one standard deviation from the mean; 95 percent lie between plus and minus two standard deviations; and 99 percent lie between plus and minus three standard deviations.
A distribution that is not symmetrical is called skewed. A return distribution with positive skew has frequent small losses and a few extreme gains. A return distribution with negative skew has frequent small gains and a few extreme losses. Figure 7 shows positively and negatively skewed distributions. The positively skewed distribution shown has a long tail on its right side; the negatively skewed distribution has a long tail on its left side. For the positively skewed unimodal distribution, the mode is less than the median, which is less than the mean. For the negatively skewed unimodal distribution, the mean is less than the median, which is less than the mode.41 Investors should be attracted by a positive skew because the mean return falls above the median. Relative to the mean return, positive skew amounts to a limited, though frequent, downside compared with a somewhat
unlimited, but less frequent, upside.
Skewness is the name given to a statistical measure of skew. (The word “skewness” is also sometimes used interchangeably for “skew.”) Like variance, skewness is computed using each observation’s deviation from its mean. Skewness (sometimes referred to as relative skewness) is computed as the average cubed deviation from the mean standardized by dividing by the standard deviation cubed to make the measure free of scale.42 A symmetric distribution has skewness of 0, a positively skewed distribution has positive skewness, and a negatively skewed distribution has negative skewness, as given by this measure.
FIGURE 6 Properties of a Normal Distribution (EV 5 Expected Value)
Source: Reprinted from Fixed Income Analysis. Copyright CFA Institute.
FIGURE 7 Properties of a Skewed Distribution
Source: Reprinted from Fixed Income Analysis. Copyright CFA Institute.
We can illustrate the principle behind the measure by focusing on the numerator. Cubing, unlike squaring, preserves the sign of the deviations from the mean. If a distribution is positively skewed with a mean greater than its median, then more than half of the deviations from the mean are negative and less than half are positive. In order for the sum to be positive, the losses must be small and likely, and the gains less likely but more extreme. Therefore, if skewness is positive, the average magnitude of positive deviations is larger than the average magnitude of negative deviations.
A simple example illustrates that a symmetrical distribution has a skewness measure equal to 0. Suppose we have the following data: 1, 2, 3, 4, 5, 6, 7, 8, and 9. The mean outcome is 5, and the deviations are −4, −3, −2, −1, 0, 1, 2, 3, and 4. Cubing the deviations yields −64, −27, −8, −1, 0, 1, 8, 27, and 64, with a sum of 0. The numerator of skewness (and so skewness itself) is thus equal to 0, supporting our claim. Below we give the formula for computing skewness from a sample.
Sample Skewness Formula. Sample skewness (also called sample relative skewness), Sk, is
(17)
where n is the number of observations in the sample and s is the sample standard deviation.43
The algebraic sign of Equation 17 indicates the direction of skew, with a negative SK indicating a negatively skewed distribution and a positive SK indicating a positively skewed distribution. Note that as n becomes large, the expression reduces to the
mean cubed deviation, . As a frame of reference, for a sample size of 100 or larger taken from a normal distribution, a skewness coefficient of ±0.5 would be considered unusually large.
TABLE 27 S&P 500 Annual and Monthly Total Returns, 1926–2012: Summary Statistics
Return Series
Number ofPeriods
ArithmeticMean (%)
Standard Deviation (%) Skewness
Excess Kurtosis
S&P 500 (Annual) 87 11.82 20.18 −0.3768 0.0100
S&P 500 (Monthly) 1,044 0.94 5.50 0.3456 9.4288
Source: Ibbotson Associates.
Table 27 shows several summary statistics for the annual and monthly returns on the S&P 500. Earlier we discussed the arithmetic mean return and standard deviation of return, and we shall shortly discuss kurtosis.
Table 27 reveals that S&P 500 annual returns during this period were negatively skewed while monthly returns were positively skewed, and the magnitude of skewness was greater for the annual series. We would find for other market series that the shape of the distribution of returns often depends on the holding period examined.
Some researchers believe that investors should prefer positive skewness, all else equal—that is, they should prefer portfolios with distributions offering a relatively large frequency of unusually large payoffs.44 Different investment strategies may tend to introduce different types and amounts of skewness into returns. Example 16 illustrates the calculation of skewness for a managed portfolio.
EXAMPLE 16 Calculating Skewness for a Mutual Fund
Table 28 presents 10 years of annual returns on the T. Rowe Price Equity Income Fund (PRFDX).
TABLE 28 Annual Rates of Return: T. Rowe Price Equity Income, 2003–2012
Year Return (%)
2003 25.78 2004 15.05 2005 4.26 2006 19.14 2007 3.30 2008 −35.75 2009 25.62 2010 15.15 2011 −0.72 2012 17.25
Source: performance.morningstar.com.
Using the information in Table 28, address the following: 1. Calculate the skewness of PRFDX showing two decimal places.
2. Characterize the shape of the distribution of PRFDX returns based on your answer to Part 1.
Solution to 1: To calculate skewness, we find the sum of the cubed deviations from the mean, divide by the standard deviation cubed, and then multiply that result by n/[(n − 1)(n − 2)]. Table 29 gives the calculations.
TABLE 29 Calculating Skewness for PRFDX
Year Rt 2003 25.78 16.87 4,801.150 2004 15.05 6.14 231.476 2005 4.26 −4.65 −100.545 2006 19.14 10.23 1,070.599 2007 3.30 −5.61 −176.558 2008 −35.75 −44.66 −89,075.067 2009 25.62 16.71 4,665.835 2010 15.15 6.24 242.971 2011 −0.72 −9.63 −893.056
2012 17.25 8.34 580.094 n = 10
8.91% Sum = −78,653.103
s = 18.12% s3 = 5,949.419 Sum/s3 = −13.2203
n/[(n − 1)(n − 2)] = 0.1389 Skewness = −1.84
Source: performance.morningstar.com.
Using Equation 17, the calculation is:
Solution to 2: Based on this small sample, the distribution of annual returns for the fund appears to be negatively skewed. In this example, four deviations are negative and six are positive. While there are more positive deviations, they are much more than offset by a huge negative deviation in 2008, when the stock markets sharply went down as a consequence of the global financial crisis. The result is that skewness is a negative number, implying that the distribution is skewed to the left.
9. KURTOSIS IN RETURN DISTRIBUTIONS In the previous section, we discussed how to determine whether a return distribution deviates from a normal distribution because of skewness. One other way in which a return distribution might differ from a normal distribution is by having more returns clustered closely around the mean (being more peaked) and more returns with large deviations from the mean (having fatter tails). Relative to a normal distribution, such a distribution has a greater percentage of small deviations from the mean return (more small surprises) and a greater percentage of extremely large deviations from the mean return (more big surprises). Most investors would perceive a greater chance of extremely large deviations from the mean as increasing risk.
Kurtosis is the statistical measure that tells us when a distribution is more or less peaked than a normal distribution. A distribution that is more peaked than normal is called leptokurtic (lepto from the Greek word for slender); a distribution that is less peaked than normal is called platykurtic (platy from the Greek word for broad); and a distribution identical to the normal distribution in this respect is called mesokurtic (meso from the Greek word for middle). The situation of more-frequent extremely large surprises that we described is one of leptokurtosis.45
Figure 8 illustrates a leptokurtic distribution. It is more peaked and has fatter tails than the normal distribution.
The calculation for kurtosis involves finding the average of deviations from the mean raised to the fourth power and then standardizing that average by dividing by the standard deviation raised to the fourth power.46 For all normal distributions, kurtosis is equal to 3. Many statistical packages report estimates of excess kurtosis, which is kurtosis minus 3.47 Excess kurtosis thus characterizes kurtosis relative to the normal distribution. A normal or other mesokurtic distribution has excess kurtosis equal to 0. A leptokurtic distribution has excess kurtosis greater than 0, and a platykurtic distribution has excess kurtosis less than 0. A return distribution with positive excess kurtosis—a leptokurtic return distribution—has more frequent extremely large deviations from the mean than a normal distribution. Below is the expression for computing kurtosis from a sample.
FIGURE 8 Leptokurtic: Fat Tailed
Source: Reprinted from Fixed Income Analysis. Copyright CFA Institute.
Sample Excess Kurtosis Formula. The sample excess kurtosis is
(18)
where n is the sample size and s is the sample standard deviation.
In Equation 18, sample kurtosis is the first term. Note that as n becomes large,
Equation 18 approximately equals . For a sample of 100 or larger taken from a normal distribution, a sample excess kurtosis of 1.0 or larger would be considered unusually large.
Most equity return series have been found to be leptokurtic. If a return distribution has positive excess kurtosis (leptokurtosis) and we use statistical models that do not account for the fatter tails, we will underestimate the likelihood of very bad or very good outcomes. For example, the return on the S&P 500 for 19 October 1987 was 20 standard deviations away from the mean daily return. Such an outcome is
possible with a normal distribution, but its likelihood is almost equal to 0. If daily returns are drawn from a normal distribution, a return four standard deviations or more away from the mean is expected once every 50 years; a return greater than five standard deviations away is expected once every 7,000 years. The return for October 1987 is more likely to have come from a distribution that had fatter tails than from a normal distribution. Looking at Table 27 given earlier, the monthly return series for the S&P 500 has very large excess kurtosis, approximately 9.4. It is extremely fat-tailed relative to the normal distribution. By contrast, the annual return series has about no excess kurtosis. The results for excess kurtosis in the table are consistent with research findings that the normal distribution is a better approximation for US equity returns for annual holding periods than for shorter ones (such as monthly).48
The following example illustrates the calculations for sample excess kurtosis for one of the two mutual funds we have been examining.
EXAMPLE 17 Calculating Sample Excess Kurtosis
Having concluded in Example 16 that the annual returns on T. Rowe Price Equity Income Fund were negatively skewed during the 2003–2012 period, what can we say about the kurtosis of the fund’s return distribution? Table 28 (repeated below) recaps the annual returns for the fund.
TABLE 28 Annual Rates of Return: T. Rowe Price Equity Income, 2003–2012 (Repeated)
Year Return (%) 2003 25.78 2004 15.05 2005 4.26 2006 19.14 2007 3.30 2008 −35.75 2009 25.62 2010 15.15 2011 −0.72 2012 17.25
Source: performance.morningstar.com.
Using the information from Table 28 repeated above, address the following: 1. Calculate the sample excess kurtosis of PRFDX showing two decimal places.
2. Characterize the shape of the distribution of PRFDX returns based on your answer to Part 1 as leptokurtic, mesokurtic, or platykurtic.
Solution to 1: To calculate excess kurtosis, we find the sum of the deviations from the mean raised to the fourth power, divide by the standard deviation raised to the fourth power, and then multiply that result by n(n + 1)/[(n − 1)(n − 2)(n − 3)]. This calculation determines kurtosis. Excess kurtosis is kurtosis minus 3(n − 1)2/[(n − 2)(n − 3)]. Table 30 gives the calculations.
TABLE 30 Calculating Kurtosis for PRFDX
Year Rt 2003 25.78 16.87 80,995.395 2004 15.05 6.14 1,421.260 2005 4.26 −4.65 467.533 2006 19.14 10.23 10,952.229 2007 3.30 −5.61 990.493 2008 −35.75 −44.66 3,978,092.479 2009 25.62 16.71 77,966.098 2010 15.15 6.24 1,516.137 2011 −0.72 −9.63 8,600.133 2012 17.25 8.34 4,837.981 n = 10
8.91% Sum = 4,165,839.738
s = 18.12% s4 = 107,803.478 Sum/s4 = 38.643
n(n + 1)/[(n − 1)(n − 2)(n −3)] = 0.2183 Kurtosis = 8.434
3(n − 1)2/[(n − 2)(n −3)] = 4.34
Excess Kurtosis = 4.09
Source: performance.morningstar.com.
Using Equation 18, the calculation is
Solution to 2: The distribution of PRFDX’s annual returns appears to be leptokurtic, based on a positive sample excess kurtosis. The fairly large excess kurtosis of 4.09 indicates that the distribution of PRFDX’s annual returns is fat- tailed relative to the normal distribution. With a negative skewness and a positive excess kurtosis, PRFDX’s annual returns do not appear to have been normally distributed during the period.49
10. USING GEOMETRIC AND ARITHMETIC MEANS With the concepts of descriptive statistics in hand, we will see why the geometric mean is appropriate for making investment statements about past performance. We will also explore why the arithmetic mean is appropriate for making investment statements in a forward-looking context.
For reporting historical returns, the geometric mean has considerable appeal because it is the rate of growth or return we would have had to earn each year to match the actual, cumulative investment performance. In our simplified Example 8, for instance, we purchased a stock for €100 and two years later it was worth €100, with an intervening year at €200. The geometric mean of 0 percent is clearly the compound rate of growth during the two years. Specifically, the ending amount is the beginning amount times (1 + RG)2. The geometric mean is an excellent measure of past performance.
Example 8 illustrated how the arithmetic mean can distort our assessment of historical performance. In that example, the total performance for the two-year period was unambiguously 0 percent. With a 100 percent return for the first year and −50 percent for the second, however, the arithmetic mean was 25 percent. As we noted previously, the arithmetic mean is always greater than or equal to the geometric mean. If we want to estimate the average return over a one-period horizon, we should use the arithmetic mean because the arithmetic mean is the average of one-period returns. If we want to estimate the average returns over more than one period, however, we should use the geometric mean of returns because the geometric mean captures how the total returns are linked over time.
As a corollary to using the geometric mean for performance reporting, the use of semilogarithmic rather than arithmetic scales is more appropriate when graphing past performance.50 In the context of reporting performance, a semilogarithmic graph has an arithmetic scale on the horizontal axis for time and a logarithmic scale on the vertical axis for the value of the investment. The vertical axis values are spaced according to the differences between their logarithms. Suppose we want to represent £1, £10, £100, and £1,000 as values of an investment on the vertical axis. Note that each successive value represents a 10-fold increase over the previous value, and each will be equally spaced on the vertical axis because the difference in their logarithms is roughly 2.30; that is, ln 10 − ln 1 = ln 100 − ln 10 = ln 1,000 − ln 100 = 2.30. On a semilogarithmic scale, equal movements on the vertical axis reflect equal percentage changes, and growth at a constant compound rate plots as a
straight line. A plot curving upward reflects increasing growth rates over time. The slopes of a plot at different points may be compared in order to judge relative growth rates.
In addition to reporting historical performance, financial analysts need to calculate expected equity risk premiums in a forward-looking context. For this purpose, the arithmetic mean is appropriate.
We can illustrate the use of the arithmetic mean in a forward-looking context with an example based on an investment’s future cash flows. In contrasting the geometric and arithmetic means for discounting future cash flows, the essential issue concerns uncertainty. Suppose an investor with $100,000 faces an equal chance of a 100 percent return or a −50 percent return, represented on the tree diagram as a 50/50 chance of a 100 percent return or a −50 percent return per period. With 100 percent return in one period and −50 percent return in the other, the geometric mean return is .
The geometric mean return of 0 percent gives the mode or median of ending wealth after two periods and thus accurately predicts the modal or median ending wealth of $100,000 in this example. Nevertheless, the arithmetic mean return better predicts the arithmetic mean ending wealth. With equal chances of 100 percent or −50 percent returns, consider the four equally likely outcomes of $400,000, $100,000, $100,000, and $25,000 as if they actually occurred. The arithmetic mean ending wealth would be $156,250 = ($400,000 + $100,000 + $100,000 + $25,000)/4. The actual returns would be 300 percent, 0 percent, 0 percent, and −75 percent for a two- period arithmetic mean return of (300 + 0 + 0 −75)/4 = 56.25 percent. This arithmetic mean return predicts the arithmetic mean ending wealth of $100,000 × 1.5625 = $156,250. Noting that 56.25 percent for two periods is 25 percent per period, we then must discount the expected terminal wealth of $156,250 at the 25 percent arithmetic mean rate to reflect the uncertainty in the cash flows.
Uncertainty in cash flows or returns causes the arithmetic mean to be larger than the geometric mean. The more uncertain the returns, the more divergence exists between the arithmetic and geometric means. The geometric mean return approximately equals the arithmetic return minus half the variance of return.51 Zero variance or zero uncertainty in returns would leave the geometric and arithmetic return approximately equal, but real-world uncertainty presents an arithmetic mean
return larger than the geometric. For example, for the nominal annual returns on S&P 500 from 1926 to 2012, Table 27 reports an arithmetic mean of 11.82 percent and standard deviation of 20.18 percent. The geometric mean of these returns is 9.84 percent. We can see the geometric mean is approximately the arithmetic mean minus half of the variance of returns: RG ≈ 0.1182 − (1/2)(0.20182) = 0.0978, or 9.78 percent.
11. SUMMARY In this reading, we have presented descriptive statistics, the set of methods that permit us to convert raw data into useful information for investment analysis.
A population is defined as all members of a specified group. A sample is a subset of a population.
A parameter is any descriptive measure of a population. A sample statistic (statistic, for short) is a quantity computed from or used to describe a sample.
Data measurements are taken using one of four major scales: nominal, ordinal, interval, or ratio. Nominal scales categorize data but do not rank them. Ordinal scales sort data into categories that are ordered with respect to some characteristic. Interval scales provide not only ranking but also assurance that the differences between scale values are equal. Ratio scales have all the characteristics of interval scales as well as a true zero point as the origin. The scale on which data are measured determines the type of analysis that can be performed on the data.
A frequency distribution is a tabular display of data summarized into a relatively small number of intervals. Frequency distributions permit us to evaluate how data are distributed.
The relative frequency of observations in an interval is the number of observations in the interval divided by the total number of observations. The cumulative relative frequency cumulates (adds up) the relative frequencies as we move from the first interval to the last, thus giving the fraction of the observations that are less than the upper limit of each interval.
A histogram is a bar chart of data that have been grouped into a frequency distribution. A frequency polygon is a graph of frequency distributions obtained by drawing straight lines joining successive points representing the class frequencies.
Sample statistics such as measures of central tendency, measures of dispersion, skewness, and kurtosis help with investment analysis, particularly in making probabilistic statements about returns.
Measures of central tendency specify where data are centered and include the (arithmetic) mean, median, and mode (most frequently occurring value). The mean is the sum of the observations divided by the number of observations. The median is the value of the middle item (or the mean of the values of the
two middle items) when the items in a set are sorted into ascending or descending order. The mean is the most frequently used measure of central tendency. The median is not influenced by extreme values and is most useful in the case of skewed distributions. The mode is the only measure of central tendency that can be used with nominal data.
A portfolio’s return is a weighted mean return computed from the returns on the individual assets, where the weight applied to each asset’s return is the fraction of the portfolio invested in that asset.
The geometric mean, G, of a set of observations X1, X2, …, Xn is G = with for i = 1, 2, …, n. The geometric mean is especially
important in reporting compound growth rates for time series data.
Quantiles such as the median, quartiles, quintiles, deciles, and percentiles are location parameters that divide a distribution into halves, quarters, fifths, tenths, and hundredths, respectively.
Dispersion measures such as the variance, standard deviation, and mean absolute deviation (MAD) describe the variability of outcomes around the arithmetic mean.
Range is defined as the maximum value minus the minimum value. Range has only a limited scope because it uses information from only two observations.
MAD for a sample is where is the sample mean and n is the number of observations in the sample.
The variance is the average of the squared deviations around the mean, and the standard deviation is the positive square root of variance. In computing sample variance (s2) and sample standard deviation, the average squared deviation is computed using a divisor equal to the sample size minus 1.
The semivariance is the average squared deviation below the mean; semideviation is the positive square root of semivariance. Target semivariance is the average squared deviation below a target level; target semideviation is its positive square root. All these measures quantify downside risk.
According to Chebyshev’s inequality, the proportion of the observations within k standard deviations of the arithmetic mean is at least 1 − 1/k2 for all k > 1. Chebyshev’s inequality permits us to make probabilistic statements about the proportion of observations within various intervals around the mean for any distribution with finite variance. As a result of Chebyshev’s inequality, a two-
standard-deviation interval around the mean must contain at least 75 percent of the observations, and a three-standard-deviation interval around the mean must contain at least 89 percent of the observations, no matter how the data are distributed.
The coefficient of variation, CV, is the ratio of the standard deviation of a set of observations to their mean value. A scale-free measure of relative dispersion, by expressing the magnitude of variation among observations relative to their average size, the CV permits direct comparisons of dispersion across different data sets.
The Sharpe ratio for a portfolio, p, based on historical returns, is defined as
,
where is the mean return to the portfolio, is the mean return to a risk-free asset, and sp is the standard deviation of return on the portfolio.
Skew describes the degree to which a distribution is not symmetric about its mean. A return distribution with positive skewness has frequent small losses and a few extreme gains. A return distribution with negative skewness has frequent small gains and a few extreme losses. Zero skewness indicates a symmetric distribution of returns.
Kurtosis measures the peakedness of a distribution and provides information about the probability of extreme outcomes. A distribution that is more peaked than the normal distribution is called leptokurtic; a distribution that is less peaked than the normal distribution is called platykurtic; and a distribution identical to the normal distribution in this respect is called mesokurtic. The calculation for kurtosis involves finding the average of deviations from the mean raised to the fourth power and then standardizing that average by the standard deviation raised to the fourth power. Excess kurtosis is kurtosis minus 3, the value of kurtosis for all normal distributions.
REFERENCES
1. Amin, Gaurav, and Harry Kat. 2003. “Stocks, Bonds, and Hedge Funds.” Journal of Portfolio Management, vol. 29, no. 4: 113–120.
2. Black, Fischer. 1993. “Estimating Expected Return.” Financial Analysts Journal, vol. 49, no. 5: 36–38.
3. Bodie, Zvi, Alex Kane, and Alan J. Marcus. 2012. Essentials of Investments, 9th edition. New York: McGraw-Hill Irwin.
4. Campbell, Stephen K. 1974. Flaws and Fallacies in Statistical Thinking. Englewood Cliffs, NJ: Prentice-Hall.
5. Campbell, John, Andrew Lo, and A. Craig MacKinlay. 1997. The Econometrics of Financial Markets. Princeton, NJ: Princeton University Press.
6. Choobinbeh, N. Fred. 2005. “Semivariance” in The Encyclopedia of Statistical Sciences. Hoboken, NJ: Wiley.
7. Dimson, Elroy, Paul Marsh, and Mike Staunton. 2011. “Equity Premiums around the World” in Rethinking the Equity Risk Premium. Charlottesville, VA: Research Foundation of CFA Institute.
8. Elton, Edwin J., Martin J. Gruber, Stephen J. Brown, and William N. Goetzmann. 2013. Modern Portfolio Theory and Investment Analysis, 9th edition. Hoboken, NJ: Wiley.
9. Estrada, Javier. 2003. “Mean–Semivariance Behavior: A Note.” Finance Letters, vol. 1, no. 1: 9–14.
10. Fabozzi, Frank J. 2007. Fixed Income Analysis, 2nd edition. Hoboken, NJ: Wiley.
11. Forbes. 16 September 2013. “The Honor Roll.” New York: Forbes Management Co., Inc.
12. Gujarati, Damodar N., and Dawn C. Porter. 2008. Basic Econometrics, 4th edition. New York: McGraw-Hill Irwin.
13. Ibbotson, Roger G., Zhiwu Chen, Daniel Y.-J. Kim, and Wendy Y. Hu. 2013. “Liquidity as an Investment Style.” Financial Analysts Journal, vol. 69, no. 3:
30–44.
14. Pinto, Jerald E., Elaine Henry, Thomas R. Robinson, and John D. Stowe. 2010. Equity Asset Valuation, 2nd edition. Charlottesville, VA: CFA Institute.
15. Reilly, Frank K., and Keith C. Brown. 2012. Investment Analysis and Portfolio Management, 10th edition. Mason, OH: Cengage South-Western.
16. Sharpe, William. 1994. “The Sharpe Ratio.” Journal of Portfolio Management, vol. 21, no. 1: 49–59.
PROBLEMS Practice Problems and Solutions: 1–19 taken from Quantitative Methods for Investment Analysis, Second Edition, by Richard A. DeFusco, CFA, Dennis W. McLeavey, CFA, Jerald E. Pinto, CFA, and David E. Runkle, CFA. Copyright © 2004 by CFA Institute. All other problems and solutions copyright © CFA Institute. 1. Identify each of the following groups as a population or a sample. If the group
is a sample, identify the population to which the sample is related.
The S&P MidCap 400 Index viewed as representing US stocks with market capitalization falling within a stated range.
UK shares that traded on 11 August 2003 and that also closed above £100/share as of the close of the London Stock Exchange on that day.
C. Marsh & McLennan Companies, Inc. (NYSE: MMC) and AON Corporation (NYSE: AOC). This group is part of Standard & Poor ’s Insurance Brokers Index.
The set of 31 estimates for Microsoft EPS for fiscal year 2003, as of the 4 June 2003 date of a First Call/Thomson Financial report.
2. State the type of scale used to measure the following sets of data.
Sales in euros.
The investment style of mutual funds.
An analyst’s rating of a stock as underweight, market weight, or overweight, referring to the analyst’s suggested weighting of the stock in a portfolio.
A measure of the risk of portfolios on a scale of whole numbers from 1 (very conservative) to 5 (very risky) where the difference between 1 and 2 represents the same increment in risk as the difference between 4 and 5.
The following information relates to Questions 3–4
The table below gives the deviations of a hypothetical portfolio’s annual total returns (gross of fees) from its benchmark’s annual returns, for a 12- year period ending in 2003.
Portfolio’s Deviations from Benchmark Return,1992–2003 (%) 1992 −7.14
1993 1.62 1994 2.48 1995 −2.59 1996 9.37 1997 −0.55 1998 −0.89 1999 −9.19 2000 −5.11 2001 −0.49 2002 6.84 2003 3.04
3.
Calculate the frequency, cumulative frequency, relative frequency, and cumulative relative frequency for the portfolio’s deviations from benchmark return, given the set of intervals in the table that follows.
Return Interval Frequency
Cumulative Frequency
Relative Frequency
(%)
Cumulative Relative Frequency (%)
−9.19 ≤ A < −4.55 −4.55 ≤ B < 0.09
0.09 ≤ C < 4.73
4.73 ≤ D ≤ 9.37
Construct a histogram using the data.
Identify the modal interval of the grouped data.
4. Tracking risk is the standard deviation of the deviation of a portfolio’s gross- of-fees total returns from benchmark return. Calculate the tracking risk of the portfolio, stated in percent (give the answer to two decimal places).
The following information relates to Questions 5–10
The table below gives the annual total returns on the MSCI Germany Index
from 1993 to 2002. The returns are in the local currency. Use the information in this table to answer Questions 5–10.
MSCI Germany Index Total Returns, 1993–2002 Year Return (%) 1993 46.21 1994 −6.18 1995 8.04 1996 22.87 1997 45.90 1998 20.32 1999 41.20 2000 −9.53 2001 −17.75 2002 −43.06
Source: Ibbotson EnCorr AnalyzerTM.
5. To describe the distribution of observations, perform the following:
Create a frequency distribution with five equally spaced classes (round up at the second decimal place in computing the width of class intervals).
Calculate the cumulative frequency of the data.
Calculate the relative frequency and cumulative relative frequency of the data.
State whether the frequency distribution is symmetric or asymmetric. If the distribution is asymmetric, characterize the nature of the asymmetry.
6. To describe the central tendency of the distribution, perform the following:
Calculate the sample mean return.
Calculate the median return.
Identify the modal interval (or intervals) of the grouped returns.
7. To describe the compound rate of growth of the MSCI Germany Index, calculate the geometric mean return.
8. To describe the values at which certain returns fall, calculate the 30th percentile.
9. To describe the dispersion of the distribution, perform the following:
Calculate the range.
Calculate the mean absolute deviation (MAD).
Calculate the sample variance.
Calculate the sample standard deviation.
Calculate the semivariance.
Calculate the semideviation.
10. To describe the degree to which the distribution may depart from normality, perform the following:
Calculate the skewness.
Explain the finding for skewness in terms of the location of the median and mean returns.
Calculate excess kurtosis.
Contrast the distribution of annual returns on the MSCI Germany Index to a normal distribution model for returns.
11.
Explain the relationship among arithmetic mean return, geometric mean return, and variability of returns.
Contrast the use of the arithmetic mean return to the geometric mean return of an investment from the perspective of an investor concerned with the investment’s terminal value.
Contrast the use of the arithmetic mean return to the geometric mean return of an investment from the perspective of an investor concerned with the investment’s average one-year return.
The following information relates to Questions 12–14
The following table repeats the annual total returns on the MSCI Germany Index previously given and also gives the annual total returns on the JP Morgan Germany five-to seven-year government bond index (JPM 5–7 Year GBI, for short). During the period given in the table, the International Monetary Fund Germany Money Market Index (IMF Germany MMI, for short) had a mean annual total return of 4.33 percent. Use that information and the information in the table to answer Questions 12–14.
Year MSCI Germany Index (%) JPM Germany 5–7 Year GBI (%) 1993 46.21 15.74 1994 −6.18 −3.40 1995 8.04 18.30 1996 22.87 8.35 1997 45.90 6.65 1998 20.32 12.45 1999 41.20 −2.19 2000 −9.53 7.44 2001 −17.75 5.55 2002 −43.06 10.27
Source: Ibbotson EnCorr AnalyzerTM.
12. Calculate the annual returns and the mean annual return on a portfolio 60 percent invested in the MSCI Germany Index and 40 percent invested in the JPM Germany GBI.
13.
Calculate the coefficient of variation for:
the 60/40 equity/bond portfolio described in Problem 12.
the MSCI Germany Index.
the JPM Germany 5–7 Year GBI.
Contrast the risk of the 60/40 equity/bond portfolio, the MSCI Germany Index, and the JPM Germany 5–7 Year GBI, as measured by the coefficient of variation.
14.
Using the IMF Germany MMI as a proxy for the risk-free return, calculate the Sharpe ratio for:
the 60/40 equity/bond portfolio described in Problem 12.
the MSCI Germany Index.
the JPM Germany 5–7 Year GBI.
Contrast the risk-adjusted performance of the 60/40 equity/bond portfolio, the MSCI Germany Index, and the JPM Germany 5–7 Year GBI, as
measured by the Sharpe ratio.
15. Suppose a client asks you for a valuation analysis on the eight-stock US common stock portfolio given in the table below. The stocks are equally weighted in the portfolio. You are evaluating the portfolio using three price multiples. The trailing 12 months (TTM) price-to-earnings ratio (P/E) is current price divided by diluted EPS over the past four quarters.1 The TTM price-to-sales ratio (P/S) is current price divided by sales per share over the last four quarters. The price-to-book ratio (P/B) is the current price divided by book value per share as given in the most recent quarterly statement. The data in the table are as of 12 September 2003.
Client Portfolio Common Stock TTM P/E TTM P/S P/B
Abercrombie & Fitch (NYSE: AFN) 13.67 1.66 3.43 Albemarle Corporation (NYSE: ALB) 14.43 1.13 1.96 Avon Products, Inc. (NYSE: AVP) 28.06 2.45 382.72
Berkshire Hathaway (NYSE: BRK.A) 18.46 2.39 1.65 Everest Re Group Ltd (NYSE: RE) 11.91 1.34 1.30 FPL Group, Inc. (NYSE: FPL) 15.80 1.04 1.70
Johnson Controls, Inc. (NYSE: JCI) 14.24 0.40 2.13 Tenneco Automotive, Inc. (NYSE: TEN) 6.44 0.07 41.31
Source: www.multexinvestor.com.
Based only on the information in the above table, calculate the following for the portfolio:
Arithmetic mean P/E.
Median P/E.
Arithmetic mean P/S.
Median P/S.
Arithmetic mean P/B.
Median P/B.
Based on your answers to Parts A, B, and C, characterize the appropriateness of using the following valuation measures:
Mean and median P/E.
Mean and median P/S.
Mean and median P/B.
16. The table below gives statistics relating to a hypothetical 10-year record of two portfolios.
Mean AnnualReturn (%)
Standard Deviation of Return (%) Skewness
Portfolio A 8.3 19.5 –1.9
Portfolio B 8.3 18.0 3.0
Based only on the information in the above table, perform the following:
Contrast the distributions of returns of Portfolios A and B.
Evaluate the relative attractiveness of Portfolios A and B.
17. The table below gives statistics relating to a hypothetical three-year record of two portfolios.
Mean Monthly Return (%)
Standard Deviation (%) Skewness
Excess Kurtosis
Portfolio A 1.1994 5.5461 −2.2603 6.2584
Portfolio B 1.1994 6.4011 −2.2603 8.0497
Based only on the information in the above table, perform the following:
Contrast the distributions of returns of Portfolios A and B.
Evaluate the relative attractiveness of Portfolios A and B.
18. The table below gives statistics relating to a hypothetical five-year record of two portfolios.
Mean Monthly Return (%)
Standard Deviation (%) Skewness
Excess Kurtosis
Portfolio
A 1.6792 5.3086 −0.1395 −0.0187
Portfolio B 1.8375 5.9047 0.4934 −0.8525
Based only on the information in the above table, perform the following:
Contrast the distributions of returns of Portfolios A and B.
Evaluate the relative attractiveness of Portfolios A and B.
19. At the UXI Foundation, portfolio managers are normally kept on only if their annual rate of return meets or exceeds the mean annual return for portfolio managers of a similar investment style. Recently, the UXI Foundation has also been considering two other evaluation criteria: the median annual return of funds with the same investment style, and two-thirds of the return performance of the top fund with the same investment style.
The table below gives the returns for nine funds with the same investment style as the UXI Foundation.
Fund Return (%) 1 17.8 2 21.0 3 38.0 4 19.2 5 2.5 6 24.3 7 18.7 8 16.9 9 12.6
With the above distribution of fund performance, which of the three evaluation criteria is the most difficult to achieve?
20. If the observations in a data set have different values, is the geometric mean for that data set less than that data set’s:
harmonic mean? arithmetic mean? A. No No B. No Yes
C. Yes No
21. Is a return distribution characterized by frequent small losses and a few large gains best described as having:
negative skew? a mean that is greater than the median? A. No No B. No Yes C. Yes No
22. An analyst gathered the following information about the return distributions for two portfolios during the same time period:
Skewness Kurtosis Portfolio A −1.3 2.2 Portfolio B 0.5 3.5
The analyst stated that the distribution for Portfolio A is more peaked than a normal distribution and that the distribution for Portfolio B has a long tail on the left side of the distribution. Which of the following is most true?
The statement is not correct in reference to either portfolio.
The statement is correct in reference to Portfolio A, but the statement is not correct in reference to Portfolio B.
The statement is not correct in reference to Portfolio A, but the statement is correct in reference to Portfolio B.
23. The coefficient of variation is useful in determining the relative degree of variability of different data sets if those data sets have different:
means or different units of measurement.
means, but not different units of measurement.
units of measurement, but not different means.
24. An analyst gathered the following information: Portfolio Mean Return (%) Standard Deviation of Returns (%)
1 9.8 19.9 2 10.5 20.3 3 13.3 33.9
If the risk-free rate of return is 3.0 percent, the portfolio that had the best risk-adjusted performance based on the Sharpe ratio is:
Portfolio 1.
Portfolio 2.
Portfolio 3.
25. An analyst gathered the following information about a portfolio’s performance over the past 10 years:
Mean annual return 11.8% Standard deviation of annual returns 15.7%
Portfolio beta 1.2
If the mean return on the risk-free asset over the same period was 5.0%, the coefficient of variation and Sharpe ratio, respectively, for the portfolio are closest to:
Coefficient of variation Sharpe ratio A. 0.75 0.43 B. 1.33 0.36 C. 1.33 0.43
Notes 1 Ibbotson Associates (www.ibbotson.com) generously provided some of the data
used in this reading. We also draw on Dimson, Marsh, and Staunton’s (2011) history and study of world markets as well as other sources.
2 This reading introduces many statistical concepts and formulas. To make it easy to locate them, we have set off some of the more important ones with bullet points.
3 Credit ratings for a bond issue gauge the bond issuer ’s ability to meet the promised principal and interest payments on the bond. For example, one rating agency, Standard & Poor ’s, assigns bond issues to one of the following ratings, given in descending order of credit quality (increasing probability of default): AAA, AA+, AA, AA−, A+, A, A−, BBB+, BBB, BBB−, BB+, BB, BB−, B+, B, B −, CCC+, CCC−, CC, C, D. For more information on credit risk and credit ratings, see Fabozzi (2007).
4 “Hedge fund” refers to investment vehicles with legal structures that result in less regulatory oversight than other pooled investment vehicles such as mutual funds. Hedge fund classification types group hedge funds by the kind of investment strategy they pursue.
5 Note, however, that if price and cash distributions in the expression for holding period return were not in one’s home currency, one would generally convert those variables to one’s home currency before calculating the holding period return. Because of exchange rate fluctuations during the holding period, holding period returns on an asset computed in different currencies would generally differ.
6 We use the total return series on the S&P 500 from January 1926 to December 2012 provided by Ibbotson Associates.
7 Intervals are also sometimes called classes, ranges, or bins.
8 The notation [−4.57 to −0.57) means −4.57 ≤ observation < −0.57. In this context, a square bracket indicates that the endpoint is included in the interval.
9 The average or arithmetic mean of a set of values equals the sum of the values divided by the number of values summed. To find the arithmetic mean of 111 annual returns, for example, we sum the 111 annual returns and then divide the total by 111. Among the most familiar of statistical concepts, the arithmetic mean
is explained in more detail laterin the reading.
10 Even though the upper limit on the interval is not a return falling in the interval, we still average it with the lower limit to determine the midpoint.
11 A wholesale club implements a store format dedicated mostly to bulk sales in warehouse-sized stores to customers who pay membership dues. As of the early 2010s, those three wholesale clubs dominated the segment in the United States.
12 Statisticians prefer the term “mean” to “average.” Some writers refer to all measures of central tendency (including the median and mode) as averages. The term “mean” avoids any possibility of confusion.
13 The term “free float adjusted” means that the weights of companies in the index reflect the value of the shares actually available for investment.
14 Other approaches to handling extreme values involve variations of the arithmetic mean. The trimmed mean is computed by excluding a stated small percentage of the lowest and highest values and then computing an arithmetic mean of the remaining values. For example, a 5 percent trimmed mean discards the lowest 2.5 percent and the largest 2.5 percent of values and computes the mean of the remaining 95 percent of values. A trimmed mean is used in sports competitions when judges’ lowest and highest scores are discarded in computing a contestant’s score. A Winsorized mean assigns a stated percent of the lowest values equal to one specified low value, and a stated percent of the highest values equal to one specified high value, then computes a mean from the restated data. For example, a 95 percent Winsorized mean sets the bottom 2.5 percent of values equal to the 2.5th percentile value and the upper 2.5 percent of values equal to the 97.5th percentile value. (Percentile values are defined later.)
15 The notation Md is occasionally used for the median. Just as for the mean, we may distinguish between a population median and a sample median. With the understanding that a population median divides a population in half while a sample median divides a sample in half, we follow general usage in using the term “median” without qualification, for the sake of brevity.
16 For more information on price multiples, see Pinto, Henry, Robinson, and Stowe (2010).
17 The notation Mo is occasionally used for the mode. Just as for the mean and the median, we may distinguish between a population mode and a sample mode. With the understanding that a population mode is the value with the greatest
probability of occurrence, while a sample mode is the most frequently occurring value in the sample, we follow general usage in using the term “mode” without qualification, for the sake of brevity.
18 For more information on credit risk and credit ratings, see Fabozzi (2007).
19 The formula for the weighted mean can be compared to the formula for the arithmetic mean. For a set of observations X1, X2, …, Xn, let the weights w1, w2, …, wn all equal 1/n. Under this assumption, the formula for the weighted mean is
. This is the formula for the arithmetic mean. Therefore, the arithmetic mean is a special case of the weighted mean in which all the weights are equal.
20 In Table 14, strategic investments include investments in property, private investments, and hedge fund investments. Bond overlay consists of derivatives used to hedge interest rate and inflation changes.
21 This statement can be proved using Jensen’s inequality that the average value of a function is less than or equal to the function evaluated at the mean if the function is concave from below—the case for ln(X).
22 For instance, suppose the return for each of the three years is 10 percent. The arithmetic mean is 10 percent. To find the geometric mean, we first express the returns as (1 + Rt) and then find the geometric mean: [(1.10)(1.10)(1.10)]1/3 − 1.0 = 10 percent. The two means are the same.
23 We will soon introduce standard deviation as a measure of variability. Holding the arithmetic mean return constant, the geometric mean return decreases for an increase in standard deviation.
24 We will introduce formal measures of variability later. But note, for example, the 71.08 percentage point swing in returns between 2008 and 2009 for SLASX versus the 61.37 percentage point for PRFDX. Similarly, note the 19.11 percentage point swing in returns between 2009 and 2010 for SLASX versus the 10.47 percentage point for PRFDX.
25 The terminology “harmonic” arises from its use relative to a type of series involving reciprocals known as a harmonic series.
26 Black (1993).
27 Another distance measure of dispersion that we may encounter, the interquartile
range, focuses on the middle rather than the extremes. The interquartile range (IQR) is the difference between the third and first quartiles of a data set: IQR = Q3 − Q1. The IQR represents the length of the interval containing the middle 50 percent of the data, with a larger interquartile range indicating greater dispersion, all else equal.
28 In some analytic work such as optimization, the calculus operation of differentiation is important. Variance as a function can be differentiated, but absolute value cannot.
29 Forbes magazine annually selects US equity mutual funds meeting certain criteria for its Honor Roll. The criteria relate to capital preservation (performance in bear markets), continuity of management (the fund must have a manager with at least six years’ tenure), diversification, accessibility (disqualifying funds that are closed to new investors), and after-tax long-term performance.
30 In fact, we could not properly use the Honor Roll funds to estimate the population variance of portfolio turnover (for example) of any other differently defined population, because the Honor Roll funds are not a random sample from any larger population of US equity mutual funds.
31 We discuss this concept further in the reading on sampling.
32 This is an informal treatment of these two measures; see the survey article by N. Fred Choobinbeh (2005) for a rigorous treatment.
33 See Estrada (2003). We discuss skewness later in this reading.
34 As discussed in the reading on probability concepts and the various readings on portfolio concepts, we can find a portfolio’s variance as a straightforward function of the variances and correlations of the component securities. There is no similar procedure for semivariance and target semivariance. We also cannot take the derivative of semivariance or target semivariance.
35 The second sample was created by adding €750 million to each observation in the first sample. Standard deviation (and variance) has the property of remaining unchanged if we add a constant amount to each observation.
36 The reader will also encounter CV defined as 100(s/ ), which states CV as a percentage.
37 The equation presents the ex post or historical Sharpe ratio. We can also think of
the Sharpe ratio for a portfolio going forward based on our expectations for mean return, the risk-free return, and the standard deviation of return; this would be the ex ante Sharpe ratio. One may also encounter an alternative calculation for the Sharpe ratio in which the denominator is the standard deviation of the series (Portfolio return – Risk-free return) rather than the standard deviation of portfolio return; in practice, the two standard deviation calculations generally yield very similar results. For more information on the Sharpe ratio (which has also been called the Sharpe measure, the reward-to-variability ratio, and the excess return to variability measure), see Elton, Gruber, Brown, and Goetzmann (2013) and Sharpe (1994).
38 If the standard deviations are equal, however, the portfolio with the negative Sharpe ratio closer to zero is superior.
39 This statement describes a return distribution with negative skewness. We discuss skewness later in this reading.
40 For more information, see Amin and Kat (2003).
41 As a mnemonic, in this case the mean, median, and mode occur in the same order as they would be listed in a dictionary.
42 We are discussing a moment coefficient of skewness. Some textbooks present the Pearson coefficient of skewness, equal to 3(Mean − Median)/Standard deviation, which has the drawback of involving the calculation of the median.
43 The term n/[(n − 1)(n − 2)] in Equation 17 corrects for a downward bias in small samples.
44 For more on the role of skewness in portfolio selection, see Reilly and Brown (2012) and Elton et al. (2013) and the references therein.
45 Kurtosis has been described as an illness characterized by episodes of extremely rude behavior.
46 This measure is free of scale. It is always positive because the deviations are raised to the fourth power.
47 Ibbotson and some software packages, such as Microsoft Excel, label “excess kurtosis” as simply “kurtosis.” This highlights the fact that one should familiarize oneself with the description of statistical quantities in any software packages that one uses.
48 See Campbell, Lo, and MacKinlay (1997) for more details.
49 It is useful to know that we can conduct a Jarque–Bera (JB) statistical test of normality based on sample size n, sample skewness, and sample excess kurtosis. We can conclude that a distribution is not normal with no more than a 5 percent chance of being wrong if the quantity is 6 or greater for a sample with at least 30 observations. In this mutual fund example, we have only 10 observations and the test described is only correct based on large samples (as a guideline, for n ≥ 30). Gujarati and Porter (2008) provides more details on this test.
50 See Campbell (1974) for more information.
51 See Bodie, Kane, and Marcus (2012).
1 In particular, diluted EPS is for continuing operations and before extraordinary items and accounting changes.
CHAPTER 4 PROBABILITY CONCEPTS Richard A. DeFusco, CFA
Dennis W. McLeavey, CFA
Jerald E. Pinto, PhD, CFA
David E. Runkle, PhD, CFA
LEARNING OUTCOMES
After completing this chapter, you will be able to do the following:
define a random variable, an outcome, an event, mutually exclusive events, and exhaustive events;
state the two defining properties of probability and distinguish among empirical, subjective, and a priori probabilities;
state the probability of an event in terms of odds for and against the event;
distinguish between unconditional and conditional probabilities;
explain the multiplication, addition, and total probability rules;
calculate and interpret 1) the joint probability of two events, 2) the probability that at least one of two events will occur, given the probability of each and the joint probability of the two events, and 3) a joint probability of any number of independent events;
distinguish between dependent and independent events;
calculate and interpret an unconditional probability using the total probability rule;
explain the use of conditional expectation in investment applications;
explain the use of a tree diagram to represent an investment problem;
calculate and interpret covariance and correlation;
calculate and interpret the expected value, variance, and standard deviation of a random variable and of returns on a portfolio;
calculate and interpret covariance given a joint probability function;
calculate and interpret an updated probability using Bayes’ formula;
identify the most appropriate method to solve a particular counting problem and solve counting problems using factorial, combination, and permutation concepts.
1. Introduction All investment decisions are made in an environment of risk. The tools that allow us to make decisions with consistency and logic in this setting come under the heading of probability. This reading presents the essential probability tools needed to frame and address many real-world problems involving risk. We illustrate how these tools apply to such issues as predicting investment manager performance, forecasting financial variables, and pricing bonds so that they fairly compensate bondholders for default risk. Our focus is practical. We explore in detail the concepts that are most important to investment research and practice. One such concept is independence, as it relates to the predictability of returns and financial variables. Another is expectation, as analysts continually look to the future in their analyses and decisions. Analysts and investors must also cope with variability. We present variance, or dispersion around expectation, as a risk concept important in investments. The reader will acquire specific skills in using portfolio expected return and variance.
The basic tools of probability, including expected value and variance, are set out in Section 2 of this reading. Section 3 introduces covariance and correlation (measures of relatedness between random quantities) and the principles for calculating portfolio expected return and variance. Two topics end the reading: Bayes’ formula and outcome counting. Bayes’ formula is a procedure for updating beliefs based on new information. In several areas, including a widely used option-pricing model, the calculation of probabilities involves defining and counting outcomes. The reading ends with a discussion of principles and shortcuts for counting.
2. Probability, Expected Value, and Variance The probability concepts and tools necessary for most of an analyst’s work are relatively few and simple but require thought to apply. This section presents the essentials for working with probability, expectation, and variance, drawing on examples from equity and fixed income analysis.
An investor ’s concerns center on returns. The return on a risky asset is an example of a random variable, a quantity whose outcomes (possible values) are uncertain. For example, a portfolio may have a return objective of 10 percent a year. The portfolio manager ’s focus at the moment may be on the likelihood of earning a return that is less than 10 percent over the next year. Ten percent is a particular value or outcome of the random variable “portfolio return.” Although we may be concerned about a single outcome, frequently our interest may be in a set of outcomes: The concept of “event” covers both.
Definition of Event. An event is a specified set of outcomes.
We may specify an event to be a single outcome—for example, the portfolio earns a return of 10 percent. (We use italics to highlight statements that define events.) We can capture the portfolio manager ’s concerns by defining the event as the portfolio earns a return below 10 percent. This second event, referring as it does to all possible returns greater than or equal to −100 percent (the worst possible return) but less than 10 percent, contains an infinite number of outcomes. To save words, it is common to use a capital letter in italics to represent a defined event. We could define A = the portfolio earns a return of 10 percent and B = the portfolio earns a return below 10 percent.
To return to the portfolio manager ’s concern, how likely is it that the portfolio will earn a return below 10 percent?
The answer to this question is a probability: a number between 0 and 1 that measures the chance that a stated event will occur. If the probability is 0.40 that the portfolio earns a return below 10 percent, there is a 40 percent chance of that event happening. If an event is impossible, it has a probability of 0. If an event is certain to happen, it has a probability of 1. If an event is impossible or a sure thing, it is not random at all. So, 0 and 1 bracket all the possible values of a probability.
Probability has two properties, which together constitute its definition.
Definition of Probability. The two defining properties of a probability are as
follows:
1. The probability of any event E is a number between 0 and 1: 0 ≤ P(E) ≤ 1.
2. The sum of the probabilities of any set of mutually exclusive and exhaustive events equals 1.
P followed by parentheses stands for “the probability of (the event in parentheses),” as in P(E) for “the probability of event E.” We can also think of P as a rule or function that assigns numerical values to events consistent with Properties 1 and 2.
In the above definition, the term mutually exclusive means that only one event can occur at a time; exhaustive means that the events cover all possible outcomes. The events A = the portfolio earns a return of 10 percent and B = the portfolio earns a return below 10 percent are mutually exclusive because A and B cannot both occur at the same time. For example, a return of 8.1 percent means that B has occurred and A has not occurred. Although events A and B are mutually exclusive, they are not exhaustive because they do not cover outcomes such as a return of 11 percent. Suppose we define a third event: C = the portfolio earns a return above 10 percent. Clearly, A, B, and C are mutually exclusive and exhaustive events. Each of P(A), P(B), and P(C) is a number between 0 and 1, and P(A) + P(B) + P(C) = 1.
The most basic kind of mutually exclusive and exhaustive events is the set of all the distinct possible outcomes of the random variable. If we know both that set and the assignment of probabilities to those outcomes—the probability distribution of the random variable—we have a complete description of the random variable, and we can assign a probability to any event that we might describe.1 The probability of any event is the sum of the probabilities of the distinct outcomes included in the definition of the event. Suppose the event of interest is D = the portfolio earns a return above the risk-free rate, and we know the probability distribution of portfolio returns. Assume the risk-free rate is 4 percent. To calculate P(D), the probability of D, we would sum the probabilities of the outcomes that satisfy the definition of the event; that is, we would sum the probabilities of portfolio returns greater than 4 percent.
Earlier, to illustrate a concept, we assumed a probability of 0.40 for a portfolio earning less than 10 percent, without justifying the particular assumption. We also talked about using a probability distribution of outcomes to calculate the probability of events, without explaining how a probability distribution might be estimated. Making actual financial decisions using inaccurate probabilities might have grave consequences. How, in practice, do we estimate probabilities? This topic is a field of study in itself, but there are three broad approaches to estimating probabilities. In investments, we often estimate the probability of an event as a relative frequency of
occurrence based on historical data. This method produces an empirical probability. For example, Thanatawee (2013) reports that of his sample of 1,927 yearly observations for nonfinancial SET (Stock Exchange of Thailand) firms during the years 2002 to 2010, 1,382 were dividend paying firms and 545 were non- dividend-paying firms. The empirical probability of a Thai firm paying a dividend is thus 1,382/1,927 = 0.72, approximately. We will point out empirical probabilities in several places as they appear in this reading.
Relationships must be stable through time for empirical probabilities to be accurate. We cannot calculate an empirical probability of an event not in the historical record or a reliable empirical probability for a very rare event. There are cases, then, in which we may adjust an empirical probability to account for perceptions of changing relationships. In other cases, we have no empirical probability to use at all. We may also make a personal assessment of probability without reference to any particular data. Each of these three types of probability is a subjective probability, one drawing on personal or subjective judgment. Subjective probabilities are of great importance in investments. Investors, in making buy and sell decisions that determine asset prices, often draw on subjective probabilities. Subjective probabilities appear in various places in this reading, notably in our discussion of Bayes’ formula.
In a more narrow range of well-defined problems, we can sometimes deduce probabilities by reasoning about the problem. The resulting probability is an a priori probability, one based on logical analysis rather than on observation or personal judgment. We will use this type of probability in Example 6. The counting methods we discuss later are particularly important in calculating an a priori probability. Because a priori and empirical probabilities generally do not vary from person to person, they are often grouped as objective probabilities.
In business and elsewhere, we often encounter probabilities stated in terms of odds —for instance, “the odds for E” or the “odds against E.” For example, as of November 2013, analysts’ fiscal year 2014 EPS forecasts for JetBlue Airways (NASDAQ: JBLU) ranged from $0.55 to $0.69. Suppose one analyst asserts that the odds for the company beating the highest estimate, $0.69, are 1 to 7. Suppose a second analyst argues that the odds against that happening are 15 to 1. What do those statements imply about the probability of the company’s EPS beating the highest estimate? We interpret probabilities stated in terms of odds as follows:
Probability Stated as Odds. Given a probability P(E),
1. Odds for E = P(E)/[1 − P(E)]. The odds for E are the probability of E divided by 1 minus the probability of E. Given odds for E of “a to b,” the
implied probability of E is a/(a + b).
In the example, the statement that the odds for the company’s EPS for FY2014 beating $0.69 are 1 to 7 means that the speaker believes the probability of the event is 1/(1 + 7) = 1/8 = 0.125.
2. Odds against E = [1 − P(E)]/P(E), the reciprocal of odds for E. Given odds
against E of “a to b,” the implied probability of E is b/(a + b).
The statement that the odds against the company’s EPS for FY2014 beating $0.69 are 15 to 1 is consistent with a belief that the probability of the event is 1/(1 + 15) = 1/16 = 0.0625.
To further explain odds for an event, if P(E) = 1/8, the odds for E are (1/8)/(7/8) = (1/8)(8/7) = 1/7, or “1 to 7.” For each occurrence of E, we expect seven cases of non-occurrence; out of eight cases in total, therefore, we expect E to happen once, and the probability of E is 1/8. In wagering, it is common to speak in terms of the odds against something, as in Statement 2. For odds of “15 to 1” against E (an implied probability of E of 1/16), a $1 wager on E, if successful, returns $15 in profits plus the $1 staked in the wager. We can calculate the bet’s anticipated profit as follows:
Win: Probability = 1/16; Profit =$15 Loss: Probability = 15/16; Profit =–$1 Anticipated profit = (1/16)($15) + (15/16)(–$1) = $0
Weighting each of the wager ’s two outcomes by the respective probability of the outcome, if the odds (probabilities) are accurate, the anticipated profit of the bet is $0.
EXAMPLE 1 Profiting from Inconsistent Probabilities
You are examining the common stock of two companies in the same industry in which an important antitrust decision will be announced next week. The first company, SmithCo Corporation, will benefit from a governmental decision that there is no antitrust obstacle related to a merger in which it is involved. You believe that SmithCo’s share price reflects a 0.85 probability of such a decision. A second company, Selbert Corporation, will equally benefit from a “go ahead” ruling. Surprisingly, you believe Selbert stock reflects only a 0.50
probability of a favorable decision. Assuming your analysis is correct, what investment strategy would profit from this pricing discrepancy?
Consider the logical possibilities. One is that the probability of 0.50 reflected in Selbert’s share price is accurate. In that case, Selbert is fairly valued but SmithCo is overvalued, as its current share price overestimates the probability of a “go ahead” decision. The second possibility is that the probability of 0.85 is accurate. In that case, SmithCo shares are fairly valued, but Selbert shares, which build in a lower probability of a favorable decision, are undervalued. You diagram the situation as shown in Table 1.
The 0.50 probability column shows that Selbert shares are a better value than SmithCo shares. Selbert shares are also a better value if a 0.85 probability is accurate. Thus SmithCo shares are overvalued relative to Selbert shares.
Your investment actions depend on your confidence in your analysis and on any investment constraints you face (such as constraints on selling stock short).2 A conservative strategy would be to buy Selbert shares and reduce or eliminate any current position in SmithCo. The most aggressive strategy is to short SmithCo stock (relatively overvalued) and simultaneously buy the stock of Selbert (relatively undervalued). This strategy is known as pairs arbitrage trade: a trade in two closely related stocks involving the short sale of one and the purchase of the other.
TABLE 1 Worksheet for Investment Problem
True Probability of a “Go Ahead” Decision 0.50 0.85
SmithCo Shares Overvalued Shares Fairly Valued Selbert Shares Fairly Valued Shares Undervalued
The prices of SmithCo and Selbert shares reflect probabilities that are not consistent. According to one of the most important probability results for investments, the Dutch Book Theorem,3 inconsistent probabilities create profit opportunities. In our example, investors, by their buy and sell decisions to exploit the inconsistent probabilities, should eliminate the profit opportunity and inconsistency.
To understand the meaning of a probability in investment contexts, we need to distinguish between two types of probability: unconditional and conditional. Both unconditional and conditional probabilities satisfy the definition of probability stated earlier, but they are calculated or estimated differently and have different
interpretations. They provide answers to different questions.
The probability in answer to the straightforward question “What is the probability of this event A?” is an unconditional probability, denoted P(A). Unconditional probability is also frequently referred to as marginal probability.4
Suppose the question is “What is the probability that the stock earns a return above the risk-free rate (event A)?” The answer is an unconditional probability that can be viewed as the ratio of two quantities. The numerator is the sum of the probabilities of stock returns above the risk-free rate. Suppose that sum is 0.70. The denominator is 1, the sum of the probabilities of all possible returns. The answer to the question is P(A) = 0.70.
Contrast the question “What is the probability of A?” with the question “What is the probability of A, given that B has occurred?” The probability in answer to this last question is a conditional probability, denoted P(A | B) (read: “the probability of A given B”).
Suppose we want to know the probability that the stock earns a return above the risk-free rate (event A), given that the stock earns a positive return (event B). With the words “given that,” we are restricting returns to those larger than 0 percent—a new element in contrast to the question that brought forth an unconditional probability. The conditional probability is calculated as the ratio of two quantities. The numerator is the sum of the probabilities of stock returns above the risk-free rate; in this particular case, the numerator is the same as it was in the unconditional case, which we gave as 0.70. The denominator, however, changes from 1 to the sum of the probabilities for all outcomes (returns) above 0 percent. Suppose that number is 0.80, a larger number than 0.70 because returns between 0 and the risk-free rate have some positive probability of occurring. Then P(A | B) = 0.70/0.80 = 0.875. If we observe that the stock earns a positive return, the probability of a return above the risk-free rate is greater than the unconditional probability, which is the probability of the event given no other information. The result is intuitive.5 To review, an unconditional probability is the probability of an event without any restriction; it might even be thought of as a stand-alone probability. A conditional probability, in contrast, is a probability of an event given that another event has occurred.
In discussing approaches to calculating probability, we gave one empirical estimate of the probability that a change in dividends is a dividend decrease. That probability was an unconditional probability. Given additional information on company characteristics, could an investor refine that estimate? Investors continually seek an information edge that will help improve their forecasts. In mathematical terms, they are attempting to frame their view of the future using probabilities conditioned on
relevant information or events. Investors do not ignore useful information; they adjust their probabilities to reflect it. Thus, the concepts of conditional probability (which we analyze in more detail below), as well as related concepts discussed further on, are extremely important in investment analysis and financial markets.
To state an exact definition of conditional probability, we first need to introduce the concept of joint probability. Suppose we ask the question “What is the probability of both A and B happening?” The answer to this question is a joint probability, denoted P(AB) (read: “the probability of A and B”). If we think of the probability of A and the probability of B as sets built of the outcomes of one or more random variables, the joint probability of A and B is the sum of the probabilities of the outcomes they have in common. For example, consider two events: the stock earns a return above the risk-free rate (A) and the stock earns a positive return (B). The outcomes of A are contained within (a subset of) the outcomes of B, so P(AB) equals P(A). We can now state a formal definition of conditional probability that provides a formula for calculating it.
Definition of Conditional Probability. The conditional probability of A given that B has occurred is equal to the joint probability of A and B divided by the probability of B (assumed not to equal 0).
(1)
Sometimes we know the conditional probability P(A | B) and we want to know the joint probability P(AB). We can obtain the joint probability from the following multiplication rule for probabilities, which is Equation 1 rearranged.
Multiplication Rule for Probability. The joint probability of A and B can be expressed as
(2)
EXAMPLE 2 Conditional Probabilities and Predictability of Mutual Fund Performance (1)
Vidal-Garcia (2013) examined whether historical performance predicts future performance for a sample of mutual funds that included 1,050 actively managed equity funds in six European countries during the period of 1988 through 2010. Funds were classified into nine investment styles based on combinations of investment focus (growth, blend, and value) and fund’s market
capitalization (small, mid, and large cap). One approach Vidal-Garcia used involved calculating each fund’s annual benchmark-adjusted return by subtracting a benchmark return from the annual return of the fund. MSCI (Morgan Stanley Capital International) style indices were used as benchmarks. For each style of fund in each country, funds were classified as winners or losers for each of two consecutive years. The top 50 percent of funds by benchmark-adjusted return for a given year were labeled winners; the bottom 50 percent were labeled losers. An excerpt from the results of the study for 135 French funds classified as large value funds is given in Table 2. It shows the percentage of those funds that were winners in two consecutive years, winner in one year and then loser in the next year, losers then winners, and losers in both years. The winner–winner entry, for example, shows that 65.5% of the first-year winner funds were also winners in the second year. Note that the four entries in the table can be viewed as conditional probabilities.
TABLE 2 Persistence of Returns for Large Value Funds in France: 1988 through 2010
Year 2 Winner Year 2 Loser Year 1 winner 65.5% 34.5% Year 1 loser 15.5% 84.5%
Source: Vidal-Garcia (2013), Table 4.
Based on the data in Table 2, answer the following questions: 1. State the four events needed to define the four conditional probabilities.
2. State the four entries of the table as conditional probabilities using the form P(this event | that event) = number.
3. Are the conditional probabilities in Part 2 empirical, a priori, or subjective probabilities?
4. Using information in the table, calculate the probability of the event a fund is a loser in both Year 1 and Year 2. (Note that because 50 percent of funds are categorized as losers in each year, the unconditional probability that a fund is labeled a loser in either year is 0.5.)
Solution to 1: The four events needed to define the conditional probabilities are as follows:
Fund is a Year 1 winner
Fund is a Year 1 loser
Fund is a Year 2 loser
Fund is a Year 2 winner
Solution to 2: From Row 1:
P(fund is a Year 2 winner | fund is a Year 1 winner) = 0.655
P(fund is a Year 2 loser | fund is a Year 1 winner) = 0.345
From Row 2:
P(fund is a Year 2 winner | fund is a Year 1 loser) = 0.155
P(fund is a Year 2 loser | fund is a Year 1 loser) = 0.845
Solution to 3: These probabilities are calculated from data, so they are empirical probabilities.
Solution to 4: The estimated probability is 0.423. With A the event that a fund is a Year 2 loser and B the event that a fund is a Year 1 loser, AB is the event that a fund is a loser in both Year 1 and Year 2. From Table 2, P(A | B) = 0.845 and P(B) = 0.50. Thus, using Equation 2, we find that
or a probability of approximately 0.423.
Equation 2 states that the joint probability of A and B equals the probability of A given B times the probability of B. Because P(AB) = P(BA), the expression P(AB) = P(BA) = P(B | A)P(A) is equivalent to Equation 2.
When we have two events, A and B, that we are interested in, we often want to know the probability that either A or B occurs. Here the word “or” is inclusive, meaning that either A or B occurs or that both A and B occur. Put another way, the probability of A or B is the probability that at least one of the two events occurs. Such probabilities are calculated using the addition rule for probabilities.
Addition Rule for Probabilities. Given events A and B, the probability that A or B occurs, or both occur, is equal to the probability that A occurs, plus the probability that B occurs, minus the probability that both A and B occur.
(3)
If we think of the individual probabilities of A and B as sets built of outcomes of one or more random variables, the first step in calculating the probability of A or B is to sum the probabilities of the outcomes in A to obtain P(A). If A and B share any outcomes, then if we now added P(B) to P(A), we would count twice the probabilities of those shared outcomes. So we add to P(A) the quantity [P(B) − P(AB)], which is the probability of outcomes in B net of the probability of any outcomes already counted when we computed P(A). Figure 1 illustrates this process; we avoid double-counting the outcomes in the intersection of A and B by subtracting P(AB). As an example of the calculation, if P(A) = 0.50, P(B) = 0.40, and P(AB) = 0.20, then P(A or B) = 0.50 + 0.40 − 0.20 = 0.70. Only if the two events A and B were mutually exclusive, so that P(AB) = 0, would it be correct to state that P(A or B) = P(A) + P(B).
FIGURE 1 Addition Rule for Probabilities
The next example shows how much useful information can be obtained using the few probability rules presented to this point.
EXAMPLE 3 Probability of a Limit Order Executing
You have two buy limit orders outstanding on the same stock. A limit order to buy stock at a stated price is an order to buy at that price or lower. A number of vendors, including an Internet service that you use, supply the estimated probability that a limit order will be filled within a stated time horizon, given the current stock price and the price limit. One buy order (Order 1) was placed at a price limit of $10. The probability that it will execute within one hour is 0.35. The second buy order (Order 2) was placed at a price limit of $9.75; it has
a 0.25 probability of executing within the same one-hour time frame. 1. What is the probability that either Order 1 or Order 2 will execute?
2. What is the probability that Order 2 executes, given that Order 1 executes?
Solution to 1: The probability is 0.35. The two probabilities that are given are P(Order 1 executes) = 0.35 and P(Order 2 executes) = 0.25. Note that if Order 2 executes, it is certain that Order 1 also executes because the price must pass through $10 to reach $9.75. Thus,
and
To answer the question, we use the addition rule for probabilities:
Note that the outcomes for which Order 2 executes are a subset of the outcomes for which Order 1 executes. After you count the probability that Order 1 executes, you have counted the probability of the outcomes for which Order 2 also executes. Therefore, the answer to the question is the probability that Order 1 executes, 0.35.
Solution to 2: If the first order executes, the probability that the second order executes is 0.714. In the solution to Part 1, you found that P(Order 1 executes and Order 2 executes) = P(Order 1 executes | Order 2 executes)P(Order 2 executes) = 1(0.25) = 0.25. An equivalent way to state this joint probability is useful here:
Because P(Order 1 executes) = 0.35 was a given, you have one equation with one unknown:
You conclude that P(Order 2 executes | Order 1 executes) = 0.25/0.35 = 5/7, or about 0.714. You can also use Equation 1 to obtain this answer.
Of great interest to investment analysts are the concepts of independence and dependence. These concepts bear on such basic investment questions as which financial variables are useful for investment analysis, whether asset returns can be predicted, and whether superior investment managers can be selected based on their past records.
Two events are independent if the occurrence of one event does not affect the probability of occurrence of the other event.
Definition of Independent Events. Two events A and B are independent if and only if P(A | B) = P(A) or, equivalently, P(B | A) = P(B).
When two events are not independent, they are dependent: The probability of occurrence of one is related to the occurrence of the other. If we are trying to forecast one event, information about a dependent event may be useful, but information about an independent event will not be useful.
When two events are independent, the multiplication rule for probabilities, Equation 2, simplifies because P(A | B) in that equation then equals P(A).
Multiplication Rule for Independent Events. When two events are independent, the joint probability of A and B equals the product of the individual probabilities of A and B.
(4)
Therefore, if we are interested in two independent events with probabilities of 0.75 and 0.50, respectively, the probability that both will occur is 0.375 = 0.75(0.50). The multiplication rule for independent events generalizes to more than two events; for example, if A, B, and C are independent events, then P(ABC) = P(A)P(B)P(C).
EXAMPLE 4 BankCorp’s Earnings per Share (1)
As part of your work as a banking industry analyst, you build models for forecasting earnings per share of the banks you cover. Today you are studying BankCorp. The historical record shows that in 55 percent of recent quarters BankCorp’s EPS has increased sequentially, and in 45 percent of quarters EPS has decreased or remained unchanged sequentially.6 At this point in your analysis, you are assuming that changes in sequential EPS are independent.
Earnings per share for 2Q:2014 (that is, EPS for the second quarter of 2014) were larger than EPS for 1Q:2014.
1. What is the probability that 3Q:2014 EPS will be larger than 2Q:2014 EPS (a
positive change in sequential EPS)?
2. What is the probability that EPS decreases or remains unchanged in the next two quarters?
Solution to 1: Under the assumption of independence, the probability that 3Q:2014 EPS will be larger than 2Q:2014 EPS is the unconditional probability of positive change, 0.55. The fact that 2Q:2014 EPS was larger than 1Q:2014 EPS is not useful information, as the next change in EPS is independent of the prior change.
Solution to 2: The probability is 0.2025 = 0.45(0.45).
The following example illustrates how difficult it is to satisfy a set of independent criteria even when each criterion by itself is not necessarily stringent.
EXAMPLE 5 Screening Stocks for Investment
You have developed a stock screen—a set of criteria for selecting stocks. Your investment universe (the set of securities from which you make your choices) is the Russell 1000 Index, an index of 1,000 large-capitalization US equities. Your criteria capture different aspects of the selection problem; you believe that the criteria are independent of each other, to a close approximation.
Criterion Fraction of Russell 1000 Stocks MeetingCriterion First valuation criterion 0.50 Second valuation criterion 0.50 Analyst coverage criterion 0.25
Profitability criterion for company 0.55 Financial strength criterion for
company 0.67
How many stocks do you expect to pass your screen?
Only 23 stocks out of 1,000 pass through your screen. If you define five events
—the stock passes the first valuation criterion, the stock passes the second valuation criterion, the stock passes the analyst coverage criterion, the company passes the profitability criterion, the company passes the financial strength criterion (say events A, B, C, D, and E, respectively)—then the probability that a stock will pass all five criteria, under independence, is
Although only one of the five criteria is even moderately strict (the strictest lets 25 percent of stocks through), the probability that a stock can pass all five is only 0.023031, or about 2 percent. The size of the list of candidate investments is 0.023031(1,000) = 23.031, or 23 stocks.
An area of intense interest to investment managers and their clients is whether records of past performance are useful in identifying repeat winners and losers. The following example shows how this issue relates to the concept of independence.
EXAMPLE 6 Conditional Probabilities and Predictability of Mutual Fund Performance (2)
The purpose of the Vidal-Garcia (2013) study, introduced in Example 2, was to address the question of repeat European mutual fund winners and losers. If the status of a fund as a winner or a loser in one year is independent of whether it is a winner in the next year, the practical value of performance ranking is questionable. Using the four events defined in Example 2 as building blocks, we can define the following events to address the issue of predictability of mutual fund performance:
Fund is a Year 1 winner and fund is a Year 2 winner
Fund is a Year 1 winner and fund is a Year 2 loser
Fund is a Year 1 loser and fund is a Year 2 winner
Fund is a Year 1 loser and fund is a Year 2 loser
In Part 4 of Example 2, you calculated that
If the ranking in one year is independent of the ranking in the next year, what will you expect P(fund is a Year 2 loser and fund is a Year 1 loser) to be?
Interpret the empirical probability 0.423.
By the multiplication rule for independent events, P(fund is a Year 2 loser and fund is a Year 1 loser) = P(fund is a Year 2 loser)P(fund is a Year 1 loser). Because 50 percent of funds are categorized as losers in each year, the unconditional probability that a fund is labeled a loser in either year is 0.50. Thus P(fund is a Year 2 loser)P(fund is a Year 1 loser) = 0.50(0.50) = 0.25. If the status of a fund as a loser in one year is independent of whether it is a loser in the prior year, we conclude that P(fund is a Year 2 loser and fund is a Year 1 loser) = 0.25. This probability is a priori because it is obtained from reasoning about the problem. You could also reason that the four events described above define categories and that if funds are randomly assigned to the four categories, there is a 1/4 probability of fund is a Year 1 loser and fund is a Year 2 loser. If the classifications in Year 1 and Year 2 were dependent, then the assignment of funds to categories would not be random. The empirical probability of 0.423 is above 0.25. Is this apparent predictability the result of chance? A test conducted by Vidal-Garcia indicated a less than1 percent chance of observing the tabled data if the Year 1 and Year 2 rankings were independent.
In investments, the question of whether one event (or characteristic) provides information about another event (or characteristic) arises in both time-series settings (through time) and cross-sectional settings (among units at a given point in time). Examples 4 and 6 examined independence in a time-series setting. Example 5 illustrated independence in a cross-sectional setting. Independence/dependence relationships are often also explored in both settings using regression analysis, a technique we discuss in a later reading.
In many practical problems, we logically analyze a problem as follows: We formulate scenarios that we think affect the likelihood of an event that interests us. We then estimate the probability of the event, given the scenario. When the scenarios (conditioning events) are mutually exclusive and exhaustive, no possible outcomes are left out. We can then analyze the event using the total probability rule. This rule explains the unconditional probability of the event in terms of probabilities conditional on the scenarios.
The total probability rule is stated below for two cases. Equation 5 gives the simplest case, in which we have two scenarios. One new notation is introduced: If we have an event or scenario S, the event not-S, called the complement of S, is written SC.7 Note that P(S) + P(SC) = 1, as either S or not-S must occur. Equation 6 states the rule for the general case of n mutually exclusive and exhaustive events or scenarios.
The Total Probability Rule.
(5)
(6)
where S1, S2, …, Sn are mutually exclusive and exhaustive scenarios or events.
Equation 6 states the following: The probability of any event [P(A)] can be expressed as a weighted average of the probabilities of the event, given scenarios [terms such P(A | S1)]; the weights applied to these conditional probabilities are the respective probabilities of the scenarios [terms such as P(S1) multiplying P(A | S1)], and the scenarios must be mutually exclusive and exhaustive. Among other applications, this rule is needed to understand Bayes’ formula, which we discuss later in the reading.
In the next example, we use the total probability rule to develop a consistent set of views about BankCorp’s earnings per share.
EXAMPLE 7 BankCorp’s Earnings per Share (2)
You are continuing your investigation into whether you can predict the direction of changes in BankCorp’s quarterly EPS. You define four events:
Event Probability A = Change in sequential EPS is positive next quarter 0.55
AC = Change in sequential EPS is 0 or negative next quarter 0.45 S = Change in sequential EPS is positive in the prior quarter 0.55
SC = Change in sequential EPS is 0 or negative in the prior quarter 0.45
On inspecting the data, you observe some persistence in EPS changes: Increases tend to be followed by increases, and decreases by decreases. The first probability estimate you develop is P(change in sequential EPS is positive next quarter | change in sequential EPS is 0 or negative in the prior quarter) = P(A | SC) = 0.40. The most recent quarter ’s EPS (2Q:2014) is announced, and the change is a positive sequential change (the event S). You are interested in
forecasting EPS for 3Q:2014. 1. Write this statement in probability notation: “the probability that the change in
sequential EPS is positive next quarter, given that the change in sequential EPS is positive the prior quarter.”
2. Calculate the probability in Part 1. (Calculate the probability that is consistent with your other probabilities or beliefs.)
Solution to 1: In probability notation, this statement is written P(A | S).
Solution to 2: The probability is 0.673 that the change in sequential EPS is positive for 3Q:2014, given the positive change in sequential EPS for 2Q:2014, as shown on the following page.
According to Equation 5, P(A) = P(A | S)P(S) + P(A | SC)P(SC). The values of the probabilities needed to calculate P(A | S) are already known: P(A) = 0.55, P(S) = 0.55, P(SC) = 0.45, and P(A | SC) = 0.40. Substituting into Equation 5,
Solving for the unknown, P(A | S) = [0.55 − 0.40(0.45)]/0.55 = 0.672727, or 0.673.
You conclude that P(change in sequential EPS is positive next quarter | change in sequential EPS is positive the prior quarter) = 0.673. Any other probability is not consistent with your other estimated probabilities. Reflecting the persistence in EPS changes, this conditional probability of a positive EPS change, 0.673, is greater than the unconditional probability of an EPS increase, 0.55.
In the reading on statistical concepts and market returns, we discussed the concept of a weighted average or weighted mean. The example highlighted in that reading was that portfolio return is a weighted average of the returns on the individual assets in the portfolio, where the weight applied to each asset’s return is the fraction of the portfolio invested in that asset. The total probability rule, which is a rule for stating an unconditional probability in terms of conditional probabilities, is also a weighted average. In that formula, probabilities of scenarios are used as weights. Part of the definition of weighted average is that the weights sum to 1. The probabilities of mutually exclusive and exhaustive events do sum to 1 (this is part of the definition of probability). The next weighted average we discuss, the expected value of a random variable, also uses probabilities as weights.
The expected value of a random variable is an essential quantitative concept in
investments. Investors continually make use of expected values—in estimating the rewards of alternative investments, in forecasting EPS and other corporate financial variables and ratios, and in assessing any other factor that may affect their financial position. The expected value of a random variable is defined as follows:
Definition of Expected Value. The expected value of a random variable is the probability-weighted average of the possible outcomes of the random variable. For a random variable X, the expected value of X is denoted E(X).
Expected value (for example, expected stock return) looks either to the future, as a forecast, or to the “true” value of the mean (the population mean, discussed in the reading on statistical concepts and market returns). We should distinguish expected value from the concepts of historical or sample mean. The sample mean also summarizes in a single number a central value. However, the sample mean presents a central value for a particular set of observations as an equally weighted average of those observations. To summarize, the contrast is forecast versus historical, or population versus sample.
EXAMPLE 8 BankCorp’s Earnings per Share (3)
You continue with your analysis of BankCorp’s EPS. In Table 3, you have recorded a probability distribution for BankCorp’s EPS for the current fiscal year.
TABLE 3 Probability Distribution for BankCorp’s EPS
Probability EPS ($) 0.15 2.60 0.45 2.45 0.24 2.20 0.16 2.00 1.00
What is the expected value of BankCorp’s EPS for the current fiscal year?
Following the definition of expected value, list each outcome, weight it by its probability, and sum the terms.
The expected value of EPS is $2.34.
An equation that summarizes your calculation in Example 8 is
(7)
where Xi is one of n possible outcomes of the random variable X.8
The expected value is our forecast. Because we are discussing random quantities, we cannot count on an individual forecast being realized (although we hope that, on average, forecasts will be accurate). It is important, as a result, to measure the risk we face. Variance and standard deviation measure the dispersion of outcomes around the expected value or forecast.
Definition of Variance. The variance of a random variable is the expected value (the probability-weighted average) of squared deviations from the random variable’s expected value:
(8)
The two notations for variance are σ2(X) and Var(X).
Variance is a number greater than or equal to 0 because it is the sum of squared terms. If variance is 0, there is no dispersion or risk. The outcome is certain, and the quantity X is not random at all. Variance greater than 0 indicates dispersion of outcomes. Increasing variance indicates increasing dispersion, all else equal. Variance of X is a quantity in the squared units of X. For example, if the random variable is return in percent, variance of return is in units of percent squared. Standard deviation is easier to interpret than variance, as it is in the same units as the random variable. If the random variable is return in percent, standard deviation of return is also in units of percent. In the following example, when the variance of returns is stated as a percent or amount of money, to conserve space the reading may suppress showing the unit squared. Note that when the variance of returns is stated as a decimal, the complication of dealing with units of “percent squared” does not arise.
Definition of Standard Deviation. Standard deviation is the positive square root of variance.
The best way to become familiar with these concepts is to work examples.
EXAMPLE 9 BankCorp’s Earnings per Share (4)
In Example 8, you calculated the expected value of BankCorp’s EPS as $2.34, which is your forecast. Now you want to measure the dispersion around your forecast. Table 4 shows your view of the probability distribution of EPS for the current fiscal year.
What are the variance and standard deviation of BankCorp’s EPS for the current fiscal year?
The order of calculation is always expected value, then variance, then standard deviation. Expected value has already been calculated. Following the definition of variance above, calculate the deviation of each outcome from the mean or expected value, square each deviation, weight (multiply) each squared deviation by its probability of occurrence, and then sum these terms.
Standard deviation is the positive square root of 0.038785:
TABLE 4 Probability Distribution for BankCorp’s EPS
Probability EPS ($) 0.15 2.60 0.45 2.45 0.24 2.20 0.16 2.00 1.00
An equation that summarizes your calculation of variance in Example 9 is
(9)
where Xi is one of n possible outcomes of the random variable X.
In investments, we make use of any relevant information available in making our forecasts. When we refine our expectations or forecasts, we are typically making adjustments based on new information or events; in these cases we are using conditional expected values. The expected value of a random variable X given an event or scenario S is denoted E(X | S). Suppose the random variable X can take on any one of n distinct outcomes X1, X2, …, Xn (these outcomes form a set of mutually exclusive and exhaustive events). The expected value of X conditional on S is the first outcome, X1, times the probability of the first outcome given S, P(X1 | S), plus the second outcome, X2, times the probability of the second outcome given S, P(X2 | S), and so forth.
(10)
We will illustrate this equation shortly.
Parallel to the total probability rule for stating unconditional probabilities in terms of conditional probabilities, there is a principle for stating (unconditional) expected values in terms of conditional expected values. This principle is the total probability rule for expected value.
The Total Probability Rule for Expected Value.
(11)
(12)
where S1, S2, …, Sn are mutually exclusive and exhaustive scenarios or events.
The general case, Equation 12, states that the expected value of X equals the expected value of X given Scenario 1, E(X | S1), times the probability of Scenario 1, P(S1), plus the expected value of X given Scenario 2, E(X | S2), times the probability of Scenario 2, P(S2), and so forth.
To use this principle, we formulate mutually exclusive and exhaustive scenarios that are useful for understanding the outcomes of the random variable. This approach was employed in developing the probability distribution of BankCorp’s EPS in
Examples 8 and 9, as we now discuss.
The earnings of BankCorp are interest rate sensitive, benefiting from a declining interest rate environment. Suppose there is a 0.60 probability that BankCorp will operate in a declining interest rate environment in the current fiscal year and a 0.40 probability that it will operate in a stable interest rate environment (assessing the chance of an increasing interest rate environment as negligible). If a declining interest rate environment occurs, the probability that EPS will be $2.60 is estimated at 0.25, and the probability that EPS will be $2.45 is estimated at 0.75. Note that 0.60, the probability of declining interest rate environment, times 0.25, the probability of $2.60 EPS given a declining interest rate environment, equals 0.15, the (unconditional) probability of $2.60 given in the table in Examples 8 and 9. The probabilities are consistent. Also, 0.60(0.75) = 0.45, the probability of $2.45 EPS given in Tables 3 and 4. The tree diagram in Figure 2 shows the rest of the analysis.
FIGURE 2 BankCorp’s Forecasted EPS
A declining interest rate environment points us to the node of the tree that branches off into outcomes of $2.60 and $2.45. We can find expected EPS given a declining interest rate environment as follows, using Equation 10:
If interest rates are stable,
Once we have the new piece of information that interest rates are stable, for example, we revise our original expectation of EPS from $2.34 downward to $2.12.
Now using the total probability rule for expected value,
So E(EPS) = $2.4875 (0.60) + $2.12 (0.40) = $2.3405 or about $2.34.
This amount is identical to the estimate of the expected value of EPS calculated directly from the probability distribution in Example 8. Just as our probabilities must be consistent, so must our expected values, unconditional and conditional; otherwise our investment actions may create profit opportunities for other investors at our expense.
To review, we first developed the factors or scenarios that influence the outcome of the event of interest. After assigning probabilities to these scenarios, we formed expectations conditioned on the different scenarios. Then we worked backward to formulate an expected value as of today. In the problem just worked, EPS was the event of interest, and the interest rate environment was the factor influencing EPS.
We can also calculate the variance of EPS given each scenario:
These are conditional variances, the variance of EPS given a declining interest rate environment and the variance of EPS given a stable interest rate environment. The relationship between unconditional variance and conditional variance is a relatively
advanced topic.9 The main points are 1) that variance, like expected value, has a conditional counterpart to the unconditional concept and 2) that we can use conditional variance to assess risk given a particular scenario.
EXAMPLE 10 BankCorp’s Earnings per Share (5)
Continuing with BankCorp, you focus now on BankCorp’s cost structure. One model you are researching for BankCorp’s operating costs is
FIGURE 3 BankCorp’s Forecasted Operating Costs
where is a forecast of operating costs in millions of dollars and X is the number of branch offices. represents the expected value of Y given X, or E(Y | X). ( is a notation used in regression analysis, which we discuss in a later reading.) You interpret the intercept a as fixed costs and b as variable costs. You estimate the equation as
BankCorp currently has 66 branch offices, and the equation estimates that 12.5 + 0.65(66) = $55.4 million. You have two scenarios for growth, pictured in the tree diagram in Figure 3.
1. Compute the forecasted operating costs given the different levels of operating
costs, using . State the probability of each level of the number of branch offices. These are the answers to the questions in the terminal boxes of
the tree diagram.
2. Compute the expected value of operating costs under the high-growth scenario. Also calculate the expected value of operating costs under the low-growth scenario.
3. Answer the question in the initial box of the tree: What are BankCorp’s expected operating costs?
Solution to 1: Using , from top to bottom, we have
Operating Costs Probability 0.80(0.50) = 0.40 0.80(0.50) = 0.40 0.20(0.85) = 0.17 0.20(0.15) = 0.03 Sum = 1.00
Solution to 2: Dollar amounts are in millions.
Solution to 3: Dollar amounts are in millions.
BankCorp’s expected operating costs are $81.205 million.
We will see conditional probabilities again when we discuss Bayes’ formula. This section has introduced a few problems that can be addressed using probability concepts. The following problem draws on these concepts, as well as on analytical skills.
EXAMPLE 11 The Default Risk Premium for a One-Period Debt Instrument
As the co-manager of a short-term bond portfolio, you are reviewing the pricing of a speculative-grade, one-year-maturity, zero-coupon bond. For this type of bond, the return is the difference between the amount paid and the principal value received at maturity. Your goal is to estimate an appropriate default risk premium for this bond. You define the default risk premium as the extra return above the risk-free return that will compensate investors for default risk. If R is the promised return (yield-to-maturity) on the debt instrument and RF is the risk-free rate, the default risk premium is R − RF. You assess the probability that the bond defaults as P(the bond defaults) = 0.06. Looking at current money market yields, you find that one-year US Treasury bills (T-bills) are offering a return of 2 percent, an estimate of RF. As a first step, you make the simplifying assumption that bondholders will recover nothing in the event of a default. What is the minimum default risk premium you should require for this instrument?
The challenge in this type of problem is to find a starting point. In many problems, including this one, an effective first step is to divide up the possible outcomes into mutually exclusive and exhaustive events in an economically logical way. Here, from the viewpoint of a bondholder, the two events that affect returns are the bond defaults and the bond does not default. These two events cover all outcomes. How do these events affect a bondholder ’s returns? A second step is to compute the value of the bond for the two events. We have no specifics on bond face value, but we can compute value per $1 or one unit of currency invested.
The Bond Defaults The Bond Does Not Default Bond value $0 $(1 + R)
The third step is to find the expected value of the bond (per $1 invested).
So E(bond) = $(1 + R)[1 − P(the bond defaults)]. The expected value of the T- bill per $1 invested is (1 + RF). In fact, this value is certain because the T-bill is risk free. The next step requires economic reasoning. You want the default premium to be large enough so that you expect to at least break even compared with investing in the T-bill. This outcome will occur if the expected value of the bond equals the expected value of the T-bill per $1 invested.
Solving for the promised return on the bond, you find R = {(1 + RF)/[1 − P(the bond defaults)]} − 1. Substituting in the values in the statement of the problem, R = [1.02/(1 − 0.06)] − 1 = 1.08511 − 1 = 0.08511 or about 8.51 percent, and default risk premium is R − RF = 8.51% − 2% = 6.51%.
You require a default risk premium of at least 651 basis points. You can state the matter as follows: If the bond is priced to yield 8.51 percent, you will earn a 651 basis-point spread and receive the bond principal with 94 percent probability. If the bond defaults, however, you will lose everything. With a premium of 651 basis points, you expect to just break even relative to an investment in T-bills. Because an investment in the zero-coupon bond has variability, if you are risk averse you will demand that the premium be larger than 651 basis points.
This analysis is a starting point. Bondholders usually recover part of their investment after a default. A next step would be to incorporate a recovery rate.
In this section, we have treated random variables such as EPS as stand-alone quantities. We have not explored how descriptors such as expected value and variance of EPS may be functions of other random variables. Portfolio return is one random variable that is clearly a function of other random variables, the random returns on the individual securities in the portfolio. To analyze a portfolio’s expected return and variance of return, we must understand these quantities are a function of characteristics of the individual securities’ returns. Looking at the dispersion or variance of portfolio return, we see that the way individual security returns move together or covary is important. To understand the significance of these movements, we need to explore some new concepts, covariance and correlation. The next section, which deals with portfolio expected return and variance of return, introduces these concepts.
3. Portfolio Expected Return and Variance of Return Modern portfolio theory makes frequent use of the idea that investment opportunities can be evaluated using expected return as a measure of reward and variance of return as a measure of risk. The calculation and interpretation of portfolio expected return and variance of return are fundamental skills. In this section, we will develop an understanding of portfolio expected return and variance of return.10 Portfolio return is determined by the returns on the individual holdings. As a result, the calculation of portfolio variance, as a function of the individual asset returns, is more complex than the variance calculations illustrated in the previous section.
We work with an example of a portfolio that is 50 percent invested in an S&P 500 Index fund, 25 percent invested in a US long-term corporate bond fund, and 25 percent invested in a fund indexed to the MSCI EAFE Index (representing equity markets in Europe, Australasia, and the Far East). Table 5 shows these weights.
We first address the calculation of the expected return on the portfolio. In the previous section, we defined the expected value of a random variable as the probability-weighted average of the possible outcomes. Portfolio return, we know, is a weighted average of the returns on the securities in the portfolio. Similarly, the expected return on a portfolio is a weighted average of the expected returns on the securities in the portfolio, using exactly the same weights. When we have estimated the expected returns on the individual securities, we immediately have portfolio expected return. This convenient fact follows from the properties of expected value.
Properties of Expected Value. Let wi be any constant and Ri be a random variable.
1. The expected value of a constant times a random variable equals the constant times the expected value of the random variable.
2. The expected value of a weighted sum of random variables equals the weighted sum of the expected values, using the same weights.
(13)
TABLE 5 Portfolio Weights
Asset Class Weights S&P 500 0.50
US long-term corporate bonds 0.25 MSCI EAFE 0.25
Suppose we have a random variable with a given expected value. If we multiply each outcome by 2, for example, the random variable’s expected value is multiplied by 2 as well. That is the meaning of Part 1. The second statement is the rule that directly leads to the expression for portfolio expected return. A portfolio with n securities is defined by its portfolio weights, w1, w2, …, wn, which sum to 1. So portfolio return, Rp, is Rp = w1R1 + w2R2 + … + wnRn. We can state the following principle:
Calculation of Portfolio Expected Return. Given a portfolio with n securities, the expected return on the portfolio is a weighted average of the expected returns on the component securities:
Suppose we have estimated expected returns on the assets in the portfolio, as given in Table 6.
We calculate the expected return on the portfolio as 11.75 percent:
In the previous section, we studied variance as a measure of dispersion of outcomes around the expected value. Here we are interested in portfolio variance of return as a measure of investment risk. Letting Rp stand for the return on the portfolio, portfolio variance is σ2(Rp) = E{[Rp − E(Rp)]2} according to Equation 8. How do we implement this definition? In the reading on statistical concepts and market returns, we learned how to calculate a historical or sample variance based on a sample of returns. Now we are considering variance in a forward-looking sense. We will use information about the individual assets in the portfolio to obtain portfolio variance of return. To avoid clutter in notation, we write ERp for E(Rp). We need the concept of covariance.
Definition of Covariance. Given two random variables Ri and Rj, the
covariance between Ri and Rj is
(14)
Alternative notations are σ(Ri,Rj) and σij.
TABLE 6 Weights and Expected Returns
Asset Class Weight Expected Return (%) S&P 500 0.50 13
US long-term corporate bonds 0.25 6 MSCI EAFE 0.25 15
Equation 14 states that the covariance between two random variables is the probability-weighted average of the cross-products of each random variable’s deviation from its own expected value. We will return to discuss covariance after we establish the need for the concept. Working from the definition of variance, we find
(15)
The last step follows from the definitions of variance and covariance.11 For the italicized covariance terms in Equation 15, we used the fact that the order of variables in covariance does not matter: Cov(R2,R1) = Cov(R1,R2), for example. As we will show, the diagonal variance terms σ2(R1), σ2(R2), and σ2(R3) can be expressed as Cov(R1,R1), Cov(R2,R2), and Cov(R3,R3), respectively. Using this fact,
the most compact way to state Equation 15 is . The double summation signs say: “Set i = 1 and let j run from 1 to 3; then set i = 2 and let j run from 1 to 3; next set i = 3 and let j run from 1 to 3; finally, add the nine terms.” This expression generalizes for a portfolio of any size n to
(16)
We see from Equation 15 that individual variances of return constitute part, but not all, of portfolio variance. The three variances are actually outnumbered by the six covariance terms off the diagonal. For three assets, the ratio is 1 to 2, or 50 percent. If there are 20 assets, there are 20 variance terms and 20(20) − 20 = 380 off- diagonal covariance terms. The ratio of variance terms to off-diagonal covariance terms is less than 6 to 100, or 6 percent. A first observation, then, is that as the number of holdings increases, covariance12 becomes increasingly important, all else equal.
What exactly is the effect of covariance on portfolio variance? The covariance terms capture how the co-movements of returns affect portfolio variance. For example, consider two stocks: One tends to have high returns (relative to its expected return) when the other has low returns (relative to its expected return). The returns on one stock tend to offset the returns on the other stock, lowering the variability or variance of returns on the portfolio. Like variance, the units of covariance are hard to interpret, and we will introduce a more intuitive concept shortly. Meanwhile, from the definition of covariance, we can establish two essential observations about covariance. 1. We can interpret the sign of covariance as follows:
Covariance of returns is negative if, when the return on one asset is above its expected value, the return on the other asset tends to be below its expected value (an average inverse relationship between returns).
Covariance of returns is 0 if returns on the assets are unrelated.
Covariance of returns is positive when the returns on both assets tend to be on the same side (above or below) their expected values at the same time (an average positive relationship between returns).
2. The covariance of a random variable with itself (own covariance) is its own variance: Cov(R,R) = E{[R − E(R)][R − E(R)]} = E{[R − E(R)]2} = σ2(R).
A complete list of the covariances constitutes all the statistical data needed to compute portfolio variance of return. Covariances are often presented in a square format called a covariance matrix. Table 7 summarizes the inputs for portfolio expected return and variance of return.
TABLE 7 Inputs to Portfolio Expected Return and Variance
A. Inputs to Portfolio Expected Return Asset A B C
E(RA) E(RB) E(RC)
B. Covariance Matrix: The Inputs to Portfolio Variance of Return Asset A B C A Cov(RA,RA) Cov(RA,RB) Cov(RA,RC)
B Cov(RB,RA) Cov(RB,RB) Cov(RB,RC)
C Cov(RC,RA) Cov(RC,RB) Cov(RC,RC)
With three assets, the covariance matrix has 32 = 3 × 3 = 9 entries, but it is customary to treat the diagonal terms, the variances, separately from the off- diagonal terms. These diagonal terms are bolded in Table 7. This distinction is natural, as security variance is a single-variable concept. So there are 9 − 3 = 6 covariances, excluding variances. But Cov(RB,RA) = Cov(RA,RB), Cov(RC,RA) = Cov(RA,RC), and Cov(RC,RB) = Cov(RB,RC). The covariance matrix below the diagonal is the mirror image of the covariance matrix above the diagonal. As a result, there are only 6/2 = 3 distinct covariance terms to estimate. In general, for n securities, there are n(n − 1)/2 distinct covariances to estimate and n variances to estimate.
Suppose we have the covariance matrix shown in Table 8. We will be working in returns stated as percents and the table entries are in units of percent squared (%2). The terms 38%2 and 400%2 are 0.0038 and 0.0400, respectively, stated as decimals; correctly working in percents and decimals leads to identical answers.
Taking Equation 15 and grouping variance terms together produces the following:
(17)
TABLE 8 Covariance Matrix
S&P 500
US Long-Term Corporate Bonds
MSCI EAFE
S&P 500 400 45 189 US long-term corporate
bonds 45 81 38
MSCI EAFE 189 38 441
The variance is 195.875. Standard deviation of return is 195.8751/2 = 14 percent. To summarize, the portfolio has an expected annual return of 11.75 percent and a standard deviation of return of 14 percent.
Let us look at the first three terms in the calculation above. Their sum, 100 + 5.0625 + 27.5625 = 132.625, is the contribution of the individual variances to portfolio variance. If the returns on the three assets were independent, covariances would be 0 and the standard deviation of portfolio return would be 132.6251/2 = 11.52 percent as compared to 14 percent before. The portfolio would have less risk. Suppose the covariance terms were negative. Then a negative number would be added to 132.625, so portfolio variance and risk would be even smaller. At the same time, we have not changed expected return. For the same expected portfolio return, the portfolio has less risk. This risk reduction is a diversification benefit, meaning a risk-reduction benefit from holding a portfolio of assets. The diversification benefit increases with decreasing covariance. This observation is a key insight of modern portfolio theory. It is even more intuitively stated when we can use the concept of correlation. Then we can say that as long as security returns are not perfectly positively correlated, diversification benefits are possible. Furthermore, the smaller the correlation between security returns, the greater the cost of not diversifying (in terms of risk-reduction benefits forgone), all else equal.
Definition of Correlation. The correlation between two random variables, Ri and Rj, is defined as ρ(Ri,Rj) = Cov(Ri,Rj)/σ(Ri)σ(Rj). Alternative notations are
Corr(Ri,Rj) and ρij.
Frequently, covariance is substituted out using the relationship Cov(Ri,Rj) = ρ(Ri,Rj)σ(Ri)σ(Rj). The division indicated in the definition makes correlation a pure number (one without a unit of measurement) and places bounds on its largest and smallest possible values. Using the above definition, we can state a correlation matrix from data in the covariance matrix alone. Table 9 shows the correlation matrix.
For example, the covariance between long-term bonds and MSCI EAFE is 38, from Table 8. The standard deviation of long-term bond returns is 811/2 = 9 percent, that of MSCI EAFE returns is 4411/2 = 21 percent, from diagonal terms in Table 8. The correlation ρ(Return on long-term bonds, Return on EAFE) is 38/(9%)(21%) = 0.201, rounded to 0.20. The correlation of the S&P 500 with itself equals 1: The calculation is its own covariance divided by its standard deviation squared.
Properties of Correlation.
1. Correlation is a number between −1 and +1 for two random variables, X and Y:
TABLE 9 Correlation Matrix of Returns
S&P 500
US Long-TermCorporate Bonds
MSCI EAFE
S&P 500 1.00 0.25 0.45 US long-term corporate
bonds 0.25 1.00 0.20
MSCI EAFE 0.45 0.20 1.00
2. A correlation of 0 (uncorrelated variables) indicates an absence of any linear (straight-line) relationship between the variables.13 Increasingly positive correlation indicates an increasingly strong positive linear relationship (up to 1, which indicates a perfect linear relationship). Increasingly negative correlation indicates an increasingly strong negative (inverse) linear relationship (down to −1, which indicates a perfect inverse linear relationship).14
EXAMPLE 12 Portfolio Expected Return and Variance of Return
You have a portfolio of two mutual funds, A and B, 75 percent invested in A, as shown in Table 10.
1. Calculate the expected return of the portfolio.
2. Calculate the correlation matrix for this problem. Carry out the answer to two decimal places.
3. Compute portfolio standard deviation of return.
Solution to 1: E(Rp) = wAE(RA) + (1 − wA)E(RB) = 0.75(20%) + 0.25(12%) = 18%. Portfolio weights must sum to 1: wB = 1 − wA.
Solution to 2: σ(RA) = 6251/2 = 25 percent σ(RB) = 1961/2 = 14 percent. There is one distinct covariance and thus one distinct correlation: ρ(RA,RB) = Cov(RA,RB)/σ(RA)σ(RB) = 120/[25(14)] = 0.342857, or 0.34. Table 11 shows the correlation matrix.
Diagonal terms are always equal to 1 in a correlation matrix.
TABLE 10 Mutual Fund Expected Returns, Return Variances, and Covariances
Fund A B E(RA) = 20% E(RB) = 12%
Covariance Matrix Fund A B A 625 120 B 120 196
TABLE 11 Correlation Matrix
A B A 1.00 0.34 B 0.34 1.00
Solution to 3:
How do we estimate return covariance and correlation? Frequently, we make forecasts on the basis of historical covariance or use other methods based on historical return data, such as a market model regression.15 We can also calculate covariance using the joint probability function of the random variables, if that can be estimated. The joint probability function of two random variables X and Y, denoted P(X, Y), gives the probability of joint occurrences of values of X and Y. For example, P(3, 2), is the probability that X equals 3 and Y equals 2.
Suppose that the joint probability function of the returns on BankCorp stock (RA) and the returns on NewBank stock (RB) has the simple structure given in Table 12.
The expected return on BankCorp stock is 0.20(25%) + 0.50(12%) + 0.30(10%) = 14%. The expected return on NewBank stock is 0.20(20%) + 0.50(16%) + 0.30(10%) = 15%. The joint probability function above might reflect an analysis based on whether banking industry conditions are good, average, or poor. Table 13 presents the calculation of covariance.
TABLE 12 Joint Probability Function of BankCorp and NewBank Returns (Entries Are Joint Probabilities)
RB = 20% RB = 16% RB = 10%
RA = 25% 0.20 0 0 RA = 12% 0 0.50 0 RA = 10% 0 0 0.30
TABLE 13 Covariance Calculations
Banking Industry Condition
Deviations BankCorp
Deviations NewBank
Product of Deviations
Probability of
Condition
Probability- Weighted Product
Good 25−14 20−15 55 0.20 11 Average 12−14 16−15 −2 0.50 −1 Poor 10−14 10−15 20 0.30 6
Cov(RA,RB) = 16
Note: Expected return for BankCorp is 14 percent and for NewBank, 15 percent.
The first and second columns of numbers show, respectively, the deviations of BankCorp and NewBank returns from their mean or expected value. The next column shows the product of the deviations. For example, for good industry conditions, (25 − 14)(20 − 15) = 11(5) = 55. Then 55 is multiplied or weighted by 0.20, the probability that banking industry conditions are good: 55(0.20) = 11. The calculations for average and poor banking conditions follow the same pattern. Summing up these probability-weighted products, we find that Cov(RA,RB) = 16.
A formula for computing the covariance between random variables RA and RB is
(18)
The formula tells us to sum all possible deviation cross-products weighted by the appropriate joint probability. In the example we just worked, as Table 12 shows, only three joint probabilities are nonzero. Therefore, in computing the covariance of returns in this case, we need to consider only three cross-products:
One theme of this reading has been independence. Two random variables are independent when every possible pair of events—one event corresponding to a value of X and another event corresponding to a value of Y—are independent events. When two random variables are independent, their joint probability function simplifies.
Definition of Independence for Random Variables. Two random variables X and Y are independent if and only if P(X, Y) = P(X)P(Y).
For example, given independence, P(3, 2) = P(3)P(2). We multiply the individual probabilities to get the joint probabilities. Independence is a stronger property than uncorrelatedness because correlation addresses only linear relationships. The following condition holds for independent random variables and, therefore, also holds for uncorrelated random variables.
Multiplication Rule for Expected Value of the Product of Uncorrelated Random Variables. The expected value of the product of uncorrelated random variables is the product of their expected values.
E(XY) = E(X)E(Y) if X and Y are uncorrelated.
Many financial variables, such as revenue (price times quantity), are the product of random quantities. When applicable, the above rule simplifies calculating expected value of a product of random variables.16
4. Topics in Probability In the remainder of the reading we discuss two topics that can be important in solving investment problems. We start with Bayes’ formula: what probability theory has to say about learning from experience. Then we move to a discussion of shortcuts and principles for counting.
4.1. Bayes’ Formula When we make decisions involving investments, we often start with viewpoints based on our experience and knowledge. These viewpoints may be changed or confirmed by new knowledge and observations. Bayes’ formula is a rational method for adjusting our viewpoints as we confront new information.17 Bayes’ formula and related concepts have been applied in many business and investment decision- making contexts, including the evaluation of mutual fund performance.18
Bayes’ formula makes use of Equation 6, the total probability rule. To review, that rule expressed the probability of an event as a weighted average of the probabilities of the event, given a set of scenarios. Bayes’ formula works in reverse; more precisely, it reverses the “given that” information. Bayes’ formula uses the occurrence of the event to infer the probability of the scenario generating it. For that reason, Bayes’ formula is sometimes called an inverse probability. In many applications, including the one illustrating its use in this section, an individual is updating his beliefs concerning the causes that may have produced a new observation.
Bayes’ Formula. Given a set of prior probabilities for an event of interest, if you receive new information, the rule for updating your probability of the event is
Updated probability of event given the new information
In probability notation, this formula can be written concisely as:
To illustrate Bayes’ formula, we work through an investment example that can be adapted to any actual problem. Suppose you are an investor in the stock of
DriveMed, Inc. Positive earnings surprises relative to consensus EPS estimates often result in positive stock returns, and negative surprises often have the opposite effect. DriveMed is preparing to release last quarter ’s EPS result, and you are interested in which of these three events happened: last quarter’s EPS exceeded the consensus EPS estimate, or last quarter’s EPS exactly met the consensus EPS estimate, or last quarter’s EPS fell short of the consensus EPS estimate. This list of the alternatives is mutually exclusive and exhaustive.
On the basis of your own research, you write down the following prior probabilities (or priors, for short) concerning these three events:
P(EPS exceeded consensus) = 0.45
P(EPS met consensus) = 0.30
P(EPS fell short of consensus) = 0.25
These probabilities are “prior” in the sense that they reflect only what you know now, before the arrival of any new information.
The next day, DriveMed announces that it is expanding factory capacity in Singapore and Ireland to meet increased sales demand. You assess this new information. The decision to expand capacity relates not only to current demand but probably also to the prior quarter ’s sales demand. You know that sales demand is positively related to EPS. So now it appears more likely that last quarter ’s EPS will exceed the consensus.
The question you have is, “In light of the new information, what is the updated probability that the prior quarter ’s EPS exceeded the consensus estimate?”
Bayes’ formula provides a rational method for accomplishing this updating. We can abbreviate the new information as DriveMed expands. The first step in applying Bayes’ formula is to calculate the probability of the new information (here: DriveMed expands), given a list of events or scenarios that may have generated it. The list of events should cover all possibilities, as it does here. Formulating these conditional probabilities is the key step in the updating process. Suppose your view is
Conditional probabilities of an observation (here: DriveMed expands) are sometimes referred to as likelihoods. Again, likelihoods are required for updating
the probability.
Next, you combine these conditional probabilities or likelihoods with your prior probabilities to get the unconditional probability for DriveMed expanding, P(DriveMed expands), as follows:
This is Equation 6, the total probability rule, in action. Now you can answer your question by applying Bayes’ formula:
Prior to DriveMed’s announcement, you thought the probability that DriveMed would beat consensus expectations was 45 percent. On the basis of your interpretation of the announcement, you update that probability to 82.3 percent. This updated probability is called your posterior probability because it reflects or comes after the new information.
The Bayes’ calculation takes the prior probability, which was 45 percent, and multiplies it by a ratio—the first term on the right-hand side of the equal sign. The denominator of the ratio is the probability that DriveMed expands, as you view it without considering (conditioning on) anything else. Therefore, this probability is unconditional. The numerator is the probability that DriveMed expands, if last quarter ’s EPS actually exceeded the consensus estimate. This last probability is larger than unconditional probability in the denominator, so the ratio (1.83 roughly) is greater than 1. As a result, your updated or posterior probability is larger than your prior probability. Thus, the ratio reflects the impact of the new information on your prior beliefs.
EXAMPLE 13 Inferring whether DriveMed’s EPS Met Consensus EPS
You are still an investor in DriveMed stock. To review the givens, your prior probabilities are P(EPS exceeded consensus) = 0.45, P(EPS met consensus) = 0.30, and P(EPS fell short of consensus) = 0.25. You also have the following conditional probabilities:
Recall that you updated your probability that last quarter ’s EPS exceeded the consensus estimate from 45 percent to 82.3 percent after DriveMed announced it would expand. Now you want to update your other priors.
1. Update your prior probability that DriveMed’s EPS met consensus.
2. Update your prior probability that DriveMed’s EPS fell short of consensus.
3. Show that the three updated probabilities sum to 1. (Carry each probability to four decimal places.)
4. Suppose, because of lack of prior beliefs about whether DriveMed would meet consensus, you updated on the basis of prior probabilities that all three possibilities were equally likely: P(EPS exceeded consensus) = P(EPS met consensus) = P(EPS fell short of consensus) = 1/3. What is your estimate of the probability P(EPS exceeded consensus | DriveMed expands)?
Solution to 1: The probability is P(EPS met consensus | DriveMed expands) =
The probability P(DriveMed expands) is found by taking each of the three conditional probabilities in the statement of the problem, such as P(DriveMed expands | EPS exceededconsensus); multiplying each one by the prior probability of the conditioning event, such as P(EPS exceeded consensus); then adding the three products. The calculation is unchanged from the problem in the text above: P(DriveMed expands) = 0.75(0.45) + 0.20(0.30) + 0.05(0.25) = 0.41, or 41 percent. The other probabilities needed, P(DriveMed expands | EPS met consensus) = 0.20 and P(EPS met consensus) = 0.30, are givens. So
After taking account of the announcement on expansion, your updated probability that last quarter ’s EPS for DriveMed just met consensus is 14.6 percent compared with your prior probability of 30 percent.
Solution to 2: P(DriveMed expands) was already calculated as 41 percent. Recall that P(DriveMed expands | EPS fell short of consensus) = 0.05 and P(EPS fell short of consensus) = 0.25 are givens.
As a result of the announcement, you have revised your probability that DriveMed’s EPS fell short of consensus from 25 percent (your prior probability) to 3 percent.
Solution to 3: The sum of the three updated probabilities is
The three events (EPS exceeded consensus, EPS met consensus, EPS fell short of consensus) are mutually exclusive and exhaustive: One of these events or statements must be true, so the conditional probabilities must sum to 1. Whether we are talking about conditional or unconditional probabilities, whenever we have a complete set of the distinct possible events or outcomes, the probabilities must sum to 1. This calculation serves as a check on your work.
Solution to 4: Using the probabilities given in the question,
Not surprisingly, the probability of DriveMed expanding is 1/3 because the decision maker has no prior beliefs or views regarding how well EPS performed relative to the consensus estimate. Now we can use Bayes’ formula to find P(EPS exceeded consensus | DriveMed expands) = [P(DriveMed expands
| EPS exceeded consensus)/P(DriveMed expands)] P(EPS exceeded consensus) = [(0.75/(1/3)](1/3) = 0.75 or 75 percent. This probability is identical to your estimate of P(DriveMed expands | EPS exceeded consensus).
When the prior probabilities are equal, the probability of information given an event equals the probability of the event given the information. When a decision-maker has equal prior probabilities (called diffuse priors), the probability of an event is determined by the information.
4.2. Principles of Counting The first step in addressing a question often involves determining the different logical possibilities. We may also want to know the number of ways that each of these possibilities can happen. In the back of our mind is often a question about probability. How likely is it that I will observe this particular possibility? Records of success and failure are an example. When we evaluate a market timer ’s record, one well-known evaluation method uses counting methods presented in this section.19 An important investment model, the binomial option pricing model, incorporates the combination formula that we will cover shortly. We can also use the methods in this section to calculate what we called a priori probabilities in Section 2. When we can assume that the possible outcomes of a random variable are equally likely, the probability of an event equals the number of possible outcomes favorable for the event divided by the total number of outcomes.
In counting, enumeration (counting the outcomes one by one) is of course the most basic resource. What we discuss in this section are shortcuts and principles. Without these shortcuts and principles, counting the total number of outcomes can be very difficult and prone to error. The first and basic principle of counting is the multiplication rule.
Multiplication Rule of Counting. If one task can be done in n1 ways, and a second task, given the first, can be done in n2 ways, and a third task, given the first two tasks, can be done in n3 ways, and so on for k tasks, then the number of ways the k tasks can be done is (n1)(n2)(n3) … (nk).
Suppose we have three steps in an investment decision process. The first step can be done in two ways, the second in four ways, and the third in three ways. Following the multiplication rule, there are (2)(4)(3) = 24 ways in which we can carry out the three steps.
Another illustration is the assignment of members of a group to an equal number of
positions. For example, suppose you want to assign three security analysts to cover three different industries. In how many ways can the assignments be made? The first analyst may be assigned in three different ways. Then two industries remain. The second analyst can be assigned in two different ways. Then one industry remains. The third and last analyst can be assigned in only one way. The total number of different assignments equals (3)(2)(1) = 6. The compact notation for the multiplication we have just performed is 3! (read: 3 factorial). If we had n analysts, the number of ways we could assign them to n tasks would be
or n factorial. (By convention, 0! = 1.) To review, in this application we repeatedly carry out an operation (here, job assignment) until we use up all members of a group (here, three analysts). With n members in the group, the multiplication formula reduces to n factorial.20
The next type of counting problem can be called labeling problems.21 We want to give each object in a group a label, to place it in a category. The following example illustrates this type of problem.
A mutual fund guide ranked 18 bond mutual funds by total returns for the year 2014. The guide also assigned each fund one of five risk labels: high risk (four funds), above-average risk (four funds), average risk (three funds), below-average risk (four funds), and low risk (three funds); as 4 + 4 + 3 + 4 + 3 = 18, all the funds are accounted for. How many different ways can we take 18 mutual funds and label 4 of them high risk, 4 above-average risk, 3 average risk, 4 below-average risk, and 3 low risk, so that each fund is labeled?
The answer is close to 13 billion. We can label any of 18 funds high risk (the first slot), then any of 17 remaining funds, then any of 16 remaining funds, then any of 15 remaining funds (now we have 4 funds in the high risk group); then we can label any of 14 remaining funds above-average risk, then any of 13 remaining funds, and so forth. There are 18! possible sequences. However, order of assignment within a category does not matter. For example, whether a fund occupies the first or third slot of the four funds labeled high risk, the fund has the same label (high risk). Thus there are 4! ways to assign a given group of four funds to the four high risk slots. Making the same argument for the other categories, in total there are (4!) (4!)(3!) (4!)(3!) equivalent sequences. To eliminate such redundancies from the 18! total, we divide 18! by (4!)(4!)(3!)(4!)(3!). We have 18!/(4!)(4!)(3!)(4!)(3!) = 18!/(24)(24)(6) (24)(6) = 12,864,852,000. This procedure generalizes as follows.
Multinomial Formula (General Formula for Labeling Problems). The
number of ways that n objects can be labeled with k different labels, with n1 of the first type, n2 of the second type, and so on, with n1 + n2 + … + nk = n, is given by
The multinomial formula with two different labels (k = 2) is especially important. This special case is called the combination formula. A combination is a listing in which the order of the listed items does not matter. We state the combination formula in a traditional way, but no new concepts are involved. Using the notation in the formula below, the number of objects with the first label is r = n1 and the number with the second label is n − r = n2 (there are just two categories, so n1 + n2 = n). Here is the formula:
Combination Formula (Binomial Formula). The number of ways that we can choose r objects from a total of n objects, when the order in which the r objects are listed does not matter, is
Here nCr and are shorthand notations for n!/(n − r)!r! (read: n choose r, or n combination r).
If we label the r objects as belongs to the group and the remaining objects as does not belong to the group, whatever the group of interest, the combination formula tells us how many ways we can select a group of size r. We can illustrate this formula with the binomial option pricing model. This model describes the movement of the underlying asset as a series of moves, price up (U) or price down (D). For example, two sequences of five moves containing three up moves, such as UUUDD and UDUUD, result in the same final stock price. At least for an option with a payoff dependent on final stock price, the number but not the order of up moves in a sequence matters. How many sequences of five moves belong to the group with three up moves? The answer is 10, calculated using the combination formula (“5 choose 3”):
A useful fact can be illustrated as follows: 5C3 = 5!/2!3! equals 5C2 = 5!/3!2!, as 3 + 2
= 5; 5C4 = 5!/1!4! equals 5C1 = 5!/4!1!, as 4 + 1 = 5. This symmetrical relationship can save work when we need to calculate many possible combinations.
Suppose jurors want to select three companies out of a group of five to receive the first-, second-, and third-place awards for the best annual report. In how many ways can the jurors make the three awards? Order does matter if we want to distinguish among the three awards (the rank within the group of three); clearly the question makes order important. On the other hand, if the question were “In how many ways can the jurors choose three winners, without regard to place of finish?” we would use the combination formula.
To address the first question above, we need to count ordered listings such as first place, New Company; second place, Fir Company; third place, Well Company. An ordered listing is known as a permutation, and the formula that counts the number of permutations is known as the permutation formula.22
Permutation Formula. The number of ways that we can choose r objects from a total of n objects, when the order in which the r objects are listed does matter, is
So the jurors have 5P3 = 5!/(5 − 3)! = (5)(4)(3)(2)(1)/(2)(1) = 120/2 = 60 ways in which they can make their awards. To see why this formula works, note that (5)(4) (3)(2)(1)/(2)(1) reduces to (5)(4)(3), after cancellation of terms. This calculation counts the number of ways to fill three slots choosing from a group of five people, according to the multiplication rule of counting. This number is naturally larger than it would be if order did not matter (compare 60 to the value of 10 for “5 choose 3” that we calculated above). For example, first place, Well Company; second place, Fir Company; third place, New Company contains the same three companies as first place, New Company; second place, Fir Company; third place, Well Company. If we were concerned only with award winners (without regard to place of finish), the two listings would count as one combination. But when we are concerned with the order of finish, the listings count as two permutations.
Answering the following questions may help you apply the counting methods we have presented in this section. 1. Does the task that I want to measure have a finite number of possible
outcomes? If the answer is yes, you may be able to use a tool in this section, and you can go to the second question. If the answer is no, the number of
outcomes is infinite, and the tools in this section do not apply.
2. Do I want to assign every member of a group of size n to one of n slots (or tasks)? If the answer is yes, use n factorial. If the answer is no, go to the third question.
3. Do I want to count the number of ways to apply one of three or more labels to each member of a group? If the answer is yes, use the multinomial formula. If the answer is no, go to the fourth question.
4. Do I want to count the number of ways that I can choose r objects from a total of n, when the order in which I list the r objects does not matter (can I give the r objects a label)? If the answer to these questions is yes, the combination formula applies. If the answer is no, go to the fifth question.
5. Do I want to count the number of ways I can choose r objects from a total of n, when the order in which I list the r objects is important? If the answer is yes, the permutation formula applies. If the answer is no, go to question 6.
6. Can the multiplication rule of counting be used? If it cannot, you may have to count the possibilities one by one, or use more advanced techniques than those presented here.23
5. Summary In this reading, we have discussed the essential concepts and tools of probability. We have applied probability, expected value, and variance to a range of investment problems.
A random variable is a quantity whose outcome is uncertain.
Probability is a number between 0 and 1 that describes the chance that a stated event will occur.
An event is a specified set of outcomes of a random variable.
Mutually exclusive events can occur only one at a time. Exhaustive events cover or contain all possible outcomes.
The two defining properties of a probability are, first, that 0 ≤ P(E) ≤ 1 (where P(E) denotes the probability of an event E), and second, that the sum of the probabilities of any set of mutually exclusive and exhaustive events equals 1.
A probability estimated from data as a relative frequency of occurrence is an empirical probability. A probability drawing on personal or subjective judgment is a subjective probability. A probability obtained based on logical analysis is an a priori probability.
A probability of an event E, P(E), can be stated as odds for E = P(E)/[1 − P(E)] or odds against E = [1 − P(E)]/P(E).
Probabilities that are inconsistent create profit opportunities, according to the Dutch Book Theorem.
A probability of an event not conditioned on another event is an unconditional probability. The unconditional probability of an event A is denoted P(A). Unconditional probabilities are also called marginal probabilities.
A probability of an event given (conditioned on) another event is a conditional probability. The probability of an event A given an event B is denoted P(A | B).
The probability of both A and B occurring is the joint probability of A and B, denoted P(AB).
P(A | B) = P(AB)/P(B), P(B) | 0.
The multiplication rule for probabilities is P(AB) = P(A | B)P(B).
The probability that A or B occurs, or both occur, is denoted by P(A or B).
The addition rule for probabilities is P(A or B) = P(A) + P(B) − P(AB).
When events are independent, the occurrence of one event does not affect the probability of occurrence of the other event. Otherwise, the events are dependent.
The multiplication rule for independent events states that if A and B are independent events, P(AB) = P(A)P(B). The rule generalizes in similar fashion to more than two events.
According to the total probability rule, if S1, S2, …, Sn are mutually exclusive and exhaustive scenarios or events, then P(A) = P(A | S1)P(S1) + P(A | S2)P(S2) + … + P(A | Sn)P(Sn).
The expected value of a random variable is a probability-weighted average of the possible outcomes of the random variable. For a random variable X, the expected value of X is denoted E(X).
The total probability rule for expected value states that E(X) = E(X | S1)P(S1) + E(X | S2)P(S2) + … + E(X | Sn)P(Sn), where S1, S2, …, Sn are mutually exclusive and exhaustive scenarios or events.
The variance of a random variable is the expected value (the probability- weighted average) of squared deviations from the random variable’s expected value E(X): σ2(X) = E{[X − E(X)]2}, where σ2(X) stands for the variance of X.
Variance is a measure of dispersion about the mean. Increasing variance indicates increasing dispersion. Variance is measured in squared units of the original variable.
Standard deviation is the positive square root of variance. Standard deviation measures dispersion (as does variance), but it is measured in the same units as the variable.
Covariance is a measure of the co-movement between random variables.
The covariance between two random variables Ri and Rj is the expected value of the cross-product of the deviations of the two random variables from their respective means: Cov(Ri,Rj) = E{[Ri − E(Ri)][Rj − E(Rj)]}. The covariance of a random variable with itself is its own variance.
Correlation is a number between −1 and +1 that measures the co-movement (linear association) between two random variables: ρ(Ri,Rj) = Cov(Ri,Rj)/[σ(Ri) σ(Rj)].
To calculate the variance of return on a portfolio of n assets, the inputs needed are the n expected returns on the individual assets, n variances of return on the individual assets, and n(n − 1)/2 distinct covariances.
Portfolio variance of return is .
The calculation of covariance in a forward-looking sense requires the specification of a joint probability function, which gives the probability of joint occurrences of values of the two random variables.
When two random variables are independent, the joint probability function is the product of the individual probability functions of the random variables.
Bayes’ formula is a method for updating probabilities based on new information.
Bayes’ formula is expressed as follows: Updated probability of event given the new information = [(Probability of the new information given event)/(Unconditional probability of the new information)] × Prior probability of event.
The multiplication rule of counting says, for example, that if the first step in a process can be done in 10 ways, the second step, given the first, can be done in 5 ways, and the third step, given the first two, can be done in 7 ways, then the steps can be carried out in (10)(5)(7) = 350 ways.
The number of ways to assign every member of a group of size n to n slots is n! = n (n − 1) (n − 2)(n − 3) … 1. (By convention, 0! = 1.)
The number of ways that n objects can be labeled with k different labels, with n1 of the first type, n2 of the second type, and so on, with n1 + n2 + … + nk = n, is given by n!/(n1!n2! … nk!). This expression is the multinomial formula.
A special case of the multinomial formula is the combination formula. The number of ways to choose r objects from a total of n objects, when the order in which the r objects are listed does not matter, is
The number of ways to choose r objects from a total of n objects, when the order in which the r objects are listed does matter, is
This expression is the permutation formula.
References
1. Bodie, Zvi, Alex Kane, and Alan J. Marcus. 2012. Essentials of Investments, 9th edition. New York: McGraw-Hill Irwin.
2. Elton, Edwin J., Martin J. Gruber, Stephen J. Brown, and William N. Goetzmann. 2013. Modern Portfolio Theory and Investment Analysis, 9th edition. Hoboken, NJ: Wiley.
3. Feller, William. 1957. An Introduction to Probability Theory and Its Applications, Vol. I, 2nd edition. New York: Wiley.
4. Henriksson, Roy D., and Robert C. Merton. 1981. “On Market Timing and Investment Performance, II. Statistical Procedures for Evaluating Forecasting Skills.” Journal of Business, vol. 54, no. 4: 513–533.
5. Huij, Joop, and Marno Verbeek. 2007. “Cross-Sectional Learning and Short- Run Persistence in Mutual Fund Performance.” Journal of Banking & Finance, vol. 31: 973–997.
6. Kemeny, John G., Arthur Schleifer, J. Laurie Snell, and Gerald L. Thompson. 1972. Finite Mathematics with Business Applications, 2nd edition. Englewood Cliffs, NJ: Prentice-Hall.
7. Kool, Clemens J.M. 2000. “International Bond Markets and the Introduction of the Euro.” Review of the Federal Reserve Bank of St. Louis. vol. 82, no. 5: 41– 56.
8. Lo, Andrew W. 1999. “The Three P’s of Total Risk Management.” Financial Analysts Journal, vol. 55, no. 1: 13–26.
9. Ramsey, Frank P. 1931. “Truth and Probability.” In The Foundations of Mathematics and Other Logical Essays, edited by R.B. Braithwaite. London: Routledge and Keegan Paul.
10. Reilly, Frank K., and Keith C. Brown. 2012. Investment Analysis and Portfolio Management, 10th edition. Mason, OH: Cengage South-Western.
11. Thanatawee, Yordying. 2013. “Ownership Structure with Dividend Policy: Evidence from Thailand.” International Journal of Economics and Finance, vol. 5, no. 1: 121–132.
12. Vidal-Garcia, Javier. 2013. “The Persistence of European Mutual Fund Performance.” Research in International Business and Finance, vol. 28: 45–67.
Problems Practice Problems and Solutions: 1–18 taken from Quantitative Methods for Investment Analysis, Second Edition, by Richard A. DeFusco, CFA, Dennis W. McLeavey, CFA, Jerald E. Pinto, CFA, and David E. Runkle, CFA. Copyright © 2004 by CFA Institute. All other problems and solutions copyright © CFA Institute. 1. Define the following terms:
1. Probability.
2. Conditional probability.
3. Event.
4. Independent events.
5. Variance.
2. State three mutually exclusive and exhaustive events describing the reaction of a company’s stock price to a corporate earnings announcement on the day of the announcement.
3. Label each of the following as an empirical, a priori, or subjective probability.
1. The probability that US stock returns exceed long-term corporate bond returns over a 10-year period, based on Ibbotson Associates data.
2. An updated (posterior) probability of an event arrived at using Bayes’ formula and the perceived prior probability of the event.
3. The probability of a particular outcome when exactly 12 equally likely possible outcomes exist.
4. A historical probability of default for double-B rated bonds, adjusted to reflect your perceptions of changes in the quality of double-B rated issuance.
4. You are comparing two companies, BestRest Corporation and Relaxin, Inc. The exports of both companies stand to benefit substantially from the removal of import restrictions on their products in a large export market. The price of BestRest shares reflects a probability of 0.90 that the restrictions will be removed within the year. The price of Relaxin stock, however, reflects a 0.50 probability that the restrictions will be removed within that time frame. By all other information related to valuation, the two stocks appear comparably valued. How would you characterize the implied probabilities reflected in share
prices? Which stock is relatively overvalued compared to the other?
5. Suppose you have two limit orders outstanding on two different stocks. The probability that the first limit order executes before the close of trading is 0.45. The probability that the second limit order executes before the close of trading is 0.20. The probability that the two orders both execute before the close of trading is 0.10. What is the probability that at least one of the two limit orders executes before the close of trading?
6. Suppose that 5 percent of the stocks meeting your stock-selection criteria are in the telecommunications (telecom) industry. Also, dividend-paying telecom stocks are 1 percent of the total number of stocks meeting your selection criteria. What is the probability that a stock is dividend paying, given that it is a telecom stock that has met your stock selection criteria?
7. You are using the following three criteria to screen potential acquisition targets from a list of 500 companies:
Criterion Fraction of the 500 CompaniesMeeting the Criterion Product lines compatible 0.20
Company will increase combined sales growth rate 0.45
Balance sheet impact manageable 0.78
If the criteria are independent, how many companies will pass the screen?
8. You apply both valuation criteria and financial strength criteria in choosing stocks. The probability that a randomly selected stock (from your investment universe) meets your valuation criteria is 0.25. Given that a stock meets your valuation criteria, the probability that the stock meets your financial strength criteria is 0.40. What is the probability that a stock meets both your valuation and financial strength criteria?
9. A report from Fitch data service states the following two facts:1
In 2002, the volume of defaulted US high-yield debt was $109.8 billion. The average market size of the high-yield bond market during 2002 was $669.5 billion.
The average recovery rate for defaulted US high-yield bonds in 2002 (defined as average price one month after default) was $0.22 on the dollar.
Address the following three tasks:
1. On the basis of the first fact given above, calculate the default rate on US high-yield debt in 2002. Interpret this default rate as a probability.
2. State the probability computed in Part A as an odds against default.
3. The quantity 1 minus the recovery rate given in the second fact above is the expected loss per $1 of principal value, given that default has occurred. Suppose you are told that an institution held a diversified high-yield bond portfolio in 2002. Using the information in both facts, what was the institution’s expected loss in 2002, per $1 of principal value of the bond portfolio?
10. You are given the following probability distribution for the annual sales of ElStop Corporation:
Probability Distribution for ElStop Annual Sales
Probability Sales ($ Millions) 0.20 275 0.40 250 0.25 200 0.10 190 0.05 180
Sum = 1.00 1. Calculate the expected value of ElStop’s annual sales.
2. Calculate the variance of ElStop’s annual sales.
3. Calculate the standard deviation of ElStop’s annual sales.
11. Suppose the prospects for recovering principal for a defaulted bond issue depend on which of two economic scenarios prevails. Scenario 1 has probability 0.75 and will result in recovery of $0.90 per $1 principal value with probability 0.45, or in recovery of $0.80 per $1 principal value with probability 0.55. Scenario 2 has probability 0.25 and will result in recovery of $0.50 per $1 principal value with probability 0.85, or in recovery of $0.40 per $1 principal value with probability 0.15.
1. Compute the probability of each of the four possible recovery amounts: $0.90, $0.80, $0.50, and $0.40.
2. Compute the expected recovery, given the first scenario.
3. Compute the expected recovery, given the second scenario.
4. Compute the expected recovery.
5. Graph the information in a tree diagram.
12. Suppose we have the expected daily returns (in terms of US dollars), standard deviations, and correlations shown in the table below.
US, German, and Italian Bond Returns
US Dollar Daily Returns in Percent US Bonds German Bonds Italian Bonds
Expected Return 0.029 0.021 0.073 Standard Deviation 0.409 0.606 0.635
Correlation Matrix US Bonds German Bonds Italian Bonds
US Bonds 1 0.09 0.10 German Bonds 1 0.70 Italian Bonds 1
Source: Kool (2000), Table 1 (excerpted and adapted). 1. Using the data given above, construct a covariance matrix for the daily
returns on US, German, and Italian bonds.
2. State the expected return and variance of return on a portfolio 70 percent invested in US bonds, 20 percent in German bonds, and 10 percent in Italian bonds.
3. Calculate the standard deviation of return for the portfolio in Part B.
13. The variance of a stock portfolio depends on the variances of each individual stock in the portfolio and also the covariances among the stocks in the portfolio. If you have five stocks, how many unique covariances (excluding variances) must you use in order to compute the variance of return on your portfolio? (Recall that the covariance of a stock with itself is the stock’s variance.)
14. Calculate the covariance of the returns on Bedolf Corporation (RB) with the returns on Zedock Corporation (RZ), using the following data.
Probability Function of Bedolf and Zedock Returns
RZ = 15% RZ = 10% RZ = 5%
RB = 30% 0.25 0 0 RB = 15% 0 0.50 0 RB = 10% 0 0 0.25
Note: Entries are joint probabilities.
15. You have developed a set of criteria for evaluating distressed credits. Companies that do not receive a passing score are classed as likely to go bankrupt within 12 months. You gathered the following information when validating the criteria:
Forty percent of the companies to which the test is administered will go bankrupt within 12 months: P(nonsurvivor) = 0.40.
Fifty-five percent of the companies to which the test is administered pass it: P(pass test) = 0.55.
The probability that a company will pass the test given that it will subsequently survive 12 months, is 0.85: P(pass test | survivor) = 0.85.
1. What is P(pass test | nonsurvivor)?
2. Using Bayes’ formula, calculate the probability that a company is a survivor, given that it passes the test; that is, calculate P(survivor | pass test).
3. What is the probability that a company is a nonsurvivor, given that it fails the test?
4. Is the test effective?
16. On one day in March, 3,292 issues traded on the NYSE: 1,303 advanced, 1,764 declined, and 225 were unchanged. In how many ways could this set of outcomes have happened? (Set up the problem but do not solve it.)
17. Your firm intends to select 4 of 10 vice presidents for the investment committee. How many different groups of four are possible?
18. As in Example 11, you are reviewing the pricing of a speculative-grade, one- year-maturity, zero-coupon bond. Your goal is to estimate an appropriate default risk premium for this bond. The default risk premium is defined as the extra return above the risk-free return that will compensate investors for default risk. If R is the promised return (yield-to-maturity) on the debt
instrument and RF is the risk-free rate, the default risk premium is R − RF. You assess that the probability that the bond defaults is 0.06: P(the bond defaults) = 0.06. One-year US T-bills are offering a return of 5.8 percent, an estimate of RF. In contrast to your approach in Example 11, you no longer make the simplifying assumption that bondholders will recover nothing in the event of a default. Rather, you now assume that recovery will be $0.35 on the dollar, given default.
1. Denote the fraction of principal and interest recovered in default as θ. Following the model of Example 11, develop a general expression for the promised return R on this bond.
2. Given your expression for R and the estimate of RF, state the minimum default risk premium you should require for this instrument.
19. An analyst developed two scenarios with respect to the recovery of $100,000 principal from defaulted loans:
Scenario Probabilityof Scenario(%) AmountRecovered
($) Probabilityof Amount
(%) 1 40 50,000 60
30,000 40 2 60 80,000 90
60,000 10
The amount of the expected recovery is closest to: 1. $36,400.
2. $63,600.
3. $81,600.
20. The correlation coefficient that indicates the weakest linear relationship between variables is:
1. −0.75.
2. −0.22.
3. 0.35.
Notes 1 In the reading on common probability distributions, we describe some of the
probability distributions most frequently used in investment applications.
2 Selling short or shorting stock means selling borrowed shares in the hope of repurchasing them later at a lower price.
3 The theorem’s name comes from the terminology of wagering. Suppose someone places a $100 bet on X at odds of 10 to 1 against X, and later he is able to place a $600 bet against X at odds of 1 to 1 against X. Whatever the outcome of X, that person makes a riskless profit (equal to $400 if X occurs or $500 if X does not occur) because the implied probabilities are inconsistent. He is said to have made a Dutch book in X. Ramsey (1931) presented the problem of inconsistent probabilities. See also Lo (1999).
4 In analyses of probabilities presented in tables, unconditional probabilities usually appear at the ends or margins of the table, hence the term marginal probability. Because of possible confusion with the way marginal is used in economics (roughly meaning incremental), we use the term unconditional probability throughout this discussion.
5 In this example, the conditional probability is greater than the unconditional probability. The conditional probability of an event may, however, be greater than, equal to, or less than the unconditional probability, depending on the facts. For instance, the probability that the stock earns a return above the risk-free rate given that the stock earns a negative return is 0.
6 Sequential comparisons of quarterly EPS are with the immediate prior quarter. A sequential comparison stands in contrast to a comparison with the same quarter one year ago (another frequent type of comparison).
7 For readers familiar with mathematical treatments of probability, S, a notation usually reserved for a concept called the sample space, is being appropriated to stand for scenario.
8 For simplicity, we model all random variables in this reading as discrete random variables, which have a countable set of outcomes. For continuous random variables, which are discussed along with discrete random variables in the reading on common probability distributions, the operation corresponding to summation is integration.
9 The unconditional variance of EPS is the sum of two terms: 1) the expected value (probability-weighted average) of the conditional variances (parallel to the total probability rules) and 2) the variance of conditional expected values of EPS. The second term arises because the variability in conditional expected value is a source of risk. Term 1 is σ2(EPS) = P(declining interest rate environment) σ2(EPS | declining interest rate environment) + P(stable interest rate environment) σ2(EPS | stable interest rate environment) = 0.60(0.004219) + 0.40(0.0096) = 0.006371. Term 2 is σ2[E(EPS | interest rate environment)] = 0.60($2.4875 − $2.34)2 + 0.40($2.12 − $2.34)2 = 0.032414. Summing the two terms, unconditional variance equals 0.006371 + 0.032414 = 0.038785.
10 Although we outline a number of basic concepts in this section, we do not present mean–variance analysis per se. For a presentation of mean–variance analysis, see the readings on portfolio concepts, as well as the extended treatments in standard investment textbooks such as Bodie, Kane, and Marcus (2012), Elton, Gruber, Brown, and Goetzmann (2013), and Reilly and Brown (2012).
11 Useful facts about variance and covariance include: 1) The variance of a constant times a random variable equals the constant squared times the variance of the random variable, or σ2(wR) = w2σ2(R); 2) The variance of a constant plus a random variable equals the variance of the random variable, or σ2(w + R) = σ2(R) because a constant has zero variance; 3) The covariance between a constant and a random variable is zero.
12 When the meaning of covariance as “off-diagonal covariance” is obvious, as it is here, we omit the qualifying words. Covariance is usually used in this sense.
13 If the correlation is 0, R1 = a + bR2 + error, with b = 0.
14 If the correlation is positive, R1 = a + bR2 + error, with b > 0. If the correlation is negative, b < 0.
15 See any of the textbooks mentioned in Footnote 10.
16 Otherwise, the calculation depends on conditional expected value; the calculation can be expressed as E(XY) = E(X) E(Y | X).
17 Named after the Reverend Thomas Bayes (1702–61).
18 See Huij and Verbeek (2007).
19 Henriksson and Merton (1981).
20 The shortest explanation of n factorial is that it is the number of ways to order n objects in a row. In all the problems to which we apply this counting method, we must use up all the members of a group (sampling without replacement).
21 This discussion follows Kemeny, Schleifer, Snell, and Thompson (1972) in terminology and approach.
22 A more formal definition states that a permutation is an ordered subset of n distinct objects.
23 Feller (1957) contains a very full treatment of counting problems and solution methods.
1 “High Yield Defaults 2002: The Perfect Storm,” 19 February, 2003.
CHAPTER 5 COMMON PROBABILITY DISTRIBUTIONS Richard A. DeFusco, CFA
Dennis W. McLeavey, CFA
Jerald E. Pinto, PhD, CFA
David E. Runkle, PhD, CFA
LEARNING OUTCOMES
After completing this chapter, you will be able to do the following:
define a probability distribution and distinguish between discrete and continuous random variables and their probability functions;
describe the set of possible outcomes of a specified discrete random variable;
interpret a cumulative distribution function;
calculate and interpret probabilities for a random variable, given its cumulative distribution function;
define a discrete uniform random variable, a Bernoulli random variable, and a binomial random variable;
calculate and interpret probabilities given the discrete uniform and the binomial distribution functions;
construct a binomial tree to describe stock price movement;
calculate and interpret tracking error;
define the continuous uniform distribution and calculate and interpret probabilities, given a continuous uniform distribution;
explain the key properties of the normal distribution;
distinguish between a univariate and a multivariate distribution and explain the role of correlation in the multivariate normal distribution;
determine the probability that a normally distributed random variable lies inside a given interval;
define the standard normal distribution, explain how to standardize a random variable, and calculate and interpret probabilities using the standard normal distribution;
define shortfall risk, calculate the safety-first ratio, and select an optimal portfolio using Roy’s safety-first criterion;
explain the relationship between normal and lognormal distributions and why the lognormal distribution is used to model asset prices;
distinguish between discretely and continuously compounded rates of return and calculate and interpret a continuously compounded rate of return, given a specific holding period return;
explain Monte Carlo simulation and describe its applications and limitations;
compare Monte Carlo simulation and historical simulation.
1. Introduction to Common Probability Distributions In nearly all investment decisions we work with random variables. The return on a stock and its earnings per share are familiar examples of random variables. To make probability statements about a random variable, we need to understand its probability distribution. A probability distribution specifies the probabilities of the possible outcomes of a random variable.
In this reading, we present important facts about four probability distributions and their investment uses. These four distributions—the uniform, binomial, normal, and lognormal—are used extensively in investment analysis. They are used in such basic valuation models as the Black–Scholes–Merton option pricing model, the binomial option pricing model, and the capital asset pricing model. With the working knowledge of probability distributions provided in this reading, you will also be better prepared to study and use other quantitative methods such as hypothesis testing, regression analysis, and time-series analysis.
After discussing probability distributions, we end the reading with an introduction to Monte Carlo simulation, a computer-based tool for obtaining information on complex problems. For example, an investment analyst may want to experiment with an investment idea without actually implementing it. Or she may need to price a complex option for which no simple pricing formula exists. In these cases and many others, Monte Carlo simulation is an important resource. To conduct a Monte Carlo simulation, the analyst must identify risk factors associated with the problem and specify probability distributions for them. Hence, Monte Carlo simulation is a tool that requires an understanding of probability distributions.
Before we discuss specific probability distributions, we define basic concepts and terms. We then illustrate the operation of these concepts through the simplest distribution, the uniform distribution. That done, we address probability distributions that have more applications in investment work but also greater complexity.
2. Discrete Random Variables A random variable is a quantity whose future outcomes are uncertain. The two basic types of random variables are discrete random variables and continuous random variables. A discrete random variable can take on at most a countable number of possible values. For example, a discrete random variable X can take on a limited number of outcomes x1, x2, …, xn (n possible outcomes), or a discrete random variable Y can take on an unlimited number of outcomes y1, y2, … (without end).1
Because we can count all the possible outcomes of X and Y (even if we go on forever in the case of Y), both X and Y satisfy the definition of a discrete random variable. By contrast, we cannot count the outcomes of a continuous random variable. We cannot describe the possible outcomes of a continuous random variable Z with a list z1, z2, … because the outcome (z1 + z2)/2, not in the list, would always be possible. Rate of return is an example of a continuous random variable.
In working with a random variable, we need to understand its possible outcomes. For example, a majority of the stocks traded on the New Zealand Stock Exchange are quoted in ticks of NZ$0.01. Quoted stock price is thus a discrete random variable with possible values NZ$0, NZ$0.01, NZ$0.02, … But we can also model stock price as a continuous random variable (as a lognormal random variable, to look ahead). In many applications, we have a choice between using a discrete or a continuous distribution. We are usually guided by which distribution is most efficient for the task we face. This opportunity for choice is not surprising, as many discrete distributions can be approximated with a continuous distribution, and vice versa. In most practical cases, a probability distribution is only a mathematical idealization, or approximate model, of the relative frequencies of a random variable’s possible outcomes.
EXAMPLE 1 The Distribution of Bond Price
You are researching a probability model for bond price, and you begin by thinking about the characteristics of bonds that affect price. What are the lowest and the highest possible values for bond price? Why? What are some other characteristics of bonds that may affect the distribution of bond price?
The lowest possible value of bond price is 0, when the bond is worthless. Identifying the highest possible value for bond price is more challenging. The promised payments on a coupon bond are the coupons (interest payments) plus the face amount (principal). The price of a bond is the present discounted value of these promised payments. Because investors require a return on their
investments, 0 percent is the lower limit on the discount rate that investors would use to discount a bond’s promised payments. At a discount rate of 0 percent, the price of a bond is the sum of the face value and the remaining coupons without any discounting. The discount rate thus places the upper limit on bond price. Suppose, for example, that face value is $1,000 and two $40 coupons remain; the interval $0 to $1,080 captures all possible values of the bond’s price. This upper limit decreases through time as the number of remaining payments decreases.
Other characteristics of a bond also affect its price distribution. Pull to par value is one such characteristic: As the maturity date approaches, the standard deviation of bond price tends to grow smaller as bond price converges to par value. Embedded options also affect bond price. For example, with bonds that are currently callable, the issuer may retire the bonds at a prespecified premium above par; this option of the issuer cuts off part of the bond’s upside. Modeling bond price distribution is a challenging problem.
Every random variable is associated with a probability distribution that describes the variable completely. We can view a probability distribution in two ways. The basic view is the probability function, which specifies the probability that the random variable takes on a specific value: P(X = x) is the probability that a random variable X takes on the value x. (Note that capital X represents the random variable and lowercase x represents a specific value that the random variable may take.) For a discrete random variable, the shorthand notation for the probability function is p(x) = P(X = x). For continuous random variables, the probability function is denoted f(x) and called the probability density function (pdf), or just the density.2
A probability function has two key properties (which we state, without loss of generality, using the notation for a discrete random variable):
0 ≤ p(x) ≤ 1, because probability is a number between 0 and 1.
The sum of the probabilities p(x) over all values of X equals 1. If we add up the probabilities of all the distinct possible outcomes of a random variable, that sum must equal 1.
We are often interested in finding the probability of a range of outcomes rather than a specific outcome. In these cases, we take the second view of a probability distribution, the cumulative distribution function (cdf). The cumulative distribution function, or distribution function for short, gives the probability that a random variable X is less than or equal to a particular value x, P(X ≤ x). For both discrete and continuous random variables, the shorthand notation is F(x) = P(X ≤ x). How
does the cumulative distribution function relate to the probability function? The word “cumulative” tells the story. To find F(x), we sum up, or cumulate, values of the probability function for all outcomes less than or equal to x. The function of the cdf is parallel to that of cumulative relative frequency, which we discussed in the reading on statistical concepts and market returns.
Next, we illustrate these concepts with examples and show how we use discrete and continuous distributions. We start with the simplest distribution, the discrete uniform.
2.1. The Discrete Uniform Distribution The simplest of all probability distributions is the discrete uniform distribution. Suppose that the possible outcomes are the integers (whole numbers) 1 to 8, inclusive, and the probability that the random variable takes on any of these possible values is the same for all outcomes (that is, it is uniform). With eight outcomes, p(x) = 1/8, or 0.125, for all values of X (X = 1, 2, 3, 4, 5, 6, 7, 8); the statement just made is a complete description of this discrete uniform random variable. The distribution has a finite number of specified outcomes, and each outcome is equally likely. Table 1 summarizes the two views of this random variable, the probability function and the cumulative distribution function.
TABLE 1 Probability Function and Cumulative Distribution Function for a Discrete Uniform Random Variable
X = x
Probability Function p(x) = P(X = x)
Cumulative Distribution Function F(x) = P(X ≤ x)
1 0.125 0.125 2 0.125 0.250 3 0.125 0.375 4 0.125 0.500 5 0.125 0.625 6 0.125 0.750 7 0.125 0.875 8 0.125 1.000
We can use Table 1 to find three probabilities: P(X ≤ 7), P(4 ≤ X ≤ 6), and P(4 < X ≤ 6). The following examples illustrate how to use the cdf to find the probability that a
random variable will fall in any interval (for any random variable, not only the uniform).
The probability that X is less than or equal to 7, P(X ≤ 7), is the next-to-last entry in the third column, 0.875 or 87.5 percent.
To find P(4 ≤ X ≤ 6), we need to find the sum of three probabilities: p(4), p(5), and p(6). We can find this sum in two ways. We can add p(4), p(5), and p(6) from the second column. Or we can calculate the probability as the difference between two values of the cumulative distribution function:
so
So we calculate the second probability as F(6) − F(3) = 3/8.
The third probability, P(4 < X ≤ 6), the probability that X is less than or equal to 6 but greater than 4, is p(5) + p(6). We compute it as follows, using the cdf:
So we calculate the third probability as F(6) − F(4) = 2/8.
Suppose we want to check that the discrete uniform probability function satisfies the general properties of a probability function given earlier. The first property is 0 ≤ p(x) ≤ 1. We see that p(x) = 1/8 for all x in the first column of the table. (Note that p(x) equals 0 for numbers x such as −14 or 12.215 that are not in that column.) The first property is satisfied. The second property is that the probabilities sum to 1. The entries in the second column of Table 1 do sum to 1.
The cdf has two other characteristic properties:
The cdf lies between 0 and 1 for any x: 0 ≤ F(x) ≤ 1.
As we increase x, the cdf either increases or remains constant.
Check these statements by looking at the third column in Table 1.
We now have some experience working with probability functions and cdfs for discrete random variables. Later in this reading, we will discuss Monte Carlo simulation, a methodology driven by random numbers. As we will see, the uniform
distribution has an important technical use: It is the basis for generating random numbers, which in turn produce random observations for all other probability distributions.3
2.2. The Binomial Distribution In many investment contexts, we view a result as either a success or a failure, or as binary (twofold) in some other way. When we make probability statements about a record of successes and failures, or about anything with binary outcomes, we often use the binomial distribution. What is a good model for how a stock price moves through time? Different models are appropriate for different uses. Cox, Ross, and Rubinstein (1979) developed an option pricing model based on binary moves, price up or price down, for the asset underlying the option. Their binomial option pricing model was the first of a class of related option pricing models that have played an important role in the development of the derivatives industry. That fact alone would be sufficient reason for studying the binomial distribution, but the binomial distribution has uses in decision-making as well.
The building block of the binomial distribution is the Bernoulli random variable, named after the Swiss probabilist Jakob Bernoulli (1654–1704). Suppose we have a trial (an event that may repeat) that produces one of two outcomes. Such a trial is a Bernoulli trial. If we let Y equal 1 when the outcome is success and Y equal 0 when the outcome is failure, then the probability function of the Bernoulli random variable Y is
where p is the probability that the trial is a success. Our next example is the very first step on the road to understanding the binomial option pricing model.
EXAMPLE 2 One-Period Stock Price Movement as a Bernoulli Random Variable
Suppose we describe stock price movement in the following way. Stock price today is S. Next period stock price can move up or down. The probability of an up move is p, and the probability of a down move is 1 − p. Thus, stock price is a Bernoulli random variable with probability of success (an up move) equal to p. When the stock moves up, ending price is uS, with u equal to 1 plus the rate of return if the stock moves up. For example, if the stock earns 0.01 or 1 percent on an up move, u = 1.01. When the stock moves down, ending price is
dS, with d equal to 1 plus the rate of return if the stock moves down. For example, if the stock earns −0.01 or −1 percent on a down move, d = 0.99. Figure 1 shows a diagram of this model of stock price dynamics.
FIGURE 1 One-Period Stock Price as a Bernoulli Random Variable
We will continue with the above example later. In the model of stock price movement in Example 2, success and failure at a given trial relate to up moves and down moves, respectively. In the following example, success is a profitable trade and failure is an unprofitable one.
EXAMPLE 3 A Trading Desk Evaluates Block Brokers (1)
You work in equities trading at an institutional money manager that regularly trades with a number of block brokers. Blocks are orders to sell or buy that are too large for the liquidity ordinarily available in dealer networks or stock exchanges. Your firm has known interests in certain kinds of stock. Block brokers call your trading desk when they want to sell blocks of stocks that they think your firm may be interested in buying. You know that these transactions have definite risks. For example, if the broker ’s client (the seller of the shares) has unfavorable information on the stock, or if the total amount he is selling through all channels is not truthfully communicated to you, you may see an immediate loss on the trade. From time to time, your firm audits the performance of block brokers. Your firm calculates the post-trade, market- risk-adjusted dollar returns on stocks purchased from block brokers. On that basis, you classify each trade as unprofitable or profitable. You have summarized the performance of the brokers in a spreadsheet, excerpted in Table 2 for November 2014. (The broker names are coded BB001 and BB002.)
TABLE 2 Block Trading Gains and Losses
November 2014
Profitable Trades Losing Trades BB001 3 9 BB002 5 3
View each trade as a Bernoulli trial. Calculate the percentage of profitable trades with the two block brokers for November 2014. These are estimates of p, the underlying probability of a successful (profitable) trade with each broker.
Your firm has logged 3 + 9 = 12 trades (the row total) with block broker BB001. Because 3 of the 12 trades were profitable, the percentage of profitable trades was 3/12 or 25 percent. With broker BB002, the percentage of profitable trades was 5/8 or 62.5 percent. A trade is a Bernoulli trial, and the above calculations provide estimates of the underlying probability of a profitable trade (success) with the two brokers. For broker BB001, your estimate is = 25; for broker BB002, your estimate is = 0.625.4
In n Bernoulli trials, we can have 0 to n successes. If the outcome of an individual trial is random, the total number of successes in n trials is also random. A binomial random variable X is defined as the number of successes in n Bernoulli trials. A binomial random variable is the sum of Bernoulli random variables Yi, i = 1, 2, …, n:
where Yi is the outcome on the ith trial (1 if a success, 0 if a failure). We know that a Bernoulli random variable is defined by the parameter p. The number of trials, n, is the second parameter of a binomial random variable. The binomial distribution makes these assumptions:
The probability, p, of success is constant for all trials.
The trials are independent.
The second assumption has great simplifying force. If individual trials were correlated, calculating the probability of a given number of successes in n trials would be much more complicated.
Under the above two assumptions, a binomial random variable is completely described by two parameters, n and p. We write
which we read as “X has a binomial distribution with parameters n and p.” You can see that a Bernoulli random variable is a binomial random variable with n = 1: Y ∼ B(1, p).
Now we can find the general expression for the probability that a binomial random variable shows x successes in n trials. We can think in terms of a model of stock price dynamics that can be generalized to allow any possible stock price movements if the periods are made extremely small. Each period is a Bernoulli trial: With probability p, the stock price moves up; with probability 1 − p, the price moves down. A success is an up move, and x is the number of up moves or successes in n periods (trials). With each period’s moves independent and p constant, the number of up moves in n periods is a binomial random variable. We now develop an expression for P(X = x), the probability function for a binomial random variable.
Any sequence of n periods that shows exactly x up moves must show n − x down moves. We have many different ways to order the up moves and down moves to get a total of x up moves, but given independent trials, any sequence with x up moves must occur with probability px(1 − p)n–x. Now we need to multiply this probability by the number of different ways we can get a sequence with x up moves. Using a basic result in counting from the reading on probability concepts, there are
different sequences in n trials that result in x up moves (or successes) and n − x down moves (or failures). Recall from the reading on probability concepts that n factorial (n!) is defined as n(n − 1)(n − 2) … 1 (and 0! = 1 by convention). For example, 5! = (5)(4)(3)(2)(1) = 120. The combination formula n!/[(n − x)!x!] is denoted by
(read “n combination x” or “n choose x”). For example, over three periods, exactly three different sequences have two up moves: UUD, UDU, and DUU. We confirm this by
If, hypothetically, each sequence with two up moves had a probability of 0.15, then the total probability of two up moves in three periods would be 3 × 0.15 = 0.45. This example should persuade you that for X distributed B(n, p), the probability of x successes in n trials is given by
(1)
Some distributions are always symmetric, such as the normal, and others are always asymmetric or skewed, such as the lognormal. The binomial distribution is symmetric when the probability of success on a trial is 0.50, but it is asymmetric or skewed otherwise.
We illustrate Equation 1 (the probability function) and the cdf through the symmetrical case. Consider a random variable distributed B(n = 5, p = 0.50). Table 3 contains a complete description of this random variable. The fourth column of Table 3 is Column 2, n combination x, times Column 3, px(1 − p)n–x; Column 4 gives the probability for each value of the number of up moves from the first column. The fifth column, cumulating the entries in the fourth column, is the cumulative distribution function.
TABLE 3 Binomial Probabilities, p = 0.50 and n = 5
Number of Up Moves,
x(1)
Number of Possible Ways to Reach x Up
Moves(2)
Probability for Each Way(3)
Probability for x, p(x)(4) = (2) ×
(3)
F(x) = P(X ≤ x)
(5)
0 1 0.500(1 − 0.50)5 = 0.03125
0.03125 0.03125
1 5 0.501(1 − 0.50)4 = 0.03125
0.15625 0.18750
2 10 0.502(1 − 0.50)3 = 0.03125
0.31250 0.50000
3 10 0.503(1 − 0.50)2 = 0.03125
0.31250 0.81250
4 5 0.504(1 − 0.50)1 = 0.03125
0.15625 0.96875
5 1 0.505(1 − 0.50)0 = 0.03125
0.03125 1.00000
What would happen if we kept n = 5 but sharply lowered the probability of success on a trial to 10 percent? “Probability for Each Way” for X = 0 (no up moves) would then be about 59 percent: 0.100(1 − 0.10)5 = 0.59049. Because zero successes could still happen one way (Column 2), p(0) = 59 percent. You may want to check that given p = 0.10, P(X ≤ 2) = 99.14 percent: The probability of two or fewer up moves would be more than 99 percent. The random variable’s probability would be massed on 0, 1, and 2 up moves, and the probability of larger outcomes would be minute. The outcomes of 3 and larger would be the long right tail, and the distribution would be right skewed. On the other hand, if we set p = 0.90, we would have the mirror image of the distribution with p = 0.10. The distribution would be left skewed.
With an understanding of the binomial probability function in hand, we can continue with our example of block brokers.
EXAMPLE 4 A Trading Desk Evaluates Block Brokers (2)
You now want to evaluate the performance of the block brokers in Example 3. You begin with two questions:
1. If you are paying a fair price on average in your trades with a broker, what
should be the probability of a profitable trade?
2. Did each broker meet or miss that expectation on probability?
You also realize that the brokers’ performance has to be evaluated in light of the sample’s size, and for that you need to use the binomial probability function (Equation 1). You thus address the following (referring to the data in Example 3):
3. Under the assumption that the prices of trades were fair,
1. calculate the probability of three or fewer profitable trades with broker BB001.
2. calculate the probability of five or more profitable trades with broker BB002.
Solution to 1 and 2: If the price you trade at is fair, 50 percent of the trades you do with a broker should be profitable.5 The rate of profitable trades with broker BB001 was 25 percent. Therefore, broker BB001 missed your performance expectation. Broker BB002, at 62.5 percent profitable trades,
exceeded your expectation.
Solution to 3: 1. For broker BB001, the number of trades (the trials) was n = 12, and 3 were
profitable. You are asked to calculate the probability of three or fewer profitable trades, F(3) = p(3) + p(2) + p(1) + p(0).
Suppose the underlying probability of a profitable trade with BB001 is p = 0.50. With n = 12 and p = 0.50, according to Equation 1 the probability of three profitable trades is
The probability of exactly three profitable trades out of 12 is 5.4 percent if broker BB001 were giving you fair prices. Now you need to calculate the other probabilities:
Adding all the probabilities, F(3) = 0.053711 + 0.016113 + 0.00293 + 0.000244 = 0.072998 or 7.3 percent. The probability of doing three or fewer profitable trades out of 12 would be 7.3 percent if your trading desk were getting fair prices from broker BB001.
2. For broker BB002, you are assessing the probability that the underlying probability of a profitable trade with this broker was 50 percent, despite the good results. The question was framed as the probability of doing five or more profitable trades if the underlying probability is 50 percent: 1 − F(4) = p(5) + p(6) + p(7) + p(8). You could calculate F(4) and subtract it from 1, but you can also calculate p(5) + p(6) + p(7) + p(8) directly.
You begin by calculating the probability that exactly five out of eight trades would be profitable if BB002 were giving you fair prices:
The probability is about 21.9 percent. The other probabilities are
So p(5) + p(6) + p(7) + p(8) = 0.21875 + 0.109375 + 0.03125 + 0.003906 = 0.363281 or 36.3 percent.6 A 36.3 percent probability is substantial; the underlying probability of executing a fair trade with BB002 might well have been 0.50 despite your success with BB002 in November 2014. If one of the trades with BB002 had been reclassified from profitable to unprofitable, exactly half the trades would have been profitable. In summary, your trading desk is getting at least fair prices from BB002; you will probably want to accumulate additional evidence before concluding that you are trading at better-than-fair prices.The magnitude of the profits and losses in these trades is another important consideration. If all profitable trades had small profits but all unprofitable trades had large losses, for example, you might lose money on your trades even if the majority of them were profitable.
In the next example, the binomial distribution helps in evaluating the performance of an investment manager.
EXAMPLE 5 Meeting a Tracking Error Objective
You work for a pension fund sponsor. You have assigned a new money manager to manage a $500 million portfolio indexed on the MSCI EAFE (Europe, Australasia, and Far East) Index, which is designed to measure developed-market equity performance excluding the United States and Canada. After research, you believe it is reasonable to expect that the manager will keep tracking error within a band of 75 basis points (bps) of the benchmark’s return, on a quarterly basis.7 Tracking error is the total return on the portfolio (gross of fees) minus the total return on the benchmark index—here, the EAFE.8 To quantify this expectation further, you will be satisfied if tracking error is within the 75 bps band 90 percent of the time. The manager meets the objective in six out of eight quarters. Of course, six out of eight quarters is a 75 percent success rate. But how does the manager ’s record precisely relate to your expectation of a 90 percent success rate and the sample size, 8 observations? To answer this question, you must find the probability that, given an assumed true or underlying success rate of 90 percent, performance could be as bad as or worse than that delivered. Calculate the probability (by hand or with a spreadsheet).
Specifically, you want to find the probability that tracking error is within the 75
bps band in six or fewer quarters out of the eight in the sample. With n = 8 and p = 0.90, this probability is F(6) = p(6) + p(5) + p(4) + p(3) + p(2) + p(1) + p(0). Start with
and work through the other probabilities:
Summing all these probabilities, you conclude that F(6) = 0.148803 + 0.033067 + 0.004593 + 0.000408 + 0.000023 + 0.00000072 + 0.00000001 = 0.186895 or 18.7 percent. There is a moderate 18.7 percent probability that the manager would show the record he did (or a worse record) if he had the skill to meet your expectations 90 percent of the time.
You can use other evaluation concepts such as tracking risk, defined as the standard deviation of tracking error, to assess the manager ’s performance. The calculation above would be only one input into any conclusions that you reach concerning the manager ’s performance. But to answer problems involving success rates, you need to be skilled in using the binomial distribution.
Two descriptors of a distribution that are often used in investments are the mean and the variance (or the standard deviation, the positive square root of variance).9 Table 4 gives the expressions for the mean and variance of binomial random variables.
Because a single Bernoulli random variable, Y ∼ B(1, p), takes on the value 1 with probability p and the value 0 with probability 1 − p, its mean or weighted-average outcome is p. Its variance is p(1 − p).10 A general binomial random variable, B(n, p), is the sum of n Bernoulli random variables, and so the mean of a B(n, p) random variable is np. Given that a B(1, p) variable has variance p(1 − p), the variance of a B(n, p) random variable is n times that value, or np(1 − p), assuming that all the trials (Bernoulli random variables) are independent. We can illustrate the calculation for two binomial random variables with differing probabilities as follows:
Random Variable Mean Variance B(n = 5, p = 0.50) 2.50 = 5(0.50) 1.25 = 5(0.50)(0.50) B(n = 5, p = 0.10) 0.50 = 5(0.10) 0.45 = 5(0.10)(0.90)
For a B(n = 5, p = 0.50) random variable, the expected number of successes is 2.5 with a standard deviation of 1.118 = (1.25)1/2; for a B(n = 5, p = 0.10) random variable, the expected number of successes is 0.50 with a standard deviation of 0.67 = (0.45)1/2.
TABLE 4 Mean and Variance of Binomial Random Variables
Mean Variance Bernoulli, B(1, p) p p(1 − p) Binomial, B(n, p) np np(1 − p)
EXAMPLE 6 The Expected Number of Defaults in a Bond Portfolio
Suppose as a bond analyst you are asked to estimate the number of bond issues expected to default over the next year in an unmanaged high-yield bond portfolio with 25 US issues from distinct issuers. The credit ratings of the bonds in the portfolio are tightly clustered around Moody’s B2/Standard & Poor ’s B, meaning that the bonds are speculative with respect to the capacity to pay interest and repay principal. The estimated annual default rate for B2/B rated bonds is 10.7 percent.
1. Over the next year, what is the expected number of defaults in the portfolio,
assuming a binomial model for defaults?
2. Estimate the standard deviation of the number of defaults over the coming year.
3. Critique the use of the binomial probability model in this context.
Solution to 1: For each bond, we can define a Bernoulli random variable equal to 1 if the bond defaults during the year and zero otherwise. With 25 bonds, the expected number of defaults over the year is np = 25(0.107) = 2.675 or approximately 3.
Solution to 2: The variance is np(1 − p) = 25(0.107)(0.893) = 2.388775. The standard deviation is (2.388775)1/2 = 1.55. Thus a two standard deviation confidence interval about the expected number of defaults would run from approximately 0 to approximately 6, for example.
Solution to 3: An assumption of the binomial model is that the trials are independent. In this context, a trial relates to whether an individual bond issue
will default over the next year. Because the issuing companies probably share exposure to common economic factors, the trials may not be independent. Nevertheless, for a quick estimate of the expected number of defaults, the binomial model may be adequate.
Earlier, we looked at a simple one-period model for stock price movement. Now we extend the model to describe stock price movement on three consecutive days. Each day is an independent trial. The stock moves up with constant probability p (the up transition probability); if it moves up, u is 1 plus the rate of return for an up move. The stock moves down with constant probability 1 − p (the down transition probability); if it moves down, d is 1 plus the rate of return for a down move. We graph stock price movement in Figure 2, where we now associate each of the n = 3 stock price moves with time indexed by t. The shape of the graph suggests why it is a called a binomial tree. Each boxed value from which successive moves or outcomes branch in the tree is called a node; in this example, a node is potential value for the stock price at a specified time.
We see from the tree that the stock price at t = 3 has four possible values: uuuS, uudS, uddS, and dddS. The probability that the stock price equals any one of these four values is given by the binomial distribution. For example, three sequences of moves result in a final stock price of uudS: These are uud, udu, and duu. These sequences have two up moves out of three moves in total; the combination formula confirms that the number of ways to get two up moves (successes) in three periods (trials) is 3!/(3 − 2)!2! = 3. Next note that each of these sequences, uud, udu, and duu, has probability p2(1 − p). So P(S3 = uudS) = 3p2(1 − p), where S3 indicates the stock’s price after three moves.
FIGURE 2 A Binomial Model of Stock Price Movement
The binomial random variable in this application is the number of up moves. Final stock price distribution is a function of the initial stock price, the number of up moves, and the size of the up moves and down moves. We cannot say that stock price
itself is a binomial random variable; rather, it is a function of a binomial random variable, as well as of u and d, and initial price. This richness is actually one key to why this way of modeling stock price is useful: It allows us to choose values of these parameters to approximate various distributions for stock price (using a large number of time periods).11 One distribution that can be approximated is the lognormal, an important continuous distribution model for stock price that we will discuss later. The flexibility extends further. In the tree shown above, the transition probabilities are the same at each node: p for an up move and 1 − p for a down move. That standard formula describes a process in which stock return volatility is constant through time. Option experts, however, sometimes model changing volatility through time using a binomial tree in which the probabilities for up and down moves differ at different nodes.
The binomial tree also supplies the possibility of testing a condition or contingency at any node. This flexibility is useful in investment applications such as option pricing. Consider an American call option on a dividend-paying stock. (Recall that an American option can be exercised at any time before expiration, at any node on the tree.) Just before an ex-dividend date, it may be optimal to exercise an American call option on stock to buy the stock and receive the dividend.12 If we model stock price with a binomial tree, we can test, at each node, whether exercising the option is optimal. Also, if we know the value of the call at the four terminal nodes at t = 3 and we have a model for discounting values by one period, we can step backward one period to t = 2 to find the call’s value at the three nodes there. Continuing back recursively, we can find the call’s value today. This type of recursive operation is easily programmed on a computer. As a result, binomial trees can value options even more complex than American calls on stock.13
3. Continuous Random Variables In the previous section, we considered discrete random variables (i.e., random variables whose set of possible outcomes is countable). In contrast, the possible outcomes of continuous random variables are never countable. If 1.250 is one possible value of a continuous random variable, for example, we cannot name the next higher or lower possible value. Technically, the range of possible outcomes of a continuous random variable is the real line (all real numbers between −∞ and +∞) or some subset of the real line.
In this section, we focus on the two most important continuous distributions in investment work, the normal and lognormal. As we did with discrete distributions, we introduce the topic through the uniform distribution.
3.1. Continuous Uniform Distribution The continuous uniform distribution is the simplest continuous probability distribution. The uniform distribution has two main uses. As the basis of techniques for generating random numbers, the uniform distribution plays a role in Monte Carlo simulation. As the probability distribution that describes equally likely outcomes, the uniform distribution is an appropriate probability model to represent a particular kind of uncertainty in beliefs in which all outcomes appear equally likely.
The pdf for a uniform random variable is
For example, with a = 0 and b = 8, f (x) = 1/8 or 0.125. We graph this density in Figure 3.
The graph of the density function plots as a horizontal line with a value of 0.125.
What is the probability that a uniform random variable with limits a = 0 and b = 8 is less than or equal to 3, or F(3) = P(X ≤ 3)? When we were working with the discrete uniform random variable with possible outcomes 1, 2, …, 8, we summed individual probabilities: p(1) + p(2) + p(3) = 0.375. In contrast, the probability that a continuous uniform random variable, or any continuous random variable, assumes any given fixed value is 0. To illustrate this point, consider the narrow interval 2.510 to 2.511. Because that interval holds an infinity of possible values, the sum of the probabilities of values in that interval alone would be infinite if each individual
value in it had a positive probability. To find the probability F(3), we find the area under the curve graphing the pdf, between 0 to 3 on the x axis. In calculus, this operation is called integrating the probability function f (x) from 0 to 3. This area under the curve is a rectangle with base 3 − 0 = 3 and height 1/8. The area of this rectangle equals base times height: 3(1/8) = 3/8 or 0.375. So F(3) = 3/8 or 0.375.
FIGURE 3 Continuous Uniform Distribution
The interval from 0 to 3 is three-eighths of the total length between the limits of 0 and 8, and F(3) is three-eighths of the total probability of 1. The middle line of the expression for the cdf captures this relationship.
For our problem, F(x) = 0 for x ≤ 0, F(x) = x/8 for 0 < x < 8, and F(x) = 1 for x ≥ 8. We graph this cdf in Figure 4.
FIGURE 4 Continuous Uniform Cumulative Distribution
The mathematical operation that corresponds to finding the area under the curve of a pdf f (x) from a to b is the integral of f (x) from a to b:
(2)
where ∫ dx is the symbol for summing ∫ over small changes dx, and the limits of integration (a and b) can be any real numbers or −∞ and +∞. All probabilities of continuous random variables can be computed using Equation 2. For the uniform distribution example considered above, F(7) is Equation 2 with lower limit a = 0 and upper limit b = 7. The integral corresponding to the cdf of a uniform distribution reduces to the three-line expression given previously. To evaluate Equation 2 for nearly all other continuous distributions, including the normal and lognormal, we rely on spreadsheet functions, computer programs, or tables of values to calculate probabilities. Those tools use various numerical methods to evaluate the integral in Equation 2.
Recall that the probability of a continuous random variable equaling any fixed point is 0. This fact has an important consequence for working with the cumulative distribution function of a continuous random variable: For any continuous random variable X, P(a ≤ X ≤ b) = P(a < X ≤ b) = P(a ≤ X < b) = P(a < X < b), because the probabilities at the endpoints a and b are 0. For discrete random variables, these relations of equality are not true, because probability accumulates at points.
EXAMPLE 7 Probability That a Lending Facility Covenant Is Breached
You are evaluating the bonds of a below-investment-grade borrower at a low point in its business cycle. You have many factors to consider, including the terms of the company’s bank lending facilities. The contract creating a bank lending facility such as an unsecured line of credit typically has clauses known as covenants. These covenants place restrictions on what the borrower can do. The company will be in breach of a covenant in the lending facility if the interest coverage ratio, EBITDA/interest, calculated on EBITDA over the four trailing quarters, falls below 2.0. EBITDA is earnings before interest, taxes, depreciation, and amortization.14 Compliance with the covenants will be checked at the end of the current quarter. If the covenant is breached, the bank can demand immediate repayment of all borrowings on the facility. That action would probably trigger a liquidity crisis for the company. With a high degree of confidence, you forecast interest charges of $25 million. Your estimate of EBITDA runs from $40 million on the low end to $60 million on the high end.
Address two questions (treating projected interest charges as a constant): 1. If the outcomes for EBITDA are equally likely, what is the probability that
EBITDA/interest will fall below 2.0, breaching the covenant?
2. Estimate the mean and standard deviation of EBITDA/interest. For a continuous uniform random variable, the mean is given by μ = (a + b)/2 and the variance is given by σ2 = (b − a)2/12.
Solution to 1: EBITDA/interest is a continuous uniform random variable because all outcomes are equally likely. The ratio can take on values between 1.6 = ($40 million)/($25 million) on the low end and 2.4 = ($60 million/$25 million) on the high end. The range of possible values is 2.4 − 1.6 = 0.8. What fraction of the possible values falls below 2.0, the level that triggers default? The distance between 2.0 and 1.6 is 0.40; the value 0.40 is one-half the total length of 0.8, or 0.4/0.8 = 0.50. So the probability that the covenant will be breached is 50 percent.
Solution to 2: In Solution 1, we found that the lower limit of EBITDA/interest is 1.6. This lower limit is a. We found that the upper limit is 2.4. This upper limit is b. Using the formula given above,
The variance of the interest coverage ratio is
The standard deviation is the positive square root of the variance, 0.230940 = (0.053333)1/2. The standard deviation is not particularly useful as a risk measure for a uniform distribution, however. The probability that lies within various standard deviation bands around the mean is sensitive to different specifications of the upper and lower limits (although Chebyshev’s inequality is always satisfied).15 Here, a one standard deviation interval around the mean of 2.0 runs from 1.769 to 2.231 and captures 0.462/0.80 = 0.5775 or 57.8 percent of the probability. A two standard deviation interval runs from 1.538 to 2.462, which extends past both the lower and upper limits of the random variable.
3.2. The Normal Distribution The normal distribution may be the most extensively used probability distribution in quantitative work. It plays key roles in modern portfolio theory and in a number of risk management technologies. Because it has so many uses, the normal distribution must be thoroughly understood by investment professionals.
The role of the normal distribution in statistical inference and regression analysis is vastly extended by a crucial result known as the central limit theorem. The central limit theorem states that the sum (and mean) of a large number of independent random variables is approximately normally distributed.16
The French mathematician Abraham de Moivre (1667–1754) introduced the normal distribution in 1733 in developing a version of the central limit theorem. As Figure 5 shows, the normal distribution is symmetrical and bell-shaped. The range of possible outcomes of the normal distribution is the entire real line: all real numbers lying between −∞ and +∞. The tails of the bell curve extend without limit to the left and to the right.
FIGURE 5 Two Normal Distributions
The defining characteristics of a normal distribution are as follows:
The normal distribution is completely described by two parameters—its mean, μ, and variance, σ2. We indicate this as X ∼ N(μ, σ2) (read “X follows a normal distribution with mean μ and variance σ2”). We can also define a normal distribution in terms of the mean and the standard deviation, σ (this is often convenient because σ is measured in the same units as X and μ). As a consequence, we can answer any probability question about a normal random variable if we know its mean and variance (or standard deviation).
The normal distribution has a skewness of 0 (it is symmetric). The normal distribution has a kurtosis (measure of peakedness) of 3; its excess kurtosis (kurtosis − 3.0) equals 0.17 As a consequence of symmetry, the mean, median, and the mode are all equal for a normal random variable.
A linear combination of two or more normal random variables is also normally distributed.
These bullet points concern a single variable or univariate normal distribution: the distribution of one normal random variable. A univariate distribution describes a single random variable. A multivariate distribution specifies the probabilities for a group of related random variables. You will encounter the multivariate normal distribution in investment work and reading and should know the following about it.
When we have a group of assets, we can model the distribution of returns on each
asset individually, or the distribution of returns on the assets as a group. “As a group” means that we take account of all the statistical interrelationships among the return series. One model that has often been used for security returns is the multivariate normal distribution. A multivariate normal distribution for the returns on n stocks is completely defined by three lists of parameters:
the list of the mean returns on the individual securities (n means in total);
the list of the securities’ variances of return (n variances in total); and
the list of all the distinct pairwise return correlations: n(n − 1)/2 distinct correlations in total.18
The need to specify correlations is a distinguishing feature of the multivariate normal distribution in contrast to the univariate normal distribution.
The statement “assume returns are normally distributed” is sometimes used to mean a joint normal distribution. For a portfolio of 30 securities, for example, portfolio return is a weighted average of the returns on the 30 securities. A weighted average is a linear combination. Thus, portfolio return is normally distributed if the individual security returns are (joint) normally distributed. To review, in order to specify the normal distribution for portfolio return, we need the means, variances, and the distinct pairwise correlations of the component securities.
With these concepts in mind, we can return to the normal distribution for one random variable. The curves graphed in Figure 5 are the normal density function:
(3)
The two densities graphed in Figure 5 correspond to a mean of μ = 0 and standard deviations of σ = 1 and σ = 2. The normal density with μ = 0 and σ = 1 is called the standard normal distribution (or unit normal distribution). Plotting two normal distributions with the same mean and different standard deviations helps us appreciate why standard deviation is a good measure of dispersion for the normal distribution: Observations are much more concentrated around the mean for the normal distribution with σ = 1 than for the normal distribution with σ = 2.
Although not literally accurate, the normal distribution can be considered an approximate model for returns. Nearly all the probability of a normal random variable is contained within three standard deviations of the mean. For realistic values of mean return and return standard deviation for many assets, the normal probability of outcomes below −100 percent is very small. Whether the
approximation is useful in a given application is an empirical question. For example, the normal distribution is a closer fit for quarterly and yearly holding period returns on a diversified equity portfolio than it is for daily or weekly returns.19 A persistent departure from normality in most equity return series is kurtosis greater than 3, the fat-tails problem. So when we approximate equity return distributions with the normal distribution, we should be aware that the normal distribution tends to underestimate the probability of extreme returns.20 Option returns are skewed. Because the normal is a symmetrical distribution, we should be cautious in using the normal distribution to model the returns on portfolios containing significant positions in options.
The normal distribution, however, is less suitable as a model for asset prices than as a model for returns. A normal random variable has no lower limit. This characteristic has several implications for investment applications. An asset price can drop only to 0, at which point the asset becomes worthless. As a result, practitioners generally do not use the normal distribution to model the distribution of asset prices. Also note that moving from any level of asset price to 0 translates into a return of −100 percent. Because the normal distribution extends below 0 without limit, it cannot be literally accurate as a model for asset returns.
Having established that the normal distribution is the appropriate model for a variable of interest, we can use it to make the following probability statements:
Approximately 50 percent of all observations fall in the interval μ ± (2/3) σ.
Approximately 68 percent of all observations fall in the interval μ ± σ.
Approximately 95 percent of all observations fall in the interval μ ± 2σ.
Approximately 99 percent of all observations fall in the interval μ ± 3σ.
One, two, and three standard deviation intervals are illustrated in Figure 6. The intervals indicated are easy to remember but are only approximate for the stated probabilities. More-precise intervals are μ ± 1.96σ for 95 percent of the observations and μ ± 2.58σ for 99 percent of the observations.
FIGURE 6 Units of Standard Deviation
In general, we do not observe the population mean or the population standard deviation of a distribution, so we need to estimate them.21 We estimate the population mean, μ, using the sample mean, (sometimes denoted as ) and estimate the population standard deviation, σ, using the sample standard deviation, s (sometimes denoted as ).
There are as many different normal distributions as there are choices for mean (μ) and variance (σ2). We can answer all of the above questions in terms of any normal distribution. Spreadsheets, for example, have functions for the normal cdf for any specification of mean and variance. For the sake of efficiency, however, we would like to refer all probability statements to a single normal distribution. The standard normal distribution (the normal distribution with μ = 0 and σ = 1) fills that role.
There are two steps in standardizing a random variable X: Subtract the mean of X from X, then divide that result by the standard deviation of X. If we have a list of observations on a normal random variable, X, we subtract the mean from each observation to get a list of deviations from the mean, then divide each deviation by the standard deviation. The result is the standard normal random variable, Z. (Z is the conventional symbol for a standard normal random variable.) If we have X ∼ N(μ, σ2) (read “X follows the normal distribution with parameters μ and σ2”), we standardize it using the formula
(4)
Suppose we have a normal random variable, X, with μ = 5 and σ = 1.5. We standardize X with Z = (X − 5)/1.5. For example, a value X = 9.5 corresponds to a standardized value of 3, calculated as Z = (9.5 − 5)/1.5 = 3. The probability that we will observe a value as small as or smaller than 9.5 for X ∼ N(5, 1.5) is exactly the same as the probability that we will observe a value as small as or smaller than 3 for Z ∼ N(0, 1). We can answer all probability questions about X using standardized
values and probability tables for Z. We generally do not know the population mean and standard deviation, so we often use the sample mean for μ and the sample standard deviation s for σ.
TABLE 5 P(Z ≤ x) = N(x) for x ≥ 0 or P(Z ≤ z) =N(z) for z ≥ 0
x or z 0 0.01 0.02 0.03 0.04 0.05 0.06 0.07 0.08 0.09 0.00 0.5000 0.5040 0.5080 0.5120 0.5160 0.5199 0.5239 0.5279 0.5319 0.5359 0.10 0.5398 0.5438 0.5478 0.5517 0.5557 0.5596 0.5636 0.5675 0.5714 0.5753 0.20 0.5793 0.5832 0.5871 0.5910 0.5948 0.5987 0.6026 0.6064 0.6103 0.6141 0.30 0.6179 0.6217 0.6255 0.6293 0.6331 0.6368 0.6406 0.6443 0.6480 0.6517 0.40 0.6554 0.6591 0.6628 0.6664 0.6700 0.6736 0.6772 0.6808 0.6844 0.6879 0.50 0.6915 0.6950 0.6985 0.7019 0.7054 0.7088 0.7123 0.7157 0.7190 0.7224
Standard normal probabilities can also be computed with spreadsheets, statistical and econometric software, and programming languages. Tables of the cumulative distribution function for the standard normal random variable are in the back of this book. Table 5 shows an excerpt from those tables. N(x) is a conventional notation for the cdf of a standard normal variable.22
To find the probability that a standard normal variable is less than or equal to 0.24, for example, locate the row that contains 0.20, look at the 0.04 column, and find the entry 0.5948. Thus, P(Z ≤ 0.24) = 0.5948 or 59.48 percent.
The following are some of the most frequently referenced values in the standard normal table:
The 90th percentile point is 1.282: P(Z ≤ 1.282) = N(1.282) = 0.90 or 90 percent, and 10 percent of values remain in the right tail.
The 95th percentile point is 1.65: P(Z ≤ 1.65) = N(1.65) = 0.95 or 95 percent, and 5 percent of values remain in the right tail. Note the difference between the use of a percentile point when dealing with one tail rather than two tails. Earlier, we used 1.65 standard deviations for the 90 percent confidence interval, where 5 percent of values lie outside that interval on each of the two sides. Here we use 1.65 because we are concerned with the 5 percent of values that lie only on one side, the right tail.
The 99th percentile point is 2.327: P(Z ≤ 2.327) = N(2.327) = 0.99 or 99 percent, and 1 percent of values remain in the right tail.
The tables that we give for the normal cdf include probabilities for x ≤ 0. Many sources, however, give tables only for x ≥ 0. How would one use such tables to find a normal probability? Because of the symmetry of the normal distribution, we can find all probabilities using tables of the cdf of the standard normal random variable, P(Z ≤ x) = N(x), for x ≥ 0. The relations below are helpful for using tables for x ≥ 0, as well as in other uses:
For a non-negative number x, use N(x) from the table. Note that for the probability to the right of x, we have P(Z ≥ x) = 1.0 − N(x).
For a negative number −x, N(−x) = 1.0 − N(x): Find N(x) and subtract it from 1. All the area under the normal curve to the left of x is N(x). The balance, 1.0 − N(x), is the area and probability to the right of x. By the symmetry of the normal distribution around its mean, the area and the probability to the right of x are equal to the area and the probability to the left of −x, N(−x).
For the probability to the right of −x, P(Z ≥ −x) = N(x).
EXAMPLE 8 Probabilities for a Common Stock Portfolio
Assume the portfolio mean return is 12 percent and the standard deviation of return estimate is 22 percent per year.
You want to calculate the following probabilities, assuming that a normal distribution describes returns. (You can use the excerpt from the table of normal probabilities to answer these questions.)
1. What is the probability that portfolio return will exceed 20 percent?
2. What is the probability that portfolio return will be between 12 percent and 20 percent? In other words, what is P(12% ≤ Portfolio return ≤ 20%)?
3. You can buy a one-year T-bill that yields 5.5 percent. This yield is effectively a one-year risk-free interest rate. What is the probability that your portfolio’s return will be equal to or less than the risk-free rate?
If X is portfolio return, standardized portfolio return is Z = (X − )/s = (X − 12%)/22%. We use this expression throughout the solutions.
Solution to 1: For X = 20%, Z = (20% − 12%)/22% = 0.363636. You want to find P(Z > 0.363636). First note that P(Z > x) = P(Z ≥ x) because the normal is a
continuous distribution. Recall that P(Z ≥ x) = 1.0 − P(Z ≤ x) or 1 − N(x). Rounding 0.363636 to 0.36, according to the table, N(0.36) = 0.6406. Thus, 1 − 0.6406 = 0.3594. The probability that portfolio return will exceed 20 percent is about 36 percent if your normality assumption is accurate.
Solution to 2: P(12% ≤ Portfolio return ≤ 20%) = N(Z corresponding to 20%) − N(Z corresponding to 12%). For the first term, Z = (20% − 12%)/22% = 0.36 approximately, and N(0.36) = 0.6406 (as in Solution 1). To get the second term immediately, note that 12 percent is the mean, and for the normal distribution 50 percent of the probability lies on either side of the mean. Therefore, N(Z corresponding to 12%) must equal 50 percent. So P(12% ≤ Portfolio return ≤ 20%) = 0.6406 − 0.50 = 0.1406 or approximately 14 percent.
Solution to 3: If X is portfolio return, then we want to find P(Portfolio return ≤ 5.5%). This question is more challenging than Parts 1 or 2, but when you have studied the solution below you will have a useful pattern for calculating other shortfall probabilities.
There are three steps, which involve standardizing the portfolio return: First, subtract the portfolio mean return from each side of the inequality: P(Portfolio return − 12% ≤ 5.5% − 12%). Second, divide each side of the inequality by the standard deviation of portfolio return: P[(Portfolio return − 12%)/22% ≤ (5.5% − 12%)/22%] = P(Z ≤ −0.295455) = N(−0.295455). Third, recognize that on the left-hand side we have a standard normal variable, denoted by Z. As we pointed out above, N(−x) = 1 − N(x). Rounding −0.29545 to −0.30 for use with the excerpted table, we have N(−0.30) = 1 − N(0.30) = 1 − 0.6179 = 0.3821, roughly 38 percent. The probability that your portfolio will underperform the one-year risk-free rate is about 38 percent.
We can get the answer above quickly by subtracting the mean portfolio return from 5.5 percent, dividing by the standard deviation of portfolio return, and evaluating the result (−0.295455) with the standard normal cdf.
3.3. Applications of the Normal Distribution Modern portfolio theory (MPT) makes wide use of the idea that the value of investment opportunities can be meaningfully measured in terms of mean return and variance of return. In economic theory, mean–variance analysis holds exactly when investors are risk averse; when they choose investments so as to maximize expected utility, or satisfaction; and when either 1) returns are normally distributed, or 2) investors have quadratic utility functions.23 Mean–variance analysis can still be useful, however—that is, it can hold approximately—when either assumption 1 or 2
is violated. Because practitioners prefer to work with observables such as returns, the proposition that returns are at least approximately normally distributed has played a key role in much of MPT.
Mean–variance analysis generally considers risk symmetrically in the sense that standard deviation captures variability both above and below the mean.24 An alternative approach evaluates only downside risk. We discuss one such approach, safety-first rules, as it provides an excellent illustration of the application of normal distribution theory to practical investment problems. Safety-first rules focus on shortfall risk, the risk that portfolio value will fall below some minimum acceptable level over some time horizon. The risk that the assets in a defined benefit plan will fall below plan liabilities is an example of a shortfall risk.
Suppose an investor views any return below a level of RL as unacceptable. Roy’s safety-first criterion states that the optimal portfolio minimizes the probability that portfolio return, RP, falls below the threshold level, RL.25 In symbols, the investor ’s objective is to choose a portfolio that minimizes P(RP < RL). When portfolio returns are normally distributed, we can calculate P(RP < RL) using the number of standard deviations that RL lies below the expected portfolio return, E(RP). The portfolio for which E(RP) − RL is largest relative to standard deviation minimizes P(RP < RL). Therefore, if returns are normally distributed, the safety-first optimal portfolio maximizes the safety-first ratio (SFRatio):
The quantity E(RP) − RL is the distance from the mean return to the shortfall level. Dividing this distance by σP gives the distance in units of standard deviation. There are two steps in choosing among portfolios using Roy’s criterion (assuming normality):26 1. Calculate each portfolio’s SFRatio.
2. Choose the portfolio with the highest SFRatio.
For a portfolio with a given safety-first ratio, the probability that its return will be less than RL is N(−SFRatio), and the safety-first optimal portfolio has the lowest such probability. For example, suppose an investor ’s threshold return, RL, is 2 percent. He is presented with two portfolios. Portfolio 1 has an expected return of 12 percent with a standard deviation of 15 percent. Portfolio 2 has an expected return of 14 percent with a standard deviation of 16 percent. The SFRatios are 0.667 = (12 −
2)/15 and 0.75 = (14 − 2)/16 for Portfolios 1 and 2, respectively. For the superior Portfolio 2, the probability that portfolio return will be less than 2 percent is N(−0.75) = 1 − N(0.75) = 1 − 0.7734 = 0.227 or about 23 percent, assuming that portfolio returns are normally distributed.
You may have noticed the similarity of SFRatio to the Sharpe ratio. If we substitute the risk-free rate, RF, for the critical level RL, the SFRatio becomes the Sharpe ratio. The safety-first approach provides a new perspective on the Sharpe ratio: When we evaluate portfolios using the Sharpe ratio, the portfolio with the highest Sharpe ratio is the one that minimizes the probability that portfolio return will be less than the risk-free rate (given a normality assumption).
EXAMPLE 9 The Safety-First Optimal Portfolio for a Client
You are researching asset allocations for a client in Canada with a C$800,000 portfolio. Although her investment objective is long-term growth, at the end of a year she may want to liquidate C$30,000 of the portfolio to fund educational expenses. If that need arises, she would like to be able to take out the C$30,000 without invading the initial capital of C$800,000. Table 6 shows three alternative allocations.
Address these questions (assume normality for Parts 2 and 3): 1. Given the client’s desire not to invade the C$800,000 principal, what is the
shortfall level, RL? Use this shortfall level to answer Part 2.
2. According to the safety-first criterion, which of the three allocations is the best?
3. What is the probability that the return on the safety-first optimal portfolio will be less than the shortfall level?
Solution to 1: Because C$30,000/C$800,000 is 3.75 percent, for any return less than 3.75 percent the client will need to invade principal if she takes out C$30,000. So RL = 3.75 percent.
Solution to 2: To decide which of the three allocations is safety-first optimal, select the alternative with the highest ratio [E(RP) − RL]/σP:
Allocation B, with the largest ratio (0.90625), is the best alternative according to the safety-first criterion.
Table 6 Mean and Standard Deviation for Three Allocations (in Percent)
A B C Expected annual return 25 11 14
Standard deviation of return 27 8 20
Solution to 3: To answer this question, note that P(RB < 3.75) = N(−0.90625). We can round 0.90625 to 0.91 for use with tables of the standard normal cdf. First, we calculate N(−0.91) = 1 − N(0.91) = 1 − 0.8186 = 0.1814 or about 18.1 percent. Using a spreadsheet function for the standard normal cdf on −0.90625 without rounding, we get 18.24 percent or about 18.2 percent. The safety-first optimal portfolio has a roughly 18 percent chance of not meeting a 3.75 percent return threshold.
Several points are worth noting. First, if the inputs were even slightly different, we could get a different ranking. For example, if the mean return on B were 10 rather than 11 percent, A would be superior to B. Second, if meeting the 3.75 percent return threshold were a necessity rather than a wish, C$830,000 in one year could be modeled as a liability. Fixed income strategies such as cash flow matching could be used to offset or immunize the C$830,000 quasi-liability.
Roy’s safety-first rule was the earliest approach to addressing shortfall risk. The standard mean–variance portfolio selection process can also accommodate a shortfall risk constraint.27
In many investment contexts besides Roy’s safety-first criterion, we use the normal distribution to estimate a probability. For example, Kolb, Gay, and Hunter (1985) developed an expression based on the standard normal distribution for the probability that a futures trader will exhaust his liquidity because of losses in a futures contract. Another arena in which the normal distribution plays an important role is financial risk management. Financial institutions such as investment banks, security dealers, and commercial banks have formal systems to measure and control financial risk at various levels, from trading positions to the overall risk for the firm.28 Two mainstays in managing financial risk are Value at Risk (VAR) and stress testing/scenario analysis. Stress testing/scenario analysis, a complement to VAR, refers to a set of techniques for estimating losses in extremely unfavorable combinations of events or scenarios. Value at Risk (VAR) is a money measure of the minimum value of losses expected over a specified time period (for example, a
day, a quarter, or a year) at a given level of probability (often 0.05 or 0.01). Suppose we specify a one-day time horizon and a level of probability of 0.05, which would be called a 95 percent one-day VAR.29 If this VAR equaled €5 million for a portfolio, there would be a 0.05 probability that the portfolio would lose €5 million or more in a single day (assuming our assumptions were correct). One of the basic approaches to estimating VAR, the variance-covariance or analytical method, assumes that returns follow a normal distribution. For more information on VAR, see Chance and Brooks (2012).
3.4. The Lognormal Distribution Closely related to the normal distribution, the lognormal distribution is widely used for modeling the probability distribution of share and other asset prices. For example, the lognormal appears in the Black–Scholes–Merton option pricing model. The Black–Scholes–Merton model assumes that the price of the asset underlying the option is lognormally distributed.
A random variable Y follows a lognormal distribution if its natural logarithm, ln Y, is normally distributed. The reverse is also true: If the natural logarithm of random variable Y, ln Y, is normally distributed, then Y follows a lognormal distribution. If you think of the term lognormal as “the log is normal,” you will have no trouble remembering this relationship.
The two most noteworthy observations about the lognormal distribution are that it is bounded below by 0 and it is skewed to the right (it has a long right tail). Note these two properties in the graphs of the pdfs of two lognormal distributions in Figure 7. Asset prices are bounded from below by 0. In practice, the lognormal distribution has been found to be a usefully accurate description of the distribution of prices for many financial assets. On the other hand, the normal distribution is often a good approximation for returns. For this reason, both distributions are very important for finance professionals.
Like the normal distribution, the lognormal distribution is completely described by two parameters. Unlike the other distributions we have considered, a lognormal distribution is defined in terms of the parameters of a different distribution. The two parameters of a lognormal distribution are the mean and standard deviation (or variance) of its associated normal distribution: the mean and variance of ln Y, given that Y is lognormal. Remember, we must keep track of two sets of means and standard deviations (or variances): the mean and standard deviation (or variance) of the associated normal distribution (these are the parameters), and the mean and standard deviation (or variance) of the lognormal variable itself.
The expressions for the mean and variance of the lognormal variable itself are
challenging. Suppose a normal random variable X has expected value μ and variance σ2. Define Y = exp(X). Remember that the operation indicated by exp(X) or eX is the opposite operation from taking logs.30 Because ln Y = ln [exp(X)] = X is normal (we assume X is normal), Y is lognormal. What is the expected value of Y = exp(X)? A guess might be that the expected value of Y is exp(μ). The expected value is actually exp(μ + 0.50σ2), which is larger than exp(μ) by a factor of exp(0.50σ2) > 1.31 To get some insight into this concept, think of what happens if we increase σ2. The distribution spreads out; it can spread upward, but it cannot spread downward past 0. As a result, the center of its distribution is pushed to the right—the distribution’s mean increases.32
The expressions for the mean and variance of a lognormal variable are summarized below, where μ and σ2 are the mean and variance of the associated normal distribution (refer to these expressions as needed, rather than memorizing them):
FIGURE 7 Two Lognormal Distributions
Mean (μL) of a lognormal random variable = exp(μ + 0.50σ2)
Variance (σL2) of a lognormal random variable = exp(2μ + σ2) × [exp(σ2) − 1]
We now explore the relationship between the distribution of stock return and stock price. In the following we show that if a stock’s continuously compounded return is normally distributed, then future stock price is necessarily lognormally distributed.33 Furthermore, we show that stock price may be well described by the lognormal distribution even when continuously compounded returns do not follow a normal distribution. These results provide the theoretical foundation for using the lognormal distribution to model prices.
To outline the presentation that follows, we first show that the stock price at some future time T, ST, equals the current stock price, S0, multiplied by e raised to power r0,T, the continuously compounded return from 0 to T; this relationship is expressed as ST = S0exp(r0,T). We then show that we can write r0,T as the sum of shorter-term continuously compounded returns and that if these shorter-period returns are normally distributed, then r0,T is normally distributed (given certain assumptions) or approximately normally distributed (not making those assumptions). As ST is proportional to the log of a normal random variable, ST is lognormal.
To supply a framework for our discussion, suppose we have a series of equally spaced observations on stock price: S0, S1, S2, …, ST. Current stock price, S0, is a known quantity and so is nonrandom. The future prices (such as S1), however, are random variables. The price relative, S1/S0, is an ending price, S1, over a beginning price, S0; it is equal to 1 plus the holding period return on the stock from t = 0 to t = 1:
For example, if S0 = $30 and S1 = $34.50, then S1/S0 = $34.50/$30 = 1.15. Therefore, R0,1 = 0.15 or 15 percent. In general, price relatives have the form
where Rt, t+1 is the rate of return from t to t + 1.
An important concept is the continuously compounded return associated with a holding period return such as R0,1. The continuously compounded return associated with a holding period is the natural logarithm of 1 plus that holding period return, or equivalently, the natural logarithm of the ending price over the beginning price (the price relative).34 For example, if we observe a one-week holding period return of 0.04, the equivalent continuously compounded return, called the one-week continuously compounded return, is ln(1.04) = 0.039221; €1.00 invested for one week at 0.039221 continuously compounded gives €1.04, equivalent to a 4 percent one-week holding period return. The continuously compounded return from t to t + 1 is
(5)
For our example, r0,1 = ln(S1/S0) = ln(1 + R0,1) = ln($34.50/$30) = ln(1.15) = 0.139762. Thus, 13.98 percent is the continuously compounded return from t = 0 to t
= 1. The continuously compounded return is smaller than the associated holding period return. If our investment horizon extends from t = 0 to t = T, then the continuously compounded return to T is
Applying the function exp to both sides of the equation, we have exp(r0,T) = exp[ln(ST/S0)] = ST/S0, so
We can also express ST/S0 as the product of price relatives:
Taking logs of both sides of this equation, we find that continuously compounded return to time T is the sum of the one-period continuously compounded returns:
(6)
Using holding period returns to find the ending value of a $1 investment involves the multiplication of quantities (1 + holding period return). Using continuously compounded returns involves addition.
A key assumption in many investment applications is that returns are independently and identically distributed (IID). Independence captures the proposition that investors cannot predict future returns using past returns (i.e., weak-form market efficiency). Identical distribution captures the assumption of stationarity.35
Assume that the one-period continuously compounded returns (such as r0,1) are IID random variables with mean μ and variance σ2 (but making no normality or other distributional assumption). Then
(7)
(we add up μ for a total of T times) and
(8)
(as a consequence of the independence assumption). The variance of the T holding period continuously compounded return is T multiplied by the variance of the one- period continuously compounded return; also, σ(r0,T) = . If the one-period continuously compounded returns on the right-hand side of Equation 6 are normally
distributed, then the T holding period continuously compounded return, r0,T, is also normally distributed with mean μT and variance σ2T. This relationship is so because a linear combination of normal random variables is also normal. But even if the one-period continuously compounded returns are not normal, their sum, r0,T, is approximately normal according to a result in statistics known as the central limit theorem.36 Now compare ST = S0exp(r0,T) to Y = exp(X), where X is normal and Y is lognormal (as we discussed above). Clearly, we can model future stock price ST as a lognormal random variable because r0,T should be at least approximately normal. This assumption of normally distributed returns is the basis in theory for the lognormal distribution as a model for the distribution of prices of shares and other assets.
Continuously compounded returns play a role in many option pricing models, as mentioned earlier. An estimate of volatility is crucial for using option pricing models such as the Black–Scholes–Merton model. Volatility measures the standard deviation of the continuously compounded returns on the underlying asset.37 In practice, we very often estimate volatility using a historical series of continuously compounded daily returns. We gather a set of daily holding period returns and then use Equation 5 to convert them into continuously compounded daily returns. We then compute the standard deviation of the continuously compounded daily returns and annualize that number using Equation 8.38 (By convention, volatility is stated as an annualized measure.)39 Example 10 illustrates the estimation of volatility for the shares of Astra International.
EXAMPLE 10 Volatility as Used in Option Pricing Models
Suppose you are researching Astra International (Indonesia Stock Exchange: ASII) and are interested in Astra’s price action in a week in which international economic news had significantly affected the Indonesian stock market. You decide to use volatility as a measure of the variability of Astra shares during that week. Table 7 shows closing prices during that week.
Use the data in Table 7 to do the following: 1. Estimate the volatility of Astra shares. (Annualize volatility based on 250 days
in a year.)
2. Identify the probability distribution for Astra share prices if continuously
compounded daily returns follow the normal distribution.
Solution to 1: First, use Equation 5 to calculate the continuously compounded daily returns; then find their standard deviation in the usual way. (In the calculation of sample variance to get sample standard deviation, use a divisor of 1 less than the sample size.)
The standard deviation of continuously compounded daily returns is 0.021261. Equation 8 states that . In this example, is the sample standard deviation of one-period continuously compounded returns. Thus, refers to 0.021261. We want to annualize, so the horizon T corresponds to one year. As is in days, we set T equal to the number of trading days in a year (250).
We find that annualized volatility for Astra stock that week was 33.6 percent, calculated as .
TABLE 7 Astra International Daily Closing Prices
Date Closing Price (IDR) 17 June 2013 6,950 18 June 2013 7,000 19 June 2013 6,850 20 June 2013 6,600 21 June 2013 6,350
Source: http://finance.yahoo.com.
Note that the sample mean, –0.022572, is a possible estimate of the mean, μ, of the continuously compounded one-period or daily returns. The sample mean can be translated into an estimate of the expected continuously compounded annual return using Equation 7: (using 250 to be consistent with the calculation of volatility). But four observations are far too few to estimate expected returns. The variability in the daily returns overwhelms any information about expected return in a series this short.
Solution to 2: Astra share prices should follow the lognormal distribution if the continuously compounded daily returns on Astra shares follow the normal distribution.
We have shown that the distribution of stock price is lognormal, given certain assumptions. What are the mean and variance of ST if ST follows the lognormal distribution? Earlier in this section, we gave bullet-point expressions for the mean and variance of a lognormal random variable. In the bullet-point expressions, the
would refer, in the context of this discussion, to the mean and variance of the T horizon (not the one-period) continuously compounded returns (assumed to follow a normal distribution), compatible with the horizon of ST.40 Related to the use of mean and variance (or standard deviation), earlier in this reading we used those quantities to construct intervals in which we expect to find a certain percentage of the observations of a normally distributed random variable. Those intervals were symmetric about the mean. Can we state similar, symmetric intervals for a lognormal random variable? Unfortunately, we cannot. Because the lognormal distribution is not symmetric, such intervals are more complicated than for the normal distribution, and we will not discuss this specialist topic here.41
Finally, we have presented the relation between the mean and variance of continuously compounded returns associated with different time horizons (see Equations 7 and 8), but how are the means and variances of holding period returns and continuously compounded returns related? As analysts, we typically think in terms of holding period returns rather than continuously compounded returns, and we may desire to convert means and standard deviations of holding period returns to means and standard deviations of continuously compounded returns for an option application, for example. To effect such conversions (and those in the other direction, from a continuous compounding to a holding period basis), we can use the expressions in Ferguson (1993).
4. Monte Carlo Simulation With an understanding of probability distributions, we are now prepared to learn about a computer-based technique in which probability distributions play an integral role. The technique is called Monte Carlo simulation. Monte Carlo simulation in finance involves the use of a computer to represent the operation of a complex financial system. A characteristic feature of Monte Carlo simulation is the generation of a large number of random samples from a specified probability distribution or distributions to represent the role of risk in the system.
Monte Carlo simulation has several quite distinct uses. One use is in planning. Stanford University researcher Sam Savage provided the following neat picture of that role: “What is the last thing you do before you climb on a ladder? You shake it, and that is Monte Carlo simulation.”42 Just as shaking a ladder helps us assess the risks in climbing it, Monte Carlo simulation allows us to experiment with a proposed policy before actually implementing it. For example, investment performance can be evaluated with reference to a benchmark or a liability. Defined benefit pension plans often invest assets with reference to plan liabilities. Pension liabilities are a complex random process. In a Monte Carlo asset-liability financial planning study, the functioning of pension assets and liabilities is simulated over time, given assumptions about how assets are invested, the work force, and other variables. A key specification in this and all Monte Carlo simulations is the probability distributions of the various sources of risk (including interest rates and security market returns, in this case). The implications of different investment policy decisions on the plan’s funded status can be assessed through simulated time. The experiment can be repeated for another set of assumptions. We can view Example 11 below as coming under this heading. In that example, market return series are not long enough to address researchers’ questions on stock market timing, so the researchers simulate market returns to find answers to their questions.
Monte Carlo simulation is also widely used to develop estimates of VAR. In this application, we simulate the portfolio’s profit and loss performance for a specified time horizon. Repeated trials within the simulation (each trial involving a draw of random observations from a probability distribution) produce a frequency distribution for changes in portfolio value. The point that defines the cutoff for the least favorable 5 percent of simulated changes is an estimate of 95 percent VAR, for example.
In an extremely important use, Monte Carlo simulation is a tool for valuing complex securities, particularly some European-style options for which no analytic pricing formula is available.43 For other securities, such as mortgage-backed
securities with complex embedded options, Monte Carlo simulation is also an important modeling resource.
Researchers use Monte Carlo simulation to test their models and tools. How critical is a particular assumption to the performance of a model? Because we control the assumptions when we do a simulation, we can run the model through a Monte Carlo simulation to examine a model’s sensitivity to a change in our assumptions.
To understand the technique of Monte Carlo simulation, let us present the process as a series of steps.44 To illustrate the steps, we take the case of using Monte Carlo simulation to value a type of option for which no analytic pricing formula is available, an Asian call option on a stock. An Asian call option is a European-style option with a value at maturity equal to the difference between the stock price at maturity and the average stock price during the life of the option, or $0, whichever is greater. For instance, if the final stock price is $34 with an average value of $31 over the life of the option, the value of the option at maturity is $3 (the greater of $34 − $31 = $3 and $0). Steps 1 through 3 of the process describe specifying the simulation; Steps 4 through 7 describe running the simulation. 1. Specify the quantities of interest (option value, for example, or the funded
status of a pension plan) in terms of underlying variables. The underlying variable or variables could be stock price for an equity option, the market value of pension assets, or other variables relating to the pension benefit obligation for a pension plan. Specify the starting values of the underlying variables.
To illustrate the steps, we are using the case of valuing an Asian call option on stock. We use CiT to represent the value of the option at maturity T. The subscript i in CiT indicates that CiT is a value resulting from the ith simulation trial, each simulation trial involving a drawing of random values (an iteration of Step 4).
2. Specify a time grid. Take the horizon in terms of calendar time and split it into a number of subperiods, say K in total. Calendar time divided by the number of subperiods, K, is the time increment, Δt.
3. Specify distributional assumptions for the risk factors that drive the underlying variables. For example, stock price is the underlying variable for the Asian call, so we need a model for stock price movement. Say we choose the following model for changes in stock price, where Zk stands for the standard normal random variable:
In the way that we are using the term, Zk is a risk factor in the simulation. Through our choice of μ and σ, we control the distribution of stock price. Although this example has one risk factor, a given simulation may have multiple risk factors.
4. Using a computer program or spreadsheet function, draw K random values of each risk factor. In our example, the spreadsheet function would produce a draw of K values of the standard normal variable Zk: Z1, Z2, Z3, …, ZK.
5. Calculate the underlying variables using the random observations generated in Step 4. Using the above model of stock price dynamics, the result is K observations on changes in stock price. An additional calculation is needed to convert those changes into K stock prices (using initial stock price, which is given). Another calculation produces the average stock price during the life of the option (the sum of K stock prices divided by K).
6. Compute the quantities of interest. In our example, the first calculation is the value of an Asian call at maturity, CiT. A second calculation discounts this terminal value back to the present to get the call value as of today, Ci0. We have completed one simulation trial. (The subscript i in Ci0 stands for the ith simulation trial, as it does in CiT.) In a Monte Carlo simulation, a running tabulation is kept of statistics relating to the distribution of the quantities of interest, including their mean value and standard deviation, over the simulation trials to that point.
7. Iteratively go back to Step 4 until a specified number of trials, I, is completed. Finally, produce statistics for the simulation. The key value for our example is the mean value of Ci0 for the total number of simulation trials. This mean value is the Monte Carlo estimate of the value of the Asian call.
How many simulation trials should be specified? In general, we need to increase the number of trials by a factor of 100 to get each extra digit of accuracy. Depending on the problem, tens of thousands of trials may be needed to obtain accuracy to two decimal places (as required for option value, for example). Conducting a large number of trials is not necessarily a problem, given today’s computing power. The number of trials needed can be reduced using variance reduction procedures, a topic outside the scope of this reading.45
In Step 4 of our example, a computer function produced a set of random observations on a standard normal random variable. Recall that for a uniform distribution, all possible numbers are equally likely. The term random number
generator refers to an algorithm that produces uniformly distributed random numbers between 0 and 1. In the context of computer simulations, the term random number refers to an observation drawn from a uniform distribution.46 For other distributions, the term “random observation” is used in this context.
It is a remarkable fact that random observations from any distribution can be produced using the uniform random variable with endpoints 0 and 1. To see why this is so, consider the inverse transformation method of producing random observations. Suppose we are interested in obtaining random observations for a random variable, X, with cumulative distribution function F(x). Recall that F(x) evaluated at x is a number between 0 and 1. Suppose a random outcome of this random variable is 3.21 and that F(3.21) = 0.25 or 25 percent. Define an inverse of F, call it F −1, that can do the following: Substitute the probability 0.25 into F −1 and it returns the random outcome 3.21. In other words, F −1(0.25) = 3.21. To generate random observations on X, the steps are 1) generate a uniform random number, r, between 0 and 1 using the random number generator and 2) evaluate F −1(r) to obtain a random observation on X. Random observation generation is a field of study in itself, and we have briefly discussed the inverse transformation method here just to illustrate a point. As a generalist you do not need to address the technical details of converting random numbers into random observations, but you do need to know that random observations from any distribution can be generated using a uniform random variable.
In Examples 11 and 12, we give an application of Monte Carlo simulation to a question of great interest to investment practice: the potential gains from market timing.
EXAMPLE 11 Potential Gains from Market Timing: A Monte Carlo Simulation (1)
All active investors want to achieve superior performance. One possible source of superior performance is market timing ability. How accurate does an investor need to be as a bull-and bear-market forecaster for market timing to be profitable? What size gains compared with a buy-and-hold strategy accrue to a given level of accuracy? Because of the variability in asset returns, a huge amount of return data is needed to find statistically reliable answers to these questions. Chua, Woodward, and To (1987) thus selected Monte Carlo simulation to address the potential gains from market timing. They were interested in the perspective of a Canadian investor.
To understand their study, suppose that at the beginning of a year, an investor
predicts that the next year will see either a bull market or bear market. If the prediction is bull market, the investor puts all her money in stocks and earns the market return for that year. On the other hand, if the prediction is bear market, the investor holds T-bills and earns the T-bill return. After the fact, a market is categorized as bull market if the stock market return, RMt, minus T- bill return, RFt, is positive for the year; otherwise, the market is classed as bear market. The investment results of a market timer can be compared with those of a buy-and-hold investor. A buy-and-hold investor earns the market return every year. For Chua et al., one quantity of interest was the gain from market timing. They defined this quantity as the market timer ’s average return minus the average return to a buy-and-hold investor.
To simulate market returns, Chua et al. generated 10,000 random standard normal observations, Zt. At the time of the study, Canadian stocks had a historical mean annual return of 12.95 percent with a standard deviation of 18.30 percent. To reflect these parameters, the simulated market returns are RMt = 0.1830Zt + 0.1295, t = 1, 2, …, 10,000. Using a second set of 10,000 random standard normal observations, historical return parameters for Canadian T- bills, as well as the historical correlation of T-bill and stock returns, the authors generated 10,000 T-bill returns.
An investor can have different skills in forecasting bull and bear markets. Chua et al. characterized market timers by accuracy in forecasting bull markets and accuracy in forecasting bear markets. For example, bull market forecasting accuracy of 50 percent means that when the timer forecasts bull market for the next year, she is right just half the time, indicating no skill. Suppose an investor has 60 percent accuracy in forecasting bull market and 80 percent accuracy in forecasting bear market (a 60–80 timer). We can simulate how an investor would fare. After generating the first observation on RMt − RFt, we know whether that observation is a bull or bear market. If the observation is bull market, then 0.60 (forecast accuracy for bull markets) is compared with a random number (between 0 and 1). If the random number is less than 0.60, which occurs with a 60 percent probability, then the market timer is assumed to have correctly predicted bull market and her return for that first observation is the market return. If the random number is greater than 0.60, then the market timer is assumed to have made an error and predicted bear market; her return for that observation is the risk-free rate. In a similar fashion, if that first observation is bear market, the timer has an 80 percent chance of being right in forecasting bear market based on a random number draw. In either case, her return is compared with the market return to record her gain versus a buy-and- hold strategy. That process is one simulation trial. The simulated mean return
earned by the timer is the average return earned by the timer over all trials in the simulation.
To increase our understanding of the process, consider a hypothetical Monte Carlo simulation with four trials for the 60–80 timer (who, to reiterate, has 60 percent accuracy in forecasting bull markets and 80 percent accuracy in forecasting bear markets). Table 8 gives data for the simulation. Let us look at Trials 1 and 2. In Trial 1, the first random number drawn leads to a market return of 0.121. Because the market return, 0.121, exceeded the T-bill return, 0.050, we have a bull market. We generate a random number, 0.531, which we then compare with the timer ’s bull market accuracy, 0.60. Because 0.531 is less than 0.60, the timer is assumed to have made a correct bull market forecast and thus to have invested in stocks. Thus the timer earns the stock market return, 0.121, for that trial. In the second trial we observe another bull market, but because the random number 0.725 is greater than 0.60, the timer is assumed to have made an error and predicted a bear market; therefore, the timer earned the T-bill return, 0.081, rather than higher stock market return.
TABLE 8 Hypothetical Simulation for a 60–80 Market Timer
After Draws for Zt andfor the T-bill Return Simulation Results
Trial RMt RFt Bull or BearMarket? Value of X
Timer’s Prediction Correct?
Return Earned by Timer
1 0.121 0.050 Bull 0.531 Yes 0.121 2 0.092 0.081 Bull 0.725 No 0.081 3 −0.020 0.034 Bear 0.786 Yes 0.034 4 0.052 0.055 A 0.901 B C
Note: is the mean return earned by the timer over the four simulation trials.
Using the data in Table 8, determine the values of A, B, C, and D.
Solution: The value of A is Bear because the stock market return was less than the T-bill return in Trial 4. The value of B is No. Because we observe a bear market, we compare the random number 0.901 with 0.80, the timer ’s bear- market forecasting accuracy. Because 0.901 is greater than 0.8, the timer is assumed to have made an error. The value of C is 0.052, the return on the stock market, because the timer made an error and invested in the stock market and
earned 0.052 rather than the higher T-bill return of 0.055. The value of D is = (0.121 + 0.081 + 0.034 + 0.052) = 0.288/4 = 0.072. Note that we could calculate other statistics besides the mean, such as the standard deviation of the returns earned by the timer over the four trials in the simulation.
EXAMPLE 12 Potential Gains from Market Timing: A Monte Carlo Simulation (2)
Having discussed the plan of the Chua et al. study and illustrated the method for a hypothetical Monte Carlo simulation with four trials, we conclude our presentation of the study.
The hypothetical simulation in Example 11 had four trials, far too few to reach statistically precise conclusions. The simulation of Chua et al. incorporated 10,000 trials. Chua et al. specified bull-and bear-market prediction skill levels of 50, 60, 70, 80, 90, and 100 percent. Table 9 presents a very small excerpt from their simulation results for the no transaction costs case (transaction costs were also examined). Reading across the row, the timer with 60 percent bull market and 80 percent bear market forecasting accuracy had a mean annual gain from market timing of −1.12 percent per year. On average, the buy-and- hold investor out-earned this skillful timer by 1.12 percentage points. There was substantial variability in gains across the simulation trials, however: The standard deviation of the gain was 14.77 percent, so in many trials (but not on average) the gain was positive. Row 3 (win/loss) is the ratio of profitable switches between stocks and T-bills to unprofitable switches. This ratio was a favorable 1.2070 for the 60–80 timer. (When transaction costs were considered, however, fewer switches are profitable: The win–loss ratio was 0.5832 for the 60–80 timer.)
The authors concluded that the cost of not being invested in the market during bull market years is high. Because a buy-and-hold investor never misses a bull market year, she has 100 percent forecast accuracy for bull markets (at the cost of 0 percent accuracy for bear markets). Given their definitions and assumptions, the authors also concluded that successful market timing requires a minimum accuracy of 80 percent in forecasting both bull and bear markets. Market timing is a continuing area of interest and study, and other perspectives exist. However, this example illustrates how Monte Carlo simulation is used to address important investment issues.
TABLE 9 Gains from Stock Market Timing (No Transaction Costs)
Bull MarketAccuracy (%) Bear Market Accuracy (%) 50 60 70 80 90 100
60 Mean (%) −2.50 −1.99 −1.57 −1.12 −0.68 −0.22 S.D. (%) 13.65 14.11 14.45 14.77 15.08 15.42 Win/Loss 0.7418 0.9062 1.0503 1.2070 1.3496 1.4986
Source: Chua, Woodward, and To (1987), Table II (excerpt).
The analyst chooses the probability distributions in Monte Carlo simulation. By contrast, historical simulation samples from a historical record of returns (or other underlying variables) to simulate a process. The concept underlying historical simulation (also called back simulation) is that the historical record provides the most direct evidence on distributions (and that the past applies to the future). For example, refer back to Step 2 in the outline of Monte Carlo simulation above and suppose the time increment is one day. Further, suppose we base the simulation on the record of daily stock returns over the last five years. In one type of historical simulation, we randomly draw K returns from that record to generate one simulation trial. We put back the observations into the sample, and in the next trial we again randomly sample with replacement. The simulation results directly reflect frequencies in the data. A drawback of this approach is that any risk not represented in the time period selected (for example, a stock market crash) will not be reflected in the simulation. Compared with Monte Carlo simulation, historical simulation does not lend itself to “what if” analyses. Nevertheless, historic simulation is an established alternative simulation methodology.
Monte Carlo simulation is a complement to analytical methods. It provides only statistical estimates, not exact results. Analytical methods, where available, provide more insight into cause-and-effect relationships. For example, the Black–Scholes– Merton option pricing model for the value of a European call option is an analytical method, expressed as a formula. It is a much more efficient method for valuing such a call than is Monte Carlo simulation. As an analytical expression, the Black– Scholes–Merton model permits the analyst to quickly gauge the sensitivity of call value to changes in current stock price and the other variables that determine call value. In contrast, Monte Carlo simulations do not directly provide such precise insights. However, only some types of options can be priced with analytical expressions. As financial product innovations proceed, the field of applications for Monte Carlo simulation continues to grow.
5. Summary In this reading, we have presented the most frequently used probability distributions in investment analysis and the Monte Carlo simulation.
A probability distribution specifies the probabilities of the possible outcomes of a random variable.
The two basic types of random variables are discrete random variables and continuous random variables. Discrete random variables take on at most a countable number of possible outcomes that we can list as x1, x2, … In contrast, we cannot describe the possible outcomes of a continuous random variable Z with a list z1, z2, … because the outcome (z1 + z2)/2, not in the list, would always be possible.
The probability function specifies the probability that the random variable will take on a specific value. The probability function is denoted p(x) for a discrete random variable and f (x) for a continuous random variable. For any probability function p(x), 0 ≤ p(x) ≤ 1, and the sum of p(x) over all values of X equals 1.
The cumulative distribution function, denoted F(x) for both continuous and discrete random variables, gives the probability that the random variable is less than or equal to x.
The discrete uniform and the continuous uniform distributions are the distributions of equally likely outcomes.
The binomial random variable is defined as the number of successes in n Bernoulli trials, where the probability of success, p, is constant for all trials and the trials are independent. A Bernoulli trial is an experiment with two outcomes, which can represent success or failure, an up move or a down move, or another binary (twofold) outcome.
A binomial random variable has an expected value or mean equal to np and variance equal to np(1 − p).
A binomial tree is the graphical representation of a model of asset price dynamics in which, at each period, the asset moves up with probability p or down with probability (1 − p). The binomial tree is a flexible method for modeling asset price movement and is widely used in pricing options.
The normal distribution is a continuous symmetric probability distribution that
is completely described by two parameters: its mean, μ, and its variance, σ2.
A univariate distribution specifies the probabilities for a single random variable. A multivariate distribution specifies the probabilities for a group of related random variables.
To specify the normal distribution for a portfolio when its component securities are normally distributed, we need the means, standard deviations, and all the distinct pairwise correlations of the securities. When we have those statistics, we have also specified a multivariate normal distribution for the securities.
For a normal random variable, approximately 68 percent of all possible outcomes are within a one standard deviation interval about the mean, approximately 95 percent are within a two standard deviation interval about the mean, and approximately 99 percent are within a three standard deviation interval about the mean.
A normal random variable, X, is standardized using the expression Z = (X − μ)/ σ, where μ and σ are the mean and standard deviation of X. Generally, we use the sample mean as an estimate of μ and the sample standard deviation s as an estimate of σ in this expression.
The standard normal random variable, denoted Z, has a mean equal to 0 and variance equal to 1. All questions about any normal random variable can be answered by referring to the cumulative distribution function of a standard normal random variable, denoted N(x) or N(z).
Shortfall risk is the risk that portfolio value will fall below some minimum acceptable level over some time horizon.
Roy’s safety-first criterion, addressing shortfall risk, asserts that the optimal portfolio is the one that minimizes the probability that portfolio return falls below a threshold level. According to Roy’s safety-first criterion, if returns are normally distributed, the safety-first optimal portfolio P is the one that maximizes the quantity [E(RP) − RL]/σP, where RL is the minimum acceptable level of return.
A random variable follows a lognormal distribution if the natural logarithm of the random variable is normally distributed. The lognormal distribution is defined in terms of the mean and variance of its associated normal distribution. The lognormal distribution is bounded below by 0 and skewed to the right (it has a long right tail).
The lognormal distribution is frequently used to model the probability
distribution of asset prices because it is bounded below by zero.
Continuous compounding views time as essentially continuous or unbroken; discrete compounding views time as advancing in discrete finite intervals.
The continuously compounded return associated with a holding period is the natural log of 1 plus the holding period return, or equivalently, the natural log of ending price over beginning price.
If continuously compounded returns are normally distributed, asset prices are lognormally distributed. This relationship is used to move back and forth between the distributions for return and price. Because of the central limit theorem, continuously compounded returns need not be normally distributed for asset prices to be reasonably well described by a lognormal distribution.
Monte Carlo simulation involves the use of a computer to represent the operation of a complex financial system. A characteristic feature of Monte Carlo simulation is the generation of a large number of random samples from specified probability distribution(s) to represent the operation of risk in the system. Monte Carlo simulation is used in planning, in financial risk management, and in valuing complex securities. Monte Carlo simulation is a complement to analytical methods but provides only statistical estimates, not exact results.
Historical simulation is an established alternative to Monte Carlo simulation that in one implementation involves repeated sampling from a historical data series. Historical simulation is grounded in actual data but can reflect only risks represented in the sample historical data. Compared with Monte Carlo simulation, historical simulation does not lend itself to “what if” analyses.
References
1. Campbell, John, Andrew Lo, and A. Craig MacKinlay. 1997. The Econometrics of Financial Markets. Princeton, NJ: Princeton University Press.
2. Chance, Don M., and Robert Brooks. 2012. An Introduction to Derivatives and Risk Management, 9th ed. Mason, OH: South-Western.
3. Chua, Jess H., Richard S. Woodward, and Eric C. To. 1987. “Potential Gains from Stock Market Timing in Canada.” Financial Analysts Journal, vol. 43, no. 5: 50–56.
4. Cox, Jonathan, Stephen Ross, and Mark Rubinstein. 1979. “Options Pricing: A Simplified Approach.” Journal of Financial Economics, vol. 7: 229–263.
5. Fama, Eugene. 1976. Foundations of Finance. New York: Basic Books.
6. Ferguson, Robert. 1993. “Some Formulas for Evaluating Two Popular Option Strategies.” Financial Analysts Journal, vol. 49, no. 5: 71–76.
7. Hillier, Frederick S., and Gerald J. Lieberman. 2010. Introduction to Operations Research, 9th edition. New York: McGraw-Hill.
8. Hull, John. 2011. Options, Futures, and Other Derivatives, 8th edition. Upper Saddle River, NJ: Pearson.
9. Kolb, Robert W., Gerald D. Gay, and William C. Hunter. 1985. “Liquidity Requirements for Financial Futures Investments.” Financial Analysts Journal, vol. 41, no. 3: 60–68.
10. Kon, Stanley J. 1984. “Models of Stock Returns—A Comparison.” Journal of Finance, vol. 39: 147–165.
11. Leibowitz, Martin, and Roy Henriksson. 1989. “Portfolio Optimization with Shortfall Constraints: A Confidence-Limit Approach to Managing Downside Risk.” Financial Analysts Journal, vol. 45, no. 2: 34–41.
12. Liang, Bing. 1999. “On the Performance of Hedge Funds.” Financial Analysts Journal, vol. 55, No. 4: 72–85.
13. Luenberger, David G. 1998. Investment Science. New York: Oxford University Press.
14. Roy, A.D. 1952. “Safety-First and the Holding of Assets.” Econometrica, vol. 20: 431–439.
15. Thomas, Rawley, and Benton E. Gup. 2010. The Valuation Handbook: Valuation Techniques from Today’s Top Practitioners. Hoboken, NJ: Wiley.
Problems Practice Problems and Solutions: 1–18 taken from Quantitative Methods for Investment Analysis, Second Edition, by Richard A. DeFusco, CFA, Dennis W. McLeavey, CFA, Jerald E. Pinto, CFA, and David E. Runkle, CFA. Copyright © 2004 by CFA Institute. All other problems and solutions copyright © CFA Institute. 1. A European put option on stock conveys the right to sell the stock at a
prespecified price, called the exercise price, at the maturity date of the option. The value of this put at maturity is (exercise price – stock price) or $0, whichever is greater. Suppose the exercise price is $100 and the underlying stock trades in ticks of $0.01. At any time before maturity, the terminal value of the put is a random variable.
1. Describe the distinct possible outcomes for terminal put value. (Think of the put’s maximum and minimum values and its minimum price increments.)
2. Is terminal put value, at a time before maturity, a discrete or continuous random variable?
3. Letting Y stand for terminal put value, express in standard notation the probability that terminal put value is less than or equal to $24. No calculations or formulas are necessary.
2. Suppose X, Y, and Z are discrete random variables with these sets of possible outcomes: X = {2, 2.5, 3}, Y = {0, 1, 2, 3}, and Z = {10, 11, 12}. For each of the functions f (X), g(Y), and h(Z), state whether the function satisfies the conditions for a probability function.
1. f (2) = −0.01 f (2.5) = −0.50 f (3) = −0.51
2. g(0) = 0.25 g(1) = 0.50 g(2) = 0.125 g(3) = 0.125
3. h(10) = 0.35 h(11) = 0.15 h(12) = 0.52
3. Define the term “binomial random variable.” Describe the types of problems for which the binomial distribution is used.
4. Over the last 10 years, a company’s annual earnings increased year over year seven times and decreased year over year three times. You decide to model the number of earnings increases for the next decade as a binomial random variable.
1. What is your estimate of the probability of success, defined as an increase
in annual earnings?
2. For Parts B, C, and D of this problem, assume the estimated probability is the actual probability for the next decade.
3. What is the probability that earnings will increase in exactly 5 of the next 10 years?
4. Calculate the expected number of yearly earnings increases during the next 10 years.
5. Calculate the variance and standard deviation of the number of yearly earnings increases during the next 10 years.
6. The expression for the probability function of a binomial random variable depends on two major assumptions. In the context of this problem, what must you assume about annual earnings increases to apply the binomial distribution in Part B? What reservations might you have about the validity of these assumptions?
5. You are examining the record of an investment newsletter writer who claims a 70 percent success rate in making investment recommendations that are profitable over a one-year time horizon. You have the one-year record of the newsletter ’s seven most recent recommendations. Four of those recommendations were profitable. If all the recommendations are independent and the newsletter writer ’s skill is as claimed, what is the probability of observing four or fewer profitable recommendations out of seven in total?
6. By definition, a down-and-out call option on stock becomes worthless and terminates if the price of the underlying stock moves down and touches a prespecified point during the life of the call. If the prespecified level is $75, for example, the call expires worthless if and when the stock price falls to $75. Describe, without a diagram, how a binomial tree can be used to value a down- and-out call option.
7. You are forecasting sales for a company in the fourth quarter of its fiscal year. Your low-end estimate of sales is €14 million, and your high-end estimate is €15 million. You decide to treat all outcomes for sales between these two values as equally likely, using a continuous uniform distribution.
1. What is the expected value of sales for the fourth quarter?
2. What is the probability that fourth-quarter sales will be less than or equal to €14,125,000?
8. State the approximate probability that a normal random variable will fall within
the following intervals:
1. Mean plus or minus one standard deviation.
2. Mean plus or minus two standard deviations.
3. Mean plus or minus three standard deviations.
9. Find the area under the normal curve up to z = 0.36; that is, find P(Z ≤ 0.36). Interpret this value.
10. In futures markets, profits or losses on contracts are settled at the end of each trading day. This procedure is called marking to market or daily resettlement. By preventing a trader ’s losses from accumulating over many days, marking to market reduces the risk that traders will default on their obligations. A futures markets trader needs a liquidity pool to meet the daily mark to market. If liquidity is exhausted, the trader may be forced to unwind his position at an unfavorable time.
Suppose you are using financial futures contracts to hedge a risk in your portfolio. You have a liquidity pool (cash and cash equivalents) of λ dollars per contract and a time horizon of T trading days. For a given size liquidity pool, λ, Kolb, Gay, and Hunter (1985) developed an expression for the probability stating that you will exhaust your liquidity pool within a T-day horizon as a result of the daily mark to market. Kolb et al. assumed that the expected change in futures price is 0 and that futures price changes are normally distributed. With σ representing the standard deviation of daily futures price changes, the standard deviation of price changes over a time horizon to day T is , given continuous compounding. With that background, the Kolb et al. expression is
11. where . Here x is a standardized value of λ. N(x) is the standard normal cumulative distribution function. For some intuition about 1 − N(x) in the expression, note that the liquidity pool is exhausted if losses exceed the size of the liquidity pool at any time up to and including T; the probability of that event happening can be shown to be proportional to an area in the right tail of a standard normal distribution, 1 − N(x).
12. Using the Kolb et al. expression, answer the following questions:
1. Your hedging horizon is five days, and your liquidity pool is $2,000 per contract. You estimate that the standard deviation of daily price changes for the contract is $450. What is the probability that you will exhaust your
liquidity pool in the five-day period?
2. Suppose your hedging horizon is 20 days, but all the other facts given in Part A remain the same. What is the probability that you will exhaust your liquidity pool in the 20-day period?
The following information relates to Questions 11–13
As reported by Liang (1999), US equity funds in three style categories had the following mean monthly returns, standard deviations of return, and Sharpe ratios during the period January 1994 to December 1996:
January 1994 to December 1996 Strategy Mean Return (%) Standard Deviation (%) Sharpe Ratio
Large-cap growth 1.15 2.89 0.26 Large-cap value 1.08 2.20 0.31 Large-cap blend 1.07 2.38 0.28
Source: Liang (1999), Table 5 (excerpt).
13. Basing your estimate of future-period monthly return parameters on the sample mean and standard deviation for the period January 1994 to December 1996, construct a 90 percent confidence interval for the monthly return on a large-cap blend fund. Assume fund returns are normally distributed.
14. Basing your estimate of future-period monthly return parameters on the sample mean and standard deviation for the period January 1994 to December 1996, calculate the probability that a large-cap growth fund will earn a monthly return of 0 percent or less. Assume fund returns are normally distributed.
15. Assuming fund returns are normally distributed, which fund category minimized the probability of earning less than the risk-free rate for the period January 1994 to December 1996?
16. A client has a portfolio of common stocks and fixed-income instruments with a current value of £1,350,000. She intends to liquidate £50,000 from the portfolio at the end of the year to purchase a partnership share in a business. Furthermore, the client would like to be able to withdraw the £50,000 without reducing the initial capital of £1,350,000. The following table shows four alternative asset allocations.
Mean and Standard Deviation for Four Allocations (in Percent) A B C D
Expected annual return 16 12 10 9
Standard deviation of return 24 17 12 11
Address the following questions (assume normality for Parts B and C): 1. Given the client’s desire not to invade the £1,350,000 principal, what is the
shortfall level, RL? Use this shortfall level to answer Part B.
2. According to the safety-first criterion, which of the allocations is the best?
3. What is the probability that the return on the safety-first optimal portfolio will be less than the shortfall level, RL?
17. A. Describe two important characteristics of the lognormal distribution.
B. Compared with the normal distribution, why is the lognormal distribution a more reasonable model for the distribution of asset prices?
C. What are the two parameters of a lognormal distribution?
18. The basic calculation for volatility (denoted σ) as used in option pricing is the annualized standard deviation of continuously compounded daily returns. Calculate volatility for Dollar General Corporation (NYSE: DG) based on its closing prices for two weeks, given in the table below. (Annualize based on 250 days in a year.)
Dollar General Corporation Daily Closing Stock Price Date Closing Price ($)
27 January 2003 10.68 28 January 2003 10.87 29 January 2003 11.00 30 January 2003 10.95 31 January 2003 11.26 3 February 2003 11.31 4 February 2003 11.23 5 February 2003 10.91 6 February 2003 10.80 7 February 2003 10.47
Source: http://finance.yahoo.com.
19. A. Define Monte Carlo simulation and explain its use in finance.
B. Compared with analytical methods, what are the strengths and
weaknesses of Monte Carlo simulation for use in valuing securities?
20. A standard lookback call option on stock has a value at maturity equal to (Value of the stock at maturity – Minimum value of stock during the life of the option prior to maturity) or $0, whichever is greater. If the minimum value reached prior to maturity was $20.11 and the value of the stock at maturity is $23, for example, the call is worth $23 − $20.11 = $2.89. Briefly discuss how you might use Monte Carlo simulation in valuing a lookback call option.
21. At the end of the current year, an investor wants to make a donation of $20,000 to charity but does not want the year-end market value of her portfolio to fall below $600,000. If the shortfall level is equal to the risk-free rate of return and returns from all portfolios considered are normally distributed, will the portfolio that minimizes the probability of failing to achieve the investor ’s objective most likely have the:
highest safety-first ratio? highest Sharpe ratio? A. No Yes B. Yes No C. Yes Yes
22. An analyst stated that normal distributions are suitable for describing asset returns and that lognormal distributions are suitable for describing distributions of asset prices. The analyst’s statement is correct in regard to:
1. both normal distributions and lognormal distributions.
2. normal distributions, but incorrect in regard to lognormal distributions.
3. lognormal distributions, but incorrect in regard to normal distributions.
Notes 1 We follow the convention that an uppercase letter represents a random variable
and a lowercase letter represents an outcome or specific value of the random variable. Thus X refers to the random variable, and x refers to an outcome of X. We subscript outcomes, as in x1 and x2, when we need to distinguish among different outcomes in a list of outcomes of a random variable.
2 The technical term for the probability function of a discrete random variable, probability mass function (pmf), is used less frequently.
3 See Hillier and Lieberman (2010). Random numbers initially generated by computers are usually random positive integer numbers that are converted to approximate continuous uniform random numbers between 0 and 1. Then the continuous uniform random numbers are used to produce random observations on other distributions, such as the normal, using various techniques. We will discuss random observation generation further in the section on Monte Carlo simulation.
4 The “hat” over p indicates that it is an estimate of p, the underlying probability of a profitable trade with the broker.
5 Of course, you need to adjust for the direction of the overall market after the trade (any broker ’s record will be helped by a bull market) and perhaps make other risk adjustments. Assume that these adjustments have been made.
6 In this example all calculations were worked through by hand, but binomial probability and cdf functions are also available in computer spreadsheet programs.
7 A basis point is one-hundredth of 1 percent (0.01 percent).
8 Some practitioners use tracking error to describe what we later call tracking risk, the standard deviation of the differences between the portfolio’s and benchmark’s returns.
9 The mean (or arithmetic mean) is the sum of all values in a distribution or dataset, divided by the number of values summed. The variance is a measure of dispersion about the mean. See the reading on statistical concepts and market returns for further details on these concepts.
10 We can show that p(1 − p) is the variance of a Bernoulli random variable as follows, noting that a Bernoulli random variable can take on only one of two values, 1 or 0: σ2(Y) = E[(Y − EY)2] = E[(Y − p)2] = (1 − p)2p + (0 − p)2(1 − p) = (1 − p)[(1 − p)p + p2] = p(1 − p).
11 For example, we can split 20 days into 100 subperiods, taking care to use compatible values for u and d.
12 Cash dividends represent a reduction of a company’s assets. Early exercise may be optimal because the exercise price of options is typically not reduced by the amount of cash dividends, so cash dividends negatively affect the position of an American call option holder.
13 See Chance and Brooks (2012) for more information on option pricing models.
14 For a detailed discussion on the use and misuse of EBITDA, see Chapter 20, EBITDA, in Thomas and Gup (2010).
15 Chebyshev’s inequality is discussed in the reading on statistical concepts and market returns.
16 The central limit theorem is discussed further in the reading on sampling.
17 If we have a sample of size n from a normal distribution, we may want to know the possible variation in sample skewness and kurtosis. For a normal random variable, the standard deviation of sample skewness is 6/n and the standard deviation of sample kurtosis is 24/n.
18 For example, a distribution with two stocks (a bivariate normal distribution) has two means, two variances, and one correlation: 2(2 − 1)/2. A distribution with 30 stocks has 30 means, 30 variances, and 435 distinct correlations: 30(30 − 1)/2. The return correlation of Dow Chemical with American Express stock is the same as the correlation of American Express with Dow Chemical stock, so these are counted as one distinct correlation.
19 See Fama (1976) and Campbell, Lo, and MacKinlay (1997).
20 Fat tails can be modeled by a mixture of normal random variables or by a Student’s t-distribution with a relatively small number of degrees of freedom. See Kon (1984) and Campbell, Lo, and MacKinlay (1997). We discuss the Student’s t-distribution in the reading on sampling and estimation.
21 A population is all members of a specified group, and the population mean is the
arithmetic mean computed for the population. A sample is a subset of a population, and the sample mean is the arithmetic mean computed for the sample. For more information on these concepts, see the reading on statistical concepts and market returns.
22 Another often-seen notation for the cdf of a standard normal variable is ϕ(x).
23 Utility functions are mathematical representations of attitudes toward risk and return.
24 We shall discuss mean–variance analysis in detail in the readings on portfolio concepts.
25 A.D. Roy (1952) introduced this criterion.
26 If there is an asset offering a risk-free return over the time horizon being considered, and if RL is less than or equal to that risk-free rate, then it is optimal to be fully invested in the risk-free asset. Holding the risk-free asset in this case eliminates the chance that the threshold return is not met.
27 See Leibowitz and Henriksson (1989), for example.
28 Financial risk is risk relating to asset prices and other financial variables. The contrast is to other, nonfinancial risks (for example, relating to operations and technology), which require different tools to manage.
29 In 95 percent one-day VAR, the 95 percent refers to the confidence in the value of VAR and is equal to 1 − 0.05; this is a traditional way to state VAR.
30 The quantity e ≈ 2.7182818.
31 Note that exp(0.50σ2) > 1 because σ2 > 0.
32 Luenberger (1998) is the source of this explanation.
33 Continuous compounding treats time as essentially continuous or unbroken, in contrast to discrete compounding, which treats time as advancing in discrete finite intervals. Continuously compounded returns are the model for returns in so-called continuous time finance models such as the Black–Scholes–Merton option pricing model. See the reading on the time value of money for more information on compounding.
34 In this reading we use lowercase r to refer specifically to continuously
compounded returns.
35 Stationarity implies that the mean and variance of return do not change from period to period.
36 We mentioned the central limit theorem earlier in our discussion of the normal distribution. To give a somewhat fuller statement of it, according to the central limit theorem the sum (as well as the mean) of a set of independent, identically distributed random variables with finite variances is normally distributed, whatever distribution the random variables follow. We discuss the central limit theorem in the reading on sampling.
37 Volatility is also called the instantaneous standard deviation, and as such is denoted σ. The underlying asset, or simply the underlying, is the asset underlying the option. For more information on these concepts, see Chance and Brooks (2012).
38 To compute the standard deviation of a set or sample of n returns, we sum the squared deviation of each return from the mean return and then divide that sum by n − 1. The result is the sample variance. Taking the square root of the sample variance gives the sample standard deviation. To review the calculation of standard deviation, see the reading on statistical concepts and market returns.
39 Annualizing is often done on the basis of 250 days in a year, the approximate number of days markets are open for trading. The 250-day number may lead to a better estimate of volatility than the 365-day number. Thus if daily volatility were 0.01, we would state volatility (on an annual basis) as .
40 The expression for the mean is E(ST) = S0 exp[E(r0,T) + 0.5σ2(r0,T)], for example.
41 See Hull (2011) for a discussion of lognormal confidence intervals.
42 Business Week, 22 January 2001.
43 A European-style option or European option is an option exercisable only at maturity.
44 The steps should be viewed as providing an overview of Monte Carlo simulation rather than as a detailed recipe for implementing a Monte Carlo simulation in its many varied applications.
45 For details on this and other technical aspects of Monte Carlo simulation, see Hillier and Lieberman (2010).
46 The numbers that random number generators produce depend on a seed or initial value. If the same seed is fed to the same generator, it will produce the same sequence. All sequences eventually repeat. Because of this predictability, the technically correct name for the numbers produced by random number generators is pseudo-random numbers. Pseudo-random numbers have sufficient qualities of randomness for most practical purposes.
CHAPTER 6 Sampling and Estimation Richard A. DeFusco, CFA
Dennis W. McLeavey, CFA
Jerald E. Pinto, PhD, CFA
David E. Runkle, PhD, CFA
LEARNING OUTCOMES
After completing this chapter, you will be able to do the following:
define simple random sampling and a sampling distribution;
explain sampling error;
distinguish between simple random and stratified random sampling;
distinguish between time-series and cross-sectional data;
explain the central limit theorem and its importance;
calculate and interpret the standard error of the sample mean;
identify and describe desirable properties of an estimator;
distinguish between a point estimate and a confidence interval estimate of a population parameter;
describe properties of Student’s t-distribution and calculate and interpret its degrees of freedom;
calculate and interpret a confidence interval for a population mean, given a normal distribution with 1) a known population variance, 2) an unknown population variance, or 3) an unknown variance and a large sample size;
describe the issues regarding selection of the appropriate sample size, data- mining bias, sample selection bias, survivorship bias, look-ahead bias, and time-period bias.
1. Introduction Each day, we observe the high, low, and close of stock market indexes from around the world. Indexes such as the S&P 500 Index and the Nikkei-Dow Jones Average are samples of stocks. Although the S&P 500 and the Nikkei do not represent the populations of US or Japanese stocks, we view them as valid indicators of the whole population’s behavior. As analysts, we are accustomed to using this sample information to assess how various markets from around the world are performing. Any statistics that we compute with sample information, however, are only estimates of the underlying population parameters. A sample, then, is a subset of the population—a subset studied to infer conclusions about the population itself.
This reading explores how we sample and use sample information to estimate population parameters. In the next section, we discuss sampling—the process of obtaining a sample. In investments, we continually make use of the mean as a measure of central tendency of random variables, such as return and earnings per share. Even when the probability distribution of the random variable is unknown, we can make probability statements about the population mean using the central limit theorem. In Section 3, we discuss and illustrate this key result. Following that discussion, we turn to statistical estimation. Estimation seeks precise answers to the question “What is this parameter ’s value?”
The central limit theorem and estimation are the core of the body of methods presented in this reading. In investments, we apply these and other statistical techniques to financial data; we often interpret the results for the purpose of deciding what works and what does not work in investments. We end this reading with a discussion of the interpretation of statistical results based on financial data and the possible pitfalls in this process.
2. Sampling In this section, we present the various methods for obtaining information on a population (all members of a specified group) through samples (part of the population). The information on a population that we try to obtain usually concerns the value of a parameter, a quantity computed from or used to describe a population of data. When we use a sample to estimate a parameter, we make use of sample statistics (statistics, for short). A statistic is a quantity computed from or used to describe a sample of data.
We take samples for one of two reasons. In some cases, we cannot possibly examine every member of the population. In other cases, examining every member of the population would not be economically efficient. Thus, savings of time and money are two primary factors that cause an analyst to use sampling to answer a question about a population. In this section, we discuss two methods of random sampling: simple random sampling and stratified random sampling. We then define and illustrate the two types of data an analyst uses: cross-sectional data and time-series data.
2.1. Simple Random Sampling Suppose a telecommunications equipment analyst wants to know how much major customers will spend on average for equipment during the coming year. One strategy is to survey the population of telecom equipment customers and inquire what their purchasing plans are. In statistical terms, the characteristics of the population of customers’ planned expenditures would then usually be expressed by descriptive measures such as the mean and variance. Surveying all companies, however, would be very costly in terms of time and money.
Alternatively, the analyst can collect a representative sample of companies and survey them about upcoming telecom equipment expenditures. In this case, the analyst will compute the sample mean expenditure, , a statistic. This strategy has a substantial advantage over polling the whole population because it can be accomplished more quickly and at lower cost.
Sampling, however, introduces error. The error arises because not all the companies in the population are surveyed. The analyst who decides to sample is trading time and money for sampling error.
When an analyst chooses to sample, he must formulate a sampling plan. A sampling plan is the set of rules used to select a sample. The basic type of sample from which we can draw statistically sound conclusions about a population is the simple random
sample (random sample, for short).
Definition of Simple Random Sample. A simple random sample is a subset of a larger population created in such a way that each element of the population has an equal probability of being selected to the subset.
The procedure of drawing a sample to satisfy the definition of a simple random sample is called simple random sampling. How is simple random sampling carried out? We need a method that ensures randomness—the lack of any pattern—in the selection of the sample. For a finite (limited) population, the most common method for obtaining a random sample involves the use of random numbers (numbers with assured properties of randomness). First, we number the members of the population in sequence. For example, if the population contains 500 members, we number them in sequence with three digits, starting with 001 and ending with 500. Suppose we want a simple random sample of size 50. In that case, using a computer random- number generator or a table of random numbers, we generate a series of three-digit random numbers. We then match these random numbers with the number codes of the population members until we have selected a sample of size 50.
Sometimes we cannot code (or even identify) all the members of a population. We often use systematic sampling in such cases. With systematic sampling, we select every kth member until we have a sample of the desired size. The sample that results from this procedure should be approximately random. Real sampling situations may require that we take an approximately random sample.
Suppose the telecommunications equipment analyst polls a random sample of telecom equipment customers to determine the average equipment expenditure. The sample mean will provide the analyst with an estimate of the population mean expenditure. Any difference between the sample mean and the population mean is called sampling error.
Definition of Sampling Error. Sampling error is the difference between the observed value of a statistic and the quantity it is intended to estimate.
A random sample reflects the properties of the population in an unbiased way, and sample statistics, such as the sample mean, computed on the basis of a random sample are valid estimates of the underlying population parameters.
A sample statistic is a random variable. In other words, not only do the original data from the population have a distribution but so does the sample statistic.
This distribution is the statistic’s sampling distribution.
Definition of Sampling Distribution of a Statistic. The sampling distribution of a statistic is the distribution of all the distinct possible values that the statistic can assume when computed from samples of the same size randomly drawn from the same population.
In the case of the sample mean, for example, we refer to the “sampling distribution of the sample mean” or the distribution of the sample mean. We will have more to say about sampling distributions later in this reading. Next, however, we look at another sampling method that is useful in investment analysis.
2.2. Stratified Random Sampling The simple random sampling method just discussed may not be the best approach in all situations. One frequently used alternative is stratified random sampling.
Definition of Stratified Random Sampling. In stratified random sampling, the population is divided into subpopulations (strata) based on one or more classification criteria. Simple random samples are then drawn from each stratum in sizes proportional to the relative size of each stratum in the population. These samples are then pooled to form a stratified random sample.
In contrast to simple random sampling, stratified random sampling guarantees that population subdivisions of interest are represented in the sample. Another advantage is that estimates of parameters produced from stratified sampling have greater precision—-that is, smaller variance or dispersion—than estimates obtained from simple random sampling.
Bond indexing is one area in which stratified sampling is frequently applied. Indexing is an investment strategy in which an investor constructs a portfolio to mirror the performance of a specified index. In pure bond indexing, also called the full-replication approach, the investor attempts to fully replicate an index by owning all the bonds in the index in proportion to their market value weights. Many bond indexes consist of thousands of issues, however, so pure bond indexing is difficult to implement. In addition, transaction costs would be high because many bonds do not have liquid markets. Although a simple random sample could be a solution to the cost problem, the sample would probably not match the index’s major risk factors—interest rate sensitivity, for example. Because the major risk factors of fixed-income portfolios are well known and quantifiable, stratified sampling offers a more effective approach. In this approach, we divide the population of index bonds into groups of similar duration (interest rate sensitivity), cash flow
distribution, sector, credit quality, and call exposure. We refer to each group as a stratum or cell (a term frequently used in this context).1 Then, we choose a sample from each stratum proportional to the relative market weighting of the stratum in the index to be replicated.
EXAMPLE 1 Bond Indexes and Stratified Sampling
Suppose you are the manager of a mutual fund indexed to the Lehman Brothers Government Index. You are exploring several approaches to indexing, including a stratified sampling approach. You first distinguish agency bonds from US Treasury bonds. For each of these two groups, you define 10 maturity intervals—1 to 2 years, 2 to 3 years, 3 to 4 years, 4 to 6 years, 6 to 8 years, 8 to 10 years, 10 to 12 years, 12 to 15 years, 15 to 20 years, and 20 to 30 years—and also separate the bonds with coupons (annual interest rates) of 6 percent or less from the bonds with coupons of more than 6 percent.
1. How many cells or strata does this sampling plan entail?
2. If you use this sampling plan, what is the minimum number of issues the indexed portfolio can have?
3. Suppose that in selecting among the securities that qualify for selection within each cell, you apply a criterion concerning the liquidity of the security’s market. Is the sample obtained random? Explain your answer.
Solution to 1: We have 2 issuer classifications, 10 maturity classifications, and 2 coupon classifications. So, in total, this plan entails 2(10)(2) = 40 different strata or cells. (This answer is an application of the multiplication rule of counting discussed in the reading on probability concepts.)
Solution to 2: You cannot have fewer than one issue for each cell, so the portfolio must include at least 40 issues.
Solution to 3: If you apply any additional criteria to the selection of securities for the cells, not every security that might be included has an equal probability of being selected. As a result, the sampling is not random. In practice, indexing using stratified sampling usually does not strictly involve random sampling because the selection of bond issues within cells is subject to various additional criteria. Because the purpose of sampling in this application is not to make an inference about a population parameter but rather to index a portfolio, lack of randomness is not in itself a problem in this application of stratified sampling.
In the next section, we discuss the kinds of data used by financial analysts in sampling and practical issues that arise in selecting samples.
2.3. Time-Series and Cross-Sectional Data Investment analysts commonly work with both time-series and cross-sectional data. A time series is a sequence of returns collected at discrete and equally spaced intervals of time (such as a historical series of monthly stock returns). Cross- sectional data are data on some characteristic of individuals, groups, geographical regions, or companies at a single point in time. The 2014 year-end book value per share for all New York Stock Exchange-listed companies is an example of cross- sectional data.
Economic or financial theory offers no basis for determining whether a long or short time period should be selected to collect a sample. As analysts, we might have to look for subtle clues. For example, combining data from a period of fixed exchange rates with data from a period of floating exchange rates would be inappropriate. The variance of exchange rates when exchange rates were fixed would certainly be less than when rates were allowed to float. As a consequence, we would not be sampling from a population described by a single set of parameters.2 Tight versus loose monetary policy also influences the distribution of returns to stocks; thus, combining data from tight-money and loose-money periods would be inappropriate. Example 2 illustrates the problems that can arise when sampling from more than one distribution.
EXAMPLE 2 Calculating Sharpe Ratios: One or Two Years of Quarterly Data
Analysts often use the Sharpe ratio to evaluate the performance of a managed portfolio. The Sharpe ratio is the average return in excess of the risk-free rate divided by the standard deviation of returns. This ratio measures the excess return earned per unit of standard deviation of return.
To compute the Sharpe ratio, suppose that an analyst collects eight quarterly excess returns (i.e., total return in excess of the risk-free rate). During the first year, the investment manager of the portfolio followed a low-risk strategy, and during the second year, the manager followed a high-risk strategy. For each of these years, the analyst also tracks the quarterly excess returns of some benchmark against which the manager will be evaluated. For each of the two years, the Sharpe ratio for the benchmark is 0.21. Table 1 gives the calculation of the Sharpe ratio of the portfolio.
For the first year, during which the manager followed a low-risk strategy, the average quarterly return in excess of the risk-free rate was 1 percent with a standard deviation of 4.62 percent. The Sharpe ratio is thus 1/4.62 = 0.22. The second year ’s results mirror the first year except for the higher average return and volatility. The Sharpe ratio for the second year is 4/18.48 = 0.22. The Sharpe ratio for the benchmark is 0.21 during the first and second years. Because larger Sharpe ratios are better than smaller ones (providing more return per unit of risk), the manager appears to have outperformed the benchmark.
TABLE 1 Calculation of Sharpe Ratios: Low-Risk and High-Risk Strategies
Quarter/Measure Year 1 Excess Returns Year 2 Excess Returns Quarter 1 −3% −12% Quarter 2 5 20 Quarter 3 −3 −12 Quarter 4 5 20
Quarterly average 1% 4% Quarterly standard deviation 4.62% 18.48%
Sharpe ratio = 0.22 = 1/4.62 = 4/18.48
Now, suppose the analyst believes a larger sample to be superior to a small one. She thus decides to pool the two years together and calculate a Sharpe ratio based on eight quarterly observations. The average quarterly excess return for the two years is the average of each year ’s average excess return. For the two-year period, the average excess return is (1 + 4)/2 = 2.5 percent per quarter. The standard deviation for all eight quarters measured from the sample mean of 2.5 percent is 12.57 percent. The portfolio’s Sharpe ratio for the two- year period is now 2.5/12.57 = 0.199; the Sharpe ratio for the benchmark remains 0.21. Thus, when returns for the two-year period are pooled, the manager appears to have provided less return per unit of risk than the benchmark and less when compared with the separate yearly results.
The problem with using eight quarters of return data is that the analyst has violated the assumption that the sampled returns come from the same population. As a result of the change in the manager ’s investment strategy, returns in Year 2 followed a different distribution than returns in Year 1. Clearly, during Year 1, returns were generated by an underlying population with lower mean and variance than the population of the second year. Combining the results for the first and second years yielded a sample that was
representative of no population. Because the larger sample did not satisfy model assumptions, any conclusions the analyst reached based on the larger sample are incorrect. For this example, she was better off using a smaller sample than a larger sample because the smaller sample represented a more homogeneous distribution of returns.
The second basic type of data is cross-sectional data.3 With cross-sectional data, the observations in the sample represent a characteristic of individuals, groups, geographical regions, or companies at a single point in time. The telecommunications analyst discussed previously is essentially collecting a cross- section of planned capital expenditures for the coming year.
Whenever we sample cross-sectionally, certain assumptions must be met if we wish to summarize the data in a meaningful way. Again, a useful approach is to think of the observation of interest as a random variable that comes from some underlying population with a given mean and variance. As we collect our sample and begin to summarize the data, we must be sure that all the data do, in fact, come from the same underlying population. For example, an analyst might be interested in how efficiently companies use their inventory assets. Some companies, however, turn over their inventory more quickly than others because of differences in their operating environments (e.g., grocery stores turn over inventory more quickly than automobile manufacturers, in general). So the distribution of inventory turnover rates may not be characterized by a single distribution with a given mean and variance. Therefore, summarizing inventory turnover across all companies might be inappropriate. If random variables are generated by different underlying distributions, the sample statistics computed from combined samples are not related to one underlying population parameter. The size of the sampling error in such cases is unknown.
In instances such as these, analysts often summarize company-level data by industry. Attempting to summarize by industry partially addresses the problem of differing underlying distributions, but large corporations are likely to be in more than one industrial sector, so analysts should be sure they understand how companies are assigned to the industry groups.
Whether we deal with time-series data or cross-sectional data, we must be sure to have a random sample that is representative of the population we wish to study. With the objective of inferring information from representative samples, we now turn to the next part of this reading, which focuses on the central limit theorem as well as point and interval estimates of the population mean.
3. Distribution of the Sample Mean Earlier in this reading, we presented a telecommunications equipment analyst who decided to sample in order to estimate mean planned capital expenditures by his customers. Supposing that the sample is representative of the underlying population, how can the analyst assess the sampling error in estimating the population mean? Viewed as a formula that takes a function of the random outcomes of a random variable, the sample mean is itself a random variable with a probability distribution. That probability distribution is called the statistic’s sampling distribution.4 To estimate how closely the sample mean can be expected to match the underlying population mean, the analyst needs to understand the sampling distribution of the mean. Fortunately, we have a result, the central limit theorem, that helps us understand the sampling distribution of the mean for many of the estimation problems we face.
3.1. The Central Limit Theorem One of the most practically useful theorems in probability theory, the central limit theorem has important implications for how we construct confidence intervals and test hypotheses. Formally, it is stated as follows:
The Central Limit Theorem. Given a population described by any probability distribution having mean μ and finite variance σ2, the sampling distribution of the sample mean computed from samples of size n from this population will be approximately normal with mean μ (the population mean) and variance σ2/n (the population variance divided by n) when the sample size n is large.
The central limit theorem allows us to make quite precise probability statements about the population mean by using the sample mean, whatever the distribution of the population (so long as it has finite variance), because the sample mean follows an approximate normal distribution for large-size samples. The obvious question is, “When is a sample’s size large enough that we can assume the sample mean is normally distributed?” In general, when sample size n is greater than or equal to 30, we can assume that the sample mean is approximately normally distributed.5
The central limit theorem states that the variance of the distribution of the sample mean is σ2/n. The positive square root of variance is standard deviation. The standard deviation of a sample statistic is known as the standard error of the statistic. The standard error of the sample mean is an important quantity in applying the central limit theorem in practice.
Definition of the Standard Error of the Sample Mean. For sample mean calculated from a sample generated by a population with standard deviation σ, the standard error of the sample mean is given by one of two expressions:
(1)
when we know σ, the population standard deviation, or by
(2)
when we do not know the population standard deviation and need to use the sample standard deviation, s, to estimate it.6
In practice, we almost always need to use Equation 2. The estimate of s is given by the square root of the sample variance, s2, calculated as follows:
(3)
We will soon see how we can use the sample mean and its standard error to make probability statements about the population mean by using the technique of confidence intervals. First, however, we provide an illustration of the central limit theorem’s force.
EXAMPLE 3 The Central Limit Theorem
It is remarkable that the sample mean for large sample sizes will be distributed normally regardless of the distribution of the underlying population. To illustrate the central limit theorem in action, we specify in this example a distinctly nonnormal distribution and use it to generate a large number of random samples of size 100. We then calculate the sample mean for each sample. The frequency distribution of the calculated sample means is an approximation of the sampling distribution of the sample mean for that sample size. Does that sampling distribution look like a normal distribution?
We return to the telecommunications analyst studying the capital expenditure plans of telecom businesses. Suppose that capital expenditures for communications equipment form a continuous uniform random variable with a lower bound equal to $0 and an upper bound equal to $100—for short, call this a uniform (0, 100) random variable. The probability function of this
continuous uniform random variable has a rather simple shape that is anything but normal. It is a horizontal line with a vertical intercept equal to 1/100. Unlike a normal random variable, for which outcomes close to the mean are most likely, all possible outcomes are equally likely for a uniform random variable.
To illustrate the power of the central limit theorem, we conduct a Monte Carlo simulation to study the capital expenditure plans of telecom businesses.7 In this simulation, we collect 200 random samples of the capital expenditures of 100 companies (200 random draws, each consisting of the capital expenditures of 100 companies with n = 100). In each simulation trial, 100 values for capital expenditure are generated from the uniform (0, 100) distribution. For each random sample, we then compute the sample mean. We conduct 200 simulation trials in total. Because we have specified the distribution generating the samples, we know that the population mean capital expenditure is equal to ($0 + $100 million)/2 = $50 million; the population variance of capital expenditures is equal to (100 − 0)2/12 = 833.33; thus, the standard deviation is $28.87 million and the standard error is under the central limit theorem.8
The results of this Monte Carlo experiment are tabulated in Table 2 in the form of a frequency distribution. This distribution is the estimated sampling distribution of the sample mean.
TABLE 2 Frequency Distribution: 200 Random Samples of a Uniform (0,100) Random Variable
Range of Sample Means ($ Million) Absolute Frequency 42.5 ≤ < 44 1 44 ≤ < 45.5 6 45.5 ≤ < 47 22 47 ≤ < 48.5 39 48.5 ≤ < 50 41 50 ≤ < 51.5 39 51.5 ≤ < 53 23 53 ≤ < 54.5 12 54.5 ≤ < 56 12 56 ≤ < 57.5 5
Note: is the mean capital expenditure for each sample.
The frequency distribution can be described as bell-shaped and centered close to the population mean of 50. The most frequent, or modal, range, with 41 observations, is 48.5 to 50. The overall average of the sample means is $49.92, with a standard error equal to $2.80. The calculated standard error is close to the value of 2.887 given by the central limit theorem. The discrepancy between calculated and expected values of the mean and standard deviation under the central limit theorem is a result of random chance (sampling error).
In summary, although the distribution of the underlying population is very nonnormal, the simulation has shown that a normal distribution well describes the estimated sampling distribution of the sample mean, with mean and standard error consistent with the values predicted by the central limit theorem.
To summarize, according to the central limit theorem, when we sample from any distribution, the distribution of the sample mean will have the following properties as long as our sample size is large:
The distribution of the sample mean will be approximately normal.
The mean of the distribution of will be equal to the mean of the population from which the samples are drawn.
The variance of the distribution of will be equal to the variance of the population divided by the sample size.
We next discuss the concepts and tools related to estimating the population parameters, with a special focus on the population mean. We focus on the population mean because analysts are more likely to meet interval estimates for the population mean than any other type of interval estimate.
4. Point and Interval Estimates of the Population Mean Statistical inference traditionally consists of two branches, hypothesis testing and estimation. Hypothesis testing addresses the question “Is the value of this parameter (say, a population mean) equal to some specific value (0, for example)?” In this process, we have a hypothesis concerning the value of a parameter, and we seek to determine whether the evidence from a sample supports or does not support that hypothesis. We discuss hypothesis testing in detail in the reading on hypothesis testing.
The second branch of statistical inference, and the focus of this reading, is estimation. Estimation seeks an answer to the question “What is this parameter ’s (for example, the population mean’s) value?” In estimating, unlike in hypothesis testing, we do not start with a hypothesis about a parameter ’s value and seek to test it. Rather, we try to make the best use of the information in a sample to form one of several types of estimates of the parameter ’s value. With estimation, we are interested in arriving at a rule for best calculating a single number to estimate the unknown population parameter (a point estimate). Together with calculating a point estimate, we may also be interested in calculating a range of values that brackets the unknown population parameter with some specified level of probability (a confidence interval). In Section 4.1 we discuss point estimates of parameters and then, in Section 4.2, the formulation of confidence intervals for the population mean.
4.1. Point Estimators An important concept introduced in this reading is that sample statistics viewed as formulas involving random outcomes are random variables. The formulas that we use to compute the sample mean and all the other sample statistics are examples of estimation formulas or estimators. The particular value that we calculate from sample observations using an estimator is called an estimate. An estimator has a sampling distribution; an estimate is a fixed number pertaining to a given sample and thus has no sampling distribution. To take the example of the mean, the calculated value of the sample mean in a given sample, used as an estimate of the population mean, is called a point estimate of the population mean. As Example 3 illustrated, the formula for the sample mean can and will yield different results in repeated samples as different samples are drawn from the population.
In many applications, we have a choice among a number of possible estimators for estimating a given parameter. How do we make our choice? We often select estimators because they have one or more desirable statistical properties. Following
is a brief description of three desirable properties of estimators: unbiasedness (lack of bias), efficiency, and consistency.9
Definition of Unbiasedness. An unbiased estimator is one whose expected value (the mean of its sampling distribution) equals the parameter it is intended to estimate.
For example, the expected value of the sample mean, , equals μ, the population mean, so we say that the sample mean is an unbiased estimator (of the population mean). The sample variance, s2, which is calculated using a divisor of n − 1 (Equation 3), is an unbiased estimator of the population variance, σ2. If we were to calculate the sample variance using a divisor of n, the estimator would be biased: Its expected value would be smaller than the population variance. We would say that sample variance calculated with a divisor of n is a biased estimator of the population variance.
Whenever one unbiased estimator of a parameter can be found, we can usually find a large number of other unbiased estimators. How do we choose among alternative unbiased estimators? The criterion of efficiency provides a way to select from among unbiased estimators of a parameter.
Definition of Efficiency. An unbiased estimator is efficient if no other unbiased estimator of the same parameter has a sampling distribution with smaller variance.
To explain the definition, in repeated samples we expect the estimates from an efficient estimator to be more tightly grouped around the mean than estimates from other unbiased estimators. Efficiency is an important property of an estimator.10 Sample mean is an efficient estimator of the population mean; sample variance s2 is an efficient estimator of σ2.
Recall that a statistic’s sampling distribution is defined for a given sample size. Different sample sizes define different sampling distributions. For example, the variance of sampling distribution of the sample mean is smaller for larger sample sizes. Unbiasedness and efficiency are properties of an estimator ’s sampling distribution that hold for any size sample. An unbiased estimator is unbiased equally in a sample of size 10 and in a sample of size 1,000. In some problems, however, we cannot find estimators that have such desirable properties as unbiasedness in small samples.11 In this case, statisticians may justify the choice of an estimator based on the properties of the estimator ’s sampling distribution in extremely large samples, the estimator ’s so-called asymptotic properties. Among such properties, the most
important is consistency.
Definition of Consistency. A consistent estimator is one for which the probability of estimates close to the value of the population parameter increases as sample size increases.
Somewhat more technically, we can define a consistent estimator as an estimator whose sampling distribution becomes concentrated on the value of the parameter it is intended to estimate as the sample size approaches infinity. The sample mean, in addition to being an efficient estimator, is also a consistent estimator of the population mean: As sample size n goes to infinity, its standard error, , goes to 0 and its sampling distribution becomes concentrated right over the value of population mean, μ. To summarize, we can think of a consistent estimator as one that tends to produce more and more accurate estimates of the population parameter as we increase the sample’s size. If an estimator is consistent, we may attempt to increase the accuracy of estimates of a population parameter by calculating estimates using a larger sample. For an inconsistent estimator, however, increasing sample size does not help to increase the probability of accurate estimates.
4.2. Confidence Intervals for the Population Mean When we need a single number as an estimate of a population parameter, we make use of a point estimate. However, because of sampling error, the point estimate is not likely to equal the population parameter in any given sample. Often, a more useful approach than finding a point estimate is to find a range of values that we expect to bracket the parameter with a specified level of probability—an interval estimate of the parameter. A confidence interval fulfills this role.
Definition of Confidence Interval. A confidence interval is a range for which one can assert with a given probability 1 − α, called the degree of confidence, that it will contain the parameter it is intended to estimate. This interval is often referred to as the 100(1 − α)% confidence interval for the parameter.
The endpoints of a confidence interval are referred to as the lower and upper confidence limits. In this reading, we are concerned only with two-sided confidence intervals—confidence intervals for which we calculate both lower and upper limits.12
Confidence intervals are frequently given either a probabilistic interpretation or a practical interpretation. In the probabilistic interpretation, we interpret a 95 percent confidence interval for the population mean as follows. In repeated sampling, 95
percent of such confidence intervals will, in the long run, include or bracket the population mean. For example, suppose we sample from the population 1,000 times, and based on each sample, we construct a 95 percent confidence interval using the calculated sample mean. Because of random chance, these confidence intervals will vary from each other, but we expect 95 percent, or 950, of these intervals to include the unknown value of the population mean. In practice, we generally do not carry out such repeated sampling. Therefore, in the practical interpretation, we assert that we are 95 percent confident that a single 95 percent confidence interval contains the population mean. We are justified in making this statement because we know that 95 percent of all possible confidence intervals constructed in the same manner will contain the population mean. The confidence intervals that we discuss in this reading have structures similar to the following basic structure:
Construction of Confidence Intervals. A 100(1 − α)% confidence interval for a parameter has the following structure.
where
The most basic confidence interval for the population mean arises when we are sampling from a normal distribution with known variance. The reliability factor in this case is based on the standard normal distribution, which has a mean of 0 and a variance of 1. A standard normal random variable is conventionally denoted by Z. The notation zα denotes the point of the standard normal distribution such that α of the probability remains in the right tail. For example, 0.05 or 5 percent of the possible values of a standard normal random variable are larger than z0.05 = 1.65.
Suppose we want to construct a 95 percent confidence interval for the population mean and, for this purpose, we have taken a sample of size 100 from a normally distributed population with known variance of σ2 = 400 (so, σ = 20). We calculate a sample mean of = 25. Our point estimate of the population mean is, therefore, 25. If we move 1.96 standard deviations above the mean of a normal distribution, 0.025 or 2.5 percent of the probability remains in the right tail; by symmetry of the normal distribution, if we move 1.96 standard deviations below the mean, 0.025 or 2.5 percent of the probability remains in the left tail. In total, 0.05 or 5 percent of the probability is in the two tails and 0.95 or 95 percent lies in between. So, z0.025 = 1.96
is the reliability factor for this 95 percent confidence interval. Note the relationship 100(1 − α)% for the confidence interval and the zα/2 for the reliability factor. The standard error of the sample mean, given by Equation 1, is = 2. The confidence interval, therefore, has a lower limit of = 25 – 1.96(2) = 25 – 3.92 = 21.08. The upper limit of the confidence interval is = 25 + 1.96(2) = 25 + 3.92 = 28.92. The 95 percent confidence interval for the population mean spans 21.08 to 28.92.
Confidence Intervals for the Population Mean (Normally Distributed Population with Known Variance). A 100(1 − α)% confidence interval for population mean μ when we are sampling from a normal distribution with known variance σ2 is given by
(4)
The reliability factors for the most frequently used confidence intervals are as follows.
Reliability Factors for Confidence Intervals Based on the Standard Normal Distribution. We use the following reliability factors when we construct confidence intervals based on the standard normal distribution:14
90 percent confidence intervals: Use z0.05 = 1.65
95 percent confidence intervals: Use z0.025 = 1.96
99 percent confidence intervals: Use z0.005 = 2.58
These reliability factors highlight an important fact about all confidence intervals. As we increase the degree of confidence, the confidence interval becomes wider and gives us less precise information about the quantity we want to estimate. “The surer we want to be, the less we have to be sure of.”15
In practice, the assumption that the sampling distribution of the sample mean is at least approximately normal is frequently reasonable, either because the underlying distribution is approximately normal or because we have a large sample and the central limit theorem applies. However, rarely do we know the population variance in practice. When the population variance is unknown but the sample mean is at least approximately normally distributed, we have two acceptable ways to calculate the confidence interval for the population mean. We will soon discuss the more conservative approach, which is based on Student’s t-distribution (the t-distribution,
for short).16 In investment literature, it is the most frequently used approach in both estimation and hypothesis tests concerning the mean when the population variance is not known, whether sample size is small or large.
A second approach to confidence intervals for the population mean, based on the standard normal distribution, is the z-alternative. It can be used only when sample size is large. (In general, a sample size of 30 or larger may be considered large.) In contrast to the confidence interval given in Equation 4, this confidence interval uses the sample standard deviation, s, in computing the standard error of the sample mean (Equation 2).
Confidence Intervals for the Population Mean—The z-Alternative (Large Sample, Population Variance Unknown). A 100(1 − α)% confidence interval for population mean μ when sampling from any distribution with unknown variance and when sample size is large is given by
(5)
Because this type of confidence interval appears quite often, we illustrate its calculation in Example 4.
EXAMPLE 4 Confidence Interval for the Population Mean of Sharpe Ratios—z-Statistic
Suppose an investment analyst takes a random sample of US equity mutual funds and calculates the average Sharpe ratio. The sample size is 100, and the average Sharpe ratio is 0.45. The sample has a standard deviation of 0.30. Calculate and interpret the 90 percent confidence interval for the population mean of all US equity mutual funds by using a reliability factor based on the standard normal distribution.
The reliability factor for a 90 percent confidence interval, as given earlier, is z0.05 = 1.65. The confidence interval will be
The confidence interval spans 0.4005 to 0.4995, or 0.40 to 0.50, carrying two decimal places. The analyst can say with 90 percent confidence that the interval includes the population mean.
In this example, the analyst makes no specific assumption about the probability
distribution describing the population. Rather, the analyst relies on the central limit theorem to produce an approximate normal distribution for the sample mean.
As Example 4 shows, even if we are unsure of the underlying population distribution, we can still construct confidence intervals for the population mean as long as the sample size is large because we can apply the central limit theorem.
We now turn to the conservative alternative, using the t-distribution, for constructing confidence intervals for the population mean when the population variance is not known. For confidence intervals based on samples from normally distributed populations with unknown variance, the theoretically correct reliability factor is based on the t-distribution. Using a reliability factor based on the t- distribution is essential for a small sample size. Using a t reliability factor is appropriate when the population variance is unknown, even when we have a large sample and could use the central limit theorem to justify using a z reliability factor. In this large sample case, the t-distribution provides more-conservative (wider) confidence intervals.
The t-distribution is a symmetrical probability distribution defined by a single parameter known as degrees of freedom (df). Each value for the number of degrees of freedom defines one distribution in this family of distributions. We will shortly compare t-distributions with the standard normal distribution, but first we need to understand the concept of degrees of freedom. We can do so by examining the calculation of the sample variance.
Equation 3 gives the unbiased estimator of the sample variance that we use. The term in the denominator, n − 1, which is the sample size minus 1, is the number of degrees of freedom in estimating the population variance when using Equation 3. We also use n − 1 as the number of degrees of freedom for determining reliability factors based on the t-distribution. The term “degrees of freedom” is used because in a random sample, we assume that observations are selected independently of each other. The numerator of the sample variance, however, uses the sample mean. How does the use of the sample mean affect the number of observations collected independently for the sample variance formula? With a sample of size 10 and a mean of 10 percent, for example, we can freely select only 9 observations. Regardless of the 9 observations selected, we can always find the value for the 10th observation that gives a mean equal to 10 percent. From the standpoint of the sample variance formula, then, there are 9 degrees of freedom. Given that we must first compute the sample mean from the total of n independent observations, only n − 1 observations can be chosen independently for the calculation of the sample variance. The concept of degrees of freedom comes up frequently in statistics, and you will
see it often in later readings.
Suppose we sample from a normal distribution. The ratio is distributed normally with a mean of 0 and standard deviation of 1; however, the ratio follows the t-distribution with a mean of 0 and n − 1 degrees of freedom. The ratio represented by t is not normal because t is the ratio of two random variables, the sample mean and the sample standard deviation. The definition of the standard normal random variable involves only one random variable, the sample mean. As degrees of freedom increase, however, the t- distribution approaches the standard normal distribution. Figure 1 shows the standard normal distribution and two t-distributions, one with df = 2 and one with df = 8.
FIGURE 1 Student’s t-Distribution versus the Standard Normal Distribution
Of the three distributions shown in Figure 1, the standard normal distribution has tails that approach zero faster than the tails of the two t-distributions. The t- distribution is also symmetrically distributed around its mean value of zero, just like the normal distribution. As the degrees of freedom increase, the t-distribution approaches the standard normal. The t-distribution with df = 8 is closer to the standard normal than the t-distribution with df = 2.
Beyond plus and minus four standard deviations from the mean, the area under the standard normal distribution appears to approach 0; both t-distributions continue to show some area under each curve beyond four standard deviations, however. The t- distributions have fatter tails, but the tails of the t-distribution with df = 8 more
closely resemble the normal distribution’s tails. As the degrees of freedom increase, the tails of the t-distribution become less fat.
Frequently referred to values for the t-distribution are presented in tables at the end of the book. For each degree of freedom, five values are given: t0.10, t0.05, t0.025, t0.01, and t0.005. The values for t0.10, t0.05, t0.025, t0.01, and t0.005 are such that, respectively, 0.10, 0.05, 0.025, 0.01, and 0.005 of the probability remains in the right tail, for the specified number of degrees of freedom.17 For example, for df = 30, t0.10 = 1.310, t0.05 = 1.697, t0.025 = 2.042, t0.01 = 2.457, and t0.005 = 2.750.
We now give the form of confidence intervals for the population mean using the t- distribution.
Confidence Intervals for the Population Mean (Population Variance Unknown)— t-Distribution. If we are sampling from a population with unknown variance and either of the conditions below holds:
the sample is large, or
the sample is small but the population is normally distributed, or approximately normally distributed,
then a 100(1 − α)% confidence interval for the population mean μ is given by
(6)
where the number of degrees of freedom for tα/2 is n − 1 and n is the sample size.
Example 5 reprises the data of Example 4 but uses the t-statistic rather than the z- statistic to calculate a confidence interval for the population mean of Sharpe ratios.
EXAMPLE 5 Confidence Interval for the Population Mean of Sharpe Ratios—t-Statistic
As in Example 4, an investment analyst seeks to calculate a 90 percent confidence interval for the population mean Sharpe ratio of US equity mutual funds based on a random sample of 100 US equity mutual funds. The sample mean Sharpe ratio is 0.45, and the sample standard deviation of the Sharpe ratios is 0.30. Now recognizing that the population variance of the distribution of Sharpe ratios is unknown, the analyst decides to calculate the confidence interval using the theoretically correct t-statistic.
Because the sample size is 100, df = 99. In the tables in the back of the book, the closest value is df = 100. Using df = 100 and reading down the 0.05 column, we find that t0.05 = 1.66. This reliability factor is slightly larger than the reliability factor z0.05 = 1.65 that was used in Example 4. The confidence interval will be
The confidence interval spans 0.4002 to 0.4998, or 0.40 to 0.50, carrying two decimal places. To two decimal places, the confidence interval is unchanged from the one computed in Example 4.
Table 3 summarizes the various reliability factors that we have used.
TABLE 3 Basis of Computing Reliability Factors
Sampling from: Statistic for SmallSample Size Statistic for Large
Sample Size Normal distribution with known
variance z z
Normal distribution with unknown variance t t*
Nonnormal distribution with known variance not available z
Nonnormal distribution with unknown variance not available t*
*Use of z also acceptable.
4.3. Selection of Sample Size What choices affect the width of a confidence interval? To this point we have discussed two factors that affect width: the choice of statistic (t or z) and the choice of degree of confidence (affecting which specific value of t or z we use). These two choices determine the reliability factor. (Recall that a confidence interval has the structure Point estimate ± Reliability factor × Standard error.)
The choice of sample size also affects the width of a confidence interval. All else equal, a larger sample size decreases the width of a confidence interval. Recall the expression for the standard error of the sample mean:
We see that the standard error varies inversely with the square root of sample size. As we increase sample size, the standard error decreases and consequently the width of the confidence interval also decreases. The larger the sample size, the greater precision with which we can estimate the population parameter.18 All else equal, larger samples are good, in that sense. In practice, however, two considerations may operate against increasing sample size. First, as we saw in Example 2 concerning the Sharpe ratio, increasing the size of a sample may result in sampling from more than one population. Second, increasing sample size may involve additional expenses that outweigh the value of additional precision. Thus three issues that the analyst should weigh in selecting sample size are the need for precision, the risk of sampling from more than one population, and the expenses of different sample sizes.
EXAMPLE 6 A Money Manager Estimates Net Client Inflows
A money manager wants to obtain a 95 percent confidence interval for fund inflows and outflows over the next six months for his existing clients. He begins by calling a random sample of 10 clients and inquiring about their planned additions to and withdrawals from the fund. The manager then computes the change in cash flow for each client sampled as a percentage change in total funds placed with the manager. A positive percentage change indicates a net cash inflow to the client’s account, and a negative percentage change indicates a net cash outflow from the client’s account. The manager weights each response by the relative size of the account within the sample and then computes a weighted average.
As a result of this process, the money manager computes a weighted average of 5.5 percent. Thus, a point estimate is that the total amount of funds under management will increase by 5.5 percent in the next six months. The standard deviation of the observations in the sample is 10 percent. A histogram of past data looks fairly close to normal, so the manager assumes the population is normal.
1. Calculate a 95 percent confidence interval for the population mean and
interpret your findings.
The manager decides to see what the confidence interval would look like if he had used a sample size of 20 or 30 and found the same mean (5.5
percent) and standard deviation (10 percent).
2. Using the sample mean of 5.5 percent and standard deviation of 10 percent, compute the confidence interval for sample sizes of 20 and 30. For the sample size of 30, use Equation 6.
3. Interpret your results from Parts 1 and 2.
Solution to 1: Because the population is unknown and the sample size is small, the manager must use the t-statistic in Equation 6 to calculate the confidence interval. Based on the sample size of 10, df = n − 1 = 10 − 1 = 9. For a 95 percent confidence interval, he needs to use the value of t0.025 for df = 9. According to the tables in Appendix B at the end of this volume, this value is 2.262. Therefore, a 95 percent confidence interval for the population mean is
The confidence interval for the population mean spans −1.65 percent to +12.65 percent.19 The manager can be confident at the 95 percent level that this range includes the population mean.
Solution to 2: Table 4 gives the calculations for the three sample sizes.
TABLE 4 The 95 Percent Confidence Interval for Three Sample Sizes
Distribution 95% ConfidenceInterval Lower Bound
Upper Bound Relative Size
t(n = 10) 5.5% ± 2.262(3.162) −1.65% 12.65% 100.0% t(n = 20) 5.5% ± 2.093(2.236) 0.82 10.18 65.5 t(n = 30) 5.5% ± 2.045(1.826) 1.77 9.23 52.2
Solution to 3: The width of the confidence interval decreases as we increase the sample size. This decrease is a function of the standard error becoming smaller as n increases. The reliability factor also becomes smaller as the number of degrees of freedom increases. The last column of Table 4 shows the relative size of the width of confidence intervals based on n = 10 to be 100 percent. Using a sample size of 20 reduces the confidence interval’s width to 65.5 percent of the interval width for a sample size of 10. Using a sample size of 30 cuts the width of the interval almost in half. Comparing these choices, the money manager would obtain the most precise results using a sample of 30.
Having covered many of the fundamental concepts of sampling and estimation, we are in a good position to focus on sampling issues of special concern to analysts. The quality of inferences depends on the quality of the data as well as on the quality of the sampling plan used. Financial data pose special problems, and sampling plans frequently reflect one or more biases. The next section of this reading discusses these issues.
5. More on Sampling We have already seen that the selection of sample period length may raise the issue of sampling from more than one population. There are, in fact, a range of challenges to valid sampling that arise in working with financial data. In this section we discuss four such sampling-related issues: data-mining bias, sample selection bias, look-ahead bias, and time-period bias. All of these issues are important for point and interval estimation and hypothesis testing. As we will see, if the sample is biased in any way, then point and interval estimates and any other conclusions that we draw from the sample will be in error.
5.1. Data-Mining Bias Data mining relates to overuse of the same or related data in ways that we shall describe shortly. Data-mining bias refers to the errors that arise from such misuse of data. Investment strategies that reflect data-mining biases are often not successful in the future. Nevertheless, both investment practitioners and researchers have frequently engaged in data mining. Analysts thus need to understand and guard against this problem.
Data-mining is the practice of determining a model by extensive searching through a dataset for statistically significant patterns (that is, repeatedly “drilling” in the same data until finding something that appears to work).20 In exercises involving statistical significance we set a significance level, which is the probability of rejecting the hypothesis we are testing when the hypothesis is in fact correct.21 Because rejecting a true hypothesis is undesirable, the investigator often sets the significance level at a relatively small number such as 0.05 or 5 percent.22 Suppose we test the hypothesis that a variable does not predict stock returns, and we test in turn 100 different variables. Let us also suppose that in truth none of the 100 variables has the ability to predict stock returns. Using a 5 percent significance level in our tests, we would still expect that 5 out of 100 variables would appear to be significant predictors of stock returns because of random chance alone. We have mined the data to find some apparently significant variables. In essence, we have explored the same data again and again until we found some after-the-fact pattern or patterns in the dataset. This is the sense in which data mining involves overuse of data. If we were to just report the significant variables, without also reporting the total number of variables that we tested that were unsuccessful as predictors, we would be presenting a very misleading picture of our findings. Our results would appear to be far more significant than they actually were, because a series of tests such as the one just described invalidates the conventional interpretation of a given
significance level (such as 5 percent), according to the theory of inference.
How can we investigate the presence of data-mining bias? With most financial data, the most ready means is to conduct out-of-sample tests of the proposed variable or strategy. An out-of-sample test uses a sample that does not overlap the time period(s) of the sample(s) on which a variable, strategy, or model, was developed. If a variable or investment strategy is the result of data mining, it should generally not be significant in out-of-sample tests. A variable or investment strategy that is statistically and economically significant in out-of-sample tests, and that has a plausible economic basis, may be the basis for a valid investment strategy. Caution is still warranted, however. The most crucial out-of-sample test is future investment success. If the strategy becomes known to other investors, prices may adjust so that the strategy, however well tested, does not work in the future. To summarize, the analyst should be aware that many apparently profitable investment strategies may reflect data-mining bias and thus be cautious about the future applicability of published investment research results.
Untangling the extent of data mining can be complex. To assess the significance of an investment strategy, we need to know how many unsuccessful strategies were tried not only by the current investigator but also by previous investigators using the same or related datasets. Much research, in practice, closely builds on what other investigators have done, and so reflects intergenerational data mining, to use the terminology of McQueen and Thorley (1999). Intergenerational data mining involves using information developed by previous researchers using a dataset to guide current research using the same or a related dataset.23 Analysts have accumulated many observations about the peculiarities of many financial datasets, and other analysts may develop models or investment strategies that will tend to be supported within a dataset based on their familiarity with the prior experience of other analysts. As a consequence, the importance of those new results may be overstated. Research has suggested that the magnitude of this type of data-mining bias may be considerable.24
With the background of the above definitions and explanations, we can understand McQueen and Thorley’s (1999) cogent exploration of data mining in the context of the popular Motley Fool “Foolish Four” investment strategy. The Foolish Four strategy, first presented in 1996, was a version of the Dow Dividend Strategy that was tuned by its developers to exhibit an even higher arithmetic mean return than the Dow Dividend Strategy over 1973 to 1993.25 From 1973 to 1993, the Foolish Four portfolio had an average annual return of 25 percent, and the claim was made in print that the strategy should have similar returns in the future. As McQueen and Thorley discussed, however, the Foolish Four strategy was very much subject to data-mining bias, including bias from intergenerational data mining, as the
strategy’s developers exploited observations about the dataset made by earlier workers. McQueen and Thorley highlighted the data-mining issues by taking the Foolish Four portfolio one step further. They mined the data to create a “Fractured Four” portfolio that earned nearly 35 percent over 1973 to 1996, beating the Foolish Four strategy by almost 8 percentage points. Observing that all of the Foolish Four stocks did well in even years but not odd years and that the second-to-lowest-priced high-yielding stock was relatively the best-performing stock in odd years, the strategy of the Fractured Four portfolio was to hold the Foolish Four stocks with equal weights in even years and hold only the second-to-lowest-priced stock in odd years. How likely is it that a performance difference between even and odd years reflected underlying economic forces, rather than a chance pattern of the data over the particular time period? Probably, very unlikely. Unless an investment strategy reflected underlying economic forces, we would not expect it to have any value in a forward-looking sense. Because the Foolish Four strategy also partook of data mining, the same issues applied to it. McQueen and Thorley found that in an out-of- sample test over the 1949–72 period, the Foolish Four strategy had about the same mean return as buying and holding the DJIA, but with higher risk. If the higher taxes and transaction costs of the Foolish Four strategy were accounted for, the comparison would have been even more unfavorable.
McQueen and Thorley presented two signs that can warn analysts about the potential existence of data mining:
Too much digging/too little confidence. The testing of many variables by the researcher is the “too much digging” warning sign of a data-mining problem. Unfortunately, many researchers do not disclose the number of variables examined in developing a model. Although the number of variables examined may not be reported, we should look closely for verbal hints that the researcher searched over many variables. The use of terms such as “we noticed (or noted) that” or “someone noticed (or noted) that,” with respect to a pattern in a dataset, should raise suspicions that the researchers were trying out variables based on their own or others’ observations of the data.
No story/no future. The absence of an explicit economic rationale for a variable or trading strategy is the “no story” warning sign of a data-mining problem. Without a plausible economic rationale or story for why a variable should work, the variable is unlikely to have predictive power. In a demonstration exercise using an extensive search of variables in an international financial database, Leinweber (1997) found that butter production in a particular country remote from the United States explained 75 percent of the variation in US stock returns as represented by the S&P 500. Such a pattern,
with no plausible economic rationale, is highly likely to be a random pattern particular to a specific time period.26 What if we do have a plausible economic explanation for a significant variable? McQueen and Thorley caution that a plausible economic rationale is a necessary but not a sufficient condition for a trading strategy to have value. As we mentioned earlier, if the strategy is publicized, market prices may adjust to reflect the new information as traders seek to exploit it; as a result, the strategy may no longer work.
5.2. Sample Selection Bias When researchers look into questions of interest to analysts or portfolio managers, they may exclude certain stocks, bonds, portfolios, or time periods from the analysis for various reasons—perhaps because of data availability. When data availability leads to certain assets being excluded from the analysis, we call the resulting problem sample selection bias. For example, you might sample from a database that tracks only companies currently in existence. Many mutual fund databases, for instance, provide historical information about only those funds that currently exist. Databases that report historical balance sheet and income statement information suffer from the same sort of bias as the mutual fund databases: Funds or companies that are no longer in business do not appear there. So, a study that uses these types of databases suffers from a type of sample selection bias known as survivorship bias.
Dimson, Marsh, and Staunton (2002) raised the issue of survivorship bias in international indexes:
An issue that has achieved prominence is the impact of market survival on estimated long-run returns. Markets can experience not only disappointing performance but also total loss of value through confiscation, hyperinflation, nationalization, and market failure. By measuring the performance of markets that survive over long intervals, we draw inferences that are conditioned on survival. Yet, as pointed out by Brown, Goetzmann, and Ross (1995) and Goetzmann and Jorion (1999), one cannot determine in advance which markets will survive and which will perish. (p. 41)
Survivorship bias sometimes appears when we use both stock price and accounting data. For example, many studies in finance have used the ratio of a company’s market price to book equity per share (i.e., the price-to-book ratio, P/B) and found that P/B is inversely related to a company’s returns (see Fama and French 1992, 1993). P/B is also used to create many popular value and growth indexes. If the database that we use to collect accounting data excludes failing companies, however, a survivorship bias might result. Kothari, Shanken, and Sloan (1995) investigated
just this question and argued that failing stocks would be expected to have low returns and low P/Bs. If we exclude failing stocks, then those stocks with low P/Bs that are included will have returns that are higher on average than if all stocks with low P/Bs were included. Kothari, Shanken, and Sloan suggested that this bias is responsible for the previous findings of an inverse relationship between average return and P/B.27 The only advice we can offer at this point is to be aware of any biases potentially inherent in a sample. Clearly, sample selection biases can cloud the results of any study.
A sample can also be biased because of the removal (or delisting) of a company’s stock from an exchange.28 For example, the Center for Research in Security Prices at the University of Chicago is a major provider of return data used in academic research. When a delisting occurs, CRSP attempts to collect returns for the delisted company, but many times, it cannot do so because of the difficulty involved; CRSP must simply list delisted company returns as missing. A study in the Journal of Finance by Shumway and Warther (1999) documented the bias caused by delisting for CRSP NASDAQ return data. The authors showed that delistings associated with poor company performance (e.g., bankruptcy) are missed more often than delistings associated with good or neutral company performance (e.g., merger or moving to another exchange). In addition, delistings occur more frequently for small companies.
Sample selection bias occurs even in markets where the quality and consistency of the data are quite high. Newer asset classes such as hedge funds may present even greater problems of sample selection bias. Hedge funds are a heterogeneous group of investment vehicles typically organized so as to be free from regulatory oversight. In general, hedge funds are not required to publicly disclose performance (in contrast to, say, mutual funds). Hedge funds themselves decide whether they want to be included in one of the various databases of hedge fund performance. Hedge funds with poor track records clearly may not wish to make their records public, creating a problem of self-selection bias in hedge fund databases. Further, as pointed out by Fung and Hsieh (2002), because only hedge funds with good records will volunteer to enter a database, in general, overall past hedge fund industry performance will tend to appear better than it really is. Furthermore, many hedge fund databases drop funds that go out of business, creating survivorship bias in the database. Even if the database does not drop defunct hedge funds, in the attempt to eliminate survivorship bias, the problem remains of hedge funds that stop reporting performance because of poor results.29
5.3. Look-Ahead Bias A test design is subject to look-ahead bias if it uses information that was not
available on the test date. For example, tests of trading rules that use stock market returns and accounting balance sheet data must account for look-ahead bias. In such tests, a company’s book value per share is commonly used to construct the P/B variable. Although the market price of a stock is available for all market participants at the same point in time, fiscal year-end book equity per share might not become publicly available until sometime in the following quarter.
5.4. Time-Period Bias A test design is subject to time-period bias if it is based on a time period that may make the results time-period specific. A short time series is likely to give period specific results that may not reflect a longer period. A long time series may give a more accurate picture of true investment performance; its disadvantage lies in the potential for a structural change occurring during the time frame that would result in two different return distributions. In this situation, the distribution that would reflect conditions before the change differs from the distribution that would describe conditions after the change.
EXAMPLE 7 Biases in Investment Research
An analyst is reviewing the empirical evidence on historical US equity returns. She finds that value stocks (i.e., those with low P/Bs) outperformed growth stocks (i.e., those with high P/Bs) in some recent time periods. After reviewing the US market, the analyst wonders whether value stocks might be attractive in the United Kingdom. She investigates the performance of value and growth stocks in the UK market for the 14-year period extending from January 2000 to December 2013. To conduct this research, the analyst does the following:
obtains the current composition of the Financial Times Stock Exchange (FTSE) All Share Index, which is a market-capitalization-weighted index;
eliminates the few companies that do not have December fiscal year-ends;
uses year-end book values and market prices to rank the remaining universe of companies by P/Bs at the end of the year;
based on these rankings, divides the universe into 10 portfolios, each of which contains an equal number of stocks;
calculates the equal-weighted return of each portfolio and the return for the FTSE All Share Index for the 12 months following the date each ranking was made; and
subtracts the FTSE returns from each portfolio’s returns to derive excess returns for each portfolio.
Describe and discuss each of the following biases introduced by the analyst’s research design:
survivorship bias;
look-ahead bias; and
time-period bias.
Survivorship Bias
A test design is subject to survivorship bias if it fails to account for companies that have gone bankrupt, merged, or otherwise departed the database. In this example, the analyst used the current list of FTSE stocks rather than the actual list of stocks that existed at the start of each year. To the extent that the computation of returns excluded companies removed from the index, the performance of the portfolios with the lowest P/B is subject to survivorship bias and may be overstated. At some time during the testing period, those companies not currently in existence were eliminated from testing. They would probably have had low prices (and low P/Bs) and poor returns.
Look-Ahead Bias
A test design is subject to look-ahead bias if it uses information unavailable on the test date. In this example, the analyst conducted the test under the assumption that the necessary accounting information was available at the end of the fiscal year. For example, the analyst assumed that book value per share for fiscal 2000 was available on 31 December 2000. Because this information is not released until several months after the close of a fiscal year, the test may have contained look-ahead bias. This bias would make a strategy based on the information appear successful, but it assumes perfect forecasting ability.
Time-Period Bias
A test design is subject to time-period bias if it is based on a time period that may make the results time-period specific. Although the test covered a period extending more than 10 years, that period may be too short for testing an anomaly. Ideally, an analyst should test market anomalies over several business cycles to ensure that results are not period specific. This bias can favor a proposed strategy if the time period chosen was favorable to the strategy.
6. Summary In this reading, we have presented basic concepts and results in sampling and estimation. We have also emphasized the challenges faced by analysts in appropriately using and interpreting financial data. As analysts, we should always use a critical eye when evaluating the results from any study. The quality of the sample is of the utmost importance: If the sample is biased, the conclusions drawn from the sample will be in error.
To draw valid inferences from a sample, the sample should be random.
In simple random sampling, each observation has an equal chance of being selected. In stratified random sampling, the population is divided into subpopulations, called strata or cells, based on one or more classification criteria; simple random samples are then drawn from each stratum.
Stratified random sampling ensures that population subdivisions of interest are represented in the sample. Stratified random sampling also produces more- precise parameter estimates than simple random sampling.
Time-series data are a collection of observations at equally spaced intervals of time. Cross-sectional data are observations that represent individuals, groups, geographical regions, or companies at a single point in time.
The central limit theorem states that for large sample sizes, for any underlying distribution for a random variable, the sampling distribution of the sample mean for that variable will be approximately normal, with mean equal to the population mean for that random variable and variance equal to the population variance of the variable divided by sample size.
Based on the central limit theorem, when the sample size is large, we can compute confidence intervals for the population mean based on the normal distribution regardless of the distribution of the underlying population. In general, a sample size of 30 or larger can be considered large.
An estimator is a formula for estimating a parameter. An estimate is a particular value that we calculate from a sample by using an estimator.
Because an estimator or statistic is a random variable, it is described by some probability distribution. We refer to the distribution of an estimator as its sampling distribution. The standard deviation of the sampling distribution of the sample mean is called the standard error of the sample mean.
The desirable properties of an estimator are unbiasedness (the expected value of the estimator equals the population parameter), efficiency (the estimator has the smallest variance), and consistency (the probability of accurate estimates increases as sample size increases).
The two types of estimates of a parameter are point estimates and interval estimates. A point estimate is a single number that we use to estimate a parameter. An interval estimate is a range of values that brackets the population parameter with some probability.
A confidence interval is an interval for which we can assert with a given probability 1 − α, called the degree of confidence, that it will contain the parameter it is intended to estimate. This measure is often referred to as the 100(1 − α)% confidence interval for the parameter.
A 100(1 − α)% confidence interval for a parameter has the following structure: Point estimate ± Reliability factor × Standard error, where the reliability factor is a number based on the assumed distribution of the point estimate and the degree of confidence (1 − α) for the confidence interval and where standard error is the standard error of the sample statistic providing the point estimate.
A 100(1 − α)% confidence interval for population mean μ when sampling from
a normal distribution with known variance σ2 is given by , where zα/2 is the point of the standard normal distribution such that α/2 remains in the right tail.
Student’s t-distribution is a family of symmetrical distributions defined by a single parameter, degrees of freedom.
A random sample of size n is said to have n − 1 degrees of freedom for estimating the population variance, in the sense that there are only n − 1 independent deviations from the mean on which to base the estimate.
The degrees of freedom number for use with the t-distribution is also n − 1.
The t-distribution has fatter tails than the standard normal distribution but converges to the standard normal distribution as degrees of freedom go to infinity.
A 100(1 − α)% confidence interval for the population mean μ when sampling from a normal distribution with unknown variance (a t-distribution confidence interval) is given by , where tα/2 is the point of the t-distribution such that α/2 remains in the right tail and s is the sample standard deviation. This confidence interval can also be used, because of the central limit theorem,
when dealing with a large sample from a population with unknown variance that may not be normal.
We may use the confidence interval as an alternative to the t- distribution confidence interval for the population mean when using a large sample from a population with unknown variance. The confidence interval based on the z-statistic is less conservative (narrower) than the corresponding confidence interval based on a t-distribution.
Three issues in the selection of sample size are the need for precision, the risk of sampling from more than one population, and the expenses of different sample sizes.
Sample data in investments can have a variety of problems. Survivorship bias occurs if companies are excluded from the analysis because they have gone out of business or because of reasons related to poor performance. Data-mining bias comes from finding models by repeatedly searching through databases for patterns. Look-ahead bias exists if the model uses data not available to market participants at the time the market participants act in the model. Finally, time- period bias is present if the time period used makes the results time-period specific or if the time period used includes a point of structural change.
References
1. Brown, Stephen, William Goetzmann, and Stephen Ross. 1995. “Survival.” Journal of Finance, vol. 50: 853–873.
2. Campbell, John, Andrew Lo, and A. Craig MacKinlay. 1997. The Econometrics of Financial Markets. Princeton, NJ: Princeton University Press.
3. Daniel, Wayne W., and James C. Terrell. 1995. Business Statistics for Management & Economics, 7th edition. Boston: Houghton-Mifflin.
4. Dimson, Elroy, Paul Marsh, and Mike Staunton. 2002. Triumphs of the Optimists: 101 Years of Global Investment Returns. Princeton, NJ: Princeton University Press.
5. Fabozzi, Frank J. 2007. Fixed Income Analysis, 2nd edition. Hoboken, NJ: Wiley.
6. Fama, Eugene F., and Kenneth R. French. 1996. “Multifactor Explanations of Asset Pricing Anomalies.” Journal of Finance, vol. 51, no. 1: 55–84.
7. Freund, John E, and Frank J. Williams. 1977. Elementary Business Statistics, 3rd edition. Englewood Cliffs, NJ: Prentice-Hall.
8. Fung, William, and David Hsieh. 2002. “Hedge-Fund Benchmarks: Information Content and Biases.” Financial Analysts Journal, vol. 58, no. 1: 22–34.
9. Goetzmann, William, and Philippe Jorion. 1999. “Re-Emerging Markets.” Journal of Financial and Quantitative Analysis, vol. 34, no. 1: 1–32.
10. Greene, William H. 2011. Econometric Analysis, 7th edition. Upper Saddle River, NJ: Prentice-Hall.
11. Kothari, S.P., Jay Shanken, and Richard G. Sloan. 1995. “Another Look at the Cross-Section of Expected Stock Returns.” Journal of Finance, vol. 50, no. 1: 185–224.
12. Leinweber, David. 1997. Stupid Data Mining Tricks: Over-Fitting the S&P 500. Monograph. Pasadena, CA: First Quadrant.
13. Lo, Andrew W., and A. Craig MacKinlay. 1990. “Data Snooping Biases in Tests of Financial Asset Pricing Models.” Review of Financial Studies, vol. 3: 175–
208.
14. McQueen, Grant, and Steven Thorley. 1999. “Mining Fools Gold.” Financial Analysts Journal, vol. 55, no. 2: 61–72.
15. Shumway, Tyler, and Vincent A. Warther. 1999. “The Delisting Bias in CRSP’s Nasdaq Data and Its Implications for the Size Effect.” Journal of Finance, vol. 54, no. 6: 2361–2379.
16. ter Horst, Jenke, and Marno Verbeek. 2007. “Fund Liquidation, Self-Selection, and Look-ahead Bias in the Hedge Fund Industry.” Review of Finance, vol. 11: 605–632.
Problems Practice Problems and Solutions: 1–20 taken from Quantitative Methods for Investment Analysis, Second Edition, by Richard A. DeFusco, CFA, Dennis W. McLeavey, CFA, Jerald E. Pinto, CFA, and David E. Runkle, CFA. Copyright © 2004 by CFA Institute. All other problems and solutions copyright © CFA Institute. 1. Peter Biggs wants to know how growth managers performed last year. Biggs
assumes that the population cross-sectional standard deviation of growth manager returns is 6 percent and that the returns are independent across managers.
1. How large a random sample does Biggs need if he wants the standard deviation of the sample means to be 1 percent?
2. How large a random sample does Biggs need if he wants the standard deviation of the sample means to be 0.25 percent?
2. Petra Munzi wants to know how value managers performed last year. Munzi estimates that the population cross-sectional standard deviation of value manager returns is 4 percent and assumes that the returns are independent across managers.
1. Munzi wants to build a 95 percent confidence interval for the mean return. How large a random sample does Munzi need if she wants the 95 percent confidence interval to have a total width of 1 percent?
2. Munzi expects a cost of about $10 to collect each observation. If she has a $1,000 budget, will she be able to construct the confidence interval she wants?
3. Assume that the equity risk premium is normally distributed with a population mean of 6 percent and a population standard deviation of 18 percent. Over the last four years, equity returns (relative to the risk-free rate) have averaged −2.0 percent. You have a large client who is very upset and claims that results this poor should never occur. Evaluate your client’s concerns.
1. Construct a 95 percent confidence interval around the population mean for a sample of four-year returns.
2. What is the probability of a −2.0 percent or lower average return over a four-year period?
4. Compare the standard normal distribution and Student’s t-distribution.
5. Find the reliability factors based on the t-distribution for the following confidence intervals for the population mean (df = degrees of freedom, n = sample size):
1. A 99 percent confidence interval, df = 20.
2. A 90 percent confidence interval, df = 20.
3. A 95 percent confidence interval, n = 25.
4. A 95 percent confidence interval, n = 16.
6. Assume that monthly returns are normally distributed with a mean of 1 percent and a sample standard deviation of 4 percent. The population standard deviation is unknown. Construct a 95 percent confidence interval for the sample mean of monthly returns if the sample size is 24.
7. Ten analysts have given the following fiscal-year earnings forecasts for a stock: Forecast (Xi) Number of Analysts (ni)
1.40 1 1.43 1 1.44 3 1.45 2 1.47 1 1.48 1 1.50 1
Because the sample is a small fraction of the number of analysts who follow this stock, assume that we can ignore the finite population correction factor. Assume that the analyst forecasts are normally distributed.
1. What are the mean forecast and standard deviation of forecasts?
2. Provide a 95 percent confidence interval for the population mean of the forecasts.
8. Thirteen analysts have given the following fiscal-year earnings forecasts for a stock: Forecast (Xi) Number of Analysts (ni)
0.70 2
0.72 4 0.74 1 0.75 3 0.76 1 0.77 1 0.82 1
Because the sample is a small fraction of the number of analysts who follow this stock, assume that we can ignore the finite population correction factor.
1. What are the mean forecast and standard deviation of forecasts?
2. What aspect of the data makes us uncomfortable about using t-tables to construct confidence intervals for the population mean forecast?
9. Explain the differences between constructing a confidence interval when sampling from a normal population with a known population variance and sampling from a normal population with an unknown variance.
10. An exchange rate has a given expected future value and standard deviation.
1. Assuming that the exchange rate is normally distributed, what are the probabilities that the exchange rate will be at least 2 or 3 standard deviations away from its mean?
2. Assume that you do not know the distribution of exchange rates. Use Chebyshev’s inequality (that at least 1 − 1/k2 proportion of the observations will be within k standard deviations of the mean for any positive integer k greater than 1) to calculate the maximum probabilities that the exchange rate will be at least 2 or 3 standard deviations away from its mean.
11. Although he knows security returns are not independent, a colleague makes the claim that because of the central limit theorem, if we diversify across a large number of investments, the portfolio standard deviation will eventually approach zero as n becomes large. Is he correct?
12. Why is the central limit theorem important?
13. What is wrong with the following statement of the central limit theorem?
Central Limit Theorem. “If the random variables X1, X2, X3, …, Xn are a random sample of size n from any distribution with finite mean μ and
variance σ2, then the distribution of will be approximately normal, with a standard deviation of .”
14. Suppose we take a random sample of 30 companies in an industry with 200 companies. We calculate the sample mean of the ratio of cash flow to total debt for the prior year. We find that this ratio is 23 percent. Subsequently, we learn that the population cash flow to total debt ratio (taking account of all 200 companies) is 26 percent. What is the explanation for the discrepancy between the sample mean of 23 percent and the population mean of 26 percent?
1. Sampling error.
2. Bias.
3. A lack of consistency.
15. Alcorn Mutual Funds is placing large advertisements in several financial publications. The advertisements prominently display the returns of 5 of Alcorn’s 30 funds for the past 1-, 3-, 5-, and 10-year periods. The results are indeed impressive, with all of the funds beating the major market indexes and a few beating them by a large margin. Is the Alcorn family of funds superior to its competitors?
16. A pension plan executive says, “One hundred percent of our portfolio managers are hired because they have above-average performance records relative to their benchmarks. We do not keep portfolio managers who have below-average records. And yet, each year about half of our managers beat their benchmarks and about half do not. What is going on?” Give a possible statistical explanation.
17. Julius Spence has tested several predictive models in order to identify undervalued stocks. Spence used about 30 company-specific variables and 10 market-related variables to predict returns for about 5,000 North American and European stocks. He found that a final model using eight variables applied to telecommunications and computer stocks yields spectacular results. Spence wants you to use the model to select investments. Should you? What steps would you take to evaluate the model?
18. Hand Associates manages two portfolios that are meant to closely track the returns of two stock indexes. One index is a value-weighted index of 500 stocks in which the weight for each stock depends on the stock’s total market value. The other index is an equal-weighted index of 500 stocks in which the weight for each stock is 1/500. Hand Associates invests in only about 50 to 100 stocks in each portfolio in order to control transactions costs. Should Hand use simple
random sampling or stratified random sampling to choose the stocks in each portfolio?
19. Give an example of each of the following:
1. Sample-selection bias.
2. Look-ahead bias.
3. Time-period bias.
20. What are some of the desirable statistical properties of an estimator, such as a sample mean?
21. An analyst stated that as degrees of freedom increase, a t-distribution will become more peaked and the tails of the t-distribution will become less fat. Is the analyst’s statement correct with respect to the t-distribution:
becoming more peaked? tails becoming less fat? A. No Yes B. Yes No C. Yes Yes
22. An analyst stated that, all else equal, increasing sample size will decrease both the standard error and the width of the confidence interval. The analyst’s statement is correct in regard to:
1. both the standard error and the confidence interval.
2. the standard error, but incorrect in regard to the confidence interval.
3. the confidence interval, but incorrect in regard to the standard error.
Notes 1 See Fabozzi (2007).
2 When the mean or variance of a time series is not constant through time, the time series is not stationary.
3 The reader may also encounter two types of datasets that have both time-series and cross-sectional aspects. Panel data consist of observations through time on a single characteristic of multiple observational units. For example, the annual inflation rate of the Eurozone countries over a five-year period would represent panel data. Longitudinal data consist of observations on characteristic(s) of the same observational unit through time. Observations on a set of financial ratios for a single company over a 10-year period would be an example of longitudinal data. Both panel and longitudinal data may be represented by arrays (matrixes) in which successive rows represent the observations for successive time periods.
4 Sometimes confusion arises because “sample mean” is also used in another sense. When we calculate the sample mean for a particular sample, we obtain a definite number, say 8. If we state that “the sample mean is 8” we are using “sample mean” in the sense of a particular outcome of sample mean as a random variable. The number 8 is of course a constant and does not have a probability distribution. In this discussion, we are not referring to “sample mean” in the sense of a constant number related to a particular sample.
5 When the underlying population is very nonnormal, a sample size well in excess of 30 may be required for the normal distribution to be a good description of the sampling distribution of the mean.
6 We need to note a technical point: When we take a sample of size n from a finite population of size N, we apply a shrinkage factor to the estimate of the standard error of the sample mean that is called the finite population correction factor (fpc). The fpc is equal to [(N − n)/(N − 1)]1/2. Thus, if N = 100 and n = 20, [(100 − 20)/(100 − 1)]1/2 = 0.898933. If we have estimated a standard error of, say, 20, according to Equation 1 or Equation 2, the new estimate is 20(0.898933) = 17.978663. The fpc applies only when we sample from a finite population without replacement; most practitioners also do not apply the fpc if sample size n is very small relative to N (say, less than 5 percent of N). For more information on the finite population correction factor, see Daniel and Terrell (1995).
7 Monte Carlo simulation involves the use of a computer to represent the operation of a system subject to risk. An integral part of Monte Carlo simulation is the generation of a large number of random samples from a specified probability distribution or distributions.
8 If a is the lower limit of a uniform random variable and b is the upper limit, then the random variable’s mean is given by (a + b)/2 and its variance is given by (b − a)2/12. The reading on common probability distributions fully describes continuous uniform random variables.
9 See Daniel and Terrell (1995) or Greene (2011) for a thorough treatment of the properties of estimators.
10 An efficient estimator is sometimes referred to as the best unbiased estimator.
11 Such problems frequently arise in regression and time-series analyses.
12 It is also possible to define two types of one-sided confidence intervals for a population parameter. A lower one-sided confidence interval establishes a lower limit only. Associated with such an interval is an assertion that with a specified degree of confidence the population parameter equals or exceeds the lower limit. An upper one-sided confidence interval establishes an upper limit only; the related assertion is that the population parameter is less than or equal to that upper limit, with a specified degree of confidence. Investment researchers rarely present one-sided confidence intervals, however.
13 The quantity (reliability factor) × (standard error) is sometimes called the precision of the estimator; larger values of the product imply lower precision in estimating the population parameter.
14 Most practitioners use values for z0.05 and z0.005 that are carried to two decimal places. For reference, more exact values for z0.05 and z0.005 are 1.645 and 2.575, respectively. For a quick calculation of a 95 percent confidence interval, z0.025 is sometimes rounded from 1.96 to 2.
15 Freund and Williams (1977), p. 266.
16 The distribution of the statistic t is called Student’s t-distribution after the pen name “Student” used by W. S. Gosset, who published his work in 1908.
17 The values t0.10, t0.05, t0.025, t0.01, and t0.005 are also referred to as one-sided critical values of t at the 0.10, 0.05, 0.025, 0.01, and 0.005 significance levels, for
the specified number of degrees of freedom.
18 A formula exists for determining the sample size needed to obtain a desired width for a confidence interval. Define E = Reliability factor × Standard error. The smaller E is, the smaller the width of the confidence interval, because 2E is the confidence interval’s width. The sample size to obtain a desired value of E at a given degree of confidence (1 − α) is n = [(tα/2s)/E]2.
19 We assumed in this example that sample size is sufficiently small compared with the size of the client base that we can disregard the finite population correction factor (mentioned in Footnote 6).
20 Some researchers use the term “data snooping” instead of data mining.
21 To convey an understanding of data mining, it is very helpful to introduce some basic concepts related to hypothesis testing. The reading on hypothesis testing contains further discussion of significance levels and tests of significance.
22 In terms of our previous discussion of confidence intervals, significance at the 5 percent level corresponds to a hypothesized value for a population statistic falling outside a 95 percent confidence interval based on an appropriate sample statistic (e.g., the sample mean, when the hypothesis concerns the population mean).
23 The term “intergenerational” comes from viewing each round of researchers as a generation. Campbell, Lo, and MacKinlay (1997) have called intergenerational data mining “data snooping.” The latter phrase, however, is commonly used as a synonym of data mining; thus McQueen and Thorley’s terminology is less ambiguous. The term “intragenerational data mining” is available when we want to highlight that the reference is to an investigator ’s new or independent data mining.
24 For example, Lo and MacKinlay (1990) concluded that the magnitude of this type of bias on tests of the capital asset pricing model was considerable.
25 The Dow Dividend Strategy, also known as Dogs of the Dow Strategy, consists of holding an equally weighted portfolio of the 10 highest-yielding DJIA stocks as of the beginning of a year. At the time of McQueen and Thorley’s research, the Foolish Four strategy was as follows: At the beginning of each year, the Foolish Four portfolio purchases a 4-stock portfolio from the 5 lowest-priced stocks of the 10 highest-yielding DJIA stocks. The lowest-priced stock of the five is excluded, and 40 percent is invested in the second-to-lowest-priced stock, with 20
percent weights in the remaining three.
26 In the finance literature, such a random but irrelevant-to-the-future pattern is sometimes called an artifact of the dataset.
27 See Fama and French (1996, p. 80) for discussion of data snooping and survivorship bias in their tests.
28 Delistings occur for a variety of reasons: merger, bankruptcy, liquidation, or migration to another exchange.
29 See Fung and Hsieh (2002) and ter Horst and Verbeek (2007) for more details on the problems of interpreting hedge fund performance. Note that an offsetting type of bias may occur if successful funds stop reporting performance because they no longer want new cash inflows.
CHAPTER 7 Hypothesis Testing Richard A. DeFusco, CFA
Dennis W. McLeavey, CFA
Jerald E. Pinto, PhD, CFA
David E. Runkle, PhD, CFA
LEARNING OUTCOMES
After completing this chapter, you will be able to do the following:
define a hypothesis, describe the steps of hypothesis testing, and describe and interpret the choice of the null and alternative hypotheses;
distinguish between one-tailed and two-tailed tests of hypotheses;
explain a test statistic, Type I and Type II errors, a significance level, and how significance levels are used in hypothesis testing;
explain a decision rule, the power of a test, and the relation between confidence intervals and hypothesis tests;
distinguish between a statistical result and an economically meaningful result;
explain and interpret the p-value as it relates to hypothesis testing;
identify the appropriate test statistic and interpret the results for a hypothesis test concerning the population mean of both large and small samples when the population is normally or approximately normally distributed and the variance is 1) known or 2) unknown;
identify the appropriate test statistic and interpret the results for a hypothesis test concerning the equality of the population means of two at least approximately normally distributed populations, based on independent random samples with 1) equal or 2) unequal assumed variances;
identify the appropriate test statistic and interpret the results for a hypothesis test concerning the mean difference of two normally distributed populations;
identify the appropriate test statistic and interpret the results for a hypothesis
test concerning 1) the variance of a normally distributed population, and 2) the equality of the variances of two normally distributed populations based on two independent random samples;
distinguish between parametric and nonparametric tests and describe situations in which the use of nonparametric tests may be appropriate.
1. Introduction Analysts often confront competing ideas about how financial markets work. Some of these ideas develop through personal research or experience with markets; others come from interactions with colleagues; and many others appear in the professional literature on finance and investments. In general, how can an analyst decide whether statements about the financial world are probably true or probably false?
When we can reduce an idea or assertion to a definite statement about the value of a quantity, such as an underlying or population mean, the idea becomes a statistically testable statement or hypothesis. The analyst may want to explore questions such as the following:
Is the underlying mean return on this mutual fund different from the underlying mean return on its benchmark?
Did the volatility of returns on this stock change after the stock was added to a stock market index?
Are a security’s bid-ask spreads related to the number of dealers making a market in the security?
Do data from a national bond market support a prediction of an economic theory about the term structure of interest rates (the relationship between yield and maturity)?
To address these questions, we use the concepts and tools of hypothesis testing. Hypothesis testing is part of statistical inference, the process of making judgments about a larger group (a population) on the basis of a smaller group actually observed (a sample). The concepts and tools of hypothesis testing provide an objective means to gauge whether the available evidence supports the hypothesis. After a statistical test of a hypothesis we should have a clearer idea of the probability that a hypothesis is true or not, although our conclusion always stops short of certainty. Hypothesis testing has been a powerful tool in the advancement of investment knowledge and science. As Robert L. Kahn of the Institute for Social Research (Ann Arbor, Michigan) has written, “The mill of science grinds only when hypothesis and data are in continuous and abrasive contact.”
The main emphases of this reading are the framework of hypothesis testing and tests concerning mean and variance, two quantities frequently used in investments. We give an overview of the procedure of hypothesis testing in the next section. We then address testing hypotheses about the mean and hypotheses about the differences
between means. In the fourth section of this reading, we address testing hypotheses about a single variance and hypotheses about the differences between variances. We end the reading with an overview of some other important issues and techniques in statistical inference.
2. Hypothesis Testing Hypothesis testing, as we have mentioned, is part of the branch of statistics known as statistical inference. Traditionally, the field of statistical inference has two subdivisions: estimation and hypothesis testing. Estimation addresses the question “What is this parameter ’s (e.g., the population mean’s) value?” The answer is in the form of a confidence interval built around a point estimate. Take the case of the mean: We build a confidence interval for the population mean around the sample mean as a point estimate. For the sake of specificity, suppose the sample mean is 50 and a 95 percent confidence interval for the population mean is 50 ± 10 (the confidence interval runs from 40 to 60). If this confidence interval has been properly constructed, there is a 95 percent probability that the interval from 40 to 60 contains the population mean’s value.1 The second branch of statistical inference, hypothesis testing, has a somewhat different focus. A hypothesis testing question is “Is the value of the parameter (say, the population mean) 45 (or some other specific value)?” The assertion “the population mean is 45” is a hypothesis. A hypothesis is defined as a statement about one or more populations.
This section focuses on the concepts of hypothesis testing. The process of hypothesis testing is part of a rigorous approach to acquiring knowledge known as the scientific method. The scientific method starts with observation and the formulation of a theory to organize and explain observations. We judge the correctness of the theory by its ability to make accurate predictions—for example, to predict the results of new observations.2 If the predictions are correct, we continue to maintain the theory as a possibly correct explanation of our observations. When risk plays a role in the outcomes of observations, as in finance, we can only try to make unbiased, probability-based judgments about whether the new data support the predictions. Statistical hypothesis testing fills that key role of testing hypotheses when chance plays a role. In an analyst’s day-to-day work, he may address questions to which he might give answers of varying quality. When an analyst correctly formulates the question into a testable hypothesis and carries out and reports on a hypothesis test, he has provided an element of support to his answer consistent with the standards of the scientific method. Of course, the analyst’s logic, economic reasoning, information sources, and perhaps other factors also play a role in our assessment of the answer ’s quality.3
We organize this introduction to hypothesis testing around the following list of seven steps.
Steps in Hypothesis Testing. The steps in testing a hypothesis are as follows:4
1. Stating the hypotheses.
2. Identifying the appropriate test statistic and its probability distribution.
3. Specifying the significance level.
4. Stating the decision rule.
5. Collecting the data and calculating the test statistic.
6. Making the statistical decision.
7. Making the economic or investment decision.
We will explain each of these steps using as illustration a hypothesis test concerning the sign of the risk premium on Canadian stocks. The steps above constitute a traditional approach to hypothesis testing. We will end the section with a frequently used alternative to those steps, the p-value approach.
The first step in hypothesis testing is stating the hypotheses. We always state two hypotheses: the null hypothesis (or null), designated H0, and the alternative hypothesis, designated Ha.
Definition of Null Hypothesis. The null hypothesis is the hypothesis to be tested. For example, we could hypothesize that the population mean risk premium for Canadian equities is less than or equal to zero.
The null hypothesis is a proposition that is considered true unless the sample we use to conduct the hypothesis test gives convincing evidence that the null hypothesis is false. When such evidence is present, we are led to the alternative hypothesis.
Definition of Alternative Hypothesis. The alternative hypothesis is the hypothesis accepted when the null hypothesis is rejected. Our alternative hypothesis is that the population mean risk premium for Canadian equities is greater than zero.
Suppose our question concerns the value of a population parameter, θ, in relation to one possible value of the parameter, θ0 (these are read, respectively, “theta” and “theta sub zero”).5 Examples of a population parameter include the population mean, μ, and the population variance, σ2. We can formulate three different sets of hypotheses, which we label according to the assertion made by the alternative hypothesis.
Formulations of Hypotheses. We can formulate the null and alternative hypotheses in three different ways:
1. H0: θ = θ0 versus Ha: θ | θ0 (a “not equal to” alternative hypothesis)
2. H0: θ ≤ θ0 versus Ha: θ > θ0 (a “greater than” alternative hypothesis)
3. H0: θ ≥ θ0 versus Ha: θ < θ0 (a “less than” alternative hypothesis)
In our Canadian example, θ = μRP and represents the population mean risk premium on Canadian equities. Also, θ0 = 0 and we are using the second of the above three formulations.
The first formulation is a two-sided hypothesis test (or two-tailed hypothesis test): We reject the null in favor of the alternative if the evidence indicates that the population parameter is either smaller or larger than θ0. In contrast, Formulations 2 and 3 are each a one-sided hypothesis test (or one-tailed hypothesis test). For Formulations 2 and 3, we reject the null only if the evidence indicates that the population parameter is respectively greater than or less than θ0. The alternative hypothesis has one side.
Notice that in each case above, we state the null and alternative hypotheses such that they account for all possible values of the parameter. With Formulation 1, for example, the parameter is either equal to the hypothesized value θ0 (under the null hypothesis) or not equal to the hypothesized value θ0 (under the alternative hypothesis). Those two statements logically exhaust all possible values of the parameter.
Despite the different ways to formulate hypotheses, we always conduct a test of the null hypothesis at the point of equality, θ = θ0. Whether the null is H0: θ = θ0, H0: θ ≤ θ0, or H0: θ ≥ θ0, we actually test θ = θ0. The reasoning is straightforward. Suppose the hypothesized value of the parameter is 5. Consider H0: θ ≤ 5, with a “greater than” alternative hypothesis, Ha: θ > 5. If we have enough evidence to reject H0: θ = 5 in favor of Ha: θ > 5, we definitely also have enough evidence to reject the hypothesis that the parameter, θ, is some smaller value, such as 4.5 or 4. To review, the calculation to test the null hypothesis is the same for all three formulations. What is different for the three formulations, we will see shortly, is how the calculation is evaluated to decide whether or not to reject the null.
How do we choose the null and alternative hypotheses? Probably most common are “not equal to” alternative hypotheses. We reject the null because the evidence indicates that the parameter is either larger or smaller than θ0. Sometimes, however,
we may have a “suspected” or “hoped for” condition for which we want to find supportive evidence.6 In that case, we can formulate the alternative hypothesis as the statement that this condition is true; the null hypothesis that we test is the statement that this condition is not true. If the evidence supports rejecting the null and accepting the alternative, we have statistically confirmed what we thought was true. For example, economic theory suggests that investors require a positive risk premium on stocks (the risk premium is defined as the expected return on stocks minus the risk-free rate). Following the principle of stating the alternative as the “hoped for” condition, we formulate the following hypotheses:
Note that “greater than” and “less than” alternative hypotheses reflect the beliefs of the researcher more strongly than a “not equal to” alternative hypothesis. To emphasize an attitude of neutrality, the researcher may sometimes select a “not equal to” alternative hypothesis when a one-sided alternative hypothesis is also reasonable.
The second step in hypothesis testing is identifying the appropriate test statistic and its probability distribution.
Definition of Test Statistic. A test statistic is a quantity, calculated based on a sample, whose value is the basis for deciding whether or not to reject the null hypothesis.
The focal point of our statistical decision is the value of the test statistic. Frequently (in all the cases that we examine in this reading), the test statistic has the form
(1)
For our risk premium example, the population parameter of interest is the population mean risk premium, μRP. We label the hypothesized value of the population mean under H0 as μ0. Restating the hypotheses using symbols, we test H0: μRP ≤ μ0 versus Ha: μRP > μ0. However, because under the null we are testing μ0 = 0, we write H0: μRP ≤ 0 versus Ha: μRP > 0.
The sample mean provides an estimate of the population mean. Therefore, we can use the sample mean risk premium calculated from historical data, RP, as the sample statistic in Equation 1. The standard deviation of the sample statistic, known as the “standard error” of the statistic, is the denominator in Equation 1. For this
example, the sample statistic is a sample mean. For a sample mean, , calculated from a sample generated by a population with standard deviation σ, the standard error is given by one of two expressions:
(2)
when we know σ (the population standard deviation), or
(3)
when we do not know the population standard deviation and need to use the sample standard deviation s to estimate it. For this example, because we do not know the population standard deviation of the process generating the return, we use Equation 3. The test statistic is thus
In making the substitution of 0 for μ0, we use the fact already highlighted that we test any null hypothesis at the point of equality, as well as the fact that μ0 = 0 here.
We have identified a test statistic to test the null hypothesis. What probability distribution does it follow? We will encounter four distributions for test statistics in this reading:
the t-distribution (for a t-test);
the standard normal or z-distribution (for a z-test);
the chi-square (χ2) distribution (for a chi-square test); and
the F-distribution (for an F-test).
We will discuss the details later, but assume we can conduct a z-test based on the central limit theorem because our Canadian sample has many observations.7 To summarize, the test statistic for the hypothesis test concerning the mean risk
premium is . We can conduct a z-test because we can plausibly assume that the test statistic follows a standard normal distribution.
The third step in hypothesis testing is specifying the significance level. When the test statistic has been calculated, two actions are possible: 1) We reject the null hypothesis or 2) we do not reject the null hypothesis. The action we take is based on comparing the calculated test statistic to a specified possible value or values. The
comparison values we choose are based on the level of significance selected. The level of significance reflects how much sample evidence we require to reject the null. Analogous to its counterpart in a court of law, the required standard of proof can change according to the nature of the hypotheses and the seriousness of the consequences of making a mistake. There are four possible outcomes when we test a null hypothesis: 1. We reject a false null hypothesis. This is a correct decision.
2. We reject a true null hypothesis. This is called a Type I error.
3. We do not reject a false null hypothesis. This is called a Type II error.
4. We do not reject a true null hypothesis. This is a correct decision.
TABLE 1 Type I and Type II Errors in Hypothesis Testing
True Situation Decision H0 True H0 False
Do not reject H0 Correct Decision Type II Error Reject H0 (accept Ha) Type I Error Correct Decision
We illustrate these outcomes in Table 1.
When we make a decision in a hypothesis test, we run the risk of making either a Type I or a Type II error. These are mutually exclusive errors: If we mistakenly reject the null, we can only be making a Type I error; if we mistakenly fail to reject the null, we can only be making a Type II error.
The probability of a Type I error in testing a hypothesis is denoted by the Greek letter alpha, α. This probability is also known as the level of significance of the test. For example, a level of significance of 0.05 for a test means that there is a 5 percent probability of rejecting a true null hypothesis. The probability of a Type II error is denoted by the Greek letter beta, ß.
Controlling the probabilities of the two types of errors involves a trade-off. All else equal, if we decrease the probability of a Type I error by specifying a smaller significance level (say 0.01 rather than 0.05), we increase the probability of making a Type II error because we will reject the null less frequently, including when it is false. The only way to reduce the probabilities of both types of errors simultaneously is to increase the sample size, n.
Quantifying the trade-off between the two types of error in practice is usually impossible because the probability of a Type II error is itself hard to quantify. Consider H0: θ ≤ 5 versus Ha: θ > 5. Because every true value of θ greater than 5 makes the null hypothesis false, each value of θ greater than 5 has a different ß (Type II error probability). In contrast, it is sufficient to state a Type I error probability for θ = 5, the point at which we conduct the test of the null hypothesis. Thus, in general, we specify only α, the probability of a Type I error, when we conduct a hypothesis test. Whereas the significance level of a test is the probability of incorrectly rejecting the null, the power of a test is the probability of correctly rejecting the null—that is, the probability of rejecting the null when it is false.8 When more than one test statistic is available to conduct a hypothesis test, we should prefer the most powerful, all else equal.9
To summarize, the standard approach to hypothesis testing involves specifying a level of significance (probability of Type I error) only. It is most appropriate to specify this significance level prior to calculating the test statistic. If we specify it after calculating the test statistic, we may be influenced by the result of the calculation, which detracts from the objectivity of the test.
We can use three conventional significance levels to conduct hypothesis tests: 0.10, 0.05, and 0.01. Qualitatively, if we can reject a null hypothesis at the 0.10 level of significance, we have some evidence that the null hypothesis is false. If we can reject a null hypothesis at the 0.05 level, we have strong evidence that the null hypothesis is false. And if we can reject a null hypothesis at the 0.01 level, we have very strong evidence that the null hypothesis is false. For the risk premium example, we will specify a 0.05 significance level.
The fourth step in hypothesis testing is stating the decision rule. The general principle is simply stated. When we test the null hypothesis, if we find that the calculated value of the test statistic is as extreme or more extreme than a given value or values determined by the specified level of significance, α, we reject the null hypothesis. We say the result is statistically significant. Otherwise, we do not reject the null hypothesis and we say the result is not statistically significant. The value or values with which we compare the calculated test statistic to make our decision are the rejection points (critical values) for the test.10
Definition of a Rejection Point (Critical Value) for the Test Statistic. A rejection point (critical value) for a test statistic is a value with which the computed test statistic is compared to decide whether to reject or not reject the null hypothesis.
For a one-tailed test, we indicate a rejection point using the symbol for the test statistic with a subscript indicating the specified probability of a Type I error, α; for example, zα. For a two-tailed test, we indicate zα/2. To illustrate the use of rejection points, suppose we are using a z-test and have chosen a 0.05 level of significance.
For a test of H0: θ = θ0 versus Ha: θ | θ0, two rejection points exist, one negative and one positive. For a two-sided test at the 0.05 level, the total probability of a Type I error must sum to 0.05. Thus, 0.05/2 = 0.025 of the probability should be in each tail of the distribution of the test statistic under the null. Consequently, the two rejection points are z0.025 = 1.96 and −z0.025 = −1.96. Let z represent the calculated value of the test statistic. We reject the null if we find that z < −1.96 or z > 1.96. We do not reject if −1.96 ≤ z ≤ 1.96.
For a test of H0: θ ≤ θ0 versus Ha: θ > θ0 at the 0.05 level of significance, the rejection point is z0.05 = 1.645. We reject the null hypothesis if z > 1.645. The value of the standard normal distribution such that 5 percent of the outcomes lie to the right is z0.05 = 1.645.
For a test of H0: θ ≥ θ0 versus Ha: θ < θ0, the rejection point is −z0.05 = −1.645. We reject the null hypothesis if z < −1.645.
Figure 1 illustrates a test H0: μ = μ0 versus Ha: μ | μ0 at the 0.05 significance level using a z-test. The “acceptance region” is the traditional name for the set of values of the test statistic for which we do not reject the null hypothesis. (The traditional name, however, is inaccurate. We should avoid using phrases such as “accept the null hypothesis” because such a statement implies a greater degree of conviction about the null than is warranted when we fail to reject it.)11 On either side of the acceptance region is a rejection region (or critical region). If the null hypothesis that μ = μ0 is true, the test statistic has a 2.5 percent chance of falling in the left rejection region and a 2.5 percent chance of falling in the right rejection region. Any calculated value of the test statistic that falls in either of these two regions causes us to reject the null hypothesis at the 0.05 significance level. The rejection points of 1.96 and −1.96 are seen to be the dividing lines between the acceptance and rejection regions.
FIGURE 1 Rejection Points (Critical Values), 0.05 Significance Level, Two-Sided Test of the Population Mean Using a z-Test
Figure 1 affords a good opportunity to highlight the relationship between confidence intervals and hypothesis tests. A 95 percent confidence interval for the population mean, μ, based on sample mean, , is given by to , where is the standard error of the sample mean (Equation 3).12
Now consider one of the conditions for rejecting the null hypothesis:
Here, μ0 is the hypothesized value of the population mean. The condition states that rejection is warranted if the test statistic exceeds 1.96. Multiplying both sides by , we have , or after rearranging, , which we can also write as
. This expression says that if the hypothesized population mean, μ0, is less than the lower limit of the 95 percent confidence interval based on the sample mean, we must reject the null hypothesis at the 5 percent significance level (the test statistic falls in the rejection region to the right).
Now, we can take the other condition for rejecting the null hypothesis:
and, using algebra as before, rewrite it as . If the hypothesized population mean is larger than the upper limit of the 95 percent confidence interval, we reject the null hypothesis at the 5 percent level (the test statistic falls in the
rejection region to the left). Thus, an α significance level in a two-sided hypothesis test can be interpreted in exactly the same way as a (1 − α) confidence interval.
In summary, when the hypothesized value of the population parameter under the null is outside the corresponding confidence interval, the null hypothesis is rejected. We could use confidence intervals to test hypotheses; practitioners, however, usually do not. Computing a test statistic (one number, versus two numbers for the usual confidence interval) is more efficient. Also, analysts encounter actual cases of one- sided confidence intervals only rarely. Furthermore, only when we compute a test statistic can we obtain a p-value, a useful quantity relating to the significance of our results (we will discuss p-values shortly).
To return to our risk premium test, we stated hypotheses H0: μRP ≤ 0 versus Ha: μRP > 0. We identified the test statistic as and stated that it follows a standard normal distribution. We are, therefore, conducting a one-sided z-test. We specified a 0.05 significance level. For this one-sided z-test, the rejection point at the 0.05 level of significance is 1.645. We will reject the null if the calculated z-statistic is larger than 1.645. Figure 2 illustrates this test.
The fifth step in hypothesis testing is collecting the data and calculating the test statistic. The quality of our conclusions depends not only on the appropriateness of the statistical model but also on the quality of the data we use in conducting the test. We first need to check for measurement errors in the recorded data. Some other issues to be aware of include sample selection bias and time-period bias. Sample selection bias refers to bias introduced by systematically excluding some members of the population according to a particular attribute. One type of sample selection bias is survivorship bias. For example, if we define our sample as US bond mutual funds currently operating and we collect returns for just these funds, we will systematically exclude funds that have not survived to the present date. Nonsurviving funds are likely to have underperformed surviving funds, on average; as a result the performance reflected in the sample may be biased upward. Time- period bias refers to the possibility that when we use a time-series sample, our statistical conclusion may be sensitive to the starting and ending dates of the sample.13
FIGURE 2 Rejection Point (Critical Value), 0.05 Significance Level, One-Sided Test of the Population Mean Using a z-Test
To continue with the risk premium hypothesis, we focus on Canadian equities. According to Dimson, Marsh, and Staunton (2011) for the period 1900 to 2010 inclusive (111 annual observations), the arithmetic mean equity risk premium for Canadian stocks relative to bond returns, RP, was 5.3 percent per year. The sample standard deviation of the annual risk premiums was 18.2 percent. Using Equation 3, the standard error of the sample mean is = 18.2%/ = 1.727%. The test statistic is = 5.3%/1.727% = 3.07.
The sixth step in hypothesis testing is making the statistical decision. For our example, because the test statistic z = 3.07 is larger than the rejection point of 1.645, we reject the null hypothesis in favor of the alternative hypothesis that the risk premium on Canadian stocks is positive. The first six steps are the statistical steps. The final decision concerns our use of the statistical decision.
The seventh and final step in hypothesis testing is making the economic or investment decision. The economic or investment decision takes into consideration not only the statistical decision but also all pertinent economic issues. In the sixth step, we found strong statistical evidence that the Canadian risk premium is positive. The magnitude of the estimated risk premium, 5.3 percent a year, is economically very meaningful as well. Based on these considerations, an investor might decide to commit funds to Canadian equities. A range of nonstatistical considerations, such as the investor ’s tolerance for risk and financial position, might also enter the decision-making process.
The preceding discussion raises an issue that often arises in this decision-making
step. We frequently find that slight differences between a variable and its hypothesized value are statistically significant but not economically meaningful. For example, we may be testing an investment strategy and reject a null hypothesis that the mean return to the strategy is zero based on a large sample. Equation 1 shows that the smaller the standard error of the sample statistic (the divisor in the formula), the larger the value of the test statistic and the greater the chance the null will be rejected, all else equal. The standard error decreases as the sample size, n, increases, so that for very large samples, we can reject the null for small departures from it. We may find that although a strategy provides a statistically significant positive mean return, the results are not economically significant when we account for transaction costs, taxes, and risk. Even if we conclude that a strategy’s results are economically meaningful, we should explore the logic of why the strategy might work in the future before actually implementing it. Such considerations cannot be incorporated into a hypothesis test.
Before leaving the subject of the process of hypothesis testing, we should discuss an important alternative approach called the p-value approach to hypothesis testing. Analysts and researchers often report the p-value (also called the marginal significance level) associated with hypothesis tests.
Definition of p-Value. The p-value is the smallest level of significance at which the null hypothesis can be rejected.
For the value of the test statistic of 3.07 in the risk premium hypothesis test, using a spreadsheet function for the standard normal distribution, we calculate a p-value of 0.00107. We can reject the null hypothesis at that level of significance. The smaller the p-value, the stronger the evidence against the null hypothesis and in favor of the alternative hypothesis. The p-value for a two-sided test that a parameter equals zero is frequently generated automatically by statistical and econometric software programs.14
We can use p-values in the hypothesis testing framework presented above as an alternative to using rejection points. If the p-value is less than our specified level of significance, we reject the null hypothesis. Otherwise, we do not reject the null hypothesis. Using the p-value in this fashion, we reach the same conclusion as we do using rejection points. For example, because 0.00107 is less than 0.05, we would reject the null hypothesis in the risk premium test. The p-value, however, provides more precise information on the strength of the evidence than does the rejection points approach. The p-value of 0.00107 indicates that the null is rejected at a far smaller level of significance than 0.05.
If one researcher examines a question using a 0.05 significance level and another
researcher uses a 0.01 significance level, the reader may have trouble comparing the findings. This concern has given rise to an approach to presenting the results of hypothesis tests that features p-values and omits specification of the significance level (Step 3). The interpretation of the statistical results is left to the consumer of the research. This has sometimes been called the p-value approach to hypothesis testing.15
3. Hypothesis Tests Concerning the Mean Hypothesis tests concerning the mean are among the most common in practice. In this section we discuss such tests for several distinct types of problems. In one type (discussed in Section 3.1), we test whether the population mean of a single population is equal to (or greater or less than) some hypothesized value. Then, in Sections 3.2 and 3.3, we address inference on means based on two samples. Is an observed difference between two sample means due to chance or different underlying (population) means? When we have two random samples that are independent of each other—no relationship exists between the measurements in one sample and the measurements in the other—the techniques of Section 3.2 apply. When the samples are dependent, the methods of Section 3.3 are appropriate.16
3.1. Tests Concerning a Single Mean An analyst who wants to test a hypothesis concerning the value of an underlying or population mean will conduct a t-test in the great majority of cases. A t-test is a hypothesis test using a statistic (t-statistic) that follows a t-distribution. The t- distribution is a probability distribution defined by a single parameter known as degrees of freedom (df). Each value of degrees of freedom defines one distribution in this family of distributions. The t-distribution is closely related to the standard normal distribution. Like the standard normal distribution, a t-distribution is symmetrical with a mean of zero. However, the t-distribution is more spread out: It has a standard deviation greater than 1 (compared to 1 for the standard normal)17 and more probability for outcomes distant from the mean (it has fatter tails than the standard normal distribution). As the number of degrees of freedom increases with sample size, the spread decreases and the t-distribution approaches the standard normal distribution as a limit.
Why is the t-distribution the focus for the hypothesis tests of this section? In practice, investment analysts need to estimate the population standard deviation by calculating a sample standard deviation. That is, the population variance (or standard deviation) is unknown. For hypothesis tests concerning the population mean of a normally distributed population with unknown variance, the theoretically correct test statistic is the t-statistic. What if a normal distribution does not describe the population? The t-test is robust to moderate departures from normality, except for outliers and strong skewness.18 When we have large samples, departures of the underlying distribution from the normal are of increasingly less concern. The sample mean is approximately normally distributed in large samples according to the central limit theorem, whatever the distribution describing the population. In
general, a sample size of 30 or more usually can be treated as a large sample and a sample size of 29 or less is treated as a small sample.19
Test Statistic for Hypothesis Tests of the Population Mean (Practical Case —Population Variance Unknown). If the population sampled has unknown variance and either of the conditions below holds:
1. the sample is large, or
2. the sample is small but the population sampled is normally distributed, or approximately normally distributed,
then the test statistic for hypothesis tests concerning a single population mean, μ, is
(4)
where
tn-1 = t-statistic with n − 1 degrees of freedom (n is the sample size)
= the sample mean
μ0 = the hypothesized value of the population mean
s = the sample standard deviation
The denominator of the t-statistic is an estimate of the sample mean standard error, .20
In Example 1, because the sample size is small, the test is called a small sample test concerning the population mean.
EXAMPLE 1 Risk and Return Characteristics of an Equity Mutual Fund (1)
You are analyzing Sendar Equity Fund, a mid-cap growth fund that has been in existence for 24 months. During this period, it has achieved a mean monthly return of 1.50 percent with a sample standard deviation of monthly returns of 3.60 percent. Given its level of systematic (market) risk and according to a pricing model, this mutual fund was expected to have earned a 1.10 percent mean monthly return during that time period. Assuming returns are normally distributed, are the actual results consistent with an underlying or population
mean monthly return of 1.10 percent? 1. Formulate null and alternative hypotheses consistent with the verbal description
of the research goal.
2. Identify the test statistic for conducting a test of the hypotheses in Part 1.
3. Identify the rejection point or points for the hypothesis tested in Part 1 at the 0.10 level of significance.
4. Determine whether the null hypothesis is rejected or not rejected at the 0.10 level of significance. (Use the tables in the back of this book.)
Solution to 1: We have a “not equal to” alternative hypothesis, where μ is the underlying mean return on Sendar Equity Fund—H0: μ = 1.10 versus Ha: μ | 1.10.
Solution to 2: Because the population variance is not known, we use a t-test with 24 − 1 = 23 degrees of freedom.
Solution to 3: Because this is a two-tailed test, we have the rejection point tα/2,n −1 = t0.05,23. In the table for the t-distribution, we look across the row for 23 degrees of freedom to the 0.05 column, to find 1.714. The two rejection points for this two-sided test are 1.714 and −1.714. We will reject the null if we find that t > 1.714 or t < −1.714.
Solution to 4:
Because 0.544 does not satisfy either t > 1.714 or t < −1.714, we do not reject the null hypothesis.
The confidence interval approach provides another perspective on this hypothesis test. The theoretically correct 100(1 − α)% confidence interval for the population mean of a normal distribution with unknown variance, based on a sample of size n, is
where tα/2 is the value of t such that α/2 of the probability remains in the right tail and where −tα/2 is the value of t such that α/2 of the probability remains in the left tail, for n − 1 degrees of freedom. Here, the 90 percent confidence interval runs from 1.5 − (1.714)(0.734847) = 0.240 to 1.5 + (1.714)(0.734847) =
2.760, compactly [0.240, 2.760]. The hypothesized value of mean return, 1.10, falls within this confidence interval, and we see from this perspective also that the null hypothesis is not rejected. At a 10 percent level of significance, we conclude that a population mean monthly return of 1.10 percent is consistent with the 24-month observed data series. Note that 10 percent is a relatively high probability of rejecting the hypothesis of a 1.10 percent population mean monthly return when it is true.
EXAMPLE 2 A Slowdown in Payments of Receivables
FashionDesigns, a supplier of casual clothing to retail chains, is concerned about a possible slowdown in payments from its customers. The controller ’s office measures the rate of payment by the average number of days in receivables.21 FashionDesigns has generally maintained an average of 45 days in receivables. Because it would be too costly to analyze all of the company’s receivables frequently, the controller ’s office uses sampling to track customers’ payment rates. A random sample of 50 accounts shows a mean number of days in receivables of 49 with a standard deviation of 8 days.
1. Formulate null and alternative hypotheses consistent with determining whether
the evidence supports the suspected condition that customer payments have slowed.
2. Identify the test statistic for conducting a test of the hypotheses in Part 1.
3. Identify the rejection point or points for the hypothesis tested in Part 1 at the 0.05 and 0.01 levels of significance.
4. Determine whether the null hypothesis is rejected or not rejected at the 0.05 and 0.01 levels of significance.
Solution to 1: The suspected condition is that the number of days in receivables has increased relative to the historical rate of 45 days, which suggests a “greater than” alternative hypothesis. With μ as the population mean number of days in receivables, the hypotheses are H0: μ ≤ 45 versus Ha: μ > 45.
Solution to 2: Because the population variance is not known, we use a t-test with 50 − 1 = 49 degrees of freedom.
Solution to 3: The rejection point is found across the row for degrees of freedom of 49. To find the one-tailed rejection point for a 0.05 significance
level, we use the 0.05 column: The value is 1.677. To find the one-tailed rejection point for a 0.01 level of significance, we use the 0.01 column: The value is 2.405. To summarize, at a 0.05 significance level, we reject the null if we find that t > 1.677; at a 0.01 significance level, we reject the null if we find that t > 2.405.
Solution to 4:
Because 3.536 > 1.677, the null hypothesis is rejected at the 0.05 level. Because 3.536 > 2.405, the null hypothesis is also rejected at the 0.01 level. We can say with a high level of confidence that FashionDesigns has experienced a slowdown in customer payments. The level of significance, 0.01, is a relatively low probability of rejecting the hypothesized mean of 45 days or less. Rejection gives us confidence that the mean has increased above 45 days.
We stated above that when population variance is not known, we use a t-test for tests concerning a single population mean. Given at least approximate normality, the t- test is always called for when we deal with small samples and do not know the population variance. For large samples, the central limit theorem states that the sample mean is approximately normally distributed, whatever the distribution of the population. So the t-test is still appropriate, but an alternative test may be more useful when sample size is large.
For large samples, practitioners sometimes use a z-test in place of a t-test for tests concerning a mean.22 The justification for using the z-test in this context is twofold. First, in large samples, the sample mean should follow the normal distribution at least approximately, as we have already stated, fulfilling the normality assumption of the z-test. Second, the difference between the rejection points for the t-test and z- test becomes quite small when sample size is large. For a two-sided test at the 0.05 level of significance, the rejection points for a z-test are 1.96 and −1.96. For a t-test, the rejection points are 2.045 and −2.045 for df = 29 (about a 4 percent difference between the z and t rejection points) and 2.009 and −2.009 for df = 50 (about a 2.5 percent difference between the z and t rejection points). Because the t-test is readily available as statistical program output and theoretically correct for unknown population variance, we present it as the test of choice.
In a very limited number of cases, we may know the population variance; in such cases, the z-test is theoretically correct.23
The z-Test Alternative.
1. If the population sampled is normally distributed with known variance σ2, then the test statistic for a hypothesis test concerning a single population mean, μ, is
(5)
2. If the population sampled has unknown variance and the sample is large, in place of a t-test, an alternative test statistic (relying on the central limit theorem) is
(6)
In the above equations,
σ = the known population standard deviation
s = the sample standard deviation
μ0 = the hypothesized value of the population mean
When we use a z-test, we most frequently refer to a rejection point in the list below.
Rejection Points for a z-Test.
Significance level of α = 0.10.
1. H0: θ = θ0 versus Ha: θ | θ0. The rejection points are z0.05 = 1.645 and −z0.05 = −1.645.
Reject the null hypothesis if z > 1.645 or if z < −1.645.
2. H0: θ ≤ θ0 versus Ha: θ > θ0. The rejection point is z0.10 = 1.28.
Reject the null hypothesis if z > 1.28.
3. H0: θ ≥ θ0 versus Ha: θ < θ0. The rejection point is −z0.10 = −1.28.
Reject the null hypothesis if z < −1.28.
Significance level of α = 0.05.
1. H0: θ = θ0 versus Ha: θ | θ0. The rejection points are z0.025 = 1.96 and −z0.025 = −1.96.
Reject the null hypothesis if z > 1.96 or if z < −1.96.
2. H0: θ ≤ θ0 versus Ha: θ > θ0. The rejection point is z0.05 = 1.645.
Reject the null hypothesis if z > 1.645.
3. H0: θ ≥ θ0 versus Ha: θ < θ0. The rejection point is −z0.05 = −1.645.
Reject the null hypothesis if z < −1.645.
Significance level of α = 0.01.
1. H0: θ = θ0 versus Ha: θ | θ0. The rejection points are z0.005 = 2.575 and −z0.005 = −2.575.
Reject the null hypothesis if z > 2.575 or if z < −2.575.
2. H0: θ ≤ θ0 versus Ha: θ > θ0. The rejection point is z0.01 = 2.33.
Reject the null hypothesis if z > 2.33.
3. H0: θ ≥ θ0 versus Ha: θ < θ0. The rejection point is −z0.01 = −2.33.
Reject the null hypothesis if z < −2.33.
EXAMPLE 3 The Effect of Control Deficiency Disclosures under the Sarbanes–Oxley Act on Share Prices
The Sarbanes–Oxley Act came into effect in 2002 and introduced major changes to the regulation of corporate governance and financial practice in the United States. One of the requirements of this Act is for firms to periodically assess and report certain types of internal control deficiencies to the audit committee, external auditors, and to the Securities and Exchange Commission (SEC). When a company makes an internal control weakness disclosure, does it convey information that affects the market value of the firm’s stock?
Gupta and Nayar (2007) addressed this question by studying a number of voluntary disclosures made in the very early days of Sarbanes–Oxley implementation. Their final sample for this study consisted of 90 firms that had made control deficiency disclosures to the SEC from March 2003 to July 2004. This 90-firm sample was termed the “full sample”. These firms were further examined to see if there were any other contemporaneous announcements, such
as earnings announcements, associated with the control deficiency disclosures. Of the 90 firms, 45 did not have any such confounding announcements, and the sample of these firms was termed the “clean sample.”
The announcement day of the internal control weakness was designated t = 0. If these announcements provide new information useful for equity valuation, the information should cause a change in stock prices and returns once it is available. Only one component of stock returns is of interest: the return in excess of that predicted given a stock’s market risk or beta, called the abnormal return. Significant negative (positive) abnormal returns indicate that investors perceive unfavorable (favorable) corporate news in the internal control weakness announcement. Although Gupta and Nayar examined abnormal returns for various time horizons or event windows, we report a selection of their findings for the window [0, +1], which includes a two-day period of the day of and the day after the announcement. The researchers chose to use z-tests for statistical significance.
Full sample (90 firms). The null hypothesis that the average abnormal stock return during [0, +1] was 0 would be true if stock investors did not find either positive or negative information in the announcement.
Mean abnormal return = –3.07 percent
z-statistic for abnormal return = –5.938
Clean sample (45 firms). The null hypothesis that the average abnormal stock return during [0, +1] was 0 would be true if stock investors did not find either positive or negative information in the announcement.
Mean abnormal return = –1.87 percent
z-statistic for abnormal return = –3.359 1. With respect to both of the cases, suppose that the null hypothesis reflects the
belief that investors do not, on average, perceive either positive or negative information in control deficiency disclosures. State one set of hypotheses (a null hypothesis and an alternative hypothesis) that covers both cases.
2. Determine whether the null hypothesis formulated in Part 1 is rejected or not rejected at the 0.05 and 0.01 levels of significance for the full sample case. Interpret the results.
3. Determine whether the null hypothesis formulated in Part 1 is rejected or not rejected at the 0.05 and 0.01 levels of significance for the clean sample case. Interpret the results.
Solution to 1: A set of hypotheses consistent with no information in control deficiency disclosures relevant to stock investors is
Solution to 2: From the information on rejection points for z-tests, we know that we reject the null hypothesis at the 0.05 significance level if z > 1.96 or if z < −1.96, and at the 0.01 significance level if z > 2.575 or if z < −2.575. The z- statistic reported by the researchers is –5.938, which is significant at the 0.05 and 0.01 levels. The null is rejected. The control deficiency disclosures appear to contain valuation-relevant information.
Because it is possible that significant results could be due to outliers, the researchers also reported the number of cases of positive and negative abnormal returns. The ratio of cases of positive to negative abnormal returns was 32:58, which tends to support the conclusion from the z-test of statistically significant negative abnormal returns.
Solution to 3: The z-statistic reported by the researchers for the clean sample is –3.359, which is significant at the 0.05 and 0.01 levels. Although both the mean abnormal return and the z-statistic are smaller in magnitude for the clean sample than for the full sample, the results continue to be statistically significant.
The ratio of cases of positive to negative abnormal returns was 16:29, which tends to support the conclusion from the z-test of statistically significant negative abnormal returns.
Nearly all practical situations involve an unknown population variance. Table 2 summarizes our discussion for tests concerning the population mean when the population variance is unknown.
TABLE 2 Test Concerning the Population Mean (Population Variance Unknown)
Large Sample (n ≥ 30) Small Sample (n < 30) Population normal t-Test (z-Test alternative) t-Test
Population non-normal t-Test (z-Test alternative) Not Available
3.2. Tests Concerning Differences between Means We often want to know whether a mean value—for example, a mean return—differs
between two groups. Is an observed difference due to chance or to different underlying values for the mean? We have two samples, one for each group. When it is reasonable to believe that the samples are from populations at least approximately normally distributed and that the samples are also independent of each other, the techniques of this section apply. We discuss two t-tests for a test concerning differences between the means of two populations. In one case, the population variances, although unknown, can be assumed to be equal. Then, we efficiently combine the observations from both samples to obtain a pooled estimate of the common but unknown population variance. A pooled estimate is an estimate drawn from the combination of two different samples. In the second case, we do not assume that the unknown population variances are equal, and an approximate t-test is then available. Letting μ1 and μ2 stand, respectively, for the population means of the first and second populations, we most often want to test whether the population means are equal or whether one is larger than the other. Thus we usually formulate the following hypotheses: 1. H0: μ1 − μ2 = 0 versus Ha: μ1 − μ2 | 0 (the alternative is that μ1 | μ2)
2. H0: μ1 − μ2 ≤ 0 versus Ha: μ1 − μ2 > 0 (the alternative is that μ1 > μ2)
3. H0: μ1 − μ2 ≥ 0 versus Ha: μ1 − μ2 < 0 (the alternative is that μ1 < μ2)
We can, however, formulate other hypotheses, such as H0: μ1 − μ2 = 2 versus Ha: μ1 − μ2 | 2. The procedure is the same.
The definition of the t-test follows.
Test Statistic for a Test of the Difference between Two Population Means (Normally Distributed Populations, Population Variances Unknown but Assumed Equal). When we can assume that the two populations are normally distributed and that the unknown population variances are equal, a t-test based on independent random samples is given by
(7)
where is a pooled estimator of the common variance.
The number of degrees of freedom is n1 + n2 − 2.
EXAMPLE 4 Mean Returns on the S&P 500: A Test of Equality across Two Halves of a Decade
The realized mean monthly return on the S&P 500 Index in the first half of the 2000s appears to have been substantially different than the mean return in the second half of the 2000s. Was the difference statistically significant? The data, shown in Table 3, indicate that assuming equal population variances for returns in the two decades is not unreasonable.
1. Formulate null and alternative hypotheses consistent with a two-sided
hypothesis test.
2. Identify the test statistic for conducting a test of the hypotheses in Part 1.
3. Identify the rejection point or points for the hypothesis tested in Part 1 at the 0.10, 0.05, and 0.01 levels of significance.
4. Determine whether the null hypothesis is rejected or not rejected at the 0.10, 0.05, and 0.01 levels of significance.
Solution to 1: Letting μ1 stand for the population mean return for the 2000 through 2004 period and μ2 stand for the population mean return for the 2005 through 2009 period, we formulate the following hypotheses:
Solution to 2: Because the two samples are drawn from two different time periods, they are independent samples. The population variances are not known but can be assumed to be equal. Given all these considerations, the t-test given in Equation 7 has 60 + 60 − 2 = 118 degrees of freedom.
Solution to 3: In the tables (Appendix B), the closest number of degrees of freedom to 118 is 120. For a two-sided test, the rejection points are ±1.658, ±1.980, and ±2.617 for, respectively, the 0.10, 0.05, and 0.01 levels for df = 120. To summarize, at the 0.10 level, we will reject the null if t < −1.658 or t > 1.658; at the 0.05 level, we will reject the null if t < −1.980 or t > 1.980; and at the 0.01 level, we will reject the null if t < −2.617 or t > 2.617.
Table 3 S&P 500 Monthly Return and Standard Deviation for Two Halves of a Decade
Time Period Number ofMonths(n) Mean MonthlyReturn
(%) Standard Deviation
2000 through 2004 60 –0.083 4.719
2005 through 2009 60 0.144 4.632
Solution to 4: In calculating the test statistic, the first step is to calculate the pooled estimate of variance:
The t value of −0.27 is not significant at the 0.10 level, so it is also not significant at the 0.05 and 0.01 levels. Therefore, we do not reject the null hypothesis at any of the three levels.
In many cases of practical interest, we cannot assume that population variances are equal. The following test statistic is often used in the investment literature in such cases:
Test Statistic for a Test of the Difference between Two Population Means (Normally Distributed Populations, Unequal and Unknown Population Variances). When we can assume that the two populations are normally distributed but do not know the population variances and cannot assume that they are equal, an approximate t-test based on independent random samples is given by
(8)
where we use tables of the t-distribution using “modified” degrees of freedom computed with the formula
(9)
A practical tip is to compute the t-statistic before computing the degrees of freedom. Whether or not the t-statistic is significant will sometimes be obvious.
EXAMPLE 5 Recovery Rates on Defaulted Bonds: A Hypothesis Test
How are the required yields on risky corporate bonds determined? Two key factors are the expected probability of default and the expected amount that will be recovered in the event of default, or the recovery rate. Jankowitsch, Nagler, and Subrahmanyam (2013) examine the recovery rates of defaulted bonds in the US corporate bond market based on an extensive set of traded prices and volumes around various types of default events. For their study period, 2002 to 2012, Jankowitsch et al. confirm that the type of default event (e.g., distressed exchanges and formal bankruptcy filings), the seniority of the bond, and the industry of the firm are important in explaining the recovery rate. In one of their analyses, they focus on non-financial firms, and find that electricity firms recover more in default than firms in the retail industry. We want to test if the difference in recovery rates between those two types of firms is statistically significant. With μ1 denoting the population mean recovery rate for the bonds of electricity firms and μ2 denoting the population mean recovery rate for the bonds of retail firms, the hypotheses are H0: μ1 − μ2 = 0 versus Ha: μ1 − μ2 | 0.
Table 4 excerpts from their findings.
We assume that the populations (recovery rates) are normally distributed and that the samples are independent. Based on the data in the table, address the following:
1. Discuss whether we should choose a test based on Equation 8 or Equation 7.
2. Calculate the test statistic to test the null hypothesis given above.
3. What is the value of the test’s modified degrees of freedom?
4. Determine whether to reject the null hypothesis at the 0.10 level.
Solution to 1: The sample standard deviation for the recovery rate on the bonds
of electricity firms ($22.67) appears much smaller than the sample standard deviation of the bonds for retail firms ($34.19). Therefore, we should not assume equal variances, and accordingly, we should employ the approximate t- test given in Equation 8.
Solution to 2: The test statistic is
TABLE 4 Recovery Rates by Industry of Firm
Electricity Retail Number of
Observations Average Pricea
Standard Deviation
Number of Observations
Average Pricea
Standard Deviation
39 $48.03 $22.67 33 $33.40 $34.19
a This is the average traded price over the default day and the following 30 days after defalt; the average price provides an indication of the amount of money that can be recovered.
Source: Jankowitsch, Nagler, and Subrahmanyam (2013), Table 2.
where
1 = sample mean recovery rate for electricity firms = 48.03
2 = sample mean recovery rate for retail firms = 33.40
= sample variance for electricity firms = 22.672 = 513.9289
= sample variance for retail firms = 34.192 = 1,168.9561
n1 = sample size of the electricity firms sample = 39
n2 = sample size of the retail firms sample = 33
Thus, t = (48.03 − 33.40)/[(513.9289/39) + (1,168.9561/33)]1/2 = 14.63/(13.177664 + 35.422912)1/2 = 14.63/6.971411 = 2.099. The calculated t- statistic is thus 2.099.
Solution to 3:
Solution to 4: The closest entry to df = 56 in the tables for the t-distribution is df = 60. For α = 0.10, we find tα/2 = 1.671. Thus, we reject the null if t < −1.671 or t > 1.671. Based on the computed value of 2.099, we reject the null hypothesis at the 0.10 level. Some evidence exists that recovery rates differ between electricity and retail industries. Why? Studies on recovery rates suggest that the higher recovery rates of electricity firms may be explained by their higher levels of tangible assets.
3.3. Tests Concerning Mean Differences In the previous section, we presented two t-tests for discerning differences between population means. The tests were based on two samples. An assumption for those tests’ validity was that the samples were independent—i.e., unrelated to each other. When we want to conduct tests on two means based on samples that we believe are dependent, the methods of this section apply.
The t-test in this section is based on data arranged in paired observations, and the test itself is sometimes called a paired comparisons test. Paired observations are observations that are dependent because they have something in common. A paired comparisons test is a statistical test for differences in dependent items. For example, we may be concerned with the dividend policy of companies before and after a change in the tax law affecting the taxation of dividends. We then have pairs of “before” and “after” observations for the same companies. We may test a hypothesis about the mean of the differences (mean differences) that we observe across companies. In other cases, the paired observations are not on the same units. For example, we may be testing whether the mean returns earned by two investment strategies were equal over a study period. The observations here are dependent in the sense that there is one observation for each strategy in each month, and both observations depend on underlying market risk factors. Because the returns to both strategies are likely to be related to some common risk factors, such as the market return, the samples are dependent. By calculating a standard error based on differences, the t-test presented below takes account of correlation between the observations.
Letting A represent “after” and B “before,” suppose we have observations for the
random variables XA and XB and that the samples are dependent. We arrange the observations in pairs. Let di denote the difference between two paired observations. We can use the notation di = xAi − xBi, where xAi and xBi are the ith pair of observations, i = 1, 2, …, n on the two variables. Let μd stand for the population mean difference. We can formulate the following hypotheses, where μd0 is a hypothesized value for the population mean difference: 1. H0: μd = μd0 versus Ha: μd | μd0 2. H0: μd ≤ μd0 versus Ha: μd > μd0
3. H0: μd ≥ μd0 versus Ha: μd < μd0
In practice, the most commonly used value for μd0 is 0.
As usual, we are concerned with the case of normally distributed populations with unknown population variances, and we will formulate a t-test. To calculate the t- statistic, we first need to find the sample mean difference:
(10)
where n is the number of pairs of observations. The sample variance, denoted by , is
(11)
Taking the square root of this quantity, we have the sample standard deviation, sd, which then allows us to calculate the standard error of the mean difference as follows:24
(12)
Test Statistic for a Test of Mean Differences (Normally Distributed Populations, Unknown Population Variances). When we have data consisting of paired observations from samples generated by normally distributed populations with unknown variances, a t-test is based on
(13)
with n − 1 degrees of freedom, where n is the number of paired observations, is the sample mean difference (as given by Equation 10), and is the standard error of (as given by Equation 12).
Table 5 reports the quarterly returns from 2008 to 2013 for two managed portfolios specializing in precious metals. The two portfolios were closely similar in risk (as measured by standard deviation of return and other measures) and had nearly identical expense ratios. A major investment services company rated Portfolio B more highly than Portfolio A in early 2014. In investigating the portfolios’ relative performance, suppose we want to test the hypothesis that the mean quarterly return on Portfolio A equaled the mean quarterly return on Portfolio B from 2008 to 2013. Because the two portfolios shared essentially the same set of risk factors, their returns were not independent, so a paired comparisons test is appropriate. Let μd stand for the population mean value of difference between the returns on the two portfolios during this period. We test H0: μd = 0 versus Ha: μd ≠ 0 at a 0.05 significance level.
TABLE 5 Quarterly Returns on Two Managed Portfolios: 2008–2013
Quarter Portfolio A (%) Portfolio B (%) Difference(Portfolio A – Portfolio B) 4Q:2013 11.40 14.64 −3.24 3Q:2013 −2.17 0.44 −2.61 2Q:2013 10.72 19.51 −8.79 1Q:2013 38.91 50.40 −11.49 4Q:2012 4.36 1.01 3.35 3Q:2012 5.13 10.18 −5.05 2Q:2012 26.36 17.77 8.59 1Q:2012 −5.53 4.76 −10.29 4Q:2011 5.27 −5.36 10.63 3Q:2011 −7.82 −1.54 −6.28 2Q:2011 2.34 0.19 2.15 1Q:2011 −14.38 −12.07 −2.31 4Q:2010 −9.80 −9.98 0.18 3Q:2010 19.03 26.18 −7.15
2Q:2010 4.11 −2.39 6.50 1Q:2010 −4.12 −2.51 −1.61 4Q:2009 −0.53 −11.32 10.79 3Q:2009 5.06 0.46 4.60 2Q:2009 −14.01 −11.56 −2.45 1Q:2009 12.50 3.52 8.98 4Q:2008 −29.05 −22.45 −6.60 3Q:2008 3.60 0.10 3.50 2Q:2008 −7.97 −8.96 0.99 1Q:2008 −8.62 −0.66 −7.96 Mean 1.87 2.52 −0.65
Sample standard deviation of differences = 6.71
The sample mean difference, , between Portfolio A and Portfolio B is −0.65 percent per quarter. The standard error of the sample mean difference is = 1.369673. The calculated test statistic is t = (−0.65 − 0)/1.369673 = −0.475 with n − 1 = 24 − 1 = 23 degrees of freedom. At the 0.05 significance level, we reject the null if t > 2.069 or if t < −2.069. Because −0.475 is not less than −2.069, we fail to reject the null. At the 0.10 significance level, we reject the null if t > 1.714 or if t < −1.714. Thus the difference in mean quarterly returns is not significant at any conventional significance level.
The following example illustrates the application of this test to evaluate two competing investment strategies.
EXAMPLE 6 A Comparison of Two Portfolios
You are investigating whether the performance of a portfolio of stocks from the entire world differs from the performance of a portfolio of only US stocks. For the worldwide portfolio, you choose to focus on Vanguard Total World Stock Index ETF (NYSE: VT). This ETF seeks to track the performance of the FTSE Global All Cap Index, which is a market-capitalization-weighted index designed to measure the market performance of stock of companies from both developed and emerging markets. For the US portfolio, you choose to focus on SPDR S&P 500 (NYSE: SPY), an ETF that seeks to track the performance of the S&P 500 Index. You analyze the monthly returns on both ETFs from August 2008 to July 2013 and prepare the following summary table.
TABLE 6 Monthly Return Summary for Vanguard Total World Stock Index ETF and SPDR S&P 500 ETF: August 2008 to July 2013 (n = 60)
Strategy Mean Return Standard Deviation Worldwide 0.61% 5.43%
US 0.39 6.50 Difference 0.22 1.86a
a Sample standard deviation of differences.
Source of data returns: finance.yahoo.com accessed 20 August 2013.
From Table 6 we have = 0.22% and sd = 1.86%. 1. Formulate null and alternative hypotheses consistent with a two-sided test that
the mean difference between the worldwide and only US strategies equals 0.
2. Identify the test statistic for conducting a test of the hypotheses in Part 1.
3. Identify the rejection point or points for the hypothesis tested in Part 1 at the 0.01 level of significance.
4. Determine whether the null hypothesis is rejected or not rejected at the 0.01 level of significance. (Use the tables in the back of this volume.)
5. Discuss the choice of a paired comparisons test.
Solution to 1: With μd as the underlying mean difference between the worldwide and US strategies, we have H0: μd = 0 versus Ha: μd | 0.
Solution to 2: Because the population variance is unknown, the test statistic is a t-test with 60 − 1 = 59 degrees of freedom.
Solution to 3: In the table for the t-distribution, the closest entry to df = 59 is df = 60. We look across the row for 60 degrees of freedom to the 0.005 column, to find 2.66. We will reject the null if we find that t > 2.66 or t < −2.66.
Solution to 4:
Because 0.92 < 2.66, we cannot reject the null hypothesis. Accordingly, we conclude that the difference in mean returns for the two strategies is not statistically significant.
Solution to 5: Several US stocks that are part of the S&P 500 index are also included in the Vanguard Total World Stock Index ETF. The profile of the World ETF indicates that nine of the top ten holdings in the ETF are US stocks. As a result, they are not independent samples; in general, the correlation of returns on the Vanguard Total World Stock Index ETF and SPDR S&P 500 ETF should be positive. Because the samples are dependent, a paired comparisons test was appropriate.
4. Hypothesis Tests Concerning Variance Because variance and standard deviation are widely used quantitative measures of risk in investments, analysts should be familiar with hypothesis tests concerning variance. The tests discussed in this section make regular appearances in investment literature. We examine two types: tests concerning the value of a single population variance and tests concerning the differences between two population variances.
4.1. Tests Concerning a Single Variance
In this section, we discuss testing hypotheses about the value of the variance, σ2, of a single population. We use to denote the hypothesized value of σ2. We can formulate hypotheses as follows: 1. (a “not equal to” alternative hypothesis)
2. (a “greater than” alternative hypothesis)
3. (a “less than” alternative hypothesis)
In tests concerning the variance of a single normally distributed population, we make use of a chi-square test statistic, denoted χ2. The chi-square distribution, unlike the normal and t-distributions, is asymmetrical. Like the t-distribution, the chi- square distribution is a family of distributions. A different distribution exists for each possible value of degrees of freedom, n − 1 (n is sample size). Unlike the t- distribution, the chi-square distribution is bounded below by 0; χ2 does not take on negative values.
Test Statistic for Tests Concerning the Value of a Population Variance (Normal Population). If we have n independent observations from a normally distributed population, the appropriate test statistic is
(14)
with n − 1 degrees of freedom. In the numerator of the expression is the sample variance, calculated as
(15)
In contrast to the t-test, for example, the chi-square test is sensitive to violations of its assumptions. If the sample is not actually random or if it does not come from a normally distributed population, inferences based on a chi-square test are likely to be faulty.
If we choose a level of significance, α, the rejection points for the three kinds of hypotheses are as follows:
Rejection Points for Hypothesis Tests on the Population Variance.
1. “Not equal to” Ha: Reject the null hypothesis if the test statistic is greater than the upper α/2 point (denoted ) or less than the lower α/2 point (denoted ) of the chi-square distribution with df = n − 1.25
2. “Greater than” Ha: Reject the null hypothesis if the test statistic is greater than the upper α point of the chi-square distribution with df = n − 1.
3. “Less than” Ha: Reject the null hypothesis if the test statistic is less than the lower α point of the chi-square distribution with df = n − 1.
EXAMPLE 7 Risk and Return Characteristics of an Equity Mutual Fund (2)
You continue with your analysis of Sendar Equity Fund, a mid-cap growth fund that has been in existence for only 24 months. Recall that during this period, Sendar Equity achieved a sample standard deviation of monthly returns of 3.60 percent. You now want to test a claim that the particular investment disciplines followed by Sendar result in a standard deviation of monthly returns of less than 4 percent.
1. Formulate null and alternative hypotheses consistent with the verbal description
of the research goal.
2. Identify the test statistic for conducting a test of the hypotheses in Part 1.
3. Identify the rejection point or points for the hypothesis tested in Part 1 at the 0.05 level of significance.
4. Determine whether the null hypothesis is rejected or not rejected at the 0.05 level of significance. (Use the tables in the back of this volume.)
Solution to 1: We have a “less than” alternative hypothesis, where σ is the
underlying standard deviation of return on Sendar Equity Fund. Being careful to square standard deviation to obtain a test in terms of variance, the hypotheses are H0: σ2 ≥ 16.0 versus Ha: σ2 < 16.0.
Solution to 2: The test statistic is χ2 with 24 − 1 = 23 degrees of freedom.
Solution to 3: The lower 0.05 rejection point is found on the line for df = 23, under the 0.95 column (95 percent probability in the right tail, to give 0.95 probability of getting a test statistic this large or larger). The rejection point is 13.091. We will reject the null if we find that χ2 is less than 13.091.
Solution to 4:
Because 18.63 (the calculated value of the test statistic) is not less than 13.091, we do not reject the null hypothesis. We cannot conclude that Sendar ’s investment disciplines result in a standard deviation of monthly returns of less than 4 percent.
4.2. Tests Concerning the Equality (Inequality) of Two Variances Suppose we have a hypothesis about the relative values of the variances of two normally distributed populations with means μ1 and μ2 and variances and . We can formulate all hypotheses as one of the choices below: 1.
2.
3.
Note that at the point of equality, the null hypothesis implies that the ratio of population variances equals 1: . Given independent random samples from these populations, tests related to these hypotheses are based on an F-test, which is the ratio of sample variances. Suppose we use n1 observations in calculating the sample variance and n2 observations in calculating the sample variance . Tests concerning the difference between the variances of two populations make use of the F-distribution. Like the chi-square distribution, the F-distribution is a family of asymmetrical distributions bounded from below by 0. Each F-distribution is defined by two values of degrees of freedom, called the numerator and denominator
degrees of freedom.26 The F-test, like the chi-square test, is not robust to violations of its assumptions.
Test Statistic for Tests Concerning Differences between the Variances of Two Populations (Normally Distributed Populations). Suppose we have two samples, the first with n1 observations and sample variance , the second with n2 observations and sample variance . The samples are random, independent of each other, and generated by normally distributed populations. A test concerning differences between the variances of the two populations is based on the ratio of sample variances
(16)
with df1 = n1 − 1 numerator degrees of freedom and df2 = n2 − 1 denominator degrees of freedom. Note that df1 and df2 are the divisors used in calculating and , respectively.
A convention, or usual practice, is to use the larger of the two ratios or as the actual test statistic. When we follow this convention, the value of the test statistic is always greater than or equal to 1; tables of critical values of F then need include only values greater than or equal to 1. Under this convention, the rejection point for any formulation of hypotheses is a single value in the right-hand side of the relevant F-distribution. Note that the labeling of populations as “1” or “2” is arbitrary in any case.
Rejection Points for Hypothesis Tests on the Relative Values of Two Population Variances. Follow the convention of using the larger of the two ratios and and consider two cases:
1. A “not equal to” alternative hypothesis: Reject the null hypothesis at the α significance level if the test statistic is greater than the upper α/2 point of the F-distribution with the specified numerator and denominator degrees of freedom.
2. A “greater than” or “less than” alternative hypothesis: Reject the null hypothesis at the α significance level if the test statistic is greater than the upper α point of the F-distribution with the specified number of numerator and denominator degrees of freedom.
Thus, if we conduct a two-sided test at the α = 0.01 level of significance, we need to
find the rejection point in F-tables at the α/2 = 0.01/2 = 0.005 significance level for a one-sided test (Case 1). But a one-sided test at 0.01 uses rejection points in F-tables for α = 0.01 (Case 2). As an example, suppose we are conducting a two-sided test at the 0.05 significance level. We calculate a value of F of 2.77 with 12 numerator and 19 denominator degrees of freedom. Using the F-tables for 0.05/2 = 0.025 in the back of the volume, we find that the rejection point is 2.72. Because the value 2.77 is greater than 2.72, we reject the null hypothesis at the 0.05 significance level.
If the convention stated above is not followed and we are given a calculated value of F less than 1, can we still use F-tables? The answer is yes; using a reciprocal property of F-statistics, we can calculate the needed value. The easiest way to present this property is to show a calculation. Suppose our chosen level of significance is 0.05 for a two-tailed test and we have a value of F of 0.11, with 7 numerator degrees of freedom and 9 denominator degrees of freedom. We take the reciprocal, 1/0.11 = 9.09. Then we look up this value in the F-tables for 0.025 (because it is a two-tailed test) with degrees of freedom reversed: F for 9 numerator and 7 denominator degrees of freedom. In other words, F9,7 = 1/F7,9 and 9.09 exceeds the critical value of 4.82, so F7,9 = 0.11 is significant at the 0.05 level.
EXAMPLE 8 Volatility and the Global Financial Crisis of the Late 2000s
You are investigating whether the population variance of returns on the KOSPI Index of the South Korean stock market changed subsequent to the global financial crisis that peaked in 2008. For this investigation, you are considering 2004 to 2006 as the pre-crisis period and 2010 to 2012 as the post-crisis period. You gather the data in Table 7 for 156 weeks of returns during 2004 to 2006 and 156 weeks of returns during 2010 to 2012. You have specified a 0.01 level of significance.
1. Formulate null and alternative hypotheses consistent with the verbal description
of the research goal.
2. Identify the test statistic for conducting a test of the hypotheses in Part 1.
3. Determine whether or not to reject the null hypothesis at the 0.01 level of significance. (Use the F-tables in the back of this volume.)
Solution to 1: We have a “not equal to” alternative hypothesis:
Solution to 2: To test a null hypothesis of the equality of two variances, we use with 156 − 1 = 155 numerator and denominator degrees of freedom.
Solution to 3: The “before” sample variance is larger, so following a convention for calculating F-statistics, the “before” sample variance goes in the numerator: F = 7.240/6.269 = 1.155. Because this is a two-tailed test, we use F-tables for the 0.005 level (= 0.01/2) to give a 0.01 significance level. In the tables in the back of the volume, the closest value to 155 degrees of freedom is 120 degrees of freedom. At the 0.01 level, the rejection point is 1.61. Because 1.155 is less than the critical value 1.61, we cannot reject the null hypothesis that the population variance of returns is the same in the pre-and post-global financial crisis periods.
TABLE 7 KOSPI Index Returns and Variance before and after the Global Financial Crisis of the Late 2000s
n Mean WeeklyReturn (%) Varianceof Returns Before crisis: 2004 to 2006 156 0.358 7.240 After crisis: 2010 to 2012 156 0.110 6.269
Source of data for returns: finance.yahoo.com accessed 27 August 2013.
EXAMPLE 9 The Volatility of Derivatives Expiration Days
Since 2001, the financial markets in the United States have seen the quadruple occurrence of stock option, index option, index futures, and single stock futures expirations on the same day during four months of the year. Such days are known as “quadruple witching days.” You are interested in investigating whether quadruple witching days exhibit greater volatility than normal days. Table 8 presents the daily standard deviation of return for normal days and options/futures expiration days during the four-year period 20X3 to 20X6. The tabled data refer to options and futures on the 30 stocks that constitute the Dow Jones Industrial Average.
1. Formulate null and alternative hypotheses consistent with the belief that
quadruple witching days display above-normal volatility.
2. Identify the test statistic for conducting a test of the hypotheses in Part 1.
3. Determine whether or not to reject the null hypothesis at the 0.05 level of
significance. (Use the F-tables in the back of this volume.)
Solution to 1: We have a “greater than” alternative hypothesis:
Solution to 2: Let represent the variance of quadruple witching days, and represent the variance of normal days, following the convention for the selection of the numerator and the denominator stated earlier. To test the null hypothesis, we use with 16 − 1 = 15 numerator and 138 − 1 = 137 denominator degrees of freedom.
Solution to 3: F = (1.217)2/(0.821)2 = 1.481/0.674 = 2.20. Because this is a one- tailed test at the 0.05 significance level, we use F-tables for the 0.05 level directly. In the tables in the back of the volume, the closest value to 137 degrees of freedom is 120 degrees of freedom. At the 0.05 level, the rejection point is 1.75. Because 2.20 is greater than 1.75, we reject the null hypothesis. It appears that quadruple witching days have above-normal volatility.
TABLE 8 Standard Deviation of Return: 20X3 to 20X6
Type of Day n Standard Deviation (%) Normal trading 138 0.821
Options/futures expiration 16 1.217
5. Other Issues: Nonparametric Inference The hypothesis-testing procedures we have discussed to this point have two characteristics in common. First, they are concerned with parameters, and second, their validity depends on a definite set of assumptions. Mean and variance, for example, are two parameters, or defining quantities, of a normal distribution. The tests also make specific assumptions—in particular, assumptions about the distribution of the population producing the sample. Any test or procedure with either of the above two characteristics is a parametric test or procedure. In some cases, however, we are concerned about quantities other than parameters of distributions. In other cases, we may believe that the assumptions of parametric tests do not hold for the particular data we have. In such cases, a nonparametric test or procedure can be useful. A nonparametric test is a test that is not concerned with a parameter, or a test that makes minimal assumptions about the population from which the sample comes.27
We primarily use nonparametric procedures in three situations: when the data we use do not meet distributional assumptions, when the data are given in ranks, or when the hypothesis we are addressing does not concern a parameter.
The first situation occurs when the data available for analysis suggest that the distributional assumptions of the parametric test are not satisfied. For example, we may want to test a hypothesis concerning the mean of a population but believe that neither a t-test nor a z-test is appropriate because the sample is small and may come from a markedly non-normally distributed population. In that case, we may use a nonparametric test. The nonparametric test will frequently involve the conversion of observations (or a function of observations) into ranks according to magnitude, and sometimes it will involve working with only “greater than” or “less than” relationships (using the signs + and − to denote those relationships). Characteristically, one must refer to specialized statistical tables to determine the rejection points of the test statistic, at least for small samples.28 Such tests, then, typically interpret the null hypothesis as a thesis about ranks or signs. In Table 9, we give examples of nonparametric alternatives to the parametric tests we have discussed in this reading.29 The reader should consult a comprehensive business statistics textbook for an introduction to such tests, and a specialist textbook for details.30
TABLE 9 Nonparametric Alternatives to Parametric Tests
Parametric Nonparametric
Tests concerning a single mean t-test Wilcoxon signed-ranktest z-test
Tests concerning differencesbetween means
t- testApproximate
t-test Mann–Whitney U test
Tests concerning mean differences(paired comparisons tests) t-test
Wilcoxon signed-rank testSign test
We pointed out that when we use nonparametric tests, we often convert the original data into ranks. In some cases, the original data are already ranked. In those cases, we also use nonparametric tests because parametric tests generally require a stronger measurement scale than ranks. For example, if our data were the rankings of investment managers, hypotheses concerning those rankings would be tested using nonparametric procedures. Ranked data also appear in many other finance contexts. For example, Heaney, Koga, Oliver, and Tran (1999) studied the relationship between the size of Japanese companies (as measured by revenue) and their use of derivatives. The companies studied used derivatives to hedge one or more of five types of risk exposure: interest rate risk, foreign exchange risk, commodity price risk, marketable security price risk, and credit risk. The researchers gave a “perceived scope of risk exposure” score to each company that was equal to the number of types of risk exposure that the company reported hedging. Although revenue is measured on a strong scale (a ratio scale), scope of risk exposure is measured on only an ordinal scale.31 The researchers thus employed nonparametric statistics to explore the relationship between derivatives usage and size.
A third situation in which we use nonparametric procedures occurs when our question does not concern a parameter. For example, if the question concerns whether a sample is random or not, we use the appropriate nonparametric test (a so- called runs test). Another type of question nonparametrics can address is whether a sample came from a population following a particular probability distribution (using the Kolmogorov–Smirnov test, for example).
We end this reading by describing in some detail a nonparametric statistic that has often been used in investment research, the Spearman rank correlation.
5.1. Tests Concerning Correlation: The Spearman Rank Correlation Coefficient In many contexts in investments, we want to assess the strength of the linear
relationship between two variables—the correlation between them. In a majority of cases, we use the correlation coefficient described in the readings on probability concepts and correlation and regression. However, the t-test of the hypothesis that two variables are uncorrelated, based on the correlation coefficient, relies on fairly stringent assumptions.32 When we believe that the population under consideration meaningfully departs from those assumptions, we can employ a test based on the Spearman rank correlation coefficient, rS. The Spearman rank correlation coefficient is essentially equivalent to the usual correlation coefficient calculated on the ranks of the two variables (say X and Y) within their respective samples. Thus it is a number between −1 and +1, where −1 (+1) denotes a perfect inverse (positive) straight-line relationship between the variables and 0 represents the absence of any straight-line relationship (no correlation). The calculation of rS requires the following steps: 1. Rank the observations on X from largest to smallest. Assign the number 1 to
the observation with the largest value, the number 2 to the observation with second-largest value, and so on. In case of ties, we assign to each tied observation the average of the ranks that they jointly occupy. For example, if the third-and fourth-largest values are tied, we assign both observations the rank of 3.5 (the average of 3 and 4). Perform the same procedure for the observations on Y.
2. Calculate the difference, di, between the ranks of each pair of observations on X and Y.
3. Then, with n the sample size, the Spearman rank correlation is given by33
(17)
Suppose an investor wants to invest in a diversified emerging markets mutual fund. He has narrowed the field to 10 such funds, which are the largest in terms of total net assets. In examining the funds, a question arises as to whether the funds’ most recent reported Sharpe ratios and expense ratios as of mid-2013 are related. Because the assumptions of the t-test on the correlation coefficient may not be met, it is appropriate to conduct a test on the rank correlation coefficient.34 Table 10 presents the calculation of rS. The first two rows contain the original data. The row of X ranks converts the Sharpe ratios to ranks; the row of Y ranks converts the expense ratios to ranks. We want to test H0: ρ = 0 versus Ha: ρ | 0, where ρ is defined in this context as the population correlation of X and Y after ranking. For small samples,
the rejection points for the test based on rS must be looked up in Table 11. For large samples (say n > 30), we can conduct a t-test using
(18)
based on n − 2 degrees of freedom.
TABLE 10 The Spearman Rank Correlation: An Example
Mutual Fund 1 2 3 4 5 6 7 8 9 10
Sharpe Ratio (X) 0.05 0.40 0.38 0.21 0.43 0.16 0.40 0.58 0.14 0.25 Expense Ratio (Y) 0.61 1.03 1.36 1.10 1.07 0.68 1.10 1.37 1.27 1.25
X Rank 10 3.5 5 7 2 8 3.5 1 9 6 Y Rank 10 8 2 5.5 7 9 5.5 1 3 4 di 0.0 –4.5 3 1.5 –5 –1 –2 0 6 2
0.0 20.25 9 2.25 25 1 4 0 36 4
Source of Sharpe and Expense Ratios: http://markets.on.nytimes.com/research/screener/mutual_funds/mutual_funds.asp accessed 20 August 2013.
TABLE 11 Spearman Rank Correlation Distribution Approximate Upper-Tail Rejection Points
Sample Size: n α = 0.05 α = 0.025 α = 0.01 5 0.8000 0.9000 0.9000 6 0.7714 0.8286 0.8857 7 0.6786 0.7450 0.8571 8 0.6190 0.7143 0.8095 9 0.5833 0.6833 0.7667 10 0.5515 0.6364 0.7333 11 0.5273 0.6091 0.7000
12 0.4965 0.5804 0.6713 13 0.4780 0.5549 0.6429 14 0.4593 0.5341 0.6220 15 0.4429 0.5179 0.6000 16 0.4265 0.5000 0.5824 17 0.4118 0.4853 0.5637 18 0.3994 0.4716 0.5480 19 0.3895 0.4579 0.5333 20 0.3789 0.4451 0.5203 21 0.3688 0.4351 0.5078 22 0.3597 0.4241 0.4963 23 0.3518 0.4150 0.4852 24 0.3435 0.4061 0.4748 25 0.3362 0.3977 0.4654 26 0.3299 0.3894 0.4564 27 0.3236 0.3822 0.4481 28 0.3175 0.3749 0.4401 29 0.3113 0.3685 0.4320 30 0.3059 0.3620 0.4251
Note: The corresponding lower-tail critical value is obtained by changing the sign of the upper-tail critical value.
In the example at hand, a two-tailed test with a 0.05 significance level, Table 11 gives the upper-tail rejection point for n = 10 as 0.6364 (we use the 0.025 column for a two-tailed test at a 0.05 significance level). Accordingly, we reject the null hypothesis if rS is less than −0.6364 or greater than 0.6364. With rS equal to 0.3848, we do not reject the null hypothesis.
In the mutual fund example, we converted observations on two variables into ranks. If one or both of the original variables were in the form of ranks, we would need to use rS to investigate correlation.
5.2. Nonparametric Inference: Summary Nonparametric statistical procedures extend the reach of inference because they make few assumptions, can be used on ranked data, and may address questions
unrelated to parameters. Quite frequently, nonparametric tests are reported alongside parametric tests. The reader can then assess how sensitive the statistical conclusion is to the assumptions underlying the parametric test. However, if the assumptions of the parametric test are met, the parametric test (where available) is generally preferred to the nonparametric test because the parametric test usually permits us to draw sharper conclusions.35 For complete coverage of all the nonparametric procedures that may be encountered in the finance and investment literature, it is best to consult a specialist textbook.36
6. Summary In this reading, we have presented the concepts and methods of statistical inference and hypothesis testing.
A hypothesis is a statement about one or more populations.
The steps in testing a hypothesis are as follows:
1. Stating the hypotheses.
2. Identifying the appropriate test statistic and its probability distribution.
3. Specifying the significance level.
4. Stating the decision rule.
5. Collecting the data and calculating the test statistic.
6. Making the statistical decision.
7. Making the economic or investment decision.
We state two hypotheses: The null hypothesis is the hypothesis to be tested; the alternative hypothesis is the hypothesis accepted when the null hypothesis is rejected.
There are three ways to formulate hypotheses:
1. H0: θ = θ0 versus Ha: θ | θ0
2. H0: θ ≤ θ0 versus Ha: θ > θ0
3. H0: θ ≥ θ0 versus Ha: θ < θ0 where θ0 is a hypothesized value of the population parameter and θ is the true value of the population parameter. In the above, Formulation 1 is a two-sided test and Formulations 2 and 3 are one-sided tests.
When we have a “suspected” or “hoped for” condition for which we want to find supportive evidence, we frequently set up that condition as the alternative hypothesis and use a one-sided test. To emphasize a neutral attitude, however, the researcher may select a “not equal to” alternative hypothesis and conduct a two-sided test.
A test statistic is a quantity, calculated on the basis of a sample, whose value is the basis for deciding whether to reject or not reject the null hypothesis. To
decide whether to reject, or not to reject, the null hypothesis, we compare the computed value of the test statistic to a critical value (rejection point) for the same test statistic.
In reaching a statistical decision, we can make two possible errors: We may reject a true null hypothesis (a Type I error), or we may fail to reject a false null hypothesis (a Type II error).
The level of significance of a test is the probability of a Type I error that we accept in conducting a hypothesis test. The probability of a Type I error is denoted by the Greek letter alpha, α. The standard approach to hypothesis testing involves specifying a level of significance (probability of Type I error) only.
The power of a test is the probability of correctly rejecting the null (rejecting the null when it is false).
A decision rule consists of determining the rejection points (critical values) with which to compare the test statistic to decide whether to reject or not to reject the null hypothesis. When we reject the null hypothesis, the result is said to be statistically significant.
The (1 − α) confidence interval represents the range of values of the test statistic for which the null hypothesis will not be rejected at an α significance level.
The statistical decision consists of rejecting or not rejecting the null hypothesis. The economic decision takes into consideration all economic issues pertinent to the decision.
The p-value is the smallest level of significance at which the null hypothesis can be rejected. The smaller the p-value, the stronger the evidence against the null hypothesis and in favor of the alternative hypothesis. The p-value approach to hypothesis testing does not involve setting a significance level; rather it involves computing a p-value for the test statistic and allowing the consumer of the research to interpret its significance.
For hypothesis tests concerning the population mean of a normally distributed population with unknown (known) variance, the theoretically correct test statistic is the t-statistic (z-statistic). In the unknown variance case, given large samples (generally, samples of 30 or more observations), the z-statistic may be used in place of the t-statistic because of the force of the central limit theorem.
The t-distribution is a symmetrical distribution defined by a single parameter: degrees of freedom. Compared to the standard normal distribution, the t-
distribution has fatter tails.
When we want to test whether the observed difference between two means is statistically significant, we must first decide whether the samples are independent or dependent (related). If the samples are independent, we conduct tests concerning differences between means. If the samples are dependent, we conduct tests of mean differences (paired comparisons tests).
When we conduct a test of the difference between two population means from normally distributed populations with unknown variances, if we can assume the variances are equal, we use a t-test based on pooling the observations of the two samples to estimate the common (but unknown) variance. This test is based on an assumption of independent samples.
When we conduct a test of the difference between two population means from normally distributed populations with unknown variances, if we cannot assume that the variances are equal, we use an approximate t-test using modified degrees of freedom given by a formula. This test is based on an assumption of independent samples.
In tests concerning two means based on two samples that are not independent, we often can arrange the data in paired observations and conduct a test of mean differences (a paired comparisons test). When the samples are from normally distributed populations with unknown variances, the appropriate test statistic is a t-statistic. The denominator of the t-statistic, the standard error of the mean differences, takes account of correlation between the samples.
In tests concerning the variance of a single, normally distributed population, the test statistic is chi-square (χ2) with n − 1 degrees of freedom, where n is sample size.
For tests concerning differences between the variances of two normally distributed populations based on two random, independent samples, the appropriate test statistic is based on an F-test (the ratio of the sample variances).
The F-statistic is defined by the numerator and denominator degrees of freedom. The numerator degrees of freedom (number of observations in the sample minus 1) is the divisor used in calculating the sample variance in the numerator. The denominator degrees of freedom (number of observations in the sample minus 1) is the divisor used in calculating the sample variance in the denominator. In forming an F-test, a convention is to use the larger of the two ratios, or , as the actual test statistic.
A parametric test is a hypothesis test concerning a parameter or a hypothesis
test based on specific distributional assumptions. In contrast, a nonparametric test either is not concerned with a parameter or makes minimal assumptions about the population from which the sample comes.
A nonparametric test is primarily used in three situations: when data do not meet distributional assumptions, when data are given in ranks, or when the hypothesis we are addressing does not concern a parameter.
The Spearman rank correlation coefficient is calculated on the ranks of two variables within their respective samples.
References
1. Bowerman, Bruce L., Richard T. O’Connell, and Emily S. Murphree. 2013. Business Statistics in Practice. New York: McGraw-Hill/Irwin.
2. Daniel, Wayne W., and James C. Terrell. 1995. Business Statistics for Management & Economics, 7th edition. Boston: Houghton-Mifflin.
3. Davidson, Russell, and James G. MacKinnon. 1993. Estimation and Inference in Econometrics. New York: Oxford University Press.
4. Dimson, Elroy, Paul Marsh, and Mike Staunton. 2011. “Equity Premiums around the World.” In Rethinking the Equity Risk Premium. Charlottesville, VA: Research Foundation of CFA Institute.
5. Freeley, Austin J., and David L. Steinberg. 2008. Argumentation and Debate: Critical Thinking for Reasoned Decision Making, 12th edition. Boston, MA: Wadsworth Cengage Learning.
6. Gupta, Parveen P, and Nandkumar Nayar. 2007. “Information Content of Control Deficiency Disclosures under the Sarbanes-Oxley Act: An Empirical Investigation.” International Journal of Disclosure and Governance, vol. 4: 3– 23.
7. Heaney, Richard, Chitoshi Koga, Barry Oliver, and Alfred Tran. 1999. “The Size Effect and Derivative Usage in Japan.” Working paper: The Australian National University.
8. Hettmansperger, Thomas P., and Joseph W. McKean. 2010. Robust Nonparametric Statistical Methods, 2nd edition. Boca Raton, FL: CRC Press.
9. Jankowitsch, Rainer, Florian Nagler, and Marti G. Subrahmanyam. 2013. “The Determinants of Recovery Rates in the US Corporate Bond Market.” Working paper. Available at SSRN: http://ssrn.com/abstract=2140793 or http://dx.doi.org/
10. Moore, David S., George P. McCabe, and Bruce Craig. 2010. Introduction to the Practice of Statistics, 7th edition. New York: W.H. Freeman.
11. Siegel, Sidney, and N. John Castellan. 1988. Nonparametric Statistics for the Behavioral Sciences, 2nd edition. New York: McGraw-Hill.
Problems Practice Problems and Solutions: 1–15 taken from Quantitative Methods for Investment Analysis, Second Edition, by Richard A. DeFusco, CFA, Dennis W. McLeavey, CFA, Jerald E. Pinto, CFA, and David E. Runkle, CFA. Copyright © 2004 by CFA Institute. All other problems and solutions copyright © CFA Institute. 1. Define the following terms:
1. Null hypothesis.
2. Alternative hypothesis.
3. Test statistic.
4. Type I error.
5. Type II error.
6. Power of a test.
7. Rejection point (critical value).
2. Suppose that, on the basis of a sample, we want to test the hypothesis that the mean debt-to-total-assets ratio of companies that become takeover targets is the same as the mean debt-to-total-assets ratio of companies in the same industry that do not become takeover targets. Explain under what conditions we would commit a Type I error and under what conditions we would commit a Type II error.
3. Suppose we are testing a null hypothesis, H0, versus an alternative hypothesis, Ha, and the p-value for the test statistic is 0.031. At which of the following levels of significance—α = 0.10, α = 0.05, and/or α = 0.01—would we reject the null hypothesis?
4. Identify the appropriate test statistic or statistics for conducting the following hypothesis tests. (Clearly identify the test statistic and, if applicable, the number of degrees of freedom. For example, “We conduct the test using an x-statistic with y degrees of freedom.”)
1. H0: μ = 0 versus Ha: μ | 0, where μ is the mean of a normally distributed population with unknown variance. The test is based on a sample of 15 observations.
2. H0: μ = 0 versus Ha: μ | 0, where μ is the mean of a normally distributed
population with unknown variance. The test is based on a sample of 40 observations.
3. H0: μ ≤ 0 versus Ha: μ > 0, where μ is the mean of a normally distributed population with known variance σ2. The sample size is 45.
4. H0: σ2 = 200 versus Ha: σ2 | 200, where σ2 is the variance of a normally distributed population. The sample size is 50.
5. , where is the variance of one normally distributed population and is the variance of a second normally distributed population. The test is based on two independent random samples.
6. H0: (Population mean 1) − (Population mean 2) = 0 versus Ha: (Population mean 1) − (Population mean 2) | 0, where the samples are drawn from normally distributed populations with unknown variances. The observations in the two samples are correlated.
7. H0: (Population mean 1) − (Population mean 2) = 0 versus Ha: (Population mean 1) − (Population mean 2) | 0, where the samples are drawn from normally distributed populations with unknown but assumed equal variances. The observations in the two samples (of size 25 and 30, respectively) are independent.
5. For each of the following hypothesis tests concerning the population mean, μ, state the rejection point condition or conditions for the test statistic (e.g., t > 1.25); n denotes sample size.
1. H0: μ = 10 versus Ha: μ | 10, using a t-test with n = 26 and α = 0.05
2. H0: μ = 10 versus Ha: μ | 10, using a t-test with n = 40 and α = 0.01
3. H0: μ ≤ 10 versus Ha: μ > 10, using a t-test with n = 40 and α = 0.01
4. H0: μ ≤ 10 versus Ha: μ > 10, using a t-test with n = 21 and α = 0.05
5. H0: μ ≥ 10 versus Ha: μ < 10, using a t-test with n = 19 and α = 0.10
6. H0: μ ≥ 10 versus Ha: μ < 10, using a t-test with n = 50 and α = 0.05
6. For each of the following hypothesis tests concerning the population mean, μ, state the rejection point condition or conditions for the test statistic (e.g., z > 1.25); n denotes sample size.
1. H0: μ = 10 versus Ha: μ | 10, using a z-test with n = 50 and α = 0.01
2. H0: μ = 10 versus Ha: μ | 10, using a z-test with n = 50 and α = 0.05
3. H0: μ = 10 versus Ha: μ | 10, using a z-test with n = 50 and α = 0.10
4. H0: μ ≤ 10 versus Ha: μ > 10, using a z-test with n = 50 and α = 0.05
7. Identify the theoretically correct test statistic to use for a hypothesis test concerning the mean of a single population under the following conditions:
1. The sample comes from a normally distributed population with known variance.
2. The sample comes from a normally distributed population with unknown variance.
3. The sample comes from a population following a non-normal distribution with unknown variance. The sample size is large.
8. Willco is a manufacturer in a mature cyclical industry. During the most recent industry cycle, its net income averaged $30 million per year with a standard deviation of $10 million (n = 6 observations). Management claims that Willco’s performance during the most recent cycle results from new approaches and that we can dismiss profitability expectations based on its average or normalized earnings of $24 million per year in prior cycles.
With μ as the population value of mean annual net income, formulate null and alternative hypotheses consistent with testing Willco management’s claim.
Assuming that Willco’s net income is at least approximately normally distributed, identify the appropriate test statistic.
Identify the rejection point or points at the 0.05 level of significance for the hypothesis tested in Part A.
Determine whether or not to reject the null hypothesis at the 0.05 significance level.
The following information relates to Questions 9–10
Performance in Forecasting Quarterly Earnings per Share Number of Forecasts
Mean Forecast Error (Predicted – Actual)
Standard Deviations of Forecast Errors
Analyst A
101 0.05 0.10
Analyst B
121 0.02 0.09
9. Investment analysts often use earnings per share (EPS) forecasts. One test of forecasting quality is the zero-mean test, which states that optimal forecasts should have a mean forecasting error of 0. (Forecasting error = Predicted value of variable − Actual value of variable.)
You have collected data (shown in the table above) for two analysts who cover two different industries: Analyst A covers the telecom industry; Analyst B covers automotive parts and suppliers.
With μ as the population mean forecasting error, formulate null and alternative hypotheses for a zero-mean test of forecasting quality.
For Analyst A, using both a t-test and a z-test, determine whether to reject the null at the 0.05 and 0.01 levels of significance.
For Analyst B, using both a t-test and a z-test, determine whether to reject the null at the 0.05 and 0.01 levels of significance.
10. Reviewing the EPS forecasting performance data for Analysts A and B, you want to investigate whether the larger average forecast errors of Analyst A are due to chance or to a higher underlying mean value for Analyst A. Assume that the forecast errors of both analysts are normally distributed and that the samples are independent.
Formulate null and alternative hypotheses consistent with determining whether the population mean value of Analyst A’s forecast errors (μ1) are larger than Analyst B’s (μ2).
Identify the test statistic for conducting a test of the null hypothesis formulated in Part A.
Identify the rejection point or points for the hypothesis tested in Part A, at the 0.05 level of significance.
Determine whether or not to reject the null hypothesis at the 0.05 level of significance.
11. Altman and Kishore (1996), in the course of a study on the recovery rates on defaulted bonds, investigated the recovery of utility bonds versus other bonds, stratified by seniority. The following table excerpts their findings.
Recovery Rates by Seniority Industry Group Ex-Utilities Sample
Industry Group/Seniority
Number of Observations
Average Pricea
Standard Deviation
Number of Observations
Average Pricea
Standard Deviation
Public Utilities Senior
Unsecured 32 $77.74 $18.06 189 $42.56
a This is the average price at default and is a measure of recovery rate.
Source: Altman and Kishore (1996, Table 5).
Assume that the populations (recovery rates of utilities, recovery rates of non-utilities) are normally distributed and that the samples are independent. The population variances are unknown; do not assume they are equal. The test hypotheses are H0: μ1 − μ2 = 0 versus Ha: μ1 − μ2 | 0, where μ1 is the population mean recovery rate for utilities and μ2 is the population mean recovery rate for non-utilities.
Calculate the test statistic.
Determine whether to reject the null hypothesis at the 0.01 significance level without reference to degrees of freedom.
Calculate the degrees of freedom.
12. The table below gives data on the monthly returns on the S&P 500 and small- cap stocks for the period January 1960 through December 1999 and provides statistics relating to their mean differences.
Measure S&P
500Return (%)
Small-CapStock Return (%)
Differences(S&P 500– Small-Cap Stock)
January 1960–December 1999, 480 months
Mean 1.0542 1.3117 –0.258 Standard deviation 4.2185 5.9570 3.752
January 1960–December 1979, 240 months
Mean 0.6345 1.2741 –0.640 Standard deviation
4.0807 6.5829 4.096
January 1980–December 1999, 240 months
Mean 1.4739 1.3492 0.125 Standard deviation 4.3197 5.2709 3.339
Let μd stand for the population mean value of difference between S&P 500 returns and small-cap stock returns. Use a significance level of 0.05 and suppose that mean differences are approximately normally distributed.
Formulate null and alternative hypotheses consistent with testing whether any difference exists between the mean returns on the S&P 500 and small- cap stocks.
Determine whether or not to reject the null hypothesis at the 0.05 significance level for the January 1960 to December 1999 period.
Determine whether or not to reject the null hypothesis at the 0.05 significance level for the January 1960 to December 1979 subperiod.
Determine whether or not to reject the null hypothesis at the 0.05 significance level for the January 1980 to December 1999 subperiod.
13. During a 10-year period, the standard deviation of annual returns on a portfolio you are analyzing was 15 percent a year. You want to see whether this record is sufficient evidence to support the conclusion that the portfolio’s underlying variance of return was less than 400, the return variance of the portfolio’s benchmark.
Formulate null and alternative hypotheses consistent with the verbal description of your objective.
Identify the test statistic for conducting a test of the hypotheses in Part A.
Identify the rejection point or points at the 0.05 significance level for the hypothesis tested in Part A.
Determine whether the null hypothesis is rejected or not rejected at the 0.05 level of significance.
14. You are investigating whether the population variance of returns on the S&P 500/BARRA Growth Index changed subsequent to the October 1987 market crash. You gather the following data for 120 months of returns before October 1987 and for 120 months of returns after October 1987. You have specified a 0.05 level of significance.
Time Period n Mean MonthlyReturn (%) Variance of Returns Before October 1987 120 1.416 22.367 After October 1987 120 1.436 15.795
Formulate null and alternative hypotheses consistent with the verbal description of the research goal.
Identify the test statistic for conducting a test of the hypotheses in Part A.
Determine whether or not to reject the null hypothesis at the 0.05 level of significance. (Use the F-tables in the back of this volume.)
15. You are interested in whether excess risk-adjusted return (alpha) is correlated with mutual fund expense ratios for US large-cap growth funds. The following table presents the sample. Mutual Fund 1 2 3 4 5 6 7 8 9 Alpha (X) −0.52 −0.13 −0.60 −1.01 −0.26 −0.89 −0.42 −0.23 −0.60
Expense Ratio (Y) 1.34 0.92 1.02 1.45 1.35 0.50 1.00 1.50 1.45
Formulate null and alternative hypotheses consistent with the verbal description of the research goal.
Identify the test statistic for conducting a test of the hypotheses in Part A.
Justify your selection in Part B.
Determine whether or not to reject the null hypothesis at the 0.05 level of significance.
16. All else equal, is specifying a smaller significance level in a hypothesis test likely to increase the probability of a:
Type I error? Type II error? A. No No B. No Yes C. Yes No
17. All else equal, is increasing the sample size for a hypothesis test likely to decrease the probability of a:
Type I error? Type II error? A. No Yes B. Yes No
C. Yes Yes
Notes 1 We discussed the construction and interpretation of confidence intervals in the
reading on sampling and estimation.
2 To be testable, a theory must be capable of making predictions that can be shown to be wrong.
3 See Freeley and Steinberg (2008) for a discussion of critical thinking applied to reasoned decision making.
4 This list is based on one in Daniel and Terrell (1995).
5 Greek letters, such as σ, are reserved for population parameters; Roman letters in italics, such as s, are used for sample statistics.
6 Part of this discussion of the selection of hypotheses follows Bowerman, O’Connell, and Murphree (2013, p. 342).
7 The central limit theorem says that the sampling distribution of the sample mean will be approximately normal with mean μ and variance σ2/n when the sample size is large. The sample we will use for this example has 111 observations.
8 The power of a test is, in fact, 1 minus the probability of a Type II error.
9 We do not always have information on the relative power of the test for competing test statistics, however.
10 “Rejection point” is a descriptive synonym for the more traditional term “critical value.”
11 The analogy in some courts of law (for example, in the United States) is that if a jury does not return a verdict of guilty (the alternative hypothesis), it is most accurate to say that the jury has failed to reject the null hypothesis, namely, that the defendant is innocent.
12 Just as with the hypothesis test, we can use this confidence interval, based on the standard normal distribution, when we have large samples. An alternative hypothesis test and confidence interval uses the t-distribution, which requires concepts that we introduce in the next section.
13 These issues are discussed further in the reading on sampling.
14 We can use spreadsheets to calculate p-values as well. In Microsoft Excel, for example, we may use the worksheet functions TTEST, NORMSDIST, CHIDIST, and FDIST to calculate p-values for t-tests, z-tests, chi-square tests, and F-tests, respectively.
15 Davidson and MacKinnon (1993) argued the merits of this approach: “The P value approach does not necessarily force us to make a decision about the null hypothesis. If we obtain a P value of, say, 0.000001, we will almost certainly want to reject the null. But if we obtain a P value of, say, 0.04, or even 0.004, we are not obliged to reject it. We may simply file the result away as information that casts some doubt on the null hypothesis, but that is not, by itself, conclusive. We believe that this somewhat agnostic attitude toward test statistics, in which they are merely regarded as pieces of information that we may or may not want to act upon, is usually the most sensible one to take.” (p. 80)
16 When we want to test whether the population means of more than two populations are equal, we use analysis of variance (ANOVA). We introduce ANOVA in its most common application, regression analysis, in the reading on correlation and regression analysis.
17 The formula for the variance of a t-distribution is df/(df − 2).
18 See Moore, McCabe, and Craig (2010). A statistic is robust if the required probability calculations are insensitive to violations of the assumptions.
19 Although this generalization is useful, we caution that the sample size needed to obtain an approximately normal sampling distribution for the sample mean depends on how non-normal the original population is. For some populations, “large” may be a sample size well in excess of 30.
20 A technical note, for reference, is required. When the sample comes from a finite population, estimates of the standard error of the mean, whether from Equation 2 or Equation 3, overestimate the true standard error. To address this, the computed standard error is multiplied by a shrinkage factor called the finite population correction factor (fpc), equal to , where N is the population size and n is the sample size. When the sample size is small relative to the population size (less than 5 percent of the population size), the fpc is usually ignored. The overestimation problem arises only in the usual situation of sampling without replacement (after an item is selected, it cannot be picked again) as opposed to sampling with replacement.
21 This measure represents the average length of time that the business must wait
after making a sale before receiving payment. The calculation is (Accounts receivable)/(Average sales per day).
22 These practitioners choose between t-tests and z-tests based on sample size. For small samples (n < 30), they use a t-test, and for large samples, a z-test.
23 For example, in Monte Carlo simulation, we prespecify the probability distributions for the risk factors. If we use a normal distribution, we know the true values of mean and variance. Monte Carlo simulation involves the use of a computer to represent the operation of a system subject to risk; we discuss Monte Carlo simulation in the reading on common probability distributions.
24 We can also use the following equivalent expression, which makes use of the correlation between the two variables: where is the sample variance of XA, is the sample variance of XB, and r(XA, XB) is the sample correlation between XA and XB.
25 Just as with other hypothesis tests, the chi-square test can be given a confidence interval interpretation. Unlike confidence intervals based on z- or t-statistics, however, chi-square confidence intervals for variance are asymmetric. A two- sided confidence interval for population variance, based on a sample of size n, has a lower limit and an upper limit . Under the null hypothesis, the hypothesized value of the population variance should fall within these two limits.
26 The relationship between the chi-square and F-distributions is as follows: If is one chi-square random variable with m degrees of freedom and is another chi- square random variable with n degrees of freedom, then follows an F-distribution with m numerator and n denominator degrees of freedom.
27 Some writers make a distinction between “nonparametric” and “distribution-free” tests. They refer to procedures that do not concern the parameters of a distribution as nonparametric and to procedures that make minimal assumptions about the underlying distribution as distribution free. We follow a commonly accepted, inclusive usage of the term nonparametric.
28 For large samples, there is often a transformation of the test statistic that permits the use of tables for the standard normal or t-distribution.
29 In some cases, there are several nonparametric alternatives to a parametric test.
30 See, for example, Hettmansperger and McKean (2010) or Siegel and Castellan
(1988).
31 We discussed scales of measurement in the reading on statistical concepts and market returns.
32 The t-test is described in the reading on correlation and regression. The assumption of the test is that each observation (x, y) on the two variables (X, Y) is a random observation from a bivariate normal distribution. Informally, in a bivariate or two-variable normal distribution, each individual variable is normally distributed and their joint relationship is completely described by the correlation, ρ, between them. For more details, see, for example, Daniel and Terrell (1995).
33 Calculating the usual correlation coefficient on the ranks would yield approximately the same result as Equation 17.
34 The expense ratio (the ratio of a fund’s operating expenses to average net assets) is bounded both from below (by zero) and from above. The Sharpe ratio is also observed within a limited range, in practice. Thus neither variable can be normally distributed, and hence jointly they cannot follow a bivariate normal distribution. In short, the assumptions of a t-test are not met.
35 To use a concept introduced in an earlier section, the parametric test is often more powerful.
36 See, for example, Hettmansperger and McKean (2010) or Siegel and Castellan (1988).
CHAPTER 8 Correlation and Regression RICHARD A. DEFUSCO, CFA
DENNIS W. MCLEAVEY, CFA
JERALD E. PINTO, PhD, CFA
DAVID E. RUNKLE, PhD, CFA
LEARNING OUTCOMES
After completing this chapter, you will be able to do the following:
calculate and interpret a sample covariance and a sample correlation coefficient and interpret a scatter plot;
describe limitations to correlation analysis;
formulate a test of the hypothesis that the population correlation coefficient equals zero and determine whether the hypothesis is rejected at a given level of significance;
distinguish between the dependent and independent variables in a linear regression;
describe the assumptions underlying linear regression and interpret regression coefficients;
calculate and interpret the standard error of estimate, the coefficient of determination, and a confidence interval for a regression coefficient;
formulate a null and alternative hypothesis about a population value of a regression coefficient and determine the appropriate test statistic and whether the null hypothesis is rejected at a given level of significance;
calculate the predicted value for the dependent variable, given an estimated regression model and a value for the independent variable;
calculate and interpret a confidence interval for the predicted value of the dependent variable;
describe the use of analysis of variance (ANOVA) in regression analysis,
interpret ANOVA results, and calculate and interpret the F-statistic;
describe limitations of regression analysis.
1. Introduction As a financial analyst, you will often need to examine the relationship between two or more financial variables. For example, you might want to know whether returns to different stock market indexes are related and, if so, in what way. Or you might hypothesize that the spread between a company’s return on invested capital and its cost of capital helps to explain the company’s value in the marketplace. Correlation and regression analysis are tools for examining these issues.
This reading1 is organized as follows. In Section 2, we present correlation analysis, a basic tool in measuring how two variables vary in relation to each other. Topics covered include the calculation, interpretation, uses, limitations, and statistical testing of correlations. Section 3 introduces basic concepts in regression analysis, a powerful technique for examining the ability of one or more variables (independent variables) to explain or predict another variable (the dependent variable).
2. Correlation Analysis We have many ways to examine how two sets of data are related. Two of the most useful methods are scatter plots and correlation analysis. We examine scatter plots first.
2.1. Scatter Plots A scatter plot is a graph that shows the relationship between the observations for two data series in two dimensions. Suppose, for example, that we want to graph the relationship between long-term money growth and long-term inflation in six industrialized countries to see how strongly the two variables are related. Table 1 shows the average annual growth rate in the money supply and the average annual inflation rate from 1980 to 2012 for the six countries.
TABLE 1 Annual Money Supply Growth Rate and Inflation Rate by Country, 1980– 2012
Country Money Supply Growth Rate (%) Inflation Rate (%) Australia 11.17 4.62 Japan 4.08 0.18
South Korea 17.81 5.31 Switzerland 5.85 1.99
United Kingdom 12.93 4.18 United States 6.53 2.93 Average 9.73 3.20
Source: International Monetary Fund.
FIGURE 1 Scatter Plot of Annual Money Supply Growth Rate and Inflation Rate by Country, 1980–2012
Source: International Monetary Fund.
To translate the data in Table 1 into a scatter plot, we use the data for each country to mark a point on a graph. For each point, the x-axis coordinate is the country’s annual average money supply growth from 1980–2012, and the y-axis coordinate is the country’s annual average inflation rate from 1980–2012. Figure 1 shows a scatter plot of the data in Table 1.
Note that each observation in the scatter plot is represented as a point, and the points are not connected. The scatter plot does not show which observation comes from which country; it shows only the actual observations of both data series plotted as pairs. For example, the rightmost point shows the data for South Korea. The data plotted in Figure 1 show a fairly strong linear relationship with a positive slope. Next we examine how to quantify this linear relationship.
2.2. Correlation Analysis In contrast to a scatter plot, which graphically depicts the relationship between two data series, correlation analysis expresses this same relationship using a single number. The correlation coefficient is a measure of how closely related two data series are. In particular, the correlation coefficient measures the direction and extent
of linear association between two variables. A correlation coefficient can have a maximum value of 1 and a minimum value of −1. A correlation coefficient greater than 0 indicates a positive linear association between the two variables: When one variable increases (or decreases), the other also tends to increase (or decrease). A correlation coefficient less than 0 indicates a negative linear association between the two variables: When one increases (or decreases), the other tends to decrease (or increase). A correlation coefficient of 0 indicates no linear relation between the two variables.2 Figure 2 shows the scatter plot of two variables with a correlation of 1.
FIGURE 2 Variables with a Correlation of 1
Note that all the points on the scatter plot in Figure 2 lie on a straight line with a positive slope. Whenever variable A increases by one unit, variable B increases by half a unit. Because all of the points in the graph lie on a straight line, an increase of one unit in A is associated with exactly the same half-unit increase in B, regardless of the level of A. Even if the slope of the line in the figure were different (but positive), the correlation between the two variables would be 1 as long as all the points lie on that straight line.
Figure 3 shows a scatter plot for two variables with a correlation coefficient of −1. Once again, the plotted observations fall on a straight line. In this graph, however, the line has a negative slope. As A increases by one unit, B decreases by half a unit, regardless of the initial value of A.
Figure 4 shows a scatter plot of two variables with a correlation of 0; they have no linear relation. This graph shows that the value of A tells us absolutely nothing
about the value of B.
2.3. Calculating and Interpreting the Correlation Coefficient To define and calculate the correlation coefficient, we need another measure of linear association: covariance. We have previously defined covariance as the expected value of the product of the deviations of two random variables from their respective population means. That was the definition of population covariance, which we would also use in a forward-looking sense. To study historical or sample correlations, we need to use sample covariance. The sample covariance of X and Y, for a sample of size n, is
FIGURE 3 Variables with a Correlation of –1
FIGURE 4 Variables with a Correlation of 0
(1)
The sample covariance is the average value of the product of the deviations of observations on two random variables from their sample means.3 If the random variables are returns, the unit of covariance would be returns squared.
The sample correlation coefficient is much easier to explain than the sample covariance. To understand the sample correlation coefficient, we need the expression for the sample standard deviation of a random variable X. We need to calculate the sample variance of X to obtain its sample standard deviation. The variance of a random variable is simply the covariance of the random variable with itself. The expression for the sample variance of X, , is
The sample standard deviation is the positive square root of the sample variance:
Both the sample variance and the sample standard deviation are measures of the dispersion of observations about the sample mean. Standard deviation uses the same units as the random variable; variance is measured in the units squared.
The formula for computing the sample correlation coefficient is
(2)
The correlation coefficient is the covariance of two variables (X and Y) divided by the product of their sample standard deviations (sX and sY). Like covariance, the correlation coefficient is a measure of linear association. The correlation coefficient, however, has the advantage of being a simple number, with no unit of measurement attached. It has no units because it results from dividing the covariance by the product of the standard deviations. Because we will be using sample variance, standard deviation, and covariance in this reading, we will repeat the calculations for these statistics.
Table 2 shows how to compute the various components of the correlation equation (Equation 2) from the data in Table 1.4 The individual observations on countries’ annual average money supply growth from 1980–2012 are denoted Xi, and individual observations on countries’ annual average inflation rate from 1980–2012 are denoted Yi. The remaining columns show the calculations for the inputs to correlation: the sample covariance and the sample standard deviations.
TABLE 2 Sample Covariance and Sample Standard Deviations: Annual Money Supply Growth Rate and Inflation Rate by Country, 1980–2012
Country Money Supply Growth Rate
Xi
Inflation Rate Yi
Cross- Product
Squared Deviations
Squared Deviations
Australia 0.1117 0.0462 0.000204 0.000208 0.000201 Japan 0.0408 0.0018 0.001707 0.003190 0.000913 South Korea 0.1781 0.0531 0.001704 0.006531 0.000445
Switzerland 0.0585 0.0199 0.000470 0.001504 0.000147 United
Kingdom 0.1293 0.0418 0.000313 0.001025 0.000096
United
States 0.0653 0.0293 0.000087 0.001023 0.000007
Sum 0.5837 0.1921 0.004485 0.013482 0.001809 Average 0.0973 0.0320
Covariance 0.000897 Variance 0.002696 0.000362 Standard deviation 0.051926 0.019019
Notes: 1 Divide the cross-product sum by n − 1 (with n = 6) to obtain the covariance of X and Y. 2 Divide the squared deviations sums by n − 1 (with n = 6) to obtain the variances of X and Y.
Source: International Monetary Fund.
Using the data shown in Table 2, we can compute the sample correlation coefficient for these two variables as follows:
The correlation coefficient of approximately 0.91 indicates a strong linear association between long-term money supply growth and long-term inflation for the countries in the sample. The correlation coefficient captures this strong association numerically, whereas the scatter plot in Figure 1 shows the information graphically.
What assumptions are necessary to compute the correlation coefficient? Correlation coefficients can be computed validly if the means and variances of X and Y, as well as the covariance of X and Y, are finite and constant. Later, we will show that when these assumptions are not true, correlations between two different variables can depend greatly on the sample that is used.
2.4. Limitations of Correlation Analysis Correlation measures the linear association between two variables, but it may not always be reliable. Two variables can have a strong nonlinear relation and still have a very low correlation. For example, the relation B = (A − 4)2 is a nonlinear relation contrasted to the linear relation B = 2A − 4. The nonlinear relation between variables A and B is shown in Figure 5. Below a level of 4 for A, variable B decreases with increasing values of A. When A is 4 or greater, however, B increases
whenever A increases. Even though these two variables are perfectly associated, the correlation between them is 0.5
FIGURE 5 Variables with a Strong Nonlinear Association
Correlation also may be an unreliable measure when outliers are present in one or both of the series. Outliers are small numbers of observations at either extreme (small or large) of a sample. Figure 6 shows a scatter plot of the monthly returns to the Standard & Poor ’s 500 Index and the monthly inflation rate in the United States from January 1990 through December 2013.
FIGURE 6 US Inflation and Stock Returns: 1990–2013
Sources: Bureau of Labor Statistics and S&P Dow Jones Indices.
In the scatter plot in Figure 6, most of the data lie clustered together with little discernible relation between the two variables. Two cases, however (the two circled observations), stand out from the rest. In one of those cases, inflation was extremely low at almost –2 percent, and in the other case, stock returns were strongly negative at almost –17 percent. These observations are outliers. If we compute the correlation coefficient for the entire data sample, that correlation is −0.0350. If we eliminate the two outliers, however, the correlation is −0.1489.
The correlation in Figure 6 is quite sensitive to excluding only two observations. Does it make sense to exclude those observations? Are they noise or news? When the outliers are excluded, there seems to be a moderately negative correlation between inflation and stock returns. One possible partial explanation of this negative correlation is that whenever inflation was very high during a month, market participants became concerned that the Federal Reserve would raise interest rates, which would cause the value of stocks to decline. This story offers one plausible explanation for how investors reacted to large inflation announcements. When the two outliers are included, there is a noticeable decrease in the magnitude of the negative correlation. A closer examination of the monthly data used in the scatter plot reveals that the two outliers correspond to the months of October and November 2008 when bad news regarding the US economy and job market caused the stock market to decline sharply. During these two months, although the inflation
was not high (in fact, inflation was negative in both months), stocks declined substantially. Therefore, inclusion of those two outliers reduces the magnitude of the negative correlation between inflation and stock returns. One could argue that while the data without the outliers provide a useful insight into the general relationship between inflation and stock returns, the outliers may provide information about the relationship during a period of market distress. Therefore, in this case, it would be reasonable to report the values of the correlation including and excluding the outliers.
As a general rule, we must determine whether a computed sample correlation changes greatly by removing a few outliers. But we must also use judgment to determine whether those outliers contain information about the two variables’ relationship (and should thus be included in the correlation analysis) or contain no information (and should thus be excluded).
Keep in mind that correlation does not imply causation. Even if two variables are highly correlated, one does not necessarily cause the other in the sense that certain values of one variable bring about the occurrence of certain values of the other. Furthermore, correlations can be spurious in the sense of misleadingly pointing towards associations between variables.
The term spurious correlation has been used to refer to 1) correlation between two variables that reflects chance relationships in a particular data set, 2) correlation induced by a calculation that mixes each of two variables with a third, and 3) correlation between two variables arising not from a direct relation between them but from their relation to a third variable. As an example of the second kind of spurious correlation, two variables that are uncorrelated may be correlated if divided by a third variable. As an example of the third kind of spurious correlation, height may be positively correlated with the extent of a person’s vocabulary, but the underlying relationships are between age and height and between age and vocabulary. Investment professionals must be cautious in basing investment strategies on high correlations. Spurious correlation may suggest investment strategies that appear profitable but actually would not be so, if implemented.
2.5. Uses of Correlation Analysis In this section, we give examples of correlation analysis for investment. Because investors’ expectations about inflation are important in determining asset prices, inflation forecast accuracy will serve as our first example.
EXAMPLE 1 Evaluating Economic Forecasts (1)
Investors closely watch economists’ forecasts of inflation, but do these forecasts contain useful information? In the euro area, the Survey of Professional Forecasters (SPF) gathers professional forecasters’ predictions about many economic variables.6 Since 1999, SPF has gathered predictions on the euro area inflation rate using the change in the Harmonised Index of Consumer Prices (HICP) for the prices of consumer goods and services acquired by households to measure inflation. If these forecasts of inflation could perfectly predict actual inflation, the correlation between forecasts and inflation would be 1.
Figure 7 shows a scatter plot of the mean forecast made in the first quarter of a year for the percentage change in HICP during that year and the actual percentage change in HICP, from 1999 through 2013.7 In this scatter plot, the forecast for each year is plotted on the x-axis and the actual change in the HICP is plotted on the y-axis.
FIGURE 7 Actual Change in Euro Area HICP versus Predicted Change
Source: European Central Bank.
Discuss whether professional forecasters’ predictions of the euro area inflation might be useful in investment decision making.
Solution: As Figure 7 shows, a fairly strong linear association exists between the forecast and the actual inflation rate, suggesting that professional forecasts of inflation might be useful in investment decision making. In fact, the correlation between the two series is 0.8913. Although there is no causal relation here, there is a direct relation because forecasters assimilate information to forecast inflation.
One important issue in evaluating a portfolio manager ’s performance is determining an appropriate benchmark for the manager. Since the early 1990s, style analysis has been an important component of benchmark selection.8
EXAMPLE 2 Style Analysis Correlations
Portfolio managers using small-cap stocks in investment portfolios may favor a growth style, a value style, or neither.
In the United States, the Russell 2000 Growth Index and the Russell 2000 Value Index are often used as benchmarks for small-cap growth and small-cap value managers, respectively. Correlation analysis shows, however, that the returns to these two indexes are very closely associated with each other. For the 15 years ending in 2013 (January 1999 to December 2013), the correlation between the monthly returns to the Russell 2000 Growth Index and the Russell 2000 Value Index was 0.8249.
What conclusions can be drawn based on this result concerning the mean returns to small-cap growth and small-cap value investment styles? Explain your answer.
Solution: The returns to the two indexes are highly positively correlated. But correlation does not provide information on variables’ mean returns, only on how their returns covary. Here, for example, even a correlation of +1 to the returns to the two styles would not imply that the mean returns to the two styles are the same. Thus, the information given is not sufficient to reach a conclusion on the relative mean returns to the small-cap growth and small-cap investment styles.
The previous examples in this reading have examined the correlation between two variables. Often, however, investment managers need to understand the correlations
among many asset returns. For example, investors who have any exposure to movements in exchange rates must understand the correlations of the returns to different foreign currencies and other assets in order to determine their optimal portfolios and hedging strategies.9 In the following example, we see how a correlation matrix shows correlation between pairs of variables when we have more than two variables. We also see one of the main challenges to investment managers: Investment return correlations can change substantially over time.
EXAMPLE 3 Exchange Rate Return Correlations
The exchange rate return measures the periodic domestic currency return to holding foreign currency. Consider a British investor with British pounds (GBP) as her domestic currency. Suppose a change in inflation rates in Canada and the United Kingdom results in the pound price of a Canadian dollar changing from £0.50 to £0.45. If this change occurred in one month, the return in that month to holding Canadian dollars would be (0.45 – 0.50)/0.50 = –10 percent, in terms of pounds.
Table 3 shows a correlation matrix of monthly returns in British pounds to holding Canadian, Japanese, Swedish, or US currencies during two seven-year periods of 2000–2006 and 2007–2013. To interpret a correlation matrix, we first examine the top panel of this table.
Table 3 Correlations of Monthly British Pound Returns to Selected Foreign Currency Returns
2000–2006 Canada Japan Sweden United States Canada 1.0000 Japan 0.4552 1.0000 Sweden 0.2686 0.1832 1.0000
United States 0.6917 0.4360 0.0074 1.0000
2007–2013 Canada Japan Sweden United States Canada 1.0000 Japan 0.3091 1.0000 Sweden 0.5278 0.1742 1.0000
United States 0.5263 0.7230 0.1862 1.0000
Source: www.oanda.com/currency/historical-rates/.
The first column of numbers of that panel shows the correlations between GBP returns to holding the Canadian dollar and GBP returns to holding Canadian, Japanese, Swedish, and US currencies during 2000–2006. Of course, any variable is perfectly correlated with itself, and so the correlation between GBP returns to holding the Canadian dollar and GBP returns to holding the Canadian dollar is 1. The second row of this column shows that the correlation between GBP returns to holding the Canadian dollar and GBP returns to holding the Japanese yen was 0.4552 during 2000–2006. The remaining correlations in the panel show how the GBP returns to other combinations of currency holdings were correlated during this period.
1. Explain why Table 3 omits many of the correlations.
Solution to 1: The formula for correlation coefficient in the earlier equation (equation 2) shows that correlations are always symmetrical: The correlation between X and Y is always the same as the correlation between Y and X. Accordingly, duplicative coefficients are excluded in Table 3. For example, Column 2 of the panels omits the correlation between GBP returns to holding yen and GBP returns to holding Canadian dollars. This correlation is omitted because it is identical to the correlation between GBP returns to holding Canadian dollars and GBP returns to holding yen shown in Row 2 of Column 1. Similarly, other omitted correlations would also have been duplicative.
2. Compare the two panels of Table 3 and discuss whether the changes in correlations from 2000–2006 to 2007–2013 show a pattern.
Solution to 2: A comparison of the two panels of Table 3 shows that that many of the currency return correlations changed dramatically between the periods of 2000–2006 and 2007–2013, but there is no pattern in these changes. During 2000–2006, for example, the correlation between the return to holding Canadian dollars and the return to holding Japanese yen (0.4552) was about the same as the correlation between the return to holding yen and the return to holding US dollars (0.4360). During 2007–2013, however, the correlation between Canadian dollar returns and yen returns dropped substantially (to 0.3091), but the correlation between yen and US dollar returns increased substantially (to 0.7230). Some other correlations also increased markedly. For example, the correlation between Canadian dollar returns and Swedish krona returns almost doubled from 0.2686 to 0.5278 and the correlation between krona and US dollar returns increased from 0.0074 to 0.1862. In contrast, the correlation between Canadian and US dollars decreased from 0.6917 to 0.5263 and the correlation between yen and krona returns hardly changed (0.1832 to
0.1742).
Optimal asset allocation depends on expectations of future correlations. With less than perfect positive correlation between two assets’ returns, there are potential risk-reduction benefits to holding both assets. Expectations of future correlation may be based on historical sample correlations, but the variability in historical sample correlations poses challenges. We discuss these issues in detail in the reading on portfolio concepts.
In the next example, we extend the discussion of the correlations of stock market indexes begun in Example 2 to indexes representing large-cap, small-cap, and broad-market returns. This type of analysis has serious diversification and asset allocation consequences because the strength of the correlations among the assets tells us how successfully the assets can be combined to diversify risk.
EXAMPLE 4 Correlations among Stock Return Series
Table 4 shows the correlation matrix of monthly returns to three UK stock indexes during the period January 1990 to December 2009 and in two subperiods (the 1990s and 2000s). The large-cap style is represented by the return to the FTSE 100 Index, the small-cap style is represented by the return to the FTSE Small Cap Excluding Investment Companies Index, and the broad- market returns are represented by the return to the FTSE All-Share Index.
TABLE 4 Correlations of Monthly Returns to Various UK Stock Indexes
1990–2009 FTSE 100 FTSE Small Cap FTSE All-Share FTSE 100 1.0000
FTSE Small Cap 0.6914 1.0000 FTSE All-Share 0.9906 0.7694 1.0000 1990–1999 FTSE 100 FTSE Small Cap FTSE All-Share FTSE 100 1.0000
FTSE Small Cap 0.6553 1.0000 FTSE All-Share 0.9873 0.7553 1.0000
FTSE 100 1.0000 FTSE Small Cap 0.7245 1.0000 FTSE All-Share 0.9937 0.7869 1.0000
Source: CompuSmart Global.
Discuss whether the correlation coefficients for the entire sample are consistent with your expectations.
Solution: The first column of numbers in the top panel of Table 4 shows nearly perfect positive correlation between returns to the FTSE 100 and returns to the FTSE All-Share: The correlation between the two return series is 0.9906. This result should not be surprising, because both the FTSE 100 and the FTSE All- Share are value-weighted indexes, and large-cap stock returns receive most of the weight in both indexes. In fact, the companies that make up the FTSE 100 have more than 80 percent of the total market value of all companies included in the FTSE All-Share.
Small-cap stocks also have a reasonably high correlation with large stocks. In the total sample, the correlation between the FTSE 100 returns and the FTSE Small Cap returns is 0.6914. The correlation between FTSE Small Cap returns and returns to the FTSE All-Share is slightly higher (0.7694). This result is also not too surprising because the FTSE All-Share contains small-cap stocks and the FTSE 100 does not.
The second and third panels of Table 4 show that correlations among the various stock market return series show some variation from decade to decade. For example, the correlation between returns to the FTSE 100 and FTSE Small Cap stocks increased from 0.6553 in the 1990s to 0.7245 in the 2000s.10
For asset allocation purposes, correlations among asset classes are studied carefully with a view toward maintaining appropriate diversification based on forecasted correlations.
EXAMPLE 5 Correlations of Debt and Equity Returns
Table 5 shows the correlation matrix for various US debt returns and US large and small company stock returns using monthly data from January 1926 to December 2012.
TABLE 5 Correlations among US Stock and Debt Returns, 1926–2012
All US Large Co.Stocks US Small Co.
Stocks US Long- Term Corp.
US Long- Term Govt.
US T- Bills
US Large Co. Stocks
1.00
US Small Co. Stocks 0.79 1.00
US Long- Term Corp. 0.16 0.06 1.00
US Long- Term Govt. 0.01 –0.08 0.89 1.00
US T-Bills –0.01 –0.09 0.15 0.18 1.00
Source: Ibbotson Associates.
The first column of numbers, in particular, shows the correlations of US large company stock returns with small company stock returns and various debt returns. As expected, large and small company stocks have a high correlation (0.79). In contrast, large company stock returns are almost completely uncorrelated (−0.01) with Treasury bill returns for this period. Long-term corporate debt returns are somewhat more correlated (0.16) with large company stock returns.
Long-term government bonds, however, have a very low correlation (0.01) with large company stock returns. We expect some correlation between these variables because interest rate increases reduce the present value of future cash flows for both bonds and stocks. The low correlation between these two return series, however, shows that other factors affect the returns on stocks besides interest rates. Without these other factors, the correlation between bond and stock returns would be higher.
The third column of numbers in Table 5 shows that the correlation between long-term government bond and corporate bond returns is quite high (0.89) for this time period. Although this correlation is the highest in the entire matrix, it is not 1. The correlation is less than 1 because the default premium for long- term corporate bonds changes, whereas US government bonds do not incorporate a default premium. As a result, changes in required yields for government bonds have a correlation less than 1 with changes in required yields for corporate bonds, and return correlations between government bonds and corporate bonds are also below 1. Note also that T-bill returns have a very low correlation with all other return series.
In the final example of this section, correlation is used in a financial statement setting to show that net income is an inadequate proxy for cash flow.
EXAMPLE 6 Correlations among Net Income, Cash Flow from Operations, and Free Cash Flow to the Firm
Net income (NI), cash flow from operations (CFO), and free cash flow to the firm (FCFF) are three measures of company performance that analysts often use to value companies. Differences in these measures for given companies would not cause differences in the relative valuation if the measures were highly correlated.
CFO equals net income plus the net noncash charges that were subtracted to obtain net income, minus the company’s investment in working capital during the same time period. FCFF equals CFO plus net-of-tax interest expense, minus the company’s investment in fixed capital over the time period. FCFF may be interpreted as the cash flow available to the company’s suppliers of capital (debtholders and shareholders) after all operating expenses have been paid and necessary investments in working and fixed capital have been made.11
Some analysts base their valuations only on NI, ignoring CFO and FCFF. If the correlations among NI, CFO, and FCFF were very high, then an analyst’s decision to ignore CFO and FCFF would be easy to understand because NI would then appear to capture everything one needs to know about cash flow.
Table 6 shows the correlations among NI, CFO, and FCFF for a group of six publicly traded US companies involved in retailing women’s clothing for 2001. Before computing the correlations, we normalized all of the data by dividing each company’s three performance measures by the company’s revenue for the year.12
Because CFO and FCFF include NI as a component (in the sense that CFO and FCFF can be obtained by adding and subtracting various quantities from NI), we might expect that the correlations between NI and CFO and between NI and FCFF would be positive. Table 6 supports that conclusion. These correlations with NI, however, are much smaller than the correlation between CFO and FCFF (0.8217). The lowest correlation in the table is between NI and FCFF (0.4045). This relatively low correlation shows that NI contained some but far from all the information in FCFF for these companies in 2001. Later in this reading, we will test whether the correlation between NI and FCFF is significantly different from zero.
TABLE 6 Correlations among Performance Measures: US Women’s Clothing Stores, 2001
NI CFO FCFF NI 1.0000 CFO 0.6959 1.0000 FCFF 0.4045 0.8217 1.0000
Source: Compustat.
2.6. Testing the Significance of the Correlation Coefficient Significance tests allow us to assess whether apparent relationships between random variables are the result of chance. If we decide that the relationships do not result from chance, we will be inclined to use this information in predictions because a good prediction of one variable will help us predict the other variable. Using the data in Table 2, we calculated 0.9083 as the sample correlation between long-term money growth and long-term inflation in six industrialized countries between 1980 and 2012. That estimated correlation seems high, but is it significantly different from 0? Before we can answer this question, we must know some details about the distribution of the underlying variables themselves. For purposes of simplicity, let us assume that both of the variables are normally distributed.13
We propose two hypotheses: the null hypothesis, H0, that the correlation in the population is 0 (ρ = 0); and the alternative hypothesis, Ha, that the correlation in the population is different from 0 (ρ ≠ 0).
The alternative hypothesis is a test that the correlation is not equal to 0; therefore, a two-tailed test is appropriate. As long as the two variables are distributed normally, we can test to determine whether the null hypothesis should be rejected using the sample correlation, r. The formula for the t-test is
(3)
This test statistic has a t-distribution with n − 2 degrees of freedom if the null hypothesis is true. One practical observation concerning Equation 3 is that the magnitude of r needed to reject the null hypothesis H0: ρ = 0 decreases as sample size n increases, for two reasons. First, as n increases, the number of degrees of freedom increases and the absolute value of the critical value tc decreases. Second, the absolute value of the numerator increases with larger n, resulting in larger- magnitude t-values. For example, with sample size n = 12, r = 0.58 results in a t- statistic of 2.252 that is just significant at the 0.05 level (tc = 2.228). With a sample
size n = 32, a smaller sample correlation r = 0.35 yields a t-statistic of 2.046 that is just significant at the 0.05 level (tc = 2.042); the r = 0.35 would not be significant with a sample size of 12 even at the 0.10 significance level. Another way to make this point is that sampling from the same population, a false null hypothesis H0: ρ = 0 is more likely to be rejected as we increase sample size, all else equal.
EXAMPLE 7 Testing the Correlation between Money Supply Growth and Inflation
Earlier in this reading, we showed that the sample correlation between long- term money supply growth and long-term inflation in six industrialized countries was 0.9083 during the 1980–2012 period. Suppose we want to test the null hypothesis, H0, that the true correlation in the population is 0 (ρ = 0) against the alternative hypothesis, Ha, that the correlation in the population is different from 0 (ρ ≠ 0).
1. Calculate the test statistic to test the null hypothesis given above.
2. Determine whether the null hypothesis is rejected or not rejected at the 0.05 level of significance.
Solution to 1: Recalling that this sample has six observations, we can compute the statistic for testing the null hypothesis as follows:
The value of the test statistic is 4.343.
Solution to 2: As the table of critical values of the t-distribution for a two-tailed test shows, for a t-distribution with n − 2 = 6 − 2 = 4 degrees of freedom at the 0.05 level of significance, we can reject the null hypothesis (that the population correlation is equal to 0) if the value of the test statistic is greater than 2.776 or less than −2.776. The fact that we can reject the null hypothesis of no correlation based on only six observations is quite unusual; it further demonstrates the strong relation between long-term money supply growth and long-term inflation in these six countries.
EXAMPLE 8 Testing the Yen–Canadian Dollar Return Correlation
The data in Table 3 showed that the sample correlation between the GBP monthly returns to Japanese yen and Canadian dollar was 0.3091 for the period from January 2007 through December 2013.
Can we reject a null hypothesis that the underlying or population correlation equals 0 at the 0.05 level of significance?
Solution: With 84 months from January 2007 through December 2013, we use the following statistic to test the null hypothesis, H0, that the true correlation in the population is 0, against the alternative hypothesis, Ha, that the correlation in the population is different from 0:
At the 0.05 significance level, the critical level for this test statistic is 1.99 (n = 84, degrees of freedom = 82). When the test statistic is either larger than 1.99 or smaller than −1.99, we can reject the hypothesis that the correlation in the population is 0. The test statistic is 2.9431, so we can reject the null hypothesis.
Note that the sample correlation coefficient in this case is significantly different from 0 at the 0.05 level, even though the coefficient is much smaller than that in the previous example. The correlation coefficient, though smaller, is still significant because the sample is much larger (84 observations instead of 6 observations).
The above example shows the importance of sample size in tests of the significance of the correlation coefficient. The following example also shows the importance of sample size and examines the relationship at the 0.01 level of significance as well as at the 0.05 level.
EXAMPLE 9 The Correlation between Bond Returns and T-Bill Returns
Table 5 showed that the sample correlation between monthly returns to US long-term government bonds and monthly returns to T-bills was 0.18 from January 1926 through December 2012.
Can we reject a null hypothesis that the underlying or population correlation coefficient equals 0 at the 0.05 and 0.01 levels of significance?
Solution: There are 1,044 months during the period January 1926 to December
2012. Therefore, to test the null hypothesis, H0 (that the true correlation in the population is 0), against the alternative hypothesis, Ha (that the correlation in the population is different from 0), we use the following test statistic:
At the 0.05 significance level, the critical value for the test statistic is approximately 1.96. At the 0.01 significance level, the critical value for the test statistic is approximately 2.58. The test statistic is 5.9069, so we can reject the null hypothesis of no correlation in the population at both the 0.05 and 0.01 levels. This example shows that, in large samples, even relatively small correlation coefficients can be significantly different from zero.
In the final example of this section, we explore another situation of small sample size.
EXAMPLE 10 Testing the Correlation between Net Income and Free Cash Flow to the Firm
Earlier in this reading, we showed that the sample correlation between NI and FCFF for six women’s clothing stores was 0.4045 in 2001. Suppose we want to test the null hypothesis, H0, that the true correlation in the population is 0 (ρ = 0) against the alternative hypothesis, Ha, that the correlation in the population is different from 0 (ρ ≠ 0). Recalling that this sample has six observations, we can compute the statistic for testing the null hypothesis as follows:
With n − 2 = 6 − 2 = 4 degrees of freedom and a 0.05 significance level, we reject the null hypothesis that the population correlation equals 0 for values of the test statistic greater than 2.776 or less than −2.776. In this case, however, the t-statistic is 0.8846, so we cannot reject the null hypothesis. Therefore, for this sample of women’s clothing stores, there is no statistically significant correlation between NI and FCFF, when each is normalized by dividing by sales for the company.14
The scatter plot creates a visual picture of the relationship between two variables, while the correlation coefficient quantifies the existence of any linear relationship. Large absolute values of the correlation coefficient indicate strong linear
relationships. Positive coefficients indicate a positive relationship and negative coefficients indicate a negative relationship between two data sets. In Examples 8 and 9, we saw that relatively small sample correlation coefficients (0.3091 and 0.18, respectively) can be statistically significant and thus might provide valuable information about the behavior of economic variables.
Next we will introduce linear regression, another tool useful in examining the relationship between two variables.
3. Linear Regression Linear regression with one independent variable, sometimes called simple linear regression, models the relationship between two variables as a straight line. When the linear relationship between the two variables is significant, linear regression provides a simple model for forecasting the value of one variable, known as the dependent variable, given the value of the second variable, known as the independent variable. The following sections explain linear regression in more detail.
3.1. Linear Regression with One Independent Variable As a financial analyst, you will often want to understand the relationship between financial or economic variables, or to predict the value of one variable using information about the value of another variable. For example, you may want to know the impact of changes in the 10-year Treasury bond yield on the earnings yield of the S&P 500 (the earnings yield is the reciprocal of the price-to-earnings ratio). If the relationship between those two variables is linear, you can use linear regression to summarize it.
Linear regression allows us to use one variable to make predictions about another, test hypotheses about the relation between two variables, and quantify the strength of the relationship between the two variables. The remainder of this reading focuses on linear regression with a single independent variable. In the next reading, we will examine regression with more than one independent variable.
Regression analysis begins with the dependent variable (denoted Y), the variable that you are seeking to explain. The independent variable (denoted X) is the variable you are using to explain changes in the dependent variable. For example, you might try to explain small-stock returns (the dependent variable) based on returns to the S&P 500 (the independent variable). Or you might try to explain inflation (the dependent variable) as a function of growth in a country’s money supply (the independent variable).
Linear regression assumes a linear relationship between the dependent and the independent variables. The following regression equation describes that relation:
(4)
This equation states that the dependent variable, Y, is equal to the intercept, b0, plus a slope coefficient, b1, times the independent variable, X, plus an error term, ε. The error term represents the portion of the dependent variable that cannot be
explained by the independent variable. We refer to the intercept b0 and the slope coefficient b1 as the regression coefficients.
Regression analysis uses two principal types of data: cross-sectional and time series. Cross-sectional data involve many observations on X and Y for the same time period. Those observations could come from different companies, asset classes, investment funds, people, countries, or other entities, depending on the regression model. For example, a cross-sectional model might use data from many companies to test whether predicted earnings-per-share growth explains differences in price-to- earnings ratios (P/Es) during a specific time period. The word “explain” is frequently used in describing regression relationships. One estimate of a company’s P/E that does not depend on any other variable is the average P/E. If a regression of a P/E on an independent variable tends to give more accurate estimates of P/E than just assuming that the company’s P/E equals the average P/E, we say that the independent variable helps explain P/Es because using that independent variable improves our estimates. Finally, note that if we use cross-sectional observations in a regression, we usually denote the observations as i = 1, 2, …, n.
Time-series data use many observations from different time periods for the same company, asset class, investment fund, person, country, or other entity, depending on the regression model. For example, a time-series model might use monthly data from many years to test whether US inflation rates determine US short-term interest rates.15 If we use time-series data in a regression, we usually denote the observations as t = 1, 2, …, T.16
Exactly how does linear regression estimate b0 and b1? Linear regression, also known as linear least squares, computes a line that best fits the observations; it chooses values for the intercept, b0, and slope, b1, that minimize the sum of the squared vertical distances between the observations and the regression line. Linear regression chooses the estimated parameters or fitted parameters and in Equation 4 to minimize17
(5)
In this equation, the term means (dependent variable – predicted value of dependent variable)2. Using this method to estimate the values of and , we can fit a line through the observations on X and Y that best explains the value that Y takes for any particular value of X.18
FIGURE 8 Fitted Regression Line Explaining the Inflation Rate Using Growth in the Money Supply by Country, 1980–2012
Source: International Monetary Fund.
Note that we never observe the population parameter values b0 and b1 in a regression model. Instead, we observe only and , which are estimates of the population parameter values. Thus predictions must be based on the parameters’ estimated values, and testing is based on estimated values in relation to the hypothesized population values.
Figure 8 gives a visual example of how linear regression works. The figure shows the linear regression that results from estimating the regression relation between the annual rate of inflation (the dependent variable) and annual rate of money supply growth (the independent variable) for six industrialized countries from 1980 to 2012 (n = 6).19 The equation to be estimated is Long-term rate of inflation = b0 + b1 (Long-term rate of money supply growth) + ε.
The distance from each of the six data points to the fitted regression line is the regression residual, which is the difference between the actual value of the dependent variable and the predicted value of the dependent variable made by the regression equation. Linear regression chooses the estimated coefficients and in Equation 4 such that the sum of the squared vertical distances is minimized. The estimated regression equation is Long-term inflation = –0.0003 + 0.3327 (Long- term money supply growth).20
According to this regression equation, if the long-term money supply growth is 0
for any particular country, the long-term rate of inflation in that country will be – 0.03 percent. For every 1-percentage-point increase in the long-term rate of money supply growth for a country, the long-term inflation rate is predicted to increase by 0.3327 percentage points. In a regression such as this one, which contains one independent variable, the slope coefficient equals Cov(Y, X)/Var(X). We can solve for the slope coefficient using data from Table 2, excerpted here:
TABLE 2 (Excerpt)
Money Supply Growth Rate Xi
Inflation Rate Yi
Cross- Product
Squared Deviations
Squared Deviations
Sum 0.5837 0.1921 0.004485 0.013481 0.001809 Average 0.0973 0.0320
Covariance 0.000897 Variance 0.002696 0.000362 Standard deviation 0.051926 0.019019
In a linear regression, the regression line fits through the point corresponding to the means of the dependent and the independent variables. As shown in Table 1 (excerpted below), from 1980 to 2012, the mean long-term growth rate of the money supply for these six countries was 9.73 percent, whereas the mean long-term inflation rate was 3.20 percent.
Because the point (9.73, 3.20) lies on the regression line , we can solve for the intercept using this point as follows:
We are showing how to solve the linear regression equation step by step to make the source of the numbers clear. Typically, an analyst will use the data analysis function on a spreadsheet or a statistical package to perform linear regression analysis. Later, we will discuss how to use regression residuals to quantify the uncertainty in a regression model.
Table 1 (Excerpt)
Money Supply Growth Rate Inflation Rate Average 9.73% 3.20%
3.2. Assumptions of the Linear Regression Model We have discussed how to interpret the coefficients in a linear regression model. Now we turn to the statistical assumptions underlying this model. Suppose that we have n observations on both the dependent variable, Y, and the independent variable, X, and we want to estimate Equation 4:
To be able to draw valid conclusions from a linear regression model with a single independent variable, we need to make the following six assumptions, known as the classic normal linear regression model assumptions: 1. The relationship between the dependent variable, Y, and the independent
variable, X is linear in the parameters b0 and b1. This requirement means that b0 and b1 are raised to the first power only and that neither b0 nor b1 is multiplied or divided by another regression parameter (as in b0/b1, for example). The requirement does not exclude X from being raised to a power other than 1.
2. The independent variable, X, is not random.21
3. The expected value of the error term is 0: E(ε) = 0.
4. The variance of the error term is the same for all observations: , i = 1, …, n.
5. The error term, ε, is uncorrelated across observations. Consequently, E(εiεj) = 0 for all i not equal to j.22
6. The error term, ε, is normally distributed.23
Now we can take a closer look at each of these assumptions.
Assumption 1 is critical for a valid linear regression. If the relationship between the independent and dependent variables is nonlinear in the parameters, then estimating that relation with a linear regression model will produce invalid results. For example, is nonlinear in b1, so we could not apply the linear regression model to it.24
Even if the dependent variable is nonlinear, linear regression can be used as long as the regression is linear in the parameters. So, for example, linear regression can be used to estimate the equation .
Assumptions 2 and 3 ensure that linear regression produces the correct estimates of b0 and b1.
Assumptions 4, 5, and 6 let us use the linear regression model to determine the distribution of the estimated parameters and and thus test whether those coefficients have a particular value.
Assumption 4, that the variance of the error term is the same for all observations, is also known as the homoskedasticity assumption. The reading on regression analysis discusses how to test for and correct violations of this assumption.
Assumption 5, that the errors are uncorrelated across observations, is also necessary for correctly estimating the variances of the estimated parameters and . The reading on multiple regression discusses violations of this assumption.
Assumption 6, that the error term is normally distributed, allows us to easily test a particular hypothesis about a linear regression model.25
EXAMPLE 11 Evaluating Economic Forecasts (2)
If economic forecasts were completely accurate, every prediction of change in an economic variable in a quarter would exactly match the actual change that occurs in that quarter. Even though forecasts can be inaccurate, we hope at least that they are unbiased—that is, that the expected value of the forecast error is zero. An unbiased forecast can be expressed as E(Actual change – Predicted change) = 0. In fact, most evaluations of forecast accuracy test whether forecasts are unbiased.26
Figure 9 repeats Figure 7 in showing a scatter plot of the mean forecast made in the first quarter of a year for the percentage change in HICP during that year and the actual percentage change in HICP, from 1999 through 2013, but it adds the fitted regression line for the equation Actual percentage change = b0 + b1 (Predicted percentage change) + ε. If the forecasts are unbiased, the intercept, b0, should be 0 and the slope, b1, should be 1. We should also find E(Actual change – Predicted change) = 0. If forecasts are actually unbiased, as long as b0
= 0 and b1 = 1, the error term [Actual change − b0 − b1(Predicted change)] will have an expected value of 0, as required by Assumption 3 of the linear regression model. With unbiased forecasts, any other values of b0 and b1 would yield an error term with an expected value different from 0.
FIGURE 9 Actual Change in Euro Area HICP versus Predicted Change
Source: European Central Bank.
If b0 = 0 and b1 = 1, our best guess of actual change in HICP would be 0 if professional forecasters’ predictions of change in HICP were 0. For every 1- percentage-point increase in the prediction of change by the professional forecasters, the regression model would predict a 1-percentage-point increase in actual change.
The fitted regression line in Figure 9 comes from the equation Actual change = −0.7006 + 1.5538(Predicted change). It seems that the estimated values of b0 and b1 are not particularly close to the values b0 = 0 and b1 = 1 that are consistent
with unbiased forecasts. Later in this reading, we discuss how to test the hypotheses that b0 = 0 and b1= 1.
3.3. The Standard Error of Estimate The linear regression model sometimes describes the relationship between two variables quite well, but sometimes it does not. We must be able to distinguish between these two cases in order to use regression analysis effectively. Therefore, in this section and the next, we discuss statistics that measure how well a given linear regression model captures the relationship between the dependent and independent variables.
FIGURE 10 Fitted Regression Line Explaining Stock Returns by Inflation during 1990–2013
Sources: Bureau of Labor Statistics and S&P Dow Jones Indices.
Figure 9, for example, shows a strong relation between predicted inflation and actual inflation. If we knew professional forecasters’ predictions for inflation in a particular quarter, we would be reasonably certain that we could use this regression model to forecast actual inflation relatively accurately.
In other cases, however, the relation between the dependent and independent variables is not strong. Figure 10 adds a fitted regression line to the data on inflation and stock returns during 1990 to 2013 from Figure 6. In this figure, the
actual observations are generally much farther from the fitted regression line than in Figure 9. Using the estimated regression equation to predict monthly stock returns assuming a particular level of inflation might result in an inaccurate forecast.
As noted, the regression relation in Figure 10 is less precise than that in Figure 9. The standard error of estimate (sometimes called the standard error of the regression) measures this uncertainty. This statistic is very much like the standard deviation for a single variable, except that it measures the standard deviation of , the residual term in the regression.
The formula for the standard error of estimate (SEE) for a linear regression model with one independent variable is
(6)
In the numerator of this equation, we are computing the difference between the dependent variable’s actual value for each observation and its predicted value
for each observation. The difference between the actual and predicted values of the dependent variable is the regression residual, .
Equation 6 looks very much like the formula for computing a standard deviation, except that n − 2 appears in the denominator instead of n − 1. We use n − 2 because the sample includes n observations and the linear regression model estimates two parameters ( and ); the difference between the number of observations and the number of parameters is n − 2. This difference is also called the degrees of freedom; it is the denominator needed to ensure that the estimated standard error of estimate is unbiased.
EXAMPLE 12 Computing the Standard Error of Estimate
Recall that the estimated regression equation for the inflation and money supply growth data shown in Figure 8 was Yi = –0.0003 + 0.3327Xi. Table 7 uses this estimated equation to compute the data needed for the standard error of estimate.
The first and second columns of numbers in Table 7 show the long-term money supply growth rates, Xi, and long-term inflations rates, Yi, for the six countries. The third column of numbers shows the predicted value of the
dependent variable from the fitted regression equation for each observation. For the United States, for example, the predicted value of long-term inflation is –0.0003 + 0.3327(0.0653) = 0.0214 or 2.14 percent. The next-to-last column contains the regression residual, which is the difference between the actual value of the dependent variable, Yi, and the predicted value of the dependent
variable, . So for the United States, the residual is equal to 0.0293 – 0.0214 = 0.0079 or 0.79 percent. The last column contains the squared regression residual. The sum of the squared residuals is 0.000316. Applying the formula for the standard error of estimate, we obtain
TABLE 7 Computing the Standard Error of Estimate
Country Money Supply Growth Rate Xi
Inflation Rate Yi
Predicted Inflation Rate
Regression Residual
Squared Residual
Australia 0.1117 0.0462 0.0368 0.0094 0.000088 Japan 0.0408 0.0018 0.0132 –0.0114 0.000131 South Korea 0.1781 0.0531 0.0589 –0.0058 0.000034
Switzerland 0.0585 0.0199 0.0191 0.0008 0.000001 United
Kingdom 0.1293 0.0418 0.0427 –0.0009 0.000001
United States 0.0653 0.0293 0.0214 0.0079 0.000063
Sum 0.000316
Source: International Monetary Fund.
Thus the standard error of estimate is about 0.89 percent.
Later, we will combine this estimate with estimates of the uncertainty about the parameters in this regression to determine confidence intervals for predicting inflation rates from money supply growth. We will see that smaller standard errors result in more accurate predictions.
3.4. The Coefficient of Determination
Although the standard error of estimate gives some indication of how certain we can be about a particular prediction of Y using the regression equation, it still does not tell us how well the independent variable explains variation in the dependent variable. The coefficient of determination does exactly this: It measures the fraction of the total variation in the dependent variable that is explained by the independent variable.
We can compute the coefficient of determination in two ways. The simpler method, which can be used in a linear regression with one independent variable, is to square the correlation coefficient between the dependent and independent variables. For example, recall that the correlation coefficient between the long-term rate of money growth and the long-term rate of inflation between 1980 and 2012 for six industrialized countries was 0.9083. Thus the coefficient of determination in the regression shown in Figure 8 is (0.9083)2 = 0.8250. So in this regression, the long- term rate of money supply growth explains approximately 82.5 percent of the variation in the long-term rate of inflation across the countries between 1980 and 2012. (Relatedly, note that the square root of the coefficient of determination in a one-independent-variable linear regression, after attaching the sign of the estimated slope coefficient, gives the correlation coefficient between the dependent and independent variables.)
The problem with this method is that it cannot be used when we have more than one independent variable.27 Therefore, we need an alternative method of computing the coefficient of determination for multiple independent variables. We now present the logic behind that alternative.
If we did not know the regression relationship, our best guess for the value of any particular observation of the dependent variable would simply be , the mean of the dependent variable. One measure of accuracy in predicting Yi based on is the
sample variance of Yi, . An alternative to using to predict a particular observation Yi is using the regression relationship to make that prediction. In that case, our predicted value would be . If the regression relationship works well, the error in predicting Yi using should be much smaller than the error in
predicting Yi using . If we call the total variation of Y and
the unexplained variation from the regression, then we can measure the explained variation from the regression using the following equation:
(7)
The coefficient of determination is the fraction of the total variation that is explained by the regression. This gives us the relationship
(8)
Note that total variation equals explained variation plus unexplained variation, as shown in Equation 7. Most regression programs report the coefficient of determination as R2.28
EXAMPLE 13 Inflation Rate and Growth in the Money Supply
Using the data in Table 7, we can see that the unexplained variation from the regression, which is the sum of the squared residuals, equals 0.000316. Table 8 shows the computation of total variation in the dependent variable, the long- term rate of inflation.
Table 8 Computing Total Variation
Country Money SupplyGrowth Rate Xi Inflation Rate Yi
Deviation from Mean
Squared Deviation
Australia 0.1117 0.0462 0.0142 0.000201 Japan 0.0408 0.0018 –0.0302 0.000913 South Korea 0.1781 0.0531 0.0211 0.000445
Switzerland 0.0585 0.0199 –0.0121 0.000147 United
Kingdom 0.1293 0.0418 0.0098 0.000096
United States 0.0653 0.0293 –0.0027 0.000007
Average: 0.0320 Sum: 0.001809
Source: International Monetary Fund.
The average inflation rate for this period is 3.20 percent. The next-to-last column shows the amount each country’s long-term inflation rate deviates from that average; the last column shows the square of that deviation. The sum of those squared deviations is the total variation in Y for the sample (0.001809),
shown in Table 8.
Compute the coefficient of determination for the regression.
Solution: The coefficient of determination for the regression is
Note that this method gives the same result that we obtained earlier. We will use this method again in the reading on multiple regression; when we have more than one independent variable, this method is the only way to compute the coefficient of determination.
3.5. Hypothesis Testing In this section, we address testing hypotheses concerning the population values of the intercept or slope coefficient of a regression model. This topic is critical in practice. For example, we may want to check a stock’s valuation using the capital asset pricing model; we hypothesize that the stock has a market-average beta or level of systematic risk. Or we may want to test the hypothesis that economists’ forecasts of the inflation rate are unbiased (not overestimates or underestimates, on average). In each case, does the evidence support the hypothesis? Questions such as these can be addressed with hypothesis tests within a regression model. Such tests are often t-tests of the value of the intercept or slope coefficient(s). To understand the concepts involved in this test, it is useful to first review a simple, equivalent approach based on confidence intervals.
We can perform a hypothesis test using the confidence interval approach if we know three things: 1) the estimated parameter value, or , 2) the hypothesized value of the parameter, b0 or b1, and 3) a confidence interval around the estimated parameter. A confidence interval is an interval of values that we believe includes the true parameter value, b1, with a given degree of confidence. To compute a confidence interval, we must select the significance level for the test and know the standard error of the estimated coefficient.
Suppose we regress a stock’s returns on a stock market index’s returns and find that the slope coefficient ( ) is 1.5 with a standard error ( ) of 0.200. Assume we used 62 monthly observations in our regression analysis. The hypothesized value of the parameter (b1) is 1.0, the market average slope coefficient. The estimated and the population slope coefficients are often called beta, because the population coefficient is often represented by the Greek symbol beta (β) rather than the b1 we
use in this reading. Our null hypothesis is that b1 = 1.0 and is the estimate for b1. We will use a 95 percent confidence interval for our test, or we could say that the test has a significance level of 0.05.
Our confidence interval will span the range to or
(9)
where tc is the critical t value.29 The critical value for the test depends on the number of degrees of freedom for the t-distribution under the null hypothesis. The number of degrees of freedom equals the number of observations minus the number of parameters estimated. In a regression with one independent variable, there are two estimated parameters, the intercept term and the coefficient on the independent variable. For 62 observations and two parameters estimated in this example, we have 60 degrees of freedom (62 − 2). For 60 degrees of freedom, the table of critical values in the back of the book shows that the critical t-value at the 0.05 significance level is 2.00. Substituting the values from our example into Equation 9 gives us the interval
Under the null hypothesis, the probability that the confidence interval includes b1 is 95 percent. Because we are testing b1 = 1.0 and because our confidence interval does not include 1.0, we can reject the null hypothesis. Therefore, we can be 95 percent confident that the stock’s beta is different from 1.0.
In practice, the most common way to test a hypothesis using a regression model is with a t-test of significance. To test the hypothesis, we can compute the statistic
(10)
This test statistic has a t-distribution with n − 2 degrees of freedom because two parameters were estimated in the regression. We compare the absolute value of the t-statistic to tc. If the absolute value of t is greater than tc, then we can reject the null hypothesis. Substituting the values from the above example into this relationship gives the t-statistic associated with the probability that the stock’s beta equals 1.0 (b1 = 1.0).
Because t > tc, we reject the null hypothesis that b1 = 1.0.
The t-statistic in the example above is 2.50, and at the 0.05 significance level, tc = 2.00; thus we reject the null hypothesis because t > tc. This statement is equivalent to saying that we are 95 percent confident that the interval for the slope coefficient does not contain the value 1.0. If we were performing this test at the 0.01 level, however, tc would be 2.66 and we would not reject the hypothesis because t would not be greater than tc at this significance level. A 99 percent confidence interval for the slope coefficient does contain the value 1.0.
The choice of significance level is always a matter of judgment. When we use higher levels of confidence, the tc increases. This choice leads to wider confidence intervals and to a decreased likelihood of rejecting the null hypothesis. Analysts often choose the 0.05 level of significance, which indicates a 5 percent chance of rejecting the null hypothesis when, in fact, it is true (a Type I error). Of course, decreasing the level of significance from 0.05 to 0.01 decreases the probability of Type I error, but it increases the probability of Type II error—failing to reject the null hypothesis when, in fact, it is false.
Often, financial analysts do not simply report whether or not their tests reject a particular hypothesis about a regression parameter. Instead, they report the p-value or probability value for a particular hypothesis. The p-value is the smallest level of significance at which the null hypothesis can be rejected. It allows the reader to interpret the results rather than be told that a certain hypothesis has been rejected or accepted. In most regression software packages, the p-values printed for regression coefficients apply to a test of null hypothesis that the true parameter is equal to 0 against the alternative that the parameter is not equal to 0, given the estimated coefficient and the standard error for that coefficient. For example, if the p-value is 0.005, we can reject the hypothesis that the true parameter is equal to 0 at the 0.5 percent significance level (99.5 percent confidence).
The standard error of the estimated coefficient is an important input for a hypothesis test concerning the regression coefficient (and for a confidence interval for the estimated coefficient). Stronger regression results lead to smaller standard errors of an estimated parameter and result in tighter confidence intervals. If the standard error ( ) in the above example were 0.100 instead of 0.200, the confidence interval range would be half as large and the t-statistic twice as large. With a standard error this small, we would reject the null hypothesis even at the 0.01
significance level because we would have t = (1.5 − 1)/0.1 = 5.00 and tc = 2.66.
With this background, we can turn to hypothesis tests using actual regression results. The next three examples illustrate hypothesis tests in a variety of typical investment contexts.
EXAMPLE 14 Estimating Beta for Westport Innovations Stock
Westport Innovations Inc. (Westport) is a Canadian company that provides low- emission engine and fuel system technologies utilizing gaseous fuels. Its stock trades on the Toronto Stock Exchange (Ticker: WPT). You are an investor in Westport’s stock and want an estimate of its beta. As in the text example, you hypothesize that Westport has an average level of market risk and that its required return in excess of the risk-free rate is the same as the market’s required excess return. One regression that summarizes these statements is
(11)
where RF is the periodic risk-free rate of return (known at the beginning of the period), RM is the periodic return on the market, R is the periodic return to the stock of the company, and β measures the sensitivity of the required excess return to the excess return to market. Estimating this equation with linear regression provides an estimate of β, , which tells us the size of the required return premium for the security, given expectations about market returns.30
Suppose we want to test the null hypothesis, H0, that β = 1 for Westport stock to see whether Westport stock has the same required return premium as the market as a whole. We need data on returns to Westport stock, a risk-free interest rate, and the returns to the market index. For this example, we use data from January 2009 through December 2013 (n = 60). The return to Westport stock is R. The monthly return to 1 month Canadian Treasury bills is RF. The return to the S&P/TSX Composite Index is RM.31 This index is the primary broad measure of the Canadian equity market. We are estimating two parameters, so the number of degrees of freedom is n − 2 = 60 − 2 = 58. Table 9 shows the results from the regression (R − RF) = α + β (RM − RF) + ε.
Table 9 Estimating Beta for Westport Stock
Regression Statistics Multiple R 0.3429
R-squared 0.1176 Standard error of estimate 0.1488
Observations 60
Coefficients Standard Error t-Statistic Alpha 0.0267 0.0273 0.9793 Beta 1.0788 0.3880 2.7800
Sources: Bank of Canada and ca.finance.yahoo.com.
1. Test the null hypothesis, H0, that β for Westport equals 1 (β = 1) against the alternative hypothesis that β does not equal 1 (β ≠ 1) using the confidence interval approach.
2. Test the above hypothesis using a t-test.
3. How much of Westport stock’s excess return variation can be attributed to company-specific risk?
Solution to 1: The estimated from the regression is 1.0788. The estimated standard error for that coefficient in the regression, is 0.3880. The regression equation has 58 degrees of freedom (60 − 2), so the critical value for the test statistic is approximately tc = 2.00 at the 0.05 significance level. Therefore, the 95 percent confidence interval for the data for any particular hypothesized value of β is shown by the range
In this case, the hypothesized parameter value is β = 1, and the value 1 falls inside this confidence interval, so we cannot reject the hypothesis at the 0.05 significance level. This means that we cannot reject the hypothesis that Westport stock has the same systematic risk as the market as a whole.
Solution to 2: The t-statistic for the Westport beta hypothesized parameter can be computed using Equation 10:
This t-statistic is less than the critical t-value of 2.00. Therefore, neither
approach allows us to reject the null hypothesis. Note that the t-statistic associated with in the regression results in Table 9 is 2.7800. Given the significance level we are using, we cannot reject the null hypothesis that β = 1, but we can reject the hypothesis that β = 0.32
Solution to 3: The R2 in this regression is only 0.1176. This result suggests that only about 12 percent of the total variation in the excess return to Westport stock (the return to Westport above the risk-free rate) can be explained by excess return to the market portfolio. The remaining 88 percent of Westport stock’s excess return variation is the nonsystematic component, which can be attributed to company-specific risk.
In the next example, we show a regression hypothesis test with a one-sided alternative.
EXAMPLE 15 Explaining Company Value Based on Returns to Invested Capital
Some financial analysts have argued that one good way to measure a company’s ability to create wealth is to compare the company’s return on invested capital (ROIC) to its weighted-average cost of capital (WACC). If a company has an ROIC greater than its cost of capital, the company is creating wealth; if its ROIC is less than its cost of capital, it is destroying wealth.33
Enterprise value (EV) is a market-price-based measure of company value defined as the market value of equity and debt minus the value of cash and investments. Invested capital (IC) is an accounting measure of company value defined as the sum of the book values of equity and debt. Higher ratios of EV to IC should reflect greater success at wealth creation in general. Mauboussin (1996) argued that the spread between ROIC and WACC helps explain the ratio of EV to IC. Using data on companies in the food-processing industry, we can test the relationship between EV/IC and (ROIC–WACC) using the regression model given in Equation 12.
(12)
Table 10 Explaining Enterprise Value/Invested Capital by the ROIC–WACC Spread
Regression Statistics Multiple R 0.9469 R-squared 0.8966
Standard error of estimate 0.7422 Observations 9
Coefficients Standard Error t-Statistic Intercept 1.3478 0.3511 3.8391 Spread 30.0169 3.8519 7.7928
Source: Nelson, Moskow, Lee, and Valentine (2003).
where the subscript i is an index to identify the company. Our null hypothesis is H0: b1 ≤ 0, and we specify a significance level of 0.05. If we reject the null hypothesis, we have evidence of a statistically significant relationship between EV/IC and (ROIC–WACC). Equation 12 is estimated using data from nine food- processing companies.34 The results of this regression are displayed in Table 10 and Figure 11.
FIGURE 11 Fitted Regression Line Explaining Enterprise Value/Invested Capital Using ROIC–WACC Spread for the Food Industry
Source: Nelson et al. (2003).
We reject the null hypothesis based on the t-statistic of approximately 7.79 on estimated slope coefficient. There is a strong positive relationship between the
return spread (ROIC–WACC) and the ratio of EV to IC in our sample of companies. Figure 11 illustrates the strong positive relationship. The R2 of 0.8966 indicates that the return spread explains about 90 percent of the variation in the ratio of EV to IC among the food-processing companies in the sample in 2001. The coefficient on the return spread of 30.0169 implies that the predicted increase in EV/IC is 0.01(30.0169) = 0.3002 or about 30 percent for a 1- percentage-point increase in the return spread, for our sample of companies.
In the final example of this section, the null hypothesis for a t-test of the slope coefficient is that the value of the slope equals 1 in contrast to the null hypothesis that it equals 0 as in prior examples.
EXAMPLE 16 Testing whether Inflation Forecasts Are Unbiased
Example 11 introduced the concept of testing for bias in forecasts. That example showed that if a forecast is unbiased, its expected error is 0. We can examine whether a time-series of forecasts for a particular economic variable is unbiased by comparing the forecast at each date with the actual value of the economic variable announced after the forecast. If the forecasts are unbiased, then, by definition, the average realized forecast error should be close to 0. In that case, the value of b0 (the intercept) should be 0 and the value of b1 (the slope) should be 1, as discussed in Example 11.
Refer once again to Figure 9, which shows the mean forecast made by professional economic forecasters in the first quarter of a year for the percentage change in euro area HICP during that year and the actual percentage change from 1999 through 2013 (n = 14). To test whether the forecasts are unbiased, we must estimate the regression shown in Example 11. We report the results of this regression in Table 11. The equation to be estimated is
This regression estimates two parameters (the intercept and the slope); therefore, the regression has n − 2 = 14 − 2 = 12 degrees of freedom.
Table 11 Testing whether Forecasts of Euro Area HICP Are Unbiased (Dependent Variable: CPI Change Expressed in Percent)
Regression Statistics Multiple R 0.9006
R-squared 0.8111 Standard error of estimate 0.3165
Observations 14 Coefficients Standard Error t-Statistic
Intercept –0.7006 0.3723 –1.8820 Forecast (slope) 1.5538 0.2079 7.4722
Source: European Central Bank.
We can now test two null hypotheses about the parameters in this regression. Our first null hypothesis is that the intercept in this regression is 0 (H0: b0 = 0). The alternative hypothesis is that the intercept does not equal 0 (Ha: b0 ≠ 0). Our second null hypothesis is that the slope coefficient in this regression is 1 (H0: b1 = 1). The alternative hypothesis is that the slope coefficient does not equal 1 (Ha: b1 ≠ 1).
To test the hypotheses about b0 and b1, we must first decide on a critical value based on a particular significance level and then construct the confidence intervals for each parameter. If we choose the 0.05 significance level, with 12 degrees of freedom, the critical value, tc, is approximately 2.18. The estimated value of the parameter is −0.7006, and the estimated value of the standard error for is 0.3723. Let B0 stand for any particular hypothesized value. Therefore, under the null hypothesis that b0 = B0, a 95 percent confidence interval for b0 is
In this case, B0 is 0. The value of 0 falls within this confidence interval, so we cannot reject the first null hypothesis that b0 = 0. We will explain how to interpret this result shortly.
Our second null hypothesis is based on the same sample as our first null hypothesis. Therefore, the critical value for testing that hypothesis is the same as the critical value for testing the first hypothesis (tc = 2.18). The estimated value of the parameter is 1.5538, and the estimated value of the standard error for , , is 0.2079. Therefore, the 95 percent confidence interval for any particular hypothesized value of b1 can be constructed as follows:
In this case, our hypothesized value of b1 is 1. The value 1 falls outside this confidence interval, so we can reject the null hypothesis that b1 = 1 at the 0.05 significance level. Because we did reject one of the two null hypotheses (b0 = 0, b1 = 1) about the parameters in this model, we can reject the hypothesis that the forecasts of HICP change were unbiased.35
As an analyst, you often will need forecasts of economic growth to help you make recommendations about asset allocation, expected returns, and other investment decisions. The hypothesis tests just conducted suggest that you can reject the hypothesis that the HICP predictions in the Survey of Professional Forecasters are unbiased. If you need an unbiased forecast of future percentage change in HICP for your asset-allocation decision, you might not want to use these forecasts.
In view of the above concern, we further explored the inflation forecasts. Figure 9 suggests that the bottommost point in the plot of actual versus realized inflations is an outlier. This point corresponds to the year 2009 when macroeconomic volatility was exceptionally high due to the financial crisis. A study of forecasts in the European Central Bank Survey of Professional Forecasters by Genre, Kenny, Meyler, and Timmerman (2010) finds that the performance of inflation forecasts is lowered when the financial crisis period is included. We re-estimated the regression equation after excluding 2009. The new equation is Actual change = –0.2513 + 1.3209(Predicted change). Under the null hypothesis for the intercept that b0 = 0, a 95 percent confidence interval for b0 is –1.2116 to 0.7090. The value of 0 falls within this confidence interval, so we cannot reject the first null hypothesis that b0 = 0. Under the null hypothesis for the slope that b1 = 1, a 95 percent confidence interval for b1 is 0.7984 to 1.8434. The value of 1 falls within this confidence interval, so we cannot reject the second null hypothesis that b1 = 1. These hypothesis tests suggest that you cannot reject the hypothesis that the HICP predictions in the Survey of Professional Forecasters are unbiased.
3.6. Analysis of Variance in a Regression with One Independent Variable Analysis of variance (ANOVA) is a statistical procedure for dividing the total
variability of a variable into components that can be attributed to different sources.36 In regression analysis, we use ANOVA to determine the usefulness of the independent variable or variables in explaining variation in the dependent variable. An important statistical test conducted in analysis of variance is the F-test. The F- statistic tests whether all the slope coefficients in a linear regression are equal to 0. In a regression with one independent variable, this is a test of the null hypothesis H0: b1 = 0 against the alternative hypothesis Ha: b1 ≠ 0.
To correctly determine the test statistic for the null hypothesis that the slope coefficient equals 0, we need to know the following:
the total number of observations (n);
the total number of parameters to be estimated (in a one-independent-variable regression, this number is two: the intercept and the slope coefficient);
the sum of squared errors or residuals, , abbreviated SSE. This value is also known as the residual sum of squares; and
the regression sum of squares, , abbreviated RSS. This value is the amount of total variation in Y that is explained in the regression equation. Total variation (TSS) is the sum of SSE and RSS.
The F-test for determining whether the slope coefficient equals 0 is based on an F- statistic, constructed using these four values. The F-statistic measures how well the regression equation explains the variation in the dependent variable. The F-statistic is the ratio of the average regression sum of squares to the average sum of the squared errors. The average regression sum of squares is computed by dividing the regression sum of squares by the number of slope parameters estimated (in this case, one). The average sum of squared errors is computed by dividing the sum of squared errors by the number of observations, n, minus the total number of parameters estimated (in this case, two: the intercept and the slope). These two divisors are the degrees of freedom for an F-test. If there are n observations, the F- test for the null hypothesis that the slope coefficient is equal to 0 is here denoted F# slope parameters,n– # parameters = F1,n−2, and the test has 1 and n − 2 degrees of freedom.
Suppose, for example, that the independent variable in a regression model explains none of the variation in the dependent variable. Then the predicted value for the regression model, , is the average value of the dependent variable . In this case,
the regression sum of squares is 0. Therefore, the F-statistic is 0. If the independent variable explains little of the variation in the dependent variable, the value of the F-statistic will be very small.
The formula for the F-statistic in a regression with one independent variable is
(13)
If the regression model does a good job of explaining variation in the dependent variable, then this ratio should be high. The explained regression sum of squares per estimated parameter will be high relative to the unexplained variation for each degree of freedom. Critical values for this F-statistic are given in Appendix D at the end of this volume.
Even though the F-statistic is commonly computed by regression software packages, analysts typically do not use ANOVA and F-tests in regressions with just one independent variable. Why not? In such regressions, the F-statistic is the square of the t-statistic for the slope coefficient. Therefore, the F-test duplicates the t-test for the significance of the slope coefficient. This relation is not true for regressions with two or more slope coefficients. Nevertheless, the one-slope coefficient case gives a foundation for understanding the multiple-slope coefficient cases.
Often, mutual fund performance is evaluated based on whether the fund has positive alpha—significantly positive excess risk-adjusted returns.37 One commonly used method of risk adjustment is based on the capital asset pricing model. Consider the regression
(14)
where RF is the periodic risk-free rate of return (known at the beginning of the period), RM is the periodic return on the market, Ri is the periodic return to Mutual Fund i, and βi is the fund’s beta. A fund has zero risk-adjusted excess return if αi = 0. If αi = 0, then (Ri − RF) = βi(RM − RF) + εi and taking expectations, E(Ri) = RF + βi(RM − RF), implying that βi completely explains the fund’s mean excess returns. If, for example, αi > 0, the fund is earning higher returns than expected given its beta.
In summary, to test whether a fund has a positive alpha, we must test the null hypothesis that the fund has no risk-adjusted excess returns (H0: α = 0) against the alternative hypothesis of nonzero risk-adjusted returns (Ha: α ≠ 0).
EXAMPLE 17 Performance Evaluation: The Dreyfus
Appreciation Fund
Table 12 presents results evaluating the excess return to the Dreyfus Appreciation Fund from January 2009 through December 2013. Note that the estimated beta in this regression, , is 0.8660. The Dreyfus Appreciation Fund was estimated to be almost 0.9 times as risky as the market as a whole.
1. Test whether the fund had a significant excess return beyond the return
associated with the market risk of the fund.
2. Based on the t-test, discuss whether the beta of the fund is likely to be zero.
3. Use Equation 13 to compute the F-statistic. Based on the F-test, determine whether the beta of the fund is likely to be zero.
Table 12 Performance Evaluation of Dreyfus Appreciation Fund, January 2009 to December 2013
Regression Statistics Multiple R 0.9633 R-squared 0.9279
Standard error of estimate 0.0111 Observations 60
ANOVA Degrees ofFreedom (df) Sum of Squares
(SS) Mean Sum of Squares
(MSS) F
Regression 1 0.0925 0.0925 746.09 Residual 58 0.0072 0.0001 Total 59 0.0997
Coefficients Standard Error t-Statistic Alpha −0.0012 0.0015 −0.8050 Beta 0.8660 0.0317 27.3147
Sources: Center for Research in Security Prices, University of Chicago; S&P Dow Jones Indices; and the Federal Reserve.
Solution to 1: The estimated alpha ( ) in this regression is negative (−0.0012).
The absolute value of the coefficient is less than the size of the standard error for that coefficient (0.0015), so the t-statistic for the coefficient is only −0.8050. Therefore, we cannot reject the null hypothesis (α = 0) that the fund did not have a significant excess return beyond the return associated with the market risk of the fund. This result means that the returns to the fund were explained by the market risk of the fund and there was no additional statistical significance to the excess returns to the fund during this period.38
Solution to 2: Because the t-statistic for the slope coefficient in this regression is 27.3147, the p-value for that coefficient is less than 0.0001 and is approximately zero. Therefore, the probability that the true value of this coefficient is actually 0 is microscopic.
Solution to 3: The ANOVA portion of Table 12 provides the data we need to compute the F-statistic. In this case:
the total number of observations (n) is 60;
the total number of parameters to be estimated is 2 (intercept and slope);
the sum of squared errors or residuals, SSE, is 0.0072; and
the regression sum of squares, RSS, is 0.0925.
Therefore, the F-statistic to test whether the slope coefficient is equal to 0 is
(The slight difference from the F-statistic in Table 12 is due to rounding.) The ANOVA output would show that the p-value for this F-statistic is less than 0.0001 and is exactly the same as the p-value for the t-statistic for the slope coefficient. Therefore, the F-test tells us nothing more than we already knew from the t-test. Note also that the F-statistic (746.09) is the square of the t- statistic (27.3147).
3.7. Prediction Intervals Financial analysts often want to use regression results to make predictions about a dependent variable. For example, we might ask, “How fast will the sales of XYZ Corporation grow this year if real GDP grows by 4 percent?” But we are not merely interested in making these forecasts; we also want to know how certain we should be about the forecasts’ results. For example, if we predicted that sales for XYZ
Corporation would grow by 6 percent this year, our prediction would mean more if we were 95 percent confident that sales growth would fall in the interval from 5 percent to 7 percent, rather than only 25 percent confident that this outcome would occur. Therefore, we need to understand how to compute confidence intervals around regression forecasts.
We must take into account two sources of uncertainty when using the regression model Yi = b0 + b1Xi + εi, i = 1, …, n and the estimated parameters, and , to make a prediction. First, the error term itself contains uncertainty. The standard deviation of the error term, σε, can be estimated from the standard error of estimate for the regression equation. A second source of uncertainty in making predictions about Y, however, comes from uncertainty in the estimated parameters and .
If we knew the true values of the regression parameters, b0 and b1, then the variance of our prediction of Y, given any particular predicted (or assumed) value of X, would simply be s2, the squared standard error of estimate. The variance would be s2 because the prediction, , would come from the equation and .
Because we must estimate the regression parameters and however, our prediction of Y, , given any particular predicted value of X, is actually . The estimated variance of the prediction error, of Y, given X, is
(15)
This estimated variance depends on:
the squared standard error of estimate, s2;
the number of observations, n;
the value of the independent variable, X, used to predict the dependent variable;
the estimated mean, ; and
variance, of the independent variable.39
Once we have this estimate of the variance of the prediction error, determining a prediction interval around the prediction is very similar to estimating a confidence interval around an estimated parameter, as shown earlier in this reading. We need to take the following four steps to determine the prediction interval for the prediction: 1. Make the prediction.
2. Compute the variance of the prediction error using Equation 15.
3. Choose a significance level, α, for the forecast. For example, the 0.05 level, given the degrees of freedom in the regression, determines the critical value for the forecast interval, tc.
4. Compute the (1 − α) percent prediction interval for the prediction, namely .
EXAMPLE 18 Predicting the Ratio of Enterprise Value to Invested Capital
We continue with the example of explaining the ratio of enterprise value to invested capital among food-processing companies by the spread between the return to invested capital and the weighted-average cost of capital (ROIC– WACC). In Example 15, we estimated the regression given in Table 10.
TABLE 10 Explaining Enterprise Value/Invested Capital by the ROIC-WACC Spread (repeated)
Regression Statistics Multiple R 0.9469 R-squared 0.8966
Standard error of estimate 0.7422 Observations 9
Coefficients Standard Error t-Statistic Intercept 1.3478 0.3511 3.8391 Spread 30.0169 3.8519 7.7928
Source: Nelson, Moskow, Lee, and Valentine (2003).
You are interested in predicting the ratio of enterprise value to invested capital for a company if the return spread between ROIC and WACC is 10 percentage points. What is the 95 percent confidence interval for the ratio of enterprise value to invested capital for that company?
Using the data provided in Table 10, take the following steps: 1. Make the prediction: Expected EV/IC = 1.3478 + 30.0169(0.10) = 4.3495. This
regression suggests that if the return spread between ROIC and WACC (Xi) is
10 percent, the EV/IC ratio will be 4.3495.
2. Compute the variance of the prediction error. To compute the variance of the forecast error, we must know:
the standard error of the estimate of the equation, s = 0.7422 (as shown in Table 10);
the mean return spread, = 0.0647 (this computation is not shown in the table); and
the variance of the mean return spread in the sample, = 0.004641 (this computation is not shown in the table).
Using these data, you can compute the variance of the forecast error ( ) for predicting EV/IC for a company with a 10 percent spread between ROIC and WACC.
In this example, the variance of the forecast error is 0.630556, and the standard deviation of the forecast error is sf = (0.630556)1/2 = 0.7941.
3. Determine the critical value of the t-statistic. Given a 95 percent confidence interval and 9 − 2 = 7 degrees of freedom, the critical value of the t-statistic, tc, is 2.365 using the tables in the back of this volume.
4. Compute the prediction interval. The 95 percent confidence interval for EV/IC extends from 4.3495 − 2.365(0.7941) to 4.3495 + 2.365(0.7941), or 2.4715 to 6.2275.
In summary, if the spread between the ROIC and the WACC is 10 percent, the 95 percent prediction interval for EV/IC will extend from 2.4715 to 6.2275. The small sample size is reflected in the relatively large prediction interval.
3.8. Limitations of Regression Analysis Although this reading has shown many of the uses of regression models for financial analysis, regression models do have limitations. First, regression relations can change over time, just as correlations can. This fact is known as the issue of parameter instability, and its existence should not be surprising as the economic, tax, regulatory, political, and institutional contexts in which financial markets
operate change. Whether considering cross-sectional or time-series regression, the analyst will probably face this issue. As one example, cross-sectional regression relationships between stock characteristics may differ between growth-led and value-led markets. As a second example, the time-series regression estimating the beta often yields significantly different estimated betas depending on the time period selected. In both cross-sectional and time-series contexts, the most common problem is sampling from more than one population, with the challenge of identifying when doing so is an issue.
A second limitation to the use of regression results specific to investment contexts is that public knowledge of regression relationships may negate their future usefulness. Suppose, for example, an analyst discovers that stocks with a certain characteristic have had historically very high returns. If other analysts discover and act upon this relationship, then the prices of stocks with that characteristic will be bid up. The knowledge of the relationship may result in the relation no longer holding in the future.
Finally, if the regression assumptions listed in Section 3.2 are violated, hypothesis tests and predictions based on linear regression will not be valid. Although there are tests for violations of regression assumptions, often uncertainty exists as to whether an assumption has been violated. This limitation will be discussed in detail in the reading on multiple regression.
4. Summary
A scatter plot shows graphically the relationship between two variables. If the points on the scatter plot cluster together in a straight line, the two variables have a strong linear relation.
The sample correlation coefficient for two variables X and Y is .
If two variables have a very strong linear relation, then the absolute value of their correlation will be close to 1. If two variables have a weak linear relation, then the absolute value of their correlation will be close to 0.
The squared value of the correlation coefficient for two variables quantifies the percentage of the variance of one variable that is explained by the other. If the correlation coefficient is positive, the two variables are directly related; if the correlation coefficient is negative, the two variables are inversely related.
If we have n observations for two variables, we can test whether the population correlation between the two variables is equal to 0 by using a t-test. This test statistic has a t-distribution with n − 2 degrees of freedom if the null hypothesis of 0 correlation is true.
Even one outlier can greatly affect the correlation between two variables. Analysts should examine a scatter plot for the variables to determine whether outliers might affect a particular correlation.
Correlations can be spurious in the sense of misleadingly pointing toward associations between variables.
The dependent variable in a linear regression is the variable that the regression model tries to explain. The independent variables are the variables that a regression model uses to explain the dependent variable.
If there is one independent variable in a linear regression and there are n observations on the dependent and independent variables, the regression model is Yi = b0 + b1Xi + εi, i = 1, …, n, where Yi is the dependent variable, Xi is the independent variable, and εi is the error term. In this model, the coefficient b0 is the intercept. The intercept is the predicted value of the dependent variable when the independent variable has a value of zero. In this model, the coefficient b1 is the slope of the regression line. If the value of the independent variable increases by one unit, then the model predicts that the value of the dependent
variable will increase by b1 units.
The assumptions of the classic normal linear regression model are the following:
A linear relation exists between the dependent variable and the independent variable.
The independent variable is not random.
The expected value of the error term is 0.
The variance of the error term is the same for all observations (homoskedasticity).
The error term is uncorrelated across observations.
The error term is normally distributed.
The estimated parameters in a linear regression model minimize the sum of the squared regression residuals.
The standard error of estimate measures how well the regression model fits the data. If the SEE is small, the model fits well.
The coefficient of determination measures the fraction of the total variation in the dependent variable that is explained by the independent variable. In a linear regression with one independent variable, the simplest way to compute the coefficient of determination is to square the correlation of the dependent and independent variables.
To calculate a confidence interval for an estimated regression coefficient, we must know the standard error of the estimated coefficient and the critical value for the t-distribution at the chosen level of significance, tc.
To test whether the population value of a regression coefficient, b1, is equal to a particular hypothesized value, B1, we must know the estimated coefficient, , the standard error of the estimated coefficient, , and the critical value for the t-distribution at the chosen level of significance, tc. The test statistic for this hypothesis is . If the absolute value of this statistic is greater than tc, then we reject the null hypothesis that b1 = B1.
In the regression model Yi = b0 + b1Xi + εi, if we know the estimated parameters, and , for any value of the independent variable, X, then the predicted value of the dependent variable Y is .
The prediction interval for a regression equation for a particular predicted value of the dependent variable is where sf is the square root of the estimated variance of the prediction error and tc is the critical level for the t- statistic at the chosen significance level. This computation specifies a (1 − α) percent confidence interval. For example, if α = 0.05, then this computation yields a 95 percent confidence interval.
Problems Practice Problems and Solutions: 1–14 taken from Quantitative Methods for Investment Analysis, Second Edition, by Richard A. DeFusco, CFA, Dennis W. McLeavey, CFA, Jerald E. Pinto, PhD, CFA, and David E. Runkle, PhD, CFA. Copyright © 2004 by CFA Institute. All other problems and solutions copyright © CFA Institute. 1. Variable X takes on the values shown in the following table for five
observations. The table also shows the values for five other variables, Y1 through Y5. Which of the variables Y1 through Y5 have a zero correlation with variable X? X Y1 Y2 Y3 Y4 Y5
1 7 2 4 4 1 2 7 4 2 1 2 3 7 2 0 0 3 4 7 4 2 1 4 5 7 2 4 4 5
2. Use the data sample below to answer the following questions.
1. Calculate the sample mean, variance, and standard deviation for X.
2. Calculate the sample mean, variance, and standard deviation for Y.
3. Calculate the sample covariance between X and Y.
4. Calculate the sample correlation between X and Y.
3. Statistics for three variables are given below. X is the monthly return for a large-stock index, Y is the monthly return for a small-stock index, and Z is the monthly return for a corporate bond index. There are 60 observations.
1. Calculate the sample variance and standard deviation for X, Y, and Z.
2. Calculate the sample covariance between X and Y, X and Z, and Y and Z.
3. Calculate the sample correlation between X and Y, X and Z, and Y and Z.
4. Home sales and interest rates should be negatively related. The following table gives the number of annual unit sales for Packard Homes and mortgage rates for four recent years. Calculate the sample correlation between sales and mortgage rates. Year Unit Sales Interest Rate (%) 2000 50 8.0 2001 70 7.0 2002 80 6.0 2003 60 7.0
5. The following table shows the sample correlations between the monthly returns for four different mutual funds and the S&P 500. The correlations are based on 36 monthly observations. The funds are as follows: Fund 1 Large-cap fund Fund 2 Mid-cap fund Fund 3 Large-cap value fund Fund 4 Emerging markets fund S&P 500 US domestic stock index
Fund 1 Fund 2 Fund 3 Fund 4 S&P 500 Fund 1 1 Fund 2 0.9231 1 Fund 3 0.4771 0.4156 1 Fund 4 0.7111 0.7238 0.3102 1 S&P 500 0.8277 0.8223 0.5791 0.7515 1
Test the null hypothesis that each of these correlations, individually, is equal to zero against the alternative hypothesis that it is not equal to zero. Use a 5 percent significance level.
6. Bouvier Co. is a Canadian company that sells forestry products to several Pacific Rim customers. Bouvier ’s sales are very sensitive to exchange rates. The following table shows recent annual sales (in millions of Canadian dollars) and the average exchange rate for the year (expressed as the units of foreign currency needed to buy one Canadian dollar). Year i Exchange Rate Xi Sales Yi 1 0.40 20 2 0.36 25 3 0.42 16 4 0.31 30 5 0.33 35 6 0.34 30
1. Calculate the sample mean and standard deviation for X (the exchange
rate) and Y (sales).
2. Calculate the sample covariance between the exchange rate and sales.
3. Calculate the sample correlation between the exchange rate and sales.
4. Calculate the intercept and coefficient for an estimated linear regression with the exchange rate as the independent variable and sales as the dependent variable.
7. Julie Moon is an energy analyst examining electricity, oil, and natural gas consumption in different regions over different seasons. She ran a regression explaining the variation in energy consumption as a function of temperature. The total variation of the dependent variable was 140.58, the explained variation was 60.16, and the unexplained variation was 80.42. She had 60 monthly observations.
1. Compute the coefficient of determination.
2. What was the sample correlation between energy consumption and temperature?
3. Compute the standard error of the estimate of Moon’s regression model.
4. Compute the sample standard deviation of monthly energy consumption.
8. You are examining the results of a regression estimation that attempts to explain the unit sales growth of a business you are researching. The analysis of variance output for the regression is given in the table below. The regression was based on five observations (n = 5). ANOVA
df SS MSS F Significance F Regression 1 88.0 88.0 36.667 0.00904 Residual 3 7.2 2.4 Total 4 95.2
1. How many independent variables are in the regression to which the
ANOVA refers?
2. Define Total SS.
3. Calculate the sample variance of the dependent variable using information in the above table.
4. Define Regression SS and explain how its value of 88 is obtained in terms of other quantities reported in the above table.
5. What hypothesis does the F-statistic test?
6. Explain how the value of the F-statistic of 36.667 is obtained in terms of other quantities reported in the above table.
7. Is the F-test significant at the 5 percent significance level?
9. The first table below contains the regression results for a regression with monthly returns on a large-cap mutual fund as the dependent variable and monthly returns on a market index as the independent variable. The analysis is performed using only 12 monthly returns (in percent). The second table provides summary statistics for the dependent and independent variables.
1. What is the predicted return on the large-cap mutual fund for a market index return of 8.00 percent?
2. Find a 95 percent prediction interval for the expected mutual fund return. Regression Statistics
Multiple R 0.776 R-squared 0.602
Standard error 4.243 Observations 12
Coefficients Standard Error t-Statistic p-Value Intercept –0.287 1.314 –0.219 0.831
Slope coefficient 0.802 0.206 3.890 0.003
Statistic Market Index Return Large-CapFund Return Mean 2.30% 1.56%
Standard deviation 6.21% 6.41% Variance 38.51 41.13 Count 12 12
10. Industry automobile sales should be related to consumer sentiment. The following table provides a regression analysis in which sales of automobiles and light trucks (in millions of vehicles) are estimated as a function of a consumer sentiment index. Regression Statistics
Multiple R 0.80113 R-squared 0.64181
Standard error 0.81325 Observations 120
Coefficients Standard Error t-Statistic p-Value Intercept 6.071 0.58432 10.389 0
Slope coefficient 0.09251 0.00636 14.541 0
For the independent variable and dependent variable, the means, standard deviations, and variances are as follows:
Sentiment IndexX Automobile Sales(Millions of Units)Y Mean 91.0983 14.4981
Standard deviation 11.7178 1.35312 Variance 137.3068 1.83094
1. Find the expected sales and a 95 percent prediction interval for sales if the
sentiment index has a value of 90.
2. Find the expected sales and a 95 percent prediction interval for sales if the sentiment index has a value of 100.
11. Use the following information to create a regression model:
1. Calculate the sample mean, variance, and standard deviation for X and for Y.
2. Calculate the sample covariance and the correlation between X and Y.
3. Calculate and for a regression of the form .
For the remaining three parts of this question, assume that the calculations shown above already incorporate the correct values for and .
4. Find the total variation, explained variation, and unexplained variation.
5. Find the coefficient of determination.
6. Find the standard error of the estimate.
12. The bid–ask spread for stocks depends on the market liquidity for stocks. One measure of liquidity is a stock’s trading volume. Below are the results of a regression analysis using the bid–ask spread at the end of 2002 for a sample of 1,819 NASDAQ-listed stocks as the dependent variable and the natural log of trading volume during December 2002 as the independent variable. Several items in the regression output have been intentionally omitted. Use the reported information to fill in the missing values. Regression Statistics
Multiple R X2 R-squared X1
Standard error X3 Observations 1819
ANOVA df SS MSS F Significance F Regression X5 14.246 X7 X9 0 Residual X6 45.893 X8 Total X4 60.139
Coefficients StandardError t-
Statistic p-
Value Lower 95% Upper95%
Intercept 0.55851 0.018707 29.85540 0 0.52182 0.59520 Slope
coefficient −0.04375 0.001842 X10 0 X11 X12
13. An economist collected the monthly returns for KDL’s portfolio and a diversified stock index. The data collected are shown below: Month Portfolio Return (%) Index Return (%)
1 1.11 −0.59 2 72.10 64.90 3 5.12 4.81 4 1.01 1.68 5 −1.72 −4.97 6 4.06 −2.06
The economist calculated the correlation between the two returns and found it to be 0.996. The regression results with the KDL return as the dependent variable and the index return as the independent variable are given as follows:
Regression Statistics Multiple R 0.996 R-squared 0.992
Standard error 2.861 Observations 6
ANOVA df SS MSS F Significance F Regression 1 4101.62 4101.62 500.79 0 Residual 4 32.76 8.19
Total 5 4134.38
Coefficients Standard Error t-Statistic p-Value Intercept 2.252 1.274 1.768 0.1518 Slope 1.069 0.0477 22.379 0
When reviewing the results, Andrea Fusilier suspected that they were unreliable. She found that the returns for Month 2 should have been 7.21 percent and 6.49 percent, instead of the large values shown in the first table. Correcting these values resulted in a revised correlation of 0.824 and the revised regression results shown as follows:
Regression Statistics Multiple R 0.824 R-squared 0.678
Standard error 2.062 Observations 6
ANOVA df SS MSS F Significance F Regression 1 35.89 35.89 8.44 0.044 Residual 4 17.01 4.25 Total 5 52.91
Coefficients Standard Error t-Statistic p-Value Intercept 2.242 0.863 2.597 0.060 Slope 0.623 0.214 2.905 0.044
Explain how the bad data affected the results.
14. Diet Partners charges its clients a small management fee plus a percentage of gains whenever portfolio returns are positive. Cleo Smith believes that strong incentives for portfolio managers produce superior returns for clients. In order to demonstrate this, Smith runs a regression with the Diet Partners’ portfolio return (in percent) as the dependent variable and its management fee (in percent) as the independent variable. The estimated regression for a 60- month period is
The calculated t-values are given in parentheses below the intercept and slope coefficients. The coefficient of determination for the regression model is 0.794.
1. What is the predicted RETURN if FEE is 0 percent? If FEE is 1 percent?
2. Using a two-tailed test, is the relationship between RETURN and FEE significant at the 5 percent level?
3. Would Smith be justified in concluding that high fees are good for clients?
The following information relates to Questions 15–20 Kenneth McCoin, CFA, is a fairly tough interviewer. Last year, he handed each job applicant a sheet of paper with the information in the following table, and he then asked several questions about regression analysis. Some of McCoin’s questions, along with a sample of the answers he received to each, are given below. McCoin told the applicants that the independent variable is the ratio of net income to sales for restaurants with a market cap of more than $100 million and the dependent variable is the ratio of cash flow from operations to sales for those restaurants. Which of the choices provided is the best answer to each of McCoin’s questions?
Regression Statistics Multiple R 0.8623 R-squared 0.7436
Standard error 0.0213 Observations 24
ANOVA df SS MSS F Significance F Regression 1 0.029 0.029000 63.81 0 Residual 22 0.010 0.000455 Total 23 0.040
Coefficients Standard Error t-Statistic p-Value Intercept 0.077 0.007 11.328 0 Slope 0.826 0.103 7.988 0
15. What is the value of the coefficient of determination?
1. 0.8261.
2. 0.7436.
3. 0.8623.
16. Suppose that you deleted several of the observations that had small residual values. If you re-estimated the regression equation using this reduced sample, what would likely happen to the standard error of the estimate and the R- squared?
Standard Error of the Estimate R-Squared A. Decrease Decrease B. Decrease Increase C. Increase Decrease
17. What is the correlation between X and Y?
1. −0.7436.
2. 0.7436.
3. 0.8623.
18. Where did the F-value in the ANOVA table come from?
1. You look up the F-value in a table. The F depends on the numerator and denominator degrees of freedom.
2. Divide the “Mean Square” for the regression by the “Mean Square” of the residuals.
3. The F-value is equal to the reciprocal of the t-value for the slope coefficient.
19. If the ratio of net income to sales for a restaurant is 5 percent, what is the predicted ratio of cash flow from operations to sales?
1. 0.007 + 0.103(5.0) = 0.524.
2. 0.077 − 0.826(5.0) = −4.054.
3. 0.077 + 0.826(5.0) = 4.207.
20. Is the relationship between the ratio of cash flow to operations and the ratio of net income to sales significant at the 5 percent level?
1. No, because the R-squared is greater than 0.05.
2. No, because the p-values of the intercept and slope are less than 0.05.
3. Yes, because the p-values for F and t for the slope coefficient are less than 0.05.
The following information relates to Questions 21–26 Howard Golub, CFA, is preparing to write a research report on Stellar Energy Corp. common stock. One of the world’s largest companies, Stellar is in the business of refining and marketing oil. As part of his analysis, Golub wants to evaluate the sensitivity of the stock’s returns to various economic factors. For example, a client recently asked Golub whether the price of Stellar Energy Corporation stock has tended to rise following increases in retail energy prices. Golub believes the association between the two variables to be negative, but he does not know the strength of the association.
Golub directs his assistant, Jill Batten, to study the relationships between Stellar monthly common stock returns versus the previous month’s percent change in the US Consumer Price Index for Energy (CPIENG), and Stellar monthly common stock returns versus the previous month’s percent change in the US Producer Price Index for Crude Energy Materials (PPICEM). Golub wants Batten to run both a correlation and a linear regression analysis. In response, Batten compiles the summary statistics shown in Exhibit 1 for the 248 months between January 1980 and August 2000. All of the data are in decimal form, where 0.01 indicates a 1 percent return. Batten also runs a regression analysis using Stellar monthly returns as the dependent variable and the monthly change in CPIENG as the independent variable. Exhibit 2 displays the results of this regression model.
EXHIBIT 1 Descriptive Statistics
Monthly Return Stellar Common Stock
Lagged Monthly Change
CPIENG PPICEM Mean 0.0123 0.0023 0.0042
Standard Deviation 0.0717 0.0160 0.0534 Covariance, Stellar vs.
CPIENG −0.00017
Covariance, Stellar vs. PPICEM −0.00048
Covariance, CPIENG vs. PPICEM 0.00044
Correlation, Stellar vs. CPIENG −0.1452
EXHIBIT 2 Regression Analysis with CPIENG
Regression Statistics Multiple R 0.1452 R-squared 0.0211
Standard error of the estimate 0.0710 Observations 248
Coefficients Standard Error t-Statistic Intercept 0.0138 0.0046 3.0275
Slope coefficient −0.6486 0.2818 −2.3014
21. Batten wants to determine whether the sample correlation between the Stellar and CPIENG variables (−0.1452) is statistically significant. The critical value for the test statistic at the 0.05 level of significance is approximately 1.96. Batten should conclude that the statistical relationship between Stellar and CPIENG is:
1. significant, because the calculated test statistic has a lower absolute value than the critical value for the test statistic.
2. significant, because the calculated test statistic has a higher absolute value than the critical value for the test statistic.
3. not significant, because the calculated test statistic has a higher absolute value than the critical value for the test statistic.
22. Did Batten’s regression analyze cross-sectional or time-series data, and what was the expected value of the error term from that regression?
Data Type Expected Value of Error Term A. Time-series 0 B. Time-series εi
C. Cross-sectional 0
23. Based on the regression, which used data in decimal form, if the CPIENG decreases by 1.0 percent, what is the expected return on Stellar common stock during the next period?
1. 0.0073 (0.73 percent).
2. 0.0138 (1.38 percent).
3. 0.0203 (2.03 percent).
24. Based on Batten’s regression model, the coefficient of determination indicates that:
1. Stellar ’s returns explain 2.11 percent of the variability in CPIENG.
2. Stellar ’s returns explain 14.52 percent of the variability in CPIENG.
3. Changes in CPIENG explain 2.11 percent of the variability in Stellar ’s returns.
25. For Batten’s regression model, the standard error of the estimate shows that the standard deviation of:
1. the residuals from the regression is 0.0710.
2. values estimated from the regression is 0.0710.
3. Stellar ’s observed common stock returns is 0.0710.
26. For the analysis run by Batten, which of the following is an incorrect conclusion from the regression output?
1. The estimated intercept coefficient from Batten’s regression is statistically significant at the 0.05 level.
2. In the month after the CPIENG declines, Stellar ’s common stock is expected to exhibit a positive return.
3. Viewed in combination, the slope and intercept coefficients from Batten’s regression are not statistically significant at the 0.05 level.
Notes 1 Examples in this reading were updated in 2014 by Professor Sanjiv Sabherwal of
the University of Texas, Arlington.
2 Later, we show that variables with a correlation of 0 can have a strong nonlinear relation.
3 The use of n − 1 in the denominator is a technical point; it ensures that the sample covariance is an unbiased estimate of population covariance.
4 We have not used full precision in the table’s calculations. We used the average value of the money supply growth rate of 0.5839/6 = 0.0973, rounded to four decimal places, in the cross-product and squared deviation calculations, and similarly, we used the mean inflation rate as rounded to 0.0320 in those calculations. We computed standard deviation as the square root of variance rounded to six decimal places, as shown in the table. Had we used full precision in all calculations, some of the table’s entries would be slightly different but would not materially affect our conclusions.
5 The perfect association is the quadratic relationship B = (A − 4)2.
6 The euro area survey is conducted by the European Central Bank (ECB). A survey of professional forecasters is also conducted by the Federal Reserve Bank of Philadelphia for the United States.
7 In this scatter plot, the actual inflation rate is from the Statistical Data Warehouse of the European Central Bank.
8 See, for example, Sharpe (1992), Buetow, Johnson, and Runkle (2000), and Chan, Dimmock, and Lakonishok (2009).
9 See, for example, Campbell, Medeiros, and Viceira (2009).
10 The correlation coefficient for the 1990s was not significantly different from that for the 1980s at the 0.10 significance level. A test for this type of hypothesis on the correlation coefficient can be conducted using Fisher ’s z-transformation. See Daniel and Terrell (1995) for information on this method.
11 For more on these three measures and their use in equity valuation, see Pinto, Henry, Robinson, and Stowe (2010). The statements in the footnoted paragraph
explain the relationships among these measures according to US GAAP. Pinto et al also discuss the relationships among these measures according to international accounting standards.
12 The results in this table are based on data for all women’s clothing stores (US Occupational Health and Safety Administration Standard Industrial Classification 5621) with a market capitalization of more than $250 million at the end of 2001. The market-cap criterion was used to eliminate microcap firms, whose performance-measure correlations may be different from those of higher-valued firms.
13 Actually, we must assume that the variables come from a bivariate normal distribution. If two variables, X and Y, come from a bivariate normal distribution, then for each value of X the distribution of Y is normal. See, for example, Ross (2012) or Greene (2011).
14 It is worth repeating that the smaller the sample, the greater the evidence in terms of the magnitude of the sample correlation needed to reject the null hypothesis of zero correlation. With a sample size of 6, the absolute value of the sample correlation would need to be greater than 0.81 (carrying two decimal places) for us to reject the null hypothesis. Viewed another way, the value of 0.4045 in the text would be significant if the sample size were 24, because 0.4045(24 − 2)1/2/(1 − 0.40452)1/2 = 2.075, which is greater than the critical t-value of 2.074 at the 0.05 significance level with 22 degrees of freedom.
15 A mix of time-series and cross-sectional data, also known as panel data, is now frequently used in financial analysis. The analysis of panel data is an advanced topic that Greene (2011) discusses in detail.
16 In this reading, we primarily use the notation i= 1, 2, …, n even for time series to prevent confusion that would be caused by switching back and forth between different notations.
17 Hats over the symbols for coefficients indicate estimated values.
18 For a discussion of the precise statistical sense in which the estimates of b0 and b1 are optimal, see Greene (2011).
19 These data appear in Table 2.
20 We entered the monthly returns as decimals. Also, we used rounded numbers in the formulas discussed later to estimate the regression equation.
21 Although we assume that the independent variable in the regression model is not random, that assumption is clearly often not true. For example, it is unrealistic to assume that the monthly returns to the S&P 500 are not random. If the independent variable is random, then is the regression model incorrect? Fortunately, no. Econometricians have shown that even if the independent variable is random, we can still rely on the results of regression models given the crucial assumption that the error term is uncorrelated with the independent variable. The mathematics underlying this reliability demonstration, however, are quite difficult. See, for example, Greene (2011) or Goldberger (1998).
22 Var(εi) = E[εi− E(εi)]2 = E(εi− 0)2 = E(εi)2. Cov(εi, εj) = E{[εi− E(εi)][εj− E(εj)]} = E[(εi− 0) (εj− 0)] = E(εiεj) = 0.
23 If the regression errors are not normally distributed, we can still use regression analysis. Econometricians who dispense with the normality assumption use chi- square tests of hypotheses rather than F-tests. This difference usually does not affect whether the test will result in a particular null hypothesis being rejected.
24 For more information on nonlinearity in the parameters, see Gujarati and Porter (2008).
25 For large sample sizes, we may be able to drop the assumption of normality by appeal to the central limit theorem; see Greene (2011). Asymptotic theory shows that, in many cases, the test statistics produced by standard regression programs are valid even if the error term is not normally distributed. Non-normality of some financial time series can be quite severe. With severe non-normality, even with a relatively large number of observations, invoking asymptotic theory to justify using test statistics from linear regression models may be inappropriate.
26 See, for example, Keane and Rumble (1990).
27 We will discuss such models in the reading on multiple regression.
28 As we illustrate in the tables of regression output later in this reading, regression programs also report multiple R, which is the correlation between the actual values and the forecast values of Y. The coefficient of determination is the square of multiple R.
29 We use the t-distribution for this test because we are using a sample estimate of the standard error, sb, rather than its true (population) value.
30 Beta (β) is typically estimated using 60 months of historical data, but the data-
sample length sometimes varies. Although monthly data is typically used, some financial analysts estimate β using daily data. For more information on methods of estimating β, see Reilly and Brown (2012). The expected excess return for Westport stock above the risk-free rate (R − RF) is β(RM − RF), given a particular excess return to the market above the risk-free rate (RM − RF). This result holds because we regress (R − RF) against (RM − RF). For example, if a stock’s beta is 1.5, its expected excess return is 1.5 times that of the market portfolio.
31 Data on Westport stock returns and S&P/TSX Composite Index returns came from ca.finance.yahoo.com. Data on Canadian T-bill returns came from the Bank of Canada.
32 The t-statistics for a coefficient automatically reported by statistical software programs assume that the null hypothesis states that the coefficient is equal to 0. If you have a different null hypothesis, as we do in this example (β = 1), then you must either construct the correct test statistic yourself or instruct the program to compute it.
33 See, for example, Stewart (1991) and Mauboussin (1996).
34 Our data come from Nelson, Moskow, Lee, and Valentine (2003) and relate to 2001. Many sell-side analysts use this type of regression. It is one of the most frequently used cross-sectional regressions in published analyst reports.
35 Jointly testing the hypothesis b0 = 0 and b1 = 1 would require us to take into account the covariance of and . For information on testing joint hypotheses of this type, see Greene (2011).
36 In this reading, we focus on regression applications of ANOVA, the most common context in which financial analysts will encounter this tool. In this context, ANOVA is used to test whether all the regression slope coefficients are equal to 0. Analysts also use ANOVA to test a hypothesis that the means of two or more populations are equal. See Daniel and Terrell (1995) for details.
37 Note that the Greek letter alpha, α, is traditionally used to represent the intercept in Equation 14 and should not be confused with another traditional usage of α to represent a significance level.
38 This example introduces a well-known investment use of regression involving the capital asset pricing model. Researchers, however, recognize qualifications to the interpretation of alpha from a linear regression. The systematic risk of a managed portfolio is controlled by the portfolio manager. If, as a consequence,
portfolio beta is correlated with the return on the market (as could result from market timing), inferences on alpha based on least-squares beta, as here, can be mistaken. This advanced subject is discussed in Dybvig and Ross (1985a) and (1985b).
39 For a derivation of this equation, see Pindyck and Rubinfeld (1998).
CHAPTER 9 Multiple Regression and Issues in Regression Analysis Richard A. DeFusco, CFA
Dennis W. McLeavey, CFA
Jerald E. Pinto, PhD, CFA
David E. Runkle, PhD, CFA
LEARNING OUTCOMES
After completing this chapter, you will be able to do the following:
formulate a multiple regression equation to describe the relation between a dependent variable and several independent variables and determine the statistical significance of each independent variable;
interpret estimated regression coefficients and their p-values;
formulate a null and an alternative hypothesis about the population value of a regression coefficient, calculate the value of the test statistic, and determine whether to reject the null hypothesis at a given level of significance;
interpret the results of hypothesis tests of regression coefficients;
calculate and interpret 1) a confidence interval for the population value of a regression coefficient and 2) a predicted value for the dependent variable, given an estimated regression model and assumed values for the independent variables;
explain the assumptions of a multiple regression model;
calculate and interpret the F-statistic, and describe how it is used in regression analysis;
distinguish between and interpret the R2 and adjusted R2 in multiple regression;
evaluate how well a regression model explains the dependent variable by analyzing the output of the regression equation and an ANOVA table;
formulate a multiple regression equation by using dummy variables to represent qualitative factors and interpret the coefficients and regression results;
explain the types of heteroskedasticity and how heteroskedasticity and serial correlation affect statistical inference;
describe multicollinearity and explain its causes and effects in regression analysis;
describe how model misspecification affects the results of a regression analysis and describe how to avoid common forms of misspecification;
describe models with qualitative dependent variables;
evaluate and interpret a multiple regression model and its results.
1. Introduction As financial analysts, we often need to use more-sophisticated statistical methods than correlation analysis or regression involving a single independent variable. For example, a trading desk interested in the costs of trading NASDAQ stocks might want information on the determinants of the bid–ask spread on the NASDAQ. A mutual fund analyst might want to know whether returns to a technology mutual fund behaved more like the returns to a growth stock index or like the returns to a value stock index. An investor might be interested in the factors that determine whether analysts cover a stock. We can answer these questions using linear regression with more than one independent variable—multiple linear regression.
In Sections 2 and 3, we introduce and illustrate the basic concepts and models of multiple regression analysis. These models rest on assumptions that are sometimes violated in practice. In Section 4, we discuss three commonly occurring violations of regression assumptions. We address practical concerns such as how to diagnose an assumption violation and what remedial steps to take when a model assumption has been violated. Section 5 outlines some guidelines for building good regression models and discusses ways that analysts sometimes go wrong in this endeavor. In Section 6, we discuss a class of models whose dependent variable is qualitative in nature. These models are useful when the concern is over the occurrence of some event, such as whether a stock has analyst coverage or not.
2. Multiple Linear Regression As investment analysts, we often hypothesize that more than one variable explains the behavior of a variable in which we are interested. The variable we seek to explain is called the dependent variable. The variables that we believe explain the dependent variable are called the independent variables.1 A tool that permits us to examine the relationship (if any) between the two types of variables is multiple linear regression. Multiple linear regression allows us to determine the effect of more than one independent variable on a particular dependent variable.
A multiple linear regression model has the general form
(1)
where
Yi = the ith observation of the dependent variable Y
Xji = the ith observation of the independent variable Xj, j = 1, 2, …, k
b0 = the intercept of the equation
b1, …, bk = the slope coefficients for each of the independent variables
εi = the error term
n = the number of observations
A slope coefficient, bj, measures how much the dependent variable, Y, changes when the independent variable, Xj, changes by one unit, holding all other independent variables constant. For example, if b1 = 1 and all of the other independent variables remain constant, then we predict that if X1 increases by one unit, Y will also increase by one unit. If b1 = −1 and all of the other independent variables are held constant, then we predict that if X1 increases by one unit, Y will decrease by one unit. Multiple linear regression estimates b0, ..., bk. In this reading, we will refer to both the intercept, b0, and the slope coefficients, b1, ..., bk, as regression coefficients. As we proceed with our discussion, keep in mind that a regression equation has k slope coefficients and k + 1 regression coefficients.
Although Equation 1 may seem to apply only to cross-sectional data because the notation for the observations is the same (i = 1, …, n), all of these results apply to
time-series data as well. For example, if we analyze data from many time periods for one company, we would typically use the notation Yt, X1t, X2t, …, Xkt, in which the first subscript denotes the variable and the second denotes the tth time period.
In practice, we use software to estimate a multiple regression model. Example 1 presents an application of multiple regression analysis in investment practice. In the course of discussing a hypothesis test, Example 1 presents typical regression output and its interpretation.
EXAMPLE 1 Explaining the Bid–Ask Spread
As the manager of the trading desk at an investment management firm, you have noticed that the average bid–ask spreads of different NASDAQ-listed stocks can vary widely. When the ratio of a stock’s bid–ask spread to its price is higher than for another stock, your firm’s costs of trading in that stock tend to be higher. You have formulated the hypothesis that NASDAQ stocks’ percentage bid–ask spreads are related to the number of market makers and the company’s stock market capitalization. You have decided to investigate your hypothesis using multiple regression analysis.
You specify a regression model in which the dependent variable measures the percentage bid–ask spread and the independent variables measure the number of market makers and the company’s stock market capitalization. The regression is estimated using data from 31 December 2013 for 2,587 NASDAQ-listed stocks. Based on earlier published research exploring bid–ask spreads, you express the dependent and independent variables as natural logarithms, a so-called log-log regression model. A log-log regression model may be appropriate when one believes that proportional changes in the dependent variable bear a constant relationship to proportional changes in the independent variable(s), as we illustrate below. You formulate the multiple regression:
(2)
where
Yi = the natural logarithm of (Bid–ask spread/Stock price) for stock i
X1i = the natural logarithm of the number of NASDAQ market makers for stock i
X2i = the natural logarithm of the market capitalization (measured in millions
of US$) of company i
In a log-log regression such as Equation 2, the slope coefficients are interpreted as elasticities, assumed to be constant. For example, a value of b2 = −0.75 would mean that for a 1 percent increase in the market capitalization, we expect Bid–ask spread/Stock price to decrease by 0.75 percent, holding all other independent variables constant.2
Reasoning that greater competition tends to lower costs, you suspect that the greater the number of market makers, the smaller the percentage bid–ask spread. Therefore, you formulate a first null hypothesis (H0) and alternative hypothesis (Ha):
The null hypothesis is the hypothesis that the “suspected” condition is not true. If the evidence supports rejecting the null hypothesis and accepting the alternative hypothesis, you have statistically confirmed your suspicion.3
You also believe that the stocks of companies with higher market capitalization may have more-liquid markets, tending to lower percentage bid–ask spreads. Therefore, you formulate a second null hypothesis and alternative hypothesis:
For both tests, we use a t-test, rather than a z-test, because we do not know the population variance of b1 and b2. Suppose that you choose a 0.01 significance level for both tests.
TABLE 1 Results from Regressing ln(Bid–Ask Spread/Price) on ln(Number of Market Makers) and ln(Market Capitalization)
Coefficient Standard Error t-Statistic Intercept 1.5949 0.2275 7.0105
ln(Number of NASDAQ market makers) −1.5186 0.0808 −18.7946 ln(Company’s market capitalization) −0.3790 0.0151 −25.0993
ANOVA df SS MSS F Significance F
Regression 2 3,728.1334 1,864.0667 2,216.75 0.00 Residual 2,584 2,172.8870 0.8409 Total 2,586 5,901.0204
Residual standard error 0.9170 Multiple R-squared 0.6318
Observations 2,587
Source: Center for Research in Security Prices, University of Chicago.
Table 1 shows the results of estimating this linear regression using data from 31 December 2013.
If the regression result is not significant, we follow the useful principle of not proceeding to interpret the individual regression coefficients. Thus the analyst might look first at the analysis of variance (ANOVA) section, which addresses the regression’s overall significance.
The ANOVA (analysis of variance) section reports quantities related to the overall explanatory power and significance of the regression. SS stands for sum of squares, and MSS stands for mean sum of squares (SS divided by df). The F-test reports the overall significance of the regression. For example, an entry of 0.01 for the significance of F means that the regression is significant at the 0.01 level. In Table 1, the regression is even more significant because the significance of F is 0 at two decimal places. Later in the reading, we will present more information on the F-test.
Having ascertained that the overall regression is highly significant, an analyst might turn to the first listed column in the first section of the regression output.
The Coefficient column gives the estimates of the intercept, b0, and the slope coefficients, b1 and b2. The estimated intercept is positive, but both estimated slope coefficients are negative. Are these estimated regression coefficients significantly different from zero? The Standard Error column gives the standard error (the standard deviation) of the estimated regression coefficients. The test statistic for hypotheses concerning the population value of a regression coefficient has the form (Estimated regression coefficient – Hypothesized population value of the regression coefficient)/(Standard error of the regression coefficient). This is a t-test. Under the null hypothesis, the hypothesized population value of the regression coefficient is 0. Thus (Estimated regression coefficient)/(Standard error of the regression
coefficient) is the t-statistic given in the third column. For example, the t- statistic for the intercept is 1.5949/0.2275 = 7.0105. To evaluate the significance of the t-statistic we need to determine a quantity called degrees of freedom (df).4 The calculation is Degrees of freedom = Number of observations – (Number of independent variables + l) = n − (k + l).
The final section of Table 1 presents two measures of how well the estimated regression fits or explains the data. The first is the standard deviation of the regression residual, the residual standard error. This standard deviation is called the standard error of estimate (SEE). The second measure quantifies the degree of linear association between the dependent variable and all of the independent variables jointly. This measure is known as multiple R2 or simply R2 (the square of the correlation between predicted and actual values of the dependent variable).5 A value of 0 for R2 indicates no linear association; a value of l indicates perfect linear association. The final item in Table 1 is the number of observations in the sample (2,587).
Having reviewed the meaning of typical regression output, we can return to complete the hypothesis tests. The estimated regression supports the hypothesis that the greater the number of market makers, the smaller the percentage bid– ask spread: We reject H0: b1 ≥ 0 in favor of Ha: b1 < 0. The results also support the belief that the stocks of companies with higher market capitalization have lower percentage bid–ask spreads: We reject H0: b2 ≥ 0 in favor of Ha: b2 < 0.
To see that the null hypothesis is rejected for both tests, we can use t-test tables. For both tests, df = 2,587 − 3 = 2,584. The tables do not give critical values for degrees of freedom that large. The critical value for a one-tailed test with df = 200 at the 0.01 significance level is 2.345; for a larger number of degrees of freedom, the critical value would be even smaller in magnitude. Therefore, in our one-sided tests, we reject the null hypothesis in favor of the alternative hypothesis if
where
= the regression estimate of bj, j = 1, 2
bj = the hypothesized value6 of the coefficient (0)
= the estimated standard error of
The t-values of −18.7946 and −25.0993 for the estimates of b1 and b2, respectively, are both less than −2.345.
Before proceeding further, we should address the interpretation of a prediction stated in natural logarithm terms. We can convert a natural logarithm to the original units by taking the antilogarithm. To illustrate this conversion, suppose that a particular stock has 20 NASDAQ market makers and a market capitalization of $100 million. The natural logarithm of the number of NASDAQ market makers is equal to ln 20 = 2.9957, and the natural logarithm of the company’s market cap (in millions) is equal to ln 100 = 4.6052. With these values, the regression model predicts that the natural log of the ratio of the bid–ask spread to the stock price will be 1.5949 + (−1.5186 × 2.9957) + (−0.3790 × 4.6052) = −4.6997. We take the antilogarithm of −4.6997 by raising e to that power: e−4.6997 = 0.0091. The predicted bid–ask spread will be 0.91 percent of the stock price.7 Later we state the assumptions of the multiple regression model; before using an estimated regression to make predictions in actual practice, we should assure ourselves that those assumptions are satisfied.
In Table 1, we presented output common to most regression software programs. Many software programs also report p-values for the regression coefficients.8 For each regression coefficient, the p-value would be the smallest level of significance at which we can reject a null hypothesis that the population value of the coefficient is 0, in a two-sided test. The lower the p-value, the stronger the evidence against that null hypothesis. A p-value quickly allows us to determine if an independent variable is significant at a conventional significance level such as 0.05, or at any other standard we believe is appropriate.
Having estimated Equation 1, we can write
where stands for the predicted value of Yi, and , , and , stand for the estimated values of b0, b1, and b2, respectively. How should we interpret the estimated slope coefficients −1.5186 and −0.3790?
Interpreting the slope coefficients in a multiple linear regression model is different than doing so in the one-independent-variable regressions explored in the reading on correlation and regression. Suppose we have a one-independent-variable
regression that we estimate as . The interpretation of the slope estimate 0.75 is that for every 1-unit increase in X1, we expect Y to increase by 0.75 units. If we were to add a second independent variable to the equation, we would
generally find that the estimated coefficient on X1 is not 0.75 unless the second independent variable were uncorrelated with X1. The slope coefficients in a multiple regression are known as partial regression coefficients or partial slope coefficients and need to be interpreted with care.9 Suppose the coefficient on X1 in a regression with the second independent variable was 0.60. Can we say that for every 1-unit increase in X1, we expect Y to increase by 0.60 units? Not without qualification. For every 1-unit increase in X1, we still expect Y to increase by 0.75 units when X2 is not held constant. We would interpret 0.60 as the expected increase in Y for a 1-unit increase X1 holding the second independent variable constant.
To explain what the shorthand reference “holding the second independent constant” refers to, if we were to regress X1 on X2, the residuals from that regression would represent the part of X1 that is uncorrelated with X2. We could then regress Y on those residuals in a one-independent-variable regression. We would find that the slope coefficient on the residuals would be 0.60; by construction, 0.60 would represent the expected effect on Y of a 1-unit increase in X1 after removing the part of X1 that is correlated with X2. Consistent with this explanation, we can view 0.60 as the expected net effect on Y of a 1-unit increase in X1, after accounting for any effects of the other independent variables on the expected value of Y. To reiterate, a partial regression coefficient measures the expected change in the dependent variable for a 1-unit increase in an independent variable, holding all the other independent variables constant.
To apply this process to the regression in Table 1, we see that the estimated coefficient on the natural logarithm of market capitalization is −0.3790. Therefore, the model predicts that an increase of 1 in the natural logarithm of the company’s market capitalization is associated with a −0.3790 change in the natural logarithm of the ratio of the bid–ask spread to the stock price, holding the natural logarithm of the number of market makers constant. We need to be careful not to expect that the natural logarithm of the ratio of the bid–ask spread to the stock price would differ by −0.3790 if we compared two stocks for which the natural logarithm of the company’s market capitalization differed by 1, because in all likelihood the number of market makers for the two stocks would differ as well, which would affect the dependent variable. The value −0.3790 is the expected net effect of difference in log market capitalizations, net of the effect of the log number of market makers on the expected value of the dependent variable.
2.1. Assumptions of the Multiple Linear Regression Model Before we can conduct correct statistical inference on a multiple linear regression
model (a model with more than one independent variable estimated using ordinary least squares), we need to know the assumptions underlying that model.10 Suppose we have n observations on the dependent variable, Y, and the independent variables, X1, X2, ..., Xk, and we want to estimate the equation Yi = b0 + b1X1i + b2X2i + ... + bkXki + εi.
In order to make a valid inference from a multiple linear regression model, we need to make the following six assumptions, which as a group define the classical normal multiple linear regression model: 1. The relationship between the dependent variable, Y, and the independent
variables, X1, X2, ..., Xk, is linear as described in Equation 1.
2. The independent variables (X1, X2, ..., Xk) are not random.11 Also, no exact linear relation exists between two or more of the independent variables.12
3. The expected value of the error term, conditioned on the independent variables, is 0: E(ε | X1, X2, …, Xk) = 0.
4. The variance of the error term is the same for all observations:13 .
5. The error term is uncorrelated across observations: E(εiεj) = 0, j ≠. i.
6. The error term is normally distributed.
Note that these assumptions are almost exactly the same as those for the single- variable linear regression model. Assumption 2 is modified such that no exact linear relation exists between two or more independent variables or combinations of independent variables. If this part of Assumption 2 is violated, then we cannot compute linear regression estimates.14 Also, even if no exact linear relationship exists between two or more independent variables, or combinations of independent variables, linear regression may encounter problems if two or more of the independent variables or combinations thereof are highly correlated. Such a high correlation is known as multicollinearity, which we will discuss later in this reading. We will also discuss the consequences of conducting regression analysis premised on Assumptions 4 and 5 being met when, in fact, they are violated.
Although Equation 1 may seem to apply only to cross-sectional data because the notation for the observations is the same (i = 1, ..., n), all of these results apply to time-series data as well. For example, if we analyze data from many time periods for one company, we would typically use the notation Yt, X1t, X2t, ..., Xkt, in which the first subscript denotes the variable and the second denotes the tth time period.
EXAMPLE 2 Factors Explaining the Valuations of Multinational Corporations
Kyaw, Manley, and Shetty (2011) examined which factors affect the valuation of a multinational corporation (MNC). Specifically, they wanted to know whether political risk, transparency, and geographic diversification affected the valuations of MNCs. They used data for 450 US MNCs from 1998 to 2003. The valuations of these corporations were measured using Tobin’s q, a commonly used measure of corporate valuation that is calculated as the ratio of the sum of the market value of a corporation’s equity and the book value of long-term debt to the sum of the book values of equity and long-term debt. The authors regressed Tobin’s q of MNCs on variables representing political risk, transparency, and geographic diversification. The authors also included some additional variables that may affect company valuation, including size, leverage, and beta.15 They used the equation
where
Table 2 shows the results of their analysis.16
TABLE 2 Results from Regressing Tobin’s q on Factors Affecting the Value of
Multinational Corporations
Coefficient Standard Error* t-Statistic Intercept 19.829 4.798 4.133 Size –0.712 0.228 –3.123
Leverage –3.897 0.987 –3.948 Beta –1.032 0.261 –3.954
Political risk –2.079 0.763 –2.725 Transparency –0.129 0.050 –2.580
Geographic diversification 0.021 0.010 2.100
* This study combines time series observations with cross-sectional observations; such data are commonly referred to as panel data. In such a setting, the standard errors need to be corrected for bias by using a clustered standard error approach as in Petersen (2009). The standard errors reported in this table are clustered standard errors.
Source: Kyaw, Manley, and Shetty (2011).
Suppose that we use the results in Table 2 to test the null hypothesis that the size of a multinational corporation has no effect on its value. Our null hypothesis is that the coefficient on the size variable equals 0 (H0: b1 = 0), and our alternative hypothesis is that the coefficient does not equal 0 (Ha: b1 | 0). The t-statistic for testing that hypothesis is
With 450 observations and seven coefficients, the t-statistic has 450 – 7 = 443 degrees of freedom. At the 0.05 significance level, the critical value for t is about 1.97. The absolute value of computed t-statistic on the size coefficient is 3.12, which suggests strongly that we can reject the null hypothesis that size is unrelated to MNC value. In fact, the critical value for t is about 2.6 at the 0.01 significance level.
Because Sizei,t is the natural (base e or 2.72) log of sales, an increase of 1 in Sizei,t is the same as a 2.72-fold increase in sales. Thus, the estimated coefficient of approximately –0.7 for Sizei,t implies that every 2.72-fold increase in sales of the MNC (an increase of 1 in Sizei,t) is associated with an expected decrease of 0.7 in Tobin’s qi,t of the MNC, holding constant the other
five independent variables in the regression.
Now suppose we want to test the null hypothesis that geographic diversification is not related to Tobin’s q; we want to test whether the coefficient on geographic diversification equals 0 (H0: b6 = 0) against the alternative hypothesis that the coefficient on geographic diversification does not equal 0 (Ha: b6 | 0). The t-statistic to test this hypothesis is
The critical value of the t-test is 1.97 at the 0.05 significance level. Therefore, at the 0.05 significance level, we can reject the null hypothesis that geographic diversification has no effect on MNC valuation. We can interpret the coefficient on geographic diversification of 0.021 as implying that an increase of 1 in the percentage of MNC’s sales that are foreign sales is associated with an expected 0.021 increase in Tobin’s q for the MNC, holding all other independent variables constant.
EXAMPLE 3 Explaining Returns to the Fidelity Select Technology Portfolio
Suppose you are considering an investment in the Fidelity Select Technology Portfolio (FSPTX), a US mutual fund specializing in technology stocks. You want to know whether the fund behaves more like a large-cap growth fund or a large-cap value fund.17 You decide to estimate the regression
where
Yt = the monthly return to the FSPTX
X1t = the monthly return to the S&P 500 Growth Index
X2t = the monthly return to the S&P 500 Value Index
The S&P 500 Growth and S&P 500 Value indices represent predominantly large-cap growth and value stocks, respectively.
Table 3 shows the results of this linear regression using monthly data from January 2009 through December 2013. The estimated intercept in the
regression is 0.0018. Thus, if both the return to the S&P 500 Growth Index and the return to the S&P 500 Value Index equal 0 in a specific month, the regression model predicts that the return to the FSPTX will be 0.18 percent. The coefficient on the large-cap growth index is 1.4697, and the coefficient on the large-cap value index return is −0.1833. Therefore, if in a given month the return to the S&P 500 Growth Index was 1 percent and the return to the S&P 500 Value Index was −2 percent, the model predicts that the return to the FSPTX would be 0.0018 + 1.4697(0.01) − 0.1833(−0.02) = 2.02 percent.
TABLE 3 Results from Regressing the FSPTX Returns on the S&P 500 Growth and S&P 500 Value Indices
Coefficient Standard Error t-Statistic Intercept 0.0018 0.0038 0.4737
S&P 500 Growth Index 1.4697 0.2479 5.9286 S&P 500 Value Index −0.1833 0.2034 −0.9012
ANOVA df SS MSS F Significance F Regression 2 0.1653 0.0826 113.7285 1.27E-20 Residual 57 0.0414 0.0007 Total 59 1.2067
Residual standard error 0.0270 Multiple R-squared 0.7996
Observations 60
Source: Bloomberg, finance.yahoo.com.
We may want to know whether the coefficient on the returns to the S&P 500 Value Index is statistically significant. Our null hypothesis states that the coefficient equals 0 (H0: b2 = 0); our alternative hypothesis states that the coefficient does not equal 0 (Ha: b2 | 0).
Our test of the null hypothesis uses a t-test constructed as follows:
where
= the regression estimate of b2
b2 = the hypothesized value18 of the coefficient (0)
= the estimated standard error of
This regression has 60 observations and three coefficients (two independent variables and the intercept); therefore, the t-test has 60 − 3 = 57 degrees of freedom. At the 0.05 significance level, the critical value for the test statistic is about 2.00. The absolute value of the test statistic is 0.9012. Because the test statistic’s absolute value is less than the critical value (0.9012 < 2.00), we fail to reject the null hypothesis that b2 = 0. (Note that the t-tests reported in Table 3, as well as the other regression tables, are tests of the null hypothesis that the population value of a regression coefficient equals 0.)
Similar analysis shows that at the 0.05 significance level, we cannot reject the null hypothesis that the intercept equals 0 (H0: b0 = 0) in favor of the alternative hypothesis that the intercept does not equal 0 (Ha: b0 | 0). Table 3 shows that the t-statistic for testing that hypothesis is 0.4737, a result smaller in absolute value than the critical value of 2.00. However, at the 0.05 significance level we can reject the null hypothesis that the coefficient on the S&P 500 Growth Index equals 0 (H0: b1 = 0) in favor of the alternative hypothesis that the coefficient does not equal 0 (Ha: b1 | 0). As Table 3 shows, the t-statistic for testing that hypothesis is 5.928, a result far above the critical value of 2.00. Thus multiple regression analysis suggests that returns to the FSPTX are very closely associated with the returns to the S&P 500 Growth Index, but they are not related to the S&P 500 Value Index (the t-statistic of 0.9012 is not statistically significant).
2.2. Predicting the Dependent Variable in a Multiple Regression Model Financial analysts often want to predict the value of the dependent variable in a multiple regression based on assumed values of the independent variables. We have previously discussed how to make such a prediction in the case of only one independent variable. The process for making that prediction with multiple linear regression is very similar.
To predict the value of a dependent variable using a multiple linear regression
model, we follow these three steps:
1. Obtain estimates , , , …, of the regression parameters b0, b1, b2, …, bk.
2. Determine the assumed values of the independent variables, , , …, .
3. Compute the predicted value of the dependent variable, , using the equation
(3)
Two practical points concerning using an estimated regression to predict the dependent variable are in order. First, we should be confident that the assumptions of the regression model are met. Second, we should be cautious about predictions based on values of the independent variables that are outside the range of the data on which the model was estimated; such predictions are often unreliable.
EXAMPLE 4 Predicting a Multinational Corporation’s Tobin’s q
In Example 2, we explained the Tobin’s q for US multinational corporations (MNC) based on the natural log of sales, leverage, beta, political risk, transparency, and geographic diversification. To review the regression equation:
Now we can use the results of the regression reported in Table 2 (excerpted here) to predict the Tobin’s q for a US MNC.
TABLE 2 (Excerpt)
Coefficient Intercept 19.829 Size –0.712
Leverage –3.897 Beta –1.032
Political risk –2.079 Transparency –0.129
Geographic diversification 0.021
1. Suppose that a particular MNC has the following data for a given year.
Total sales of $7,600 million. The natural log of total sales in millions of US$ equals ln(7,600) = 8.94.
Leverage (Total debt/Total assets) of 0.45.
Beta of 1.30.
Political risk of 0.47, implying that the ratio of the number of safe countries to the total number of foreign countries in which the MNC has operations is 0.53.
Transparency score of 65, indicating 65 percent “yes” answers to survey questions related to the corporation’s transparency.
Geographic diversification of 30, indicating that 30 percent of the corporation’s sales are in foreign countries.
What is the predicted Tobin’s q for the above MNC?
Solution to 1: The predicted Tobin’s q for the MNC, based on the regression, is:
When predicting the dependent variable using a linear regression model, we encounter two types of uncertainty: uncertainty in the regression model itself, as reflected in the standard error of estimate, and uncertainty about the estimates of the regression model’s parameters. In the reading on correlation and regression, we presented procedures for constructing a prediction interval for linear regression with one independent variable. For multiple regression, however, computing a prediction interval to properly incorporate both types of uncertainty requires matrix algebra, which is outside the scope of this reading.19
2.3. Testing whether All Population Regression Coefficients Equal Zero Earlier, we illustrated how to conduct hypothesis tests on regression coefficients individually. What if we now want to test the significance of the regression as a whole? As a group, do the independent variables help explain the dependent variable? To address this question, we test the null hypothesis that all the slope
coefficients in a regression are simultaneously equal to 0. In this section, we further discuss ANOVA with regard to a regression’s explanatory power and the inputs for an F-test of the above null hypothesis.
If none of the independent variables in a regression model helps explain the dependent variable, the slope coefficients should all equal 0. In a multiple regression, however, we cannot test the null hypothesis that all slope coefficients equal 0 based on t-tests that each individual slope coefficient equals 0, because the individual tests do not account for the effects of interactions among the independent variables. For example, a classic symptom of multicollinearity is that we can reject the hypothesis that all the slope coefficients equal 0 even though none of the t- statistics for the individual estimated slope coefficients is significant. Conversely, we can construct unusual examples in which the estimated slope coefficients are significantly different from 0 although jointly they are not.
To test the null hypothesis that all of the slope coefficients in the multiple regression model are jointly equal to 0 (H0: b1 = b2 = ... = bk = 0) against the alternative hypothesis that at least one slope coefficient is not equal to 0 we must use an F-test. The F-test is viewed as a test of the regression’s overall significance.
To correctly calculate the test statistic for the null hypothesis, we need four inputs:
total number of observations, n;
total number of regression coefficients to be estimated, k + 1, where k is the number of slope coefficients;
sum of squared errors or residuals, , abbreviated SSE, also known as the residual sum of squares (unexplained variation);20 and
regression sum of squares, , abbreviated RSS.21 This amount is the variation in Y from its mean that the regression equation explains (explained variation).
The F-test for determining whether the slope coefficients equal 0 is based on an F- statistic calculated using the four values listed above. The F-statistic measures how well the regression equation explains the variation in the dependent variable; it is the ratio of the mean regression sum of squares to the mean squared error.
We compute the mean regression sum of squares by dividing the regression sum of squares by the number of slope coefficients estimated, k. We compute the mean squared error by dividing the sum of squared errors by the number of observations,
n, minus (k + 1). The two divisors in these computations are the degrees of freedom for calculating an F-statistic. For n observations and k slope coefficients, the F-test for the null hypothesis that the slope coefficients are all equal to 0 is denoted Fk,n– (k+1). The subscript indicates that the test should have k degrees of freedom in the numerator (numerator degrees of freedom) and n − (k + 1) degrees of freedom in the denominator (denominator degrees of freedom).
The formula for the F-statistic is
(4)
where MSR is the mean regression sum of squares and MSE is the mean squared error. In our regression output tables, MSR and MSE are the first and second quantities under the MSS (mean sum of squares) column in the ANOVA section of the output. If the regression model does a good job of explaining variation in the dependent variable, then the ratio MSR/MSE will be large.
What does this F-test tell us when the independent variables in a regression model explain none of the variation in the dependent variable? In this case, each predicted value in the regression model, , has the average value of the dependent variable,
, and the regression sum of squares, is 0. Therefore, the F-statistic for testing the null hypothesis (that all the slope coefficients are equal to 0) has a value of 0 when the independent variables do not explain the dependent variable at all.
To specify the details of making the statistical decision when we have calculated F, we reject the null hypothesis at the α significance level if the calculated value of F is greater than the upper α critical value of the F distribution with the specified numerator and denominator degrees of freedom. Note that we use a one-tailed F- test.22
We can illustrate the test using Example 1, in which we investigated whether the natural log of the number of NASDAQ market makers and the natural log of the stock’s market capitalization explained the natural log of the bid–ask spread divided by price. Assume that we set the significance level for this test to α = 0.05 (i.e., a 5 percent probability that we will mistakenly reject the null hypothesis if it is true). Table 1 (excerpted here) presents the results of variance computations for this regression.
This model has two slope coefficients (k = 2), so there are two degrees of freedom in the numerator of this F-test. With 2,587 observations in the sample, the number of
degrees of freedom in the denominator of the F-test is n − (k + 1) = 2,587 − 3 = 2,584. The sum of the squared errors is 2,172.8870. The regression sum of squares is 3,728.1334. Therefore, the F-test for the null hypothesis that the two slope coefficients in this model equal 0 is
TABLE 1 (Excerpt)
ANOVA df SS MSS F Significance F Regression 2 3,728.1334 1,864.0667 2,216.7505 0.00 Residual 2,584 2,172.8870 0.8409 Total 2,586 5,901.0204
This test statistic is distributed as an F2,2,584 random variable under the null hypothesis that the slope coefficients are equal to 0. In the table for the 0.05 significance level, we look at the second column, which shows F-distributions with two degrees of freedom in the numerator. Near the bottom of the column, we find that the critical value of the F-test needed to reject the null hypothesis is between 3.00 and 3.07.23 The actual value of the F-test statistic at 2,216.75 is much greater, so we reject the null hypothesis that coefficients of both independent variables equal 0. In fact, Table 1 under “Significance F,” reports a p-value of 0. This p-value means that the smallest level of significance at which the null hypothesis can be rejected is practically 0. The large value for this F-statistic implies a minuscule probability of incorrectly rejecting the null hypothesis (a mistake known as a Type I error).
2.4. Adjusted R2
In the reading on correlation and regression, we presented the coefficient of determination, R2, as a measure of the goodness of fit of an estimated regression to the data. In a multiple linear regression, however, R2 is less appropriate as a measure of whether a regression model fits the data well (goodness of fit). Recall that R2 is defined as
The numerator equals the regression sum of squares, RSS. Thus R2 states RSS as a
fraction of the total sum of squares, . If we add regression variables to the model, the amount of unexplained variation will decrease, and RSS will increase, if
the new independent variable explains any of the unexplained variation in the model. Such a reduction occurs when the new independent variable is even slightly correlated with the dependent variable and is not a linear combination of other independent variables in the regression.24 Consequently, we can increase R2 simply by including many additional independent variables that explain even a slight amount of the previously unexplained variation, even if the amount they explain is not statistically significant.
Some financial analysts use an alternative measure of goodness of fit called adjusted R2, or . This measure of fit does not automatically increase when another variable is added to a regression; it is adjusted for degrees of freedom. Adjusted R2 is typically part of the multiple regression output produced by statistical software packages.
The relation between R2 and is
where n is the number of observations and k is the number of independent variables (the number of slope coefficients). Note that if k ≥ 1, then R2 is strictly greater than adjusted R2.
When a new independent variable is added, can decrease if adding that variable results in only a small increase in R2. In fact, can be negative, although R2 is always nonnegative.25 If we use to compare regression models, it is important that the dependent variable be defined the same way in both models and that the sample sizes used to estimate the models are the same.26 For example, it makes a difference for the value of if the dependent variable is GDP (gross domestic product) or ln(GDP), even if the independent variables are identical. Furthermore, we should be aware that a high does not necessarily indicate that the regression is well specified in the sense of including the correct set of variables.27 One reason for caution is that a high may reflect peculiarities of the dataset used to estimate the regression. To evaluate a regression model, we need to take many other factors into account, as we discuss in Section 5.1.
3. Using Dummy Variables in Regressions Often, financial analysts need to use qualitative variables as independent variables in a regression. One type of qualitative variable, called a dummy variable, takes on a value of 1 if a particular condition is true and 0 if that condition is false.28 For example, suppose we want to test whether stock returns were different in January than during the remaining months of a particular year. We include one independent variable in the regression, X1t, that has a value of 1 for each January and a value of 0 for every other month of the year. We estimate the regression model
In this equation, the coefficient b0 is the average value of Yt in months other than January, and b1 is the difference between the average value of Yt in January and the average value of Yt in months other than January.
We need to exercise care in choosing the number of dummy variables in a regression. The rule is that if we want to distinguish among n categories, we need n − 1 dummy variables. For example, to distinguish between during January and not during January above (n = 2 categories), we used one dummy variable (n − 1 = 2 − 1 = 1). If we want to distinguish between each of the four quarters in a year, we would include dummy variables for three of the four quarters in a year. If we make the mistake of including dummy variables for four rather than three quarters, we have violated Assumption 2 of the multiple regression model and cannot estimate the regression. The next example illustrates the use of dummy variables in a regression with monthly data.
EXAMPLE 5 Month-of-the-Year Effects on Japanese Small- Stock Returns]
For many years, financial analysts have been concerned about seasonality in stock returns.29 In particular, analysts have researched whether returns to small stocks differ during various months of the year. Suppose we want to test whether total returns to one small-stock index, the MSCI Japan Small Cap Index, differ by month. Using data from January 2001 (the first available date for these data) through the end of 2013, we can estimate a regression including an intercept and 11 dummy variables, one for each of the first 11 months of the year. The equation that we estimate is
where each monthly dummy variable has a value of 1 when the month occurs (e.g., Jan1 = Jan13 = 1, as the first observation is a January) and a value of 0 for the other months. Table 4 shows the results of this regression.
The intercept, b0, measures the average return for stocks in December because there is no dummy variable for December.30 This equation estimates that the average return in December is 2.73 percent ( = 0.0273). Each of the estimated coefficients for the dummy variables shows the estimated difference between returns in that month and returns for December. So, for example, the estimated additional return in January is 2.13 percent lower than December ( = –0.0213). This gives a January return prediction of 0.60 percent (2.73 in December – 2.13 corresponding to the January coefficient).
TABLE 4 Results from Regressing MSCI Japan Small Cap Index Returns on Monthly Dummy Variables
Coefficient Standard Error t-Statistic Intercept 0.0273 0.0149 1.8322 January –0.021 0.0210 1.0143 February –0.0112 0.0210 –0.5533 March 0.0101 0.0210 0.4810 April –0.0012 0.0210 –0.0571 May –0.0425 0.0210 –2.0238 June –0.0065 0.0210 –0.31095 July –0.0481 0.0210 –2.2905
August –0.0367 0.0210 –1.7476 September –0.0285 0.0210 –1.3571 October –0.0429 0.0210 –2.0429 November –0.0339 0.0210 –1.6153
ANOVA df SS MSS F Significance F Regression 11 0.0551 0.0050 1.7421 0.0698 Residual 144 0.4142 0.0029 Total 155 0.4693
Residual standard error 0.0536
Multiple R-squared 0.1174 Observations 156
Source: Morgan Stanley Capital International.
The low R2 in this regression (0.1174), however, suggests that a month-of-the- year effect in small-stock returns may not be very important for explaining small-stock returns. We can use the F-test to analyze the null hypothesis that jointly, the monthly dummy variables all equal 0 (H0: b1 = b2 = ... = b11 = 0). We are testing for significant monthly variation in small-stock returns. Table 4 shows the data needed to perform an analysis of variance. The number of degrees of freedom in the numerator of the F-test is 11; the number of degrees of freedom in the denominator is [156 − (11 + 1)] = 144. The regression sum of squares equals 0.0551, and the sum of squared errors equals 0.4142. Therefore, the F-statistic to determine whether all of the regression slope coefficients are jointly equal to 0 is
Appendix D (the F-distribution table) at the end of this volume shows the critical values for this F-test. If we choose a significance level of 0.05 and look in Column 11 (because the numerator has 11 degrees of freedom), we see that the critical value is 1.87 when the denominator has 120 degrees of freedom. The denominator actually has 144 degrees of freedom, so the critical value of the F-statistic is smaller than 1.87 (for df = 120) but larger than 1.79 (for an infinite number of degrees of freedom). The value of the test statistic is 1.74, so we cannot reject the null hypothesis that all of the coefficients jointly are equal to 0.
The p-value of 0.0698 shown for the F-test in Table 4 means that the smallest level of significance at which we can reject the null hypothesis is roughly 0.07, or 7 percent—which is above the conventional level of 5 percent. Among the 11 monthly dummy variables, May, July, and October have a t-statistic with an absolute value greater than 2. Although the coefficients for these dummy variables are statistically significant, we have so many insignificant estimated coefficients that we cannot reject the null hypothesis that returns are equal across the months. This test suggests that the significance of a few coefficients in this regression model may be the result of random variation. We may thus want to avoid portfolio strategies calling for differing investment weights for small stocks in different months.
EXAMPLE 6 Determinants of Short-Term Stock Return Performance in Mergers and Acquisitions by Chinese Companies
Bhabra and Huang (2013) examined short-term market reaction to mergers and acquisition deals initiated by Chinese companies listed on the Shanghai Stock Exchange and the Shenzhen Stock Exchange. They examined those deals during 1997 to 2007 in which the acquirer gained complete control of the target. As the measure of short-term stock return performance around the announcement day, they used the cumulative abnormal return on the acquirer ’s stock during a five-day window from day –2 to day +2, where day 0 is the acquisition announcement day. Cumulative abnormal return is the excess return achieved over a stated period measured in relation to the return expected given a security’s risk. The independent variables in their model included the following firm-and deal-related factors that may affect short-term stock return performance:
profit margin: ratio of acquiring firm’s net income to revenue prior to the acquisition;
sales growth: acquiring firm’s annual sales growth rate prior to the acquisition;
change in leverage: change in the ratio of debt to total assets for the acquiring firm due to the acquisition;
firm value: natural logarithm of the market value of the acquirer in the announcement year;
same industry: Dummy variable (1 = the acquiring and target firms are in the same industry, 0 = in different industries);
state-owned enterprise: Dummy variable (1 = acquiring firm is a state-owned firm, 0 = not a state-owned firm);
cash: Dummy variable (1 = the form of payment in the transaction is cash, 0 = other forms of payment);
cross-border: Dummy variable (1 = cross-border deal, 0 = domestic deal);
private: Dummy variable (1 = target firm is a stand-alone private firm, 0 = not a stand-alone private firm);
missing method of payment: Dummy variable (1 = the form of payment information is not available, 0 = form of payment information is available).
Table 5 shows the authors’ results.
TABLE 5 Multiple Regression Model of Cumulative Abnormal Returns for Chinese Acquisitions, 1997–2007
Coefficient p-Value Intercept −0.0543 0.2316
Profit margin 0.0000 0.1522 Sales growth −0.0180 0.2774
Change in leverage −0.0136 0.5887 Firm value 0.0024 0.2348
Same industry 0.0012 0.4296 State-owned enterprise 0.0333 0.0435
Cash 0.0041 0.7912 Cross-border −0.0311 0.3376
Private −0.0336 0.0204 Missing method of payment 0.0013 0.9399
R-squared 0.2194 Observations 87
Source: Bhabra and Huang (2013).
We can summarize Bhabra and Huang’s findings as follows:
The coefficient of state-owned enterprise in this regression model is positive and statistically significant at the 0.05 level as the p-value is less than 0.05. State-owned firms play a very important role in the Chinese economy. Non- state-owned firms are relatively new entrants in the Chinese market and are smaller firms. The statistically significant coefficient of state-owned enterprise suggests that the short-term increase in firm value is greater when the acquiring firm is state owned.
The coefficient of private targets is also statistically significant at the 0.05 level. The sign of this coefficient is negative. Bhabra and Huang point out that the vast majority of target firms in Chinese M&As are unlisted firms, either
stand-alone private firms or subsidiaries of listed firms. All the firms included in their sample are either subsidiaries or stand-alone private firms. The significantly negative coefficient of the dummy for private targets tends to result in lower cumulative abnormal returns compared to acquisitions of unlisted subsidiaries. The authors point out that a possible reason could be relatively limited data accessibility for stand-alone private firms as compared with subsidiaries of listed parents. Because of the challenges faced by acquirers when estimating the value and prospects of the private firms, acquisitions of subsidiaries elicit a more positive stock price response.
Although none of the other coefficients are statistically significant in the above model, in some of the other models estimated in the study (not included in this reading), the authors find some evidence that the stock price response is more positive when the target is in the same industry as the acquirer and the change in leverage is low.
4. Violations of Regression Assumptions In Section 2.1, we presented the assumptions of the multiple linear regression model. Inference based on an estimated regression model rests on those assumptions being satisfied. In applying regression analysis to financial data, analysts need to be able to diagnose violations of regression assumptions, understand the consequences of violations, and know the remedial steps to take. In the following sections we discuss three regression violations: heteroskedasticity, serial correlation, and multicollinearity.
4.1. Heteroskedasticity So far, we have made an important assumption that the variance of error in a regression is constant across observations. In statistical terms, we assumed that the errors were homoskedastic. Errors in financial data, however, are often heteroskedastic: the variance of the errors differs across observations. In this section, we discuss how heteroskedasticity affects statistical analysis, how to test for heteroskedasticity, and how to correct for it.
We can see the difference between homoskedastic and heteroskedastic errors by comparing two graphs. Figure 1 shows the values of the dependent and independent variables and a fitted regression line for a model with homoskedastic errors. There is no systematic relationship between the value of the independent variable and the regression residuals (the vertical distance between a plotted point and the fitted regression line). Figure 2 shows the values of the dependent and independent variables and a fitted regression line for a model with heteroskedastic errors. Here, a systematic relationship is visually apparent: On average, the regression residuals grow much larger as the size of the independent variable increases.
4.1.1. The Consequences of Heteroskedasticity
What are the consequences when the assumption of constant error variance is violated? Although heteroskedasticity does not affect the consistency31 of the regression parameter estimators, it can lead to mistakes in inference. When errors are heteroskedastic, the F-test for the overall significance of the regression is unreliable.32 Furthermore, t-tests for the significance of individual regression coefficients are unreliable because heteroskedasticity introduces bias into estimators of the standard error of regression coefficients. If a regression shows significant heteroskedasticity, the standard errors and test statistics computed by regression programs will be incorrect unless they are adjusted for heteroskedasticity.
FIGURE 1 Regression with Homoskedasticity
FIGURE 2 Regression with Heteroskedasticity
In regressions with financial data, the most likely result of heteroskedasticity is that the estimated standard errors will be underestimated and the t-statistics will be
inflated. When we ignore heteroskedasticity, we tend to find significant relationships where none actually exist.33 The consequences in practice may be serious if we are using regression analysis in the development of investment strategies. As Example 7 shows, the issue impinges even on our understanding of financial models.
EXAMPLE 7 Heteroskedasticity and Tests of an Asset Pricing Model
MacKinlay and Richardson (1991) examined how heteroskedasticity affects tests of the capital asset pricing model (CAPM).34 These authors argued that if the CAPM is correct, they should find no significant differences between the risk-adjusted returns for holding small stocks versus large stocks. To implement their test, MacKinlay and Richardson grouped all stocks on the New York Stock Exchange and the American Stock Exchange (now called NYSE MKT) by market-value decile with annual reassignment. They then tested for systematic differences in risk-adjusted returns across market-capitalization- based stock portfolios. They estimated the following regression:
where
ri,t = excess return (return above the risk-free rate) to portfolio i in period t
rm,t = excess return to the market as a whole in period t
The CAPM formulation hypothesizes that excess returns on a portfolio are explained by excess returns on the market as a whole. That hypothesis implies that αi = 0 for every portfolio i; on average, no excess return accrues to any portfolio after taking into account its systematic (market) risk.
Using data from January 1926 to December 1988 and a market index based on equal-weighted returns, MacKinlay and Richardson failed to reject the CAPM at the 0.05 level when they assumed that the errors in the regression model are normally distributed and homoskedastic. They found, however, that they could reject the CAPM when they corrected their test statistics to account for heteroskedasticity. They rejected the hypothesis that there are no size-based, risk-adjusted excess returns in historical data.35
We have stated that effects of heteroskedasticity on statistical inference can be severe. To be more precise about this concept, we should distinguish between two
broad kinds of heteroskedasticity: unconditional and conditional.
Unconditional heteroskedasticity occurs when heteroskedasticity of the error variance is not correlated with the independent variables in the multiple regression. Although this form of heteroskedasticity violates Assumption 4 of the linear regression model, it creates no major problems for statistical inference.
The type of heteroskedasticity that causes the most problems for statistical inference is conditional heteroskedasticity—heteroskedasticity in the error variance that is correlated with (conditional on) the values of the independent variables in the regression. Fortunately, many statistical software packages easily test and correct for conditional heteroskedasticity.
4.1.2. Testing for Heteroskedasticity
Because of conditional heteroskedasticity’s consequences on inference, the analyst must be able to diagnose its presence. The Breusch–Pagan test is widely used in finance research because of its generality.36
Breusch and Pagan (1979) suggested the following test for conditional heteroskedasticity: Regress the squared residuals from the estimated regression equation on the independent variables in the regression. If no conditional heteroskedasticity exists, the independent variables will not explain much of the variation in the squared residuals. If conditional heteroskedasticity is present in the original regression, however, the independent variables will explain a significant portion of the variation in the squared residuals. The independent variables can explain the variation because each observation’s squared residual will be correlated with the independent variables if the independent variables affect the variance of the errors.
Breusch and Pagan showed that under the null hypothesis of no conditional heteroskedasticity, nR2 (from the regression of the squared residuals on the independent variables from the original regression) will be a χ2 random variable with the number of degrees of freedom equal to the number of independent variables in the regression.37 Therefore, the null hypothesis states that the regression’s squared error term is uncorrelated with the independent variables. The alternative hypothesis states that the squared error term is correlated with the independent variables. Example 8 illustrates the Breusch–Pagan test for conditional heteroskedasticity.
EXAMPLE 8 Testing for Conditional Heteroskedasticity in the Relation between Interest Rates and Expected Inflation
Suppose an analyst wants to know how closely nominal interest rates are related to expected inflation to determine how to allocate assets in a fixed income portfolio. The analyst wants to test the Fisher effect, the hypothesis suggested by Irving Fisher that nominal interest rates increase by 1 percentage point for every 1 percentage point increase in expected inflation.38 The Fisher effect assumes the following relation between nominal interest rates, real interest rates, and expected inflation:
where
i = the nominal rate
r = the real interest rate (assumed constant)
πe = the expected rate of inflation
To test the Fisher effect using time-series data, we could specify the following regression model for the nominal interest rate:
(5)
Noting that the Fisher effect predicts that the coefficient on the inflation variable is 1, we can state the null and alternative hypotheses as
We might also specify a 0.05 significance level for the test. Before we estimate Equation 5, we must decide how to measure expected inflation ( ) and the nominal interest rate (it).
The Survey of Professional Forecasters (SPF) has compiled data on the quarterly inflation expectations of professional forecasters.39 We use those data as our measure of expected inflation. We use three-month Treasury bill returns as our measure of the (risk-free) nominal interest rate.40 We use quarterly data from the fourth quarter of 1968 to the fourth quarter of 2013 to estimate Equation 5. Table 6 shows the regression results.
To make the statistical decision on whether the data support the Fisher effect, we calculate the following t-statistic, which we then compare to its critical value.
With a t-statistic of 2.29 and 181 − 2 = 179 degrees of freedom, the critical t- value is about 1.97. If we have conducted a valid test, we cannot reject at the 0.05 significance level the hypothesis that the true coefficient in this regression is 1 and that the Fisher effect holds. The t-test assumes that the errors are homoskedastic. Before we accept the validity of the t-test, therefore, we should test whether the errors are conditionally heteroskedastic. If those errors prove to be conditionally heteroskedastic, then the test is invalid.
TABLE 6 Results from Regressing T-Bill Returns on Predicted Inflation
Coefficient Standard Error t-Statistic Intercept 0.0116 0.0033 3.5152
Inflation prediction 1.1744 0.0761 15.4323 Residual standard error 0.0233 Multiple R-squared 0.5708
Observations 181 Durbin–Watson statistic* 0.2980
Note: The Durbin–Watson statistic will be explained in Section 4.2.2.
Source: Federal Reserve Bank of Philadelphia, US Department of Commerce.
We can perform the Breusch–Pagan test for conditional heteroskedasticity on the squared residuals from the Fisher effect regression. The test regresses the squared residuals on the predicted inflation rate. The R2 in the squared residuals regression (not shown here) is 0.0666. The test statistic from this regression, nR2, is 181 × 0.0666 = 12.0546. Under the null hypothesis of no conditional heteroskedasticity, this test statistic is a χ2 random variable with one degree of freedom (because there is only one independent variable).
We should be concerned about heteroskedasticity only for large values of the test statistic. Therefore, we should use a one-tailed test to determine whether we can reject the null hypothesis. The critical value of the test statistic for a variable from a χ2 distribution with one degree of freedom at the 0.05 significance level is 3.84. The test statistic from the Breusch–Pagan test is 12.0546, so we can reject the hypothesis of no conditional heteroskedasticity at the 0.05 level. In fact, we can even reject the hypothesis of no conditional
heteroskedasticity at the 0.01 significance level, because the critical value of the test statistic in the case is 6.63. As a result, we conclude that the error term in the Fisher effect regression is conditionally heteroskedastic. The standard errors computed in the original regression are not correct, because they do not account for heteroskedasticity. Therefore, we cannot accept the t-test as valid.
In Example 8, we concluded that a t-test that we might use to test the Fisher effect was not valid. Does that mean that we cannot use a regression model to investigate the Fisher effect? Fortunately, no. A methodology is available to adjust regression coefficients’ standard error to correct for heteroskedasticity. Using an adjusted standard error for , we can reconduct the t-test. As we shall see in the next section, using this valid t-test we will not reject the null hypothesis in Example 8. That is, our statistical conclusion will change after we correct for heteroskedasticity.
4.1.3. Correcting for Heteroskedasticity
Financial analysts need to know how to correct for heteroskedasticity, because such a correction may reverse the conclusions about a particular hypothesis test—and thus affect a particular investment decision. In Example 7, for instance, MacKinlay and Richardson reversed their investment conclusions after correcting their model’s significance tests for heteroskedasticity.
We can use two different methods to correct the effects of conditional heteroskedasticity in linear regression models. The first method, computing robust standard errors, corrects the standard errors of the linear regression model’s estimated coefficients to account for the conditional heteroskedasticity. The second method, generalized least squares, modifies the original equation in an attempt to eliminate the heteroskedasticity. The new, modified regression equation is then estimated under the assumption that heteroskedasticity is no longer a problem.41 The technical details behind these two methods ofcorrecting for conditional heteroskedasticity are outside the scope of this reading.42 Many statistical software packages can easily compute robust standard errors, however, and we recommend using them.43
Returning to the subject of Example 8 concerning the Fisher effect, recall that we concluded that the error variance was heteroskedastic. If we correct the regression coefficients’ standard errors for conditional heteroskedasticity, we get the results shown in Table 7. In comparing the standard errors in Table 7 with those in Table 6, we see that the standard error for the intercept changes very little, but the standard error for the coefficient on predicted inflation (the slope coefficient) increases by about 22 percent (from 0.0761 to 0.0931). Note also that the regression coefficients are the same in both tables, because the results in Table 7 correct only the standard
errors in Table 6.
We can now conduct a valid t-test of the null hypothesis that the slope coefficient has a true value of 1, using the robust standard error for . We find that t = (1.1744 − 1)/0.0931 = 1.8733. This number is smaller than the critical value of 1.97 needed to reject the null hypothesis that the slope equals 1.44 So, we can no longer reject the null hypothesis that the slope equals 1. Thus, in this particular example, correcting for the statistically significant conditional heteroskedasticity had an effect on the result of the hypothesis test about the slope of the predicted inflation coefficient. Example 7 concerning tests of the CAPM is a similar case. In other cases, however, our statistical decision might not change based on using robust standard errors in the t-test.
TABLE 7 Results from Regressing T-Bill Returns on Predicted Inflation (Standard Errors Corrected for Conditional Heteroskedasticity)
Coefficients Standard Error t-Statistic Intercept 0.0116 0.0034 3.4118
Inflation prediction 1.1744 0.0931 12.6144 Residual standard error 0.0233 Multiple R-squared 0.5708
Observations 181
Source: Federal Reserve Bank of Philadelphia, US Department of Commerce.
4.2. Serial Correlation A more common—and potentially more serious—problem than violation of the homoskedasticity assumption is the violation of the assumption that regression errors are uncorrelated across observations. Trying to explain a particular financial relation over a number of periods is risky, because errors in financial regression models are often correlated through time.
When regression errors are correlated across observations, we say that they are serially correlated (or autocorrelated). Serial correlation most typically arises in time-series regressions. In this section, we discuss three aspects of serial correlation: its effect on statistical inference, tests for it, and methods to correct for it.
4.2.1. The Consequences of Serial Correlation
As with heteroskedasticity, the principal problem caused by serial correlation in a
linear regression is an incorrect estimate of the regression coefficient standard errors computed by statistical software packages. As long as none of the independent variables is a lagged value of the dependent variable (a value of the dependent variable from a previous period), then the estimated parameters themselves will be consistent and need not be adjusted for the effects of serial correlation. If, however, one of the independent variables is a lagged value of the dependent variable—for example, if the T-bill return from the previous month was an independent variable in the Fisher effect regression—then serial correlation in the error term will cause all the parameter estimates from linear regression to be inconsistent and they will not be valid estimates of the true parameters.45
In none of the regressions examined in this reading is an independent variable a lagged value of the dependent variable. Thus, in these tregressions, any effect of serial correlation appears in the regression coefficient standard errors. We will examine here the positive serial correlation case, because that case is so common. Positive serial correlation is serial correlation in which a positive error for one observation increases the chance of a positive error for another observation. Positive serial correlation also means that a negative error for one observation increases the chance of a negative error for another observation.46 In examining positive serial correlation, we make the common assumption that serial correlation takes the form of first-order serial correlation, or serial correlation between adjacent observations. In a time-series context, that assumption means the sign of the error term tends to persist from one period to the next.
Although positive serial correlation does not affect the consistency of the estimated regression coefficients, it does affect our ability to conduct valid statistical tests. First, the F-statistic to test for overall significance of the regression may be inflated because the mean squared error (MSE) will tend to underestimate the population error variance. Second, positive serial correlation typically causes the ordinary least squares (OLS) standard errors for the regression coefficients to underestimate the true standard errors. As a consequence, if positive serial correlation is present in the regression, standard linear regression analysis will typically lead us to compute artificially small standard errors for the regression coefficient. These small standard errors will cause the estimated t-statistics to be inflated, suggesting significance where perhaps there is none. The inflated t-statistics may, in turn, lead us to incorrectly reject null hypotheses about population values of the parameters of the regression model more often than we would if the standard errors were correctly estimated. This Type I error could lead to improper investment recommendations.47
4.2.2. Testing for Serial Correlation
We can choose from a variety of tests for serial correlation in a regression model,48 but the most common is based on a statistic developed by Durbin and Watson (1951); in fact, many statistical software packages compute the Durbin– Watson statistic automatically. The equation for the Durbin–Watson test statistic is
(6)
where is the regression residual for period t. We can rewrite this equation as
If the variance of the error is constant through time, then we expect for all t, where we use to represent the estimate of the constant error variance. If, in addition, the errors are also not serially correlated, then we expect . In that case, the Durbin–Watson statistic is approximately equal to
This equation tells us that if the errors are homoskedastic and not serially correlated, then the Durbin–Watson statistic will be close to 2. Therefore, we can test the null hypothesis that the errors are not serially correlated by testing whether the Durbin–Watson statistic differs significantly from 2.
If the sample is very large, the Durbin–Watson statistic will be approximately equal to 2(1 − r), where r is the sample correlation between the regression residuals from one period and those from the previous period. This approximation is useful because it shows the value of the Durbin–Watson statistic for differing levels of serial correlation. The Durbin–Watson statistic can take on values ranging from 0 (in the case of serial correlation of +1) to 4 (in the case of serial correlation of −1):
If the regression has no serial correlation, then the regression residuals will be uncorrelated through time and the value of the Durbin–Watson statistic will be equal to 2(1 − 0) = 2.
If the regression residuals are positively serially correlated, then the Durbin– Watson statistic will be less than 2. For example, if the serial correlation of the errors is 1, then the value of the Durbin–Watson statistic will be 0.
If the regression residuals are negatively serially correlated, then the Durbin– Watson statistic will be greater than 2. For example, if the serial correlation of the errors is −1, then the value of the Durbin–Watson statistic will be 4.
Returning to Example 8, which explored the Fisher effect, as shown in Table 6 the Durbin–Watson statistic for the OLS regression is 0.2980. This result means that the regression residuals are positively serially correlated:
This outcome raises the concern that OLS standard errors may be incorrect because of positive serial correlation. Does the observed Durbin–Watson statistic (0.2980) provide enough evidence to warrant rejecting the null hypothesis of no positive serial correlation?
We should reject the null hypothesis of no serial correlation if the Durbin–Watson statistic is below a critical value, d*. Unfortunately, Durbin and Watson also showed that, for a given sample, we cannot know the true critical value, d*. Instead, we can determine only that d* lies either between two values, du (an upper value) and dl (a lower value), or outside those values. Figure 3 depicts the upper and lower values of d* as they relate to the results of the Durbin–Watson statistic.
From Figure 3, we learn the following:
When the Durbin–Watson (DW) statistic is less than dl, we reject the null hypothesis of no positive serial correlation.
When the DW statistic falls between dl and du, the test results are inconclusive.
When the DW statistic is greater than du, we fail to reject the null hypothesis of no positive serial correlation.49
FIGURE 3 Value of the Durbin–Watson Statistic
Returning to Example 8, the Fisher effect regression has one independent variable and 181 observations. The Durbin–Watson statistic is 0.2980. We can reject the null hypothesis of no correlation in favor of the alternative hypothesis of positive serial correlation at the 0.05 level because the Durbin–Watson statistic is far below dl for k = 1 and n = 100 (1.65). The level of dl would be even higher for a sample of 181 observations. This finding of significant positive serial correlation suggests that the OLS standard errors in this regression probably significantly underestimate the true standard errors.
4.2.3. Correcting for Serial Correlation
We have two alternative remedial steps when a regression has significant serial correlation. First, we can adjust the coefficient standard errors for the linear regression parameter estimates to account for the serial correlation. Second, we can modify the regression equation itself to eliminate the serial correlation. We recommend using the first method for dealing with serial correlation; the second method may result in inconsistent parameter estimates unless implemented with extreme care.
Two of the most prevalent methods for adjusting standard errors were developed by Hansen (1982) and Newey and West (1987). These methods are standard features in many statistical software packages.50 An additional advantage of these methods is that they simultaneously correct for conditional heteroskedasticity.51
Table 8 shows the results of correcting the standard errors from Table 6 for serial correlation and heteroskedasticity using the Newey–West method. Note that the coefficients for both the intercept and the slope are exactly the same as in the original regression. The robust standard errors are now much larger, however— more than twice the OLS standard errors in Table 6. Because of the severe serial correlation in the regression error, OLS greatly underestimates the uncertainty about the estimated parameters in the regression.
Note also that the serial correlation has not been eliminated, but the standard error has been corrected to account for the serial correlation.
Now suppose we want to test our original null hypothesis (the Fisher effect) that the coefficient on the predicted inflation term equals 1 (H0: b1 = 1) against the alternative that the coefficient on the inflation term is not equal to 1 (Ha: b1 | 1). With the corrected standard errors, the value of the test statistic for this null hypothesis is
TABLE 8 Results from Regressing T-Bill Returns on Predicted Inflation (Standard Errors Corrected for Conditional Heteroskedasticity and Serial Correlation)
Coefficient Standard Error t-Statistic Intercept 0.0116 0.0067 1.7313
Inflation prediction 1.1744 0.1751 6.7070 Residual standard error 0.0233 Multiple R-squared 0.5708
Observations 181
Source: Federal Reserve Bank of Philadelphia, US Department of Commerce.
The critical values for both the 0.05 and 0.01 significance level are much larger than 0.996 (the t-test statistic), so we cannot reject the null hypothesis. This conclusion is the same as that reached in Example 7 where the correction was only for heteroskedasticity; but it is the opposite of the conclusion in Example 6 where there were no correlations.
This shows that for some hypotheses, serial correlation and conditional heteroskedasticity could have a big effect on whether we accept or reject those hypotheses.52
4.3. Multicollinearity The second assumption of the multiple linear regression model is that no exact linear relationship exists between two or more of the independent variables. When one of the independent variables is an exact linear combination of other independent variables, it becomes mechanically impossible to estimate the regression. That case, known as perfect collinearity, is much less of a practical concern than multicollinearity.53 Multicollinearity occurs when two or more independent variables (or combinations of independent variables) are highly (but not perfectly) correlated with each other. With multicollinearity we can estimate the regression, but the interpretation of the regression output becomes problematic.
Multicollinearity is a serious practical concern because approximate linear relationships among financial variables are common.
4.3.1. The Consequences of Multicollinearity
Although the presence of multicollinearity does not affect the consistency of the OLS estimates of the regression coefficients, the estimates become extremely imprecise and unreliable. Furthermore, it becomes practically impossible to distinguish the individual impacts of the independent variables on the dependent variable. These consequences are reflected in inflated OLS standard errors for the regression coefficients. With inflated standard errors, t-tests on the coefficients have little power (ability to reject the null hypothesis).
4.3.2. Detecting Multicollinearity
In contrast to the cases of heteroskedasticity and serial correlation, we shall not provide a formal statistical test for multicollinearity. In practice, multicollinearity is often a matter of degree rather than of absence or presence.54
The analyst should be aware that using the magnitude of pairwise correlations among the independent variables to assess multicollinearity, as has occasionally been suggested, is generally not adequate. Although very high pairwise correlations among independent variables can indicate multicollinearity, it is not necessary for such pairwise correlations to be high for there to be a problem of multicollinearity.55 Stated another way, high pairwise correlations among the independent variables are not a necessary condition for multicollinearity, and low pairwise correlations do not mean that multicollinearity is not a problem. The only case in which correlation between independent variables may be a reasonable indicator of multicollinearity occurs in a regression with exactly two independent variables.
The classic symptom of multicollinearity is a high R2 (and significant F-statistic) even though the t-statistics on the estimated slope coefficients are not significant. The insignificant t-statistics reflect inflated standard errors. Although the coefficients might be estimated with great imprecision, as reflected in low t- statistics, the independent variables as a group may do a good job of explaining the dependent variable, and a high R2 would reflect this effectiveness. Example 9 illustrates this diagnostic.
EXAMPLE 9 Multicollinearity in Explaining Returns to the Fidelity Select Technology Portfolio
In Example 3 we regressed returns to the Fidelity Select Technology Portfolio (FSPTX) on returns to the S&P 500 Growth Index and the S&P 500 Value Index. Table 9 shows the results of our regression, which uses data from January 2009 through December 2013. The t-statistic of 5.9286 on the growth index return is greater than 2, indicating that the coefficient on the growth index differs significantly from 0 at standard significance levels. On the other hand, the t-statistic on the value index return is −0.9012 and thus is not statistically significant. This result suggests that the returns to the FSPTX are linked to the returns to the growth index and not closely associated with the returns to the value index. The coefficient on the growth index, however, is 1.4697. This result implies that returns on the FSPTX are more volatile than are returns on the growth index.
TABLE 9 Results from Regressing the FSPTX Returns on the S&P 500 Growth and Value Indices
Coefficient Standard Error t-Statistic Intercept 0.0018 0.0038 0.4737
S&P 500 Growth Index 1.4697 0.2479 5.9286 S&P 500 Value Index −0.1833 0.2034 −0.9012
ANOVA df SS MSS F Significance F Regression 2 0.1653 0.0826 113.7285 1.27E-20 Residual 57 0.0414 0.0007 Total 59 0.2067
Residual standard error 0.0270 Multiple R-squared 0.8084
Observations 60
Source: Bloomberg, finance.yahoo.com.
Note also that this regression explains a significant amount of the variation in the returns to the FSPTX. Specifically, the R2 from this regression is 0.7996. Thus approximately 80 percent of the variation in the returns to the FSPTX is explained by returns to the S&P 500 Growth and S&P 500 Value indices.
Now suppose we run another linear regression that adds returns to the S&P 500 itself to the returns to the S&P 500 Growth and S&P 500 Value indices. The S&P 500 includes the component stocks of these two style indices, so we are
introducing a severe multicollinearity problem.
Table 10 shows the results of that regression. Note that the R2 in this regression has changed almost imperceptibly from the R2 in the previous regression (increasing from 0.7996 to 0.8084), but now the standard errors of the coefficients of the independent variables are much larger. Adding the return to the S&P 500 to the previous regression does not explain any more of the variance in the returns to the FSPTX than the previous regression did, but now none of the coefficients is statistically significant. This is the classic case of multicollinearity mentioned in the reading.
TABLE 10 Results from Regressing the FSPTX Returns on Returns to the S&P 500 Growth and S&P 500 Value Indices and the S&P 500 Index
Coefficient Standard Error t-Statistic Intercept 0.0008 0.0038 0.2105
S&P 500 Growth Index 14.2444 7.9783 1.7854 S&P 500 Value Index 11.6955 7.4180 1.5766
S&P 500 Index –24.6734 15.4022 –0.6019
ANOVA df SS MSS F Significance F Regression 3 0.1671 0.0557 78.7577 4.14E-20 Residual 56 0.0396 0.0007 Total 59 0.2067
Residual standard error 0.0266 Multiple R-squared 0.8084
Observations 60
Source: Bloomberg, finance.yahoo.com, S&P Dow Jones Indices.
Multicollinearity may be a problem even when we do not observe the classic symptom of insignificant t-statistics but a highly significant F-test. Advanced textbooks provide further tools to help diagnose multicollinearity.56
4.3.3. Correcting for Multicollinearity
The most direct solution to multicollinearity is excluding one or more of the regression variables. In the example above, we can see that the S&P 500 total returns should not be included if both the S&P 500 Growth and S&P 500 Value indices are
included, because the returns to the entire S&P 500 Index are a weighted average of the return to growth stocks and value stocks. In many cases, however, no easy solution is available to the problem of multicollinearity, and you will need to experiment with including or excluding different independent variables to determine the source of multicollinearity.
TABLE 11 Problems in Linear Regression and Their Solutions
Problem Effect Solution
Heteroskedasticity Incorrect standard errors
Use robust standard errors (corrected for
conditional heteroskedasticity)
Serial correlation
Incorrect standard errors (additional problems if a lagged value of the dependent variable is used as an
independent variable)
Use robust standard errors (corrected for serial correlation)
Multicollinearity High R2 and low t-statistics
Remove one or more independent variables; often no solution based
in theory
4.4. Heteroskedasticity, Serial Correlation, Multicollinearity: Summarizing the Issues We have discussed some of the problems that heteroskedasticity, serial correlation, and multicollinearity may cause in interpreting regression results. These violations of regression assumptions, we have noted, all lead to problems in making valid inferences. The analyst should check that model assumptions are fulfilled before interpreting statistical tests.
Table 11 gives a summary of these problems, the effect they have on the linear regression results (an analyst can see these effects using regression software), and the solutions to these problems.
5. Model Specification and Errors in Specification Until now, we have assumed that whatever regression model we estimate is correctly specified. Model specification refers to the set of variables included in the regression and the regression equation’s functional form. In the following, we first give some broad guidelines for correctly specifying a regression. Then we turn to three types of model misspecification: misspecified functional form, regressors that are correlated with the error term, and additional time-series misspecification. Each of these types of misspecification invalidates statistical inference using OLS; most of these misspecifications will cause the estimated regression coefficients to be inconsistent.
5.1. Principles of Model Specification In discussing the principles of model specification, we need to acknowledge that there are competing philosophies about how to approach model specification. Furthermore, our purpose for using regression analysis may affect the specification we choose. The following principles have fairly broad application, however.
The model should be grounded in cogent economic reasoning. We should be able to supply the economic reasoning behind the choice of variables, and the reasoning should make sense. When this condition is fulfilled, we increase the chance that the model will have predictive value with new data. This approach contrasts to the variable-selection process known as data mining. With data mining, the investigator essentially develops a model that maximally exploits the characteristics of a specific dataset.
The functional form chosen for the variables in the regression should be appropriate given the nature of the variables. As one illustration, consider studying mutual fund market timing based on fund and market returns alone. One might reason that for a successful timer, a plot of mutual fund returns against market returns would show curvature, because a successful timer would tend to increase (decrease) beta when market returns were high (low). The model specification should reflect the expected nonlinear relationship.57 In other cases, we may transform the data such that a regression assumption is better satisfied.
The model should be parsimonious. In this context, “parsimonious” means accomplishing a lot with a little. We should expect each variable included in a regression to play an essential role.
The model should be examined for violations of regression assumptions before being accepted. We have already discussed detecting the presence of heteroskedasticity, serial correlation, and multicollinearity. As a result of such diagnostics, we may conclude that we need to revise the set of included variables and/or their functional form.
The model should be tested and be found useful out of sample before being accepted. The term “out of sample” refers to observations outside the dataset on which the model was estimated. A plausible model may not perform well out of sample because economic relationships have changed since the sample period. That possibility is itself useful to know. A second explanation, however, may be that relationships have not changed but that the model explains only a specific dataset.
Having given some broad guidance on model specification, we turn to a discussion of specific model specification errors. Understanding these errors will help an analyst develop better models and be a more informed consumer of investment research.
5.2. Misspecified Functional Form Whenever we estimate a regression, we must assume that the regression has the correct functional form. This assumption can fail in several ways:
One or more important variables could be omitted from regression.
One or more of the regression variables may need to be transformed (for example, by taking the natural logarithm of the variable) before estimating the regression.
The regression model pools data from different samples that should not be pooled.
First, consider the effects of omitting an important independent variable from a regression (omitted variable bias). If the true regression model was
(7)
but we estimate the model58
then our regression model would be misspecified. What is wrong with the model?
If the omitted variable (X2) is correlated with the remaining variable (X1), then the error term in the model will be correlated with (X1), and the estimated values of the regression coefficients a0 and a1 would be biased and inconsistent. In addition, the estimates of the standard errors of those coefficients will also be inconsistent, so we can use neither the coefficients estimates nor the estimated standard errors to make statistical tests.
EXAMPLE 10 Omitted Variable Bias and the Bid–Ask Spread
In this example, we extend our examination of the bid–ask spread to show the effect of omitting an important variable from a regression. In Example 1, we showed that the natural logarithm of the ratio [(Bid–ask spread)/Price] was significantly related to both the natural logarithm of the number of market makers and the natural logarithm of the market capitalization of the company. We repeat Table 1 from Example 1 below.
If we did not include the natural log of market capitalization as an independent variable in the regression, and we regressed the natural logarithm of the ratio [(Bid–ask spread)/Price] only on the natural logarithm of the number of market makers for the stock, the results would be as shown in Table 12.
TABLE 12 Results from Regressing ln(Bid–Ask Spread/Price) on ln(Number of Market Makers)
Coefficients Standard Error t-Statistic Intercept 5.0707 0.2009 25.2399
ln(Number of NASDAQ market makers) −3.1027 0.0561 −55.3066
ANOVA df SS MSS F Significance F Regression 1 3,200.3918 3,200.3918 3,063.3655 0.00 Residual 2,585 2,700.6287 1.0447 Total 2,586 5,901.0204
Residual standard error 1.0221 Multiple R-squared 0.5423
Observations 2,587
Source: Center for Research in Security Prices, University of Chicago.
Note that the coefficient on ln(Number of NASDAQ market makers) changed from −1.5186 in the original (correctly specified) regression to −3.1027 in the misspecified regression. Also, the intercept changed from 1.5949 in the correctly specified regression to 5.0707 in the misspecified regression. These results illustrate that omitting an independent variable that should be in the regression can cause the remaining regression coefficients to be inconsistent.
A second common cause of misspecification in regression models is the use of the wrong form of the data in a regression, when a transformed version of the data is appropriate. For example, sometimes analysts fail to account for curvature or nonlinearity in the relationship between the dependent variable and one or more of the independent variables, instead specifying a linear relation among variables. When we are specifying a regression model, we should consider whether economic theory suggests a nonlinear relation. We can often confirm the nonlinearity by plotting the data, as we will illustrate in Example 11 below. If the relationship between the variables becomes linear when one or more of the variables is represented as a proportional change in the variable, we may be able to correct the misspecification by taking the natural logarithm of the variable(s) we want to represent as a proportional change. Other times, analysts use unscaled data in regressions, when scaled data (such as dividing net income or cash flow by sales) are more appropriate. In Example 1, we scaled the bid–ask spread by stock price because what a given bid–ask spread means in terms of transactions costs for a given size investment depends on the price of the stock; if we had not scaled the bid– ask spread, the regression would have been misspecified.
EXAMPLE 11 Nonlinearity and the Bid–Ask Spread
In Example 1, we showed that the natural logarithm of the ratio [(Bid–ask spread)/Price] was significantly related to both the natural logarithm of the number of market makers and the natural logarithm of the company’s market capitalization. But why did we take the natural logarithm of each of the variables in the regression? We began a discussion of this question in Example 1, which we continue now.
What does theory suggest about the nature of the relationship between the ratio (Bid–ask spread)/Price, or the percentage bid–ask spread, and its determinants (the independent variables)? Stoll (1978) builds a theoretical model of the determinants of percentage bid–ask spread in a dealer market. In his model, the
determinants enter multiplicatively in a particular fashion. In terms of the independent variables introduced in Example 1, the functional form assumed is
where c is a constant. The relationship of the percentage bid–ask spread with the number of market makers and market capitalization is not linear in the original variables.59 If we take the natural log of both sides of the above model, however, we have a log-log regression that is linear in the transformed variables:60
where
Yi = the natural logarithm of the ratio (Bid–ask spread)/Price for stock i
b0 = a constant that equals ln(c)
X1i = the natural logarithm of the number of market makers for stock i
X2i = the natural logarithm of the market capitalization of company i
εi = the error term
As mentioned in Example 1, a slope coefficient in the log-log model is interpreted as an elasticity, precisely, the partial elasticity of the dependent variable with respect to the independent variable (“partial” means holding the other independent variables constant).
We can plot the data to assess whether the variables are linearly related after the logarithmic transformation. For example Figure 4 shows a scatterplot of the natural logarithm of the number of market makers for a stock (on the X axis) and the natural logarithm of (Bid–ask spread)/Price (on the Y axis), as well as a regression line showing the linear relation between the two transformed variables. The relation between the two transformed variables is clearly linear.
FIGURE 4 Linear Regression When Two Variables Have a Linear Relation
FIGURE 5 Linear Regression When Two Variables Have a Nonlinear Relation
If we do not take log of the ratio (Bid–ask spread)/Price, the plot is not linear. Figure 5 shows a plot of the natural logarithm of the number of market makers for a stock (on the X axis) and the ratio (Bid–ask spread)/Price expressed as a percentage (on the Y axis), as well as a regression line that attempts to show a
linear relation between the two variables. We see that the relation between the two variables is very nonlinear.61 Consequently, we should not estimate a regression with (Bid–ask spread)/Price as the dependent variable. Consideration of the need to ensure that predicted bid–ask spreads are positive would also lead us to not use (Bid–ask spread)/Price as the dependent variable. If we use the non-transformed ratio (Bid–ask spread)/Price as the dependent variable, the estimated model could predict negative values of the bid–ask spread. This result would be nonsensical; in reality, no bid–ask spread is negative (it is hard to motivate traders to simultaneously buy high and sell low), so a model that predicts negative bid–ask spreads is certainly misspecified.62 We illustrate the problem of negative values of the predicted bid–ask spreads now.
Table 13 shows the results of a regression with (Bid–ask spread)/ Price as the dependent variable and the natural logarithm of the number of market makers and the natural logarithm of the company’s market capitalization as the independent variables.
1. Suppose that for a particular NASDAQ-listed stock, the number of market
makers is 50 and the market capitalization is $6 billion. What is the predicted ratio of bid–ask spread to price for this stock based on the above model?
Solution to 1: The natural log of the number of market makers equals ln 50 = 3.9120 and the natural log of the stock’s market capitalization (in millions) is ln 6,000 = 8.6995. In this case, the predicted ratio of bid–ask spread to price is 0.0674 + (−0.0142 × 3.9120) + (−0.0016 × 8.6995) = −0.0021. Therefore, the model predicts that the ratio of bid–ask spread to stock price is −0.0021 or −0.21 percent of the stock price.
TABLE 13 Results from Regressing Bid–Ask Spread/Price on ln(Number of Market Makers) and ln(Market Cap)
Coefficients StandardError t-
Statistic Intercept 0.0674 0.0035 19.2571
ln(Number of NASDAQ market makers) −0.0142 0.0012 −11.8333
ln(Company’s market cap) −0.0016 0.0002 −8.0000
ANOVA df SS MSS F
Significance F
Regression 2 0.1539 0.0770 392.3338 0.00 Residual 2,584 0.5068 0.0002 Total 2,586 0.6607
Residual standard error 0.0140 Multiple R-squared 0.2329
Observations 2,587
Source: Center for Research in Security Prices, University of Chicago.
2. Does the predicted bid–ask spread for the above stock make sense? If not, how could this problem be avoided?
Solution to 2: The predicted bid–ask spread is negative, which does not make economic sense. This problem could be avoided by using log of (Bid–ask spread)/Price as the dependent variable.63
Often, analysts must decide whether to scale variables before they compare data across companies. For example, in financial statement analysis, analysts often compare companies using common size statements. In a common size income statement, all the line items in a company’s income statement are divided by the company’s revenues.64 Common size statements make comparability across companies much easier. An analyst can use common size statements to quickly compare trends in gross margins (or other income statement variables) for a group of companies.
Issues of comparability also appear for analysts who want to use regression analysis to compare the performance of a group of companies. Example 12 illustrates this issue.
EXAMPLE 12 Scaling and the Relation between Cash Flow from Operations and Free Cash Flow
Suppose an analyst wants to explain free cash flow to the firm as a function of cash flow from operations in 2001 for 11 family clothing stores in the United States with market capitalizations of more than $100 million as of the end of 2001.
To investigate this issue, the analyst might use free cash flow as the dependent variable and cash flow from operations as the independent variable in single-
independent-variable linear regression. Table 14 shows the results of that regression. Note that the t-statistic for the slope coefficient for cash flow from operations is quite high (6.5288), the significance level for the F-statistic for the regression is very low (0.0001), and the R-squared is quite high. We might be tempted to believe that this regression is a success and that for a family clothing store, if cash flow from operations increased by $1.00, we could confidently predict that free cash flow to the firm would increase by $0.3579.
TABLE 14 Results from Regressing the Free Cash Flow on Cash Flow from Operations for Family Clothing Stores
Coefficients Standard Error t-Statistic Intercept 0.7295 27.7302 0.0263
Cash flow from operations 0.3579 0.0548 6.5288
ANOVA df SS MSS F Significance F Regression 1 245,093.7836 245,093.7836 42.6247 0.0001 Residual 9 51,750.3139 5,750.0349 Total 10 296,844.0975
Residual standard error 75.8290 Multiple R-squared 0.8257
Observations 11
Source: Compustat.
But is this specification correct? The regression does not account for size differences among the companies in the sample.
We can account for size differences by using common size cash flow results across companies. We scale the variables by dividing cash flow from operations and free cash flow to the firm by the company’s sales before using regression analysis. We will use (Free cash flow to the firm/Sales) as the dependent variable and (Cash flow from operations/Sales) as the independent variable. Table 15 shows the results of this regression. Note that the t-statistic for the slope coefficient on (Cash flow from operations/Sales) is 1.6262, so it is not significant at the 0.05 level. Note also that the significance level of the F- statistic is 0.1383, so we cannot reject at the 0.05 level the hypothesis that the regression does not explain variation in (Free cash flow/Sales) among family
clothing stores. Finally, note that the R-squared in this regression is much lower than that of the previous regression.
TABLE 15 Results from Regressing the Free Cash Flow/Sales on Cash Flow from Operations/Sales for Family Clothing Stores
Coefficient Standard Error t-Statistic Intercept −0.0121 0.0221 −0.5497
Cash flow from operations/Sales 0.4749 0.2920 1.6262
ANOVA df SS MSS F Significance F Regression 1 0.0030 0.0030 2.6447 0.1383 Residual 9 0.0102 0.0011 Total 10 0.0131
Residual standard error 0.0336 Multiple R-squared 0.2271
Observations 11
Source: Compustat.
Which regression makes more sense? Usually, the scaled regression makes more sense. We want to know what happens to free cash flow (as a fraction of sales) if a change occurs in cash flow from operations (as a fraction of sales). Without scaling, the results of the regression can be based solely on scale differences across companies, rather than based on the companies’ underlying economics.
A third common form of misspecification in regression models is pooling data from different samples that should not be pooled. This type of misspecification can best be illustrated graphically. Figure 6 shows two clusters of data on variables X and Y, with a fitted regression line. The data could represent the relationship between two financial variables at two different time periods, for example.
In each cluster of data on X and Y, the correlation between the two variables is virtually 0. Because the means of both X and Y are different for the two clusters of data in the combined sample, X and Y are highly correlated. The correlation is spurious (misleading), however, because it reflects differences in the relationship between X and Y during two different time periods.
5.3. Time-Series Misspecification (Independent Variables Correlated with Errors) In the previous section, we discussed the misspecification that arises when a relevant independent variable is omitted from a regression. In this section, we discuss problems that arise from the kinds of variables included in the regression, particularly in a time-series context. In models that use time-series data to explain the relations among different variables, it is particularly easy to violate Regression Assumption 3, that the error term has mean 0, conditioned on the independent variables. If this assumption is violated, the estimated regression coefficients will be biased and inconsistent.
FIGURE 6 Plot of Two Series with Changing Means
Three common problems that create this type of time-series misspecification are:
including lagged dependent variables as independent variables in regressions with serially correlated errors;
including a function of a dependent variable as an independent variable, sometimes as a result of the incorrect dating of variables; and
independent variables that are measured with error.
The next examples demonstrate these problems.
Suppose that an analyst includes the first lagged value of the dependent variable in a multiple regression that, as a result, has significant serial correlation in the errors. For example, the analyst might use the regression equation
(8)
Because we assume that the error term is serially correlated, by definition the error term is correlated with the dependent variable. Consequently, the lagged dependent variable, Yt−1, will be correlated with the error term, violating the assumption that the independent variables are uncorrelated with the error term. As a result, the estimates of the regression coefficients will be biased and inconsistent.
EXAMPLE 13 Fisher Effect with a Lagged Dependent Variable
In our discussion of serial correlation, we concluded from a test using the Durbin–Watson test that the error term in the Fisher effect equation (Equation 5) showed positive (first-order) serial correlation, using three-month T-bill returns as the dependent variable and inflation expectations of professional forecasters as the independent variable. Observations on the dependent and independent variables were quarterly. Table 16 modifies that regression by including the previous quarter ’s three-month T-bill returns as an additional independent variable.
TABLE 16 Results from Regressing T-Bill Returns on Predicted Inflation and Lagged T-Bill Returns
Coefficient Standard Error t-Statistic Intercept –0.0005 0.0014 –0.3571
Inflation prediction 0.1843 0.0455 4.0505 Lagged T-bill return 0.8796 0.0295 29.8169
Residual standard error 0.0095 Multiple R-squared 0.9285
Observations 181
Source: Federal Reserve Bank of Philadelphia, US Department of Commerce.
At first glance, these regression results look very interesting—the coefficient on the lagged T-bill return appears to be highly significant. But on closer consideration, we must ignore these regression results, because the regression
is fundamentally misspecified. As long as the error term is serially correlated, including lagged T-bill returns as an independent variable in the regression will cause all the coefficient estimates to be biased and inconsistent. Therefore, this regression is not usable for either testing a hypothesis or for forecasting.
A second common time-series misspecification in investment analysis is to forecast the past. What does that mean? If we forecast the future (say we predict at time t the value of variable Y in period t + 1), we must base our predictions on information we knew at time t. We could use a regression to make that forecast using the equation
(9)
In this equation, we predict the value of Y in time t + 1 using the value of X in time t. The error term, εt+1, is unknown at time t and thus should be uncorrelated with X1t.
Unfortunately, analysts sometimes use regressions that try to forecast the value of a dependent variable at time t + 1 based on independent variable(s) that are functions of the value of the dependent variable at time t + 1. In such a model, the independent variable(s) would be correlated with the error term, so the equation would be misspecified. As an example, an analyst may try to explain the cross-sectional returns for a group of companies during a particular year using the market-to-book ratio and the market capitalization for those companies at the end of the year.65 If the analyst believes that such a regression predicts whether companies with high market-to-book ratios or high market capitalizations will have high returns, the analyst is mistaken. This is because for any given period, the higher the return during the period, the higher the market capitalization and the market-to-book period will be at the end of the period. So in this case, if all the cross-sectional data come from period t + 1, a high value of the dependent variable (returns) actually causes a high value of the independent variables (market capitalization and the market-to-book ratio), rather than the other way around. In this type of misspecification, the regression model effectively includes the dependent variable on both the right-and left-hand sides of the regression equation.
The third common time-series misspecification arises when an independent variable is measured with error. Suppose a financial theory tells us that a particular variable Xt, such as expected inflation, should be included in the regression model. But we cannot directly observe Xt; instead, we can observe actual inflation, Zt = Xt + ut, where we assume ut is an error term that is uncorrelated with Xt. Even in this best of circumstances, using Zt in the regression instead of Xt will cause the regression coefficient estimates to be biased and inconsistent. To see why, assume we want to estimate the regression
but we substitute Zt for Xt. Then we would estimate
But Zt = Xt + ut, Zt is correlated with the error term (–b1ut + εt). Therefore, our estimated model violates the assumption that the error term is uncorrelated with the independent variable. Consequently, the estimated regression coefficients will be biased and inconsistent.
EXAMPLE 14 The Fisher Effect with Measurement Error
Recall from Example 8 on the Fisher effect that based on our initial analysis in which we did not correct for heteroskedasticity and serial correlation, we rejected the hypothesis that three-month T-bill returns moved one-for-one with expected inflation.
What if we used actual inflation instead of expected inflation as the independent variable? Note first that
where
π = actual rate of inflation
πe = expected rate of inflation
v = the difference between actual and expected inflation
Because actual inflation measures expected inflation with error, the estimators of the regression coefficients using T-bill yields as the dependent variable and actual inflation as the independent variable will not be consistent.66
TABLE 6 Results from Regressing T-Bill Returns on Predicted Inflation (repeated)
Coefficient Standard Error t-Statistic Intercept 0.0116 0.0033 3.5152
Inflation prediction 1.1744 0.0761 15.4323 Residual standard error 0.0223 Multiple R-squared 0.5708
Observations 181 Durbin–Watson statistic 0.2980
Source: Federal Reserve Bank of Philadelphia, US Department of Commerce.
TABLE 17 Results from Regressing T-Bill Returns on Actual Inflation
Coefficient Standard Error t-Statistic Intercept 0.0227 0.0034 6.6765
Actual inflation 0.8946 0.0761 11.7556 Residual standard error 0.0267 Multiple R-squared 0.4356
Observations 181
Source: Federal Reserve Bank of Philadelphia, US Department of Commerce.
Table 17 shows the results of using actual inflation as the independent variable. The estimates in this table are quite different from those presented in the previous table. Note that the slope coefficient on actual inflation is much lower than the slope coefficient on predicted inflation in the previous regression. This result is an illustration of a general proposition: In a single-independent- variable regression, if we select a version of that independent variable that is measured with error, the estimated slope coefficient on that variable will be biased toward 0.67
5.4. Other Types of Time-Series Misspecification By far the most frequent source of misspecification in linear regressions that use time series from two or more different variables is nonstationarity. Very roughly, nonstationarity means that a variable’s properties, such as mean and variance, are not constant through time. We will postpone our discussion about stationarity to the reading on time-series analysis, but we can list some examples in which we need to use stationarity tests before we use regression statistical inference.68
Relations among time series with trends (for example, the relation between consumption and GDP).
Relations among time series that may be random walks (time series for which
the best predictor of next period’s value is this period’s value). Exchange rates are often random walks.
The time-series examples in this reading were carefully chosen such that nonstationarity was unlikely to be an issue for any of them. But nonstationarity can be a very severe problem for analyzing the relations among two or more time series in practice. Analysts must understand these issues before they apply linear regression to analyzing the relations among time series. Otherwise, they may rely on invalid statistical inference.
6. Models with Qualitative Dependent Variables Financial analysts often need to be able to explain the outcomes of a qualitative dependent variable. Qualitative dependent variables are dummy variables used as dependent variables instead of as independent variables.
For example, to predict whether or not a company will go bankrupt, we need to use a qualitative dependent variable (bankrupt or not) as the dependent variable and use data on the company’s financial performance (e.g., return on equity, debt-to-equity ratio, or debt rating) as independent variables. Unfortunately, linear regression is not the best statistical method to use for estimating such a model. If we use the qualitative dependent variable bankrupt (1) or not bankrupt (0) as the dependent variable in a regression with financial variables as the independent variables, the predicted value of the dependent variable could be much greater than 1 or much lower than 0. Of course, these results would be invalid. The probability of bankruptcy (or of anything, for that matter) cannot be greater than 100 percent or less than 0 percent. Instead of a linear regression model, we should use probit, logit, or discriminant analysis for this kind of estimation.
Probit and logit models estimate the probability of a discrete outcome given the values of the independent variables used to explain that outcome. The probit model, which is based on the normal distribution, estimates the probability that Y = 1 (a condition is fulfilled) given the value of the independent variable X. The logit model is identical, except that it is based on the logistic distribution rather than the normal distribution.69 Both models must be estimated using maximum likelihood methods.70
Another technique to handle qualitative dependent variables is discriminant analysis. In his Z-score and Zeta analysis, Altman (1968) and Altman, Halderman, and Narayanan (1977) reported on the results of discriminant analysis. Altman uses financial ratios to predict the qualitative dependent variable bankruptcy. Discriminant analysis yields a linear function, similar to a regression equation, which can then be used to create an overall score. Based on the score, an observation can be classified into the bankrupt or not bankrupt category.
Qualitative dependent variable models can be useful not only for portfolio management but also for business management. For example, we might want to predict whether a client is likely to continue investing in a company or to withdraw assets from the company. We might also want to explain how particular demographic characteristics might affect the probability that a potential investor will sign on as a new client, or evaluate the effectiveness of a particular direct-mail advertising campaign based on the demographic characteristics of the target
audience. These issues can be analyzed with either probit or logit models.
EXAMPLE 15 Explaining Analyst Coverage
Suppose we want to investigate what factors determine whether at least one analyst covers a company. We can employ a probit model to address the question. The sample consists of 4,619 observations on public companies in 2013.
The variables in the probit model are as follows:
ANALYSTS = the discrete dependent variable, which takes on a value of 1 if at least one analyst covers the company and a value of 0 if no analysts cover the company
LNVOLUME = the natural log of the company’s trading volume in the last month of the year
LNMV = the natural log of the market value of the company’s equity
MATURITY = the mix of the company’s earned and contributed capital, i.e., retained earnings as a proportion of total equity (RE/TE)
DIVPAYER = a dummy independent variable that takes on a value of 1 if the company paid a dividend
In this attempt to explain analyst coverage, we are examining whether more liquid companies, as captured by the trading volume in their shares, and larger companies, as captured by their market values, are more likely to be followed by at least one analyst. We also examine whether more mature and well established firms, as reflected in their mix of earned and contributed capital and their status as dividend payers, are more likely to be followed by an analyst. Table 18 shows the results of the probit estimation.
As Table 18 shows, two coefficients (besides the intercept) have t-statistics with an absolute value greater than 2.0. The coefficient on LNMV has a t-statistic of 13.7910. That value is far above the critical value at the 0.05 level for the t- statistic (1.96), so we can reject at the 0.05 level of significance the null hypothesis that the coefficient on LNMV equals 0, in favor of the alternative hypothesis that the coefficient is not equal to 0. The second coefficient with an absolute value greater than 2 is DIVPAYER, which has a t-statistic of 5.9786. We can also reject at the 0.05 level of significance the null hypothesis that the coefficient on DIVPAYER is equal to 0, in favor of the alternative hypothesis
that the coefficient is not equal to 0.
Neither of the two remaining independent variables is statistically significant at the 0.05 level in this probit analysis. That is, neither one reaches the critical value of 1.96 needed to reject the null hypothesis (that the associated coefficient is significantly different from 0). This result shows that once we take into account a company’s market value and whether it pays dividends, the other factors—trading volume and maturity—have no power to explain whether at least one analyst will cover the company.
TABLE 18 Explaining Analyst Coverage Using a Probit Model
Coefficient Standard Error t-Statistic Intercept 2.5066 0.1005 24.9413
LNVOLUME –0.0221 0.0173 –1.2775 LNMV 0.2441 0.0177 13.7910
MATURITY 0.0011 0.0007 1.5714 DIVPAYER 0.2798 0.0468 5.9786
Percent correctly predicted 78.00
Source: I/B/E/S from Thomson Reuters, Center for Research in Security Prices at the University of Chicago, and S&P Capital IQ/Compustat.
7. Summary In this reading, we have presented the multiple linear regression model and discussed violations of regression assumptions, model specification and misspecification, and models with qualitative variables.
The general form of a multiple linear regression model is Yi = b0 + b1X1i + b2X2i + … + bkXki + εi
We conduct hypothesis tests concerning the population values of regression coefficients using t-tests of the form
The lower the p-value reported for a test, the more significant the result.
The assumptions of classical normal multiple linear regression model are as follows:
1. A linear relation exists between the dependent variable and the independent variables.
2. The independent variables are not random. Also, no exact linear relation exists between two or more of the independent variables.
3. The expected value of the error term, conditioned on the independent variables, is 0.
4. The variance of the error term is the same for all observations.
5. The error term is uncorrelated across observations.
6. The error term is normally distributed.
To make a prediction using a multiple linear regression model, we take the following three steps:
1. Obtain estimates of the regression coefficients.
2. Determine the assumed values of the independent variables.
3. Compute the predicted value of the dependent variable.
When predicting the dependent variable using a linear regression model, we
encounter two types of uncertainty: uncertainty in the regression model itself, as reflected in the standard error of estimate, and uncertainty about the estimates of the regression coefficients.
The F-test is reported in an ANOVA table. The F-statistic is used to test whether at least one of the slope coefficients on the independent variables is significantly different from 0.
Under the null hypothesis that all the slope coefficients are jointly equal to 0, this test statistic has a distribution of Fk,n−(k+1), where the regression has n observations and k independent variables. The F-test measures the overall significance of the regression.
R2 is nondecreasing in the number of independent variables, so it is less reliable as a measure of goodness of fit in a regression with more than one independent variable than in a one-independent-variable regression.
Analysts often choose to use adjusted R2 because it does not necessarily increase when one adds an independent variable.
Dummy variables in a regression model can help analysts determine whether a particular qualitative independent variable explains the model’s dependent variable. A dummy variable takes on the value of 0 or 1. If we need to distinguish among n categories, the regression should include n − 1 dummy variables. The intercept of the regression measures the average value of the dependent variable of the omitted category, and the coefficient on each dummy variable measures the average incremental effect of that dummy variable on the dependent variable.
If a regression shows significant conditional heteroskedasticity, the standard errors and test statistics computed by regression programs will be incorrect unless they are adjusted for heteroskedasticity.
One simple test for conditional heteroskedasticity is the Breusch–Pagan test. Breusch and Pagan showed that, under the null hypothesis of no conditional heteroskedasticity, nR2 (from the regression of the squared residuals on the independent variables from the original regression) will be a χ2 random variable with the number of degrees of freedom equal to the number of independent variables in the regression.
The principal effect of serial correlation in a linear regression is that the standard errors and test statistics computed by regression programs will be
incorrect unless adjusted for serial correlation. Positive serial correlation typically inflates the t-statistics of estimated regression coefficients as well as the F-statistic for the overall significance of the regression.
The most commonly used test for serial correlation is based on the Durbin– Watson statistic. If the Durbin–Watson statistic differs sufficiently from 2, then the regression errors have significant serial correlation.
Multicollinearity occurs when two or more independent variables (or combinations of independent variables) are highly (but not perfectly) correlated with each other. With multicollinearity, the regression coefficients may not be individually statistically significant even when the overall regression is significant as judged by the F-statistic.
Model specification refers to the set of variables included in the regression and the regression equation’s functional form. The following principles can guide model specification:
The model should be grounded in cogent economic reasoning.
The functional form chosen for the variables in the regression should be appropriate given the nature of the variables.
The model should be parsimonious.
The model should be examined for violations of regression assumptions before being accepted.
The model should be tested and be found useful out of sample before being accepted.
If a regression is misspecified, then statistical inference using OLS is invalid and the estimated regression coefficients may be inconsistent.
Assuming that a model has the correct functional form, when in fact it does not, is one example of misspecification. There are several ways this assumption may be violated:
One or more important variables could be omitted from the regression.
One or more of the regression variables may need to be transformed before estimating the regression.
The regression model pools data from different samples that should not be pooled.
Another type of misspecification occurs when independent variables are correlated with the error term. This is a violation of Regression Assumption 3,
that the error term has a mean of 0, and causes the estimated regression coefficients to be biased and inconsistent. Three common problems that create this type of time-series misspecification are:
including lagged dependent variables as independent variables in regressions with serially correlated errors;
including a function of dependent variable as an independent variable, sometimes as a result of the incorrect dating of variables; and
independent variables that are measured with error.
Probit and logit models estimate the probability of a discrete outcome (the value of a qualitative dependent variable, such as whether a company enters bankruptcy) given the values of the independent variables used to explain that outcome. The probit model, which is based on the normal distribution, estimates the probability that Y = 1 (a condition is fulfilled) given the values of the independent variables. The logit model is identical, except that it is based on the logistic distribution rather than the normal distribution.
References
Altman, Edward I. 1968. “Financial Ratios, Discriminant Analysis and the Prediction of Corporate Bankruptcy.” Journal of Finance, vol. 23: 589–609.
Altman, Edward I., R. Halderman, and P. Narayanan. 1977. “Zeta Analysis: A New Model to Identify Bankruptcy Risk of Corporations.” Journal of Banking & Finance, vol. 1: 29–54.
Bhabra, Harjeet S., and Jiayin Huang. 2013. “An Empirical Investigation of Mergers and Acquisitions by Chinese Listed Companies, 1997–2007.” Journal of Multinational Financial Management, vol. 23: 186–207.
Bodie, Zvi, Alex Kane, and Alan J. Marcus. 2014. Investments, 10th edition. New York: McGraw-Hill Irwin.
Breusch, T., and A. Pagan. 1979. “A Simple Test for Heteroscedasticity and Random Coefficient Variation.” Econometrica, vol. 47: 1287–1294.
Buetow, Gerald W., Robert R. Johnson, and David E. Runkle. 2000. “The Inconsistency of Return-Based Style Analysis.” Journal of Portfolio Management, vol. 26, no. 3: 61–77.
Durbin, J., and G.S. Watson. 1951. “Testing for Serial Correlation in Least Squares Regression, II.” Biometrika, vol. 38: 159–178.
Goldberger, Arthur S. 1998. Introductory Econometrics. Cambridge, MA: Harvard University Press.
Greene, William H. 2011. Economic Analysis, 7th edition. Upper Saddle River, NJ: Prentice-Hall.
Gujarati, Damodar N., and Dawn C. Porter. 2008. Basic Econometrics, 4th edition. New York: McGraw-Hill Irwin.
Hansen, Lars Peter. 1982. “Large Sample Properties of Generalized Method of Moments Estimators.” Econometrica, vol. 50, no. 4: 1029–1054.
Keane, Michael P., and David E. Runkle. 1998. “Are Financial Analysts’ Forecasts of Corporate Profits Rational?” Journal of Political Economy, vol. 106, no. 4: 768–805.
Kmenta, Jan. 1986. Elements of Econometrics, 2nd edition. New York: Macmillan.
Kyaw, NyoNyo A., John Manley, and Anand Shetty. 2011. “Factors in Multinational Valuations: Transparency, Political Risk, and Diversification.” Journal of Multinational Financial Management, vol. 21: 55–67.
MacKinlay, A. Craig, and Matthew P. Richardson. 1991. “Using Generalized Methods of Moments to Test Mean–Variance Efficiency.” Journal of Finance, vol. 46, no. 2: 511–527.
Mankiw, N. Gregory. 2012. Macroeconomics, 8th edition. New York: Worth Publishers.
Mayer, Thomas. 1980. “Economics as a Hard Science: Realistic Goal or Wishful Thinking?” Economic Inquiry, vol. 18, no. 2: 165–178.
Newey, Whitney K., and Kenneth D. West. 1987. “A Simple, Positive Semi- definite, Heteroskedasticity and Autocorrelation Consistent Covariance Matrix.” Econometrica, vol. 55, no. 3: 703–708.
Petersen, Mitchell A. 2009. “Estimating Standard Errors in Finance Panel Data Sets: Comparing Approaches.” Review of Financial Studies, vol. 22, no. 1: 435– 480.
Pinto, Jerald E., Elaine Henry, Thomas R. Robinson, and John D. Stowe. 2010. Equity Asset Valuation, 2nd edition. Charlottesville, VA: CFA Institute.
Sharpe, William F. 1988. “Determining a Fund’s Effective Asset Mix.” Investment Management Review, November/December: 59–69.
Siegel, Jeremy J. 2014. Stocks for the Long Run, 5th edition. New York: McGraw-Hill.
Stoll, Hans R. 1978. “The Pricing of Security Dealer Services: An Empirical Study of Nasdaq Stocks.” Journal of Finance, vol. 33, no. 4: 1153–1172.
Treynor, Jack L., and Kay Mazuy. 1966. “Can Mutual Funds Outguess the Market?” Harvard Business Review, vol. 44: 131–136.
White, Gerald I., Ashwinpaul C. Sondhi, and Dov Fried. 2003. The Analysis and Use of Financial Statements, 3rd edition. Hoboken, NJ: Wiley.
Problems Practice Problems and Solutions: 1–16 taken from Quantitative Methods for Investment Analysis, Second Edition, by Richard A. DeFusco, CFA, Dennis W. McLeavey, CFA, Jerald E. Pinto, PhD, CFA, and David E. Runkle, PhD, CFA. Copyright © 2004 by CFA Institute. All other problems and solutions copyright © CFA Institute. 1. With many US companies operating globally, the effect of the US dollar ’s
strength on a US company’s returns has become an important investment issue. You would like to determine whether changes in the US dollar ’s value and overall US equity market returns affect an asset’s returns. You decide to use the S&P 500 Index to represent the US equity market.
Write a multiple regression equation to test whether changes in the value of the dollar and equity market returns affect an asset’s returns. Use the notations below.
You estimate the regression for Archer Daniels Midland Company (NYSE: ADM). You regress its monthly returns for the period January 1990 to December 2002 on S&P 500 Index returns and changes in the log of the trade-weighted exchange value of the US dollar. The table below shows the coefficient estimates and their standard errors.
Coefficient Estimates from Regressing ADM’s Returns:
Monthly Data, January 1990–December 2002
Coefficient Standard Error Intercept 0.0045 0.0062 RMt 0.5373 0.1332 ΔXt −0.5768 0.5121
n = 156
Source: FactSet, Federal Reserve Bank of Philadelphia.
Determine whether S&P 500 returns affect ADM’s returns. Then determine whether changes in the value of the US dollar affect ADM’s returns. Use a 0.05 significance level to make your decisions.
Based on the estimated coefficient on RMt, is it correct to say that “for a 1 percentage point increase in the return on the S&P 500 in period t, we expect a 0.5373 percentage point increase in the return on ADM”?
2. One of the most important questions in financial economics is what factors determine the cross-sectional variation in an asset’s returns. Some have argued that book-to-market ratio and size (market value of equity) play an important role.
Write a multiple regression equation to test whether book-to-market ratio and size explain the cross-section of asset returns. Use the notations below.
The table below shows the results of the linear regression for a cross- section of 66 companies. The size and book-to-market data for each company are for December 2001. The return data for each company are for January 2002.
Results from Regressing Returns on the Book-to-Market Ratio and Size
Coefficient Standard Error Intercept 0.0825 0.1644 (B/M)i −0.0541 0.0588 Sizei −0.0164 0.0350 n = 66
Source: FactSet.
Determine whether the book-to-market ratio and size are each useful for explaining the cross-section of asset returns. Use a 0.05 significance level to make your decision.
3. There is substantial cross-sectional variation in the number of financial analysts who follow a company. Suppose you hypothesize that a company’s size (market cap) and financial risk (debt-to-equity ratios) influence the number of financial analysts who follow a company. You formulate the following
regression model:
where
(Analyst following)i = the natural log of (1 + n), where ni is the number of analysts following company i
Sizei = the natural log of the market capitalization of company i in millions of dollars
(D/E)i = the debt-to-equity ratio for company i
In the definition of Analyst following, 1 is added to the number of analysts following a company because some companies are not followed by any analysts, and the natural log of 0 is indeterminate. The following table gives the coefficient estimates of the above regression model for a randomly selected sample of 500 companies. The data are for the year 2002.
Coefficient Estimates from Regressing Analyst Following on Size and Debt-to-Equity Ratio
Coefficient Standard Error t-Statistic Intercept −0.2845 0.1080 −2.6343 Sizei 0.3199 0.0152 21.0461 (D/E)i −0.1895 0.0620 −3.0565 n = 500
Source: First Call/Thomson Financial, Compustat.
Consider two companies, both of which have a debt-to-equity ratio of 0.75. The first company has a market capitalization of $100 million, and the second company has a market capitalization of $1 billion. Based on the above estimates, how many more analysts will follow the second company than the first company?
Suppose the p-value reported for the estimated coefficient on (D/E)i is 0.00236. State the interpretation of 0.00236.
4. In early 2001, US equity marketplaces started trading all listed shares in minimal increments (ticks) of $0.01 (decimalization). After decimalization,
bid–ask spreads of stocks traded on the NASDAQ tended to decline. In response, spreads of NASDAQ stocks cross-listed on the Toronto Stock Exchange (TSE) tended to decline as well. Researchers Oppenheimer and Sabherwal (2003) hypothesized that the percentage decline in TSE spreads of cross-listed stocks was related to company size, the predecimalization ratio of spreads on NASDAQ to those on the TSE, and the percentage decline in NASDAQ spreads. The following table gives the regression coefficient estimates from estimating that relationship for a sample of 74 companies. Company size is measured by the natural logarithm of the book value of company’s assets in thousands of Canadian dollars.
Coefficient Estimates from Regressing Percentage Decline in TSE Spreads on Company Size, Predecimalization Ratio of NASDAQ to TSE Spreads, and Percentage Decline in NASDAQ Spreads
Coefficient t-Statistic Intercept −0.45 −1.86 Sizei 0.05 2.56
(Ratio of spreads)i −0.06 −3.77 (Decline in NASDAQ spreads)i 0.29 2.42
n = 74
Source: Oppenheimer and Sabherwal (2003).
The average company in the sample has a book value of assets of C$900 million and a predecimalization ratio of spreads equal to 1.3. Based on the above model, what is the predicted decline in spread on the TSE for a company with these average characteristics, given a 1 percent decline in NASDAQ spreads?
5. The “neglected-company effect” claims that companies that are followed by fewer analysts will earn higher returns on average than companies that are followed by many analysts. To test the neglected-company effect, you have collected data on 66 companies and the number of analysts providing earnings estimates for each company. You decide to also include size as an independent variable, measuring size as the log of the market value of the company’s equity, to try to distinguish any small-company effect from a neglected- company effect. The small-company effect asserts that small-company stocks may earn average higher risk-adjusted returns than large-company stocks.
The table below shows the results from estimating the model Ri = b0 +
b1Sizei + b2(Number of analysts)i + εi for a cross-section of 66 companies. The size and number of analysts for each company are for December 2001. The return data are for January 2002.
Results from Regressing Returns on Size and Number of Analysts
Coefficient Standard Error t-Statistic Intercept 0.0388 0.1556 0.2495 Sizei −0.0153 0.0348 −0.4388
(Number of analysts)i 0.0014 0.0015 0.8995
ANOVA df SS MSS Regression 2 0.0094 0.0047 Residual 63 0.6739 0.0107 Total 65 0.6833
Residual standard error 0.1034 R-squared 0.0138
Observations 66
Source: First Call/Thomson Financial, FactSet.
What test would you conduct to see whether the two independent variables are jointly statistically related to returns (H0: b1 = b2 = 0)?
What information do you need to conduct the appropriate test?
Determine whether the two variables jointly are statistically related to returns at the 0.05 significance level.
Explain the meaning of adjusted R2 and state whether adjusted R2 for the regression would be smaller than, equal to, or larger than 0.0138.
6. Some developing nations are hesitant to open their equity markets to foreign investment because they fear that rapid inflows and outflows of foreign funds will increase volatility. In July 1993, India implemented substantial equity market reforms, one of which allowed foreign institutional investors into the Indian equity markets. You want to test whether the volatility of returns of stocks traded on the Bombay Stock Exchange (BSE) increased after July 1993, when foreign institutional investors were first allowed to invest in India. You have collected monthly return data for the BSE from February 1990 to
December 1997. Your dependent variable is a measure of return volatility of stocks traded on the BSE; your independent variable is a dummy variable that is coded 1 if foreign investment was allowed during the month and 0 otherwise.
You believe that market return volatility actually decreases with the opening up of equity markets. The table below shows the results from your regression.
Results from Dummy Regression for Foreign Investment in India with a Volatility Measure as the Dependent Variable
Coefficient Standard Error t-Statistic Intercept 0.0133 0.0020 6.5351 Dummy −0.0075 0.0027 −2.7604 n = 95
Source: FactSet.
State null and alternative hypotheses for the slope coefficient of the dummy variable that are consistent with testing your stated belief about the effect of opening the equity markets on stock return volatility.
Determine whether you can reject the null hypothesis at the 0.05 significance level (in a one-sided test of significance).
According to the estimated regression equation, what is the level of return volatility before and after the market-opening event?
7. Both researchers and the popular press have discussed the question as to which of the two leading US political parties, Republicans or Democrats, is better for the stock market.
Write a regression equation to test whether overall market returns, as measured by the annual returns on the S&P 500 Index, tend to be higher when the Republicans or the Democrats control the White House. Use the notations below.
The table below shows the results of the linear regression from Part A using annual data for the S&P 500 and a dummy variable for the party that controlled the White House. The data are from 1926 to 2002.
Results from Regressing S&P 500 Returns on a Dummy Variable for the Party That Controlled the White House, 1926-2002
Coefficient Standard Error t-Statistic Intercept 0.1494 0.0323 4.6270 Partyt −0.0570 0.0466 −1.2242
ANOVA df SS MSS F Significance F Regression 1 0.0625 0.0625 1.4987 0.2247 Residual 75 3.1287 0.0417 Total 76 3.1912
Residual standard error 0.2042 R-squared 0.0196
Observations 77
Source: FactSet.
Based on the coefficient and standard error estimates, verify to two decimal places the t-statistic for the coefficient on the dummy variable reported in the table.
Determine at the 0.05 significance level whether overall US equity market returns tend to differ depending on the political party controlling the White House.
8. Problem 3 addressed the cross-sectional variation in the number of financial analysts who follow a company. In that problem, company size and debt-to- equity ratios were the independent variables. You receive a suggestion that membership in the S&P 500 Index should be added to the model as a third independent variable; the hypothesis is that there is greater demand for analyst coverage for stocks included in the S&P 500 because of the widespread use of the S&P 500 as a benchmark.
Write a multiple regression equation to test whether analyst following is systematically higher for companies included in the S&P 500 Index. Also include company size and debt-to-equity ratio in this equation. Use the notations below.
In the above specification for analyst following, 1 is added to the number of analysts following a company because some companies are not followed by any analyst, and the natural log of 0 is indeterminate.
State the appropriate null hypothesis and alternative hypothesis in a two- sided test of significance of the dummy variable.
The following table gives estimates of the coefficients of the above regression model for a randomly selected sample of 500 companies. The data are for the year 2002. Determine whether you can reject the null hypothesis at the 0.05 significance level (in a two-sided test of significance).
Coefficient Estimates from Regressing Analyst Following on Size, Debt-to-Equity Ratio, and S&P 500 Membership, 2002
Coefficient Standard Error t-Statistic Intercept −0.0075 0.1218 −0.0616 Sizei 0.2648 0.0191 13.8639 (D/E)i −0.1829 0.0608 −3.0082 S&Pi 0.4218 0.0919 4.5898 n = 500
Source: First Call/Thomson Financial, Compustat.
Consider a company with a debt-to-equity ratio of 2/3 and a market capitalization of $10 billion. According to the estimated regression equation, how many analysts would follow this company if it were not included in the S&P 500 Index, and how many would follow if it were included in the index?
In Problem 3, using the sample, we estimated the coefficient on the size variable as 0.3199, versus 0.2648 in the above regression. Discuss whether there is an inconsistency in these results.
9. You believe there is a relationship between book-to-market ratios and subsequent returns. The output from a cross-sectional regression and a graph of the actual and predicted relationship between the book-to-market ratio and return are shown below.
Results from Regressing Returns on the Book-to-Market Ratio
Coefficient Standard Error t-Statistic Intercept 12.0130 3.5464 3.3874
−9.2209 8.4454 −1.0918
ANOVA df SS MSS F Significance F Regression 1 154.9866 154.9866 1.1921 0.2831 Residual 32 4162.1895 130.0684 Total 33 4317.1761
Residual standard error 11.4048 R-squared 0.0359
Observations 34
You are concerned with model specification problems and regression assumption violations. Focusing on assumption violations, discuss symptoms of conditional heteroskedasticity based on the graph of the actual and predicted relationship.
Describe in detail how you could formally test for conditional heteroskedasticity in this regression.
Describe a recommended method for correcting for conditional heteroskedasticity.
10. You are examining the effects of the January 2001 NYSE implementation of the trading of shares in minimal increments (ticks) of $0.01 (decimalization). In particular, you are analyzing a sample of 52 Canadian companies cross-listed on both the NYSE and the Toronto Stock Exchange (TSE). You find that the bid–ask spreads of these shares decline on both exchanges after the NYSE decimalization. You run a linear regression analyzing the decline in spreads on the TSE, and find that the decline on the TSE is related to company size, predecimalization ratio of NYSE to TSE spreads, and decline in the NYSE spreads. The relationships are statistically significant. You want to be sure, however, that the results are not influenced by conditional heteroskedasticity. Therefore, you regress the squared residuals of the regression model on the three independent variables. The R2 for this regression is 14.1 percent. Perform a statistical test to determine if conditional heteroskedasticity is present.
11. You are analyzing if institutional investors such as mutual funds and pension funds prefer to hold shares of companies with less volatile returns. You have the percentage of shares held by institutional investors at the end of 1998 for a random sample of 750 companies. For these companies, you compute the standard deviation of daily returns during that year. Then you regress the institutional holdings on the standard deviation of returns. You find that the regression is significant at the 0.01 level and the F-statistic is 12.98. The R2 for this regression is 1.7 percent. As expected, the regression coefficient of the standard deviation of returns is negative. Its t-statistic is −3.60, which is also significant at the 0.01 level. Before concluding that institutions prefer to hold shares of less volatile stocks, however, you want to be sure that the regression results are not influenced by conditional heteroskedasticity. Therefore, you regress the squared residuals of the regression model on the standard deviation of returns. The R2 for this regression is 0.6 percent.
Perform a statistical test to determine if conditional heteroskedasticity is present.
In view of your answer to Part A, what remedial action, if any, is appropriate?
12. In estimating a regression based on monthly observations from January 1987 to December 2002 inclusive, you find that the coefficient on the independent variable is positive and significant at the 0.05 level. You are concerned, however, that the t-statistic on the independent variable may be inflated because of serial correlation between the error terms. Therefore, you examine the Durbin–Watson statistic, which is 1.8953 for this regression.
Based on the value of the Durbin–Watson statistic, what can you say about the serial correlation between the regression residuals? Are they positively correlated, negatively correlated, or not correlated at all?
Compute the sample correlation between the regression residuals from one period and those from the previous period.
Perform a statistical test to determine if serial correlation is present. Assume that the critical values for 192 observations when there is a single independent variable are about 0.09 above the critical values for 100 observations.
13. The book-to-market ratio and the size of a company’s equity are two factors that have been asserted to be useful in explaining the cross-sectional variation in subsequent returns. Based on this assertion, you want to estimate the following regression model:
where
Ri = Return of company i’s shares (in the following period)
= company i’s book-to-market ratio
Sizei = Market value of company i’s equity
A colleague suggests that this regression specification may be erroneous, because he believes that the book-to-market ratio may be strongly related to (correlated with) company size.
To what problem is your colleague referring, and what are its consequences for regression analysis?
With respect to multicollinearity, critique the choice of variables in the regression model above.
Regression of Return on Book-to-Market and Size
Coefficient Standard Error t-Statistic Intercept 14.1062 4.220 3.3427
−12.1413 9.0406 −1.3430
Sizei −0.00005502 0.00005977 −0.92047 R-squared 0.06156
Observations 34 Correlation Matrix
Book-to-Market Ratio Size Book-to-Market Ratio 1.0000
Size −0.3509 1.0000
State the classic symptom of multicollinearity and comment on that basis whether multicollinearity appears to be present, given the additional fact that the F-test for the above regression is not significant.
14. You are analyzing the variables that explain the returns on the stock of the Boeing Company. Because overall market returns are likely to explain a part of the returns on Boeing, you decide to include the returns on a value-weighted index of all the companies listed on the NYSE, AMEX, and NASDAQ as an independent variable. Further, because Boeing is a large company, you also decide to include the returns on the S&P 500 Index, which is a value-weighted index of the larger market-capitalization companies. Finally, you decide to include the changes in the US dollar ’s value. To conduct your test, you have collected the following data for the period 1990–2002.
Rt = monthly return on the stock of Boeing in month t
RALLt = monthly return on a value-weighted index of all the companies listed on the NYSE, AMEX, and NASDAQ in month t
RSPt = monthly return on the S&P 500 Index in month t
ΔXt = change in month t in the log of a trade-weighted index of the foreign exchange value of the US dollar against the currencies of a broad group of major US trading partners
The following table shows the output from regressing the monthly return on Boeing stock on the three independent variables.
Regression of Boeing Returns on Three Explanatory Variables: Monthly Data, January 1990–December 2002
Coefficient Standard Error t-Statistic Intercept 0.0026 0.0066 0.3939 RALLt −0.1337 0.6219 −0.2150 RSPt 0.8875 0.6357 1.3961 ΔXt 0.2005 0.5399 0.3714
ANOVA df SS MSS Regression 3 0.1720 0.0573 Residual 152 0.8947 0.0059 Total 155 1.0667
Residual standard error 0.0767 R-squared 0.1610
Observations 156
Source: FactSet, Federal Reserve Bank of Philadelphia.
From the t-statistics, we see that none of the explanatory variables is statistically significant at the 5 percent level or better. You wish to test, however, if the three variables jointly are statistically related to the returns on Boeing.
Your null hypothesis is that all three population slope coefficients equal 0 —that the three variables jointly are statistically not related to the returns on Boeing. Conduct the appropriate test of that hypothesis.
Examining the regression results, state the regression assumption that may be violated in this example. Explain your answer.
State a possible way to remedy the violation of the regression assumption identified in Part B.
15. You are analyzing the cross-sectional variation in the number of financial analysts that follow a company (also the subject of Problems 3 and 8). You believe that there is less analyst following for companies with a greater debt-
to-equity ratio and greater analyst following for companies included in the S&P 500 Index. Consistent with these beliefs, you estimate the following regression model.
where
(Analysts following)i = natural log of (1 + Number of analysts following company i)
(D/E)i = debt-to-equity ratio for company i
S&Pi = inclusion of company i in the S&P 500 Index (1 if included; 0 if not included)
In the preceding specification, 1 is added to the number of analysts following a company because some companies are not followed by any analysts, and the natural log of 0 is indeterminate. The following table gives the coefficient estimates of the above regression model for a randomly selected sample of 500 companies. The data are for the year 2002.
Coefficient Estimates from Regressing Analyst Following on Debt-to- Equity Ratio and S&P 500 Membership, 2002
Coefficient Standard Error t-Statistic Intercept 1.5367 0.0582 26.4038 (D/E)i −0.1043 0.0712 −1.4649 S&Pi 1.2222 0.0841 14.5327 n = 500
Source: First Call/Thomson Financial, Compustat.
You discuss your results with a colleague. She suggests that this regression specification may be erroneous, because analyst following is likely to be also related to the size of the company.
What is this problem called, and what are its consequences for regression analysis?
To investigate the issue raised by your colleague, you decide to collect data on company size also. You then estimate the model after including an
additional variable, Size i, which is the natural log of the market capitalization of company i in millions of dollars. The following table gives the new coefficient estimates.
Coefficient Estimates from Regressing Analyst Following on Size, Debt- to-Equity Ratio, and S&P 500 Membership, 2002
Coefficient Standard Error t-Statistic Intercept −0.0075 0.1218 −0.0616 Sizei 0.2648 0.0191 13.8639 (D/E)i −0.1829 0.0608 −3.0082 S&Pi 0.4218 0.0919 4.5898 n =500
Source: First Call/Thomson Financial, Compustat.
What do you conclude about the existence of the problem mentioned by your colleague in the original regression model you had estimated?
16. You have noticed that hundreds of non-US companies are listed not only on a stock exchange in their home market but also on one of the exchanges in the United States. You have also noticed that hundreds of non-US companies are listed only in their home market and not in the United States. You are trying to predict whether or not a non-US company will choose to list on a US exchange. One of the factors that you think will affect whether or not a company lists in the United States is its size relative to the size of other companies in its home market.
What kind of a dependent variable do you need to use in the model?
What kind of a model should be used?
The following information relates to Questions 17–22
Gary Hansen is a securities analyst for a mutual fund specializing in small-capitalization growth stocks. The fund regularly invests in initial public offerings (IPOs). If the fund subscribes to an offer, it is allocated shares at the offer price. Hansen notes that IPOs frequently are underpriced, and the price rises when open market trading begins. The initial return for an IPO is calculated as the change in price on the first day of trading divided by the offer price. Hansen is developing a regression model to predict the initial return for IPOs. Based on past research, he
selects the following independent variables to predict IPO initial returns:
Underwriter rank = 1–10, where 10 is highest rank Pre-offer price adjustmenta
= (Offer price – Initial filing price)/Initial filing price
Offer size ($ millions) = Shares sold × Offer price
Fraction retaineda = Fraction of total company shares retained byinsiders
aExpressed as a decimal.
Hansen collects a sample of 1,725 recent IPOs for his regression model. Regression results appear in Exhibit 1, and ANOVA results appear in Exhibit 2.
Hansen wants to use the regression results to predict the initial return for an upcoming IPO. The upcoming IPO has the following characteristics:
underwriter rank = 6;
pre-offer price adjustment = 0.04;
offer size = $40 million;
fraction retained = 0.70.
Exhibit 1 Hansen’s Regression Results Dependent Variable: IPO Initial Return (Expressed in Decimal Form, i.e., 1% = 0.01)
Variable Coefficient (bj) Standard Error t-Statistic Intercept 0.0477 0.0019 25.11
Underwriter rank 0.0150 0.0049 3.06 Pre-offer price adjustment 0.4350 0.0202 21.53
Offer size −0.0009 0.0011 −0.82 Fraction retained 0.0500 0.0260 1.92
Exhibit 2 Selected ANOVA Results for Hansen’s Regression
Degrees of Freedom (df) Sum of Squares (SS) Regression 4 51.433 Residual 1,720 91.436
Total 1,724 142.869 Multiple R-squared = 0.36
Exhibit 3Selected Values for the t-Distribution (df = ∞)
Area in Right Tail t-Value 0.050 1.645 0.025 1.960 0.010 2.326 0.005 2.576
Exhibit 4 Selected Values for the F-Distribution (α = 0.01) (df1/df2: Numerator/Denominator Degrees of Freedom)
df1 4 ∞
df2 4 16.00 13.50 ∞ 3.32 1.00
Because he notes that the pre-offer price adjustment appears to have an important effect on initial return, Hansen wants to construct a 95 percent confidence interval for the coefficient on this variable. He also believes that for each 1 percent increase in pre-offer price adjustment, the initial return will increase by less than 0.5 percent, holding other variables constant. Hansen wishes to test this hypothesis at the 0.05 level of significance.
Before applying his model, Hansen asks a colleague, Phil Chang, to review its specification and results. After examining the model, Chang concludes that the model suffers from two problems: 1) conditional heteroskedasticity, and 2) omitted variable bias. Chang makes the following statements:
Statement 1 “Conditional heteroskedasticity will result in consistent coefficient estimates, but both the t-statistics and F-statistic will be biased, resulting in false inferences.”
Statement 2 “If an omitted variable is correlated with variables already
included in the model, coefficient estimates will be biased and inconsistent and standard errors will also be inconsistent.”
Selected values for the t-distribution and F-distribution appear in Exhibits 3 and 4, respectively.
17. Based on Hansen’s regression, the predicted initial return for the upcoming IPO is closest to:
0.0943.
0.1064.
0.1541.
18. The 95 percent confidence interval for the regression coefficient for the pre- offer price adjustment is closest to:
0.156 to 0.714.
0.395 to 0.475.
0.402 to 0.468.
19. The most appropriate null hypothesis and the most appropriate conclusion regarding Hansen’s belief about the magnitude of the initial return relative to that of the pre-offer price adjustment (reflected by the coefficient bj) are:
Null Hypothesis Conclusion about bj(0.05 Level of Significance)
A. H0: bj = 0.5 Reject H0 B. H0: bj ≥ 0.5 Fail to reject H0 C. H0: bj ≥ 0.5 Reject H0
20. The most appropriate interpretation of the multiple R-squared for Hansen’s model is that:
unexplained variation in the dependent variable is 36 percent of total variation.
correlation between predicted and actual values of the dependent variable is 0.36.
correlation between predicted and actual values of the dependent variable is 0.60.
21. Is Chang’s Statement 1 correct?
Yes.
No, because the model’s F-statistic will not be biased.
No, because the model’s t-statistics will not be biased.
22. Is Chang’s Statement 2 correct?
Yes.
No, because the model’s coefficient estimates will be unbiased.
No, because the model’s coefficient estimates will be consistent.
The following information relates to Questions 23–28
Adele Chiesa is a money manager for the Bianco Fund. She is interested in recent findings showing that certain business condition variables predict excess US stock market returns (one-month market return minus one- month T-bill return). She is also familiar with evidence showing how US stock market returns differ by the political party affiliation of the US President. Chiesa estimates a multiple regression model to predict monthly excess stock market returns accounting for business conditions and the political party affiliation of the US President:
Default spread is equal to the yield on Baa bonds minus the yield on Aaa bonds. Term spread is equal to the yield on a 10-year constant-maturity US Treasury index minus the yield on a 1-year constant-maturity US Treasury index. Pres party dummy is equal to 1 if the US President is a member of the Democratic Party and 0 if a member of the Republican Party.
Chiesa collects 432 months of data (all data are in percent form, i.e., 0.01 = 1 percent). The regression is estimated with 431 observations because the independent variables are lagged one month. The regression output is in Exhibit 1. Exhibits 2 through 5 contain critical values for selected test statistics.
Exhibit 1 Multiple Regression Output (the Dependent Variable Is the One- Month Market Return in Excess of the One-Month T-Bill Return)
Coefficient t-Statistic p-Value Intercept −4.60 −4.36 <0.01
Default spreadt−1 3.04 4.52 <0.01
Term spreadt−1 0.84 3.41 <0.01 Pres party dummyt−1 3.17 4.97 <0.01
Number of observations 431 Test statistic from Breusch–Pagan (BP) test 7.35
R2 0.053 Adjusted R2 0.046
Durbin–Watson (DW) 1.65 Sum of squared errors (SSE) 19,048
Regression sum of squares (SSR) 1,071
An intern working for Chiesa has a number of questions about the results in Exhibit 1:
Question 1 How do you test to determine whether the overall regression model is significant?
Question 2 Does the estimated model conform to standard regression assumptions? For instance, is the error term serially correlated, or is there conditional heteroskedasticity?
Question 3 How do you interpret the coefficient for the Pres party dummy variable?
Question 4 Default spread appears to be quite important. Is there some way to assess the precision of its estimated coefficient? What is the economic interpretation of this variable?
After responding to her intern’s questions, Chiesa concludes with the following statement: “Predictions from Exhibit 1 are subject to parameter estimate uncertainty, but not regression model uncertainty.”
Exhibit 2 Critical Values for the Durbin–Watson Statistic (α = 0.05)
K = 3 N dl du 420 1.825 1.854 430 1.827 1.855 440 1.829 1.857
Exhibit 3 Table of the Student’s t-Distribution (One-Tailed Probabilities for df = ∞)
P t 0.10 1.282 0.05 1.645 0.025 1.960 0.01 2.326 0.005 2.576
Exhibit 4 Values of χ2
Probability in Right Tail df 0.975 0.95 0.05 0.025 1 0.0001 0.0039 3.841 5.024 2 0.0506 0.1026 5.991 7.378 3 0.2158 0.3518 7.815 9.348 4 0.4840 0.7110 9.488 11.14
Exhibit 5 Table of the F-Distribution (Critical Values for Right-Hand Tail Area Equal to 0.05) Numerator: df1 and Denominator: df2
df1 df2 1 2 3 4 427 1 161 200 216 225 254 2 18.51 19.00 19.16 19.25 19.49 3 10.13 9.55 9.28 9.12 8.53 4 7.71 6.94 6.59 6.39 5.64 427 3.86 3.02 2.63 2.39 1.17
23. Regarding the intern’s Question 1, is the regression model as a whole significant at the 0.05 level?
No, because the calculated F-statistic is less than the critical value for F.
B. Yes, because the calculated F-statistic is greater than the critical value for F.
Yes, because the calculated χ2 statistic is greater than the critical value for χ2.
24. Which of the following is Chiesa’s best response to Question 2 regarding serial correlation in the error term? At a 0.05 level of significance, the test for serial correlation indicates that there is:
no serial correlation in the error term.
positive serial correlation in the error term.
negative serial correlation in the error term.
25. Regarding Question 3, the Pres party dummy variable in the model indicates that the mean monthly value for the excess stock market return is:
1.43 percent larger during Democratic presidencies than Republican presidencies.
3.17 percent larger during Democratic presidencies than Republican presidencies.
3.17 percent larger during Republican presidencies than Democratic presidencies.
26. In response to Question 4, the 95 percent confidence interval for the regression coefficient for the default spread is closest to:
0.13 to 5.95.
1.72 to 4.36.
1.93 to 4.15.
27. With respect to the default spread, the estimated model indicates that when business conditions are:
strong, expected excess returns will be higher.
weak, expected excess returns will be lower.
weak, expected excess returns will be higher.
28. Is Chiesa’s concluding statement correct regarding parameter estimate uncertainty and regression model uncertainty?
Yes.
No, predictions are not subject to parameter estimate uncertainty.
No, predictions are subject to regression model uncertainty and parameter estimate uncertainty.
Notes 1 Independent variables are also called explanatory variables or regressors.
2 Note that Δ(ln X) ≈ ΔX/X, where Δ represents “change in” and ΔX/X is a proportional change in X. We discuss the model further in Example 11.
3 An alternative valid formulation is a two-sided test (H0: b1 = 0 versus Ha: b1 | 0) which reflects the beliefs of the researcher less strongly. A two-sided test could also be conducted for the hypothesis on market capitalization that we discuss next.
4 To calculate the degrees of freedom lost in the regression, we add l to the number of independent variables to account for the intercept term.
5 Multiple R2 is also known as the multiple coefficient of determination, or simply the coefficient of determination.
6 To economize on notation in stating test statistics, in this context we use bjto represent the hypothesized value of the parameter (elsewhere we use it to represent the unknown population parameter).
7 The operation illustrated (taking the antilogarithm) recovers the value of a variable in the original units as elnX = X.
8 The entry 0.00 for the significance of F was a p-value for the F-test.
9 The terminology comes from the fact that they correspond to the partial derivatives of Y with respect to the independent variables. Note that in this usage, the term “regression coefficients” refers just to the slope coefficients.
10 Ordinary least squares (OLS) is an estimation method based on the criterion of minimizing the sum of the squared residuals of a regression.
11 As discussed in the reading on correlation and regression, even though we assume that independent variables in the regression model are not random, often that assumption is clearly not true. For example, the monthly returns to the S&P 500 are not random. If the independent variable is random, then is the regression model incorrect? Fortunately, no. Even if the independent variable is random but uncorrelated with the error term, we can still rely on the results of regression models. See, for example, Greene (2011) or Goldberger (1998).
12 No independent variable can be expressed as a linear combination of any set of the other independent variables. Technically, a constant equal to 1 is included as an independent variable associated with the intercept in this condition.
13 Var(ε) = E(ε2) and Cov(εiεj) = E(εiεj) because E(ε) = 0.
14 When we encounter this kind of linear relationship (called perfect collinearity), we cannot compute the matrix inverse needed to compute the linear regression estimates. See Greene (2011) for a further description of this issue.
15 As mentioned in an earlier footnote, technically a constant equal to 1 is included as an independent variable associated with the intercept term in a regression. Because all the regressions reported in this reading include an intercept term, we will not separately mention a constant as an independent variable in the remainder of this reading.
16 Size is the natural log of total sales. A log transformation (either natural log or log base 10) is commonly used for independent variables that can take a wide range of values; company size and fund size are two such variables. One reason to use the log transformation is to improve the statistical properties of the residuals. If the authors had not taken the log of sales and instead used sales as the independent variable, the regression model probably would not have explained Tobin’s q as well.
17 This regression is related to return-based style analysis, one of the most frequent applications of regression analysis in the investment profession. For more information, see Sharpe (1988), who pioneered this field, and Buetow, Johnson, and Runkle (2000).
18 To economize on notation in stating test statistics, in this context we use b2 to represent the hypothesized value of the parameter (elsewhere we use it to represent the unknown population parameter).
19 For more information, see Greene (2011).
20 In a table of regression output, this is the number under “SS” column in the row “Residual.”
21 In a table of regression output, this is the number under the “SS” column in the row “Regression.”
22 We use a one-tailed test because MSR necessarily increases relative to MSE as the
explanatory power of the regression increases.
23 We see a range of values because the denominator has more than 120 degrees of freedom but less than an infinite number of degrees of freedom.
24 We say that variable y is a linear combination of variables x and z if y = ax + bz for some constants a and b. A variable can also be a linear combination of more than two variables.
25 When is negative, we can effectively consider its value to be 0.
26 See Gujarati and Porter (2008). The value of adjusted R2 depends on sample size. These points hold if we are using R2 to compare two regression models.
27 See Mayer (1980).
28 Not all qualitative variables are simple dummy variables. For example, in a trinomial choice model (a model with three choices), a qualitative variable might have the value 0, 1, or 2.
29 For a discussion of this issue, see Siegel (2014).
30 When Jan t = Febt = ... = Novt = 0, the return is not associated with January through November so the month is December and the regression equation simplifies to Returns1 = b0 + εt. Because E (Returns1) = b0 + E (εt) = b0, the intercept b0 represents the mean return for December.
31 Informally, an estimator of a regression parameter is consistent if the probability that estimates of a regression parameter differ from the true value of the parameter decreases as the number of observations used in the regression increases. The regression parameter estimates from ordinary least squares are consistent regardless of whether the errors are heteroskedastic or homoskedastic. For a more advanced discussion, see Greene (2011).
32 This unreliability occurs because the mean squared error is a biased estimator of the true population variance given heteroskedasticity.
33 Sometimes, however, failure to adjust for heteroskedasticity results in standard errors that are too large (and t-statistics that are too small).
34 For more on the CAPM, see Bodie, Kane, and Marcus (2014), for example.
35 MacKinlay and Richardson also show that when using value-weighted returns,
one can reject the CAPM whether or not one assumes normally distributed returns and homoskedasticity.
36 Some other tests require more-specific assumptions about the functional form of the heteroskedasticity. For more information, see Greene (2011).
37 The Breusch–Pagan test is distributed as a χ2 random variable in large samples. The constant 1 technically associated with the intercept term in a regression is not counted here in computing the number of independent variables. For more on the Breusch–Pagan test, see Greene (2011).
38 For more on the Fisher effect, see, for example, Mankiw (2012).
39 For this example, we use the annualized median SPF prediction of current-quarter growth in the GDP deflator (GNP deflator before 1992).
40 Our data on Treasury bill returns are based on three-month T-bill yields in the secondary market. Because those yields are stated on a discount basis, we convert them to a compounded annual rate so they will be measured on the same basis as our data on inflation expectations. These returns are risk-free because they are known at the beginning of the quarter and there is no default risk.
41 Generalized least squares requires econometric expertise to implement correctly on financial data. See Greene (2011), Hansen (1982), and Keane and Runkle (1998).
42 For more details on both methods, see Greene (2011).
43 Robust standard errors are also known as heteroskedasticity-consistent standard errors or White-corrected standard errors.
44 Remember, this is a two-tailed test.
45 We address this issue in the reading on time-series analysis.
46 In contrast, with negative serial correlation, a positive error for one observation increases the chance of a negative error for another observation, and a negative error for one observation increases the chance of a positive error for another.
47 OLS standard errors need not be underestimates of actual standard errors if negative serial correlation is present in the regression.
48 See Greene (2011) for a detailed discussion of tests of serial correlation.
49 Of course, sometimes serial correlation in a regression model is negative rather than positive. For a null hypothesis of no serial correlation, the null hypothesis is rejected if DW < dl (indicating significant positive serial correlation) or if DW > 4 − dl (indicating significant negative serial correlation).
50 This correction is known by various names, including serial-correlation consistent standard errors, serial correlation and heteroskedasticity adjusted standard errors, and robust standard errors. The Hansen standard errors are also known as Hansen–White standard errors.
51 We do not always use Hansen’s method or Newey–West method to correct for serial correlation and heteroskedasticity because sometimes the errors of a regression are not serially correlated.
52 Serial correlation can also affect forecast accuracy.
53 To give an example of perfect collinearity, suppose we tried to explain a company’s credit ratings with a regression that included net sales, cost of goods sold, and gross profit as independent variables. Because Gross profit = Net sales – Cost of goods sold by definition, there is an exact linear relationship between these variables. This type of blunder is relatively obvious (and easy to avoid).
54 See Kmenta (1986).
55 Even if pairs of independent variables have low correlation, there may be linear combinations of the independent variables that are very highly correlated, creating a multicollinearity problem.
56 See Greene (2011).
57 This example is based on Treynor and Mazuy (1966), an early regression study of mutual fund timing. To capture curvature, they included a term in the squared market excess return, which does not violate the assumption of the multiple linear regression model that relationship between the dependent and independent variables is linear in the coefficients.
58 We use a different regression coefficient notation when X2i is omitted, because the intercept term and slope coefficient on X1i will generally not be the same as when X2i is included.
59 The form of the model is analogous to the Cobb–Douglas production function in economics.
60 We have added an error term to the model.
61 The relation between (Bid–ask spread)/Price and ln(Market cap) is also nonlinear, while the relation between ln(Bid–ask spread)/Price and ln(Market cap) is linear. We omit these scatterplots to save space.
62 In our data sample, the bid–ask spread for each of the 2,587 companies is positive.
63 Whether the natural log of the percentage bid–ask spread, Y, is positive or negative, the percentage bid–ask spread found as eY is positive, because a positive number raised to any power is positive. The constant e is positive (e ≈ 2.7183).
64 For more on common size statements, see White, Sondhi, and Fried (2003). Free cash flow and cash flow from operations are discussed in Pinto, Henry, Robinson, and Stowe (2010).
65 “Market-to-book ratio” is the ratio of price per share divided by book value per share.
66 A consistent estimator is one for which the probability of estimates close to the value of the population parameter increases as sample size increases.
67 This proposition does not generalize to regressions with more than one independent variable. Of course, we ignore serially-correlated errors in this example, but because the regression coefficients are inconsistent (due to measurement error), testing or correcting for serial correlation is not worthwhile.
68 We include both unit root tests and tests for cointegration in the term “stationarity tests.”
69 The logistic distribution is easier to compute than the cumulative normal distribution. Consequently, logit models gained popularity when computing power was expensive.
70 For more on probit and logit models, see Greene (2011).
CHAPTER 10 TIME-SERIES ANALYSIS Richard A. DeFusco, CFA
Dennis W. McLeavey, CFA
Jerald E. Pinto, PhD, CFA
David E. Runkle, PhD, CFA
LEARNING OUTCOMES
After completing this chapter, you will be able to do the following:
calculate and evaluate the predicted trend value for a time series, modeled as either a linear trend or a log-linear trend, given the estimated trend coefficients;
describe factors that determine whether a linear or a log-linear trend should be used with a particular time series and evaluate limitations of trend models;
explain the requirement for a time series to be covariance stationary and describe the significance of a series that is not stationary;
describe the structure of an autoregressive (AR) model of order p and calculate one-and two-period-ahead forecasts given the estimated coefficients;
explain how autocorrelations of the residuals can be used to test whether the autoregressive model fits the time series;
explain mean reversion and calculate a mean-reverting level;
contrast in-sample and out-of-sample forecasts and compare the forecasting accuracy of different time-series models based on the root mean squared error criterion;
explain the instability of coefficients of time-series models;
describe characteristics of random walk processes and contrast them to covariance stationary processes;
describe implications of unit roots for time-series analysis, explain when unit roots are likely to occur and how to test for them, and demonstrate how a time
series with a unit root can be transformed so it can be analyzed with an AR model;
describe the steps of the unit root test for nonstationarity and explain the relation of the test to autoregressive time-series models;
explain how to test and correct for seasonality in a time-series model and calculate and interpret a forecasted value using an AR model with a seasonal lag;
explain autoregressive conditional heteroskedasticity (ARCH) and describe how ARCH models can be applied to predict the variance of a time series;
explain how time-series variables should be analyzed for nonstationarity and/or cointegration before use in a linear regression;
determine an appropriate time-series model to analyze a given investment problem and justify that choice.
1. Introduction to Time-Series Analysis As financial analysts, we often use time-series data to make investment decisions. A time series is a set of observations on a variable’s outcomes in different time periods: the quarterly sales for a particular company during the past five years, for example, or the daily returns on a traded security. In this reading, we explore the two chief uses of time-series models: to explain the past and to predict the future of a time series. We also discuss how to estimate time-series models, and we examine how a model describing a particular time series can change over time. The following two examples illustrate the kinds of questions we might want to ask about time series.
Suppose it is the beginning of 2014 and we are managing a US-based investment portfolio that includes Swiss stocks. Because the value of this portfolio would decrease if the Swiss franc depreciates with respect to the dollar, and vice-versa, holding all else constant, we are considering whether to hedge the portfolio’s exposure to changes in the value of the franc. To help us in making this decision, we decide to model the time series of the franc/dollar exchange rate. Figure 1 shows monthly data on the franc/dollar exchange rate. (The data are monthly averages of daily exchange rates.) Has the exchange rate been more stable since 1987 than it was in previous years? Has the exchange rate shown a long-term trend? How can we best use past exchange rates to predict future exchange rates?
FIGURE 1 Swiss Franc/US Dollar Exchange Rate, Monthly Average of Daily Data
Source: Board of Governors of the Federal Reserve System.
FIGURE 2Monthly US Retail Sales
Source: US Department of Commerce, Census Bureau.
As another example, suppose it is the beginning of 2014. We cover retail stores for a sell-side firm and want to predict retail sales for the coming year. Figure 2 shows monthly data on US retail sales. The data are not seasonally adjusted, hence the spikes around the holiday season at the turn of each year. Because the reported sales in the stores’ financial statements are not seasonally adjusted, we model seasonally unadjusted retail sales. How can we model the trend in retail sales? How can we adjust for the extreme seasonality reflected in the peaks and troughs occurring at regular intervals? How can we best use past retail sales to predict future retail sales?
Some fundamental questions arise in time-series analysis: How do we model trends? How do we predict the future value of a time series based on its past values? How do we model seasonality? How do we choose among time-series models? And how do we model changes in the variance of time series over time? We address each of these issues in this reading.
The reading1 is organized as follows. Section 2 describes typical challenges in applying the linear regression model to time series data. Section 3 presents linear and log-linear trend models, which describe, respectively, the value and the natural log of the value of a time series as a linear function of time. Section 4 presents autoregressive time series models—which explain the current value of a time series in terms of one or more lagged values of the series. Such models are among the most commonly used in investments, and the section addresses many related concepts and issues. Section 5 addresses time series that are random walks. Because
such time series are not covariance stationary, they cannot be modeled using autoregressive models unless they can be transformed into stationary series. The section explores appropriate transformations and tests of stationarity. Section 6 addresses moving-average time-series models, and Section 7 discusses the problem of seasonality in time series and how to address it. Section 8 covers autoregressive moving-average models, a more complex alternative to autoregressive models. Section 9 addresses modeling changing variance of the error term in a time series. Section 10 examines the consequences of regression of one time series on another when one or both time series may not be covariance stationary.
2. Challenges of Working with Time Series Throughout the reading, our objective will be to apply linear regression to a given time series. Unfortunately, in working with time series we often find that the assumptions of the linear regression model are not satisfied. To apply time-series analysis, we need to assure ourselves that the linear regression model assumptions are met. When those assumptions are not satisfied, in many cases we can transform the time series, or specify the regression model differently, so that the assumptions of the linear regression model are met.
We can illustrate assumption difficulties in the context of a common time-series model, an autoregressive model. Informally, an autoregressive model is one in which the independent variable is a lagged (that is, past) value of the dependent variable, such as the model xt = b0 + b1xt −1 + εt.2 Specific problems that we often encounter in dealing with time series include the following:
The residual errors are correlated instead of being uncorrelated. In the calculated regression, the difference between xt and b0 + b1xt −1 is called the residual error (εt). The linear regression assumes that this error term is not correlated across observations. The violation of that assumption is frequently more critical in terms of its consequences in the case of time-series models involving past values of the time series as independent variables than for other models (such as cross-sectional) in which the dependent and independent variables are distinct. As we discussed in the reading on multiple regression, in a regression in which the dependent and independent variables are distinct, serial correlation of the errors in this model does not affect the consistency of our estimates of intercept or slope coefficients. By contrast, in an autoregressive time-series regression such as xt = b0 + b1xt −1 + εt, serial correlation in the error term causes estimates of the intercept (b0) and slope coefficient (b1) to be inconsistent.
The mean and/or variance of the time series changes over time. Regression results are invalid if we estimate an autoregressive model for a time series with mean and/or variance that changes over time.
Before we try to use time series for forecasting, we may need to transform the time- series model so that it is well specified for linear regression. With this objective in mind, you will observe that time-series analysis is relatively straightforward and logical.
3. Trend Models Estimating a trend in a time series and using that trend to predict future values of the time series is the simplest method of forecasting. For example, we saw in Figure 2 that monthly US retail sales show a long-term pattern of upward movement—that is, a trend. In this section, we examine two types of trends—linear trends and log- linear trends—and discuss how to choose between them.
3.1. Linear Trend Models The simplest type of trend is a linear trend, one in which the dependent variable changes at a constant rate with time. If a time series, yt, has a linear trend, then we can model the series using the following regression equation:
(1)
where
In Equation 1, the trend line, b0 + b1t, predicts the value of the time series at time t (where t takes on a value of 1 in the first period of the sample and increases by 1 in each subsequent period). Because the coefficient b1 is the slope of the trend line, we refer to b1 as the trend coefficient. We can estimate the two coefficients, b0 and b1, using ordinary least squares, denoting the estimated coefficients as and .3
Now we demonstrate how to use these estimates to predict the value of the time series in a particular period. Recall that t takes on a value of 1 in Period 1. Therefore, the predicted or fitted value of yt in Period 1 is . Similarly, in
a subsequent period, say the sixth period, the fitted value is . Now suppose that we want to predict the value of the time series for a period outside the sample, say period T + 1. The predicted value of yt for period T + 1 is
. For example, if is 5.1 and is 2, then at t = 5 the predicted value of y5 is 15.1 and at t = 6 the predicted value of y6 is 17.1. Note that each consecutive observation in this time series increases by = 2 irrespective of the level of the series in the previous period.
EXAMPLE 1 The Trend in the US Consumer Price Index
It is January 2014. As a fixed income analyst in the trust department of a bank, Lisette Miller is concerned about the future level of inflation and how it might affect portfolio value. Therefore, she wants to predict future inflation rates. For this purpose, she first needs to estimate the linear trend in inflation. To do so, she uses the monthly US Consumer Price Index (CPI) inflation data, expressed as an annual percentage rate,4 shown in Figure 3. The data include 228 months from January 1995 through December 2013, and the model to be estimated is yt = b0 + b1t + εt, t = 1, 2, …, 228. Table 1 shows the results of estimating this equation. With 228 observations and two parameters, this model has 226 degrees of freedom. At the 0.05 significance level, the critical value for a t- statistic is 1.97. The intercept is statistically significant because the value of the t-statistic for the coefficient is well above the critical value. The trend coefficient is negative , suggesting a declining trend in inflation during the sample time period. However, the trend is not statistically significant because the absolute value of the t-statistic for the coefficient is well below the critical value. The estimated regression equation can be written as
FIGURE 3 Monthly CPI Inflation, Not Seasonally Adjusted
Source: Bureau of Labor Statistics.
Because the trend line slope is estimated to be −0.0038, Miller concludes that
the linear trend model’s best estimate is that the annualized rate of inflation declined at a rate of about 38 basis points per month during the sample time period. The decline is not statistically significantly different from zero.
TABLE 1 Estimating a Linear Trend in Inflation Monthly Observations, January 1995–December 2013
Regression Statistics R-squared 0.0033
Standard error 4.3297 Observations 228 Durbin–Watson 1.09
Coefficient Standard Error t-Statistic Intercept 2.8853 0.5754 5.0144 t (Trend) −0.0038 0.0044 −0.8636
Source: US Bureau of Labor Statistics.
In January 1995, the first month of the sample, the predicted value of inflation is = 2.8853 − 0.0038(1) = 2.8815 percent. In December 2013, the 228th or last month of the sample, the predicted value of inflation is = 2.8853 − 0.0038(228) = 2.0189 percent. Note, though, that these predicted values are for in-sample periods. A comparison of these values with the actual values indicates how well Miller ’s model fits the data; however, a main purpose of the estimated model is to predict the level of inflation for out-of-sample periods. For example, for December 2014 (12 months after the end of the sample), t = 228 + 12 = 240, and the predicted level of inflation is = 2.8853 − 0.0038(240) = 1.9733 percent.
Figure 4 shows the inflation data along with the fitted trend. Consistent with the negative but small and statistically insignificant trend coefficient, the fitted trend line is slightly downward sloping. Note that inflation does not appear to be above or below the trend line for a long period of time. No persistent differences exist between the trend and actual inflation. The residuals (actual minus trend values) appear to be unpredictable and uncorrelated in time. Therefore, using a linear trend line to model inflation rates from 1995 through 2013 does not appear to violate the assumptions of the linear regression model. Note also that the R2 in this model is quite low, indicating great uncertainty in the inflation forecasts from this model. In fact, the estimated model explains
only 0.33 percent of the variation in monthly inflation. Although linear trend models have their uses, they are often inappropriate for economic data. Most economic time series reflect trends with changing slopes and/or intercepts over time. The linear trend model identifies the slope and intercept that provides the best linear fit for all past data. The model’s deviation from the actual data can be greatest near the end of a data series which can compromise forecasting accuracy. Later in this reading, we will examine whether we can build a better model of inflation than a model that uses only a trend line.
FIGURE 4Monthly CPI Inflation with Trend
Source: US Bureau of Labor Statistics.
3.2. Log-Linear Trend Models Sometimes a linear trend does not correctly model the growth of a time series. In those cases, we often find that fitting a linear trend to a time series leads to persistent rather than uncorrelated errors. If the residuals from a linear trend model are persistent, we then need to employ an alternative model satisfying the conditions of linear regression. For financial time series, an important alternative to a linear trend is a log-linear trend. Log-linear trends work well in fitting time series that have exponential growth.
Exponential growth means constant growth at a particular rate. For example, annual growth at a constant rate of 5 percent is exponential growth. How does exponential growth work? Suppose we describe a time series by the following equation:
(2)
Exponential growth is growth at a constant rate with continuous compounding. For instance, consider values of the time series in two consecutive periods. In Period 1, the time series has the value y1 = , and in Period 2, it
has the value y2 = . The resulting ratio of the values of the time series in
the first two periods is y2/y1 = = .Generally, in any period t, the time series has the value yt = . In period t + 1, the time series has the
value yt +1 = . The ratio of the values in the periods (t + 1) and t is yt +1/yt = = . Thus, the proportional rate of growth in the time series over two consecutive periods is always the same: − 1.5 Therefore, exponential growth is growth at a constant rate. Continuous compounding is a mathematical convenience that allows us to restate the equation in a form that is easy to estimate.
If we take the natural log of both sides of Equation 2, the result is the following equation:
Therefore, if a time series grows at an exponential rate, we can model the natural log of that series using a linear trend.6 Of course, no time series grows exactly at a constant rate. Consequently, if we want to use a log-linear model, we must estimate the following equation:
(3)
Note that this equation is linear in the coefficients b0 and b1. In contrast to a linear trend model, in which the predicted trend value of yt is , the predicted trend value of yt in a log-linear trend model is because eln yt = yt.
Examining Equation 3, we see that a log-linear model predicts that ln yt will increase by b1 from one time period to the next. The model predicts a constant growth rate in yt of .For example, if b1 = 0.05, then the predicted growth rate of yt in each period is e0.05 − 1 = 0.051271 or 5.13 percent. In contrast, the linear trend model (Equation 1) predicts that yt grows by a constant amount from one period to
the next.
Example 2 illustrates the problem of nonrandom residuals in a linear trend model, and Example 3 shows a log-linear regression fit to the same data.
EXAMPLE 2 A Linear Trend Regression for Quarterly Sales at Starbucks
In October 2013, technology analyst Ray Benedict wants to use Equation 1 to fit the data on quarterly sales for Starbucks Corporation shown in Figure 5. Starbucks’ fiscal year ends in September. Benedict uses 76 observations on Starbucks’ sales from the first quarter of fiscal year 1995 (starting in October 1994) to the fourth quarter of fiscal year 2013 (ending in September 2013) to estimate the linear trend regression model yt = b0 + b1t + εt, t = 1, 2, …, 76.7
Table 2 shows the results of estimating this equation.
FIGURE 5Starbucks Quarterly Sales by Fiscal Year
Source: Compustat.
TABLE 2 Estimating a Linear Trend in Starbucks Sales
Regression Statistics R-squared 0.9595
Standard error 233.21 Observations 76
Durbin–Watson 0.32 Coefficient Standard Error t-Statistic
Intercept −428.5380 54.0345 −7.9308 t (Trend) 51.0866 1.2194 41.8949
Source: Compustat.
FIGURE 6 Starbucks Quarterly Sales with Trend
Source: Compustat.
At first glance, the results shown in Table 2 seem quite reasonable: Both the intercept and the trend coefficient are highly statistically significant. When Benedict plots the data on Starbucks’ sales and the trend line, however, he sees a different picture. As Figure 6 shows, before 1998 the trend line is persistently below sales. Subsequently, until 2006, the trend line is persistently above sales and then varies somewhat thereafter.
Recall a key assumption underlying the regression model: that the regression errors are not correlated across observations. If a trend is persistently above or below the value of the time series, however, the residuals (the difference between the time series and the trend) are serially correlated. Figure 7 shows the residuals (the difference between sales and the trend) from estimating a linear trend model with the raw sales data. The figure shows that the residuals
are persistent: they are consistently positive from 1995 to 1998, 2007 to 2009, and after 2012 and consistently negative from 1999 to 2006.
Because of this persistent serial correlation in the errors of the trend model, using a linear trend to fit sales at Starbucks would be inappropriate, even though the R2 of the equation is high (0.96). The assumption of uncorrelated residual errors has been violated. Because the dependent and independent variables are not distinct, as in cross-sectional regressions, this assumption violation is serious and causes us to search for a better model.
FIGURE 7 Residual from Predicting Starbucks Sales with a Trend
Source: Compustat.
EXAMPLE 3 A Log-Linear Regression for Quarterly Sales at Starbucks
Having rejected a linear trend model in Example 2, technology analyst Benedict now tries a different model for the quarterly sales for Starbucks Corporation from the first quarter of 1995 to the fourth quarter of 2013. The curvature in the data plot shown in Figure 5 is a hint that an exponential curve may fit the data. Consequently, he estimates the following linear equation:
This equation seems to fit the sales data well. As Table 3 shows, the R2 for this equation is 0.95. An R2 of 0.95 means that 95 percent of the variation in the natural log of Starbucks’ sales is explained solely by a linear trend.
TABLE 3 Estimating a Linear Trend in Lognormal Starbucks Sales
Regression Statistics R-squared 0.9453
Standard error 0.2480 Observations 76 Durbin–Watson 0.12
Coefficient Standard Error t-Statistic Intercept 5.1304 0.0575 89.2243 t (Trend) 0.0464 0.0013 35.6923
Source: Compustat.
Although both Equations 1 and 3 have a high R2, Figure 8 shows how well a linear trend fits the natural log of Starbucks’ sales (Equation 3). The natural logs of the sales data lie very close to the linear trend during the sample period, and log sales are not substantially above or below the trend for long periods of time. Thus, a log-linear trend model seems better suited for modeling Starbucks’ sales than does a linear trend model.
1. Benedict wants to use the results of estimating Equation 3 to predict Starbucks’
sales in the future. What is the predicted value of Starbucks’ sales for the first quarter of 2014?
Solution to 1: The estimated value is 5.1304, and the estimated value is 0.0464. Therefore, for the first quarter of 2014 (t = 77), the estimated model predicts that ln = 5.1304 + 0.0464(77) = 8.7032 and that predicted sales are = e8.7032 = $6,022.15 million.8
FIGURE 8Natural Log of Starbucks Quarterly Sales
Source: Compustat.
2. How much different is the above forecast from the prediction of the linear trend model?
Solution to 2: Table 2 showed that for the linear trend model, the estimated value of is −428.5380 and the estimated value of is 51.0866. Thus, if we predict Starbucks’ sales for the first quarter of 2014 (t = 77) using the linear trend model, the forecast is = −428.5380 + 51.0866(77) = $3,505.13 million. This forecast is far below the prediction made by the log-linear regression model. Later in this reading, we will examine whether we can build a better model of Starbucks’ quarterly sales than a model that uses only a log-linear trend.
3.3. Trend Models and Testing for Correlated Errors Both the linear trend model and the log-linear trend model are single-variable regression models. If they are to be correctly specified, the regression-model assumptions must be satisfied. In particular, the regression error for one period must be uncorrelated with the regression error for all other periods.9 In Example 2 in the previous section, we could infer an obvious violation of that assumption from a visual inspection of a plot of residuals (Figure 7). The log-linear trend model of Example 3 appeared to fit the data much better, but we still need to confirm that the
uncorrelated errors assumption is satisfied. To address that question formally, we must carry out a Durbin–Watson test on the residuals.
In the reading on regression analysis, we showed how to test whether regression errors are serially correlated using the Durbin–Watson statistic. For example, if the trend models shown in Examples 1 and 3 really capture the time-series behavior of inflation and the log of Starbucks’ sales, then the Durbin–Watson statistic for both of those models should not differ significantly from 2.0. Otherwise, the errors in the model are either positively or negatively serially correlated, and that correlation can be used to build a better forecasting model for those time series.
In Example 1, estimating a linear trend in the monthly CPI inflation yielded a Durbin–Watson statistic of 1.09. Is this result significantly different from 2.0? To find out, we need to test the null hypothesis of no positive serial correlation. For a sample with 228 observations and one independent variable, the critical value, dl, for the Durbin–Watson test statistic at the 0.05 significance level is above 1.77. Because the value of the Durbin–Watson statistic (1.09) is below this critical value, we can reject the hypothesis of no positive serial correlation in the errors. We can conclude that a regression equation that uses a linear trend to model inflation has positive serial correlation in the errors.10 We will need a different kind of regression model because this one violates the least-squares assumption of no serial correlation in the errors.
In Example 3, estimating a linear trend with the natural logarithm of sales for the Starbucks example yielded a Durbin–Watson statistic of 0.12. Suppose we wish to test the null hypothesis of no positive serial correlation. The critical value, dl, is above 1.60 at the 0.05 significance level. The value of the Durbin–Watson statistic (0.12) is below this critical value, so we can reject the null hypothesis of no positive serial correlation in the errors. We can conclude that a regression equation that uses a trend to model the log of Starbucks’ quarterly sales has positive serial correlation in the errors. So, for this series as well, we need to build a different kind of model.
Overall, we conclude that the trend models sometimes have the limitation that errors are serially correlated. Existence of serial correlation suggests that we can build better forecasting models for such time series than trend models.
4. Autoregressive (AR) Time-Series Models A key feature of the log-linear model’s depiction of time series and a key feature of time series in general is that current-period values are related to previous-period values. For example, Starbucks’ sales for the current period are related to its sales in the previous period. An autoregressive model (AR), a time series regressed on its own past values, represents this relationship effectively. When we use this model, we can drop the normal notation of y as the dependent variable and x as the independent variable because we no longer have that distinction to make. Here we simply use xt. For example, Equation 4 shows a first-order autoregression, AR(1), for the variable xt:
(4)
Thus, in an AR(1) model, we use only the most recent past value of xt to predict the current value of xt. In general, a pth-order autoregression, AR(p), for the variable xt is shown by
(5)
In this equation, p past values of xt are used to predict the current value of xt. In the next section we discuss a key assumption of time-series models that include lagged values of the dependent variable as independent variables.
4.1. Covariance-Stationary Series Note that the independent variable (xt −1) in Equation 4 is a random variable. This fact may seem like a mathematical subtlety, but it is not. If we use ordinary least squares to estimate Equation 4 when we have a randomly distributed independent variable that is a lagged value of the dependent variable, our statistical inference may be invalid. To conduct valid statistical inference, we must make a key assumption in time-series analysis: We must assume that the time series we are modeling is covariance stationary.11
What does it mean for a time series to be covariance stationary? The basic idea is that a time series is covariance stationary if its properties, such as mean and variance, do not change over time. A covariance stationary series must satisfy three principal requirements.12 First, the expected value of the time series must be constant and finite in all periods: E(yt) = μ and |μ| < ∞, t = 1, 2, …, T. Second, the variance of the time series must be constant and finite in all periods. Third, the
covariance of the time series with itself for a fixed number of periods in the past or future must be constant and finite in all periods. The second and third requirements can be summarized as follows:13
where λ signifies a constant. What happens if a time series is not covariance stationary but we model it using Equation 4? The estimation results will have no economic meaning. For a non-covariance-stationary time series, estimating the regression in Equation 4 will yield spurious results. In particular, the estimate of b1 will be biased, and any hypothesis tests will be invalid.
How can we tell if a time series is covariance stationary? We can often answer this question by looking at a plot of the time series. If the plot shows roughly the same mean and variance through time without any significant seasonality, then we may want to assume that the time series is covariance stationary.
Some of the time series we looked at in Figures 1 to 4 appear to be covariance stationary. For example, the inflation data shown in Figure 3 appear to have roughly the same mean and variance over the sample period. Many of the time series one encounters in business and investments, however, are not covariance stationary. For example, many time series appear to grow (or decline) steadily through time and so have a mean that is nonconstant, which implies that they are nonstationary. As an example, the time series of quarterly sales in Figure 5 clearly shows the mean increasing as time passes. Thus Starbucks’ quarterly sales are not covariance stationary.14 Macroeconomic time series such as those relating to income and consumption are often strongly trending as well. A time series with seasonality (regular patterns of movement with the year) also has a nonconstant mean, as do other types of time series that we discuss later.15
Figure 2 showed that monthly retail sales (not seasonally adjusted) are also not covariance stationary. Sales in December are always much higher than sales in other months (these are the regular large peaks), and sales in January are always much lower (these are the regular large drops after the December peaks). On average, sales also increase over time, so the mean of sales is not constant.
Later in the reading, we will show that we can often transform a nonstationary time series into a stationary time series. But whether a stationary time series is original or transformed, a caution applies: Stationarity in the past does not guarantee stationarity in the future. There is always the possibility that a well-specified model will fail when the state of the world changes and yields a different underlying model that generates the time series.
4.2. Detecting Serially Correlated Errors in an Autoregressive Model We can estimate an autoregressive model using ordinary least squares if the time series is covariance stationary and the errors are uncorrelated. Unfortunately, our previous test for serial correlation, the Durbin–Watson statistic, is invalid when the independent variables include past values of the dependent variable. Therefore, for most time-series models, we cannot use the Durbin–Watson statistic. Fortunately, we can use other tests to determine whether the errors in a time-series model are serially correlated. One such test reveals whether the autocorrelations of the error term are significantly different from 0. This test is a t-test involving a residual autocorrelation and the standard error of the residual autocorrelation. As background for the test, we next discuss autocorrelation in general before moving to residual autocorrelation.
The autocorrelations of a time series are the correlations of that series with its own past values. The order of the correlation is given by k where k represents the number of periods lagged. When k = 1, the autocorrelation shows the correlation of the variable in one period to its occurrence in the previous period. For example, the kth order autocorrelation (ρk) is
where E stands for the expected value. Note that we have the relationship Cov(xt, xt −k) ≤ with equality holding when k = 0. This means that the absolute value of ρk is less than or equal to 1.
Of course, we can never directly observe the autocorrelations, ρk. Instead, we must estimate them. Thus we replace the expected value of xt, μ, with its estimated value, , to compute the estimated autocorrelations. The kth order estimated autocorrelation of the time series xt, which we denote , is
Analogous to the definition of autocorrelations for a time series, we can define the autocorrelations of the error term for a time-series model as16
We assume that the expected value of the error term in a time-series model is 0.17
We can determine whether we are using the correct time-series model by testing whether the autocorrelations of the error term (error autocorrelations) differ significantly from 0. If they do, the model is not specified correctly. We estimate the error autocorrelation using the sample autocorrelations of the residuals (residual autocorrelations) and their sample variance.
A test of the null hypothesis that an error autocorrelation at a specified lag equals 0 is based on the residual autocorrelation for that lag and the standard error of the residual correlation, which is equal to , where T is the number of observations in the time series.18 Thus, if we have 100 observations in a time series, the standard error for each of the estimated autocorrelations is 0.1. We can compute the t-test of the null hypothesis that the error correlation at a particular lag equals 0, by dividing the residual autocorrelation at that lag by its standard error .
How can we use information about the error autocorrelations to determine whether an autoregressive time-series model is correctly specified? We can use a simple three-step method. First, estimate a particular autoregressive model, say an AR(1) model. Second, compute the autocorrelations of the residuals from the model.19 Third, test to see whether the residual autocorrelations differ significantly from 0. If significance tests show that the residual autocorrelations differ significantly from 0, the model is not correctly specified; we may need to modify it in ways that we will discuss shortly.20 We now present an example to demonstrate how this three-step method works.
EXAMPLE 4 Predicting Gross Margins for Intel Corporation
Analyst Melissa Jones decides to use a time-series model to predict Intel Corporation’s gross margin [(Sales – Cost of goods sold)/Sales] using quarterly data from the second quarter of 1999 through the fourth quarter of 2013. She does not know the best model for gross margin but believes that the current-period value will be related to the previous-period value. She decides
to start out with a first-order autoregressive model, AR(1): Gross margint = b0 + b1(Gross margint −1) + εt. Her observations on the dependent variable are 2Q:1999 through 4Q:2013. Table 4 shows the results of estimating this AR(1) model, along with the autocorrelations of the residuals from that model.
Table 4 Autoregression: AR(1) Model Gross Margin of Intel Quarterly Observations, April 1999–December 2013
Regression Statistics R-squared 0.5429
Standard error 0.0337 Observations 59 Durbin–Watson 2.0987
Coefficient Standard Error t-Statistic Intercept 0.1795 0.0635 2.8268
Gross margint −1 0.7449 0.0905 8.2309 Autocorrelations of the Residual
Lag Autocorrelation Standard Error t-Statistic 1 −0.0495 0.1302 −0.3802 2 −0.0392 0.1302 −0.3011 3 0.0524 0.1302 0.4025 4 0.1450 0.1302 1.1137
Source: Compustat.
The first thing to note about Table 4 is that both the intercept ( = 0.1795) and the coefficient on the first lag ( = 0.7449) of the gross margin are highly significant in the regression equation.21 The t-statistic for the intercept is about 2.8, whereas the t-statistic for the first lag of the gross margin is more than 8. With 59 observations and two parameters, this model has 57 degrees of freedom. At the 0.05 significance level, the critical value for a t-statistic is about 2.0. Therefore, Jones must reject the null hypotheses that the intercept is equal to 0 (b0 = 0) and the coefficient on the first lag is equal to 0 (b1 = 0) in favor of the alternative hypothesis that the coefficients, individually, are not equal to 0. But are these statistics valid? Although the Durbin–Watson statistic is presented in Table 4, it cannot be used to test serial correlation when the
independent variables include past values of the dependent variable. The correct approach is to test whether the residuals from this model are serially correlated.
At the bottom of Table 4, the first four autocorrelations of the residual are displayed along with the standard error and the t-statistic for each of those autocorrelations.22 The sample has 59 observations, so the standard error for each of the autocorrelations is = 0.1302. Table 4 shows that none of the first four autocorrelations has a t-statistic larger than 1.11 in absolute value. Therefore, Jones can conclude that none of these autocorrelations differs significantly from 0. Consequently, she can assume that the residuals are not serially correlated and that the model is correctly specified, and she can validly use ordinary least squares to estimate the parameters and the parameters’ standard errors in the autoregressive model.23
Now that Jones has concluded that this model is correctly specified, how can she use it to predict Intel’s gross margin in the next period? The estimated equation is Gross margint = 0.1795 + 0.7449(Gross margint −1) + εt. The expected value of the error term is 0 in any period. Therefore, this model predicts that gross margin in period t + 1 will be Gross margint +1 = 0.1795 + 0.7449(Gross margint). For example, if gross margin is 65 percent in this quarter (0.65), the model predicts that in the next quarter gross margin will increase to 0.1795 + 0.7449(0.65) = 0.6637 or 66.37 percent. On the other hand, if gross margin is currently 75 percent (0.75), the model predicts that in the next quarter, gross margin will fall to 0.1795 + 0.7449(0.75) = 0.7382 or 73.82 percent. As we show in the following section, the model predicts that gross margin will increase if it is below a certain level (70.36 percent) and decrease if it is above that level.
4.3. Mean Reversion We say that a time series shows mean reversion if it tends to fall when its level is above its mean and rise when its level is below its mean. Much like the temperature in a room controlled by a thermostat, a mean-reverting time series tends to return to its long-term mean. How can we determine the value that the time series tends toward? If a time series is currently at its mean-reverting level, then the model predicts that the value of the time series will be the same in the next period. At its mean-reverting level, we have the relationship xt +1 = xt. For an AR(1) model (xt +1 = b0 + b1xt), the equality xt +1 = xt implies the level xt = b0 + b1xt, or that the mean- reverting level, xt, is given by
So the AR(1) model predicts that the time series will stay the same if its current value is b0/(1 − b1), increase if its current value is below b0/(1 − b1), and decrease if its current value is above b0/(1 − b1).
In the case of gross margins for Intel, the mean-reverting level for the model shown in Table 4 is 0.1795/(1 − 0.7449) = 0.7036. If the current gross margin is above 0.7036, the model predicts that the gross margin will fall in the next period. If the current gross margin is below 0.7036, the model predicts that the gross margin will rise in the next period. As we will discuss later, all covariance-stationary time series have a finite mean-reverting level.
4.4. Multiperiod Forecasts and the Chain Rule of Forecasting Often, financial analysts want to make forecasts for more than one period. For example, we might want to use a quarterly sales model to predict sales for a company for each of the next four quarters. To use a time-series model to make forecasts for more than one period, we must examine how to make multiperiod forecasts using an AR(1) model. The one-period-ahead forecast of xt from an AR(1) model is as follows:
(6)
If we want to forecast xt +2 using an AR(1) model, our forecast will be based on
(7)
Unfortunately, we do not know xt +1 in period t, so we cannot use Equation 7 directly to make a two-period-ahead forecast. We can, however, use our forecast of xt +1 and the AR(1) model to make a prediction of xt +2. The chain rule of forecasting is a process in which the next period’s value, predicted by the forecasting equation, is substituted into the equation to give a predicted value two periods ahead. Using the chain rule of forecasting, we can substitute the predicted value of xt +1 into Equation 7 to get . . We already know from our one-period-ahead forecast in Equation 6. Now we have a simple way of predicting xt +2.
Multiperiod forecasts are more uncertain than single-period forecasts because each forecast period has uncertainty. For example, in forecasting xt +2, we first have the
uncertainty associated with forecasting xt +1 using xt, and then we have the uncertainty associated with forecasting xt +2 using the forecast of xt +1. In general, the more periods a forecast has, the more uncertain it is.24
EXAMPLE 5 Multiperiod Prediction of Intel’s Gross Margin
Suppose that at the beginning of 2014, we want to predict Intel’s gross margin in two periods using the model shown in Table 4. Assume that Intel’s gross margin in the current period is 65 percent. The one-period-ahead forecast of Intel’s gross margin from this model is 0.6637 = 0.1795 + 0.7449(0.65). By substituting the one-period-ahead forecast, 0.6637, back into the regression equation, we can derive the following two-period-ahead forecast: 0.6739 = 0.1795 + 0.7449(0.6637). Therefore, if the current gross margin for Intel is 65 percent, the model predicts that Intel’s gross margin in two quarters will be 67.39 percent.
EXAMPLE 6 Modeling US CPI Inflation
Analyst Lisette Miller has been directed to build a time-series model for monthly US inflation. Inflation and expectations about inflation, of course, have a significant effect on bond returns. For a 30-year period beginning with January 1984 and ending with December 2013, she selects as data the annualized monthly percentage change in the CPI. Which model should Miller use?
The process of model selection parallels that of Example 4 relating to Intel’s gross margins. The first model Miller estimates is an AR(1) model, using the previous month’s inflation rate as the independent variable: Inflationt = b0 + b1. Inflationt −1 + εt, t = 1, 2, …, 359. To estimate this model, she uses monthly CPI inflation data from January 1984 to December 2013 (t = 1 denotes February 1984). Table 5 shows the results of estimating this model.
TABLE 5 Monthly CPI Inflation at an Annual Rate: AR(1) Model Monthly Observations, February 1984–December 2013
Regression Statistics R-squared 0.2038
Standard error 3.4250 Observations 359 Durbin–Watson 1.8201
Coefficient Standard Error t-Statistic Intercept 1.5703 0.2266 6.9298
Inflationt −1 0.4510 0.0472 9.5551 Autocorrelations of the Residual
Lag Autocorrelation Standard Error t-Statistic 1 0.0898 0.0528 1.7008 2 −0.1205 0.0528 −2.2822 3 −0.1571 0.0528 −2.9754 4 −0.0316 0.0528 −0.5985
Source: US Bureau of Labor Statistics.
As Table 5 shows, both the intercept ( = 1.5703) and the coefficient on the first lagged value of inflation ( = 0.4510) are highly statistically significant, with large t-statistics. With 359 observations and two parameters, this model has 357 degrees of freedom. The critical value for a t-statistic at the 0.05 significance level is about 1.97. Therefore, Miller can reject the individual null hypotheses that the intercept is equal to 0 (b0 = 0) and the coefficient on the first lag is equal to 0 (b1 = 0) in favor of the alternative hypothesis that the coefficients, individually, are not equal to 0.
Are these statistics valid? Miller will know when she tests whether the residuals from this model are serially correlated. With 359 observations in this sample, the standard error for each of the estimated autocorrelations is = 0.0528. The critical value for the t-statistic is 1.97. Because both the second and the third estimated autocorrelation have t-statistics larger than 1.97 in absolute value, Miller concludes that the autocorrelations are significantly different from 0. This model is thus misspecified because the residuals are serially correlated.
If the residuals in an autoregressive model are serially correlated, Miller can eliminate the correlation by estimating an autoregressive model with more lags of the dependent variable as explanatory variables. Table 6 shows the result of estimating a second time-series model, an AR(2) model using the same data as in the analysis shown in Table 5.25 With 358 observations and three parameters, this model has 355 degrees of freedom. Because the degrees of freedom are
almost the same as those for the estimates shown in Table 5, the critical value of the t-statistic at the 0.05 significance level also is almost the same (1.97). If she estimates the equation with two lags, Inflationt = b0 + b1 Inflationt −1 + b2 Inflationt −2 + εt, Miller finds that all three of the coefficients in the regression model (an intercept and the coefficients on two lags of the dependent variable) differ significantly from 0. The bottom portion of Table 6 shows that none of the first four autocorrelations of the residual has a t-statistic greater in absolute value than the critical value of 1.97. Therefore, Miller fails to reject the hypothesis that the individual autocorrelations of the residual equal 0. She concludes that this model is correctly specified because she finds no evidence of serial correlation in the residuals.
TABLE 6 Monthly CPI Inflation at an Annual Rate: AR(2) Model Monthly Observations, March 1984–December 2013
Regression Statistics R-squared 0.2349
Standard error 3.3637 Observations 358 Durbin–Watson 2.0273
Coefficient Standard Error t-Statistic Intercept 1.8953 0.2379 7.9668
Inflationt −1 0.5406 0.0520 10.3962 Inflationt −2 −0.2015 0.0520 −3.8750
Autocorrelations of the Residual Lag Autocorrelation Standard Error t-Statistic 1 −0.0140 0.0529 −0.2647 2 0.0335 0.0529 0.6333 3 −0.0726 0.0529 −1.3724 4 −0.0056 0.0529 −0.1059
Source: US Bureau of Labor Statistics.
1. The analyst selected an AR(2) model because the residuals from the AR(1) model were serially correlated. Suppose that in a given month, inflation had been 4 percent at an annual rate in the previous month and 3 percent in the
month before that. What would be the difference in the analyst forecast of the inflation for that month if she had used an AR(1) model instead of the AR(2) model?
Solution to 1: The AR(1) model shown in Table 5 predicted that inflation in the next month would be 1.5703 + 0.4510(4) = 3.37 percent approximately, whereas the AR(2) model shown in Table 6 predicts that inflation in the next month will be 1.8953 + 0.5406(4) − 0.2015(3) = 3.45 percent approximately. If the analyst had used the incorrect AR(1) model, she would have predicted inflation to be 8 basis points lower (3.37 percent versus 3.45 percent) than using the AR(2) model. This incorrect forecast could have adversely affected the quality of her company’s investment choices.
4.5. Comparing Forecast Model Performance One way to compare the forecast performance of two models is to compare the variance of the forecast errors that the two models make. The model with the smaller forecast error variance will be the more accurate model, and it will also have the smaller standard error of the time-series regression. (This standard error usually is reported directly in the output for the time-series regression.)
In comparing forecast accuracy among models, we must distinguish between in- sample forecast errors and out-of-sample forecast errors. In-sample forecast errors are the residuals from a fitted time-series model. For example, when we estimated a linear trend with raw inflation data from January 1984 to December 2013, the in-sample forecast errors were the residuals from January 1984 to December 2013. If we use this model to predict inflation outside this period, the differences between actual and predicted inflation are out-of-sample forecast errors.
EXAMPLE 7 In-Sample Forecast Comparisons of US CPI Inflation
In Example 6, the analyst compared an AR(1) forecasting model of monthly US inflation with an AR(2) model of monthly US inflation and decided that the AR(2) model was preferable. Table 5 showed that the standard error from the AR(1) model of inflation is 3.4250, and Table 6 showed that the standard error from the AR(2) model is 3.3637. Therefore, the AR(2) model had a lower in- sample forecast error variance than the AR(1) model, which is consistent with our belief that the AR(2) model was preferable. Its standard error is
3.3637/3.4250 = 98.21 percent of the forecast error of the AR(1) model.
Often, we want to compare the forecasting accuracy of different models after the sample period for which they were estimated. We wish to compare the out-of- sample forecast accuracy of the models. Out-of-sample forecast accuracy is important because the future is always out of sample. Although professional forecasters distinguish between out-of-sample and in-sample forecasting performance, many articles that analysts read contain only in-sample forecast evaluations. Analysts should be aware that out-of-sample performance is critical for evaluating a forecasting model’s real-world contribution.
Typically, we compare the out-of-sample forecasting performance of forecasting models by comparing their root mean squared error (RMSE), which is the square root of the average squared error. The model with the smallest RMSE is judged most accurate. The following example illustrates the computation and use of RMSE in comparing forecasting models.
EXAMPLE 8 Out-of-Sample Forecast Comparisons of US CPI Inflation
Suppose we want to compare the forecasting accuracy of the AR(1) and AR(2) models of US inflation estimated over 1984 to 2013, using data on US inflation from January 2014 to September 2014.
Table 7 Out-of-Sample Forecast Error Comparisons: January 2014–September 2014 US CPI Inflation (Annualized)
Date Infl(t) Infl(t−1) Infl(t−2) AR(1)Error Squared Error
AR(2) Error
Squared Error
2014 January 4.5568 −0.1029 −2.4236 3.0329 9.1986 2.2288 4.9675 February 4.5289 4.5568 −0.1029 0.9034 0.8162 0.1494 0.0223 March 8.0077 4.5289 4.5568 4.3949 19.3154 4.5823 20.9978 April 4.0286 8.0077 4.5289 −1.1532 1.3298 −1.2831 1.6463 May 4.2726 4.0286 8.0077 0.8854 0.7839 1.8130 3.2869 June 2.2576 4.2726 4.0286 −1.2397 1.5368 −1.1357 1.2898 July −0.4672 2.2576 4.2726 −3.0557 9.3373 −2.7221 7.4096
August −1.9863 −0.4672 2.2576 −3.3459 11.1949 −3.1741 10.0750
September 0.9068 −1.9863 −0.4672 0.2324 0.0540 −0.0088 0.0001 Average 5.9519 Average 5.5217 RMSE 2.4396 RMSE 2.3498
Note: Any apparent discrepancies between error and squared error results are due to rounding.
Source: US Bureau of Labor Statistics.
For each month from January 2014 to September 2014, the first column of numbers in Table 7 shows the actual annualized inflation rate during the month. The second and third columns show the rate of inflation in the previous two months. The fourth column shows the out-of-sample errors (Actual – Forecast) from the AR(1) model shown in Table 5. The fifth column shows the squared errors from the AR(1) model. The sixth column shows the out-of-sample errors from the AR(2) model shown in Table 6. The final column shows the squared errors from the AR(2) model. The bottom of the table displays the average squared error and the RMSE. According to these measures, the AR(2) model was slightly more accurate than the AR(1) model in its out-of-sample forecasts of inflation from January 2014 to September 2014. The RMSE from the AR(2) model was only 2.3498/2.4396 = 96.32 percent as large as the RMSE from the AR(1) model. Therefore, the AR(2) model was more accurate both in- sample and out of sample. Of course, this was a small sample to use in evaluating out-of-sample forecasting performance. Sometimes, an analyst may have conflicting information about whether to choose an AR(1) or an AR(2) model. We must also consider regression coefficient stability. We will continue the comparison between these two models in the following section.
4.6. Instability of Regression Coefficients One of the important issues an analyst faces in modeling a time series is the sample period to use. The estimates of regression coefficients of the time-series model can change substantially across different sample periods used for estimating the model. Often, the regression coefficient estimates of a time-series model estimated using an earlier sample period can be quite different from those of a model estimated using a later sample period. Similarly, the estimates can be different between models estimated using relatively shorter and longer sample periods. Further, the choice of model for a particular time series can also depend on the sample period. For example, an AR(1) model may be appropriate for the sales of a company in one particular sample period, but an AR(2) model may be necessary for an earlier or later sample period (or for a longer or shorter sample period). Thus the choice of a
sample period is an important decision in modeling a financial time series.
Unfortunately, there is usually no clear-cut basis in economic or financial theory for determining whether to use data from a longer or shorter sample period to estimate a time-series model. We can get some guidance, however, if we remember that our models are valid only for covariance-stationary time series. For example, we should not combine data from a period when exchange rates were fixed with data from a period when exchange rates were floating. The exchange rates in these two periods would not likely have the same variance because exchange rates are usually much more volatile under a floating-rate regime than when rates are fixed. Similarly, many US analysts consider it inappropriate to model US inflation or interest-rate behavior since the 1960s as a part of one sample period, because the Federal Reserve had distinct policy regimes during this period. A simple way to determine appropriate samples for time-series estimation is to look at graphs of the data to see if the time series looks stationary before estimation begins. If we know that a government policy changed on a specific date, we might also test whether the time-series relation was the same before and after that date.
In the following example, we illustrate how the choice of a longer versus a shorter period can affect the decision of whether to use, for example, a first-or second- order time-series model. We then show how the choice of the time-series model (and the associated regression coefficients) affects our forecast. Finally, we discuss which sample period, and accordingly which model and corresponding forecast, is appropriate for the time series analyzed in the example.
EXAMPLE 9 Instability in Time-Series Models of US Inflation
In Example 6, analyst Lisette Miller concluded that US CPI inflation should be modeled as an AR(2) time series. A colleague examined her results and questioned estimating one time-series model for inflation in the United States since 1984, given that the Federal Reserve responded aggressively to the financial crisis that emerged in 2007. He argues that the inflation time series from 1984 to 2013 has two regimes or underlying models generating the time series: one running from 1984 through 2006, and another starting in 2007. Therefore, the colleague suggests that Miller estimate a new time-series model for US inflation starting in 2007. Because of his suggestion, Miller first estimates an AR(1) model for inflation using data for a shorter sample period from 2007 to 2013. Table 8 shows her AR(1) estimates.
Table 8 Autoregression: AR(1) Model Monthly CPI Inflation at an Annual Rate, February 2007–December 2013
Regression Statistics R-squared 0.3070
Standard error 4.4749 Observations 83 Durbin–Watson 1.8164
Coefficient Standard Error t-Statistic Intercept 0.9585 0.5337 1.7960
Inflationt −1 0.5544 0.0926 5.9870 Autocorrelations of the Residual
Lag Autocorrelation Standard Error t-Statistic 1 0.0878 0.1098 0.7999 2 −0.0091 0.1098 −0.0829 3 −0.0355 0.1098 −0.3234 4 0.0020 0.1098 0.0182
Source: US Bureau of Labor Statistics.
The bottom part of Table 8 shows that the first four autocorrelations of the residuals from the AR(1) model are quite small. None of these autocorrelations has a t-statistic larger than 1.99, the critical value for significance. Consequently, Miller cannot reject the null hypothesis that the residuals are serially uncorrelated. The AR(1) model is correctly specified for the sample period from 2007 to 2013, so there is no need to estimate the AR(2) model. This conclusion is very different from that reached in Example 6 using data from 1984 to 2013. In that example, Miller initially rejected the AR(1) model because its residuals exhibited serial correlation. When she used a larger sample, an AR(2) model initially appeared to fit the data much better than did an AR(1) model.
How deeply does our choice of sample period affect our forecast of future inflation? Suppose that in a given month, inflation was 4 percent at an annual rate, and the month before that it was 3 percent. The AR(1) model shown in Table 8 predicts that inflation in the next month will be 0.9585 + 0.5544(4) = approximately 3.17 percent. Therefore, the forecast of the next month’s inflation using the 2007 to 2013 sample is 3.17 percent. Remember from the analysis following Example 6 that the AR(2) model for the 1984 to 2013 sample predicts inflation of 3.45 percent in the next month. Thus, using the
correctly specified model for the shorter sample produces an inflation forecast 0.28 percentage points below the forecast made from the correctly specified model for the longer sample period. Such a difference might substantially affect a particular investment decision.
Which model is correct? Figure 9 suggests an answer. Monthly US inflation was so much more volatile during the latter part of the study period than in the earlier years that inflation is probably not a covariance-stationary time series from 1984 to 2013. Therefore, we can reasonably believe that the data have more than one regime and Miller should estimate a separate model for inflation from 2007 to 2013, as shown above. In fact, the standard deviation of annualized monthly inflation rates is just 3.24 percent for the period of 1984 to 2006 but 5.32 percent for the period of 2007 to 2013. As the example shows, experience (such as knowledge of government policy changes) and judgment play a vital role in determining how to model a time series. Simply relying on autocorrelations of the residuals from a time-series model cannot tell us the correct sample period for our analysis.
FIGURE 9 Monthly CPI Inflation
Source: US Bureau of Labor Statistics.
5. Random Walks and Unit Roots So far, we have examined those time series in which the time series has a tendency to revert to its mean level as the change in a variable from one period to the next follows a mean-reverting pattern. In contrast, there are many financial time series in which the changes follow a random pattern. We discuss these “random walks” in the following section.
5.1. Random Walks A random walk is one of the most widely studied time-series models for financial data. A random walk is a time series in which the value of the series in one period is the value of the series in the previous period plus an unpredictable random error. A random walk can be described by the following equation:
(8)
Equation 8 means that the time series xt is in every period equal to its value in the previous period plus an error term, εt, that has constant variance and is uncorrelated with the error term in previous periods. Note two important points. First, this equation is a special case of an AR(1) model with b0 = 0 and b1 = 1.26 Second, the expected value of εt is zero. Therefore, the best forecast of xt that can be made in period t − 1 is xt−1. In fact, in this model, xt −1 is the best forecast of x in every period after t − 1.
Random walks are quite common in financial time series. For example, many studies have tested and found that currency exchange rates follow a random walk. Consistent with the second point made above, some studies have found that sophisticated exchange rate forecasting models cannot outperform forecasts made using the random walk model, and that the best forecast of the future exchange rate is the current exchange rate.
Unfortunately, we cannot use the regression methods we have discussed so far to estimate an AR(1) model on a time series that is actually a random walk. To see why this is so, we must determine why a random walk has no finite mean-reverting level or finite variance. Recall that if xt is at its mean-reverting level, then xt = b0 + b1xt, or xt = b0/(1 − b1). In a random walk, however, b0 = 0 and b1 = 1, so b0/(1 − b1) = 0/0. Therefore, a random walk has an undefined mean-reverting level.
What is the variance of a random walk? Suppose that in Period 1, the value of x1 is 0. Then we know that x2 = 0 + ε2. Therefore, the variance of x2 = Var(ε2) = σ2. Now x3 = x2 + ε3 = ε2 + ε3. Because the error term in each period is assumed to be uncorrelated with the error terms in all other periods, the variance of x3 = Var(ε2) + Var(ε3) = 2σ2. By a similar argument, we can show that for any period t, the variance of xt = (t − 1)σ2. But this means that as t grows large, the variance of xt grows without an upper bound: It approaches infinity. This lack of upper bound, in turn, means that a random walk is not a covariance-stationary time series, because a covariance-stationary time series must have a finite variance.
What is the practical implication of these issues? We cannot use standard regression analysis on a time series that is a random walk. We can, however, attempt to convert the data to a covariance-stationary time series if we suspect that the time series is a random walk. In statistical terms, we can difference it.
We difference a time series by creating a new time series, say yt, that in each period is equal to the difference between xt and xt −1. This transformation is called first- differencing because it subtracts the value of the time series in the first prior period from the current value of the time series. Sometimes the first difference of xt is written as Δxt = xt − xt −1. Note that the first difference of the random walk in Equation 8 yields
The expected value of εt is 0. Therefore, the best forecast of yt that can be made in period t − 1 is 0. This implies that the best forecast is that there will be no change in the value of the current time series, xt −1.
The first-differenced variable, yt, is covariance stationary. How is this so? First, note that this model (yt = εt) is an AR(1) model with b0 = 0 and b1 = 0. We can compute the mean-reverting level of the first-differenced model as b0/(1 − b1) = 0/1 = 0. Therefore, a first-differenced random walk has a mean-reverting level of 0. Note also that the variance of yt in each period is Var(εt) = σ2. Because the variance and the mean of yt are constant and finite in each period, yt is a covariance- stationary time series and we can model it using linear regression.27 Of course, modeling the first-differenced series with an AR(1) model does not help us predict the future, as b0 = 0 and b1 = 0. We simply conclude that the original time series is, in fact, a random walk.
Had we tried to estimate an AR(1) model for a time series that was a random walk, our statistical conclusions would have been incorrect because AR models cannot be used to estimate random walks or any time series that is not covariance stationary. The following example illustrates this issue with exchange rates.
EXAMPLE 10 The Yen/US Dollar Exchange Rate
Financial analysts often assume that exchange rates are random walks. Consider an AR(1) model for the Japanese yen/US dollar exchange rate (JPY/USD). Table 9 shows the results of estimating the model using month-end observations from January 1980 through December 2013.
The results in Table 9 suggest that the yen/US dollar exchange rate is a random walk because the estimated intercept does not appear to be significantly different from 0 and the estimated coefficient on the first lag of the exchange rate is very close to 1. Can we use the t-statistics in Table 9 to test whether the exchange rate is a random walk? Unfortunately, no, because the standard errors in an AR model are invalid if the model is estimated using a data series that is a random walk (remember, a random walk is not covariance stationary). If the exchange rate is, in fact, a random walk, we might come to an incorrect conclusion based on faulty statistical tests and then invest incorrectly. We can use a test presented in the next section to test whether the time-series is a random walk.
Suppose the exchange rate is a random walk, as we now suspect. If so, the first- differenced series, yt = xt − xt −1, will be covariance stationary. We present the results from estimating yt = b0 + b1yt −1 + εt in Table 10. If the exchange rate is a random walk, then b0 = 0 and b1 = 0 and the error term will not be serially correlated.
TABLE 9 Yen/US Dollar Exchange Rate: AR(1) Model Month-End Observations, January 1980–December 2013
Regression Statistics R-squared 0.9902
Standard error 4.9437 Observations 408 Durbin–Watson 1.8981
Coefficient Standard Error t-Statistic Intercept 0.9958 0.7125 1.3976
JPY/USDt −1 0.9903 0.0049 202.1020 Autocorrelations of the Residual
Lag Autocorrelation Standard Error t-Statistic 1 0.0687 0.0495 1.3879 2 0.0384 0.0495 0.7758 3 0.0686 0.0495 1.3859 4 0.0407 0.0495 0.8222
Source: US Federal Reserve Board of Governors.
TABLE 10 First-Differenced Yen/US Dollar Exchange Rate: AR(1) Model Month-End Observations, January 1980–December 2013
Regression Statistics R-squared 0.0026
Standard error 4.9611 Observations 408 Durbin–Watson 2.0010
Coefficient Standard Error t-Statistic Intercept −0.3128 0.2463 −1.2700
JPY/USDt −1 − JPY/USDt −2 0.0506 0.0494 1.0243 Autocorrelations of the Residual
Lag Autocorrelation Standard Error t-Statistic 1 0.0193 0.0495 0.3899 2 0.0345 0.0495 0.6970 3 0.0680 0.0495 1.3737 4 0.0399 0.0495 0.8061
Source: US Federal Reserve Board of Governors.
In Table 10, neither the intercept nor the coefficient on the first lag of the first- differenced exchange rate differs significantly from 0, and no residual
autocorrelations differ significantly from 0.28 These findings are consistent with the yen/US dollar exchange rate being a random walk.
We have concluded that the differenced regression is the model to choose. Now we can see that we would have been seriously misled if we had based our model choice on an R2 comparison. In Table 9, the R2 is 0.9902, whereas in Table 10 the R2 is 0.0026. How can this be, if we just concluded that the model in Table 10 is the one that we should use? In Table 9, the R2 measures how well the exchange rate in one period predicts the exchange rate in the next period. If the exchange rate is a random walk, its current value will be an extremely good predictor of its value in the next period, and thus the R2 will be extremely high. At the same time, if the exchange rate is a random walk, then changes in the exchange rate should be completely unpredictable. Table 10 estimates whether changes in the exchange rate from one month to the next can be predicted by changes in the exchange rate over the previous month. If they cannot be predicted, the R2 in Table 10 should be very low. In fact, it is low (0.0026). This comparison provides a good example of the general rule that we cannot necessarily choose which model is correct solely by comparing the R2 from the two models.
The exchange rate is a random walk, and changes in a random walk are by definition unpredictable. Therefore, we cannot profit from an investment strategy that predicts changes in the exchange rate.
To this point, we have discussed only simple random walks; that is, random walks without drift. In a random walk without drift, the best predictor of the time series in the next period is its current value. A random walk with drift, however, should increase or decrease by a constant amount in each period. The equation describing a random walk with drift is a special case of the AR(1) model:
(9)
A random walk with drift has b0 | 0 compared to a simple random walk, which has b0 = 0.
We have already seen that b1 = 1 implies an undefined mean-reversion level and thus nonstationarity. Consequently, we cannot use an AR model to analyze a time series that is a random walk with drift until we transform the time series by taking first differences. If we first-difference Equation 9, the result is yt = xt − xt −1, yt = b0 + εt,
b0 | 0.
5.2. The Unit Root Test of Nonstationarity In this section, we discuss how to use random walk concepts to determine whether a time series is covariance stationary. This approach focuses on the slope coefficient in the random-walk-with-drift case of an AR(1) model in contrast with the traditional autocorrelation approach which we discuss first.
The examination of the autocorrelations of a time series at various lags is a well- known prescription for inferring whether or not a time series is stationary. Typically, for a stationary time series, either autocorrelations at all lags are statistically indistinguishable from zero, or the autocorrelations drop off rapidly to zero as the number of lags becomes large. Conversely, the autocorrelations of a nonstationary time series do not exhibit those characteristics. However, this approach is less definite than a currently more popular test for nonstationarity known as the Dickey–Fuller test for a unit root.
We can explain what is known as the unit root problem in the context of an AR(1) model. If a time series comes from an AR(1) model, then to be covariance stationary the absolute value of the lag coefficient, b1, must be less than 1.0. We could not rely on the statistical results of an AR(1) model if the absolute value of the lag coefficient were greater than or equal to 1.0 because the time series would not be covariance stationary. If the lag coefficient is equal to 1.0, the time series has a unit root: it is a random walk and is not covariance stationary.29 By definition, all random walks, with or without a drift term, have unit roots.
How do we test for unit roots in a time series? If we believed that a time series, xt, was a random walk with drift, it would be tempting to estimate the parameters of the AR(1) model xt = b0 + b1xt −1 + εt using linear regression and conduct a t-test of the hypothesis that b1 = 1. Unfortunately, if b1 = 1, then xt is not covariance stationary and the t-value of the estimated coefficient, , does not actually follow the t- distribution; consequently, a t-test would be invalid.
Dickey and Fuller (1979) developed a regression-based unit root test based on a transformed version of the AR(1) model xt = b0 + b1xt −1 + εt. Subtracting xt−1 from both sides of the AR(1) model produces
or
(10)
where g1 = (b1 − 1). If b1 = 1, then g1 = 0 and thus a test of g1 = 0 is a test of b1 = 1. If there is a unit root in the AR(1) model, then g1 will be 0 in a regression where the dependent variable is the first difference of the time series and the independent variable is the first lag of the time series. The null hypothesis of the Dickey–Fuller test is H0: g1 = 0—that is, that the time series has a unit root and is nonstationary— and the alternative hypothesis is Ha: g1 < 0, that the time series does not have a unit root and is stationary.
To conduct the test, one calculates a t-statistic in the conventional manner for but instead of using conventional critical values for a t-test, one uses a revised set of values computed by Dickey and Fuller; the revised set of critical values are larger in absolute value than the conventional critical values. A number of software packages incorporate Dickey–Fuller tests.30
EXAMPLE 11 AstraZeneca’s Quarterly Sales (1)
In January 2012, equity analyst Aron Berglin is building a time-series model for the quarterly sales of AstraZeneca, a British-Swedish biopharmaceutical company headquartered in London, UK He is using AstraZeneca’s quarterly sales in US dollars during January 2000 to December 2011 and any lagged sales data that he may need prior to 2000 to build this model. He finds that a log-linear trend model seems better suited for modeling AstraZeneca’s sales than does a linear trend model. However, the Durbin–Watson statistic from the log-linear regression is just 0.7064, which causes him to reject the hypothesis that the errors in the regression are serially uncorrelated. He concludes that he cannot model the log of AstraZeneca’s quarterly sales using only a time-trend line. He decides to model the log of AstraZeneca’s quarterly sales using an AR(1) model. He uses ln Salest = b0 + b1 ln Salest −1 + εt.
Before he estimates this regression, the analyst should use the Dickey–Fuller test to determine whether there is a unit root in the log of AstraZeneca’s quarterly sales. If he uses the sample of quarterly data on AstraZeneca’s sales from the first quarter of 2000 through the fourth quarter of 2011, takes the natural log of each observation, and computes the Dickey–Fuller t-test statistic, the value of that statistic might cause him to fail to reject the null hypothesis that there is a unit root in the log of AstraZeneca’s quarterly sales.
If a time series appears to have a unit root, how should we model it? One method
that is often successful is to model the first-differenced series as an autoregressive time series. The following example demonstrates this method.
EXAMPLE 12 AstraZeneca’s Quarterly Sales (2)
The plot of the log of AstraZeneca’s quarterly sales is shown as Figure 10. By looking at the plot, Berglin is convinced that the log of quarterly sales is not covariance stationary (that it has a unit root).
So he creates a new series, yt, that is the first difference of the log of AstraZeneca’s quarterly sales. Figure 11 shows that series.
Berglin compares Figure 11 to Figure 10 and notices that first-differencing the log of AstraZeneca’s quarterly sales eliminates the strong upward trend that was present in the log of AstraZeneca’s sales. Because the first-differenced series has no strong trend, Berglin is better off assuming that the differenced series is covariance stationary rather than assuming that AstraZeneca’s sales or the log of AstraZeneca’s sales is a covariance-stationary time series.
Now suppose Berglin decides to model the new series using an AR(1) model. Berglin uses ln (Salest) – ln (Salest −1) = b0 + b1[ln (Salest −1) – ln (Salest −2)] + εt. Table 11 shows the results of that regression.
FIGURE 10 Log of AstraZeneca’s Quarterly Sales
Source: Compustat.
FIGURE 11 Log Difference, AstraZeneca’s Quarterly Sales
Source: Compustat.
TABLE 11 Log Differenced Sales: AR(1) Model of AstraZeneca Quarterly Observations, January 2000–December 2011
Regression Statistics R-squared 0.3005
Standard error 0.0475 Observations 48 Durbin–Watson 1.6874
Coefficient Standard Error t-Statistic Intercept 0.0222 0.0071 3.1268
ln Salest −1 − ln Salest −2 ln Salest −2 −0.5493 0.1236 −4.4442 Autocorrelations of the Residual
Lag Autocorrelation Standard Error t-Statistic 1 0.2809 0.1443 1.9466 2 −0.0466 0.1443 −0.3229 3 0.0081 0.1443 0.0561 4 0.2647 0.1443 1.8344
Source: Compustat.
The lower part of Table 11 suggests that the first four autocorrelations of residuals in this model are not statistically significant. With 48 observations and two parameters, this model has 46 degrees of freedom. The critical value for a t-statistic in this model is above 2.0 at the 0.05 significance level. None of the t-statistics for these autocorrelations has an absolute value larger than 2.0. Therefore, we fail to reject the null hypotheses that each of these autocorrelations is equal to 0 and conclude instead that no significant autocorrelation is present in the residuals.
This result suggests that the model is well specified and that we could use the estimates. Both the intercept ( = 0.0222) and the coefficient ( = −0.5493) on the first lag of the new first-differenced series are statistically significant.
1. Explain how to interpret the estimated coefficients in the model.
Solution to 1: The value of the intercept (0.0222) implies that if sales have not changed in the current quarter (yt = ln Salest – ln Salest −1 = 0), sales will grow by 2.22 percent next quarter.31 If sales have changed during this quarter, however, the model predicts that sales will grow by 2.22 percent minus 0.5493 times the sales growth in this quarter.
2. Astra Zeneca’s sales in the third and fourth quarters of 2011 were $8,405 million and $8,872 million, respectively. If we use the above model soon after the end of the fourth quarter of 2011, what will be the predicted value of AstraZeneca’s sales for the first quarter of 2012?
Solution to 2: Let us say that t is the fourth quarter of 2011, so t − 1 is the third quarter of 2011 and t + 1 is the first quarter of 2012. Then we would have to compute = 0.0222 − 0.5493yt. To compute , we need to know yt = ln Salest − ln Salest−1. In the third quarter of 2011, AstraZeneca’s sales were $8,405 million, so ln (Salest −1) = ln 8,405 =9.0366. In the fourth quarter of 2011, AstraZeneca’s sales were $8,872 million, so ln (Salest) = ln 8,872 =
9.0907. Thus yt = 9.0907 − 9.0366 = 0.0541. Therefore, = 0.0222 −
0.5493(0.0541) = −0.0075. If = −0.0075, then −0.0075 = ln (Salest +1) − ln (Salest) = ln (Salest +1/Salest). If we exponentiate both sides of this equation, the result is
Thus, based on fourth quarter sales for 2011, this model would have predicted that AstraZeneca’s sales in the first quarter of 2012 would be $8,805 million. This sales forecast might have affected our decision to buy AstraZeneca’s stock at the time.
6. Moving-Average Time-Series Models So far, many of the forecasting models we have used have been autoregressive models. Because most financial time series have the qualities of an autoregressive process, autoregressive time-series models are probably the most frequently used time-series models in financial forecasting. Some financial time series, however, seem to follow more closely another kind of time-series model called a moving- average model. For example, as we will see later, returns on the S&P BSE 100 Index can be better modeled as a moving-average process than as an autoregressive process.
In this section, we present the fundamentals of moving-average models so that you can ask the right questions when considering their use. We first discuss how to smooth past values with a moving average and then how to forecast a time series using a moving-average model. Even though both methods include the words “moving average” in the name, they are very different.
6.1. Smoothing Past Values with an n-Period Moving Average Suppose you are analyzing the long-term trend in the past sales of a company. In order to focus on the trend, you may find it useful to remove short-term fluctuations or noise by smoothing out the time series of sales. One technique to smooth out period-to-period fluctuations in the value of a time series is an n-period moving average. An n-period moving average of the current and past n − 1 values of a time series, xt, is calculated as
(11)
The following example demonstrates how to compute a moving average of AstraZeneca’s quarterly sales.
EXAMPLE 13 AstraZeneca’s Quarterly Sales (3)
Suppose we want to compute the four-quarter moving average of AstraZeneca’s sales as of the beginning of the first quarter of 2012. AstraZeneca’s sales in the previous four quarters were 1Q:2011, $8,490 million; 2Q:2011, $8,601 million; 3Q:2011, $8,405 million; and 4Q:2011, $8,872 million. The four-quarter moving average of sales as of the beginning
of the first quarter of 2012 is thus (8,490 + 8,601 + 8,405 + 8,872)/4 = $8,592 million.
We often plot the moving average of a series with large fluctuations to help discern any patterns in the data. Figure 12 shows monthly retail sales for the United States from January 1995 to December 2013, along with a 12-month moving average of the data.32
FIGURE 12 Monthly US Real Retail Sales and 12-Month Moving Average of Retail Sales
Source: US Department of Commerce, Census Bureau.
As Figure 12 shows, each year has a very strong peak in retail sales (December) followed by a sharp drop in sales (January). Because of the extreme seasonality in the data, a 12-month moving average can help us focus on the long-term movements in retail sales instead of seasonal fluctuations. Note that the moving average does not have the sharp seasonal fluctuations of the original retail sales data. Rather, the moving average of retail sales grows steadily, for example, from 1995 through the second half of 2008, then declines for about a year, and grows steadily thereafter. We can see that trend more easily by looking at a 12-month moving average than by looking at the time series itself.
Figure 13 shows monthly Europe Brent Crude Oil spot prices along with a 12-
month moving average of oil prices. Although these data do not have the same sharp regular seasonality displayed in the retail sales data in Figure 12, the moving average smooths out the monthly fluctuations in oil prices to show the longer-term movements.
Figure 13 also shows one weakness with a moving average: It always lags large movements in the actual data. For example, when oil prices rose quickly in late 2007 and the first half of 2008, the moving average rose only gradually. When oil prices fell sharply toward the end of 2008, the moving average also lagged. Consequently, a simple moving average of the recent past, though often useful in smoothing out a time series, may not be the best predictor of the future. A main reason for this is that a simple moving average gives equal weight to all the periods in the moving average. In order to forecast the future values of a time series, it is often better to use a more sophisticated moving-average time-series model. We discuss such models below.
FIGURE 13 Monthly Europe Brent Crude Oil Price and 12-Month Moving Average of Prices
Source: US Energy Information Administration.
6.2. Moving-Average Time-Series Models for Forecasting Suppose that a time series, xt, is consistent with the following model:
(12)
This equation is called a moving-average model of order 1, or simply an MA(1) model. Theta (θ) is the parameter of the MA(1) model.33
Equation 12 is a moving-average model because in each period, xt is a moving average of εt and εt −1, two uncorrelated random variables that each have an expected value of zero. Unlike the simple moving-average model of Equation 11, this moving-average model places different weights on the two terms in the moving average (1 on εt, and θ on εt −1).
We can see if a time series fits an MA(1) model by looking at its autocorrelations to determine whether xt is correlated only with its preceding and following values. First, we examine the variance of xt in Equation 12 and its first two autocorrelations. Because the expected value of xt is 0 in all periods and εt is uncorrelated with its own past values, the first autocorrelation is not equal to 0, but the second and higher autocorrelations are equal to 0. Further analysis shows that all autocorrelations except for the first will be equal to 0 in an MA(1) model. Thus for an MA(1) process, any value xt is correlated with xt −1 and xt +1 but with no other time-series values; we could say that an MA(1) model has a memory of one period.
Of course, an MA(1) model is not the most complex moving-average model. A qth order moving-average model, denoted MA(q) and with varying weights on lagged terms, can be written as
(13)
How can we tell whether an MA(q) model fits a time series? We examine the autocorrelations. For an MA(q) model, the first q autocorrelations will be significantly different from 0, and all autocorrelations beyond that will be equal to 0; an MA(q) model has a memory of q periods. This result is critical for choosing the right value of q for an MA model. We discussed this result above for the specific case of q = 1 that all autocorrelations except for the first will be equal to 0 in an MA(1) model.
How can we distinguish an autoregressive time series from a moving-average time series? Once again, we do so by examining the autocorrelations of the time series itself. The autocorrelations of most autoregressive time series start large and
decline gradually, whereas the autocorrelations of an MA(q) time series suddenly drop to 0 after the first q autocorrelations. We are unlikely to know in advance whether a time series is autoregressive or moving average. Therefore, the autocorrelations give us our best clue about how to model the time series. Most time series, however, are best modeled with an autoregressive model.
EXAMPLE 14 A Time-Series Model for Monthly Returns on the S&P BSE 100 Index
The S&P BSE 100 index is designed to reflect the performance of India’s top 100 large-cap companies listed on the BSE Ltd. (formerly Bombay Stock Exchange). Are monthly returns on the S&P BSE 100 index autocorrelated? If so, we may be able to devise an investment strategy to exploit the autocorrelation. What is an appropriate time-series model for S&P BSE 100 monthly returns?
Table 12 shows the first six autocorrelations of returns to the S&P BSE 100 using monthly data from January 2000 through December 2013. Note that all of the autocorrelations are quite small. Do they reach significance? With 168 observations, the critical value for a t-statistic in this model is about 1.98 at the 0.05 significance level. None of the autocorrelations has a t-statistic larger in absolute value than the critical value of 1.98. Consequently, we fail to reject the null hypothesis that those autocorrelations, individually, differ significantly from 0.
TABLE 12 Annualized Monthly Returns to the S&P BSE 100 January 2000– December 2013
Autocorrelations Lag Autocorrelation Standard Error t-Statistic 1 0.1103 0.0772 1.4288 2 −0.0045 0.0772 −0.0583 3 0.0327 0.0772 0.4236 4 0.0370 0.0772 0.4793 5 −0.0218 0.0772 −0.2824 6 0.0191 0.0772 0.2474
Observations 168
Source: BSE Ltd.
If returns on the S&P BSE 100 were an MA(q) time series, then the first q autocorrelations would differ significantly from 0. None of the autocorrelations is statistically significant, however, so returns to the S&P BSE 100 appear to come from an MA(0) time series. An MA(0) time series in which we allow the mean to be nonzero takes the following form:34
(14)
which means that the time series is not predictable. This result should not be too surprising, as most research suggests that short-term returns to stock indexes are difficult to predict.
We can see from this example how examining the autocorrelations allowed us to choose between the AR and MA models. If returns to the S&P BSE 100 had come from an AR(1) time series, the first autocorrelation would have differed significantly from 0 and the autocorrelations would have declined gradually. Not even the first autocorrelation is significantly different from 0, however. Therefore, we can be sure that returns to the S&P BSE 100 do not come from an AR(1) model—or from any higher-order AR model, for that matter. This finding is consistent with our conclusion that the S&P BSE 100 series is MA(0).
7. Seasonality in Time-Series Models As we analyze the results of the time-series models in this reading, we encounter complications. One common complication is significant seasonality, a case in which the series shows regular patterns of movement within the year. At first glance, seasonality might appear to rule out using autoregressive time-series models. After all, autocorrelations will differ by season. This problem can often be solved, however, by using seasonal lags in an autoregressive model.
A seasonal lag is usually the value of the time series one year before the current period, included as an extra term in an autoregressive model. Suppose, for example, that we model a particular quarterly time series using an AR(1) model, xt = b0 + b1xt −1 + εt. If the time series had significant seasonality, this model would not be correctly specified. The seasonality would be easy to detect because the seasonal autocorrelation (in the case of quarterly data, the fourth autocorrelation) of the error term would differ significantly from 0. Suppose this quarterly model has significant seasonality. In this case, we might include a seasonal lag in the autoregressive model and estimate
(15)
to test whether including the seasonal lag would eliminate statistically significant autocorrelation in the error term.
In Examples 15 and 16, we illustrate how to test and adjust for seasonality in a time- series model. We also illustrate how to compute a forecast using an autoregressive model with a seasonal lag.
EXAMPLE 15 Seasonality in Sales at Starbucks
Earlier, we concluded that we could not model the log of Starbucks’ quarterly sales using only a time-trend line (as shown in Example 3) because the Durbin– Watson statistic from the regression provided evidence of positive serial correlation in the error term. Based on methods presented in this reading, we might next investigate using the first difference of log sales to remove an exponential trend from the data to obtain a covariance stationary time series.
Using quarterly data from the first quarter of 1995 to the last quarter of 2012, we estimate the following AR(1) model using ordinary least squares: (ln Salest
− ln Salest −1) = b0 + b1(ln Salest −1 − ln Salest −2) + εt. Table 13 shows the results of the regression.
TABLE 13 Log Differenced Sales: AR(1) Model Starbucks, Quarterly Observations, 1995–2013
Regression Statistics R-squared 0.1548
Standard error 0.0762 Observations 74 Durbin–Watson 1.9165
Coefficient Standard Error t-Statistic Intercept 0.0669 0.0101 6.6238
ln Salest −1 − ln Salest −2 −0.3813 0.1050 −3.6314 Autocorrelations of the Residual
Lag Autocorrelation Standard Error t-Statistic 1 −0.0141 0.1162 −0.1213 2 −0.0390 0.1162 −0.3356 3 0.0294 0.1162 0.2530 4 0.7667 0.1162 6.5981
Source: Compustat.
The first thing to note in Table 13 is the strong seasonal autocorrelation of the residuals. The bottom portion of the table shows that the fourth autocorrelation has a value of 0.7667 and a t-statistic of 6.60. With 74 observations and two parameters, this model has 72 degrees of freedom.35 The critical value for a t- statistic is about 1.99 at the 0.05 significance level. Given this value of the t- statistic, we must reject the null hypothesis that the fourth autocorrelation is equal to 0 because the t-statistic is larger than the critical value of 1.99.
In this model, the fourth autocorrelation is the seasonal autocorrelation because this AR(1) model is estimated with quarterly data. Table 13 shows the strong and statistically significant seasonal autocorrelation that occurs when a time series with strong seasonality is modeled without taking the seasonality into account. Therefore, the AR(1) model is misspecified, and we should not use it for forecasting.
Suppose we decide to use an autoregressive model with a seasonal lag because
of the seasonal autocorrelation. We are modeling quarterly data, so we estimate Equation 15: (ln Salest − ln Salest −1) = b0 + b1(ln Salest −1 − ln Salest −2) + b2(ln Salest −4 − ln Salest−5) + εt. Adding the seasonal difference ln Salest −4 − ln Salest −5 is an attempt to remove a consistent quarterly pattern in the data and could also eliminate a seasonal nonstationarity if one existed. The estimates of this equation appear in Table 14.
TABLE 14 Log Differenced Sales: AR(1) Model with Seasonal Lag Starbucks, Quarterly Observations, 1995–2013
Regression Statistics R-squared 0.8163
Standard error 0.03405 Observations 71 Durbin–Watson 2.0791
Coefficient Standard Error t-Statistic Intercept 0.0084 0.0059 1.4237
ln Salest −1 − ln Salest −2 −0.0602 0.0540 −1.1148 ln Salest −4 − ln Salest −5 0.8048 0.0524 15.3588
Autocorrelations of the Residual Lag Autocorrelation Standard Error t-Statistic 1 −0.0441 0.1187 −0.3715 2 0.0675 0.1187 0.5687 3 0.0749 0.1187 0.6310 4 −0.2091 0.1187 −1.7616
Source: Compustat.
Note the autocorrelations of the residual shown at the bottom of Table 14. When we include a seasonal lag in the regression, none of the t-statistics on the first four autocorrelations remains significant.
Now that we know that the residuals of this model do not have significant serial correlation, we can assume that the model is correctly specified. How can we interpret the coefficients in this model? To predict the current quarter ’s sales growth at Starbucks, we need to know two things: sales growth in the previous quarter and sales growth four quarters ago. If sales remained constant in each
of those two quarters, the model in Table 14 predicts that sales will grow by 0.0084 (0.84 percent) in the current quarter. If sales grew by 1 percent last quarter and by 2 percent four quarters ago, then the model predicts that sales growth this quarter will be 0.0084 − 0.0602(0.01) + 0.8048(0.02) = 0.0239 or 2.39 percent.36 Notice also that the R2 in the model with the seasonal lag (0.8163 in Table 14) was more than five times higher than the R2 in the model without the seasonal lag (0.1548 in Table 13). Again, the seasonal lag model does a much better job of explaining the data.
EXAMPLE 16 Retail Sales Growth
We want to predict the growth in monthly retail sales of Canadian furniture and home furnishing stores so that we can decide whether to recommend the shares of these stores. We decide to use non-seasonally adjusted data on retail sales. To begin with, we estimate an AR(1) model with observations on the annualized monthly growth in retail sales from January 1995 to December 2012. We estimate the following equation: Sales growtht = b0 + b1 Sales growtht −1 + εt. Table 15 shows the results from this model.
The autocorrelations of the residuals from this model, shown at the bottom of Table 15,indicate that seasonality is extremely significant in this model. With 216 observations and two parameters, this model has 214 degrees of freedom. At the 0.05 significance level, the critical value for a t-statistic is about 1.97. The 12th-lag autocorrelation (the seasonal autocorrelation, because we are using monthly data) has a value of 0.7620 and a t-statistic of 11.21. The t- statistic on this autocorrelation is larger than the critical value (1.97) implying that we can reject the null hypothesis that the 12th autocorrelation is 0. Note also that many of the other t-statistics for autocorrelations shown in the table differ significantly from 0. Consequently, the model shown in Table 15 is misspecified, so we cannot rely on it to forecast sales growth.
Suppose we add the seasonal lag of sales growth (the 12th lag) to the AR(1) model to estimate the equation Sales growtht = b0 + b1(Sales growtht −1) + b2(Sales growtht −12) + εt.37 Table 16 presents the results of estimating this equation. The estimated value of the seasonal autocorrelation (the 12th autocorrelation) has fallen to −0.1168. None of the first 12 autocorrelations has a t-statistic with an absolute value greater than the critical value of 1.97 at the 0.05 significance level. We can conclude that there is no significant serial correlation in the residuals from this model. Because we can reasonably
believe that the model is correctly specified, we can use it to predict retail sales growth. Note that the R2 in Table 16 is 0.6724, much larger than the R2 in Table 15 (computed by the model without the seasonal lag).
TABLE 15 Monthly Retail Sales Growth of Canadian Furniture and Home Furnishing Stores: AR(1) Model January 1995–December 2012
Regression Statistics R-squared 0.0509
Standard error 1.8198 Observations 216 Durbin–Watson 2.0956
Coefficient Standard Error t-Statistic Intercept 1.0518 0.1365 7.7055
Sales growtht −1 −0.2252 0.0665 −3.3865 Autocorrelations of the Residual
Lag Autocorrelation Standard Error t-Statistic 1 −0.0109 0.0680 −0.1603 2 −0.1949 0.0680 −2.8662 3 0.1173 0.0680 1.7250 4 −0.0756 0.0680 −1.1118 5 −0.1270 0.0680 −1.8676 6 −0.1384 0.0680 −2.0353 7 −0.1374 0.0680 −2.0206 8 −0.0325 0.0680 −0.4779 9 0.1207 0.0680 1.7750 10 −0.2197 0.0680 −3.2309 11 −0.0342 0.0680 −0.5029 12 0.7620 0.0680 11.2059
Source: Statistics Canada (Government of Canada).
How can we interpret the coefficients in the model? To predict growth in retail sales in this month, we need to know last month’s retail sales growth and retail sales growth 12 months ago. If retail sales remained constant both last month and 12 months ago, the model in Table 16 predicts that retail sales will grow at
an annual rate of about 23.7 percent this month. If retail sales grew at an annual rate of 10 percent last month and at an annual rate of 5 percent 12 months ago, the model in Table 16 predicts that retail sales will grow in the current month at an annual rate of 0.2371 − 0.0792(0.10) + 0.7798(0.05) = 0.2682 or 26.8 percent.
TABLE 16 Monthly Retail Sales Growth of Canadian Furniture and Home Furnishing Stores: AR(1) Model with Seasonal Lag January 1995–December 2012
Regression Statistics R-squared 0.6724
Standard error 1.0717 Observations 216 Durbin–Watson 2.1784
Coefficient Standard Error t-Statistic Intercept 0.2371 0.0900 2.6344
Sales growtht −1 −0.0792 0.0398 −1.9899 Sales growtht −12 0.7798 0.0388 20.0979
Autocorrelations of the Residual Lag Autocorrelation Standard Error t-Statistic 1 −0.0770 0.0680 −1.1324 2 −0.0374 0.0680 −0.5500 3 0.0292 0.0680 0.4294 4 −0.0358 0.0680 −0.5265 5 −0.0399 0.0680 −0.5868 6 0.0227 0.0680 0.3338 7 −0.0967 0.0680 −1.4221 8 0.1241 0.0680 1.8250 9 0.0499 0.0680 0.7338 10 −0.0631 0.0680 −0.9279 11 0.0231 0.0680 0.3397 12 −0.1168 0.0680 −1.7176
Source: Statistics Canada (Government of Canada).
8. Autoregressive Moving-Average Models So far, we have presented autoregressive and moving-average models as alternatives for modeling a time series. The time series we have considered in examples have usually been explained quite well with a simple autoregressive model (with or without seasonal lags).38 Some statisticians, however, have advocated using a more general model, the autoregressive moving-average (ARMA) model. The advocates of ARMA models argue that these models may fit the data better and provide better forecasts than do plain autoregressive (AR) models. However, as we discuss later in this section, there are severe limitations to estimating and using these models. Because you may encounter ARMA models, we provide a brief overview below.
An ARMA model combines both autoregressive lags of the dependent variable and moving-average errors. The equation for such a model with p autoregressive terms and q moving-average terms, denoted ARMA(p, q), is
(16)
where b1, b2, …, bp are the autoregressive parameters and θ1, θ2, …, θq are the moving-average parameters.
Estimating and using ARMA models has several limitations. First, the parameters in ARMA models can be very unstable. In particular, slight changes in the data sample or the initial guesses for the values of the ARMA parameters can result in very different final estimates of the ARMA parameters. Second, choosing the right ARMA model is more of an art than a science. The criteria for deciding on p and q for a particular time series are far from perfect. Moreover, even after a model is selected, that model may not forecast well.
To reiterate, ARMA models can be very unstable, depending on the data sample used and the particular ARMA model estimated. Therefore, you should be skeptical of claims that a particular ARMA model provides much better forecasts of a time series than any other ARMA model. In fact, in most cases, you can use an AR model to produce forecasts that are just as accurate as those from ARMA models without nearly as much complexity. Even some of the strongest advocates of ARMA models admit that these models should not be used with fewer than 80 observations, and they do not recommend using ARMA models for predicting quarterly sales or gross margins for a company using even 15 years of quarterly data.
9. Autoregressive Conditional Heteroskedasticity Models Up to now, we have ignored any issues of heteroskedasticity in time-series models and have assumed homoskedasticity. Heteroskedasticity is the dependence of the error term variance on the independent variable; homoskedasticity is the independence of the error term variance from the independent variable. We have assumed that the error term’s variance is constant and does not depend on the value of the time series itself or on the size of previous errors. At times, however, this assumption is violated and the variance of the error term is not constant. In such a situation, the standard errors of the regression coefficients in AR, MA, or ARMA models will be incorrect, and our hypothesis tests would be invalid. Consequently, we can make poor investment decisions based on those tests.
For example, suppose you are building an autoregressive model of a company’s sales. If heteroskedasticity is present, then the standard errors of the regression coefficients of your model are incorrect. It is likely that due to heteroskedasticity, one or more of the lagged sales terms may appear statistically significant when in fact they are not. Therefore, if you use this model for your decision making, you may make some suboptimal decisions.
In work responsible in part for his shared Nobel Prize in Economics for 2003, Robert F. Engle in 1982 first suggested a way of testing whether the variance of the error in a particular time-series model in one period depends on the variance of the error in previous periods. He called this type of heteroskedasticity autoregressive conditional heteroskedasticity (ARCH).
As an example, consider the ARCH(1) model
(17)
where the distribution of εt, conditional on its value in the previous period, εt −1, is normal with mean 0 and variance . If a1 = 0, the variance of the error in every period is just a0. The variance is constant over time and does not depend on past errors. Now suppose thata1 > 0. Then the variance of the error in one period depends on how large the squared error was in the previous period. If a large error occurs in one period, the variance of the error in the next period will be even larger.
Engle shows that we can test whether a time series is ARCH(1) by regressing the squared residuals from a previously estimated time-series model (AR, MA, or
ARMA) on a constant and one lag of the squared residuals. We can estimate the linear regression equation
(18)
where ut is an error term. If the estimate of a1 is statistically significantly different from zero, we conclude that the time series is ARCH(1). If a time-series model has ARCH(1) errors, then the variance of the errors in period t + 1 can be predicted in period t using the formula
EXAMPLE 17 Testing for ARCH(1) in Monthly Inflation
Analyst Lisette Miller wants to test whether monthly data on CPI inflation contain autoregressive conditional heteroskedasticity. She could estimate Equation 18 using the residuals from the time-series model. Based on the analyses in Examples 6 through 9, she has concluded that if she modeled monthly CPI inflation from 1984 to 2013, there is not much difference in the performance of AR(1) and AR(2) models in forecasting inflation. The AR(1) model is clearly better for the period of 2007–2013. She decides to further explore the AR(1) model for the entire period of 1984 to 2013. Table 17 shows the results of testing whether the errors in that model are ARCH(1). Because the test involves the first lag of residuals of the estimated time series model, the number of observations in the test is one less than that in the model.
TABLE 17 Test for ARCH(1) in an AR(1) Model Residuals from Monthly CPI Inflation at an Annual Rate February 1984–December 2013
Regression Statistics R-squared 0.0467
Standard error 24.3630 Observations 358 Durbin–Watson 2.1046
Coefficient Standard Error t-Statistic Intercept 9.1666 1.4228 6.4426
0.2161 0.0518 4.1718
Source: US Bureau of Labor Statistics.
The t-statistic for the coefficient on the previous period’s squared residuals is greater than 4.1. Therefore, Miller easily rejects the null hypothesis that the variance of the error does not depend on the variance of previous errors. Consequently, the test statistics she computed in Table 5 are not valid, and she should not use them in deciding her investment strategy.
It is possible Miller ’s conclusion—that the AR(1) model for monthly inflation has ARCH in the errors—may have been due to the sample period employed (1984 to 2013). In Example 9, she used a shorter sample period of 2007 to 2013 and concluded that monthly CPI inflation follows an AR(1) process. (These results were shown in Table 8.) Table 17 shows that errors for a time-series model of inflation for the entire sample (1984 to 2013) have ARCH errors. Do the errors estimated with a shorter sample period (2007 to 2013) also display ARCH? For the shorter sample period, Miller estimated an AR(1) model using monthly inflation data. Now she tests to see whether the errors display ARCH. Table 18 shows the results.
In this sample, the coefficient on the previous period’s squared residual is quite small and has a t-statistic of only 1.3713. Consequently, Miller fails to reject the null hypothesis that the errors in this regression have no autoregressive conditional heteroskedasticity. This is additional evidence that the AR(1) model for 2007 to 2013 is a good fit. The error variance appears to be homoskedastic, and Miller can rely on the t-statistics. This result again confirms that a single AR process for the entire 1984–2013 period is misspecified (it does not describe the data well).
TABLE 18 Test for ARCH(1) in an AR(1) Model Monthly CPI Inflation at an Annual Rate February 2007–December 2013
Regression Statistics R-squared 0.0230
Standard error 35.218 Observations 82 Durbin–Watson 2.0521
Coefficient Standard Error t-Statistic Intercept 16.6213 4.4632 3.7241
0.1518 0.1107 1.3713
Source: US Bureau of Labor Statistics.
Suppose a model contains ARCH(1) errors. What are the consequences of that fact? First, if ARCH exists, the standard errors for the regression parameters will not be correct. We will need to use generalized least squares39 or other methods that correct for heteroskedasticity to correctly estimate the standard error of the parameters in the time-series model. Second, if ARCH exists and we have it modeled, for example as ARCH(1), we can predict the variance of the errors. Suppose, for instance, that we want to predict the variance of the error in inflation using the estimated parameters from Table 17: . If the error in one period were 0 percent, the predicted variance of the error in the next period would be 9.1666 + 0.2161(0) = 9.1666. If the error in one period were 1 percent, the predicted variance of the error in the next period would be 9.1666 + 0.2161(12) = 9.3827.
Engle and other researchers have suggested many generalizations of the ARCH(1) model, including ARCH(p) and generalized autoregressive conditional heteroskedasticity (GARCH) models. In an ARCH(p) model, the variance of the error term in the current period depends linearly on the squared errors from the
previous p periods: . GARCH models are similar to ARMA models of the error variance in a time series. Just like ARMA models, GARCH models can be finicky and unstable: Their results can depend greatly on the sample period and the initial guesses of the parameters in the GARCH model. Financial analysts who use GARCH models should be well aware of how delicate these models can be, and they should examine whether GARCH estimates are robust to changes in the sample and the initial guesses about the parameters.40
10. Regressions with More than One Time Series Up to now, we have discussed time-series models only for one time series. Although in the readings on correlation and regression and on multiple regression we used linear regression to analyze the relationship among different time series, in those readings we completely ignored unit roots. A time series that contains a unit root is not covariance stationary. If any time series in a linear regression contains a unit root, ordinary least squares estimates of regression test statistics may be invalid.
To determine whether we can use linear regression to model more than one time series, let us start with a single independent variable; that is, there are two time series, one corresponding to the dependent variable and one corresponding to the independent variable. We will then extend our discussion to multiple independent variables.
We first use a unit root test, such as the Dickey–Fuller test, for each of the two time series to determine whether either of them has a unit root.41 There are several possible scenarios related to the outcome of these tests. One possible scenario is that we find that neither of the time series has a unit root. Then we can safely use linear regression to test the relations between the two time series. Otherwise, we may have to use additional tests, as we discuss later in this section.
EXAMPLE 18 Unit Roots and the Fisher Effect
In Example 8 in the reading on multiple regression, we examined the Fisher effect by estimating the regression relation between expected inflation and US Treasury bill (T-bill) returns. We used 181 quarterly observations on expected inflation rates and T-Bill returns from the sample period extending from the fourth quarter of 1968 through the fourth quarter of 2013. We used linear regression to analyze the relationship between the two time series. The results of this regression would be valid if both the time series are covariance stationary; that is, neither of the two time series has a unit root. So, if we compute the Dickey–Fuller t-test statistic of the hypothesis of a unit root separately for each time series and find that we can reject the null hypothesis that the T-bill return series has a unit root and the null hypothesis that the expected inflation time series has a unit root, then we can use linear regression to analyze the relation between the two series. In that case, the results of our analysis of the Fisher effect would be valid.
A second possible scenario is that we reject the hypothesis of a unit root for the
independent variable but fail to reject the hypothesis of a unit root for the dependent variable. In this case, the error term in the regression would not be covariance stationary. Therefore, one or more of the following linear regression assumptions would be violated: 1) that the expected value of the error term is 0, 2) that the variance of the error term is constant for all observations, and 3) that the error term is uncorrelated across observations. Consequently, the estimated regression coefficients and standard errors would be inconsistent. The regression coefficients might appear significant, but those results would be spurious.42 Thus we should not use linear regression to analyze the relation between the two time series in this scenario.
A third possible scenario is the reverse of the second scenario: We reject the hypothesis of a unit root for the dependent variable but fail to reject the hypothesis of a unit root for the independent variable. In this case also, like the second scenario, the error term in the regression would not be covariance stationary, and we cannot use linear regression to analyze the relation between the two time series.
EXAMPLE 19 Unit Roots and Predictability of Stock Market Returns by Price-to-Earnings Ratio
Johann de Vries is analyzing the performance of the South African stock market. He examines whether the percentage change in the Johannesburg Stock Exchange (JSE) All Share Index can be predicted by the price-to-earnings ratio (P/E) for the index. Using monthly data from January 1994 to December 2013, he runs a regression using (Pt − Pt −1)/Pt −1 as the dependent variable and Pt −1/Et −2 as the independent variable, where Pt is the value of the JSE index at time t and Et is the earnings on the index. De Vries finds that the regression coefficient is negative and statistically significant and the value of the R- squared for the regression is quite high. What additional analysis should he perform before accepting the regression as valid?
De Vries needs to perform unit root tests for each of the two time series. If one of the two time series has a unit root, implying that it is not stationary, the results of the linear regression are not meaningful and cannot be used to conclude that stock market returns are predictable by P/E.43
The next possibility is that both time series have a unit root. In this case, we need to establish whether the two time series are cointegrated before we can rely on regression analysis.44 Two time series are cointegrated if a long-term financial or
economic relationship exists between them such that they do not diverge from each other without bound in the long run. For example, two time series are cointegrated if they share a common trend.
In the fourth scenario, both time series have a unit root but are not cointegrated. In this scenario, as in the second and third scenarios above, the error term in the linear regression will not be covariance stationary, some regression assumptions will be violated, the regression coefficients and standard errors will not be consistent, and we cannot use them for hypothesis tests. Consequently, linear regression of one variable on the other would be meaningless.
Finally, the fifth possible scenario is that both time series have a unit root, but they are cointegrated. In this case, the error term in the linear regression of one time series on the other will be covariance stationary. Accordingly, the regression coefficients and standard errors will be consistent, and we can use them for hypothesis tests. However, we should be very cautious in interpreting the results of a regression with cointegrated variables. The cointegrated regression estimates the long-term relation between the two series but may not be the best model of the short-term relation between the two series. Short-term models of cointegrated series (error correction models) are discussed in Engle and Granger (1987) and Tsay (2010), but these are specialist topics.
Now let us look at how we can test for cointegration between two time series that each have a unit root as in the last two scenarios above.45 Engle and Granger suggest this test: If yt and xt are both time series with a unit root, we should do the following: 1. Estimate the regression yt = b0 + b1xt + εt.
2. Test whether the error term from the regression in Step 1 has a unit root using a Dickey–Fuller test. Because the residuals are based on the estimated coefficients of the regression, we cannot use the standard critical values for the Dickey–Fuller test. Instead, we must use the critical values computed by Engle and Granger, which take into account the effect of uncertainty about the regression parameters on the distribution of the Dickey–Fuller test.
3. If the (Engle–Granger) Dickey–Fuller test fails to reject the null hypothesis that the error term has a unit root, then we conclude that the error term in the regression is not covariance stationary. Therefore, the two time series are not cointegrated. In this case any regression relation between the two series is spurious.
4. If the (Engle–Granger) Dickey–Fuller test rejects the null hypothesis that the
error term has a unit root, then we may assume that the error term in the regression is covariance stationary and that the two time series are cointegrated. The parameters and standard errors from linear regression will be consistent and will let us test hypotheses about the long-term relation between the two series.
EXAMPLE 20 Testing for Cointegration between Intel Sales and Nominal GDP
Suppose we want to test whether the natural log of Intel's sales and the natural log of GDP are cointegrated (that is, whether there is a long-term relation between GDP and Intel sales). We want to test this hypothesis using quarterly data from the first quarter of 1995 through the fourth quarter of 2013. Here are the steps:
1. Test whether the two series each have a unit root. If we cannot reject the null
hypothesis of a unit root for both series, implying that both series are nonstationary, we must then test whether the two series are cointegrated.
2. Having established that each series has a unit root, we estimate the regression ln (Intel Salest) = b0 + b1 ln GDPt + εt, then conduct the (Engle–Granger) Dickey–Fuller test of the hypothesis that there is a unit root in the error term of this regression using the residuals from the estimated regression. If we reject the null hypothesis of a unit root in the error term of the regression, we reject the null hypothesis of no cointegration. That is, the two series would be cointegrated. If the two series are cointegrated, we can use linear regression to estimate the long-term relation between the natural log of Intel Sales and the natural log of GDP.
We have so far discussed models with a single independent variable. We now extend the discussion to a model with two or more independent variables, so that there are three or more time series. The simplest possibility is that none of the time series in the model has a unit root. Then, we can safely use multiple regression to test the relation among the time series.
EXAMPLE 21 Unit Roots and Returns to the Fidelity Select Technology Fund
In Example 3 in the reading on multiple regression, we used multiple linear
regression to examine whether returns to either the S&P 500 Growth Index or the S&P 500 Value Index explain returns to the Fidelity Select Technology Portfolio using 60 monthly observations between January 2009 and December 2013. Of course, if any of the three time series has a unit root, then the results of our regression analysis may be invalid. Therefore, we could use a Dickey– Fuller test to determine whether any of these series has a unit root.
If we reject the hypothesis of unit roots for all three series, we can use linear regression to analyze the relation among the series. In that case the results of our analysis of the factors affecting returns to the Fidelity Select Technology Portfolio would be valid.
If at least one time series (the dependent variable or one of the independent variables) has a unit root while at least one time series (the dependent variable or one of the independent variables) does not, the error term in the regression cannot be covariance stationary. Consequently, we should not use multiple linear regression to analyze the relation among the time series in this scenario.
Another possibility is that each time series, including the dependent variable and each of the independent variables, has a unit root. If this is the case, we need to establish whether the time series are cointegrated. To test for cointegration, the procedure is similar to that for a model with a single independent variable. First, estimate the regression yt = b0 + b1x1t + b2x2t + … + bkxkt + εt. Then conduct the (Engle–Granger) Dickey–Fuller test of the hypothesis that there is a unit root in the errors of this regression using the residuals from the estimated regression.
If we cannot reject the null hypothesis of a unit root in the error term of the regression, we cannot reject the null hypothesis of no cointegration. In this scenario, the error term in the multiple regression will not be covariance stationary, so we cannot use multiple regression to analyze the relationship among the time series.
If we can reject the null hypothesis of a unit root in the error term of the regression, we can reject the null hypothesis of no cointegration. However, modeling three or more time series that are cointegrated may be difficult. For example, an analyst may want to predict a retirement services company’s sales based on the country’s GDP and the total population over age 65. Although the company’s sales, GDP, and the population over 65 may each have a unit root and be cointegrated, modeling the cointegration of the three series may be difficult, and doing so is beyond the scope of this volume. Analysts who have not mastered all these complex issues should avoid forecasting models with multiple time series that have unit roots: The regression coefficients may be inconsistent and may produce incorrect forecasts.
11. Other Issues in Time Series Time-series analysis is an extensive topic and includes many highly complex issues. Our objective in this reading has been to present those issues in time series that are the most important for financial analysts and can also be handled with relative ease. In this section, we briefly discuss some of the issues that we have not covered but could be useful for analysts.
In this reading, we have shown how to use time-series models to make forecasts. We have also introduced the RMSE as a criterion for comparing forecasting models. However, we have not discussed measuring the uncertainty associated with forecasts made using time-series models. The uncertainty of these forecasts can be very large, and should be taken into account when making investment decisions. Fortunately, the same techniques apply to evaluating the uncertainty of time-series forecasts as apply to evaluating the uncertainty about forecasts from linear regression models. To accurately evaluate forecast uncertainty, we need to consider both the uncertainty about the error term and the uncertainty about the estimated parameters in the time- series model. Evaluating this uncertainty is fairly complicated when using regressions with more than one independent variable.
In this reading, we used the US CPI inflation series to illustrate some of the practical challenges analysts face in using time-series models. We used information on US Federal Reserve policy to explore the consequences of splitting the inflation series in two. In financial time-series work, we may suspect that a time series has more than one regime but lack the information to attempt to sort the data into different regimes. If you face such a problem, you may want to investigate other methods, especially switching regression models, to identify multiple regimes using only the time series itself.
If you are interested in these and other advanced time-series topics, you can learn more in Diebold (2007) and Tsay (2010).
12. Suggested Steps in Time-Series Forecasting The following is a step-by-step guide to building a model to predict a time series. 1. Understand the investment problem you have, and make an initial choice of
model. One alternative is a regression model that predicts the future behavior of a variable based on hypothesized causal relationships with other variables. Another is a time-series model that attempts to predict the future behavior of a variable based on the past behavior of the same variable.
2. If you have decided to use a time-series model, compile the time series and plot it to see whether it looks covariance stationary. The plot might show important deviations from covariance stationarity, including the following:
a linear trend;
an exponential trend;
seasonality; or
a significant shift in the time series during the sample period (for example, a change in mean or variance).
3. If you find no significant seasonality or shift in the time series, then perhaps either a linear trend or an exponential trend will be sufficient to model the time series. In that case, take the following steps:
Determine whether a linear or exponential trend seems most reasonable (usually by plotting the series).
Estimate the trend.
Compute the residuals.
Use the Durbin–Watson statistic to determine whether the residuals have significant serial correlation. If you find no significant serial correlation in the residuals, then the trend model is sufficient to capture the dynamics of the time series and you can use that model for forecasting.
4. If you find significant serial correlation in the residuals from the trend model, use a more complex model, such as an autoregressive model. First, however, reexamine whether the time series is covariance stationary. Following is a list of violations of stationarity, along with potential methods to adjust the time series to make it covariance stationary:
If the time series has a linear trend, first-difference the time series.
If the time series has an exponential trend, take the natural log of the time series and then first-difference it.
If the time series shifts significantly during the sample period, estimate different time-series models before and after the shift.
If the time series has significant seasonality, include seasonal lags (discussed in Step 7).
5. After you have successfully transformed a raw time series into a covariance- stationary time series, you can usually model the transformed series with a short autoregression.46 To decide which autoregressive model to use, take the following steps:
Estimate an AR(1) model.
Test to see whether the residuals from this model have significant serial correlation.
If you find no significant serial correlation in the residuals, you can use the AR(1) model to forecast.
6. If you find significant serial correlation in the residuals, use an AR(2) model and test for significant serial correlation of the residuals of the AR(2) model.
If you find no significant serial correlation, use the AR(2) model.
If you find significant serial correlation of the residuals, keep increasing the order of the AR model until the residual serial correlation is no longer significant.
7. Your next move is to check for seasonality. You can use one of two approaches:
Graph the data and check for regular seasonal patterns.
Examine the data to see whether the seasonal autocorrelations of the residuals from an AR model are significant (for example, the fourth autocorrelation for quarterly data) and whether the autocorrelations before and after the seasonal autocorrelations are significant. To correct for seasonality, add seasonal lags to your AR model. For example, if you are using quarterly data, you might add the fourth lag of a time series as an additional variable in an AR(1) or an AR(2) model.
8. Next, test whether the residuals have autoregressive conditional heteroskedasticity. To test for ARCH(1), for example, do the following:
Regress the squared residual from your time-series model on a lagged value of the squared residual.
Test whether the coefficient on the squared lagged residual differs significantly from 0.
If the coefficient on the squared lagged residual does not differ significantly from 0, the residuals do not display ARCH and you can rely on the standard errors from your time-series estimates.
If the coefficient on the squared lagged residual does differ significantly from 0, use generalized least squares or other methods to correct for ARCH.
9. Finally, you may also want to perform tests of the model’s out-of-sample forecasting performance to see how the model’s out-of-sample performance compares to its in-sample performance.
Using these steps in sequence, you can be reasonably sure that your model is correctly specified.
13. Summary
The predicted trend value of a time series in period t is t in a linear trend model; the predicted trend value of a time series in a log-linear trend model is
.
Time series that tend to grow by a constant amount from period to period should be modeled by linear trend models, whereas time series that tend to grow at a constant rate should be modeled by log-linear trend models.
Trend models often do not completely capture the behavior of a time series, as indicated by serial correlation of the error term. If the Durbin–Watson statistic from a trend model differs significantly from 2, indicating serial correlation, we need to build a different kind of model.
An autoregressive model of order p, denoted AR(p), uses p lags of a time series to predict its current value: xt = b0 + b1xt −1 + b2 xt −2 + … + bp xt −p + εt.
A time series is covariance stationary if the following three conditions are satisfied: First, the expected value of the time series must be constant and finite in all periods. Second, the variance of the time series must be constant and finite in all periods. Third, the covariance of the time series with itself for a fixed number of periods in the past or future must be constant and finite in all periods. Inspection of a nonstationary time-series plot may reveal an upward or downward trend (nonconstant mean) and/or nonconstant variance. The use of linear regression to estimate an autoregressive time-series model is not valid unless the time series is covariance stationary.
For a specific autoregressive model to be a good fit to the data, the autocorrelations of the error term should be 0 at all lags.
A time series is mean reverting if it tends to fall when its level is above its long-run mean and rise when its level is below its long-run mean. If a time series is covariance stationary, then it will be mean reverting.
The one-period-ahead forecast of a variable xt from an AR(1) model made in period t for period t + 1 is . This forecast can be used to create the
two-period-ahead forecast from the model made in period t, . Similar results hold for AR (p) models.
In-sample forecasts are the in-sample predicted values from the estimated time- series model. Out-of-sample forecasts are the forecasts made from the
estimated time-series model for a time period different from the one for which the model was estimated. Out-of-sample forecasts are usually more valuable in evaluating the forecasting performance of a time-series model than are in- sample forecasts. The root mean squared error (RMSE), defined as the square root of the average squared forecast error, is a criterion for comparing the forecast accuracy of different time-series models; a smaller RMSE implies greater forecast accuracy.
Just as in regression models, the coefficients in time-series models are often unstable across different sample periods. In selecting a sample period for estimating a time-series model, we should seek to assure ourselves that the time series was stationary in the sample period.
A random walk is a time series in which the value of the series in one period is the value of the series in the previous period plus an unpredictable random error. If the time series is a random walk, it is not covariance stationary. A random walk with drift is a random walk with a nonzero intercept term. All random walks have unit roots. If a time series has a unit root, then it will not be covariance stationary.
If a time series has a unit root, we can sometimes transform the time series into one that is covariance stationary by first-differencing the time series; we may then be able to estimate an autoregressive model for the first-differenced series.
An n-period moving average of the current and past (n − 1) values of a time series, xt, is calculated as [xt + xt −1 + … + xt −(n −1)]/n.
A moving-average model of order q, denoted MA(q), uses q lags of a random error term to predict its current value.
The order q of a moving average model can be determined using the fact that if a time series is a moving-average time series of order q, its first q autocorrelations are nonzero while autocorrelations beyond the first q are zero.
The autocorrelations of most autoregressive time series start large and decline gradually, whereas the autocorrelations of an MA(q) time series suddenly drop to 0 after the first q autocorrelations. This helps in distinguishing between autoregressive and moving-average time series.
If the error term of a time-series model shows significant serial correlation at seasonal lags, the time series has significant seasonality. This seasonality can often be modeled by including a seasonal lag in the model, such as adding a
term lagged four quarters to an AR(1) model on quarterly observations.
The forecast made in time t for time t + 1 using a quarterly AR(1) model with a seasonal lag would be .
ARMA models have several limitations: the parameters in ARMA models can be very unstable; determining the AR and MA order of the model can be difficult; and even with their additional complexity, ARMA models may not forecast well.
The variance of the error in a time-series model sometimes depends on the variance of previous errors, representing autoregressive conditional heteroskedasticity (ARCH). Analysts can test for first-order ARCH in a time- series model by regressing the squared residual on the squared residual from the previous period. If the coefficient on the squared residual is statistically significant, the time-series model has ARCH(1) errors.
If a time-series model has ARCH(1) errors, then the variance of the errors in period t + 1 can be predicted in period t using the formula .
If linear regression is used to model the relationship between two time series, a test should be performed to determine whether either time series has a unit root:
If neither of the time series has a unit root, then we can safely use linear regression.
If one of the two time series has a unit root, then we should not use linear regression.
If both time series have a unit root and the time series are cointegrated, we may safely use linear regression; however, if they are not cointegrated, we should not use linear regression. The (Engle–Granger) Dickey–Fuller test can be used to determine if time series are cointegrated.
Problems Practice Problems and Solutions: Quantitative Methods for Investment Analysis, Second Edition, by Richard A. DeFusco, CFA, Dennis W. McLeavey, CFA, Jerald E. Pinto, PhD, CFA, and David E. Runkle, PhD, CFA. Copyright © 2004 by CFA Institute.
Note: In the Problems and Solutions for this reading, we use the hat ( ) to indicate an estimate if we are trying to differentiate between an estimated and an actual value. However, we suppress the hat when we are clearly showing regression output. 1. The civilian unemployment rate (UER) is an important component of many
economic models. Table 1 gives regression statistics from estimating a linear trend model of the unemployment rate: UERt = b0 + b1t + εt.
TABLE 1 Estimating a Linear Trend in the Civilian Unemployment Rate Monthly Observations, January 1996–December 2000
Regression Statistics R-squared 0.9314
Standard error 0.1405 Observations 60 Durbin–Watson 0.9099
Coefficient Standard Error t-Statistic Intercept 5.5098 0.0367 150.0363 Trend –0.0294 0.0010 –28.0715
1. Using the regression output in the above table, what is the model’s
prediction of the unemployment rate for July 1996, midway through the first year of the sample period?
2. How should we interpret the Durbin–Watson (DW) statistic for this regression? What does the value of the DW statistic say about the validity of a t-test on the coefficient estimates?
2. Figure 1 compares the predicted civilian unemployment rate (PRED) with the actual civilian unemployment rate (UER) from January 1996 to December 2000. The predicted results come from estimating the linear time trend model UERt = b0 + b1t + εt.
What can we conclude about the appropriateness of this model?
FIGURE 1 Predicted and Actual Civilian Unemployment Rates
3. You have been assigned to analyze automobile manufacturers and as a first step in your analysis, you decide to model monthly sales of lightweight vehicles to determine sales growth in that part of the industry. Figure 2 gives lightweight vehicle monthly sales (annualized) from January 1992 to December 2000.
FIGURE 2 Lightweight Vehicle Sales
Monthly sales in the lightweight vehicle sector, Salest, have been increasing over time, but you suspect that the growth rate of monthly sales is relatively constant. Write the simplest time-series model for Salest that is consistent with your perception.
4. Figure 3 shows a plot of the first differences in the civilian unemployment rate (UER) between January 1996 and December 2000, ΔUERt = UERt − UERt−1.
FIGURE 3 Change in Civilian Unemployment Rate 1. Has differencing the data made the new series, ΔUERt, covariance
stationary? Explain your answer.
2. Given the graph of the change in the unemployment rate shown in the figure, describe the steps we should take to determine the appropriate autoregressive time-series model specification for the series ΔUERt.
5. Table 2 gives the regression output of an AR(1) model on first differences in the unemployment rate. Describe how to interpret the DW statistic for this regression.
TABLE 2 Estimating an AR(1) Model of Changes in the Civilian Unemployment Rate Monthly Observations, March 1996–December 2000
Regression Statistics R-squared 0.2184
Standard error 0.1202 Observations 58 Durbin–Watson 2.1852
Coefficient Standard Error t-Statistic
Intercept −0.0405 0.0161 −2.5110 ΔUERt −1 −0.4674 0.1181 −3.9562
6. Assume that changes in the civilian unemployment rate are covariance stationary and that an AR(1) model is a good description for the time series of changes in the unemployment rate. Specifically, we have ΔUERt = −0.0405 − 0.4674ΔUERt −1 (using the coefficient estimates given in the previous problem). Given this equation, what is the mean-reverting level to which changes in the unemployment rate converge?
7. Suppose the following model describes changes in the civilian unemployment rate: ΔUERt = −0.0405 − 0.4674ΔUERt −1. The current change (first difference) in the unemployment rate is 0.0300. Assume that the mean-reverting level for changes in the unemployment rate is −0.0276.
1. What is the best prediction of the next change?
2. What is the prediction of the change following the next change?
3. Explain your answer to Part B in terms of equilibrium.
8. Table 3 gives the actual sales, log of sales, and changes in the log of sales of Cisco Systems for the period 1Q:2001 to 4Q:2001.
TABLE 3
Date Quarter: Year
Actual Sales($ Millions)
Log of Sales
Changes in Log of Sales Δln (Salest)
1Q:2001 6,519 8.7825 0.1308 2Q:2001 6,748 8.8170 0.0345 3Q:2001 4,728 8.4613 −0.3557 4Q:2001 4,298 8.3659 −0.0954 1Q:2002 2Q:2002
Forecast the first-and second-quarter sales of Cisco Systems for 2002 using the regression Δln (Salest) = 0.0661 + 0.4698Δln (Salest −1).
9. Table 4 gives the actual change in the log of sales of Cisco Systems from 1Q:2001 to 4Q:2001, along with the forecasts from the regression model Δln (Salest) = 0.0661 + 0.4698Δln (Salest −1) estimated using data from 3Q:1991 to
4Q:2000. (Note that the observations after the fourth quarter of 2000 are out of sample.)
TABLE 4
Date Actual Values of Changes in theLog of Sales Δln (Salest) Forecast Values of Changes in the Log of Sales Δln (Salest)
1Q:2001 0.1308 0.1357 2Q:2001 0.0345 0.1299 3Q:2001 –0.3557 0.1271 4Q:2001 –0.0954 0.1259 1. Calculate the RMSE for the out-of-sample forecast errors.
2. Compare the forecasting performance of the model given with that of another model having an out-of-sample RMSE of 20 percent.
10. A. The AR(1) model for the civilian unemployment rate, ΔUERt = −0.0405 − 0.4674ΔUERt −1, was developed with five years of data. What would be the drawback to using the AR(1) model to predict changes in the civilian unemployment rate 12 months or more ahead, as compared with one month ahead?
B.For purposes of estimating a predictive equation, what would be the drawback to using 30 years of civilian unemployment data rather than only five years?
11. Figure 4 shows monthly observations on the natural log of lightweight vehicle sales, ln (Salest), for the period January 1992 to December 2000.
FIGURE 4 Lightweight Vehicle Sales 1. Using the figure, comment on whether the specification ln (Salest) = b0 + b1[ln (Salest −1)] + εt is appropriate.
2. State an appropriate transformation of the time series.
12. Figure 5 shows a plot of first differences in the log of monthly lightweight vehicle sales over the same period as in Problem 11. Has differencing the data made the resulting series, Δln (Salest) = ln (Salest) − ln (Salest −1), covariance stationary?
FIGURE 5 Change in Natural Log of Lightweight Vehicle Sales
13. Using monthly data from January 1992 to December 2000, we estimate the following equation for lightweight vehicle sales: Δln (Salest) = 2.7108 + 0.3987Δln (Salest −1) + εt. Table 5 gives sample autocorrelations of the errors from this model.
TABLE 5 Different Order Autocorrelations of Differences in the Logs of Vehicle Sales
Lag Autocorrelation Standard Error t-Statistic 1 0.9358 0.0962 9.7247 2 0.8565 0.0962 8.9005 3 0.8083 0.0962 8.4001 4 0.7723 0.0962 8.0257 5 0.7476 0.0962 7.7696
6 0.7326 0.0962 7.6137 7 0.6941 0.0962 7.2138 8 0.6353 0.0962 6.6025 9 0.5867 0.0962 6.0968 10 0.5378 0.0962 5.5892 11 0.4745 0.0962 4.9315 12 0.4217 0.0962 4.3827
1. Use the information in the table to assess the appropriateness of the
specification given by the equation.
2. If the residuals from the AR(1) model above violate a regression assumption, how would you modify the AR(1) specification?
14. Figure 6 shows the quarterly sales of Cisco Systems from 1Q:1991 to 4Q:2000.
FIGURE 6 Quarterly Sales at Cisco
Table 6 gives the regression statistics from estimating the model Δln (Salest) = b0 + b1Δln (Salest −1) + εt.
TABLE 6 Change in the Natural Log of Sales for Cisco Systems Quarterly Observations, 3Q:1991–4Q:2000
Regression Statistics R-squared 0.2899
Standard error 0.0408
Observations 38 Durbin–Watson 1.5707
Coefficient Standard Error t-Statistic Intercept 0.0661 0.0175 3.7840
Δln (Salest −1) 0.4698 0.1225 3.8339 1. Describe the salient features of the quarterly sales series.
2. Describe the procedures we should use to determine whether the AR(1) specification is correct.
3. Assuming the model is correctly specified, what is the long-run change in the log of sales toward which the series will tend to converge?
15. Figure 7 shows the quarterly sales of Avon Products from 1Q:1992 to 2Q:2002. Describe the salient features of the data shown.
FIGURE 7 Quarterly Sales at Avon
16. Table 7 following shows the autocorrelations of the residuals from an AR(1) model fit to the changes in the gross profit margin (GPM) of The Home Depot, Inc.
TABLE 7 Autocorrelations of the Residuals from Estimating the Regression ΔGPMt = 0.0006 − 0.33301ΔGPMt −1 + εt 1Q:1992–4Q:2001 (40 Observations)
Lag Autocorrelation
1 −0.1106 2 −0.5981 3 −0.1525 4 0.8496 5 −0.1099
Table 8 shows the output from a regression on changes in the GPM for Home Depot, where we have changed the specification of the AR regression.
TABLE 8 Change in Gross Profit Margin for Home Depot 1Q:1992– 4Q:2001
Regression Statistics R-squared 0.9155
Standard error 0.0057 Observations 40 Durbin–Watson 2.6464
Coefficient Standard Error t-Statistic Intercept −0.0001 0.0009 −0.0610 ΔGPMt −1 −0.0608 0.0687 −0.8850 ΔGPMt −4 0.8720 0.0678 12.8683
1. Identify the change that was made to the regression model.
2. Discuss the rationale for changing the regression specification.
17. Suppose we decide to use an autoregressive model with a seasonal lag because of the seasonal autocorrelation in the previous problem. We are modeling quarterly data, so we estimate Equation 15: (ln Salest − ln Salest −1) = b0 + b1(ln Salest −1 – ln Salest −2) + b2(ln Salest −4 − ln Salest −5) + εt. Table 9 shows the regression statistics from this equation.
TABLE 9 Log Differenced Sales: AR(1) Model with Seasonal Lag Johnson & Johnson Quarterly Observations, January 1985–December 2001
Regression Statistics
R-squared 0.4220 Standard error 0.0318 Observations 68 Durbin–Watson 1.8784
Coefficient Standard Error t-Statistic Intercept 0.0121 0.0053 2.3055 Lag 1 −0.0839 0.0958 −0.8757 Lag 4 0.6292 0.0958 6.5693
Autocorrelations of the Residual Lag Autocorrelation Standard Error t-Statistic 1 0.0572 0.1213 0.4720 2 −0.0700 0.1213 −0.5771 3 0.0065 0.1213 −0.0532 4 −0.0368 0.1213 −0.3033
1. Using the information in Table 9, determine if the model is correctly
specified.
2. If sales grew by 1 percent last quarter and by 2 percent four quarters ago, use the model to predict the sales growth for this quarter.
18. Describe how to test for autoregressive conditional heteroskedasticity (ARCH) in the residuals from the AR(1) regression on first differences in the civilian unemployment rate, ΔUERt = b0 + b1ΔUERt −1 + εt.
19. Suppose we want to predict the annualized return of the five-year T-bill using the annualized return of the three-month T-bill with monthly observations from January 1993 to December 2002. Our analysis produces the data shown in Table 10.
TABLE 10 Regression with 3-Month T-Bill as the Independent Variable and 5-Year Treasury Bill as the Dependent Variable Monthly Observations, January 1993 to December 2002
Regression Statistics R-squared 0.5829
Standard error 0.6598 Observations 120
Durbin–Watson 0.1130 Coefficient Standard Error t-Statistic
Intercept 3.0530 0.2060 14.8181 Three-month 0.5722 0.0446 12.8408
Can we rely on the regression model in Table 10 to produce meaningful predictions? Specify what problem might be a concern with this regression.
Notes 1 Examples in this reading were updated in 2014 by Professor Sanjiv Sabherwal of
the University of Texas, Arlington.
2 We could also write the equation as yt = b0 + b1 yt −1 + εt.
3 Recall that ordinary least squares is an estimation method based on the criterion of minimizing the sum of a regression’s squared residuals.
4 In these data, 1 percent is represented as 1.0.
5 For example, if we use annual periods and = 1.04 for a particular series, then that series grows by 1.04 − 1 = 0.04, or 4 percent a year.
6 An exponential growth rate is a compound growth rate with continuous compounding.
7 In discussions of Starbucks’ sales in this reading, year refers to Starbucks’ fiscal year.
8 Note that = 0.0464 implies that the exponential growth rate per quarter in Starbucks’ sales will be 4.75 percent (e0.0464 − 1 = 0.0475).
9 Note that time-series observations, in contrast to cross-sectional observations, have a logical ordering: They must be processed in chronological order of the time periods involved. For example, we should not make a prediction of the inflation rate using a CPI series in which the order of the observations had been scrambled, because time patterns such as growth in the independent variables can negatively affect the statistical properties of the estimated regression coefficients.
10 Significantly small values of the Durbin–Watson statistic indicate positive serial correlation; significantly large values point to negative serial correlation. Here the DW statistic of 1.09 indicates positive serial correlation. For more information, see the readings on regression analysis.
11 “Weakly stationary” is a synonym for covariance stationary. Note that the terms “stationary” or “stationarity” are often used to mean “covariance stationary” or “covariance stationarity,” respectively. You may also encounter the more restrictive concept of “strictly” stationary, which has little practical application.
For details, see Diebold (2007).
12 In the first requirement, we will use the absolute value to rule out the case in which the mean is negative without limit (minus infinity).
13 When s in this equation equals 0, then this equation imposes the condition that the variance of the time series is finite. This is so because the covariance of a random variable with itself is its variance: Cov(yt,yt) =Var(yt).
14 In general, any time series accurately described with a linear or log-linear trend model is not covariance stationary, although a transformation of the original series might be covariance stationary.
15 In particular, random walks are not covariance stationary.
16 Whenever we refer to autocorrelation without qualification, we mean autocorrelation of the time series itself rather than autocorrelation of the error term or residuals.
17 This assumption is similar to the one made in the previous two readings about the expected value of the error term.
18 This calculation is derived in Diebold (2007).
19 We can compute these residual autocorrelations easily with most statistical software packages. In Microsoft Excel, for example, to compute the first-order residual autocorrelation, we compute the correlation of the residuals from observations 1 through T – 1 with the residuals from observations 2 through T.
20 Often, econometricians use additional tests for the significance of residual autocorrelations. For example, the Box–Pierce Q-statistic is frequently used to test the joint hypothesis that all autocorrelations of the residuals are equal to 0. For further discussion, see Diebold (2007).
21 The first lag of a time series is the value of the time series in the previous period.
22 For seasonally unadjusted data, analysts often compute the same number of autocorrelations as there are observations in a year (for example, four for quarterly data). The number of autocorrelations computed also often depends on sample size, as discussed in Diebold (2007).
23 Statisticians have many other tests for serial correlation of the residuals in a time- series model. For details, see Diebold (2007).
24 If a forecasting model is well specified, the prediction errors from the model will not be serially correlated. If the prediction errors for each period are not serially correlated, then the variance of a multiperiod forecast will be higher than the variance of a single-period forecast.
25 Note that Table 6 shows only 358 observations in the regression because the extra lag of inflation requires the estimation sample to start one month later than the regression in Table 5. (With two lags, inflation for January and February 1984 must be known in order to estimate the equation starting in March 1984.)
26 Equation 8 with a nonzero intercept added (as in Equation 9 given later) is sometimes referred to as a random walk with drift.
27 All the covariances are finite, for two reasons: The variance is finite, and the covariance of a time series with its own past value can be no greater than the variance of the series.
28 See Greene (2011) for a test of the joint hypothesis that both regression coefficients are equal to 0.
29 When b1 is greater than 1 in absolute value, we say that there is an explosive root. For details, see Diebold (2007).
30 Dickey and Fuller developed three separate tests of the hypothesis that g1 = 0 assuming the following models: random walk, random walk with drift, or random walk with drift and trend. The critical values for the Dickey–Fuller tests for the three models are different. For more on this topic, see Greene (2011) or Tsay (2010).
31 Note that 2.24 percent is the exponential growth rate, not [(Current quarter sales/Previous quarter sales) − 1]. The difference between these two methods of computing growth is usually small.
32 A 12-month moving average is the average value of a time series over each of the last 12 months. Although the sample period starts in 1995, data from 1994 are used to compute the 12-month moving average for the months of 1994.
33 Note that a moving-average time-series model is very different from a simple moving average, as discussed in Section 6.1. The simple moving average is based on observed values of a time series. In a moving-average time-series model, we never directly observe, εt or any other εt −j, but we can infer how a particular moving-average model will imply a particular pattern of serial
correlation for a time series, as we discuss below.
34 On the basis of investment theory and evidence, we expect that the mean monthly return on the S&P BSE 100 is positive (μ > 0). We can also generalize Equation 13 for an MA(q) time series by adding a constant term, μ. Including a constant term in a moving-average model does not change the expressions for the variance and autocovariances of the time series. A number of early studies of weak-form market efficiency used Equation 14 as the model for stock returns. See Garbade (1982).
35 In this example, we restrict the start of the sample period to the beginning of 1995, and we do not use prior observations for the lags. Accordingly, the number of observations decreases with an increase in the number of lags. In Table 13, the first observation is for the third quarter of 1995 because we use up to two lags. In Table 14, the first observation is for the second quarter of 1996 because we use up to five lags.
36 Note that all of these growth rates are exponential growth rates.
37 In this example, although we state that the sample period begins in 1995, we use prior observations for the lags. This results in the same number of observations irrespective of the number of lags.
38 For the returns on the S&P BSE 100 (see Example 14), we chose a moving- average model over an autoregressive model.
39 See Greene (2011).
40 For more on ARCH, GARCH, and other models of time-series variance, see Hamilton (1994).
41 For theoretical details of unit root tests, see Greene (2011) or Tsay (2010). Unit root tests are available in some econometric software packages, such as EViews.
42 The problem of spurious regression for nonstationary time series was first discussed by Granger and Newbold (1974).
43 Barr and Kantor (1999) contains evidence that the P/E time series is nonstationary.
44 Engle and Granger (1987) first discussed cointegration.
45 Consider a time series, xt, that has a unit root. For many such financial and
economic time series, the first difference of the series, xt − xt −1, is stationary. We say that such a series, whose first difference is stationary, has a single unit root. However, for some time series, even the first difference may not be stationary and further differencing may be needed to achieve stationarity. Such a time series is said to have multiple unit roots. In this section, we consider only the case in which each nonstationary series has a single unit root (which is quite common).
46 Most financial time series can be modeled using an autoregressive process. For a few time series, a moving-average model may fit better. To see if this is the case, examine the first five or six autocorrelations of the time series. If the autocorrelations suddenly drop to 0 after the first q autocorrelations, a moving- average model (of order q) is appropriate. If the autocorrelations start large and decline gradually, an autoregressive model is appropriate.
CHAPTER 11 AN INTRODUCTION TO MULTIFACTOR MODELS Jerald E. Pinto, PhD, CFA
Eugene L. Podkaminer, CFA
LEARNING OUTCOMES
After completing this chapter, you will be able to do the following:
describe arbitrage pricing theory (APT), including its underlying assumptions and its relation to multifactor models;
define arbitrage opportunity and determine whether an arbitrage opportunity exists;
calculate the expected return on an asset given an asset’s factor sensitivities and the factor risk premiums;
describe and compare macroeconomic factor models, fundamental factor models, and statistical factor models;
explain sources of active risk and interpret tracking risk and the information ratio;
describe uses of multifactor models and interpret the output of analyses based on multifactor models;
describe the potential benefits for investors in considering multiple risk dimensions when modeling asset returns.
1. Introduction As used in investments, a factor is a variable or a characteristic with which individual asset returns are correlated. Models using multiple factors are used by asset owners, asset managers, investment consultants, and risk managers for a variety of portfolio construction, portfolio management, risk management, and general analytical purposes. In comparison to single-factor models (typically based on a market risk factor), multifactor models offer increased explanatory power and flexibility. These comparative strengths of multifactor models allow practitioners to
build portfolios that replicate or modify in a desired way the characteristics of a particular index;
establish desired exposures to one or more risk factors, including those that express specific macro expectations (such as views on inflation or economic growth), in portfolios;
perform granular risk and return attribution on actively managed portfolios;
understand the comparative risk exposures of equity, fixed-income, and other asset class returns;
identify active decisions relative to a benchmark and measure the sizing of those decisions; and
ensure that an investor ’s aggregate portfolio is meeting active risk and return objectives commensurate with active fees.
Multifactor models have come to dominate investment practice, having demonstrated their value in helping asset managers and asset owners address practical tasks in measuring and controlling risk. This reading explains and illustrates the various practical uses of multifactor models.
The reading is organized as follows. Section 2 describes the modern portfolio theory background of multifactor models. Section 3 describes arbitrage pricing theory and provides a general expression for multifactor models. Section 4 describes the types of multifactor models, and Section 5 describes selected applications. Section 6 summarizes major points.
2. Multifactor Models and Modern Portfolio Theory In 1952, Markowitz introduced a framework for constructing portfolios of securities by quantitatively considering each investment in the context of a portfolio rather than in isolation; that framework is widely known today as modern portfolio theory (MPT). Markowitz simplified modeling asset returns using a multivariate normal distribution, which completely defines the distribution of returns in terms of mean returns, return variances, and return correlations. One of the key insights of MPT is that any value of correlation among asset returns of less than one offers the potential for risk reduction by means of diversification. In 1964, Sharpe introduced the capital asset pricing model (CAPM), a model for the expected return of assets in equilibrium based on a mean–variance foundation. The CAPM and the literature that developed around it has provided investors with useful and influential concepts— such as alpha, beta, and systematic risk—for thinking about investing. The concept of systematic risk, for example, is critical to understanding multifactor models. There are potentially many different types of risks to which an investment may be subject, but they are generally not equally important so far as investment valuation is concerned. Risk that can be avoided by holding an asset in a portfolio, where the risk might be offset by the various risks of other assets, should not be compensated by higher expected return, according to theory. By contrast, investors would expect compensation for bearing an asset’s non-diversifiable risk—systematic risk; only this risk, theory indicates, should be priced risk. In the CAPM, an asset’s systematic risk is a positive function of its beta, which measures the sensitivity of an asset’s return to the market’s return.1 According to the CAPM, differences in mean return are explained by a single factor, the market portfolio return. Greater risk with respect to the market factor—represented by higher beta—is expected to be associated with higher return.
The accumulation of evidence from the equity markets during the decades following the CAPM’s development have provided clear indications that the CAPM provides an incomplete description of risk and that models incorporating multiple sources of systematic risk more effectively model asset returns.2 There are, however, various perspectives in practice on how to model risk in the context of multifactor models. This reading will examine some of these, focusing on macroeconomic factor models and fundamental factor models, in subsequent sections.
3. Arbitrage Pricing Theory In the 1970s, Ross (1976) developed the arbitrage pricing theory (APT) as an alternative to the CAPM. APT introduced a framework that explains the expected return of an asset (or portfolio) in equilibrium as a linear function of the risk of the asset (or portfolio) with respect to a set of factors capturing systematic risk. Unlike the CAPM, the APT does not indicate the identity or even the number of risk factors. Rather, for any multifactor model assumed to generate returns (“return-generating process”), the theory gives the associated expression for the asset’s expected return.
Suppose that K factors are assumed to generate returns. Then the simplest expression for a multifactor model for the return of asset i is given by
(1)
where
Ri = the return to asset i
ai = an intercept term
Ik = the return to factor k, k = 1, 2, ..., K
bik = the sensitivity of the return on asset i to the return to factor k, k = 1, 2, ..., K
εi = an error term with a zero mean that represents the portion of the return to asset i not explained by the factor model
The intercept term ai is the expected return of asset i, given that all the factors take on a value of zero. Equation 1 presents a multifactor return-generating process (a time-series model for returns). In any given period, the model may not account fully for the asset’s return, as indicated by the error term. But error is assumed to average to zero. Another common formulation subtracts the risk-free rate from both sides of Equation 1 so that the dependent variable is the return in excess of the risk-free rate and one of the explanatory variables is a factor return in excess of the risk-free rate. (The Carhart model described below is an example.)
Based on Equation 1, the APT provides an expression for the expected return of asset i assuming that financial markets are in equilibrium. The APT is similar to the CAPM, but the APT makes less strong assumptions than the CAPM.3 The APT makes just three key assumptions:
1. A factor model describes asset returns.
2. There are many assets, so investors can form well-diversified portfolios that eliminate asset-specific risk.
3. No arbitrage opportunities exist among well-diversified portfolios.
Arbitrage is a risk-free operation that requires no net investment of money but earns an expected positive net profit.4 An arbitrage opportunity is an opportunity to conduct an arbitrage—an opportunity to earn an expected positive net profit without risk and with no net investment of money.
In the first assumption, the number of factors is not specified. The second assumption allows investors to form portfolios with factor risk but without asset- specific risk. The third assumption is the condition of financial market equilibrium.
Empirical evidence indicates that Assumption 2 is reasonable. When a portfolio contains many stocks, the asset-specific or non-systematic risk of individual stocks makes almost no contribution to the variance of portfolio returns. Roll and Ross (2001) found that only 1 percent–3 percent of a well-diversified portfolio’s variance comes from the non-systematic variance of the individual stocks in the portfolio, as Exhibit 1 shows.
According to the APT, if the above three assumptions hold, the following equation holds:5
(2)
EXHIBIT 1 Sources of Volatility: The Case of a Well-Diversified Portfolio
Source: “What Is the Arbitrage Pricing Theory.” Retrieved May 25, 2001, from the World Wide Web: www.rollross.com/apt.html. Reprinted with permission of Richard
Roll.
where
E(Rp) = the expected return to portfolio p
RF = the risk-free rate
λj = the expected reward for bearing the risk of factor j
βp,j = the sensitivity of the portfolio to factor j
K = the number of factors
The APT equation, Equation 2, says that the expected return on any well-diversified portfolio is linearly related to the factor sensitivities of that portfolio.
The factor risk premium (or factor price), λj, represents the expected reward for bearing the risk of a portfolio with a sensitivity of 1 to factor j and a sensitivity of 0 to all other factors. The exact interpretation of “expected reward” depends on the multifactor model that is the basis for Equation 2. For example, in the Carhart four- factor model shown later as Equation 3, the risk premium for the market factor is the expected return of the market in excess of the risk-free rate; the factor risk premiums for the other three factors are the mean returns of the specific portfolios held long (e.g., the portfolio of small-cap stocks for the “small minus big” factor) minus the mean return for a related but opposite portfolio (e.g., a portfolio of large- cap stocks, in the case of that factor). A portfolio with a sensitivity of 1 to factor j and a sensitivity of 0 to all other factors is called a pure factor portfolio for factor j (or simply the factor portfolio for factor j).
For example, suppose we have a portfolio with a sensitivity of 1 with respect to Factor 1 and a sensitivity of 0 to all other factors. Using Equation 2, the expected return on this portfolio is E1 = RF + λ1 × 1. If E1 = 0.12 and RF = 0.04, then the risk premium for Factor 1 is
EXAMPLE 1 Determining the Parameters in a One-Factor APT Model
Suppose we have three well-diversified portfolios that are each sensitive to the same single factor. Exhibit 2 shows the expected returns and factor sensitivities
of these portfolios. Assume that the expected returns reflect a one-year investment horizon. To keep the analysis simple, all investors are assumed to agree upon the expected returns of the three portfolios as shown in the exhibit.
Exhibit 2 Sample Portfolios for a One-Factor Model
Portfolio Expected Return Factor Sensitivity A 0.075 0.5 B 0.150 2.0 C 0.070 0.4
We can use these data to determine the parameters of the APT equation. According to Equation 2, for any well-diversified portfolio and assuming a single factor explains returns, we have E(Rp) = RF + λ1βp,1. The factor sensitivities and expected returns are known; thus there are two unknowns, the parameters RF and λ1. Because two points define a straight line, we need to set up only two equations. Selecting Portfolios A and B, we have
and
From the equation for Portfolio A, we have RF = 0.075 – 0.5λ1. Substituting this expression for the risk-free rate into the equation for Portfolio B gives
So we have λ1 = (0.15 – 0.075)/1.5 = 0.05. Substituting this value for λ1 back into the equation for the expected return to Portfolio A yields
So the risk-free rate is 0.05 or 5%, and the factor premium for the common factor is also 0.05 or 5%. The APT equation is
From Exhibit 2, Portfolio C has a factor sensitivity of 0.4. Therefore, according to the APT, the expected return of Portfolio C should be
which is consistent with the expected return for Portfolio C given in Exhibit 2.
EXAMPLE 2 Checking Whether Portfolio Returns Are Consistent with No Arbitrage
In this example, we examine how to tell whether expected returns and factor sensitivities for a set of well-diversified portfolios may indicate the presence of an arbitrage opportunity. Exhibit 3 provides data on four hypothetical portfolios. The data for Portfolios A, B, and C are repeated from Exhibit 2. Portfolio D is a new portfolio. The factor sensitivities given relate to the one- factor APT model E(Rp) = 0.05 + 0.05βp,1 derived in Example 1. As in Example 1, all investors are assumed to agree upon the expected returns of the portfolios. The question raised by the addition of this new Portfolio D is, Has the addition of this portfolio created an arbitrage opportunity? If there exists a portfolio that can be formed from Portfolios A, B, and C that has the same factor sensitivity as Portfolio D but a different expected return, then an arbitrage opportunity exists: Portfolio D would be either undervalued (if it offers a relatively high expected return) or overvalued (if it offers a relatively low expected return).
EXHIBIT 3 Sample Portfolios for a One-Factor Model
Portfolio Expected Return Factor Sensitivity A 0.0750 0.50 B 0.1500 2.00 C 0.0700 0.40 D 0.0800 0.45
0.5A + 0.5C 0.0725 0.45
Exhibit 3 gives data for an equally weighted portfolio of A and C. The expected return and factor sensitivity of this new portfolio are calculated as weighted averages of the expected returns and factor sensitivities of A and C. Expected return is thus (0.50)(0.0750) + (0.50)(0.07) = 0.0725, or 7.25%. The factor sensitivity is (0.50)(0.50) + (0.50)(0.40) = 0.45. Note that the factor sensitivity of 0.45 matches the factor sensitivity of Portfolio D. In this case, the configuration of expected returns in relation to factor risk presents an arbitrage opportunity involving Portfolios A, C, and D. Portfolio D offers, at 8 percent, an expected
return that is too high given its factor sensitivity. According to the assumed APT model, the expected return on Portfolio D should be E(RD) = 0.05 + 0.05βD,1 = 0.05 + (0.05 × 0.45) = 0.0725, or 7.25%. Portfolio D is undervalued relative to its factor risk. We will buy D (hold it long) in the portfolio that exploits the arbitrage opportunity (the arbitrage portfolio). We purchase D using the proceeds from selling short an equally weighted portfolio of A and C with exactly the same 0.45 factor sensitivity as D.
The arbitrage thus involves the following strategy: Invest $10,000 in Portfolio D and fund that investment by selling short an equally weighted portfolio of Portfolios A and C; then close out the investment position at the end of one year (the investment horizon for expected returns). Exhibit 4 demonstrates the arbitrage profits to the arbitrage strategy. The final row of the exhibit shows the net cash flow to the arbitrage portfolio.
EXHIBIT 4 Arbitrage Opportunity within Sample Portfolios
Initial Cash Flow Final Cash Flow Factor Sensitivity Portfolio D –$10,000.00 $10,800.00 0.45
Portfolios A and C $10,000.00 –$10,725.00 –0.45 Sum $0.00 $75.00 0.00
As Exhibit 4 shows, if we buy $10,000 of Portfolio D and sell $10,000 of an equally weighted portfolio of Portfolios A and C, we have an initial net cash flow of $0. The expected value of our investment in Portfolio D at the end of one year is $10,000(1 + 0.08) = $10,800. The expected value of our short position in Portfolios A and C at the end of one year is −$10,000(1.0725) = – $10,725. So, the combined expected cash flow from our investment position in one year is $75.
What about the risk? Exhibit 4 shows that the factor risk has been eliminated: Purchasing D and selling short an equally weighted portfolio of A and C creates a portfolio with a factor sensitivity of 0.45 – 0.45 = 0. The portfolios are well diversified, and we assume any asset-specific risk is negligible.
Because an arbitrage is possible, Portfolios A, C, and D cannot all be consistent with the same equilibrium. If Portfolio D actually had an expected return of 8 percent, investors would bid up its price until the expected return fell and the arbitrage opportunity vanished. Thus, arbitrage restores equilibrium relationships among expected returns.
The Carhart four-factor model, also known as the four-factor model or simply the Carhart model, is a frequently referenced multifactor model in current equity portfolio management practice. Presented in Carhart (1997), it is an extension of the three-factor model developed by Fama and French (1992) to include a momentum factor. According to the model, there are three groups of stocks that tend to have higher returns than those predicted solely by their sensitivity to the market return:
Small-capitalization stocks
Low price-to-book-ratio stocks, commonly referred to as “value” stocks
Stocks whose prices have been rising, commonly referred to as “momentum” stocks
On the basis of that evidence, the Carhart model posits the existence of three systematic risk factors beyond the market risk factor. They are named, in the same order as above, the following:
Small minus big (SMB)
High minus low (HML)
Winners minus losers (WML)
Equation 3a is the Carhart model, in which the excess return on the portfolio is explained as a function of the portfolio’s sensitivity to a market index (RMRF), a market capitalization factor (SMB), a book-value-to-price factor (HML), and a momentum factor (WML).
(3a)
where
Rp and RF = the return on the portfolio and the risk-free rate of return, respectively
ap = “alpha” or return in excess of that expected given the portfolio’s level of systematic risk (assuming the four factors capture all systematic risk)
bp = the sensitivity of the portfolio to the given factor
RMRF = the return on a value-weighted equity index in excess of the one- month T-bill rate
SMB = small minus big, a size (market capitalization) factor; SMB is the average return on three small-cap portfolios minus the average return on three large-cap portfolios
HML = high minus low, the average return on two high book-to-market portfolios minus the average return on two low book-to-market portfolios
WML = winners minus losers, a momentum factor; WML is the return on a portfolio of the past year ’s winners minus the return on a portfolio of the past year ’s losers6
εp = an error term that represents the portion of the return to the portfolio, p, not explained by the model
Following Equation 2, the Carhart model can be stated as giving equilibrium expected return as
(3b)
because the expected value of alpha is zero.
The Carhart model can be viewed as a multifactor extension of the CAPM that explicitly incorporates drivers of differences in expected returns among assets variables that are viewed as anomalies from a pure CAPM perspective. (The term “anomaly” in this context refers to an observed capital market regularity that is not explained by, or contradicts, a theory of asset pricing.) From the perspective of the CAPM, there are size, value, and momentum anomalies. From the perspective of the Carhart model, however, size, value, and momentum represent systematic risk factors; exposure to them is expected to be compensated in the marketplace in the form of differences in mean return.
Size, value, and momentum are common themes in equity portfolio construction, and all three factors continue to have robust uses in active management risk decomposition and return attribution.
4. Multifactor Models: Types Having introduced the APT, it is appropriate to examine the diversity of multifactor models in current use.
In the following sections, we explain the basic principles of multifactor models and discuss various types of models and their application. We also expand on the APT, which relates the expected return of investments to their risk with respect to a set of factors.
4.1. Factors and Types of Multifactor Models Many varieties of multifactor models have been proposed and researched. We can categorize most of them into three main groups according to the type of factor used:
In a macroeconomic factor model, the factors are surprises in macroeconomic variables that significantly explain returns. In the example of equities, the factors can be understood as affecting either the expected future cash flows of companies or the interest rate used to discount these cash flows back to the present. Among macroeconomic factors that have been used are interest rates, inflation risk, business cycle risk, and credit spreads.
In a fundamental factor model, the factors are attributes of stocks or companies that are important in explaining cross-sectional differences in stock prices. Among the fundamental factors that have been used are the book-value- to-price ratio, market capitalization, the price-to-earnings ratio, and financial leverage.
In a statistical factor model, statistical methods are applied to historical returns of a group of securities to extract factors that can explain the observed returns of securities in the group. In statistical factor models, the factors are actually portfolios of the securities in the group under study and are therefore defined by portfolio weights. Two major types of factor models are factor analysis models and principal components models. In factor analysis models, the factors are the portfolios of securities that best explain (reproduce) historical return covariances. In principal components models, the factors are portfolios of securities that best explain (reproduce) the historical return variances.
A potential advantage of statistical factor models is that they make minimal assumptions. But the interpretation of statistical factors is generally difficult, in
contrast to macroeconomic and fundamental factors. A statistical factor that is a portfolio with weights that are similar to market index weights might be interpreted as “the market factor,” for example. But in general, associating a statistical factor with economic meaning may not be possible. Because understanding statistical factor models requires substantial preparation in quantitative methods, a detailed discussion of statistical factor models is outside the scope of this reading.
Our discussion concentrates on macroeconomic factor models and fundamental factor models. Industry use has generally favored fundamental and macroeconomic models, perhaps because such models are much more easily interpreted and rely less on data-mining approaches; nevertheless, statistical factor models have proponents and are also used in practical applications.
4.2. The Structure of Macroeconomic Factor Models The representation of returns in macroeconomic factor models assumes that the returns to each asset are correlated with only the surprises in some factors related to the aggregate economy, such as inflation or real output.7 We can define surprise in general as the actual value minus predicted (or expected) value. A factor ’s surprise is the component of the factor ’s return that was unexpected, and the factor surprises constitute the model’s independent variables. This idea contrasts with the representation of independent variables as returns in Equation 2, reflecting the fact that how the independent variables are represented varies across different types of models.
Suppose that K macro factors explain asset returns. Then in a macroeconomic factor model, Equation 4 expresses the return of asset i:
(4)
where
Ri = the return to asset i
ai = the expected return to asset i
bik = the sensitivity of the return on asset i to a surprise in factor k, k = 1, 2, ..., K
Fk = the surprise in the factor k, k = 1, 2, ..., K
εi = an error term with a zero mean that represents the portion of the return to asset i not explained by the factor model
Surprise in a macroeconomic factor can be illustrated as follows: Suppose we are analyzing monthly returns for stocks. At the beginning of each month, we have a prediction of inflation for the month. The prediction may come from an econometric model or a professional economic forecaster, for example. Suppose our forecast at the beginning of the month is that inflation will be 0.4 percent during the month. At the end of the month, we find that inflation was actually 0.5 percent during the month. During any month,
In this case, actual inflation was 0.5 percent and predicted inflation was 0.4 percent. Therefore, the surprise in inflation was 0.5% – 0.4% = 0.1%.
What is the effect of defining the factors in terms of surprises? Suppose we believe that inflation and gross domestic product (GDP) growth are two factors that carry risk premiums; that is, inflation and GDP represent priced risk. (GDP is a money measure of the goods and services produced within a country’s borders.) We do not use the predicted values of these variables because the predicted values should already be reflected in stock prices and thus in their expected returns. The intercept ai, the expected return to asset i, reflects the effect of the predicted values of the macroeconomic variables on expected stock returns. The surprise in the macroeconomic variables during the month, however, contains new information about the variable. As a result, this model structure analyzes the return to an asset in three components: the asset’s expected return, its unexpected return resulting from new information about the factors, and an error term.
Consider a factor model in which the returns to each asset are correlated with two factors. For example, we might assume that the returns for a particular stock are correlated with surprises in inflation rates and surprises in GDP growth. For stock i, the return to the stock can be modeled as
where
Ri = the return to stock i
ai = the expected return to stock i
bi1 = the sensitivity of the return to stock i to inflation rate surprises
FINFL = the surprise in inflation rates
bi2 = the sensitivity of the return to stock i to GDP growth surprises
FGDP = the surprise in GDP growth (assumed to be uncorrelated with FINT)
εi = an error term with a zero mean that represents the portion of the return to asset i not explained by the factor model
Consider first how to interpret bi1. The factor model predicts that a 1 percentage point surprise in inflation rates will contribute bi1 percentage points to the return to stock i. The slope coefficient bi2 has a similar interpretation relative to the GDP growth factor. Thus, slope coefficients are naturally interpreted as the factor sensitivities of the asset. A factor sensitivity is a measure of the response of return to each unit of increase in a factor, holding all other factors constant. (Factor sensitivities are sometimes called factor betas or factor loadings.)
Now consider how to interpret the intercept ai. Recall that the error term has a mean or average value of zero. If the surprises in both inflation rates and GDP growth are zero, the factor model predicts that the return to asset i will be ai. Thus, ai is the expected value of the return to stock i.
Finally, consider the error term, εi. The intercept ai represents the asset’s expected return. The term (bi1FINT + bi2FGDP) represents the return resulting from factor surprises, and we have interpreted these as the sources of risk shared with other assets. The term εi is the part of return that is unexplained by expected return or the factor surprises. If we have adequately represented the sources of common risk (the factors), then εi must represent an asset-specific risk. For a stock, it might represent the return from an unanticipated company-specific event.
The risk premium for the GDP growth factor is typically positive. The risk premium for the inflation factor, however, is typically negative. Thus, an asset with a positive sensitivity to the inflation factor—an asset with returns that tend to be positive in response to unexpectedly high inflation—would have a lower required return than if its inflation sensitivity were negative; an asset with positive sensitivity to inflation would be in demand for its inflation-hedging ability.
This discussion has broader applications: It can be used for various asset classes, including fixed income and commodities, and can also be used in asset allocation where asset classes can be examined in relation to inflation and GDP growth, as illustrated below. In Exhibit 5, each quadrant reflects a unique mix of inflation and economic growth expectations. Certain asset classes or securities can be expected to perform differently in various inflation and GDP growth regimes and can be plotted in the appropriate quadrant, thus forming a concrete illustration of a two-factor model, as shown below.
In macroeconomic factor models, the time series of factor surprises are constructed first. Regression analysis is then used to estimate assets’ sensitivities to the factors. In practice, estimated sensitivities and intercepts are often acquired from one of the many consulting companies that specialize in factor models. When we have the parameters for the individual assets in a portfolio, we can calculate the portfolio’s parameters as a weighted average of the parameters of individual assets. An individual asset’s weight in that calculation is the proportion of the total market value of the portfolio that the individual asset represents.
Exhibit 5 Growth and Inflation Factor Matrix
Inflation
Growth
Low Inflation/Low Growth
Cash
Government bonds
High Inflation/Low Growth
Inflation-linked bonds
Commodities
Infrastructure
Low Inflation/High Growth
Equity
Corporate debt
High Inflation/High Growth
Real assets (real estate, timberland, farmland, energy)
Note: Entries are assets likely to benefit from the specified combination of growth and inflation.
Example 3 Estimating Returns for a Two-Stock Portfolio Given Factor Sensitivities
Suppose that stock returns are affected by two common factors: surprises in inflation and surprises in GDP growth. A portfolio manager is analyzing the returns on a portfolio of two stocks, Manumatic (MANM) and Nextech (NXT). The following equations describe the returns for those stocks, where the factors FINFL and FGDP represent the surprise in inflation and GDP growth, respectively:
One-third of the portfolio is invested in Manumatic stock, and two-thirds is invested in Nextech stock.
1. Formulate an expression for the return on the portfolio.
2. State the expected return on the portfolio.
3. Calculate the return on the portfolio given that the surprises in inflation and GDP growth are 1 percent and 0 percent, respectively, assuming that the error terms for MANM and NXT both equal 0.5 percent.
In evaluating the equations for surprises in inflation and GDP, amounts stated in percentage terms need to be converted to decimal form.
Solution to 1: The portfolio’s return is the following weighted average of the returns to the two stocks:
Solution to 2: The expected return on the portfolio is 11 percent, the value of the intercept in the expression obtained in the solution to 1.
Solution to 3:
4.3. The Structure of Fundamental Factor Models We earlier gave the equation of a macroeconomic factor model as
We can also represent the structure of fundamental factor models with this equation, but we need to interpret the terms differently.
In fundamental factor models, the factors are stated as returns rather than return surprises in relation to predicted values, so they do not generally have expected values of zero. This approach changes the meaning of the intercept, which is no longer interpreted as the expected return.8
Factor sensitivities are also interpreted differently in most fundamental factor
models. In fundamental factor models, the factor sensitivities are attributes of the security. An asset’s sensitivity to a factor is expressed using a standardized beta, the value of the attribute for the asset minus the average value of the attribute across all stocks divided by the standard deviation of the attribute’s values across all stocks.
(5)
Consider a fundamental model for equities that uses a dividend yield factor. After standardization, a stock with an average dividend yield will have a factor sensitivity of 0, a stock with a dividend yield one standard deviation above the average will have a factor sensitivity of 1, and a stock with a dividend yield one standard deviation below the average will have a factor sensitivity of −1. Suppose, for example, that an investment has a dividend yield of 3.5 percent and that the average dividend yield across all stocks being considered is 2.5 percent. Further, suppose that the standard deviation of dividend yields across all stocks is 2 percent. The investment’s sensitivity to dividend yield is (3.5% − 2.5%)/2% = 0.50, or one-half standard deviation above average. The scaling permits all factor sensitivities to be interpreted similarly, despite differences in units of measure and scale in the variables. The exception to this interpretation is factors for binary variables, such as industry membership. A company either participates in an industry or does not. Industry factor sensitivities are typically modeled by 0−1 dummy variables.9 The sensitivity is 1 if the stock belongs to the industry and 0 if it does not.
A second distinction between macroeconomic multifactor models and fundamental factor models is that with the former, we develop the factor (surprise) series first and then estimate the factor sensitivities through regressions; with the latter, we generally specify the factor sensitivities (attributes) first and then estimate the factor returns through regressions.
Financial analysts use fundamental factor models for a variety of purposes, including portfolio performance attribution and risk analysis. (Performance attribution consists of return attribution and risk attribution. Return attribution is a set of techniques used to identify the sources of the excess return of a portfolio against its benchmark. Risk attribution addresses the sources of risk, identifying the sources of portfolio volatility for absolute mandates and the sources of tracking risk for relative mandates.) Fundamental factor models focus on explaining the returns to individual stocks using observable fundamental factors that describe either attributes of the securities themselves or attributes of the securities’ issuers. Industry membership, price-to-earnings ratio, book-value-to-price ratio, size, and financial leverage are examples of fundamental factors.
Example 4 discusses a study that examined macroeconomic, fundamental, and
statistical factor models.
EXAMPLE 4 Comparing Types of Factor Models
Connor (1995) contrasted a macroeconomic factor model with a fundamental factor model to compare how well the models explain stock returns.
Connor reported the results of applying a macroeconomic factor model to the returns for 779 large-cap US stocks based on monthly data from January 1985 through December 1993. Using five macroeconomic factors, Connor was able to explain approximately 11% of the variance of return on these stocks.10 Exhibit 6 shows his results.
Connor also reported a fundamental factor analysis of the same companies. The factor model employed was the BARRA US-E2 model (as of 2015, the current version is E4). Exhibit 7 shows these results. In the exhibit, “variability in markets” represents the stock’s volatility, “success” is a price momentum variable, “trade activity” distinguishes stocks by how often their shares trade, and “growth” distinguishes stocks by past and anticipated earnings growth.11
Exhibit 6 The Explanatory Power of the Macroeconomic Factors
Factor Explanatory Power
from Using Each Factor Alone
Increase in Explanatory Power from Adding Each Factor to All the
Others Inflation 1.3% 0.0%
Term structure 1.1% 7.7% Industrial production 0.5% 0.3%
Default premium 2.4% 8.1% Unemployment –0.3% 0.1% All factors (total explanatory power)
10.9%
Source: Connor (1995).
Exhibit 7 The Explanatory Power of the Fundamental Factors
Factor Explanatory Power
from Using Each Factor Alone
Increase in Explanatory Power from Adding Each Factor to All the
Others Industries 16.3% 18.0%
Variability in markets 4.3% 0.9%
Success 2.8% 0.8% Size 1.4% 0.6%
Trade activity 1.4% 0.5% Growth 3.0% 0.4%
Earnings to price 2.2% 0.6% Book to price 1.5% 0.6% Earnings variability 2.5% 0.4%
Financial leverage 0.9% 0.5%
Foreign investment 0.7% 0.4%
Labor intensity 2.2% 0.5% Dividend yield 2.9% 0.4% All factors (total explanatory power)
42.6%
Source: Connor (1995).
As Exhibit 7 shows, the most important fundamental factor is “industries,” represented by 55 industry dummy variables. The fundamental factor model explained approximately 43 percent of the variation in stock returns, compared with approximately 11 percent for the macroeconomic factor model. Because “industries” must sum to the market and the market portfolio is not incorporated in the macroeconomic factor model, some advantage to the explanatory power of the fundamental factor may be built into the specific models being compared. Connor ’s article also does not provide tests of the statistical significance of the various factors in either model; however, Connor ’s research is strong evidence for the usefulness of fundamental factor models, and this evidence is mirrored by the wide use of those models in the
investment community. Fundamental factor models are frequently used in portfolio performance attribution, for example. Typically, fundamental factor models employ many more factors than macroeconomic factor models, giving a more detailed picture of the sources of an investment manager ’s returns.
We cannot conclude from this study, however, that fundamental factor models are inherently superior to macroeconomic factor models. Each major type of model has its uses. The factors in various macroeconomic factor models are individually backed by statistical evidence that they represent systematic risk (i.e., risk that cannot be diversified away). The same may not be true of each factor in a fundamental factor model; for example, a portfolio manager can easily construct a portfolio that excludes a particular industry, so exposure to a particular industry is not systematic risk. The two types of factors, macroeconomic and fundamental, have different implications for measuring and managing risk, in general. The macroeconomic factor set is parsimonious (five variables in the model studied) and allows a portfolio manager to incorporate economic views into portfolio construction by adjustments to portfolio exposures to macro factors. The fundamental factor set examined by Connor is large (67 variables, including the 55 industry dummy variables), and at the expense of greater complexity, it can give a more detailed picture of risk in terms that are easily related to company and security characteristics. Connor found that the macroeconomic factor model had no marginal explanatory power when added to the fundamental factor model, implying that the fundamental risk attributes capture all the risk characteristics represented by the macroeconomic factor betas. Because the fundamental factors supply such a detailed description of the characteristics of a stock and its issuer, however, this finding is not necessarily surprising.
We encounter a range of distinct representations of risk in the fundamental models that are currently used in practical applications. Diversity exists in both the identity and exact definition of factors as well as in the underlying functional form and estimation procedures. Despite the diversity, we can place the factors of most fundamental factor models for equities into three broad groups:
Company fundamental factors. These are factors related to the company’s internal performance. Examples are factors relating to earnings growth, earnings variability, earnings momentum, and financial leverage.
Company share-related factors. These factors include valuation measures and other factors related to share price or the trading characteristics of the shares. In contrast to the previous category, these factors directly incorporate
investors’ expectations concerning the company. Examples include price multiples such as earnings yield, dividend yield, and book to market. Market capitalization falls under this heading. Various models incorporate variables relating to share price momentum, share price volatility, and trading activity that fall in this category.
Macroeconomic factors. Sector or industry membership factors fall under this heading. Various models include factors such as CAPM beta, other similar measures of systematic risk, and yield curve level sensitivity, all of which can be placed in this category.
For global factor models in particular, a classification of country, industry, and style factors is often used. In that classification, country and industry factors are dummy variables for country and industry membership, respectively. Style factors include those related to earnings, risk, and valuation that define types of securities typical of various styles of investing.
5. Multifactor Models: Selected Applications The following sections present selected applications of multifactor models in investment practice. The applications discussed are return attribution, risk attribution, portfolio construction, and strategic portfolio decisions.
We begin by discussing portfolio return attribution and risk attribution, focusing on the analysis of benchmark-relative returns.
After discussing performance attribution and risk analysis, we explain the use of multifactor models in creating a portfolio with a desired set of risk exposures.
Additionally, multifactor models can be used for asset allocation purposes. Some large, sophisticated asset owners have chosen to define their asset allocation opportunity sets in terms of macroeconomic or thematic factors and aggregate factor exposures (represented by pure factor portfolios as defined earlier). Many others are examining their traditionally derived asset allocation policies using factor models to map asset class exposure to factor sensitivities. The trend toward factor-based asset allocation has two chief causes: first is the increasing availability of sophisticated factor models (like the BARRA models used in the examples below), and second is the more intense focus by asset owners on the many dimensions of risk.
5.1. Factor Models in Return Attribution Multifactor models can help us understand in detail the sources of a manager ’s returns relative to a benchmark. For simplicity, in this section we analyze the sources of the returns of a portfolio fully invested in the equities of a single national equity market,12 though the same methodology can be applied across asset classes and geographies.
Analysts often favor fundamental multifactor models in decomposing (separating into basic elements) the sources of returns. In contrast to statistical factor models, fundamental factor models allow the sources of portfolio performance to be described using commonly understood terms. Fundamental factors are also thematically understandable and can be incorporated into simple narratives for clients concerning return or risk attribution.
Also, in contrast to macroeconomic factor models, fundamental models express investment style choices and security characteristics more directly and often in greater detail.
We first need to understand the objectives of active managers. As mentioned
previously, managers are commonly evaluated relative to a specified benchmark. Active portfolio managers hold securities in different-from-benchmark weights in an attempt to add value to their portfolios relative to a passive investment approach. Securities held in different-from-benchmark weights reflect portfolio manager expectations that differ from consensus expectations. For an equity manager, those expectations may relate to common factors driving equity returns or to considerations unique to a company. Thus, when we evaluate an active manager, we want to ask questions such as, Did the manager have insights that were effectively translated into returns in excess of those that were available from a passive alternative? Analyzing the sources of returns using multifactor models can help answer these questions.
The return on a portfolio, Rp, can be viewed as the sum of the benchmark’s return, RB, and the active return (portfolio return minus benchmark return):
(6)
With the help of a factor model, we can analyze a portfolio manager ’s active return as the sum of two components. The first component is the product of the portfolio manager ’s factor tilts (over-or underweights relative to the benchmark factor sensitivities) and the factor returns; we call that component the return from factor tilts. The second component of active return reflects the manager ’s skill in individual asset selection (ability to overweight securities that outperform the benchmark or underweight securities that underperform the benchmark); we call that component security selection. Equation 7 shows the decomposition of active return into those two components, where k represents the factor or factors represented in the benchmark portfolio:
(7)
In Equation 7, the portfolio’s and benchmark’s sensitivities to each factor are calculated as of the beginning of the evaluation period.
EXAMPLE 5 Four-Factor Model Active Return Decomposition
As an equity analyst at a pension fund sponsor, Ronald Service uses the Carhart four-factor multifactor model of Equation 3a to evaluate US equity portfolios:
Service’s current task is to evaluate the performance of the most recently hired US equity manager. That manager ’s benchmark is an index representing the performance of the 1,000 largest US stocks by market value. The manager describes himself as a “stock picker” and points to his performance in beating the benchmark as evidence that he is successful. Exhibit 8 presents an analysis based on the Carhart model of the sources of that manager ’s active return during the year, given an assumed set of factor returns. In Exhibit 8, the entry titled “A. Return from Factor Tilts,” equal to 2.1241 percent, is the sum of the four numbers above it. The exhibit (“B. Security Selection”) gives security selection as equal to −0.05 percent. Active return is found as the sum of these two components: 2.1241% + (−0.05%) = 2.0741%.
Exhibit 8 Active Return Decomposition
Factor Sensitivity Contribution to ActiveReturn
Factor Portfolio(1) Benchmark(2) Difference(3)= (1) − (2) Factor
Return(4) Absolute(3)
× (4)
Proportion of Total Active
RMRF 0.95 1.00 −0.05 5.52% −0.2760% −13.3% SMB −1.05 −1.00 −0.05 −3.35% 0.1675% 8.1% HML 0.40 0.00 0.40 5.10% 2.0400% 98.4% WML 0.05 0.03 0.02 9.63% 0.1926% 9.3%
A. Return from Factor Tilts = 2.1241% 102.4%
B. Security Selection = −0.0500% −2.4% C. Active Return (A + B)
= 2.0741% 100.0%
From his previous work, Service knows that the returns to growth-style portfolios often have a positive sensitivity to the momentum factor (WML). By contrast, the returns to certain value-style portfolios, in particular those following a contrarian strategy, often have a negative sensitivity to the momentum factor. Using the information given, address the following questions (assume the benchmark chosen for the manager is appropriate):
1. Determine the manager ’s investment mandate and his actual investment style.
2. Evaluate the sources of the manager ’s active return for the year.
3. What concerns might Service discuss with the manager as a result of the return decomposition?
Solution to 1: The benchmarks chosen for the manager should reflect the baseline risk characteristics of the manager ’s investment opportunity set and his mandate. We can ascertain whether the manager ’s actual style follows the mandate by examining the portfolio’s actual factor exposures:
The sensitivities of the benchmark are consistent with the description in the text. The sensitivity to RMRF of 1 indicates that the assigned benchmark has average market risk, consistent with it being a broad-based index; the negative sensitivity to SMB indicates a large-cap orientation. The mandate might be described as large-cap without a value/growth bias (HML is zero) or a momentum bias (WML is close to zero).
Stocks with high book-to-market ratios are generally viewed as value stocks. Because the equity manager has a positive sensitivity to HML (0.40), it appears that the manager has a value orientation. The manager is approximately neutral to the momentum factor, so the equity manager is not a momentum investor and probably not a contrarian value investor. In summary, the above considerations suggest that the manager has a large-cap value orientation.
Solution to 2: The dominant source of the manager ’s positive active return was his positive active exposure to the HML factor. The bet contributed approximately 98 percent of the realized active return of about 2.07 percent. The manager ’s active exposure to the overall market (RMRF) was unprofitable, but his active exposures to small stocks (SMB) and to momentum (WML) were profitable; however, the magnitudes of the manager ’s active exposures to RMRF, SMB, and WML were relatively small, so the effects of those bets on active return were minor compared with his large and successful bet on HML.
Solution to 3: Although the manager is a self-described “stock picker,” his active return from security selection in this period was actually negative. His positive active return resulted from the concurrence of a large active bet on HML and a high return to that factor during the period. If the market had favored growth rather than value without the manager doing better in individual security selection, the manager ’s performance would have been unsatisfactory. Service’s conversations with the manager should focus on evidence that he can predict changes in returns to the HML factor and on the manager ’s stock selection discipline.
5.2. Factor Models in Risk Attribution Building on the discussion of active returns, this section explores the analysis of active risk. A few key terms are important to the understanding of how factor models are used to build an understanding of a portfolio manager ’s risk exposures. We will describe them briefly before moving on to the detailed discussion of risk attribution.
Active risk can be represented by the standard deviation of active returns. A traditional term for that standard deviation is tracking error (TE). Tracking risk is a synonym for tracking error that is often used in the CFA Program curriculum. We will use the abbreviation TE for the concept of active risk and refer to it usually as tracking error:
(8)
In Equation 8, s(Rp – RB) indicates that we take the sample standard deviation (indicated by s) of the time series of differences between the portfolio return, Rp, and the benchmark return, RB. We should be careful that active return and tracking error are stated on the same time basis.13
As a broad indication of the range for tracking error, in US equity markets a well- executed passive investment strategy can often achieve a tracking error on the order of 0.10 percent or less per annum. A low-risk active or enhanced index investment strategy, which makes tightly controlled use of managers’ expectations, often has a tracking error goal of 2 percent per annum. A diversified active large-cap equity strategy that might be benchmarked to the S&P 500 Index would commonly have a tracking error in the range of 2 percent–6 percent per annum. An aggressive active equity manager might have a tracking error in the range of 6 percent–10 percent or more.
Somewhat analogous to the use of the traditional Sharpe measure in evaluating absolute returns, the information ratio (IR) is a tool for evaluating mean active returns per unit of active risk. The historical or ex post IR is expressed as follows:
(9)
In the numerator of Equation 9, and stand for the sample mean return on the portfolio and the sample mean return on the benchmark, respectively.14 To illustrate the calculation, if a portfolio achieved a mean return of 9 percent during the same period that its benchmark earned a mean return of 7.5 percent and the portfolio’s tracking error (the denominator) was 6%, we would calculate an information ratio
of (9% − 7.5%)/6% = 0.25. Setting guidelines for acceptable active risk or tracking error is one of the methods that some investors use to ensure that the overall risk and style characteristics of their investments are in line with their chosen benchmark.
Note that in addition to focusing exclusively on active risk, multifactor models can also be used to decompose and attribute sources of total risk. For instance, a multi- asset class multi-strategy long/short fund can be evaluated with an appropriate multifactor model to reveal insights on sources of total risk.
EXAMPLE 6 Creating Active Manager Guidelines
The framework of active return and active risk is appealing to investors who want to manage the risk of investments. The benchmark serves as a known and continuously observable reference standard in relation to which quantitative risk and return objectives may be stated and communicated. For example, a US public employee retirement system invited investment managers to submit proposals to manage a “low-active-risk US large-cap equity fund” that would be subject to the following constraints:
Shares must be components of the S&P 500.
The portfolio should have a minimum of 200 issues. At time of purchase, the maximum amount that may be invested in any one issuer is 5 percent of the portfolio at market value or 150 percent of the issuers’ weight within the S&P 500, whichever is greater.
The portfolio must have a minimum information ratio of 0.30 either since inception or over the last seven years.
The portfolio must also have tracking risk of less than 3 percent with respect to the S&P 500 either since inception or over the last seven years.
Once a suitable active manager is found and hired, these requirements can be written into the manager ’s guidelines. The retirement system’s individual mandates would be set such that the sum of mandates across managers would equal the desired risk exposures.
Analysts use multifactor models to understand a portfolio manager ’s risk exposures in detail. By decomposing active risk, the analyst’s objective is to measure the portfolio’s active exposure along each dimension of risk—in other words, to understand the sources of tracking error.15 Among the questions analysts will want
to answer are the following:
What active exposures contributed most to the manager ’s tracking error?
Was the portfolio manager aware of the nature of his active exposures, and if so, can he articulate a rationale for assuming them?
Are the portfolio’s active risk exposures consistent with the manager ’s stated investment philosophy?
Which active bets earned adequate returns for the level of active risk taken?
In addressing these questions, analysts often choose fundamental factor models because they can be used to relate active risk exposures to a manager ’s portfolio decisions in a fairly direct and intuitive way. In this section, we explain how to decompose or explain a portfolio’s active risk using a multifactor model.
We previously addressed the decomposition of active return; now we address the decomposition of active risk. In analyzing risk, it is more convenient to use variances rather than standard deviations because the variances of uncorrelated variables are additive. We refer to the variance of active risk as active risk squared:
(10)
We can separate a portfolio’s active risk squared into two components:
Active factor risk is the contribution to active risk squared resulting from the portfolio’s different-from-benchmark exposures relative to factors specified in the risk model.
Active specific risk or security selection risk measures the active non-factor or residual risk assumed by the manager. Portfolio managers attempt to provide a positive average return from security selection as compensation for assuming active specific risk.
As we use the terms, “active specific risk” and “active factor risk” refer to variances rather than standard deviations. When applied to an investment in a single asset class, active risk squared has two components:
(11)
Active factor risk represents the part of active risk squared explained by the portfolio’s active factor exposures. Active factor risk can be found indirectly as the risk remaining after active specific risk is deducted from active risk squared. Active
specific risk can be expressed as16
where is the ith asset’s active weight in the portfolio (that is, the difference between the asset’s weight in the portfolio and its weight in the benchmark) and is the residual risk of the ith asset (the variance of the ith asset’s returns left unexplained by the factors).
EXAMPLE 7 A Comparison of Active Risk
Richard Gray is comparing the risk of four US equity managers who share the same benchmark. He uses a fundamental factor model, the BARRA US-E4 model, which incorporates 12 style factors and a set of 60 industry factors. The style factors measure various fundamental aspects of companies and their shares, such as size, liquidity, leverage, and dividend yield. In the model, companies have non-zero exposures to all industries in which the company operates. Exhibit 9 presents Gray’s analysis of the active risk squared of the four managers, based on Equation 11.17 In Exhibit 9, the column labeled “Industry” gives the portfolio’s active factor risk associated with the industry exposures of its holdings; the “Style Factor” column gives the portfolio’s active factor risk associated with the exposures of its holdings to the 12 style factors.
Using the information in Exhibit 9, address the following: 1. Contrast the active risk decomposition of Portfolios A and B.
2. Contrast the active risk decomposition of Portfolios B and C.
3. Characterize the investment approach of Portfolio D.
Solution to 1: Exhibit 10 restates the information in Exhibit 9 to show the proportional contributions of the various sources of active risk. (e.g., Portfolio A’s active risk related to industry exposures is 25 percent of active risk squared, calculated as 12.25/49 = 0.25, or 25%).
The last column of Exhibit 10 now shows the square root of active risk squared —that is, active risk or tracking error.
Exhibit 9 Active Risk Squared Decomposition
Active Factor
Portfolio Industry StyleFactor Total Factor
Active Specific
Active Risk Squared
A 12.25 17.15 29.40 19.60 49 B 1.25 13.75 15.00 10.00 25 C 1.25 17.50 18.75 6.25 25 D 0.03 0.47 0.50 0.50 1
Note: Entries are in % squared.
EXHIBIT 10 Active Risk Decomposition (restated)
Active Factor(% of total active)
Portfolio Industry StyleFactor Total Factor
Active Specific(% of total active)
Active Risk
A 25% 35% 60% 40% 7% B 5% 55% 60% 40% 5% C 5% 70% 75% 25% 5% D 3% 47% 50% 50% 1%
Portfolio A has assumed a higher level of active risk than B (7 percent versus 5 percent). Portfolios A and B assumed the same proportions of active factor and active specific risk, but a sharp contrast exists between the two in the types of active factor risk exposure. Portfolio A assumed substantial active industry risk, whereas Portfolio B was approximately industry neutral relative to the benchmark. By contrast, Portfolio B had higher active bets on the style factors representing company and share characteristics.
Solution to 2: Portfolios B and C were similar in their absolute amounts of active risk. Furthermore, both Portfolios B and C were both approximately industry neutral relative to the benchmark. Portfolio C assumed more active factor risk related to the style factors, but B assumed more active specific risk. It is also possible to infer from the greater level of B’s active specific risk that B is somewhat less diversified than C.
Solution to 3: Portfolio D appears to be a passively managed portfolio, judging by its negligible level of active risk. Referring to Exhibit 10, Portfolio D’s
active factor risk of 0.50, equal to 0.707 percent expressed as a standard deviation, indicates that the portfolio’s risk exposures very closely match the benchmark.
The discussion of performance attribution and risk analysis has used examples related to common stock portfolios. Multifactor models have also been effectively used in similar roles for portfolios of bonds and other asset classes. For example, factors such as duration and spread can be used to decompose the risk and return of a fixed-income manager.
5.3. Factor Models in Portfolio Construction Equally as important to the use of multifactor models in analyzing a portfolio’s active returns and active risk is the use of such multifactor models in portfolio construction. At this stage of the portfolio management process, multifactor models permit the portfolio manager to make focused bets or to control portfolio risk relative to the benchmark’s risk. This greater level of detail in modeling risk that multifactor models afford is useful in both passive and active management.
Passive management. In managing a fund that seeks to track an index with many component securities, portfolio managers may need to select a sample of securities from the index. Analysts can use multifactor models to replicate an index fund’s factor exposures, mirroring those of the index tracked.
Active management. Many quantitative investment managers rely on multifactor models in predicting alpha (excess risk-adjusted returns) or relative return (the return on one asset or asset class relative to that of another) as part of a variety of active investment strategies. In constructing portfolios, analysts use multifactor models to establish desired risk profiles.
Rules-based active management (alternative indexes). These strategies routinely tilt toward factors such as size, value, quality, or momentum when constructing portfolios. As such, alternative index approaches aim to capture some systematic exposure traditionally attributed to manager skill, or “alpha,” in a transparent, mechanical, rules-based manner at low cost. Alternative index strategies rely heavily on factor models to introduce intentional factor and style biases versus capitalization-weighted indexes.
In the following, we explore some of these uses in more detail. As indicated, an important use of multifactor models is to establish a specific desired risk profile for a portfolio. In the simplest instance, the portfolio manager may want to create a portfolio with sensitivity to a single factor. This particular (pure) factor portfolio
would have a sensitivity of 1 for that factor and a sensitivity (or weight) of 0 for all other factors. It is thus a portfolio with exposure to only one risk factor and exactly represents the risk of that factor. As a pure bet on a source of risk, factor portfolios are of interest to a portfolio manager who wants to hedge that risk (offset it) or speculate on it. This simple case can be expanded to multiple factors where a factor replication portfolio can be built based either on an existing target portfolio or on a set of desired exposures. Example 8 illustrates the use of factor portfolios.
EXAMPLE 8 Factor Portfolios
Analyst Wanda Smithfield has constructed six portfolios for possible use by portfolio managers in her firm. The portfolios are labeled A, B, C, D, E, and F in Exhibit 11. Smithfield adapts a macroeconomic factor model based on research presented in Burmeister, Roll, and Ross (1994). The model includes five factors:
Confidence risk, based on the yield spread between corporate bonds and government bonds. A positive surprise in the spread suggests that investors are willing to accept a smaller reward for bearing default risk and so that confidence is high.
Time horizon risk, based on the yield spread between 20-year government bonds and 30-day Treasury bills. A positive surprise indicates increased investor willingness to invest for the long term.
Inflation risk, measured by the unanticipated change in the inflation rate.
Business cycle risk, measured by the unexpected change in the level of real business activity.
Market timing risk, measured as the portion of the return on a broad-based equity index that is unexplained by the first four risk factors.
Exhibit 11 Factor Portfolios
Portfolios Risk Factor A B C D E F
Confidence risk 0.50 0.00 1.00 0.00 0.00 0.80 Time horizon risk 1.92 0.00 1.00 1.00 1.00 1.00
Inflation risk 0.00 0.00 1.00 0.00 0.00 −1.05 Business cycle risk 1.00 1.00 0.00 0.00 1.00 0.30 Market timing risk 0.90 0.00 1.00 0.00 0.00 0.75
Note: Entries are factor sensitivities.
1. A portfolio manager wants to place a bet that real business activity will increase.
1. Determine and justify the portfolio among the six given that would be most useful to the manager.
2. Would the manager take a long or short position in the portfolio chosen in Part A?
2. A portfolio manager wants to hedge an existing positive (long) exposure to time horizon risk.
1. Determine and justify the portfolio among the six given that would be most useful to the manager.
2. What type of position would the manager take in the portfolio chosen in Part A?
Solution to 1A: Portfolio B is the most appropriate choice. Portfolio B is the factor portfolio for business cycle risk because it has a sensitivity of 1 to business cycle risk and a sensitivity of 0 to all other risk factors. Portfolio B is thus efficient for placing a pure bet on an increase in real business activity.
Solution to 1B: The manager would take a long position in Portfolio B to place a bet on an increase in real business activity.
Solution to 2A: Portfolio D is the appropriate choice. Portfolio D is the factor portfolio for time horizon risk because it has a sensitivity of 1 to time horizon risk and a sensitivity of 0 to all other risk factors. Portfolio D is thus efficient for hedging an existing positive exposure to time horizon risk.
Solution to 2B: The manager would take a short position in Portfolio D to hedge the positive exposure to time horizon risk.
5.4. How Factor Considerations Can Be Useful in Strategic Portfolio Decisions Multifactor models can help investors recognize considerations that are relevant in
making various strategic decisions. For example, given a sound model of the systematic risk factors that affect assets’ mean returns, the investor can ask, relative to other investors,18
What types of risk do I have a comparative advantage in bearing?
What types of risk am I at a comparative disadvantage in bearing?
For example, university endowments, because they typically have very long investment horizons, may have a comparative advantage in bearing business cycle risk of traded equities or the liquidity risk associated with many private equity investments. They may tilt their strategic asset allocation or investments within an asset class to capture the associated risk premiums for risks that do not much affect them. However, such investors may be at a comparative disadvantage in bearing inflation risk to the extent that the activities they support have historically been subject to cost increases running above the average rate of inflation.
This is a richer framework than that afforded by the CAPM, according to which all investors optimally should invest in two funds: the market portfolio and a risk-free asset. Practically speaking, a CAPM-oriented investor might hold a money market fund and a portfolio of capitalization-weighted broad market indexes across many asset classes, varying the weights in these two in accordance with risk tolerance. These types of considerations are also relevant to individual investors. An individual investor who depends on income from salary or self-employment is sensitive to business cycle risk, in particular to the effects of recessions. If this investor compared two stocks with the same CAPM beta, given his concern about recessions, he might be very sensitive to receiving an adequate premium for investing in procyclical assets. In contrast, an investor with independent wealth and no job-loss concerns would have a comparative advantage in bearing business cycle risk; his optimal risky asset portfolio might be quite different from that of the investor with job-loss concerns in tilting toward greater-than-average exposure to the business cycle factor, all else being equal. Investors should be aware of which priced risks they face and analyze the extent of their exposure.
A multifactor approach can help investors achieve better-diversified and possibly more efficient portfolios. For example, the characteristics of a portfolio can be better explained by a combination of SMB, HML, and WML factors in addition to the market factor than by using the market factor alone.
Thus, compared with single-factor models, multifactor models offer a richer context for investors to search for ways to improve portfolio selection.
6. Summary In this reading, we have presented a set of concepts, models, and tools that are key ingredients to quantitative portfolio management and are used to both construct portfolios as well as to attribute sources of risk and return.
Multifactor models permit a nuanced view of risk that is more granular than the single-factor approach allows.
Multifactor models describe the return on an asset in terms of the risk of the asset with respect to a set of factors. Such models generally include systematic factors, which explain the average returns of a large number of risky assets. Such factors represent priced risk—risk for which investors require an additional return for bearing.
The arbitrage pricing theory (APT) describes the expected return on an asset (or portfolio) as a linear function of the risk of the asset with respect to a set of factors. Like the CAPM, the APT describes a financial market equilibrium, but the APT makes less strong assumptions.
The major assumptions of the APT are as follows:
Asset returns are described by a factor model.
There are many assets, so asset-specific risk can be eliminated.
Assets are priced such that there are no arbitrage opportunities.
Multifactor models are broadly categorized according to the type of factor used as follows:
Macroeconomic factor models
Fundamental factor models
Statistical factor models
In macroeconomic factor models, the factors are surprises in macroeconomic variables that significantly explain asset class (equity in our examples) returns. Surprise is defined as actual minus forecasted value and has an expected value of zero. The factors can be understood as affecting either the expected future cash flows of companies or the interest rate used to discount these cash flows back to the present and are meant to be uncorrelated.
In fundamental factor models, the factors are attributes of stocks or companies that are important in explaining cross-sectional differences in stock prices.
Among the fundamental factors are book-value-to-price ratio, market capitalization, price-to-earnings ratio, and financial leverage.
In contrast to macroeconomic factor models, in fundamental models the factors are calculated as returns rather than surprises. In fundamental factor models, we generally specify the factor sensitivities (attributes) first and then estimate the factor returns through regressions, in contrast to macroeconomic factor models, in which we first develop the factor (surprise) series and then estimate the factor sensitivities through regressions. The factors of most fundamental factor models may be classified as company fundamental factors, company share-related factors, or macroeconomic factors.
In statistical factor models, statistical methods are applied to a set of historical returns to determine portfolios that explain historical returns in one of two senses. In factor analysis models, the factors are the portfolios that best explain (reproduce) historical return covariances. In principal-components models, the factors are portfolios that best explain (reproduce) the historical return variances.
Multifactor models have applications to return attribution, risk attribution, portfolio construction, and strategic investment decisions.
A factor portfolio is a portfolio with unit sensitivity to a factor and zero sensitivity to other factors.
Active return is the return in excess of the return on the benchmark.
Active risk is the standard deviation of active returns. Active risk is also called tracking error or tracking risk. Active risk squared can be decomposed as the sum of active factor risk and active specific risk.
The information ratio (IR) is mean active return divided by active risk (tracking error). The IR measures the increment in mean active return per unit of active risk.
Factor models have uses in constructing portfolios that track market indexes and in alternative index construction.
Traditionally, the CAPM approach would allocate assets between the risk-free asset and a broadly diversified index fund. Considering multiple sources of systematic risk may allow investors to improve on that result by tilting away from the market portfolio. Generally, investors would gain from accepting above average (below average) exposures to risks that they have a comparative advantage (comparative disadvantage) in bearing.
References
1. Bodie, Zvi, Alex Kane, and Alan J. Marcus. 2014. Investments, 10th ed. Boston: Irwin/McGraw-Hill.
2. Burmeister, Edwin, Richard Roll, and Stephen A. Ross. 1994. “A Practitioner ’s Guide to Arbitrage Pricing Theory.” In A Practitioner’s Guide to Factor Models. Charlottesville, VA: Research Foundation of the Institute of Chartered Financial Analysts.
3. Carhart, Mark M. 1997. “On Persistence in Mutual Fund Performance.” Journal of Finance, vol. 52, no. 1: 57–82.
4. Connor, Gregory. 1995. “The Three Types of Factor Models: A Comparison of Their Explanatory Power.” Financial Analysts Journal, vol. 51, no. 3: 42–46.
5. Fama, Eugene F., and Kenneth R. French. 1992. “The Cross-Section of Expected Stock Returns.” Journal of Finance, vol. 47, no. 2: 427–465.
6. Fischer, B., and R. Wermers. 2013. Performance Evaluation and Attribution of Security Portfolios. Oxford, UK: Elsevier.
7. Grinold, Richard, and Ronald N. Kahn. 1994. “MultiFactor Models for Portfolio Risk.” In A Practitioner’s Guide to Factor Models. Charlottesville, VA: Research Foundation of the Institute of Chartered Financial Analysts.
8. Roll, Richard, and Stephen A. Ross. 2001. “What Is the Arbitrage Pricing Theory?” Retrieved 25 May 2001 from www.rollross.com/apt.html.
9. Ross, S. A. 1976. “The Arbitrage Theory of Capital Asset Pricing.” Journal of Economic Theory, vol. 13, no. 3: 341–360.
Problems
1. Compare the assumptions of the arbitrage pricing theory (APT) with those of the capital asset pricing model (CAPM).
2. Last year the return on Harry Company stock was 5 percent. The portion of the return on the stock not explained by a two-factor macroeconomic factor model was 3 percent. Using the data given below, calculate Harry Company stock’s expected return.
Macroeconomic Factor Model for Harry Company Stock
Variable Actual Value(%) Expected Value
(%) Stock’s Factor Sensitivity
Change in interest rate 2.0 0.0 –1.5
Growth in GDP 1.0 4.0 2.0
3. Assume that the following one-factor model describes the expected return for portfolios:
Also assume that all investors agree on the expected returns and factor sensitivity of the three highly diversified Portfolios A, B, and C given in the following table:
Portfolio Expected Return Factor Sensitivity A 0.20 0.80 B 0.15 1.00 C 0.24 1.20
Assuming the one-factor model is correct and based on the data provided for Portfolios A, B, and C, determine if an arbitrage opportunity exists and explain how it might be exploited.
4. Which type of factor model is most directly applicable to an analysis of the style orientation (for example, growth vs. value) of an active equity investment manager? Justify your answer.
5. Suppose an active equity manager has earned an active return of 110 basis
points, of which 80 basis points is the result of security selection ability. Explain the likely source of the remaining 30 basis points of active return.
6. Address the following questions about the information ratio.
1. What is the information ratio of an index fund that effectively meets its investment objective?
2. What are the two types of risk an active investment manager can assume in seeking to increase his information ratio?
7. A wealthy investor has no other source of income beyond her investments and that income is expected to reliably meet all her needs. Her investment advisor recommends that she tilt her portfolio to cyclical stocks and high-yield bonds. Explain the advisor ’s advice in terms of comparative advantage in bearing risk.
1 Beta can be analyzed as the correlation of the asset’s returns with the market return multiplied by a constant that increases with the asset’s return standard deviation.
2 See Bodie, Kane, and Marcus (2014) for an introduction to the empirical evidence.
3 The CAPM can be viewed as a special case of the APT that results from making a set of strong additional assumptions about investors and markets.
4 The word “arbitrage” or the phrase “risk arbitrage” is also sometimes used in practice to describe investment operations in which significant risk is present.
5 A risk-free asset is assumed. If no risk-free asset exists, in place of RF we write λ0 to represent the expected return on a risky portfolio with zero sensitivity to all the factors. The number of factors is not specified but must be much lower than the number of assets, a condition fulfilled in practice.
6 WML is an equally weighted average of the stocks with the highest 30 percent month returns lagged 1 month minus the equally weighted average of the stocks with the lowest 30 percent 11-month returns lagged 1 month. WML has the label PR1YR in the original paper by Carhart.
7 See, for example, Burmeister, Roll, and Ross (1994).
8 If the coefficients were not standardized as described in the following paragraph, the intercept could be interpreted as the risk-free rate, because it would be the return to an asset with no factor risk (zero factor betas) and no asset-specific risk. With standardized coefficients, the intercept is not interpreted beyond being an intercept in a regression included so that the expected asset-specific risk
equals zero.
9 To further explain 0–1 variables, industry membership is measured on a nominal scale because measurement consists only in identifying the industry to which a company belongs. A nominal variable can be represented in a regression by a dummy variable (a variable that takes on the value of 0 or 1). For more on dummy variables, see the reading on multiple regression.
10 The explanatory power of a given model was computed as 1 – [(Average asset- specific variance of return across stocks)/(Average total variance of return across stocks)]. The variance estimates were corrected for degrees of freedom, so the marginal contribution of a factor to explanatory power can be zero or negative. Explanatory power captures the proportion of the total variance of return that a given model explains for the average stock.
11 The explanations of the variables are from Grinold and Kahn (1994); Connor did not supply definitions.
12 This assumption allows us to ignore the roles of country selection, asset allocation, market timing, and currency hedging, greatly simplifying the analysis. However, we can perform similar analyses using multifactor models in a more general context.
13 As an approximation assuming returns are serially uncorrelated, to annualize a daily TE based on daily returns, we multiply daily TE by (250)1/2 based on 250 trading days in a year; to annualize a monthly TE based on monthly returns, we multiply monthly TE by (12)1/2.
14 The expression for IR given here assumes that the portfolio being evaluated has the same systematic risk as its benchmark. There is also a more precise form of the information ratio that corrects for any differences in systematic risk relative to the benchmark. This form has alpha, rather than active return, in the numerator and the portfolio’s residual (non-systematic) risk in the denominator. See Fischer and Wermers (2013, pp. 75–80) for a detailed treatment of the IR. The IR, especially in the more precise form, is also known as the Treynor–Black appraisal ratio.
15 The portfolio’s active risks are weighted averages of the component securities’ active risk. Therefore, we may also perform the analysis at the level of individual holdings. A portfolio manager may find this approach useful in making adjustments to his active risk profile.
16 The direct procedure for calculating active factor risk is as follows. A portfolio’s active factor exposure to a given factor j, , is found by weighting each asset’s
sensitivity to factor j by its active weight and summing the terms: . Then
active factor risk equals .
17 There is a covariance term in active factor risk, reflecting the correlation of industry membership and the risk indexes, which we assume is negligible in this example.
18 Passive management is a distinct issue from holding a single portfolio. There are efficient-markets arguments for holding indexed investments that are separate from the CAPM. However, an index fund is reasonable for this investor.
Appendices Appendices
Appendix A Cumulative Probabilities for a Standard Normal Distribution Appendix B Table of the Student’s t-Distribution (One-Tailed Probabilities) Appendix C Values of X2 (Degrees of Freedom, Level of Significance) Appendix D Table of the F-Distribution Appendix E Critical Values for the Durbin-Watson Statistic (α = .05)
Appendix A Cumulative Probabilities for a Standard Normal Distribution P(Z ≤ x) = N(x) for x ≥ 0 or P(Z ≤ z) = N(z) for z ≥ 0
Appendix B Table of the Student’s t-Distribution (One-Tailed Probabilities)
Appendix C Values of χ2 (Degrees of Freedom, Level of Significance)
Appendix DTable of the F-Distribution
Appendix E Critical Values for the Durbin-Watson Statistic (α = .05)
Glossary A priori probability A probability based on logical analysis rather than on observation or personal judgment.
Absolute dispersion The amount of variability present without comparison to any reference point or benchmark.
Absolute frequency The number of observations in a given interval (for grouped data).
Accrued interest Interest earned but not yet paid.
Active factor risk The contribution to active risk squared resulting from the portfolio’s different-than-benchmark exposures relative to factors specified in the risk model.
Active return The return on a portfolio minus the return on the portfolio’s benchmark.
Active risk The standard deviation of active returns.
Active risk squared The variance of active returns; active risk raised to the second power.
Active specific risk The contribution to active risk squared resulting from the portfolio’s active weights on individual assets as those weights interact with assets’ residual risk.
Addition rule for probabilities A principle stating that the probability that A or B occurs (both occur) equals the probability that A occurs, plus the probability that B occurs, minus the probability that both A and B occur.
Adjusted R2 A measure of goodness-of-fit of a regression that is adjusted for degrees of freedom and hence does not automatically increase when another independent variable is added to a regression.
Analysis of variance (ANOVA) The analysis of the total variability of a dataset (such as observations on the dependent variable in a regression) into components representing different sources of variation; with reference to regression, ANOVA provides the inputs for an F-test of the significance of the regression as a whole.
Annual percentage rate The cost of borrowing expressed as a yearly rate.
Annuity due An annuity having a first cash flow that is paid immediately.
Annuity A finite set of level sequential cash flows.
Arbitrage 1) The simultaneous purchase of an undervalued asset or portfolio and sale of an overvalued but equivalent asset or portfolio, in order to obtain a riskless profit on the price differential. Taking advantage of a market inefficiency in a risk- free manner. 2) The condition in a financial market in which equivalent assets or combinations of assets sell for two different prices, creating an opportunity to profit at no risk with no commitment of money. In a well-functioning financial market, few arbitrage opportunities are possible. 3) A risk-free operation that earns an expected positive net profit but requires no net investment of money.
Arbitrage opportunity An opportunity to conduct an arbitrage; an opportunity to earn an expected positive net profit without risk and with no net investment of money.
Arbitrage portfolio The portfolio that exploits an arbitrage opportunity.
Arithmetic mean The sum of the observations divided by the number of observations.
Asian call option A European-style option with a value at maturity equal to the difference between the stock price at maturity and the average stock price during the life of the option, or $0, whichever is greater.
Autocorrelation The correlation of a time series with its own past values.
Autoregressive model (AR) A time series regressed on its own past values, in which the independent variable is a lagged value of the dependent variable.
Back simulation Another term for the historical method of estimating VAR. This term is somewhat misleading in that the method involves not a simulation of the past but rather what actually happened in the past, sometimes adjusted to reflect the fact that a different portfolio may have existed in the past than is planned for the future.
Bank discount basis A quoting convention that annualizes, on a 360-day year, the discount as a percentage of face value.
Bernoulli random variable A random variable having the outcomes 0 and 1.
Bernoulli trial An experiment that can produce one of two outcomes.
Binomial model A model for pricing options in which the underlying price can move to only one of two possible new prices.
Binomial random variable The number of successes in n Bernoulli trials for which the probability of success is constant for all trials and the trials are independent.
Binomial tree The graphical representation of a model of asset price dynamics in
which, at each period, the asset moves up with probability p or down with probability (1 – p).
Bond equivalent yield A calculation of yield that is annualized using the ratio of 365 to the number of days to maturity. Bond equivalent yield allows for the restatement and comparison of securities with different compounding periods.
Breusch–Pagan test A test for conditional heteroskedasticity in the error term of a regression.
Capital budgeting The allocation of funds to relatively long-range projects or investments.
Capital structure The mix of debt and equity that a company uses to finance its business; a company’s specific mixture of long-term financing.
Cash flow additivity principle The principle that dollar amounts indexed at the same point in time are additive.
CD equivalent yield A yield on a basis comparable to the quoted yield on an interest-bearing money market instrument that pays interest on a 360-day basis; the annualized holding period yield, assuming a 360-day year.
Chain rule of forecasting A forecasting process in which the next period’s value as predicted by the forecasting equation is substituted into the right-hand side of the equation to give a predicted value two periods ahead.
Coefficient of variation (CV) The ratio of a set of observations’ standard deviation to the observations’ mean value.
Cointegrated Describes two time series that have a long-term financial or economic relationship such that they do not diverge from each other without bound in the long run.
Combination A listing in which the order of the listed items does not matter.
Common size statements Financial statements in which all elements (accounts) are stated as a percentage of a key figure such as revenue for an income statement or total assets for a balance sheet.
Company fundamental factors Factors related to the company’s internal performance, such as factors relating to earnings growth, earnings variability, earnings momentum, and financial leverage.
Company share-related factors Valuation measures and other factors related to share price or the trading characteristics of the shares, such as earnings yield, dividend yield, and book-to-market value.
Complements Said of goods which tend to be used together; technically, two goods whose cross-price elasticity of demand is negative.
Compounding The process of accumulating interest on interest.
Conditional expected value The expected value of a stated event given that another event has occurred.
Conditional heteroskedasticity Heteroskedasticity in the error variance that is correlated with the values of the independent variable(s) in the regression.
Conditional probability The probability of an event given (conditioned on) another event.
Conditional variances The variance of one variable, given the outcome of another.
Consistent With reference to estimators, describes an estimator for which the probability of estimates close to the value of the population parameter increases as sample size increases.
Continuous random variable A random variable for which the range of possible outcomes is the real line (all real numbers between −∞ and +∞ or some subset of the real line).
Continuous time Time thought of as advancing in extremely small increments.
Continuously compounded return The natural logarithm of 1 plus the holding period return, or equivalently, the natural logarithm of the ending price over the beginning price.
Correlation analysis The analysis of the strength of the linear relationship between two data series.
Correlation A number between −1 and +1 that measures the comovement (linear association) between two random variables.
Cost averaging The periodic investment of a fixed amount of money.
Covariance matrix A matrix or square array whose entries are covariances; also known as a variance–covariance matrix.
Covariance stationary Describes a time series when its expected value and variance are constant and finite in all periods and when its covariance with itself for a fixed number of periods in the past or future is constant and finite in all periods.
Cross-sectional data Observations over individual units at a point in time, as opposed to time-series data.
Cumulative distribution function A function giving the probability that a random
variable is less than or equal to a specified value.
Cumulative relative frequency For data grouped into intervals, the fraction of total observations that are less than the value of the upper limit of a stated interval.
Data mining The practice of determining a model by extensive searching through a dataset for statistically significant patterns. Also called data snooping.
Default risk premium An extra return that compensates investors for the possibility that the borrower will fail to make a promised payment at the contracted time and in the contracted amount.
Degree of confidence The probability that a confidence interval includes the unknown population parameter.
Degrees of freedom (df) The number of independent observations used.
Dependent With reference to events, the property that the probability of one event occurring depends on (is related to) the occurrence of another event.
Dependent variable The variable whose variation about its mean is to be explained by the regression; the left-hand-side variable in a regression equation.
Descriptive statistics The study of how data can be summarized effectively.
Diffuse prior The assumption of equal prior probabilities.
Discount To reduce the value of a future payment in allowance for how far away it is in time; to calculate the present value of some future amount. Also, the amount by which an instrument is priced below its face (par) value.
Discrete random variable A random variable that can take on at most a countable number of possible values.
Discriminant analysis A multivariate classification technique used to discriminate between groups, such as companies that either will or will not become bankrupt during some time frame.
Dispersion The variability around the central tendency.
Down transition probability The probability that an asset’s value moves down in a model of asset price dynamics.
Dummy variable A type of qualitative variable that takes on a value of 1 if a particular condition is true and 0 if that condition is false.
Dutch Book theorem A result in probability theory stating that inconsistent probabilities create profit opportunities.
Economic growth The expansion of production possibilities that results from
capital accumulation and technological change.
Effective annual rate The amount by which a unit of currency will grow in a year with interest on interest included.
Effective annual yield (EAY) An annualized return that accounts for the effect of interest on interest; EAY is computed by compounding 1 plus the holding period yield forward to one year, then subtracting 1.
Empirical probability The probability of an event estimated as a relative frequency of occurrence.
Error autocorrelation The autocorrelation of the error term.
Error term The portion of the dependent variable that is not explained by the independent variable(s) in the regression.
Estimate The particular value calculated from sample observations using an estimator.
Estimated parameters With reference to a regression analysis, the estimated values of the population intercept and population slope coefficient(s) in a regression.
Estimation With reference to statistical inference, the subdivision dealing with estimating the value of a population parameter.
Estimator An estimation formula; the formula used to compute the sample mean and other sample statistics are examples of estimators.
Event Any outcome or specified set of outcomes of a random variable.
Excess kurtosis Degree of peakedness (fatness of tails) in excess of the peakedness of the normal distribution.
Exhaustive Covering or containing all possible outcomes.
Expected value The probability-weighted average of the possible outcomes of a random variable.
Face value The amount of cash payable by a company to the bondholders when the bonds mature; the promised payment at maturity separate from any coupon payment.
Factor A common or underlying element with which several variables are correlated.
Factor portfolio See pure factor portfolio.
Factor price The expected return in excess of the risk-free rate for a portfolio with a sensitivity of 1 to one factor and a sensitivity of 0 to all other factors.
Factor risk premium The expected return in excess of the risk-free rate for a portfolio with a sensitivity of 1 to one factor and a sensitivity of 0 to all other factors. Also called factor price.
First-differencing A transformation that subtracts the value of the time series in period t – 1 from its value in period t.
First-order serial correlation Correlation between adjacent observations in a time series.
Fitted parameters With reference to a regression analysis, the estimated values of the population intercept and population slope coefficient(s) in a regression.
Fractile A value at or below which a stated fraction of the data lies.
Frequency distribution A tabular display of data summarized into a relatively small number of intervals.
Frequency polygon A graph of a frequency distribution obtained by drawing straight lines joining successive points representing the class frequencies.
Full price The price of a security with accrued interest; also called the invoice or dirty price.
Fundamental factor models A multifactor model in which the factors are attributes of stocks or companies that are important in explaining cross-sectional differences in stock prices.
Future value (FV) The amount to which a payment or series of payments will grow by a stated future date.
Generalized least squares A regression estimation technique that addresses heteroskedasticity of the error term.
Geometric mean A measure of central tendency computed by taking the nth root of the product of n non-negative values.
Harmonic mean A type of weighted mean computed by averaging the reciprocals of the observations, then taking the reciprocal of that average.
Heteroskedastic With reference to the error term of regression, having a variance that differs across observations.
Heteroskedasticity-consistent standard errors Standard errors of the estimated parameters of a regression that correct for the presence of heteroskedasticity in the regression’s error term.
Heteroskedasticity The property of having a nonconstant variance; refers to an error term with the property that its variance differs across observations.
Histogram A bar chart of data that have been grouped into a frequency distribution.
Historical simulation Another term for the historical method of estimating VAR. This term is somewhat misleading in that the method involves not a simulation of the past but rather what actually happened in the past, sometimes adjusted to reflect the fact that a different portfolio may have existed in the past than is planned for the future.
Holding period return The return that an investor earns during a specified holding period; a synonym for total return.
Holding period yield (HPY) The return that an investor earns during a specified holding period; holding period return with reference to a fixed-income instrument.
Homoskedasticity The property of having a constant variance; refers to an error term that is constant across observations.
Hurdle rate The rate of return that must be met for a project to be accepted.
Hypothesis testing With reference to statistical inference, the subdivision dealing with the testing of hypotheses about one or more populations.
Hypothesis With reference to statistical inference, a statement about one or more populations.
In-sample forecast errors The residuals from a fitted time-series model within the sample period used to fit the model.
Incremental cash flow The cash flow that is realized because of a decision; the changes or increments to cash flows resulting from a decision or action.
Independent With reference to events, the property that the occurrence of one event does not affect the probability of another event occurring.
Independent variable A variable used to explain the dependent variable in a regression; a right-hand-side variable in a regression equation.
Independently and identically distributed (IID) With respect to random variables, the property of random variables that are independent of each other but follow the identical probability distribution.
Indexing An investment strategy in which an investor constructs a portfolio to mirror the performance of a specified index.
Inflation premium An extra return that compensates investors for expected inflation.
Information ratio (IR) Mean active return divided by active risk; or alpha divided by the standard deviation of diversifiable risk.
Interest rate A rate of return that reflects the relationship between differently dated cash flows; a discount rate.
Intergenerational data mining A form of data mining that applies information developed by previous researchers using a dataset to guide current research using the same or a related dataset.
Internal rate of return (IRR) The discount rate that makes net present value equal 0; the discount rate that makes the present value of an investment’s costs (outflows) equal to the present value of the investment’s benefits (inflows).
Interquartile range The difference between the third and first quartiles of a dataset.
Interval With reference to grouped data, a set of values within which an observation falls.
Interval scale A measurement scale that not only ranks data but also gives assurance that the differences between scale values are equal.
IRR rule An investment decision rule that accepts projects or investments for which the IRR is greater than the opportunity cost of capital.
Joint probability function A function giving the probability of joint occurrences of values of stated random variables.
Joint probability The probability of the joint occurrence of stated events.
kth order autocorrelation The correlation between observations in a time series separated by k periods.
Kurtosis The statistical measure that indicates the peakedness of a distribution.
Leptokurtic Describes a distribution that is more peaked than a normal distribution.
Level of significance The probability of a Type I error in testing a hypothesis.
Likelihood The probability of an observation, given a particular set of conditions.
Linear association A straight-line relationship, as opposed to a relationship that cannot be graphed as a straight line.
Linear interpolation The estimation of an unknown value on the basis of two known values that bracket it, using a straight line between the two known values.
Linear regression Regression that models the straight-line relationship between the dependent and independent variable(s).
Linear trend A trend in which the dependent variable changes at a constant rate with time.
Liquidity premium An extra return that compensates investors for the risk of loss relative to an investment’s fair value if the investment needs to be converted to cash quickly.
Log-linear model With reference to time-series models, a model in which the growth rate of the time series as a function of time is constant.
Log-log regression model A regression that expresses the dependent and independent variables as natural logarithms.
Logit model A qualitative-dependent-variable multiple regression model based on the logistic probability distribution.
Longitudinal data Observations on characteristic(s) of the same observational unit through time.
Look-ahead bias A bias caused by using information that was unavailable on the test date.
Lower bound The lowest possible value of an option.
Macroeconomic factor model A multifactor model in which the factors are surprises in macroeconomic variables that significantly explain equity returns.
Macroeconomic factors Factors related to the economy, such as the inflation rate, industrial production, or economic sector membership.
Marginal probability The probability of an event not conditioned on another event.
Market timing Asset allocation in which the investment in the market is increased if one forecasts that the market will outperform T-bills.
Maturity premium An extra return that compensates investors for the increased sensitivity of the market value of debt to a change in market interest rates as maturity is extended.
Mean absolute deviation With reference to a sample, the mean of the absolute values of deviations from the sample mean.
Mean excess return The average rate of return in excess of the risk-free rate.
Mean reversion The tendency of a time series to fall when its level is above its mean and rise when its level is below its mean; a mean-reverting time series tends to return to its long-term mean.
Mean–variance analysis An approach to portfolio analysis using expected means, variances, and covariances of asset returns.
Measure of central tendency A quantitative measure that specifies where data are
centered.
Measurement scales A scheme of measuring differences. The four types of measurement scales are nominal, ordinal, interval, and ratio.
Measures of location A quantitative measure that describes the location or distribution of data; includes not only measures of central tendency but also other measures such as percentiles.
Median The value of the middle item of a set of items that has been sorted into ascending or descending order; the 50th percentile.
Mesokurtic Describes a distribution with kurtosis identical to that of the normal distribution.
Modal interval With reference to grouped data, the most frequently occurring interval.
Mode The most frequently occurring value in a set of observations.
Model specification With reference to regression, the set of variables included in the regression and the regression equation’s functional form.
Monetary policy Actions taken by a nation’s central bank to affect aggregate output and prices through changes in bank reserves, reserve requirements, or its target interest rate.
Money market The market for short-term debt instruments (one-year maturity or less).
Money market yield A yield on a basis comparable to the quoted yield on an interest-bearing money market instrument that pays interest on a 360-day basis; the annualized holding period yield, assuming a 360-day year.
Monte Carlo simulation An approach to estimating a probability distribution of outcomes to examine what might happen if particular risks are faced. This method is widely used in the sciences as well as in business to study a variety of problems.
Multicollinearity A regression assumption violation that occurs when two or more independent variables (or combinations of independent variables) are highly but not perfectly correlated with each other.
Multiple linear regression Linear regression involving two or more independent variables.
Multiple linear regression model A linear regression model with two or more independent variables.
Multiplication rule for probabilities The rule that the joint probability of events A
and B equals the probability of A given B times the probability of B.
Multivariate distribution A probability distribution that specifies the probabilities for a group of related random variables.
Multivariate normal distribution A probability distribution for a group of random variables that is completely defined by the means and variances of the variables plus all the correlations between pairs of the variables.
Mutually exclusive projects Mutually exclusive projects compete directly with each other. For example, if Projects A and B are mutually exclusive, you can choose A or B, but you cannot choose both.
n Factorial For a positive integer n, the product of the first n positive integers; 0 factorial equals 1 by definition. n factorial is written as n!.
n-Period moving average The average of the current and immediately prior n – 1 values of a time series.
Negative serial correlation Serial correlation in which a positive error for one observation increases the chance of a negative error for another observation, and vice versa.
Net present value (NPV) The present value of an investment’s cash inflows (benefits) minus the present value of its cash outflows (costs).
Node Each value on a binomial tree from which successive moves or outcomes branch.
Nominal risk-free interest rate The sum of the real risk-free interest rate and the inflation premium.
Nominal scale A measurement scale that categorizes data but does not rank them.
Nonlinear relation An association or relationship between variables that cannot be graphed as a straight line.
Nonparametric test A test that is not concerned with a parameter, or that makes minimal assumptions about the population from which a sample comes.
Nonstationarity With reference to a random variable, the property of having characteristics such as mean and variance that are not constant through time.
NPV rule An investment decision rule that states that an investment should be undertaken if its NPV is positive but not undertaken if its NPV is negative.
Objective probabilities Probabilities that generally do not vary from person to person; includes a priori and objective probabilities.
One-sided hypothesis test A test in which the null hypothesis is rejected only if the evidence indicates that the population parameter is greater than (smaller than) θ0. The alternative hypothesis also has one side.
One-tailed hypothesis test A test in which the null hypothesis is rejected only if the evidence indicates that the population parameter is greater than (smaller than) θ0. The alternative hypothesis also has one side.
Opportunity cost The value that investors forgo by choosing a particular course of action; the value of something in its best alternative use.
Ordinal scale A measurement scale that sorts data into categories that are ordered (ranked) with respect to some characteristic.
Ordinary annuity An annuity with a first cash flow that is paid one period from the present.
Out-of-sample forecast errors The differences between actual and predicted value of time series outside the sample period used to fit the model.
Out-of-sample test A test of a strategy or model using a sample outside the time period on which the strategy or model was developed.
Outcome A possible value of a random variable.
Paired comparisons test A statistical test for differences based on paired observations drawn from samples that are dependent on each other.
Paired observations Observations that are dependent on each other.
Pairs arbitrage trade A trade in two closely related stocks involving the short sale of one and the purchase of the other.
Panel data Observations through time on a single characteristic of multiple observational units.
Parameter A descriptive measure computed from or used to describe a population of data, conventionally represented by Greek letters.
Parameter instability The problem or issue of population regression parameters that have changed over time.
Parametric test Any test (or procedure) concerned with parameters or whose validity depends on assumptions concerning the population generating the sample.
Partial regression coefficients The slope coefficients in a multiple regression. Also called partial slope coefficients.
Partial slope coefficients The slope coefficients in a multiple regression. Also
called partial regression coefficients.
Percentiles Quantiles that divide a distribution into 100 equal parts.
Performance appraisal The evaluation of risk-adjusted performance; the evaluation of investment skill.
Performance measurement The calculation of returns in a logical and consistent manner.
Permutation An ordered listing.
Perpetuity A perpetual annuity, or a set of never-ending level sequential cash flows, with the first cash flow occurring one period from now. A bond that does not mature.
Platykurtic Describes a distribution that is less peaked than the normal distribution.
Point estimate A single numerical estimate of an unknown quantity, such as a population parameter.
Population All members of a specified group.
Population mean The arithmetic mean value of a population; the arithmetic mean of all the observations or values in the population.
Population standard deviation A measure of dispersion relating to a population in the same unit of measurement as the observations, calculated as the positive square root of the population variance.
Population variance A measure of dispersion relating to a population, calculated as the mean of the squared deviations around the population mean.
Positive serial correlation Serial correlation in which a positive error for one observation increases the chance of a positive error for another observation, and a negative error for one observation increases the chance of a negative error for another observation.
Posterior probability An updated probability that reflects or comes after new information.
Power of a test The probability of correctly rejecting the null—that is, rejecting the null hypothesis when it is false.
Present value (PV) The present discounted value of future cash flows: For assets, the present discounted value of the future net cash inflows that the asset is expected to generate; for liabilities, the present discounted value of the future net cash outflows that are expected to be required to settle the liabilities.
Price relative A ratio of an ending price over a beginning price; it is equal to 1 plus the holding period return on the asset.
Priced risk Risk for which investors demand compensation for bearing (e.g., equity risk, company-specific factors, macroeconomic factors).
Principal The amount of funds originally invested in a project or instrument; the face value to be paid at maturity.
Prior probabilities Probabilities reflecting beliefs prior to the arrival of new information.
Probability density function A function with non-negative values such that probability can be described by areas under the curve graphing the function.
Probability distribution A distribution that specifies the probabilities of a random variable’s possible outcomes.
Probability function A function that specifies the probability that the random variable takes on a specific value.
Probability A number between 0 and 1 describing the chance that a stated event will occur.
Probit model A qualitative-dependent-variable multiple regression model based on the normal distribution.
Pseudo-random numbers Numbers produced by random number generators.
Pure discount instruments Instruments that pay interest as the difference between the amount borrowed and the amount paid back.
Pure factor portfolio A portfolio with sensitivity of 1 to the factor in question and a sensitivity of 0 to all other factors.
Qualitative dependent variables Dummy variables used as dependent variables rather than as independent variables.
Quantile A value at or below which a stated fraction of the data lies. Also called fractile.
Quartiles Quantiles that divide a distribution into four equal parts.
Quintiles Quantiles that divide a distribution into five equal parts.
Quoted interest rate A quoted interest rate that does not account for compounding within the year. Also called stated annual interest rate.
Random number generator An algorithm that produces uniformly distributed random numbers between 0 and 1.
Random number An observation drawn from a uniform distribution.
Random variable A quantity whose future outcomes are uncertain.
Random walk A time series in which the value of the series in one period is the value of the series in the previous period plus an unpredictable random error.
Range The difference between the maximum and minimum values in a dataset.
Ratio scales A measurement scale that has all the characteristics of interval measurement scales as well as a true zero point as the origin.
Real risk-free interest rate The single-period interest rate for a completely risk- free security if no inflation were expected.
Regime With reference to a time series, the underlying model generating the times series.
Regression coefficients The intercept and slope coefficient(s) of a regression.
Relative dispersion The amount of dispersion relative to a reference value or benchmark.
Relative frequency With reference to an interval of grouped data, the number of observations in the interval divided by the total number of observations in the sample.
Residual autocorrelations The sample autocorrelations of the residuals.
Risk premium An extra return expected by investors for bearing some specified risk.
Robust standard errors Standard errors of the estimated parameters of a regression that correct for the presence of heteroskedasticity in the regression’s error term.
Robust The quality of being relatively unaffected by a violation of assumptions.
Root mean squared error (RMSE) The square root of the average squared forecast error; used to compare the out-of-sample forecasting performance of forecasting models.
Rule of 72 The principle that the approximate number of years necessary for an investment to double is 72 divided by the stated interest rate.
Safety-first rules Rules for portfolio selection that focus on the risk that portfolio value will fall below some minimum acceptable level over some time horizon.
Sample A subset of a population.
Sample excess kurtosis A sample measure of the degree of a distribution’s peakedness in excess of the normal distribution’s peakedness.
Sample kurtosis A sample measure of the degree of a distribution’s peakedness.
Sample mean The sum of the sample observations, divided by the sample size.
Sample selection bias Bias introduced by systematically excluding some members of the population according to a particular attribute—for example, the bias introduced when data availability leads to certain observations being excluded from the analysis.
Sample skewness A sample measure of degree of asymmetry of a distribution.
Sample standard deviation The positive square root of the sample variance.
Sample statistic A quantity computed from or used to describe a sample.
Sample variance A sample measure of the degree of dispersion of a distribution, calculated by dividing the sum of the squared deviations from the sample mean by the sample size minus 1.
Sampling The process of obtaining a sample.
Sampling distribution The distribution of all distinct possible values that a statistic can assume when computed from samples of the same size randomly drawn from the same population.
Sampling error The difference between the observed value of a statistic and the quantity it is intended to estimate.
Sampling plan The set of rules used to select a sample.
Scatter plot A two-dimensional plot of pairs of observations on two data series.
Scenario analysis Analysis that shows the changes in key financial quantities that result from given (economic) events, such as the loss of customers, the loss of a supply source, or a catastrophic event; a risk management technique involving examination of the performance of a portfolio under specified situations. Closely related to stress testing.
Seasonality A characteristic of a time series in which the data experiences regular and predictable periodic changes, e.g., fan sales are highest during the summer months.
Security selection risk See active specific risk.
Semideviation The positive square root of semivariance (sometimes called semistandard deviation).
Semilogarithmic Describes a scale constructed so that equal intervals on the vertical scale represent equal rates of change, and equal intervals on the horizontal scale represent equal amounts of change.
Semivariance The average squared deviation below the mean.
Serially correlated With reference to regression errors, errors that are correlated across observations.
Sharpe ratio The average return in excess of the risk-free rate divided by the standard deviation of return; a measure of the average excess return earned per unit of standard deviation of return.
Shortfall risk The risk that portfolio value will fall below some minimum acceptable level over some time horizon.
Simple interest The interest earned each period on the original investment; interest calculated on the principal only.
Simple random sample A subset of a larger population created in such a way that each element of the population has an equal probability of being selected to the subset.
Simple random sampling The procedure of drawing a sample to satisfy the definition of a simple random sample.
Simulation trial A complete pass through the steps of a simulation.
Skewed Not symmetrical.
Skewness A quantitative measure of skew (lack of symmetry); a synonym of skew.
Spearman rank correlation coefficient A measure of correlation applied to ranked data.
Spurious correlation A correlation that misleadingly points toward associations between variables.
Standard deviation The positive square root of the variance; a measure of dispersion in the same units as the original data.
Standard normal distribution The normal density with mean (μ) equal to 0 and standard deviation (σ) equal to 1.
Standardizing A transformation that involves subtracting the mean and dividing the result by the standard deviation.
Standardized beta With reference to fundamental factor models, the value of the attribute for an asset minus the average value of the attribute across all stocks,
divided by the standard deviation of the attribute across all stocks.
Stated annual interest rate A quoted interest rate that does not account for compounding within the year. Also called quoted interest rate.
Statistic A quantity computed from or used to describe a sample of data.
Statistical factor model A multifactor model in which statistical methods are applied to a set of historical returns to determine portfolios that best explain either historical return covariances or variances.
Statistical inference Making forecasts, estimates, or judgments about a larger group from a smaller group actually observed; using a sample statistic to infer the value of an unknown population parameter.
Statistically significant A result indicating that the null hypothesis can be rejected; with reference to an estimated regression coefficient, frequently understood to mean a result indicating that the corresponding population regression coefficient is different from 0.
Stress testing A specific type of scenario analysis that estimates losses in rare and extremely unfavorable combinations of events or scenarios.
Subjective probability A probability drawing on personal or subjective judgment.
Survivorship bias The bias resulting from a test design that fails to account for companies that have gone bankrupt, merged, or are otherwise no longer reported in a database.
Systematic risk Risk that affects the entire market or economy; it cannot be avoided and is inherent in the overall market. Systematic risk is also known as non- diversifiable or market risk.
Systematic sampling A procedure of selecting every kth member until reaching a sample of the desired size. The sample that results from this procedure should be approximately random.
t-Test A hypothesis test using a statistic (t-statistic) that follows a t-distribution.
Target semideviation The positive square root of target semivariance.
Target semivariance The average squared deviation below a target value.
Time series A set of observations on a variable’s outcomes in different time periods.
Time value of money The principles governing equivalence relationships between cash flows with different dates.
Time-period bias The possibility that when we use a time-series sample, our statistical conclusion may be sensitive to the starting and ending dates of the sample.
Time-series data Observations of a variable over time.
Time-weighted rate of return The compound rate of growth of one unit of currency invested in a portfolio during a stated measurement period; a measure of investment performance that is not sensitive to the timing and amount of withdrawals or additions to the portfolio.
Total probability rule for expected value A rule explaining the expected value of a random variable in terms of expected values of the random variable conditional on mutually exclusive and exhaustive scenarios.
Total probability rule A rule explaining the unconditional probability of an event in terms of probabilities of the event conditional on mutually exclusive and exhaustive scenarios.
Tracking error The standard deviation of the differences between a portfolio’s returns and its benchmark’s returns; a synonym of active risk. Also called tracking risk.
Tracking risk The standard deviation of the differences between a portfolio’s returns and its benchmark’s returns; a synonym of active risk. Also called tracking error.
Tree diagram A diagram with branches emanating from nodes representing either mutually exclusive chance events or mutually exclusive decisions.
Trend A long-term pattern of movement in a particular direction.
Trimmed mean A mean computed after excluding a stated small percentage of the lowest and highest observations.
Two-sided hypothesis test A test in which the null hypothesis is rejected in favor of the alternative hypothesis if the evidence indicates that the population parameter is either smaller or larger than a hypothesized value.
Type I error The error of rejecting a true null hypothesis.
Type II error The error of not rejecting a false null hypothesis.
Unconditional heteroskedasticity Heteroskedasticity of the error term that is not correlated with the values of the independent variable(s) in the regression.
Unconditional probability The probability of an event not conditioned on another event.
Unit normal distribution The normal density with mean (μ) equal to 0 and standard
deviation (σ) equal to 1.
Unit root A time series that is not covariance stationary is said to have a unit root.
Univariate distribution A distribution that specifies the probabilities for a single random variable.
Up transition probability The probability that an asset’s value moves up.
Value at risk (VaR) A money measure of the minimum value of losses expected during a specified time period at a given level of probability.
Variance The expected value (the probability-weighted average) of squared deviations from a random variable’s expected value.
Volatility As used in option pricing, the standard deviation of the continuously compounded returns on the underlying asset.
Weighted average cost of capital A weighted average of the aftertax required rates of return on a company’s common stock, preferred stock, and long-term debt, where the weights are the fraction of each source of financing in the company’s target capital structure.
Weighted mean An average in which each observation is weighted by an index of its relative importance.
White-corrected standard errors A synonym for robust standard errors.
Winsorized mean A mean computed after assigning a stated percent of the lowest values equal to one specified low value, and a stated percent of the highest values equal to one specified high value.
Working capital management The management of a company’s short-term assets (such as inventory) and short-term liabilities (such as money owed to suppliers).
About the Editors and Authors Richard A. DeFusco, CFA, is a Professor of Finance at the University of Nebraska– Lincoln (UNL). He earned his CFA charter in 1999. A member of CFA Society Nebraska, he also serves on committees for CFA Institute. Dr. DeFusco’s primary teaching interest is investments, and he coordinates the Cornhusker Fund, UNL’s student-managed investment fund. He has published a number of journal articles, primarily in the field of finance. Dr. DeFusco earned his bachelor ’s degree in management science from the University of Rhode Island and his doctoral degree in finance from the University of Tennessee–Knoxville.
Dennis W. McLeavey, CFA, is Faculty Advisor to the Ram Fund and Professor Emeritus of Finance and OR/MS at the University of Rhode Island. Previously, he served as Vice President of Curriculum Development at CFA Institute and co- authored several texts for use in the CFA Program. He earned his CFA charter in 1990 and his Stanford University Certificate in Strategic Decisions and Risk Management in 2012. His research articles have appeared in Management Science, Journal of Operations Research, Journal of Portfolio Management, and other journals. He served as chairperson of the CFA Institute Retirement Investment Policy Committee and as a New York Stock Exchange Arbitrator. After studying economics to earn his bachelor ’s degree at Western University in 1968, he earned a doctorate in production management and industrial engineering, minor in mathematics, from Indiana University, 1972, with the late John F. Muth as his dissertation chair.
Jerald E. Pinto, CFA, has been at CFA Institute since 2002 as Visiting Scholar, Vice President, and now Director, Curriculum Projects in the Education Division for the CFA and CIPM Programs. Prior to joining CFA Institute, he worked for nearly two decades in the investment and banking industries in New York City, including as a consultant in investment planning. He was an assistant professor at NYU’s Stern School from 1992 to 1994, where he taught MBA and undergraduate investments and financial institution management and received a Schools of Business award for teaching. He was a co-author of Quantitative Investment Analysis and Equity Asset Valuation, and a co-editor and co-author of chapters of the third edition of Managing Investment Portfolios: A Dynamic Process, and co-editor of Investments: Principles of Portfolio and Equity Analysis (2011) and Economics for Investment Decision Makers (2013) (all published by John Wiley & Sons). He holds an MBA from Baruch College, a PhD in finance from the Stern School, and is a member of CFA Virginia.
Eugene L. Podkaminer, CFA, is a Senior Vice President in the Capital Markets Research group at Callan Associates. He is responsible for assisting clients with their strategic investment planning, conducting asset allocation studies, developing optimal investment manager structures, and providing custom research on a variety of investment topics. Prior to joining Callan in 2010, Mr. Podkaminer spent nearly a decade with Barclays Global Investors. As a Senior Strategist in the Client Advisory Group, he advised some of the world’s largest and most sophisticated pension plans, non-profits, and sovereign wealth funds in the areas of strategic asset allocation, liability driven investing, manager structure optimization, and risk budgeting. As Chief Strategist of Barclays’ CIO-outsourcing platform Mr. Podkaminer executed CIO-level functions for corporate pension plans and endowments.
Mr. Podkaminer received a BA in Economics from the University of San Francisco and an MBA from Yale University. He is a member of CFA Institute and CFA Society of San Francisco. He is a frequent speaker on investment related topics, and his work on risk factor allocation and alternative indexes has been featured in numerous publications.
David E. Runkle, CFA, is Director of Quantitative Research for Trilogy Global Advisors, LP. He joined Trilogy in 2007 from Piper Jaffray, where he was an Investment Research Manager. Mr. Runkle has been an award-winning professor in finance, accounting, and economics at Brown University and the University of Minnesota, and has served as a Research Officer at the Federal Reserve Bank of Minneapolis. In addition to holding a PhD in economics from the Massachusetts Institute of Technology, Mr. Runkle has served on the CFA Institute Curriculum Committee and Corporate Disclosure Policy Committee and on the Financial Accounting Standards Advisory Council.
Sanjiv Sabherwal, PhD, is an Associate Professor of Finance and Distinguished Teaching Professor at the University of Texas at Arlington, where he teaches graduate and undergraduate courses in international corporate finance, investments, and financial modeling. He has received several teaching awards, and was inducted into the Academy of Distinguished Teachers of the university in 2012. He has been working with CFA Institute as a consultant since 2000. He has published several research articles in journals such as the Journal of Finance, Journal of Financial and Quantitative Analysis,Journal of Banking and Finance, Financial Review, Journal of Investment Management, and Decision Sciences. He holds a B. Tech. in Chemical Engineering from the Indian Institute of Technology, New Delhi, an MBA from the University of Miami, and a PhD in Finance from Georgia Institute of Technology, Atlanta. Prior to his PhD studies, he worked at an international management consulting firm and contributed to projects sponsored by international funding institutions such as the World Bank and the Danish International
Development Agency.
About the CFA Program If the subject matter of this book interests you, and you are not already a CFA Charterholder, we hope you will consider registering for the CFA Program and starting progress toward earning the Chartered Financial Analyst designation. The CFA designation is a globally recognized standard of excellence for measuring the competence and integrity of investment professionals. To earn the CFA charter, candidates must successfully complete the CFA Program, a global graduate-level self-study program that combines a broad curriculum with professional conduct requirements as preparation for a career as an investment professional.
Anchored by a practice-based curriculum, the CFA Program Body of Knowledge reflects the knowledge, skills, and abilities identified by professionals as essential to the investment decision-making process. This body of knowledge maintains its relevance through a regular, extensive survey of practicing CFA charterholders across the globe. The curriculum covers 10 general topic areas, ranging from equity and fixed-income analysis to portfolio management to corporate finance—all with a heavy emphasis on the application of ethics in professional practice. Known for its rigor and breadth, the CFA Program curriculum highlights principles common to every market so that professionals who earn the CFA designation have a thoroughly global investment perspective and a profound understanding of the global marketplace.
www.cfainstitute.org
Index Note: An n after a page number indicates a footnote.
A
1. Absolute dispersion
2. Absolute frequency
3. Absolute value differentiation
4. Acceptance region
5. Accrued interest
6. Active return
7. Active risk 1. active factor risk
2. active manager guidelines
3. active risk squared
4. active specific risk
5. decomposition of
6. information ratio
7. tracking error
8. Addition rule for probabilities
9. Adjusted R2
10. α (alpha)
11. Alternative hypothesis
12. American options
13. Analysis of variance (ANOVA) 1. bid–ask spread explained
2. definition
3. F-tests
4. population means and
5. regression analysis (see Regression analysis)
6. regression sum of squares
7. sum of squared errors
14. Annualizing 1. bank discount basis
2. holding period yield
3. IRR periodic rate
4. semiannual yield to maturity
5. time-weighted return
6. tracking error
7. volatility estimation
15. Annual percentage rate (APR)
16. Annual percentage yield (APY)
17. Annuities 1. annuities due
2. lump sum as
3. ordinary annuities
4. perpetuities
5. solving for payment size
18. ANOVA. See Analysis of variance
19. Antilogarithm conversion
20. A priori probability
21. Arbitrage pricing theory (APT) 1. APT definition
2. APT equation
3. arbitrage definition
4. arbitrage opportunity
5. arbitrage portfolio
6. CAPM as special case
7. factor risk premium
8. pairs arbitrage trade
9. parameters
10. pure factor portfolios
22. ARCH (autoregressive conditional heteroskedasticity)
23. Arithmetic average. See Sample mean
24. Arithmetic mean. See also Mean 1. of Bernoulli random variables
2. of binomial random variables
3. Chebyshev’s inequality
4. computation of
5. cross-sectional mean
6. definition
7. extreme-value sensitivity
8. forward-looking context
9. geometric mean versus
10. harmonic mean versus
11. historical returns
12. investment style comparison
13. population mean,
14. properties of
15. sample mean
16. standard deviation correlation
17. trimmed mean
18. weighted mean versus
19. Winsorized mean
25. Arithmetic mean returns
26. Artifacts of dataset
27. Asian call options
28. Asset allocation 1. correlations and
2. multifactor models for
29. Autocorrelation. See also Serial correlation 1. consequences of
2. moving-average fitting time-series
3. number computed
4. residual autocorrelations
5. residual error
6. time-series model
30. Autoregressive (AR) models 1. autoregressive conditional heteroskedasticity
2. autoregressive moving-average models
3. chain rule of forecasting
4. covariance stationary
5. CPI model (see also Consumer Price Index)
6. definition
7. forecasting
8. mean reversion
9. model specification check
10. moving-average versus
11. random walk as
12. regression coefficient instability
13. sample period length
14. seasonality
15. serially correlated errors
16. time-series data challenges
31. Average. See also Arithmetic mean; Mean
B
1. Back simulations
2. Bank discount basis
3. Bank discount yield
4. Basis points
5. Bayes’ formula
6. Bernoulli random variables 1. binomial random variables
2. definition
3. mean of
4. price movements as
5. standard deviation
6. variance of
7. Bernoulli trial
8. Beta 1. capital asset pricing model
2. definition
3. estimating via regression
4. standardized beta
9. β (beta)
10. Biases 1. data-mining bias
2. evaluation of forecast bias
3. hypothesis testing
4. in investment research
5. look-ahead bias
6. omitted variable bias
7. sample selection bias
8. survivorship bias
9. time-period bias
11. Bid–ask spreads 1. ANOVA explaining
2. multiple linear regression
3. negative
4. nonlinearity and
5. omitted variable bias
12. Bimodal distributions
13. Binomial distributions 1. Bernoulli random variable
2. binomial option pricing model and
3. binomial probability function
4. binomial random variables
5. block broker evaluation
6. combination formula
7. independence assumption
8. price movement modeling
9. as probability distributions
10. skew
14. Binomial formula
15. Binomial option pricing model 1. binomial distributions and
2. combination formula
3. price movement modeling
4. as probability distribution
16. Binomial random variables
17. Binomial trees 1. earnings per share analysis
2. option valuing
3. price movement as binomial model
18. Bivariate normal distributions
19. Black–Scholes–Merton option pricing model 1. continuously compounded returns
2. as continuous time finance model
3. lognormal distribution assumption
4. Monte Carlo simulations versus
5. probability distribution use
6. volatility and
20. Block broker evaluation
21. Bond funds 1. equity funds versus
2. portfolio expected return and variance
22. Bonds 1. bond equivalent yield
2. bond indexing
3. correlations
4. coupon-bearing bonds
5. credit ratings
6. defaults via binomial distribution
7. IRR as yield to maturity
8. pricing of
9. recovery rates on defaulted
10. zero coupon default risk
23. Book-value-to-price ratio
24. Box–Pierce Q-statistic
25. Breusch–Pagan test
C
1. Calculators 1. geometric mean
2. internal rate of return
3. money-weighted rate of return
4. precision of
5. rounding errors
2. Call options
3. Canada Treasury bills
4. Capital asset pricing model (CAPM) 1. as arbitrage pricing theory special case
2. beta
3. data mining bias
4. heteroskedasticity and
5. market portfolio return
6. mean–variance foundation
7. multifactor models versus
8. as probability distribution
9. regression of mutual fund performance
10. systematic risk
5. Capital budgeting
6. Capital structure
7. Carhart four-factor model
8. Cash flow from operations
9. Cash flows
1. additivity principle
2. annuity future values
3. annuity types
4. correlations among
5. estimation
6. future value of single
7. internal rate of return
8. macroeconomic factor models
9. money-weighted rate of return
10. net present value
11. present value of
12. regression analysis
13. retirement annuity payment size
14. series of equal cash flows
15. series of unequal cash flows
16. weighted average cost of capital
10. Cdf. See Cumulative distribution function
11. Cells of stratified sampling
12. Center for Research in Security Prices (CRSP)
13. Central limit theorem 1. continuously compounded return normality
2. definition
3. nonnormal underlying
4. sample size and
14. Central tendency 1. arithmetic mean
2. definition
3. dispersion (see also Dispersion)
4. geometric mean
5. harmonic mean
6. median
7. mode
8. weighted mean
15. Certificates of deposit (CDs)
16. Chain rule of forecasting
17. “Change in” (Δ)
18. Chebyshev’s inequality
19. Chi-square distributions
20. Chi-square tests 1. confidence intervals and
2. critical values for
3. F-tests versus
4. hypothesis tests on single variance
5. p-value via spreadsheet function
21. Classic normal linear regression model assumptions
22. Coefficient of determination (R2) 1. adjusted R2
2. linear regression
3. multicollinearity symptom
4. multiple linear regression
23. Coefficient of variation
24. Cointegrated time-series data
25. Combination formula
26. Common size statements
27. Comparisons 1. common size statements for
2. exchange traded funds vs. T-bills
3. factor models
4. forecasting model performance
5. geometric vs. arithmetic means
6. global vs. US stocks
7. growth and inflation rates
8. investment styles
9. Monte Carlo simulations for
10. negative Sharpe ratios
11. paired comparisons test
12. portfolio managers
13. returns across countries
14. sequential comparisons
15. stock return and price
16. time-vs. money-weighted returns
17. variance of returns pre-and post-crisis
28. Compounding 1. annual
2. compound growth rate
3. continuous (see also Continuous compounding)
4. daily
5. definition
6. frequency
7. future value of single cash flow
8. monthly
9. quarterly
10. semiannual
11. “stated annual interest rate,”
12. time-weighted rate of return
29. Compound returns
30. Conditional expected values
31. Conditional heteroskedasticity 1. autoregressive
2. Breusch–Pagan test
3. serial correlation corrections
32. Conditional probability 1. definition
2. historical vs. future performance
3. likelihoods
4. limit order execution
5. as ratio
6. total probability rule
7. unconditional probability and
33. Conditional variances
34. Confidence intervals 1. chi-square test and
2. confidence limits
3. construction of
4. definition
5. degree of confidence
6. hypothesis testing and
7. one-sided
8. population mean
9. population variance unknown
10. practical interpretations
11. prediction intervals as
12. probabilistic interpretations
13. reliability factors
14. sample size selection
15. Sharpe ratio population mean
16. Student’s t-distribution
17. two-sided
18. underlying distribution unknown
19. for variance
20. z-alternative
35. Consistent probabilities
36. Consol bonds
37. Constant-proportions strategy
38. Consumer Price Index (CPI) 1. autoregressive conditional heteroskedasticity
2. in-vs. out-of-sample
3. modeling
4. time-series model instability
5. trend,
39. Continuous compounding 1. compounding frequency
2. continuously compounded returns
3. discrete versus
4. effective annual rate
5. exponential growth
6. future value of lump sum
40. Continuous random variables
41. Continuous time finance models
42. Continuous uniform distributions
43. Corporate valuation
44. Correlation 1. bivariate normal distributions
2. correlation analysis definition
3. correlation coefficient calculation
4. correlation coefficient definition
5. correlation coefficient significance tests
6. correlation coefficient variables
7. correlation matrix
8. definition
9. diversification and
10. estimating
11. factor definition
12. first-order serial correlation
13. Fisher ’s z-transformation
14. independence vs. uncorrelatedness
15. joint normal distributions
16. multivariate normal distributions
17. negative serial correlation
18. nonlinear relation and
19. outliers
20. pairwise and multicollinearity
21. positive serial correlation
22. properties of
23. residual error
24. scatter plots
25. serial correlation, (see also Serial correlation)
26. Spearman rank correlation coefficient
27. spurious correlation
28. uses of
45. Cost averaging via harmonic mean
46. Counting 1. Bernoulli trials
2. combination formula
3. enumeration
4. factorials
5. labeling problems
6. multinomial formula
7. multiplication rule of
8. permutations
47. Coupon-bearing bonds
48. Covariance 1. correlation coefficient calculation
2. covariance matrix
3. definition
4. diversification correlation
5. estimating
6. joint probability function
7. “off-diagonal covariance
8. portfolio variance
9. of random variable with itself
10. sample covariance
11. sign of
12. stationarity, (see also Covariance stationarity)
13. variance affected by
49. Covariance stationarity 1. definition
2. forecasting model steps
3. random walk determination
4. random walks
5. trend models
6. unit root test
50. Credit risk
51. Critical thinking decision making
52. Critical values. See also Rejection points 1. chi-square
2. Durbin–Watson (DW) statistic
3. F-distributions
4. one-sided t-distributions
53. Cross-sectional data 1. cross-sectional mean
2. cross-sectional standard deviation
3. cross-sectional variance
4. definition
5. fundamental factor models
6. linear regression
7. longitudinal data
8. multiple linear regression
9. observation notation
10. ordering vs. time-series
11. panel data
12. parameter instability
13. sampling
14. stock screening and independence
54. Cumulative distribution function (cdf) 1. continuous uniform cumulative distribution
2. definition
3. discrete uniform distribution
4. normal cumulative distribution function
5. normal cumulative probabilities
6. random observation generation
55. Cumulative frequency distributions
56. Cumulative relative frequency
57. Currencies. See Exchange rates
D
1. Data mining 1. bias
2. definition
3. intergenerational
4. model specification versus
5. warning signs of
2. Data snooping. See Data mining
3. Debt instruments 1. correlations of debt and equity returns
2. holding period yield
3. long-term debt markets
4. pure discount instruments
5. short-term debt markets
4. Deciles
5. Decision making via critical thinking
6. Default risk premium
7. Defaults
8. Degree of confidence
9. Degrees of freedom (df) 1. chi-square and F-distributions
2. definition
3. F-tests
4. linear regression
5. multiple linear regression
6. sample variance
7. t-distribution parameter
10. Delistings
11. Density. See also Probability density function
12. Dependent events
13. Dependent variables 1. ANOVA
2. autoregressive models
3. linear regression
4. log-log regression models
5. multiple linear regression
6. population regression coefficient testing
7. predictions about
8. qualitative dependent variables
9. time-series misspecifications
14. Derivatives of semivariance
15. Descriptive statistics. See also Statistical methods
16. Deviations. See Dispersion
17. Differencing random walks
18. Diffuse priors
19. Discount
20. Discounted cash flow. See Time value of money
21. Discount rate 1. “interest rate” equivalence
2. internal rate of return
3. net present value computation
4. present value and
22. Discount yield
23. Discrete random variables
24. Discrete uniform distributions
25. Discrete vs. continuous compounding
26. Dispersion 1. absolute dispersion
2. Chebyshev’s inequality
3. coefficient of variation
4. definition
5. interquartile range
6. mean absolute deviation
7. population standard deviation
8. population variance
9. range of data
10. relative dispersion definition
11. sample standard deviation
12. sample variance
13. semideviation
14. semivariance
15. Sharpe ratio
16. standard deviation definition (see also Standard deviation)
17. sum around mean
18. variance definition (see also Variance)
19. “Distribution-free” vs. “nonparametric,”
27. Distribution function
28. Diversification 1. arbitrage and
2. asset-specific risk elimination
3. correlation and
4. risk reduction via
29. Dividend yield factor
30. Dogs of the Dow Strategy
31. Dollar-weighted rate of return
32. Dow Dividend Strategy
33. Down transition probability
34. Dummy variables
35. Durbin–Watson (DW) statistic 1. autoregression invalidity
2. critical values for
3. serial correlation in regression
4. trend model correlated errors
36. Dutch Book Theorem
37. DW statistic. See Durbin–Watson (DW) statistic
E
1. Earnings per share (EPS) 1. Bayes’ formula
2. conditional variances
3. dispersion around forecast
4. as independent events
5. operating cost expected value
6. price–earnings ratio analysis
7. as random variables
8. total probability rule
9. tree diagram
10. unconditional variance
2. Earnings yield
3. EBITDA analysis
4. Economic forecast evaluation
5. Effective annual rate (EAR)
6. Effective annual yield (EAY)
7. Empirical probability
8. Enterprise value (EV)
9. Enumeration shortcuts
10. Equal cash flows
11. Equity funds 1. active risk comparison
2. arithmetic mean returns
3. bond funds versus
4. correlations
5. exchange traded funds vs. T-bills
6. Forbes Magazine Honor Roll
7. geometric mean returns
8. performance predictability
9. population variance hypothesis testing
10. sample standard deviation
11. sample variance
12. skewness
13. t-test of population mean
12. Equity markets 1. coefficient of variation
2. frequency distribution
3. global vs. US
4. median
5. risk premium hypothesis testing
6. sample mean return
7. stock return and price relationship
13. Equivalent annual rate (EAR)
14. Errors 1. error autocorrelation
2. heteroskedasticity
3. homoskedasticity
4. in-vs. out-of-sample forecast errors
5. prediction errors
6. root mean squared error
7. rounding errors
15. Error term of regression
1. assumptions
2. definition
3. error autocorrelation
4. uncertainty of
16. Estimation 1. asymptotic properties
2. confidence intervals
3. consistency of estimator
4. definition
5. efficiency of estimator
6. estimate definition
7. estimators definition
8. hats over symbols as estimated
9. linear regression parameters
10. normal distributions for
11. ordinary least squares (see also Ordinary least squares)
12. point estimates
13. precision of estimator
14. Student’s t-distribution
15. unbiasedness of estimator
17. European-style options
18. EURO STOXX
19. Events 1. Bayes’ formula
2. complements
3. definition
4. dependent events
5. exhaustive events
6. independent events (see also Independence
7. probabilities of outcomes
8. total probability rule
20. Excess kurtosis 1. definition
2. as “kurtosis,”
3. normal distribution
4. sample excess kurtosis
5. S&P 500 Index
21. Excess return 1. Carhart four-factor model
2. mean excess return
3. performance evaluation regression
4. return attribution
5. Sharpe ratio
22. Exchange rates 1. exchange rate return correlations
2. holding period return formula
3. as random walks
4. time-series data for
23. Exchange traded fund comparison
24. Exhaustive events
25. Expected returns 1. arbitrage pricing theory
2. calculation of
3. Carhart four-factor model
4. portfolio (see Portfolio expected return)
5. risk premium definition
26. Expected values 1. conditional expected values
2. default risk premium
3. dispersion around
4. forward-looking data
5. multiplication rule
6. operating costs
7. properties of
8. of random variable
9. total probability rule
10. variance definition
27. Expense ratio
28. Explanatory variables
29. Explosive roots
30. Exponential growth 1. log-linear trends
2. Starbucks’ seasonality
3. time series with unit root
F
1. Face value of US T-bills
2. Factorials
3. Factors 1. arbitrage pricing theory
2. company fundamental factors
3. company share-related factors
4. definition
5. factor analysis models
6. factor price
7. factor risk premium
8. factor sensitivity
9. fundamental
10. macroeconomic
11. principal components models
12. pure factor portfolios
13. statistical
4. Fat tails
5. F-distributions
6. Financial calculators. See Calculators
7. Financial crisis 1. inflation forecasts and
2. mutual fund skewness
3. returns and inflation rate outliers
4. volatility pre-and post-crisis
8. Finite population correction factor (fpc)
9. First-order serial correlation
10. Fisher effect 1. corrected standard errors
2. Durbin–Watson statistic
3. lagged dependent variable
4. measurement error
5. testing
6. unit roots and
11. Fisher ’s z-transformation
12. Fitted parameters of linear regression
13. Foolish Four investment strategy
14. Forbes Magazine Honor Roll
15. Forecasting 1. autoregressive models (see also Autoregressive (AR) models)
2. autoregressive moving-average models
3. bias evaluation
4. chain rule of forecasting
5. CPI (see Consumer Price Index)
6. economic forecast evaluation
7. EPS dispersion around forecast
8. forecasting the past
9. future as out-of-sample
10. inflation rate forecast evaluation
11. in-sample forecast errors
12. model performance comparison
13. moving-average models
14. multiperiod forecasts
15. out-of-sample forecast errors
16. root mean squared error
17. seasonality
18. Survey of Professional Forecasters
19. time-series steps
20. uncertainty
16. Four-factor model
17. France Treasury bills
18. Free cash flow to the firm
19. Frequency distributions 1. absolute frequency
2. construction of
3. cumulative relative frequency
4. definition
5. histograms
6. holding period returns
7. intervals
8. Monte Carlo simulations
9. relative frequency
20. Frequency polygons
21. F-statistic
22. F-tests 1. analysis of variance
2. chi-square tests versus
3. critical values for
4. degrees of freedom
5. heteroskedasticity and
6. p-value for
7. p-value via spreadsheet function
23. Full-replication approach
24. Fundamental factor models
25. Future value 1. cash flow additivity principle
2. compounding formula
3. compounding frequency and
4. present value equivalence
5. series of cash flows future value
6. single cash flow future value
7. single cash flow present value
8. solving for annuity payment size
9. solving for growth rate
10. solving for interest rate
11. solving for number of periods
G
1. GARCH
2. Generalized least squares
3. Geometric mean 1. arithmetic mean versus
2. compound growth rate
3. computation of
4. definition
5. harmonic mean versus
6. historical returns
7. standard deviation correlation
8. time-weighted rate of return
4. Geometric mean returns 1. arithmetic mean versus
2. formula
3. investment style comparison
4. mutual fund returns
5. standard deviation correlation
5. Germany Treasury bills
6. Graphical representations of data 1. continuous uniform distributions
2. cumulative frequency distributions
3. frequency polygons
4. hetero-vs. homoskedasticity
5. histograms
6. linear regression
7. lognormal distributions
8. normal distribution
9. outliers
10. scatter plots (see also Scatter plots)
11. semilogarithmic scales
12. standard deviation
13. Student’s t-distribution
14. time-series
15. time series plotting
7. Growth rates 1. compound growth rate
2. exponential growth
3. geometric mean for
4. growth and inflation rate correlation
5. growth factor risk premium
6. interest rates as
7. linear regression with inflation rate
8. solving for
9. time-weighted rate of return
H
1. Hansen standard errors
2. Harmonic mean
3. Hats over symbols
4. Hedge funds
5. Hedging via time-series data
6. Heteroskedasticity 1. autoregressive conditional
2. conditional
3. consequences of
4. correcting for
5. definition
6. standard error of estimate and
7. testing for
8. unconditional
7. Histograms
8. Historical returns 1. beta estimation
2. correlation estimation
3. covariance estimation
4. empirical probability
5. geometric vs. arithmetic means
6. historical mean vs. expected value
7. historical simulations
8. mutual fund performance
9. statistical factor models
10. survivorship bias
11. volatility estimation
9. Historical simulations
10. Holding period return (HPR) 1. computation of
2. continuously compounded returns
3. money market yield
4. money-weighted rate of return
5. price relative
6. rates of return formula
7. time-weighted rate of return
11. Holding period yield (HPY)
12. Homoskedasticity
13. Hurdle rate
14. Hypothesis testing 1. acceptance region
2. alternative hypothesis
3. beta estimation via
4. confidence intervals and
5. critical thinking decision making
6. data mining
7. definition
8. equity risk premiums example
9. hypothesis definition
10. hypothesis statements
11. inflation forecast evaluation
12. joint hypothesis testing
13. linear regression
14. Mann–Whitney U test
15. mean oriented
16. nonparametric tests
17. null hypothesis formulation
18. one-sided hypothesis tests
19. parametric tests definition
20. parametric vs. nonparametric tests
21. population variance known
22. population variance unknown
23. power of a test
24. p-value approach to
25. rejecting the null (see also Rejecting the null hypothesis)
26. rejection points
27. scientific method
28. significance level
29. significance tests
30. sign test
31. as statistical inference
32. statistical significance
33. steps of
34. Student’s t-distribution
35. test statistic calculation
36. t-tests
37. two-sided hypothesis tests
38. Type I errors
39. Type II errors
40. variance oriented
41. Wilcoxon signed-rank test
42. z-tests
I
1. Inconsistent probabilities
2. Incremental cash flows
3. Independence 1. binomial distribution assumption
2. central limit theorem
3. definition
4. earnings per share as
5. independence vs. uncorrelatedness
6. independently and identically distributed
7. multiplication rule for
8. mutual fund performance
9. portfolio variance
10. random variable independence
11. screening stocks and
4. Independently and identically distributed (IID)
5. Independent variables 1. ANOVA
2. autoregressive models
3. linear combinations of
4. linear regression
5. log-log regression models
6. log transformations
7. macroeconomic factor models
8. multicollinearity
9. multiple linear regression
10. pairwise correlations
11. perfect collinearity
12. population regression coefficient testing
13. randomness of
14. time-series misspecifications
6. Indexes 1. correlations among stock return series
2. S&P 500 Index (see S&P 500 Index)
3. survivorship bias
7. Indexing bonds
8. Inferential statistics. See Statistical inference
9. Inflation premium and interest rates
10. Inflation rates 1. CPI (see Consumer Price Index)
2. Fisher effect
3. forecast evaluation
4. growth and inflation rate correlation
5. inflation factor risk premium
6. interest rate relation heteroskedasticity
7. linear regression of forecast bias
8. linear regression with growth
9. linear regression with stock returns
10. surveys of
11. time-series model instability
11. Information ratio
12. In-sample forecast errors
13. Integration
14. Interest coverage ratio
15. Interest rates 1. annual percentage rate
2. definition
3. effective annual rate
4. Fisher effect
5. future value lump sum formulas
6. as growth rates
7. inflation relation heteroskedasticity
8. macroeconomic factor models
9. periodic interest rate
10. Rule of
11. simple interest
12. solving for
13. stated annual interest rate
14. tables of factors vs. calculators
15. time value of money definitions
16. Intergenerational data mining
17. Internal rate of return (IRR) 1. as bond yield to maturity
2. caveats
3. computation of,
4. definition
5. IRR rule
6. money-weighted rate of return
18. Interpolation for quantiles
19. Interquartile range
20. Intervals 1. Chebyshev’s inequality
2. confidence intervals definition
3. confidence intervals of population mean
4. frequency distributions
5. lognormal distributions
6. modal interval
7. normal distribution standard deviations
8. prediction intervals
21. Interval scales
22. Inverse probability
23. Invested capital (IC)
24. IRR. See Internal rate of return
J
1. Japan Treasury bills
2. Jarque–Bera (JB) statistical test
3. Joint normal distributions
4. Joint probability function
K
1. Kurtosis 1. definition
2. excess kurtosis
3. fat tails
4. leptokurtic distributions
5. mesokurtic distributions
6. normal distributions
7. platykurtic distributions
8. sample excess kurtosis
9. sample kurtosis
L
1. Labeling problems
2. Large-capitalization stocks 1. Carhart four-factor model
2. factor model comparison
3. growth vs. value
4. mean returns
5. returns of
6. Russell 1000 Index
7. S&P 500 Index
3. Leptokurtic distributions
4. Level of significance. See Significance level
5. Likelihoods
6. Limit order execution
7. Linear association
8. Linear combinations
9. Linear interpolation for quantiles
10. Linear least squares. See Linear regression
11. Linear regression 1. assumptions
2. beta estimation
3. coefficient of determination
4. cross-sectional data
5. definition
6. dependent variable
7. Durbin–Watson statistic
8. economic forecast evaluation
9. error term
10. estimated parameters
11. hypothesis testing
12. independent variable
13. inflation and growth rates
14. limits of models
15. linear trend models
16. measurement error
17. multiple independent variables (see also Multiple linear regression)
18. one independent variable
19. parameter instability
20. prediction intervals
21. random walk differencing
22. random walk with drift
23. regression coefficients
24. regression residual
25. standard error of estimate
26. stationarity tests
27. time series, more than one
28. time-series covariance stationary
29. time-series data
30. uncertainty
12. Linear trend models
13. Liquidity investment style
14. Liquidity premium and interest rates
15. Logistic distributions
16. Logit models
17. Log-linear trend models
18. Log-log regression models
19. Lognormal distributions
20. Lognormal variables
21. Log transformations for wide ranges
22. Longitudinal data
23. Long-term debt markets
24. Look-ahead bias
25. Lump sums
M
1. Macroeconomic factor models
2. Mann–Whitney U test
3. Marginal probability
4. Market model regression
5. Market portfolio return
6. Market timing,
7. Market-to-book ratio
8. Maturity date pull to par value
9. Maturity premium and interest rates
10. Mean. See also Arithmetic mean 1. “average” versus
2. of Bernoulli random variables
3. of binomial random variables
4. Chebyshev’s inequality
5. geometric mean
6. geometric vs. arithmetic
7. harmonic mean
8. harmonic vs. geometric vs. arithmetic
9. hypothesis testing of differences between
10. hypothesis testing of single means
11. hypothesis tests concerning
12. linear regression line
13. lognormal distributions
14. mean absolute deviation
15. mean excess return
16. mean of deviations
17. mean reversion
18. normal distributions
19. of populations equal
20. skewness
21. stationarity (see also Stationarity)
22. uniform random variable with limits
23. variance
24. weighted mean
11. Mean absolute deviation (MAD)
12. Mean returns 1. across equity markets
2. arithmetic mean returns
3. constant-proportions strategy
4. deviation sum of zero
5. dispersion (see also Dispersion)
6. geometric mean returns
7. multivariate normal distributions
8. mutual funds
9. risk evaluation
10. skewness
11. as typical outcome measure
12. weighted mean
13. Mean reversion
14. Mean squared error (MSE)
15. Mean–variance analysis
16. Measurement errors
17. Measurement scales
18. Median 1. as 50th percentile
2. definition
3. extreme values and
4. normal distributions
5. skewness
19. Mergers and acquisitions
20. Mesokurtic distributions
21. Mode 1. calculating
2. definition
3. modal interval
4. normal distributions
5. skewness
22. Modeling. See also Forecasting 1. arbitrage pricing theory
2. CAPM (see also Capital asset pricing model)
3. CPI model (see also Consumer Price Index)
4. data mining definition
5. forecast model comparison
6. heteroskedasticity and
7. model specification
8. Monte Carlo simulations for
9. multifactor models (see also Multifactor models)
10. out-of-sample tests
11. probit models
12. root mean squared error
13. time-series forecasting steps
14. time series with unit root
15. trend models
23. Modern portfolio theory (MPT) 1. definition
2. diversification benefit
3. expected return as reward measure
4. mean–variance analysis
5. multifactor models and
6. normal distribution
7. variance of return as risk measure
24. Momentum 1. Carhart four-factor model
2. investment style
3. “momentum” stocks
25. Money market yields
26. Money-weighted rate of return
27. Monte Carlo simulations 1. analytical methods versus
2. applications of
3. central limit theorem test
4. definition
5. historical simulation versus
6. market timing vs. buy and hold
7. normal distribution
8. overview of
9. probability distributions and
10. simulation trials
11. VAR estimation
28. Moody’s Investors Service
29. Mortgage-backed security simulations
30. Mortgages
31. Motley Fool “Foolish Four,”
32. Moving-average time-series models 1. autoregressive moving-average models
2. autoregressive versus
3. lagging large movements
4. plotting
5. simple moving average
6. time-series models
7. 12-month moving averages
33. MSCI
34. Multicollinearity
35. Multifactor models 1. arbitrage pricing theory
2. for asset allocation
3. Carhart four-factor model
4. factor analysis models
5. factor definition
6. factor model comparison
7. factor sensitivity
8. fundamental factor models
9. macroeconomic factor models
10. modern portfolio theory and
11. for portfolio construction
12. principal components models
13. pure factor portfolios
14. for return attribution
15. for risk attribution
16. standardized beta
17. statistical factor models
18. for strategic decisions
19. strengths of
20. total risk attribution
36. Multinational corporation valuation
37. Multinomial formula
38. Multiple coefficient of determination (multiple R2) 1. adjusted R2
2. bid–ask spread explained
3. definition
4. multicollinearity
39. Multiple linear regression 1. adjusted R2
2. assumptions
3. assumption violations
4. bid–ask spread
5. cross-sectional data
6. data nonlinearity
7. data scaling
8. data transformations
9. definition
10. degrees of freedom
11. dependent variable
12. dependent variable prediction
13. dummy variables for
14. Durbin–Watson statistic
15. forecasting the past
16. formula
17. functional form of
18. heteroskedasticity
19. independent variables
20. linear trend models
21. log-log regression models
22. mergers and acquisitions and returns
23. model specification
24. month-of-year effects on returns
25. multicollinearity
26. mutual fund regression analysis
27. partial regression coefficients
28. perfect collinearity
29. population regression coefficient testing
30. random walk differencing
31. random walk with drift
32. regression coefficients
33. samples pooled poorly
34. serial correlation
35. standard error of estimate
36. stationarity tests
37. time series, more than one
38. time-series covariance stationary
39. time-series data
40. time-series misspecification
41. uncertainty
42. valuation of corporations
40. Multiplication rules 1. counting
2. expected value of uncorrelated variables
3. independent events
4. probability
41. Multivariate normal distributions
42. Mutual funds 1. arithmetic mean returns
2. Bayes’ formula for evaluation
3. geometric mean returns
4. hypothesis test on population variance
5. market timing model specification
6. performance predictability
7. regression of performance
8. sample standard deviation
9. sample variance
10. Sharpe ratios
11. skewness
12. Spearman rank correlation coefficient
13. t-test of population mean
43. Mutually exclusive projects
N
1. Natural logarithm 1. antilogarithm conversion
2. geometric mean
3. independent variables with wide ranges
4. log-linear trend models
5. log-log regression models
6. lognormal distributions
7. nonlinearity of data
8. rules for
2. Negative serial correlation
3. Negative skew
4. Net income
5. Net present value (NPV)
6. Newey–West serial correlation correction
7. Nodes of tree diagrams
8. Nominal risk-free interest rate
9. Nominal scales
10. Nonlinear relation
11. Nonparametric hypothesis testing
12. Nonstationarity 1. definition
2. price-to-earnings time series
3. spurious regression
4. unit root test
13. Normal density functions
14. Normal distributions 1. applications of
2. bivariate normal distributions
3. central limit theorem
4. Chebyshev’s inequality
5. error term of linear regression
6. excess kurtosis
7. Jarque–Bera statistical test
8. joint normal distribution
9. kurtosis
10. lognormal distributions and
11. multivariate
12. normal density functions
13. option return modeling
14. parameters defining
15. as probability distributions
16. probability estimation via
17. probit models
18. properties of
19. as returns model
20. sample mean and sample size
21. skewness
22. standard deviation
23. standard normal distributions
24. univariate
15. Normal linear regression model assumptions
16. Normal random variables
17. Notation 1. cdf of standard normal variable
2. “change in,”
3. complements of events
4. correlation
5. double summation signs
6. factorials
7. Greek for population parameters
8. hats over symbols as estimated
9. intervals
10. observations
11. outcomes
12. random variables
13. rejection points
14. reliability factors
15. Roman italics for sample statistics
16. unconditional probability
18. NPV. See Net present value
19. Null hypothesis 1. acceptance region
2. definition
3. formulation
4. power of a test
5. p-value
6. rejection of (see also Rejecting the null hypothesis)
7. sample size and
8. significance level
9. Type I and II errors
O
1. Objective probabilities
2. Observations 1. autocorrelations computed and
2. notation of
3. paired observations
4. random observations
5. to ranks for nonparametric tests
6. time-series logical ordering
7. t-tests
3. Odds of probability
4. Off-diagonal covariance
5. Omitted variable bias
6. One-tailed hypothesis tests
7. Opportunity costs 1. definition
2. interest rates as
3. IRR rule
4. net present value discount rate
8. Optimization
9. Option pricing models 1. Binomial (see Binomial option pricing model)
2. Black–Scholes–Merton (see Black–Scholes–Merton option pricing model)
3. continuously compounded returns
4. Monte Carlo simulations for
5. volatility
10. Options 1. skew of returns
2. underlying assets
3. volatility
11. Ordinal scales
12. Ordinary annuities
13. Ordinary least squares (OLS)
14. definition 1. linear trend models
2. multicollinearity and
3. multiple linear regression model
4. negative serial correlation and
5. positive serial correlation and
15. Outcomes 1. counting
2. cumulative distribution function
3. definition
4. discrete uniform distribution
5. event probability
6. notation for
7. possible for random variables
8. probability distribution definition
9. probability function
10. random variable definition
11. sum to one
16. Outliers
17. Out-of-sample forecast errors
18. Out-of-sample tests
19. Own covariance
P
1. Paired comparisons test
2. Paired observations
3. Pairs arbitrage trade
4. Pairwise correlations and multicollinearity
5. Panel data
6. Parameters 1. autoregressive moving-average models
2. confidence intervals
3. definition
4. estimated of linear regression. See also Estimation
5. examples of
6. Greek letters for
7. lognormal distributions
8. multivariate normal distributions
9. normal distributions
10. one-factor APT model
11. parameter instability
12. parametric tests definition
13. parametric vs. nonparametric tests
14. point estimators
15. t-distributions
7. Partial regression coefficients
8. Partial slope coefficients
9. Pearson coefficient of skewness
10. Percentiles
11. Perfect collinearity
12. Performance appraisal 1. analyst coverage probit model
2. binomial distribution for
3. block broker evaluation
4. definition
5. fundamental factor models
6. investment manager
7. investment styles
8. money-weighted rate of return
9. risk-adjusted performance
10. sample selection bias
11. style analysis correlations
12. time-weighted rate of return
13. Performance attribution
14. Performance measurement 1. Bayes’ formula for
2. common size statements
3. correlations among measures
4. free cash flow explained
5. holding period return
6. linear regression hypothesis testing
7. money market yields
8. money-weighted rate of return
9. MSCI EAFE Index
10. mutual fund regression analysis
11. Sharpe ratio for
12. time-series data for
13. time-weighted rate of return
14. tracking error
15. Periodic rate of return
16. Permutations
17. Perpetuities
18. Persistence of returns
19. Platykurtic distributions
20. Point estimators
21. Populations 1. definition
2. hypothesis definition
3. nonparametric tests
4. parameter confidence intervals
5. parameter point estimators
6. parameters
7. parameters in Greek letters
8. parameters via samples
9. parametric tests definition
10. population mean,
11. population mean and ANOVA
12. population mean confidence intervals
13. population mean hypothesis tests
14. population median
15. population mode
16. population regression coefficient
17. population standard deviation
18. population variance
19. population variance known
20. population variance unknown
21. samples versus
22. sampling more than one
23. underlying samples
22. Portfolio expected return 1. arbitrage pricing theory
2. calculation of
3. Carhart four-factor model
4. as function of individual securities
5. as measure of reward
6. variance and
23. Portfolio returns 1. hypothesis testing comparisons
2. joint normal distribution
3. performance (see Performance measurement)
4. safety-first rules
24. Portfolio standard deviation of return
25. Portfolio variance of return 1. covariance
2. as function of individual securities
3. as measure of risk
4. well-diversified portfolios
26. Positive serial correlation
27. Positive skew
28. Posterior probability
29. Power of a test
30. Prediction errors
31. Prediction intervals
32. Present value 1. compounding frequency and
2. discount rate and
3. future value equivalence
4. money-weighted rate of return
5. series of cash flows
6. single cash flow future value
7. single cash flow present value
8. solving for annuity payment size
9. solving for growth rate
10. solving for interest rate
11. solving for number of periods
33. Price 1. arbitrage pricing theory
2. Asian call options
3. binomial model for
4. bond pricing
5. capital asset pricing model (see also Capital asset pricing model)
6. Center for Research in Security Prices
7. fundamental factor models
8. heteroskedasticity and models
9. lognormal distributions for
10. as random variable
11. stock return and price relationship
34. Priced risk
35. Price relatives
36. Price-to-book ratio (P/B)
37. Price-to-earnings ratios (P/Es) 1. arithmetic mean analysis
2. definition
3. earnings per share correlation
4. fundamental factor models
5. linear regression explaining
6. median analysis
7. nonstationary
8. stock market returns and
38. Principal
39. Principal components models
40. Prior probabilities
41. Probability 1. addition rule
2. a priori probability
3. Bayes’ formula
4. conditional probability
5. correlation (see also Correlation)
6. counting
7. covariance (see also Covariance)
8. default risk premium
9. definition
10. Dutch Book Theorem
11. empirical probability
12. estimation via normal distributions (see also Estimation)
13. events (see also Events)
14. expected value (see also Expected values)
15. inconsistent probabilities
16. inverse probability
17. joint probability function
18. likelihoods
19. limit order execution
20. marginal probability
21. Monte Carlo simulations (see also Monte Carlo simulations)
22. multiplication rules
23. objective probabilities
24. odds of
25. outcomes (see also Outcomes)
26. posterior probability
27. prior probabilities
28. probability distributions (see also Probability distributions)
29. random variables
30. subjective probability
31. sum to one
32. total probability rule
33. tree diagrams
34. unconditional probability
35. variance
42. Probability density function (pdf) 1. definition
2. lognormal distributions
3. normal distributions
4. uniform random variable
43. Probability distributions 1. binomial
2. bond price modeling
3. cdf (see also Cumulative distribution function)
4. continuous uniform
5. continuous vs. discrete
6. definition
7. discrete uniform
8. discrete vs. continuous
9. historical simulations
10. lognormal
11. Monte Carlo simulations
12. normal (see also Normal distributions)
13. probability function (see also Probability function)
14. random numbers vs. observations
15. t-distribution (see also t-distributions)
16. test statistics
17. uniform
44. Probability function 1. Bernoulli random variables
2. binomial random variables
3. continuous uniform random variables
4. definition
5. discrete uniform distributions
45. Probability mass function (pmf)
46. Probit models
47. Pseudo-random numbers
48. Pure discount instruments
49. p-value 1. approach to hypothesis testing
2. definition
3. for F-test
4. for null rejection
5. for regression coefficient
6. spreadsheet calculation
Q
1. Q-statistic
2. Quadruple witching days
3. Qualitative dependent variables
4. Quantiles
5. Quartiles
6. Quintiles
7. Quoted interest rate
R
1. R2. See Coefficient of determination; Multiple coefficient of determination
2. Random number generation
3. Random numbers
4. Random observation generation
5. Random observations
6. Random sampling
7. Random variables 1. Bernoulli random variables
2. binomial random variables
3. central limit theorem
4. continuous (see also Continuous random variables)
5. covariance and
6. covariance with itself
7. definition
8. discrete
9. discrete uniform distributions
10. earnings per share as
11. expected value
12. independence definition
13. joint probability function
14. lognormal variables
15. multiplication rule for expected value
16. normal random variables
17. notation for
18. possible outcomes
19. price as
20. probability distribution of (see also Probability distributions)
21. return as
22. sample mean as
23. sample statistics as
24. significance tests
25. standardizing
26. uniform
27. variance of
8. Random walks 1. as autoregression
2. covariance stationarity
3. covariance stationarity determination
4. definition
5. differencing
6. exchange rates as
7. mean-reversion level undefined
8. random walk with drift
9. random walk with drift and trend
10. unit root test
9. Range of data 1. bond pricing
2. definition
3. frequency distribution construction
4. interquartile range
5. log transformations for wide ranges
6. risk evaluation
10. Ranked data
11. Rates of return 1. as continuous random variables
2. frequency distributions for
3. geometric mean for compound
4. interest rates
5. internal (see Internal rate of return [IRR])
6. money-weighted
7. stock return and price relationship
8. time-weighted
12. Ratio averaging
13. Ratio scales
14. Real risk free interest rate
15. Receivables mean number of days
16. Regimes in time series
17. Regression analysis 1. analysis of variance
2. assumptions
3. assumption violations
4. autoregression (see Autoregressive [AR] models)
5. beta estimation
6. coefficient of determination
7. confidence intervals
8. cross-sectional data
9. data nonlinearity
10. data scaling
11. data transformations
12. dependent variables
13. dummy variables for
14. Durbin–Watson statistic
15. error term
16. forecasting the past
17. functional form of
18. heteroskedasticity
19. hypothesis testing
20. independent variables
21. limits of models
22. linear regression (see also Linear regression; Multiple linear regression)
23. linear trend models
24. model specification
25. multicollinearity
26. operating cost expected value
27. parameter instability
28. partial regression coefficients
29. perfect collinearity
30. random walk differencing
31. random walk with drift
32. regression coefficient instability
33. regression coefficients
34. regression coefficients of population
35. regression residual
36. residual error
37. samples pooled poorly
38. serial correlation
39. small sample unbiasedness
40. standard error of estimate
41. stationarity tests
42. time series, more than one
43. time-series covariance stationary
44. time-series data
45. time-series misspecification
46. uncertainty
18. Regression sum of squares (RSS)
19. Regressors
20. Reinvestment risk
21. Rejecting the null hypothesis 1. DW statistic and serial correlation
2. hypothesis testing procedure
3. one-vs. two-tailed hypothesis tests
4. positive serial correlation
5. p-value indicating
6. sample size and
7. Type I and II errors
22. Rejection points
23. Relative dispersion
24. Relative frequency
25. Relative skewness. See Skewness
26. Residual autocorrelations
27. Residual error
28. Return distributions 1. annual vs. shorter holding periods
2. central tendency
3. dispersion (see also Dispersion)
4. frequency distributions
5. Jarque–Bera statistical test
6. kurtosis
7. normal distributions for
8. option returns skew
9. Sharpe ratio and symmetry
10. skewness
11. variance vs. semivariance
29. Return on invested capital (ROIC)
30. Returns 1. across countries
2. active return
3. arbitrage pricing theory
4. attribution of
5. compound returns
6. continuously compounded returns
7. correlations
8. excess (see Excess return)
9. expected returns calculation (see also Expected returns)
10. histograms of
11. historical (see Historical returns)
12. holding period return (see Holding period return)
13. independently and identically distributed
14. linear regression of inflation and stock return
15. macroeconomic factor models
16. mean returns (see Mean returns)
17. mergers and acquisitions and
18. modeling via normal distribution (see also Modeling)
19. month-of-year effects
20. mutual fund regression analysis
21. persistence of
22. portfolio (see Portfolio returns)
23. price-to-earnings ratio and
24. as random variables
25. rates of (see Rates of return)
26. stationarity
27. stock return and price relationship
28. time-series autocorrelations
31. Risk 1. active risk
2. arbitrage
3. attribution of
4. capital asset pricing model
5. priced risk
6. risk premium
7. systematic risk
8. total risk attribution
32. Risk-adjusted performance
33. Risk assessment 1. coefficient of variation
2. conditional variance for
3. diversification
4. frequency distributions for
5. Monte Carlo simulations for
6. normal distribution for
7. portfolio variance for
8. range and mean absolute deviation
9. safety-first rules
10. semideviation measure
11. semivariance measure
12. shortfall risk
13. standard deviation measure
14. stress testing/scenario analysis
15. uniform distributions
16. value at risk (VAR)
17. variance measure
34. Risk-free rate 1. probability overview
2. real risk-free interest rate
3. risk premium definition
4. Sharpe ratio
5. T-bills for
35. Risk premiums
36. Robust standard errors
37. Robust t-tests
38. ROIC. See Return on invested capital
39. Root mean squared error
40. Rounding errors
41. Rule of 72
42. Russell 1000 Index
43. Russell 2000 Growth Index
44. Russell 2000 Index
45. Russell 2000 Value Index
S
1. Safety-first rules
2. Sample mean 1. central limit theorem
2. cross-sectional mean
3. definition
4. distribution of
5. expected value versus
6. population mean estimation
7. sample size
8. sampling distribution normality
9. standard error of
3. Samples 1. definition
2. in-vs. out-of-sample forecast errors
3. out-of-sample tests
4. populations versus
5. sample period length
6. samples pooled poorly
7. sampling (see also Sampling)
8. sampling error
9. statistics (see Sample statistics)
4. Sample selection bias
5. Sample size 1. autocorrelations computed and
2. autoregressive moving-average models
3. central limit theorem and
4. correlation coefficient significance
5. estimator consistency
6. estimator unbiasedness and efficiency
7. finite population correction factor
8. “large” meaning
9. nonnormal underlying
10. sampling distribution and
11. selection of
12. t-distribution
13. Type I and II errors and
14. z-alternative
15. zero correlation rejection
16. z-test vs. t-test
6. Sample statistics 1. definition
2. as random variables
3. Roman italic letters for
4. sample correlation coefficient
5. sample covariance
6. sample excess kurtosis
7. sample kurtosis
8. sample mean
9. sample median
10. sample mode
11. sample skewness
12. sample standard deviation,
13. sample variance
14. sample variance and degrees of freedom
15. sampling distribution of
16. standard error of
7. Sampling 1. bias in investment research
2. bond indexing
3. cross-sectional data
4. data-mining bias
5. definition
6. finite population correction factor
7. look-ahead bias
8. out-of-sample tests
9. sample selection bias
10. sample size, (see also Sample size)
11. sampling distribution of statistic (see also Central limit theorem)
12. sampling error
13. sampling plan
14. simple random sampling
15. stratified random sampling
16. survivorship bias
17. systematic sampling
18. time-period bias
19. time-series data
20. underlying nonnormal
21. underlying populations differ
8. S&P 500 Index 1. Chebyshev’s inequality
2. cumulative frequency distributions
3. excess kurtosis
4. expected value
5. histograms of returns
6. holding period returns
7. mean return comparison
8. means vs. standard deviation
9. portfolio expected return and variance
10. positive excess kurtosis
11. returns and inflation rate
12. as sample
13. skewness
9. Sarbanes–Oxley Act (2002)
10. Scales of measurement
11. Scatter plots 1. correlated data
2. definition
3. economic forecast evaluation
4. EV/IC fitted regression line
5. hetero-vs. homoskedasticity
6. inflation rates and stock returns
7. outliers
8. regression of linear vs. nonlinear data
9. uncorrelated data
12. Scientific method
13. Seasonality
14. Selected American Shares (SLASX) 1. arithmetic mean
2. geometric mean
3. mean absolute deviation
4. range
5. sample standard deviation
6. sample variance
7. Sharpe ratio
15. Selling short
16. Semideviation
17. Semilogarithmic scales
18. Semistandard deviation. See Semideviation
19. Semivariance
20. Serial correlation 1. autoregressive models
2. consequences of
3. consistent standard errors
4. correcting for
5. Durbin–Watson statistic
6. first-order serial correlation
7. forecasting model steps
8. negative
9. positive
10. prediction errors
11. regression analysis
12. residuals in time-series models
13. testing for
14. trend models
21. Sharpe ratio 1. confidence interval for population mean of
2. definition
3. exchange traded funds vs. T-bills
4. negative
5. performance measurement
6. safety-first ratio similarity
22. Shortfall risk
23. Shorting stock
24. Short-term debt markets
25. Significance level 1. linear regression hypothesis testing
2. p-value
3. specifying
26. Significance tests 1. correlation coefficient
2. heteroskedasticity and
3. t-test of significance
27. Sign test
28. Simple interest
29. Simple linear regression. See Linear regression
30. Simple moving average
31. Simple random sampling
32. Simulation
33. Skewness
1. binomial distributions
2. calculating
3. definition
4. lognormal distributions
5. normal distributions
6. option returns
7. Pearson coefficient of skewness
8. sample skewness
9. S&P 500 Index
10. variance vs. semivariance
34. Small-capitalization stocks 1. Carhart four-factor model
2. mean annual returns
3. month-of-year effects
4. returns of
5. Russell 2000 Index
35. Spearman rank correlation coefficient
36. Spreadsheets 1. internal rate of return
2. money-weighted rate of return
3. normal cumulative distribution function
4. precision of
5. p-value calculation
6. series of cash flows
37. Spurious correlation
38. Standard deviation
1. Bernoulli random variables
2. bond pull to par value
3. Chebyshev’s inequality
4. coefficient of variation
5. definition
6. earnings per share forecast
7. geometric mean return correlation
8. investment style comparison
9. mean absolute deviation versus
10. normal distributions
11. normal random variables
12. population standard deviation
13. portfolio standard deviation of return
14. of return
15. as risk measure
16. sample standard deviation,
17. semideviation
18. tracking risk
19. volatility
39. Standard error of estimate 1. adjusted
2. heteroskedasticity and
3. linear regression
4. multiple linear regression
5. robust
6. serial correlation and
7. serial-correlation consistent standard errors
40. Standard error of sample statistics
41. Standard error of time-series regression
42. Standardized beta
43. Standard normal distributions 1. definition
2. hypothesis testing
3. probability estimation
4. p-value via
5. standardizing random variables
6. standard normal probabilities
44. Standard normal random variable (Z)
45. Standard & Poor ’s bond ratings
46. Starbucks’ sales 1. Durbin–Watson statistic
2. linear trend regression
3. log-linear regression
4. seasonality
47. Stated annual interest rate
48. Stationarity 1. covariance stationary (see also Covariance stationarity)
2. definition
3. differencing random walks
4. forecasting model steps
5. nonstationarity definition
6. past vs. future
7. stationarity tests
49. Statistical factor models
50. Statistical inference 1. ARCH
2. definition
3. estimation (see Estimation)
4. heteroskedasticity and
5. hypothesis testing (see Hypothesis testing)
6. multicollinearity and
7. nonparametric
8. serial correlation and
9. serial correlation of regression errors
10. time series as covariance stationary
51. Statistical methods 1. ANOVA, (see also Analysis of variance)
2. central tendency (see also Central tendency)
3. dispersion, (see also Dispersion)
4. frequency distributions, (see also Frequency distributions)
5. geometric vs. arithmetic means
6. graphing data (see also Graphical representations of data)
7. kurtosis, (see also Kurtosis)
8. measurement scales
9. populations, (see also Populations)
10. quantiles
11. samples, (see also Samples)
12. skewness, (see also Skewness)
13. statistical inference, (see also Statistical inference)
14. statistics definition
52. Statistical significance
53. Statistics. See Sample statistics
54. Stocks 1. beta estimation via regression
2. bid–ask spreads explained
3. binomial model for price movement
4. correlations
5. dividends as distributions
6. fundamental factor models
7. global vs. US
8. inflation rates and stock returns
9. large-cap (see Large-capitalization stocks)
10. lognormal model for price
11. market timing,
12. mergers and acquisitions and returns
13. “momentum” stocks
14. month-of-year effects
15. mutual fund regression analysis
16. return and price relationship
17. return attribution
18. returns and inflation rate
19. returns and price-to-earnings ratio
20. screening criteria as independent
21. selling short
22. small-cap (see Small-capitalization stocks)
23. T-bills versus
24. time-series returns model
25. “value” stocks
55. Stratified sampling
56. Stratum of stratified sampling
57. Stress testing/scenario analysis
58. Student’s t-distribution. See also t-distributions
59. Subjective probability
60. Summation signs doubled
61. Sum of squared errors (SSE)
62. Sunk costs
63. Survey of Professional Forecasters (SPF)
64. Survivorship bias
65. Systematic risk 1. capital asset pricing model and
2. Carhart four-factor model
3. information ratio
66. Systematic sampling
T
1. Target semideviation
2. Target semivariance
3. Tax issues in cash flow
4. T-bills. See US Treasury bills
5. t-distributions 1. degrees of freedom
2. fat tail modeling
3. frequently referred to values
4. hypothesis testing
5. one-sided critical values of t
6. population mean confidence interval
7. “Student” pen name
6. Test statistic of hypothesis testing 1. definition
2. differences between means
3. on population variance
4. population variance unknown
5. power of a test
6. p-values from
7. rejection points
8. single mean
9. statistical significance
7. Time-period bias
8. Time-series data
1. ARCH
2. autocorrelations of
3. challenges of
4. cointegrated
5. covariance stationary
6. current and previous period relation
7. definition
8. earnings per share independence
9. exponential growth
10. Fisher effect
11. forecasting the past
12. GARCH
13. heteroskedasticity
14. linear regression
15. logical ordering of
16. longitudinal data
17. mean reversion
18. multiple linear regression
19. mutual fund performance
20. nonnormality of
21. observation notation
22. panel data
23. parameter instability
24. random walk differencing
25. random walks (see also Random walks)
26. regimes
27. regression misspecification
28. sample period length
29. sampling
30. seasonality
31. seasonal lag
32. Sharpe ratio computation
33. small sample unbiasedness
34. stationarity tests
35. time series mean
36. unit root test
9. Time-series models 1. advanced topics in
2. arbitrage pricing theory
3. ARCH
4. autocorrelations of error term
5. autoregression challenges
6. autoregressive models
7. autoregressive moving-average models
8. chain rule of forecasting
9. CPI model (see also Consumer Price Index)
10. forecasting model performance (see also Forecasting)
11. forecasting model steps
12. GARCH
13. heteroskedasticity
14. linear trend models
15. log-linear trend models
16. more than one time series
17. moving average
18. multiperiod forecasts
19. regression coefficient instability
20. sample period length
21. seasonality
22. serial correlation of residuals
23. Starbucks’ sales
24. trend models
25. with unit root
10. Time value of money 1. calculators vs. tables of factors
2. definition
3. future value of series of cash flows
4. future value of single cash flow
5. holding period return
6. interest rates
7. internal rate of return
8. money market yields
9. money-weighted rate of return
10. net present value
11. present and future value equivalence
12. present value of series of cash flows
13. present value of single cash flow
14. Rule of 72
15. solving for annuity payment size
16. solving for growth rate
17. solving for interest rate
18. solving for number of periods
19. time-weighted rate of return
20. weighted average cost of capital
11. Time-weighted rate of return
12. Tobin’s q
13. Total probability rule
14. Total probability rule for expected value
15. Tracking error 1. active risk
2. annualizing
3. binomial distribution of
4. definition
5. tracking risk
16. Tree diagrams 1. earnings per share analysis
2. option valuing
3. price movement as binomial model
17. Trend models 1. correlated errors
2. Durbin–Watson statistic
3. forecasting model steps
4. linear
5. log-linear trend models
6. moving average
7. serial correlation
8. trend definition
18. Treynor–Black appraisal ratio
19. Trials
20. Trimmed mean
21. Trimodal distributions
22. T. Rowe Price Equity Income Fund (PRFDX) 1. arithmetic mean
2. geometric mean
3. mean absolute deviation
4. range
5. sample excess kurtosis
6. sample standard deviation
7. sample variance
8. Sharpe ratio
9. skewness
23. t-tests 1. bivariate normal distributions
2. heteroskedasticity and
3. hypothesis testing of differences between means
4. hypothesis testing of single mean
5. hypothesis testing with linear regression
6. paired comparisons test
7. p-values via spreadsheet function
8. robustness
9. t-test of significance
10. z-tests versus
24. Two-tailed hypothesis tests 1. definition
2. instead of one-tailed
3. p-value
4. rejection points
5. significance level and confidence interval
25. Type I errors (α)
26. Type II errors (β)
U
1. Uncertainty 1. error term of regression
2. forecasting via time-series models
3. multiperiod forecasts
4. regression analysis
2. Unconditional heteroskedasticity
3. Unconditional probability
4. Underlying assets
5. Underlying populations 1. nonnormal
2. samples from differing
3. sample statistics estimates of
6. Unequal cash flows
7. Uniform distributions
8. Uniform random variables,
9. Unimodal distributions
10. United Kingdom 1. consol bonds
2. equivalent annual rate
3. Treasury bills
11. United States (US) 1. annual percentage yield
2. CPI model (see also Consumer Price Index)
3. T-bills (see US Treasury bills)
12. US Treasury bills 1. bank discount basis
2. correlation of bonds and T-bills
3. discount
4. effective annual yield
5. exchange traded funds versus
6. face value
7. holding period yield
8. international equivalents
9. liquidity premium
10. as pure discount instruments
11. regressing on predicted inflation
12. risk-free interest rate
13. Unit normal distribution
14. Unit root test
15. Univariate normal distributions
16. Up transition probability
17. Utility functions
V
1. Valuation 1. correlations among performance measures
2. linear regression hypothesis testing
3. multiple linear regression for
4. predicting EV/IC ratio
5. Tobin’s q
2. Value/growth investment style
3. “Value” stocks
4. VAR (value at risk)
5. Variance 1. Bernoulli random variables
2. binomial random variables
3. conditional variances
4. confidence intervals
5. constant plus random variable
6. constant times random variable
7. covariance
8. covariance effect on
9. definition,
10. diagonal vs. off-diagonal
11. differentiation of
12. earnings per share forecast
13. forecast model errors
14. hypothesis testing on single variance
15. hypothesis testing on two variances
16. lognormal distributions
17. multiperiod vs. single-period forecasts
18. multivariate normal distributions
19. normal distributions
20. own covariance as
21. population variance
22. population variance known
23. population variance unknown
24. portfolio expected return and variance
25. random variables
26. returns pre-and post-crisis
27. as risk measure
28. sample variance
29. semivariance
30. stationarity (see also Stationarity)
31. unconditional variance
6. Volatility 1. definition
2. option pricing models and
3. pre-and post-crisis
4. quadruple witching days
5. well-diversified portfolios
W
1. Weakly stationary. See Covariance stationary
2. Weighted average cost of capital (WACC)
3. Weighted averages of portfolio return
4. Weighted mean 1. arithmetic mean as
2. expected value
3. harmonic mean as
4. market indexes as
5. mean returns
6. portfolio expected return as
7. total probability rule
8. weights sum to one
5. White-corrected standard errors
6. Wholesale clubs
7. Wilcoxon signed-rank test
8. Winsorized mean
9. Working capital management
Y
1. Yield to maturity (YTM)
Z
1. z-alternative
2. z-distributions. See Standard normal distributions
3. Zero coupon bonds
4. z-tests 1. hypothesis testing
2. p-value via spreadsheet function
3. rejection points
4. t-tests versus
WILEY END USER LICENSE AGREEMENT Go to www.wiley.com/go/eula to access Wiley’s ebook EULA.
Table of Contents Foreword Preface Acknowledgment About the CFA Institute Investment Series CHAPTER 1: The Time Value of Money
1. Introduction 2. Interest Rates: Interpretation 3. The Future Value of a Single Cash Flow 4. The Future Value of a Series of Cash Flows 5. The Present Value of a Single Cash Flow 6. The Present Value of a Series of Cash Flows 7. Solving for Rates, Number of Periods, or Size of Annuity Payments 8. Summary Problems Notes
CHAPTER 2: Discounted Cash Flow Applications 1. Introduction 2. Net Present Value and Internal Rate of Return 3. Portfolio Return Measurement 4. Money Market Yields 5. Summary References Problems Notes
CHAPTER 3: Statistical Concepts and Market Returns 1. Introduction 2. Some Fundamental Concepts 3. Summarizing Data Using Frequency Distributions 4. The Graphic Presentation of Data 5. Measures of Central Tendency 6. Other Measures of Location: Quantiles 7. Measures of Dispersion 8. Symmetry and Skewness in Return Distributions 9. Kurtosis in Return Distributions 10. Using Geometric and Arithmetic Means 11. Summary References
Problems Notes
CHAPTER 4: Probability Concepts 1. Introduction 2. Probability, Expected Value, and Variance 3. Portfolio Expected Return and Variance of Return 4. Topics in Probability 5. Summary References Problems Notes
CHAPTER 5: Common Probability Distributions 1. Introduction to Common Probability Distributions 2. Discrete Random Variables 3. Continuous Random Variables 4. Monte Carlo Simulation 5. Summary References Problems Notes
CHAPTER 6: Sampling and Estimation 1. Introduction 2. Sampling 3. Distribution of the Sample Mean 4. Point and Interval Estimates of the Population Mean 5. More on Sampling 6. Summary References Problems Notes
CHAPTER 7: Hypothesis Testing 1. Introduction 2. Hypothesis Testing 3. Hypothesis Tests Concerning the Mean 4. Hypothesis Tests Concerning Variance 5. Other Issues: Nonparametric Inference 6. Summary References Problems Notes
CHAPTER 8: Correlation and Regression 1. Introduction 2. Correlation Analysis 3. Linear Regression 4. Summary Problems Notes
CHAPTER 9: Multiple Regression and Issues in Regression Analysis 1. Introduction 2. Multiple Linear Regression 3. Using Dummy Variables in Regressions 4. Violations of Regression Assumptions 5. Model Specification and Errors in Specification 6. Models with Qualitative Dependent Variables 7. Summary References Problems Notes
CHAPTER 10 : Time-Series Analysis 1. Introduction to Time-Series Analysis 2. Challenges of Working with Time Series 3. Trend Models 4. Autoregressive (AR) Time-Series Models 5. Random Walks and Unit Roots 6. Moving-Average Time-Series Models 7. Seasonality in Time-Series Models 8. Autoregressive Moving-Average Models 9. Autoregressive Conditional Heteroskedasticity Models 10. Regressions with More than One Time Series 11. Other Issues in Time Series 12. Suggested Steps in Time-Series Forecasting 13. Summary Problems Notes
CHAPTER 11: An Introduction to Multifactor Models 1. Introduction 2. Multifactor Models and Modern Portfolio Theory 3. Arbitrage Pricing Theory 4. Multifactor Models: Types 5. Multifactor Models: Selected Applications
6. Summary References Problems
Appendices Glossary About the Editors and Authors About the CFA Program Index Advert EULA
- Foreword
- Preface
- Acknowledgment
- About the CFA Institute Investment Series
- CHAPTER 1: The Time Value of Money
- 1. Introduction
- 2. Interest Rates: Interpretation
- 3. The Future Value of a Single Cash Flow
- 4. The Future Value of a Series of Cash Flows
- 5. The Present Value of a Single Cash Flow
- 6. The Present Value of a Series of Cash Flows
- 7. Solving for Rates, Number of Periods, or Size of Annuity Payments
- 8. Summary
- Problems
- Notes
- CHAPTER 2: Discounted Cash Flow Applications
- 1. Introduction
- 2. Net Present Value and Internal Rate of Return
- 3. Portfolio Return Measurement
- 4. Money Market Yields
- 5. Summary
- References
- Problems
- Notes
- CHAPTER 3: Statistical Concepts and Market Returns
- 1. Introduction
- 2. Some Fundamental Concepts
- 3. Summarizing Data Using Frequency Distributions
- 4. The Graphic Presentation of Data
- 5. Measures of Central Tendency
- 6. Other Measures of Location: Quantiles
- 7. Measures of Dispersion
- 8. Symmetry and Skewness in Return Distributions
- 9. Kurtosis in Return Distributions
- 10. Using Geometric and Arithmetic Means
- 11. Summary
- References
- Problems
- Notes
- CHAPTER 4: Probability Concepts
- 1. Introduction
- 2. Probability, Expected Value, and Variance
- 3. Portfolio Expected Return and Variance of Return
- 4. Topics in Probability
- 5. Summary
- References
- Problems
- Notes
- CHAPTER 5: Common Probability Distributions
- 1. Introduction to Common Probability Distributions
- 2. Discrete Random Variables
- 3. Continuous Random Variables
- 4. Monte Carlo Simulation
- 5. Summary
- References
- Problems
- Notes
- CHAPTER 6: Sampling and Estimation
- 1. Introduction
- 2. Sampling
- 3. Distribution of the Sample Mean
- 4. Point and Interval Estimates of the Population Mean
- 5. More on Sampling
- 6. Summary
- References
- Problems
- Notes
- CHAPTER 7: Hypothesis Testing
- 1. Introduction
- 2. Hypothesis Testing
- 3. Hypothesis Tests Concerning the Mean
- 4. Hypothesis Tests Concerning Variance
- 5. Other Issues: Nonparametric Inference
- 6. Summary
- References
- Problems
- Notes
- CHAPTER 8: Correlation and Regression
- 1. Introduction
- 2. Correlation Analysis
- 3. Linear Regression
- 4. Summary
- Problems
- Notes
- CHAPTER 9: Multiple Regression and Issues in Regression Analysis
- 1. Introduction
- 2. Multiple Linear Regression
- 3. Using Dummy Variables in Regressions
- 4. Violations of Regression Assumptions
- 5. Model Specification and Errors in Specification
- 6. Models with Qualitative Dependent Variables
- 7. Summary
- References
- Problems
- Notes
- CHAPTER 10 : Time-Series Analysis
- 1. Introduction to Time-Series Analysis
- 2. Challenges of Working with Time Series
- 3. Trend Models
- 4. Autoregressive (AR) Time-Series Models
- 5. Random Walks and Unit Roots
- 6. Moving-Average Time-Series Models
- 7. Seasonality in Time-Series Models
- 8. Autoregressive Moving-Average Models
- 9. Autoregressive Conditional Heteroskedasticity Models
- 10. Regressions with More than One Time Series
- 11. Other Issues in Time Series
- 12. Suggested Steps in Time-Series Forecasting
- 13. Summary
- Problems
- Notes
- CHAPTER 11: An Introduction to Multifactor Models
- 1. Introduction
- 2. Multifactor Models and Modern Portfolio Theory
- 3. Arbitrage Pricing Theory
- 4. Multifactor Models: Types
- 5. Multifactor Models: Selected Applications
- 6. Summary
- References
- Problems
- Appendices
- Glossary
- About the Editors and Authors
- About the CFA Program
- Index
- Advert
- EULA