probability,random variables

profileVAMSI
SchaumsOutlineofProbabilityRandomVariablesandRandomProcessesFourthEditionbyHweiP.Hsuz-lib.org.pdf

HWEI P. HSU received his BS from National Taiwan University and his MS and PhD from Case Institute of Technology. He has published several books which include Schaum’s Outline of Analog and Digital Communications and Schaum’s Outline of Signals and Systems.

Copyright © 2020, 2013, 2009, 1998 by McGraw-Hill Education. All rights reserved. Except as permitted under the United States Copyright Act of 1976, no part of this publication may be reproduced or distributed in any form or by any means, or stored in a database or retrieval system, without the prior written permission of the publisher.

ISBN: 978-1-26-045382-9 MHID: 1-26-045382-0

The material in this eBook also appears in the print version of this title: ISBN: 978-1-26-045381-2, MHID: 1-26-045381-2.

eBook conversion by codeMantra Version 1.0

All trademarks are trademarks of their respective owners. Rather than put a trademark symbol after every occurrence of a trademarked name, we use names in an editorial fashion only, and to the benefit of the trademark owner, with no intention of infringement of the trademark. Where such designations appear in this book, they have been printed with initial caps.

McGraw-Hill Education eBooks are available at special quantity discounts to use as premiums and sales promotions or for use in corporate training programs. To contact a representative, please visit the Contact Us page at www.mhprofessional.com.

McGraw-Hill Education, the McGraw-Hill Education logo, Schaum’s, and related trade dress are trademarks or registered trademarks of McGraw-Hill Education and/or its affiliates in the United States and other countries and may not be used without written permission. All other trademarks are the

property of their respective owners. McGraw-Hill Education is not associated with any product or vendor mentioned in this book.

TERMS OF USE

This is a copyrighted work and McGraw-Hill Education and its licensors reserve all rights in and to the work. Use of this work is subject to these terms. Except as permitted under the Copyright Act of 1976 and the right to store and retrieve one copy of the work, you may not decompile, disassemble, reverse engineer, reproduce, modify, create derivative works based upon, transmit, distribute, disseminate, sell, publish or sublicense the work or any part of it without McGraw-Hill Education’s prior consent. You may use the work for your own noncommercial and personal use; any other use of the work is strictly prohibited. Your right to use the work may be terminated if you fail to comply with these terms.

THE WORK IS PROVIDED “AS IS.” McGRAW-HILL EDUCATION AND ITS LICENSORS MAKE NO GUARANTEES OR WARRANTIES AS TO THE ACCURACY, ADEQUACY OR COMPLETENESS OF OR RESULTS TO BE OBTAINED FROM USING THE WORK, INCLUDING ANY INFORMATION THAT CAN BE ACCESSED THROUGH THE WORK VIA HYPERLINK OR OTHERWISE, AND EXPRESSLY DISCLAIM ANY WARRANTY, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO IMPLIED WARRANTIES OF MERCHANTABILITY OR FITNESS FOR A PARTICULAR PURPOSE. McGraw-Hill Education and its licensors do not warrant or guarantee that the functions contained in the work will meet your requirements or that its operation will be uninterrupted or error free. Neither McGraw-Hill Education nor its licensors shall be liable to you or anyone else for any inaccuracy, error or omission, regardless of cause, in the work or for any damages resulting therefrom. McGraw-Hill Education has no responsibility for the content of any information accessed through the work. Under no circumstances shall McGraw-Hill Education and/or its licensors be liable for any indirect, incidental, special, punitive, consequential or similar damages that result from the use of or inability to use the work, even if any of them has been advised of the possibility of such damages. This limitation

of liability shall apply to any claim or cause whatsoever whether such claim or cause arises in contract, tort or otherwise.

Preface to The Second Edition

The purpose of this book, like its previous edition, is to provide an introduction to the principles of probability, random variables, and random processes and their applications.

The book is designed for students in various disciplines of engineering, science, mathematics, and management. The background required to study the book is one year of calculus, elementary differential equations, matrix analysis, and some signal and system theory, including Fourier transforms. The book can be used as a self-contained textbook or for self-study. Each topic is introduced in a chapter with numerous solved problems. The solved problems constitute an integral part of the text.

This new edition includes and expands the contents of the first edition. In addition to refinement through the text, two new sections on probability- generating functions and martingales have been added and a new chapter on information theory has been added.

I wish to thank my granddaughter Elysia Ann Krebs for helping me in the preparation of this revision. I also wish to express my appreciation to the editorial staff of the McGraw-Hill Schaum’s Series for their care, cooperation, and attention devoted to the preparation of the book.

HWEI P. HSU Shannondell at Valley Forge, Audubon, Pennsylvania

Preface to The First Edition

The purpose of this book is to provide an introduction to the principles of probability, random variables, and random processes and their applications.

The book is designed for students in various disciplines of engineering, science, mathematics, and management. It may be used as a textbook and/or a supplement to all current comparable texts. It should also be useful to those interested in the field of self-study. The book combines the advantages of both the textbook and the so-called review book. It provides the textual explanations of the textbook, and in the direct way characteristic of the review book, it gives hundreds of completely solved problems that use essential theory and techniques. Moreover, the solved problems are an integral part of the text. The background required to study the book is one year of calculus, elementary differential equations, matrix analysis, and some signal and system theory, including Fourier transforms.

I wish to thank Dr. Gordon Silverman for his invaluable suggestions and critical review of the manuscript. I also wish to express my appreciation to the editorial staff of the McGraw-Hill Schaum Series for their care, cooperation, and attention devoted to the preparation of the book. Finally, I thank my wife, Daisy, for her patience and encouragement.

HWEI P. HSU Montville, New Jersey

Contents

CHAPTER 1 Probability

1.1 Introduction 1.2 Sample Space and Events 1.3 Algebra of Sets 1.4 Probability Space 1.5 Equally Likely Events 1.6 Conditional Probability 1.7 Total Probability 1.8 Independent Events

Solved Problems

CHAPTER 2 Random Variables

2.1 Introduction 2.2 Random Variables 2.3 Distribution Functions 2.4 Discrete Random Variables and Probability Mass

Functions 2.5 Continuous Random Variables and Probability

Density Functions 2.6 Mean and Variance 2.7 Some Special Distributions 2.8 Conditional Distributions

Solved Problems

CHAPTER 3 Multiple Random Variables

3.1 Introduction 3.2 Bivariate Random Variables

3.3 Joint Distribution Functions 3.4 Discrete Random Variables—Joint Probability Mass

Functions 3.5 Continuous Random Variables—Joint Probability

Density Functions 3.6 Conditional Distributions 3.7 Covariance and Correlation Coefficient 3.8 Conditional Means and Conditional Variances 3.9 N-Variate Random Variables 3.10 Special Distributions

Solved Problems

CHAPTER 4 Functions of Random Variables, Expectation, Limit Theorems

4.1 Introduction 4.2 Functions of One Random Variable 4.3 Functions of Two Random Variables 4.4 Functions of n Random Variables 4.5 Expectation 4.6 Probability Generating Functions 4.7 Moment Generating Functions 4.8 Characteristic Functions 4.9 The Laws of Large Numbers and the Central Limit

Theorem Solved Problems

CHAPTER 5 Random Processes

5.1 Introduction 5.2 Random Processes 5.3 Characterization of Random Processes 5.4 Classification of Random Processes 5.5 Discrete-Parameter Markov Chains

5.6 Poisson Processes 5.7 Wiener Processes 5.8 Martingales

Solved Problems

CHAPTER 6 Analysis and Processing of Random Processes

6.1 Introduction 6.2 Continuity, Differentiation, Integration 6.3 Power Spectral Densities 6.4 White Noise 6.5 Response of Linear Systems to Random Inputs 6.6 Fourier Series and Karhunen-Loéve Expansions 6.7 Fourier Transform of Random Processes

Solved Problems

CHAPTER 7 Estimation Theory

7.1 Introduction 7.2 Parameter Estimation 7.3 Properties of Point Estimators 7.4 Maximum-Likelihood Estimation 7.5 Bayes’ Estimation 7.6 Mean Square Estimation 7.7 Linear Mean Square Estimation

Solved Problems

CHAPTER 8 Decision Theory

8.1 Introduction 8.2 Hypothesis Testing 8.3 Decision Tests

Solved Problems

CHAPTER 9 Queueing Theory

9.1 Introduction 9.2 Queueing Systems 9.3 Birth-Death Process 9.4 The M/M/1

Queueing System 9.5 The M/M/s Queueing System 9.6 The M/M/1/K Queueing System 9.7 The M/M/s/K Queueing System

Solved Problems

CHAPTER 10 Information Theory

10.1 Introduction 10.2 Measure of

Information 10.3 Discrete Memoryless Channels 10.4 Mutual Information 10.5 Channel Capacity 10.6 Continuous Channel 10.7 Additive White Gaussian Noise Channel 10.8 Source Coding 10.9 Entropy Coding

Solved Problems

APPENDIX A Normal Distribution

APPENDIX B Fourier Transform

B.1 Continuous-Time Fourier Transform

B.2 Discrete-Time Fourier Transform

INDEX

*The laptop icon next to an exercise indicates that the exercise is also available as a video with step- by-step instructions. These videos are available on the Schaums.com website by following the instructions on the inside front cover.

CHAPTER 1

Probability

1.1 Introduction

The study of probability stems from the analysis of certain games of chance, and it has found applications in most branches of science and engineering. In this chapter the basic concepts of probability theory are presented.

1.2 Sample Space and Events

A. Random Experiments:

In the study of probability, any process of observation is referred to as an experiment. The results of an observation are called the outcomes of the experiment. An experiment is called a random experiment if its outcome cannot be predicted. Typical examples of a random experiment are the roll of a die, the toss of a coin, drawing a card from a deck, or selecting a message signal for transmission from several messages.

B. Sample Space: The set of all possible outcomes of a random experiment is called the sample space (or universal set), and it is denoted by S. An element in S is called a sample point. Each outcome of a random experiment corresponds to a sample point.

EXAMPLE 1.1 Find the sample space for the experiment of tossing a coin (a) once and (b) twice.

(a) There are two possible outcomes, heads or tails. Thus:

where H and T represent head and tail, respectively.

(b) There are four possible outcomes. They are pairs of heads and tails. Thus:

EXAMPLE 1.2 Find the sample space for the experiment of tossing a coin repeatedly and of counting the number of tosses required until the first head appears.

Clearly all possible outcomes for this experiment are the terms of the sequence 1, 2, 3, … Thus:

Note that there are an infinite number of outcomes.

EXAMPLE 1.3 Find the sample space for the experiment of measuring (in hours) the lifetime of a transistor.

Clearly all possible outcomes are all nonnegative real numbers. That is,

where τ represents the life of a transistor in hours. Note that any particular experiment can often have many different sample

spaces depending on the observation of interest (Probs. 1.1 and 1.2). A sample space S is said to be discrete if it consists of a finite number of sample points (as in Example 1.1) or countably infinite sample points (as in Example 1.2). A set is called countable if its elements can be placed in a one-to-one correspondence with the positive integers. A sample space S is said to be continuous if the sample points constitute a continuum (as in Example 1.3).

C. Events: Since we have identified a sample space S as the set of all possible outcomes of a random experiment, we will review some set notations in the following.

If ζ is an element of S (or belongs to S), then we write

If S is not an element of S (or does not belong to S), then we write

A set A is called a subset of B, denoted by

if every element of A is also an element of B. Any subset of the sample space S is called an event. A sample point of S is often referred to as an elementary event. Note that the sample space S is the subset of itself: that is, S ⊂ S. Since S is the set of all possible outcomes, it is often called the certain event.

EXAMPLE 1.4 Consider the experiment of Example 1.2. Let A be the event that the number of tosses required until the first head appears is even. Let B be the event that the number of tosses required until the first head appears is odd. Let C be the event that the number of tosses required until the first head appears is less than 5. Express events A, B, and C.

1.3 Algebra of Sets

A. Set Operations:

1. Equality: Two sets A and B are equal, denoted A = B, if and only if A ⊂ B and B ⊂ A.

2. Complementation: Suppose A ⊂ S. The complement of set A, denoted , is the set containing all elements in S but not in A.

3. Union:

The union of sets A and B, denoted A ∪ B, is the set containing all elements in either A or B or both.

4. Intersection: The intersection of sets A and B, denoted A ∩ B, is the set containing all elements in both A and B.

5. Difference: The difference of sets A and B, denoted A\ B, is the set containing all elements in A but not in B.

Note that .

6. Symmetrical Difference: The symmetrical difference of sets A and B, denoted A Δ B, is the set of all elements that are in A or B but not in both.

Note that .

7. Null Set: The set containing no element is called the null set, denoted Ø. Note that

8. Disjoint Sets: Two sets A and B are called disjoint or mutually exclusive if they contain no common element, that is, if A ∩ B = Ø.

The definitions of the union and intersection of two sets can be extended to any finite number of sets as follows:

Note that these definitions can be extended to an infinite number of sets:

In our definition of event, we state that every subset of S is an event, including S and the null set Ø Then

If A and B are events in S, then

Similarly, if A1, A2, …, An are a sequence of events in S, then

9. Partition of S :

If Ai ∩ Aj = Ø for i ≠ j and , then the collection {Ai; 1 ≤ i ≤ k} is said to

form a partition of S.

10. Size of Set: When sets are countable, the size (or cardinality) of set A, denoted |A|, is the number of elements contained in A. When sets have a finite number of elements, it is easy to see that size has the following properties:

Note that the property (iv) can be easily seen if A and B are subsets of a line with length |A| and |B|, respectively.

11. Product of Sets: The product (or Cartesian product) of sets A and B, denoted by A × B, is the set of ordered pairs of elements from A and B.

Note that A × B ≠ B × A, and |C| = |A × B| = |A| × |B|.

EXAMPLE 1.5 Let A = {a1, a2, a3} and B = {b1, b2}. Then

B. Venn Diagram: A graphical representation that is very useful for illustrating set operation is the Venn diagram. For instance, in the three Venn diagrams shown in Fig. 1-1, the shaded areas represent, respectively, the events A ∪ B, A ∩ B, and . The Venn

diagram in Fig. 1-2(a) indicates that B ⊂ A, and the event A ∩ = A\B is shown as the shaded area. In Fig. 1-2(b), the shaded area represents the event A Δ B.

Fig. 1-1

Fig. 1-2

C. Identities: By the above set definitions or reference to Fig. 1-1, we obtain the following identities:

The union and intersection operations also satisfy the following laws:

Commutative Laws:

Associative Laws:

Distributive Laws:

De Morgan’s Laws:

These relations are verified by showing that any element that is contained in the set on the left side of he equality sign is also contained in the set on the right side, and vice versa. One way of showing this is by means of a Venn diagram (Prob. 1.14). The distributive laws can be extended as follows:

Similarly, De Morgan’s laws also can be extended as follows (Prob. 1.21):

1.4 Probability Space

A. Event Space:

We have defined that events are subsets of the sample space S. In order to be precise, we say that a subset A of S can be an event if it belongs to a collection F of subsets of S, satisfying the following conditions:

The collection F is called an event space. In mathematical literature, event space is known as sigma field (σ-field) or σ-algebra.

Using the above conditions, we can show that if A and B are in F, then so are

EXAMPLE 1.6 Consider the experiment of tossing a coin once in Example 1.1. We have S = {H, T}. The set {S, Ø}, {S, Ø, H, T} are event spaces, but {S, Ø, H} is not an event space, since = T is not in the set.

B. Probability Space: An assignment of real numbers to the events defined in an event space F is known as the probability measure P. Consider a random experiment with a sample space S, and let A be a particular event defined in F. The probability of the event A is denoted by P(A). Thus, the probability measure is a function defined over F. The triplet (S, F, P) is known as the probability space.

C. Probability Measure a. Classical Definition: Consider an experiment with equally likely finite outcomes. Then the classical definition of probability of event A, denoted P(A), is defined by

If A and B are disjoint, i.e., A ∩ B = Ø, then, |A ∪ B| = |A| + |B|. Hence, in this case

We also have

EXAMPLE 1.7 Consider an experiment of rolling a die. The outcome is

Define: A: the event that outcome is even, i.e., A = {2, 4, 6} B: the event that outcome is odd, i.e., B = {1, 3, 5} C: the event that outcome is prime, i.e., C = {1, 2, 3, 5}

Then

Note that in the classical definition, P(A) is determined a priori without actual experimentation and the definition can be applied only to a limited class of problems such as only if the outcomes are finite and equally likely or equally probable.

b. Relative Frequency Definition: Suppose that the random experiment is repeated n times. If event A occurs n(A) times, then the probability of event A, denoted P(A), is defined as

where n(A)/n is called the relative frequency of event A. Note that this limit may not exist, and in addition, there are many situations in which the concepts of repeatability may not be valid. It is clear that for any event A, the relative frequency of A will have the following properties:

1. 0 ≤ n(A)/n ≤ 1, where n(A)/n = 0 if A occurs in none of the n repeated trials and n(A)/n = 1 if A occurs in all of the n repeated trials.

2. If A and B are mutually exclusive events, then

and

c. Axiomatic Definition: Consider a probability space (S, F, P). Let A be an event in F. Then in the axiomatic definition, the probability P(A) of the event A is a real number assigned to A which satisfies the following three axioms:

If the sample space S is not finite, then axiom 3 must be modified as follows: Axiom 3′: If A1, A2, … is an infinite sequence of mutually exclusive events in S (Ai ∩ Aj = Ø for i ≠ j), then

These axioms satisfy our intuitive notion of probability measure obtained from the notion of relative frequency.

d. Elementary Properties of Probability: By using the above axioms, the following useful properties of probability can be obtained:

where the sum of the second term is over all distinct pairs of events, that of the third term is over all distinct triples of events, and so forth.

8. If A1, A2, …, An is a finite sequence of mutually exclusive events in S (Ai ∩ Aj = Ø for i ≠ j), then

and a similar equality holds for any subcollection of the events. Note that property 4 can be easily derived from axiom 2 and property 3.

Since A ⊂ S, we have

Thus, combining with axiom 1, we obtain

Property 5 implies that

since P(A ∩ B) ≥ 0 by axiom 1.

Property 6 implies that

since A ∩ B = B if B ⊂ A.

1.5 Equally Likely Events

A. Finite Sample Space:

Consider a finite sample space S with n finite elements

where ζi’s are elementary events. Let P(ζi) = pi. Then

B. Equally Likely Events: When all elementary events ζi(i = 1, 2, …, n) are equally likely, that is,

then from Eq. (1.51), we have

and

where n(A) is the number of outcomes belonging to event A and n is the number of sample points in S. [See classical definition (1.29).]

1.6 Conditional Probability

A. Definition:

The conditional probability of an event A given event B, denoted by P(A | B), is defined as

where P(A ∩ B) is the joint probability of A and B. Similarly,

is the conditional probability of an event B given event A. From Eqs. (1.55) and (1.56), we have

Equation (1.57) is often quite useful in computing the joint probability of events.

B. Bayes’ Rule: From Eq. (1.57) we can obtain the following Bayes’ rule:

1.7 Total Probability

The events A1, A2, …, An are called mutually exclusive and exhaustive if

Let B be any event in S. Then

which is known as the total probability of event B (Prob. 1.57). Let A = Ai in Eq. (1.58); then, using Eq. (1.60), we obtain

Note that the terms on the right-hand side are all conditioned on events Ai, while the term on the left is conditioned on B. Equation (1.61) is sometimes referred to as Bayes’ theorem.

1.8 Independent Events

Two events A and B are said to be (statistically) independent if and only if

It follows immediately that if A and B are independent, then by Eqs. (1.55) and (1.56),

If two events A and B are independent, then it can be shown that A and are also independent; that is (Prob. 1.63),

Then

Thus, if A is independent of B, then the probability of A‘s occurrence is unchanged by information as to whether or not B has occurred. Three events A, B, C are said to be independent if and only if

We may also extend the definition of independence to more than three events. The events A1, A2, …, An are independent if and only if for every subset {Ai1, Ai2, … Aik} (2 ≤ k ≤ n) of these events,

Finally, we define an infinite set of events to be independent if and only if every finite subset of these events is independent.

To distinguish between the mutual exclusiveness (or disjointness) and independence of a collection of events, we summarize as follows:

1. {Ai, i = 1, 2, …, n} is a sequence of mutually exclusive events, then

2. If {Ai, i = 1, 2, …, n} is a sequence of independent events, then

and a similar equality holds for any subcollection of the events.

SOLVED PROBLEMS

Sample Space and Events

1.1. Consider a random experiment of tossing a coin three times. (a) Find the sample space S1 if we wish to observe the exact sequences of

heads and tails obtained. (b) Find the sample space S2 if we wish to observe the number of heads in

the three tosses.

(a) The sampling space S1 is given by

where, for example, HTH indicates a head on the first and third throws and a tail on the second throw. There are eight sample points in S1.

(b) The sampling space S2 is given by

where, for example, the outcome 2 indicates that two heads were obtained in the three tosses. The sample space S2 contains four sample points.

1.2. Consider an experiment of drawing two cards at random from a bag containing four cards marked with the integers 1 through 4. (a) Find the sample space S1 of the experiment if the first card is replaced

before the second is drawn. (b) Find the sample space S2 of the experiment if the first card is not

replaced.

(a) The sample space S1 contains 16 ordered pairs (i, j), 1 ≤ i ≤ 4, 1 ≤ j ≤ 4, where the first number indicates the first number drawn. Thus,

(b) The sample space S2 contains 12 ordered pairs (i, j), i ≠ j, 1 ≤ i ≤ 4, 1 ≤ j ≤ 4, where the first number indicates the first number drawn. Thus,

1.3. An experiment consists of rolling a die until a 6 is obtained. (a) Find the sample space S1 if we are interested in all possibilities. (b) Find the sample space S2 if we are interested in the number of throws

needed to get a 6.

(a) The sample space S1 would be

where the first line indicates that a 6 is obtained in one throw, the second line indicates that a 6 is obtained in two throws, and so forth.

(b) In this case, the sample space S2 is

where i is an integer representing the number of throws needed to get a 6.

1.4. Find the sample space for the experiment consisting of measurement of the voltage output v from a transducer, the maximum and minimum of which are + 5 and − 5 volts, respectively.

A suitable sample space for this experiment would be

1.5. An experiment consists of tossing two dice. (a) Find the sample space S. (b) Find the event A that the sum of the dots on the dice equals 7. (c) Find the event B that the sum of the dots on the dice is greater than 10. (d) Find the event C that the sum of the dots on the dice is greater than 12.

(a) For this experiment, the sample space S consists of 36 points (Fig. 1-3):

Fig. 1-3

where i represents the number of dots appearing on one die and j represents the number of dots appearing on the other die.

(b) The event A consists of 6 points (see Fig. 1-3):

(c) The event B consists of 3 points (see Fig. 1-3):

(d) The event C is an impossible event, that is, C = Ø.

1.6. An automobile dealer offers vehicles with the following options: (a) With or without automatic transmission

(b) With or without air-conditioning (c) With one of two choices of a stereo system (d) With one of three exterior colors If the sample space consists of the set of all possible vehicle types, what is the number of outcomes in the sample space?

The tree diagram for the different types of vehicles is shown in Fig. 1-4. From Fig. 1-4 we see that the number of sample points in S is 2 × 2 × 2 × 3 = 24.

Fig. 1-4

1.7. State every possible event in the sample space S = {a, b, c, d}.

There are 24= 16 possible events in S. They are Ø; {a}, {b}, {c}, {d}; {a, b}, {a, c}, {a, d}, {b, c}, {b, d}, {c, d}; {a, b, c}, {a, b, d}, {a, c, d}, {b, c, d}; S = {a, b, c, d}.

1.8. How many events are there in a sample space S with n elementary events?

Let S = {s1, s2, …, sn}. Let Ω be the family of all subsets of S. (Ω is sometimes referred to as the power set of S.) Let Si be the set consisting of two statements, that is,

Then Ω can be represented as the Cartesian product

Since each subset of S can be uniquely characterized by an element in the above Cartesian product, we obtain the number of elements in Ω by

where n(Si) = number of elements in Si = 2.

An alternative way of finding n(Ω) is by the following summation:

The last sum is an expansion of (1+1)n = 2n.

Algebra of Sets

1.9. Consider the experiment of Example 1.2. We define the events

where k is the number of tosses required until the first H (head) appears. Determine the events , , , A ∪ B, B ∪ C, A ∩ B, A ∩ C, B ∩ C, and ∩ B.

1.10. Consider the experiment of Example 1.7 of rolling a die. Express

From Example 1.7, we have S = {1, 2, 3, 4, 5, 6}, A = {2, 4, 6}, B = {1, 3, 5}, and C = {1, 2, 3, 5}. Then

1.11. The sample space of an experiment is the real line express as

(a) Consider the events

Determine the events

(b) Consider the events

Determine the events

(a) It is clear that

Noting that the Ai’s are mutually exclusive, we have

(b) Noting that B1 ⊃ B2 ⊃ … ⊃ Bi ⊃ …, we have

1.12. Consider the switching networks shown in Fig. 1-5. Let A1, A2, and A3 denote the events that the switches s1, s2, and s3 are closed, respectively. Let Aab denote the event that there is a closed path between terminals a and b. Express Aab in terms of A1, A2, and A3 for each of the networks shown.

Fig. 1-5

(a) From Fig. 1-5(a), we see that there is a closed path between a and b only if all switches s1, s2, and s3 are closed. Thus,

(b) From Fig. 1-5(b), we see that there is a closed path between a and b if at least one switch is closed. Thus,

(c) From Fig. 1-5(c), we see that there is a closed path between a and b if s1 and either s2 or s3 are closed. Thus,

Using the distributive law (1.18), we have

which indicates that there is a closed path between a and b if s1 and s2 or s1 and s3 are closed.

(d) From Fig. 1-5(d), we see that there is a closed path between a and b if either s1 and s2 are closed or s3 is closed. Thus

1.13. Verify the distributive law (1.18).

Let s ∈ [A ∩ (B ∪ C)]. Then s ∈ A and s ∈ (B ∪ C). This means either that s ∈ A and s ∈ B or that s ∈ A and s ∈ C; that is, s ∈ (A ∩ B) or s ∈ (A ∩ C). Therefore,

Next, let s ∈ [(A ∩ B) ∪ (A ∩ C)]. Then s ∈ A and s ∈ B or s ∈ A and s ∈ C. Thus, s ∈ A and (s ∈ B or s ∈ C). Thus,

Thus, by the definition of equality, we have

1.14. Using a Venn diagram, repeat Prob. 1.13.

Fig. 1-6 shows the sequence of relevant Venn diagrams. Comparing Fig. 1- 6(b) and 1.6(e), we conclude that

Fig. 1-6

1.15 Verify De Morgan’s law (1.24)

Suppose that , then s ∉ A ∩ B. So s ∉ {both A and B}. This means that either s ∉ A or s ∉ B or s ∉ A and s ∉ B. This implies that . Conversely, suppose that , that is either or or s ∉ {both

and }. Then it follows that s ∉ A or s ∉ B or s ∉ {both A and B}; that is, s ∉ A ∩ B or . Thus, we conclude that .

Note that De Morgan’s law can also be shown by using Venn diagram.

1.16. Let A and B be arbitrary events. Show that A ⊂ B if and only if A ∩ B = A.

“If” part: We show that if A ∩ B = A, then A ⊂ B. Let s ∈ A. Then s ∈ (A ∩ B), since A = A ∩ B. Then by the definition of intersection, s ∈ B. Therefore, A ⊃ B.

“Only if” part: We show that if A ⊂ B, then A ∩ B = A. Note that from the definition of the intersection, (A = B) ⊂ A. Suppose s ∈ A. If A ⊂ B, then s ∈ B. So s ∈ A and s ∈ B; that is, s ∈(A ∩ B). Therefore, it follows that A ⊂ (A ∩ B). Hence, A = A ∩ B. This completes the proof.

1.17. Let A be an arbitrary event in S and let Ø be the null event. Verify Eqs. (1.8) and (1.9), i.e.

(a) A ∪ Ø = {s: s ∈ A or s ∈ Ø} But, by definition, there are no s ∈ Ø. Thus,

(b) A ∩ Ø {s: s ∈ A and s ∈ Ø} But, since there are no s ∈ Ø, there cannot be an s such that s ∈ A and s ∈ Ø. Thus,

Note that Eq. (1.9) shows that Ø is mutually exclusive with every other event and including with itself.

1.18. Show that the null (or empty) set Ø is a subset of every set A.

From the definition of intersection, it follows that

for any pair of events, whether they are mutually exclusive or not. If A and B are mutually exclusive events, that is, A ∩ B = Ø, then by Eq. (1.70) we obtain

Therefore, for any event A,

that is, Ø is a subset of every set A.

1.19. Show that A and B are disjoint if and only if A\B = A.

First, if A \ B = A ∩ = A, then

and A and B are disjoint. Next, if A and B are disjoint, then A ∩ B = Ø, and A \ B = A ∩ = A. Thus, A and B are disjoint if and only if A\B = A.

1.20. Show that there is a distribution law also for difference; that is,

By Eq. (1.8) and applying commutative and associated laws, we have

Next,

Thus, we have

1.21. Verify Eqs. (1.24) and (1.25).

(a) Suppose first that .

That is, if s is not contained in any of the events Ai, i = 1, 2, …, n, then s is contained in for all i = 1, 2, …, n. Thus

Next, we assume that

Then s is contained in for all i = 1, 2, …, n, which means that s is not contained in Ai for any i = 1, 2, …, n, implying that

Thus,

This proves Eq. (1.24).

(b) Using Eqs. (1.24) and (1.3), we have

Taking complements of both sides of the above yields

which is Eq. (1.25).

Probability Space

1.22. Consider a probability space (S, F, P). Show that if A and B are in an event space (σ-field) F, so are A ∩ B, A\B, and A Δ B.

By condition (ii) of F, Eq. (1.27), if A, B ∈ F, then , ∈ F. Now by De Morgan’s law (1.21), we have

Similarly, we see that

Now by Eq. (1.10), we have

Finally, by Eq. (1.13), and Eq. (1.28), we see that

1.23. Consider the experiment of Example 1.7 of rolling a die. Show that {S, Ø, A, B} are event spaces but {S, Ø, A} and {S, Ø, A, B, C} are not event spaces.

Let F = {S, Ø, A, B}. Then we see that

and

Thus, we conclude that {S, Ø, A, B} is an event space (σ-field). Next, let F = {S, Ø, A}. Now . Thus {S, Ø, A} is not an event space. Finally, let F = {S, Ø, A, B, C}, but . Hence, {S, Ø, A, B, C} is not an event space.

1.24. Using the axioms of probability, prove Eq. (1.39).

We have

Thus, by axioms 2 and 3, it follows that

from which we obtain

1.25. Verify Eq. (1.40).

From Eq. (1.39), we have

Let A = Ø. Then, by Eq. (1.2), = Ø = S, and by axiom 2 we obtain

1.26. Verify Eq. (1.41).

Let A ⊂ B. Then from the Venn diagram shown in Fig. 1-7, we see that

Fig. 1-7

Hence, from axiom 3,

However, by axiom 1, . Thus, we conclude that

1.27. Verify Eq. (1.43).

From the Venn diagram of Fig. 1-8, each of the sets A ∈ B and B can be represented, respectively, as a union of mutually exclusive sets as follows:

Fig. 1-8

Thus, by axiom 3,

and

From Eq. (1.76), we have

Substituting Eq. (1.77) into Eq. (1.75), we obtain

1.28. Let P(A) = 0.9 and P(B) = 0.8. Show that P(A ∩ B) ≥ 0.7.

From Eq. (1.43), we have

By Eq. (1.47), 0 ≤ P(A ∪ B) ≤ 1. Hence,

Substituting the given values of P(A) and P(B) in Eq. (1.78), we get

Equation (1.77) is known as Bonferroni’s inequality.

1.29. Show that

From the Venn diagram of Fig. 1-9, we see that

Fig. 1-9

Thus, by axiom 3, we have

1.30. Given that P(A) = 0.9, P(B) = 0.8, and P(A ∩ B) = 0.75, find (a) P(A ∪ B); (b) P(A ∩ ); and (c) P( ∩ ).

(a) By Eq. (1.43), we have

(b) By Eq. (1.79) (Prob. 1.29), we have

(c) By De Morgan’s law, Eq. (1.20), and Eq. (1.39) and using the result from part (a), we get

1.31. For any three events A1, A2, and A3, show that

Let B = A2 ∪ A3. By Eq. (1.43), we have

Using distributive law (1.18), we have

Applying Eq. (1.43) to the above event, we obtain

Applying Eq. (1.43) to the set B = A2 ∪ A3, we have

Substituting Eqs. (1.84) and (1.83) into Eq. (1.82), we get

1.32. Prove that

which is known as Boole’s inequality.

We will prove Eq. (1.85) by induction. Suppose Eq. (1.85) is true for n = k.

Then

Thus, Eq. (1.85) is also true for n = k + 1. By Eq. (1.48), Eq. (1.85) is true for n = 2. Thus, Eq. (1.85) is true or n ≥ 2.

1.33. Verify Eq. (1.46).

Again we prove it by induction. Suppose Eq. (1.46) is true for n = k.

Then

Using the distributive law (1.22), we have

since Ai ∩ Aj = Ø for i ≠ j. Thus, by axiom 3, we have

which indicates that Eq. (1.46) is also true for n = k + 1. By axiom 3, Eq. (1.46) is true for n = 2. Thus, it is true for n ≥ 2.

1.34. A sequence of events {An, n ≥ 1} is said to be an increasing sequence if [Fig. 1-10(a)]

Fig. 1-10

whereas it is said to be a decreasing sequence if [Fig. 1-10(b)]

If {An, n ≥ 1} is an increasing sequence of events, we define a new event A∞ by

Similarly, if {An, n ≥ 1} is a decreasing sequence of events, we define a new event A∞ by

Show that if {An, n ≥ 1 } is either an increasing or a decreasing sequence of events, then

which is known as the continuity theorem of probability.

If {An, n ≥ 1} is an increasing sequence of events, then by definition

Now, we define the events Bn, n ≥ 1, by

Thus, Bn consists of those elements in An that are not in any of the earlier Ak, k < n. From the Venn diagram shown in Fig. 1-11, it is seen that Bn are mutually exclusive events such that

Fig. 1-11

Thus, using axiom 3′, we have

Next, if {An, n ≥ 1} is a decreasing sequence, then { n n < 1} is an increasing sequence. Hence, by Eq. (1.89), we have

From Eq. (1.25),

Thus,

Using Eq. (1.39), Eq. (1.91) reduces to

Combining Eqs. (190) and (1.92), we obtain Eq. (1.89).

Equally Likely Events

1.35. Consider a telegraph source generating two symbols, dots and dashes. We observed that the dots were twice as likely to occur as the dashes. Find the probabilities of the dots occurring and the dashes occurring.

From the observation, we have

Then, by Eq. (1.51),

Thus,

1.36. The sample space S of a random experiment is given by

with probabilities P(a) = 0.2, P(b) = 0.3, P(c) = 0.4, and P(d) = 0.1. Let A denote the event {a, b}, and B the event {b, c, d}. Determine the following probabilities: (a) P(A); (b) P(B); (c) P( ); (d) P(A ∪ B); and (e) P(A ∩ B).

Using Eq. (1.52), we obtain

1.37. An experiment consists of observing the sum of the dice when two fair dice are thrown (Prob. 1.5). Find (a) the probability that the sum is 7 and (b) the probability that the sum is greater than 10.

(a) Let ξij denote the elementary event (sampling point) consisting of the following outcome: ξij = (i, j), where i represents the number appearing on one die and j represents the number appearing on the other die. Since the dice are fair, all the outcomes are equally likely. So . Let A denote the event that the sum is 7. Since the events ξij are mutually exclusive and from Fig. 1-3 (Prob. 1.5), we have

(b) Let B denote the event that the sum is greater than 10. Then from Fig. 1- 3, we obtain

1.38. There are n persons in a room.

(a) What is the probability that at least two persons have the same birthday? (b) Calculate this probability for n = 50. (c) How large need n be for this probability to be greater than 0.5?

(a) As each person can have his or her birthday on any one of 365 days (ignoring the possibility of February 29), there are a total of (365)n possible outcomes. Let A be the event that no two persons have the same birthday. Then the number of outcomes belonging to A is

Assuming that each outcome is equally likely, then by Eq. (1.54),

Let B be the event that at least two persons have the same birthday. Then B = and by Eq.(1.39), P(B) = 1 – P(A).

(b) Substituting n = 50 in Eq. (1.93), we have

(c) From Eq. (1.93), when n = 23, we have

That is, if there are 23 persons in a room, the probability that at least two of them have the same birthday exceeds 0.5.

1.39. A committee of 5 persons is to be selected randomly from a group of 5 men and 10 women. (a) Find the probability that the committee consists of 2 men and 3 women. (b) Find the probability that the committee consists of all women.

(a) The number of total outcomes is given by

It is assumed that “random selection” means that each of the outcomes is equally likely. Let A be the event that the committee consists of 2 men and 3 women. Then the number of outcomes belonging to A is given by

Thus, by Eq. (1.54),

(b) Let B be the event that the committee consists of all women. Then the number of outcomes belonging to B is

Thus, by Eq. (1.54),

1.40. Consider the switching network shown in Fig. 1-12. It is equally likely that a switch will or will not work. Find the probability that a closed path will exist between terminals a and b.

Fig. 1-12

Consider a sample space S of which a typical outcome is (1, 0, 0, 1), indicating that switches 1 and 4 are closed and switches 2 and 3 are open. The sample space contains 24= 16 points, and by assumption, they are equally likely (Fig. 1-13).

Fig. 1-13

Let Ai, i = 1, 2, 3, 4 be the event that the switch si is closed. Let A be the event that there exists a closed path between a and b. Then

Applying Eq. (1.45), we have

Now, for example, the event A2 ∩ A3 contains all elementary events with a 1 in the second and third places. Thus, from Fig. 1-13, we see that

Thus,

1.41. Consider the experiment of tossing a fair coin repeatedly and counting the number of tosses required until the first head appears. (a) Find the sample space of the experiment. (b) Find the probability that the first head appears on the kth toss. (c) Verify that P(S) = 1.

(a) The sample space of this experiment is

where ek is the elementary event that the first head appears on the kth toss.

(b) Since a fair coin is tossed, we assume that a head and a tail are equally likely to appear. Then . Let

Since there are 2k equally likely ways of tossing a fair coin k times, only one of which consists of (k − 1) tails following a head we observe that

(c) Using the power series summation formula, we have

1.42. Consider the experiment of Prob. 1.41. (a) Find the probability that the first head appears on an even-numbered

toss. (b) Find the probability that the first head appears on an odd-numbered toss.

(a) Let A be the event “the first head appears on an even-numbered toss.” Then, by Eq. (1.52) and using Eq. (1.94) of Prob. 1.41, we have

(b) Let B be the event “the first head appears on an odd-numbered toss.” Then it is obvious that B = . Then, by Eq. (1.39), we get

As a check, notice that

Conditional Probability

1.43. Show that P(A | B) defined by Eq. (1.55) satisfies the three axioms of a probability, that is,

(a) From definition (1.55),

By axiom 1, P(A ∩ B) ≥ 0. Thus,

(b) By Eq. (1.5), S ∩ B = B. Then

(c) By definition (1.55),

Now by Eqs. (1.14) and (1.17), we have

and A1 ∩ A2 = Ø implies that (A1 ∩ B) ∩ (A2 ∩ B) = Ø. Thus, by axiom 3 we get

1.44. Find P(A | B) if (a) A ∩ B = Ø, (b) A ⊂ B, and (c) B ⊂ A.

(a) If A ∩ B = Ø, then P(A ∩ B) = P(Ø) = 0. Thus,

(b) If A ⊂ B, then A ∩ B = A and

(c) If B ⊂ A, then A ∩ B = B and

1.45. Show that if P(A | B) > P(A), then P(B | A) > P(B).

1.46. Show that

By Eqs. (1.6) and (1.7), we have

Then by Eqs. (1.97) and (1.98), we get

Thus we obtain

note that Eq. (1.99) is similar to property 1 (Eq. (1.39)).

1.47. Show that

Using a Venn diagram, we see that

and

By Eq. (1.98) we have

Again, using a Venn diagram, we see that

and

By Eq. (1.98) we have

Thus,

Substituting Eq. (1.101) into Eq. (1.100), we obtain

Note that Eq. (1.100) is similar to property 5 (Eq. (1.43)).

1.48. Consider the experiment of throwing the two fair dice of Prob. 1.37 behind you; you are then informed that the sum is not greater than 3. (a) Find the probability of the event that two faces are the same without the

information given. (b) Find the probability of the same event with the information given.

(a) Let A be the event that two faces are the same. Then from Fig. 1-3 (Prob. 1.5) and by Eq. (1.54), we have

and

(b) Let B be the event that the sum is not greater than 3. Again from Fig. 1- 3, we see that

and

Now A ∩ B is the event that two faces are the same and also that their sum is not greater than 3. Thus,

Then by definition (1.55), we obtain

Note that the probability of the event that two faces are the same doubled from to with the information given.

Alternative Solution:

There are 3 elements in B, and 1 of them belongs to A. Thus, the probability of the same event with the information given is .

1.49. Two manufacturing plants produce similar parts. Plant 1 produces 1,000 parts, 100 of which are defective. Plant 2 produces 2,000 parts, 150 of which

are defective. A part is selected at random and found to be defective. What is the probability that it came from plant 1?

Let B be the event that “the part selected is defective,” and let A be the event that “the part selected came from plant 1.” Then A ∩ B is the event that the item selected is defective and came from plant 1. Since a part is selected at random, we assume equally likely events, and using Eq. (1.54), we have

Similarly, since there are 3000 parts and 250 of them are defective, we have

By Eq. (1.55), the probability that the part came from plant 1 is

Alternative Solution: There are 250 defective parts, and 100 of these are from plant 1. Thus, the probability that the defective part came from plant 1 is .

1.50. Mary has two children. One child is a boy. What is the probability that the other child is a girl?

Let S be the sample space of all possible events S = {B B, B G, G B, G G} where B B denotes the events the first child is a boy and the second is also a boy, and BG denotes the event the first child is a boy and the second is a girl, and so on.

Now we have

Let B be the event that there is at least one boy; B = {B B, B G, G B}, and A be the event that there is at least one girl; A = {B G, G B, G G}

Then

Now, by Eq. (1.55) we have

Note that one would intuitively think the answer is 1/2 because the second event looks independent of the first. This problem illustrates that the initial intuition can be misleading.

1.51. A lot of 100 semiconductor chips contains 20 that are defective. Two chips are selected at random, without replacement, from the lot. (a) What is the probability that the first one selected is defective?

(b) What is the probability that the second one selected is defective given that the first one was defective?

(c) What is the probability that both are defective?

(a) Let A denote the event that the first one selected is defective. Then, by Eq. (1.54),

(b) Let B denote the event that the second one selected is defective. After the first one selected is defective, there are 99 chips left in the lot, with 19 chips that are defective. Thus, the probability that the second one selected is defective given that the first one was defective is

(c) By Eq. (1.57), the probability that both are defective is

1.52. A number is selected at random from {1, 2, …, 100). Given that the number selected is divisible by 2, find the probability that it is divisible by 3 or 5.

Let

Then the desired probability is

Now

and

Thus,

1.53. Show that

By definition Eq. (1.55) we have

since P(B ∩ A) = P(A ∩ B) and P(C ∩ A ∩ B) = P(A ∩ B ∩ C).

1.54. Let A1, A2, …, An be events in a sample space S. Show that

We prove Eq. (1.103) by induction. Suppose Eq. (1.103) is true for n = k:

Multiplying both sides by P(Ak+1 | A1 ∩ A2 ∩ … ∩ Ak), we have

and

Thus, Eq. (1.103) is also true for n = k + 1. By Eq. (1.57), Eq. (1.103) is true for n = 2. Thus, Eq. (1.103) is true for n ≤ 2.

1.55. Two cards are drawn at random from a deck. Find the probability that both are aces.

Let A be the event that the first card is an ace, and let B be the event that the second card is an ace. The desired probability is P(B ∩ A). Since a card is drawn at random, . Now if the first card is an ace, then there will be 3 aces left in the deck of 51 cards. Thus, . By Eq. (1.57),

Check:

By counting technique, we have

1.56. There are two identical decks of cards, each possessing a distinct symbol so that the cards from each deck can be identified. One deck of cards is laid out in a fixed order, and the other deck is shuffled and the cards laid out one by one on top of the fixed deck. Whenever two cards with the same symbol occur in the same position, we say that a match has occurred. Let the number of cards in the deck be 10. Find the probability of getting a match at the first four positions.

Let Ai, i = 1, 2, 3, 4, be the events that a match occurs at the ith position. The required probability is

By Eq. (1.103),

There are 10 cards that can go into position 1, only one of which matches. Thus, is the conditional probability of a match at position 2 given a match at position 1. Now there are 9 cards left to go into

position 2, only one of which matches. Thus, . In a similar fashion, we obtain and . Thus,

Total Probability

1.57. Verify Eq. (1.60).

Since B ∩ S = B [and using Eq. (1.59)], we have

Now the events B ∩ Ai, i = 1, 2, …, n, are mutually exclusive, as seen from the Venn diagram of Fig. 1-14. Then by axiom 3 of probability and Eq. (1.57), we obtain

Fig. 1-14

1.58. Show that for any events A and B in S,

From Eq. (1.78) (Prob. 1.29), we have

Using Eq. (1.55), we obtain

Note that Eq. (1.105) is the special case of Eq. (1.60).

1.59. Suppose that a laboratory test to detect a certain disease has the following statistics. Let

It is known that

and 0.1 percent of the population actually has the disease. What is the probability that a person has the disease given that the test result is positive?

From the given statistics, we have

The desired probability is P(A | B). Thus, using Eqs. (1.58) and (1.105), we obtain

Note that in only 16.5 percent of the cases where the tests are positive will the person actually have the disease even though the test is 99 percent effective in detecting the disease when it is, in fact, present.

1.60. A company producing electric relays has three manufacturing plants producing 50, 30, and 20 percent, respectively, of its product. Suppose that the probabilities that a relay manufactured by these plants is defective are 0.02, 0.05, and 0.01, respectively.

(a) If a relay is selected at random from the output of the company, what is the probability that it is defective?

(b) If a relay selected at random is found to be defective, what is the probability that it was manufactured by plant 2?

(a) Let B be the event that the relay is defective, and let Ai be the event that the relay is manufactured by plant i (i = 1, 2, 3). The desired probability is P(B). Using Eq. (1.60), we have

(b) The desired probability is P(A2 | B). Using Eq. (1.58) and the result from part (a), we obtain

1.61. Two numbers are chosen at random from among the numbers 1 to 10 without replacement. Find the probability that the second number chosen is 5.

Let At, i = 1, 2, …, 10 denote the event that the first number chosen is i. Let B be the event that the second number chosen is 5. Then by Eq. (1.60),

Now is the probability that the second number chosen is 5, given that the first is i. If i = 5, then P(B | Ai) = 0. If i ≠ 5, then

. Hence,

1.62. Consider the binary communication channel shown in Fig. 1-15. The channel input symbol X may assume the state 0 or the state 1, and, similarly, the channel output symbol Y may assume either the state 0 or the state 1. Because of the channel noise, an input 0 may convert to an output 1 and vice versa. The channel is characterized by the channel transition probabilities p0, q0, p1, and q1, defined by

Fig. 1-15

where x0 and x1 denote the events (X = 0) and (X = 1), respectively, and y0 and y1 denote the events (Y = 0) and (Y = 1), respectively. Note that p0 + q0 = 1 = p1 + q1. Let P(x0) = 0.5, p0 = 0.1, and p1 = 0.2. (a) Find P(y0) and P(y1). (b) If a 0 was observed at the output, what is the probability that a 0 was the

input state? (c) If a 1 was observed at the output, what is the probability that a 1 was the

input state?

(d) Calculate the probability of error Pe.

(a) We note that

Using Eq. (1.60), we obtain

(b) Using Bayes’ rule (1.58), we have

(c) Similarly,

(d) The probability of error is

Independent Events

1.63. Let A and B be events in an event space F. Show that if A and B are independent, then so are (a) A and , (b) A– and B, and (c) and .

(a) From Eq. (1.79) (Prob. 1.29), we have

Since A and B are independent, using Eqs. (1.62) and (1.39), we obtain

Thus, by definition (1.62), A and are independent.

(b) Interchanging A and B in Eq. (1.106), we obtain

which indicates that and B are independent.

(c) We have

Hence, and are independent.

1.64. Let A and B be events defined in an event space F. Show that if both P(A) and P(B) are nonzero, then events A and B cannot be both mutually exclusive and independent.

Let A and B be mutually exclusive events and P(A) ≠ 0, P(B) ≠ 0. Then P(A ∩ B) = P(Ø) = 0 but P(A)P(B) ≠ 0. Since

A and B cannot be independent.

1.65. Show that if three events A, B, and C are independent, then A and (B ∪ C) are independent.

We have

Thus, A and (B ∪ C) are independent.

1.66. Consider the experiment of throwing two fair dice (Prob. 1.37). Let A be the event that the sum of the dice is 7, B be the event that the sum of the dice is 6, and C be the event that the first die is 4. Show that events A and C are independent, but events B and C are not independent.

From Fig. 1-3 (Prob. 1.5), we see that

and

Now

and

Thus, events A and C are independent. But

Thus, events B and C are not independent.

1.67. In the experiment of throwing two fair dice, let A be the event that the first die is odd, B be the event that the second die is odd, and C be the event that the sum is odd. Show that events A, B, and C are pairwise independent, but A, B, and C are not independent.

From Fig. 1-3 (Prob. 1.5), we see that

Thus

which indicates that A, B, and C are pairwise independent. However, since the sum of two odd numbers is even, {A ∩ B ∩ C) = Ø and

which shows that A, B, and C are not independent.

1.68. A system consisting of n separate components is said to be a series system if it functions when all n components function (Fig. 1-16). Assume that the components fail independently and that the probability of failure of component i is pi, i = 1, 2, …, n. Find the probability that the system functions.

Fig. 1-16. Series system.

Let Ai be the event that component si functions. Then

Let A be the event that the system functions. Then, since Ai’s are independent, we obtain

1.69. A system consisting of n separate components is said to be a parallel system if it functions when at least one of the components functions (Fig. 1-17). Assume that the components fail independently and that the probability of failure of component i is pi, i = 1, 2, …, n. Find the probability that the system functions.

Fig. 1-17 Parallel system.

Let A be the event that component si functions. Then

Let A be the event that the system functions. Then, since are independent, we obtain

1.70. Using Eqs. (1.107) and (1.108), redo Prob. 1.40.

From Prob. 1.40, , i = 1, 2, 3, 4, where pi is the probability of failure of switch si. Let A be the event that there exists a closed path between a and b. Using Eq. (1.108), the probability of failure for the parallel combination of switches 3 and 4 is

Using Eq. (1.107), the probability of failure for the combination of switches 2, 3, and 4 is

Again, using Eq. (1.108), we obtain

1.71. A Bernoulli experiment is a random experiment, the outcome of which can be classified in but one of two mutually exclusive and exhaustive ways, say success or failure. A sequence of Bernoulli trials occurs when a Bernoulli experiment is performed several independent times so that the probability of success, say p, remains the same from trial to trial. Now an infinite sequence of Bernoulli trials is performed. Find the probability that (a) at least 1 success occurs in the first n trials; (b) exactly k successes occur in the first n trials; (c) all trials result in successes.

(a) In order to find the probability of at least 1 success in the first n trials, it is easier to first compute the probability of the complementary event, that of no successes in the first n trials. Let Ai denote the event of a failure on the ith trial. Then the probability of no successes is, by independence,

Hence, the probability that at least 1 success occurs in the first n trials is 1 – (1 – p)n.

(b) In any particular sequence of the first n outcomes, if k successes occur, where k = 0, 1, 2, …, n, then n – k failures occur. There are such sequences, and each one of these has probability pk(1 – p)n–k. Thus, the probability that exactly k successes occur in the first n trials is given by

.

(c) Since denotes the event of a success on the ith trial, the probability that all trials resulted in successes in the first n trials is, by independence,

Hence, using the continuity theorem of probability (1.89) (Prob. 1.34), the probability that all trials result in successes is given by

1.72. Let S be the sample space of an experiment and S = {A, B, C}, where P(A) = p, P(B) = q, and P(C) = r. The experiment is repeated infinitely, and it is assumed that the successive experiments are independent. Find the probability of the event that A occurs before B.

Suppose that A occurs for the first time at the nth trial of the experiment. If A is to have occurred before B, then C must have occurred on the first (n – 1) trials. Let D be the event that A occurs before B.

Then

where Dn is the event that C occurs on the first (n – 1) trials and A occurs on the nth trial. Since Dn’s are mutually exclusive, we have

Since the trials are independent, we have

Thus,

or

since p + q + r = 1.

1.73. In a gambling game, craps, a pair of dice is rolled and the outcome of the experiment is the sum of the dice. The player wins on the first roll if the sum is 7 or 11 and loses if the sum is 2, 3, or 12. If the sum is 4, 5, 6, 8, 9, or 10, that number is called the player’s “point.” Once the point is established, the rule is: If the player rolls a 7 before the point, the player loses; but if the point is rolled before a 7, the player wins. Compute the probability of winning in the game of craps.

Let A, B, and C be the events that the player wins, the player wins on the first roll, and the player gains point, respectively. Then P(A) = P(B) + P(C). Now from Fig. 1-3 (Prob. 1.5),

Let Ak be the event that point of k occurs before 7. Then

By Eq. (1.111) (Prob. 1.72),

Again from Fig. 1-3,

Now by Eq. (1.112),

Using these values, we obtain

SUPPLEMENTARY PROBLEMS

1.74. Consider the experiment of selecting items from a group consisting of three items {a, b, c}.

(a) Find the sample space S1 of the experiment in which two items are selected without replacement.

(b) Find the sample space S2 of the experiment in which two items are selected with replacement.

1.75. Let A and B be arbitrary events. Then show that A ⊂ B if and only if A ∪ B = B.

1.76. Let A and B be events in the sample space S. Show that if A ⊂ B, then ⊂ .

1.77. Verify Eq. (1.19).

1.78. Show that

1.79. Let A and B be any two events in S. Express the following events in terms of A and B. (a) At least one of the events occurs. (b) Exactly one of two events occurs.

1.80. Show that A and B are disjoint if and only if

1.81. Let A, B, and C be any three events in S. Express the following events in terms of these events. (a) Either B or C occurs, but not A. (b) Exactly one of the events occurs. (c) Exactly two of the events occur.

1.82. Show that F = {S, Ø} is an event space.

1.83. Let S = {1, 2, 3, 4} and F1 = {S, Ø, {1, 3}, {2, 4}}, F2 = {S, Ø, {1, 3}}. Show that F1 is an event space, and F2 is not an event space.

1.84. In an experiment one card is selected from an ordinary 52-card deck. Define the events: A = select a King, B = select a Jack or a Queen, C = select a

Heart. Find P(A), P(B), and P(C).

1.85. A random experiment has sample space S = {a, b, c}. Suppose that P({a, c}) = 0.75 and P({b, c)} = 0.6. Find the probabilities of the elementary events.

1.86. Show that

1.87. Let A, B, and C be three events in S. If , and P(B ∩ C) = 0,

find P(A ∪ B ∪ C).

1.88. Verify Eq. (1.45).

1.89. Show that

1.90. In an experiment consisting of 10 throws of a pair of fair dice, find the probability of the event that at least one double 6 occurs.

1.91. Show that if P(A) > P(B), then P(A | B) > P(B | A).

1.92. Show that

1.93. Show that

1.94. An urn contains 8 white balls and 4 red balls. The experiment consists of drawing 2 balls from the urn without replacement. Find the probability that both balls drawn are white.

1.95. There are 100 patients in a hospital with a certain disease. Of these, 10 are selected to undergo a drug treatment that increases the percentage cured rate from 50 percent to 75 percent. What is the probability that the patient received a drug treatment if the patient is known to be cured?

1.96. Two boys and two girls enter a music hall and take four seats at random in a row. What is the probability that the girls take the two end seats?

1.97. Let A and B be two independent events in S. It is known that P(A ∩ B) = 0.16 and P(A ∪ B) = 0.64. Find P(A) and P(B).

1.98. Consider the random experiment of Example 1.7 of rolling a die. Let A be the event that the outcome is an odd number and B the event that the outcome is less than 3. Show that events A and B are independent.

1.99. The relay network shown in Fig. 1-18 operates if and only if there is a closed path of relays from left to right. Assume that relays fail independently and that the probability of failure of each relay is as shown. What is the probability that the relay network operates?

Fig. 1-18

ANSWERS TO SUPPLEMENTARY PROBLEMS

1.74. (a) S1 = {ab, ac, ba, bc, ca, cb} (b) S2 = {aa, ab, ac, ba, bb, bc, ca, cb, cc}

1.75. Hint: Draw a Venn diagram.

1.76. Hint: Draw a Venn diagram.

1.77. Hint: Draw a Venn diagram.

1.78. Hint: Use Eqs. (1.10) and (1.17).

1.79. (a) A ∪ B; (b) A Δ B

1.80. Hint: Follow Prob. 1.19.

1.81.

1.82. Hint: Follow Prob. 1.23.

1.83. Hint: Follow Prob. 1.23.

1.84. P(A) = 1/13, P(B) = 2/13, P(C) = 13/52

1.85. P(a) = 0.4, P(b) = 0.25, P(c) = 0.35

1.86. Hint: (a) Use Eqs. (1.21) and (1.39). (b) Use Eqs. (1.43), (1.39), and (1.42). (c) Use a Venn diagram.

1.87.

1.88. Hint: Prove by induction.

1.89. Hint: Use induction to generalize Bonferroni’s inequality (1.77) (Prob. 1.28).

1.90. 0.246

1.91. Hint: Use Eqs. (1.55) and (1.56).

1.92. Hint: Use definition Eq.(1.55).

1.93. Hint: Follow Prob. 1.53.

1.94. 0.424

1.95. 0.143

1.96.

1.97. P(A) = P(B) = 0.4

1.98. Hint: Show that P(A ∩ B) = P(A)P(B) = 1/6.

1.99. 0.865

CHAPTER 2

Random Variables

2.1 Introduction

In this chapter, the concept of a random variable is introduced. The main purpose of using a random variable is so that we can define certain probability functions that make it both convenient and easy to compute the probabilities of various events.

2.2 Random Variables

A. Definitions:

Consider a random experiment with sample space S. A random variable X(ζ) is a single-valued real function that assigns a real number called the value of X(ζ) to each sample point ζ of S. Often, we use a single letter X for this function in place of X(ζ) and use r.v. to denote the random variable.

Note that the terminology used here is traditional. Clearly a random variable is not a variable at all in the usual sense, and it is a function.

The sample space S is termed the domain of the r.v. X, and the collection of all numbers [values of X(ζ)] is termed the range of the r.v. X. Thus, the range of X is a certain subset of the set of all real numbers (Fig. 2-1).

Fig. 2-1 Random variable X as a function.

Note that two or more different sample points might give the same value of X(ζ), but two different numbers in the range cannot be assigned to the same sample point.

EXAMPLE 2.1 In the experiment of tossing a coin once (Example 1.1), we might define the r.v. X as (Fig. 2-2)

Fig. 2-2 One random variable associated with coin tossing.

Note that we could also define another r.v., say Y or Z, with

B. Events Defined by Random Variables: If X is a r.v. and x is a fixed real number, we can define the event (X = x) as

Similarly, for fixed numbers x, x1, and x2, we can define the following events:

These events have probabilities that are denoted by

EXAMPLE 2.2 In the experiment of tossing a fair coin three times (Prob. 1.1), the sample space S1 consists of eight equally likely sample points S1 = {HHH, …, TTT}. If X is the r.v. giving the number of heads obtained, find (a) P(X = 2); (b) P(X < 2).

(a) Let A ⊂ S1 be the event defined by X = 2. Then, from Prob. 1.1, we have

Since the sample points are equally likely, we have

(b) Let B ⊂ S1 be the event defined by X < 2. Then

2.3 Distribution Functions

A. Definition:

The distribution function [or cumulative distribution function (cdf)] of X is the function defined by

Most of the information about a random experiment described by the r.v. X is determined by the behavior of FX(x).

B. Properties of FX(x):

Several properties of FX(x) follow directly from its definition (2.4).

Property 1 follows because FX(x) is a probability. Property 2 shows that FX(x) is a nondecreasing function (Prob. 2.5). Properties 3 and 4 follow from Eqs. (1.22) and (1.26):

Property 5 indicates that FX(x) is continuous on the right. This is the consequence of the definition (2.4).

EXAMPLE 2.3 Consider the r.v. X defined in Example 2.2. Find and sketch the cdf FX(x) of X.

Table 2-1 gives FX(x) = P(X ≤ x) for x = −1, 0, 1, 2, 3, 4. Since the value of X must be an integer, the value of FX(x) for noninteger values of x must be the same as the value of FX(x) for the nearest smaller integer value of x. The FX(x) is sketched in Fig. 2-3. Note that FX(x) has jumps at x = 0, 1, 2, 3, and that at each jump the upper value is the correct value for FX(x).

TABLE 2-1

Fig. 2-3

C. Determination of Probabilities from the Distribution Function: From definition (2.4), we can compute other probabilities, such as P(a < X ≤ b), P(X > a), and P(X < b) (Prob. 2.6):

2.4 Discrete Random Variables and Probability Mass Functions

A. Definition:

Let X be a r.v. with cdf FX(x). If FX(x) changes values only in jumps (at most a countable number of them) and is constant between jumps—that is, FX(x) is a staircase function (see Fig. 2-3)—then X is called a discrete random variable. Alternatively, X is a discrete r.v. only if its range contains a finite or countably infinite number of points. The r.v. X in Example 2.3 is an example of a discrete r.v.

B. Probability Mass Functions: Suppose that the jumps in FX(x) of a discrete r.v. X occur at the points x1, x2, …, where the sequence may be either finite or countably infinite, and we assume xi < xj if i < j. Then

Let

The function pX(x) is called the probability mass function (pmf) of the discrete r.v. X.

Properties of pX(x):

1.

2.

3.

The cdf FX(x) of a discrete r.v. X can be obtained by

2.5 Continuous Random Variables and Probability Density Functions

A. Definition:

Let X be a r.v. with cdf FX(x). If FX(x) is continuous and also has a derivative dFX(x)/dx which exists everywhere except at possibly a finite number of points and is piecewise continuous, then X is called a continuous random variable. Alternatively, X is a continuous r.v. only if its range contains an interval (either finite or infinite) of real numbers. Thus, if X is a continuous r.v., then (Prob. 2.18)

Note that this is an example of an event with probability 0 that is not necessarily the impossible event ∅.

In most applications, the r.v. is either discrete or continuous. But if the cdf FX(x) of a r.v. X possesses features of both discrete and continuous r.v.’s, then the r.v. X is called the mixed r.v. (Prob. 2.10).

B. Probability Density Functions: Let

The function fX(x) is called the probability density function (pdf) of the continuous r.v. X.

Properties of ƒX(x):

1.

2.

3. ƒX(x) is piecewise continuous. 4.

The cdf FX(x) of a continuous r.v. X can be obtained by

By Eq. (2.19), if X is a continuous r.v., then

2.6 Mean and Variance

A. Mean:

The mean (or expected value) of a r.v. X, denoted by μX or E(X), is defined by

B. Moment: The nth moment of a r.v. X is defined by

Note that the mean of X is the first moment of X.

C. Variance: The variance of a r.v. X, denoted by or Var(X), is defined by

Thus,

Note from definition (2.28) that

The standard deviation of a r.v. X, denoted by σX, is the positive square root of Var(X).

Expanding the right-hand side of Eq. (2.28), we can obtain the following relation:

which is a useful formula for determining the variance.

2.7 Some Special Distributions

In this section we present some important special distributions.

A. Bernoulli Distribution: A r.v. X is called a Bernoulli r.v. with parameter p if its pmf is given by

where 0 ≤ p ≤ 1. By Eq. (2.18), the cdf FX(x) of the Bernoulli r.v. X is given by

Fig. 2-4 illustrates a Bernoulli distribution.

Fig. 2-4 Bernoulli distribution.

The mean and variance of the Bernoulli r.v. X are

A Bernoulli r.v. X is associated with some experiment where an outcome can be classified as either a “success” or a “failure,” and the probability of a success is p and the probability of a failure is 1 − p. Such experiments are often called Bernoulli trials (Prob. 1.71).

B. Binomial Distribution: A r.v. X is called a binomial r.v. with parameters (n, p) if its pmf is given by

where 0 ≤ p ≤ 1 and

which is known as the binomial coefficient. The corresponding cdf of X is

Fig. 2-5 illustrates the binomial distribution for n = 6 and p = 0.6.

Fig. 2-5 Binomial distribution with n = 6, p = 0.6.

The mean and variance of the binomial r.v. X are (Prob. 2.28)

A binomial r.v. X is associated with some experiments in which n independent Bernoulli trials are performed and X represents the number of successes that occur in the n trials. Note that a Bernoulli r.v. is just a binomial r.v. with parameters (1, p).

C. Geometric Distribution: A r.v. X is called a geometric r.v. with parameter p if its pmf is given by

where 0 ≤ p ≤ 1. The cdf FX(x) of the geometric r.v. X is given by

Fig. 2-6 illustrates the geometric distribution with p = 0.25.

Fig. 2-6 Geometric distribution with p = 0.25.

The mean and variance of the geometric r.v. X are (Probs. 2.29 and 4.55)

A geometric r.v. X is associated with some experiments in which a sequence of Bernoulli trials with probability p of success is obtained. The sequence is observed until the first success occurs. The r.v. X denotes the trial number on which the first success occurs. Memoryless property of the geometric distribution:

If X is a geometric r.v., then it has the following significant property (Probs. 2.17, 2.57).

Equation (2.44) indicates that suppose after i flips of a coin, no “head" has turned up yet, then the probability for no “head" to turn up for the next j flips of the coin is exactly the same as the probability for no “head" to turn up for the first i flips of the coin.

Equation (2.44) is known as the memoryless property. Note that memoryless property Eq. (2.44) is only valid when i, j are integers. The geometric distribution is the only discrete distribution that possesses this property.

D. Negative Binomial Distribution: A r.v. X is called a negative binomial r.v. with parameters p and k if its pmf is given by

where 0 ≤ p ≤ 1. Fig. 2-7 illustrates the negative binomial distribution for k = 2 and p = 0.25.

Fig. 2-7 Negative binomial distribution with k = 2 and p = 0.25.

The mean and variance of the negative binomial r.v. X are (Probs. 2.80, 4.56).

A negative binomial r.v. X is associated with sequence of independent Bernoulli trials with the probability of success p, and X represents the number of trials until the kth success is obtained. In the experiment of flipping a coin, if X = x, then it must be true that there were exactly k − 1 heads thrown in the first x − 1 flippings,

and a head must have been thrown on the xth flipping. There are

sequences of length x with these properties, and each of them is assigned the same probability of pk−1 (1 − p)x−k.

Note that when k = 1, X is a geometrical r.v. A negative binomial r.v. is sometimes called a Pascal r.v.

E. Poisson Distribution: A r.v. X is called a Poisson r.v. with parameter λ(> 0) if its pmf is given by

The corresponding cdf of X is

Fig. 2-8 illustrates the Poisson distribution for λ = 3.

Fig. 2-8 Poisson distribution with λ = 3.

The mean and variance of the Poisson r.v. X are (Prob. 2.31)

The Poisson r.v. has a tremendous range of applications in diverse areas because it may be used as an approximation for a binomial r.v. with parameters (n, p) when n is large and p is small enough so that np is of a moderate size (Prob. 2.43).

Some examples of Poisson r.v.’s include

1. The number of telephone calls arriving at a switching center during various intervals of time

2. The number of misprints on a page of a book 3. The number of customers entering a bank during various intervals of time

F. Discrete Uniform Distribution: A r.v. X is called a discrete uniform r.v. if its pmf is given by

The cdf FX(x) of the discrete uniform r.v. X is given by

where ⌊x⌋ denotes the integer less than or equal to x. Fig. 2-9 illustrates the discrete uniform distribution for n = 6.

Fig. 2-9 Discrete uniform distribution with n = 6.

The mean and variance of the discrete uniform r.v. X are (Prob. 2.32)

The discrete uniform r.v. X is associated with cases where all finite outcomes of an experiment are equally likely. If the sample space is a countably infinite set, such as the set of positive integers, then it is not possible to define a discrete uniform r.v. X. If the sample space is an uncountable set with finite length such as the interval (a, b), then a continuous uniform r.v. X will be utilized.

G. Continuous Uniform Distribution: A r.v. X is called a continuous uniform r.v. over (a, b) if its pdf is given by

The corresponding cdf of X is

Fig. 2-10 illustrates a continuous uniform distribution.

Fig. 2-10 Continuous uniform distribution over (a, b).

The mean and variance of the uniform r.v. X are (Prob. 2.34)

A uniform r.v. X is often used where we have no prior knowledge of the actual pdf and all continuous values in some range seem equally likely (Prob. 2.75).

H. Exponential Distribution: A r.v. X is called an exponential r.v. with parameter λ(> 0) if its pdf is given by

which is sketched in Fig. 2-11(a). The corresponding cdf of X is

Fig. 2-11 Exponential distribution.

which is sketched in Fig. 2-11(b). The mean and variance of the exponential r.v. X are (Prob. 2.35)

Memoryless property of the exponential distribution: If X is an exponential r.v., then it has the following interesting memoryless

property (cf. Eq. (2.44)) (Prob. 2.58)

Equation (2.64) indicates that if X represents the lifetime of an item, then the item that has been in use for some time is as good as a new item with respect to the amount of time remaining until the item fails. The exponential distribution is the only continuous distribution that possesses this property. This memoryless property is a fundamental property of the exponential distribution and is basic for the theory of Markov processes (see Sec. 5.5).

I. Gamma Distribution: A r.v. X is called a gamma r.v. with parameter (α, λ) (α > 0 and λ > 0) if its pdf is given by

where Γ(α) is the gamma function defined by

and it satisfies the following recursive formula (Prob. 2.26)

The pdf ƒX(x) with (α, λ) = (1, 1), (2, 1), and (5, 2) are plotted in Fig. 2-12.

Fig. 2-12 Gamma distributions for selected values of α and λ

The mean and variance of the gamma r.v. are (Prob. 4.65)

Note that when α = 1, the gamma r.v. becomes an exponential r.v. with parameter λ [Eq. (2.60)], and when α = n/2, λ = 1/2, the gamma pdf becomes

which is the chi-squared r.v. pdf with n degrees of freedom (Prob. 4.40). When α = n (integer), the gamma distribution is sometimes known as the Erlang distribution.

J. Normal (or Gaussian) Distribution: A r.v. X is called a normal (or Gaussian) r.v. if its pdf is given by

The corresponding cdf of X is

This integral cannot be evaluated in a closed form and must be evaluated numerically. It is convenient to use the function Φ(z), defined as

to help us to evaluate the value of FX(x). Then Eq. (2.72) can be written as

Note that

The function Φ(z) is tabulated in Table A (Appendix A). Fig. 2-13 illustrates a normal distribution.

Fig. 2-13 Normal distribution.

The mean and variance of the normal r.v. X are (Prob. 2.36)

We shall use the notation N(μ; σ2) to denote that X is normal with mean μ and variance σ2. A normal r.v. Z with zero mean and unit variance—that is, Z = N(0; 1) —is called a standard normal r.v. Note that the cdf of the standard normal r.v. is given by Eq. (2.73). The normal r.v. is probably the most important type of continuous r.v. It has played a significant role in the study of random phenomena in nature. Many naturally occurring random phenomena are approximately normal. Another reason for the importance of the normal r.v. is a remarkable theorem called the central limit theorem. This theorem states that the sum of a large number of independent r.v.’s, under certain conditions, can be approximated by a normal r.v. (see Sec. 4.8C).

2.8 Conditional Distributions

In Sec. 1.6 the conditional probability of an event A given event B is defined as

The conditional cdf FX(x | B) of a r.v. X given event B is defined by

The conditional cdf FX(x | B) has the same properties as FX(x). (See Prob. 1.43 and Sec. 2.3.) In particular,

If X is a discrete r.v., then the conditional pmf pX(xk | B) is defined by

If X is a continuous r.v., then the conditional pdf fX(x | B) is defined by

SOLVED PROBLEMS

Random Variables

2.1. Consider the experiment of throwing a fair die. Let X be the r.v. which assigns 1 if the number that appears is even and 0 if the number that appears is odd. (a) What is the range of X? (b) Find P(X = 1) and P(X = 0).

The sample space S on which X is defined consists of 6 points which are equally likely:

(a) The range of X is RX = {0, 1}. (b) (X = 1) = {2, 4, 6}. Thus, . Similarly, (X = 0) = {1, 3,

5}, and .

2.2. Consider the experiment of tossing a coin three times (Prob. 1.1). Let X be the r.v. giving the number of heads obtained. We assume that the tosses are independent and the probability of a head is p. (a) What is the range of X? (b) Find the probabilities P(X = 0), P(X = 1), P(X = 2), and P(X = 3).

The sample space S on which X is defined consists of eight sample points (Prob. 1.1):

(a) The range of X is RX = {0, 1, 2, 3}. (b) If P(H) = p, then P(T) = 1 − p. Since the tosses are independent, we

have

2.3. An information source generates symbols at random from a four-letter alphabet {a, b, c, d} with probabilities , and

. A coding scheme encodes these symbols into binary codes as follows:

Let X be the r.v. denoting the length of the code—that is, the number of binary symbols (bits). (a) What is the range of X? (b) Assuming that the generations of symbols are independent, find the

probabilities P(X = 1), P(X = 2), P(X = 3), and P(X > 3).

(a) The range of X is RX = {1, 2, 3}.

(b)

2.4. Consider the experiment of throwing a dart onto a circular plate with unit radius. Let X be the r.v. representing the distance of the point where the dart lands from the origin of the plate. Assume that the dart always lands on the plate and that the dart is equally likely to land anywhere on the plate. (a) What is the range of X? (b) Find (i) P(X < a) and (ii) P(a < X < b), where a < b ≤ 1.

(a) The range of X is RX = {x: 0 ≤ x < 1}. (b) (i) (X < a) denotes that the point is inside the circle of radius a. Since

the dart is equally likely to fall anywhere on the plate, we have (Fig. 2- 14)

Fig. 2-14

(ii) (a < X < b) denotes the event that the point is inside the annular ring with inner radius a and outer radius b. Thus, from Fig. 2-14, we have

Distribution Function

2.5. Verify Eq. (2.6). Let x1 < x2. Then (X ≤ x1) is a subset of (X ≤ x2); that is, (X ≤ x1) ⊂ (X ≤ x2). Then, by Eq. (1.41), we have

2.6. Verify (a) Eq. (2.10); (b) Eq. (2.11); (c) Eq. (2.12). (a) Since (X ≤ b) = (X ≤ a) ∪ (a < X ≤ b) and (X ≤ a) ∩ (a < X ≤ b) = ∅, we

have

(b) Since (X ≤ a) ∪ (X > a) = S and (X ≤ a) ∩ (X > a) = ∅, we have

(c) Now

2.7. Show that (a)

(b)

(c)

(a) Using Eqs. (1.37) and (2.10), we have

(b) We have

Again using Eq. (2.10), we obtain

(c) Similarly,

Using Eq. (2.83), we obtain

2.8. Let X be the r.v. defined in Prob. 2.3. (a) Sketch the cdf FX(x) of X and specify the type of X. (b) Find (i) P(X ≤ 1), (ii) P(1 < X ≤ 2), (iii) P(X > 1), and (iv) P(1 ≤ X ≤ 2).

(a) From the result of Prob. 2.3 and Eq. (2.18), we have which is sketched in Fig. 2-15. The r.v. X is a discrete r.v.

Fig. 2-15

(b) (i) We see that

(ii) By Eq. (2.10),

(iii) By Eq. (2.11),

(iv) By Eq. (2.83),

2.9. Sketch the cdf FX(x) of the r.v. X defined in Prob. 2.4 and specify the type of X.

From the result of Prob. 2.4, we have

which is sketched in Fig. 2-16. The r.v. X is a continuous r.v.

Fig. 2-16

2.10. Consider the function given by

(a) Sketch F(x) and show that F(x) has the properties of a cdf discussed in Sec. 2.3B.

(b) If X is the r.v. whose cdf is given by F(x), find (i) , (ii) , (iii) P(X = 0), and (iv) .

(c) Specify the type of X.

(a) The function F(x) is sketched in Fig. 2-17. From Fig. 2-17, we see that 0 ≤ F(x) ≤ 1 and F(x) is a nondecreasing function, F(−∞) = 0, F(∞) = 1,

, and F(x) is continuous on the right. Thus, F(x) satisfies all the properties [Eqs. (2.5) to (2.9)] required of a cdf.

Fig. 2-17

(b) (i) We have

(ii) By Eq. (2.10),

(iii) By Eq. (2.12),

(iv) By Eq. (2.83),

(c) The r.v. X is a mixed r.v.

2.11. Find the values of constants a and b such that

is a valid cdf.

To satisfy property 1 of FX(x) [0 ≤ FX(x) ≤ 1], we must have 0 ≤ a ≤ 1 and b > 0. Since b > 0, property 3 of FX(x) [FX(∞) = 1] is satisfied. It is seen that property 4 of FX(x) [FX (−∞) = 0] is also satisfied. For 0 ≤ a ≤ 1 and b > 0, F(x) is sketched in Fig. 2-18. From Fig. 2-18, we see that F(x) is a nondecreasing function and continuous on the right, and properties 2 and 5 of FX(x) are satisfied. Hence, we conclude that F(x) given is a valid cdf if 0 ≤ a ≤ 1 and b > 0. Note that if a = 0, then the r.v. X is a discrete r.v.; if a = 1, then X is a continuous r.v.; and if 0 < a < 1, then X is a mixed r.v.

Fig. 2-18

Discrete Random Variables and Pmf’s

2.12. Suppose a discrete r.v. X has the following pmfs:

(a) Find and sketch the cdf FX(x) of the r.v. X. (b) Find (i) P(X ≤ 1), (ii) P(1 < X ≤ 3), (iii) P(1 ≤ X ≤ 3).

(a) By Eq. (2.18), we obtain

which is sketched in Fig. 2-19.

Fig. 2-19

(b) (i) By Eq. (2.12), we see that

(ii) By Eq. (2.10),

(iii) By Eq. (2.83),

2.13. (a) Verify that the function p(x) defined by

is a pmf of a discrete r.v. X. (b) Find (i) P(X = 2), (ii) P(X ≤ 2), (iii) P(X ≥ 1).

(a) It is clear that 0 ≤ p(x) ≤ 1 and

Thus, p(x) satisfies all properties of the pmf [Eqs. (2.15) to (2.17)] of a discrete r.v. X.

(b) (i) By definition (2.14),

(ii) By Eq. (2.18),

(iii) By Eq. (1.39),

2.14. Consider the experiment of tossing an honest coin repeatedly (Prob. 1.41). Let the r.v. X denote the number of tosses required until the first head appears. (a) Find and sketch the pmf pX(x) and the cdf FX(x) of X. (b) Find (i) P(1 < X ≤ 4), (ii) P(X > 4).

(a) From the result of Prob. 1.41, the pmf of X is given by

Then by Eq. (2.18),

where |x| is the integer part of x or

These functions are sketched in Fig. 2-20.

Fig. 2-20

(b) (i) By Eq. (2.10),

(ii) By Eq. (1.39),

2.15 Let X be a binomial r.v. with parameters (n, p). (a) Show that pX(x) given by Eq. (2.36) satisfies Eq. (2.17). (b) Find P(X > 1) if n = 6 and p = 0.1.

(a) Recall that the binomial expansion formula is given by

Thus, by Eq. (2.36),

(b) Now

2.16. Let X be a geometric r.v. with parameter p. (a) Show that pX(x) given by Eq. (2.40) satisfies Eq. (2.17). (b) Find the cdf FX(x) of X.

(a) Recall that for a geometric series, the sum is given by

Thus,

(b) Using Eq. (2.87), we obtain

Thus,

and

Note that the r.v. X of Prob. 2.14 is the geometric r.v. with .

2.17. Verify the memoryless property (Eq. 2.44) for the geometric distribution.

From Eq. (2.41) and using Eq. (2.11), we obtain

Now, by definition of conditional probability, Eq. (1.55), we obtain

2.18. Let X be a negative binomial r.v. with parameters p and k. Show that pX(x) given by Eq. (2.45) satisfies Eq. (2.17) for k = 2.

From Eq. (2.45)

Let k = 2 and 1 − p = q. Then

Now let

Then

Subtracting the second series from the first series, we obtain

and we have

Thus,

2.19. Let X be a Poisson r.v. with parameter λ. (a) Show that pX(x) given by Eq. (2.48) satisfies Eq. (2.17). (b) Find P(X > 2) with λ = 4.

(a) By Eq. (2.48),

(b) With λ = 4, we have

and

Thus,

Continuous Random Variables and Pdf’s

2.20. Verify Eq. (2.19).

From Eqs. (1.41) and (2.10), we have

for any ε ≥ 0. As FX(x) is continuous, the right-hand side of the above expression approaches 0 as ε → 0.

Thus, P(X = x) = 0.

2.21. The pdf of a continuous r.v. X is given by

Find the corresponding cdf FX(x) and sketch fX(x) and FX(x).

By Eq. (2.24), the cdf of X is given by

The functions fX(x) and FX(x) are sketched in Fig. 2-21.

Fig. 2-21

2.22. Let X be a continuous r.v. X with pdf

where k is a constant. (a) Determine the value of k and sketch fX(x). (b) Find and sketch the corresponding cdf FX(x). (c) Find .

(a) By Eq. (2.21), we must have k > 0, and by Eq. (2.22),

Thus, k = 2 and

which is sketched in Fig. 2-22(a).

Fig. 2-22

(b) By Eq. (2.24), the cdf of X is given by

which is sketched in Fig. 2-22(b). (c) By Eq. (2.25),

2.23. Show that the pdf of a normal r.v. X given by Eq. (2.71) satisfies Eq. (2.22).

From Eq. (2.71),

Let . Then and

Let

Then

Letting x = r cos θ and y = r sin θ (that is, using polar coordinates), we have

Thus,

and

2.24. Consider a function

Find the value of a such that f(x) is a pdf of a continuous r.v. X.

If f(x) is a pdf of a continuous r.v. X, then by Eq. (2.22), we must have

Now by Eq. (2.52), the pdf of . Thus,

from which we obtain .

2.25. A r.v. X is called a Rayleigh r.v. if its pdf is given by

(a) Determine the corresponding cdf FX(x). (b) Sketch fX(x) and FX(x) for σ = 1.

(a) By Eq. (2.24), the cdf of X is

Let y = ξ2/(2σ2). Then dy = (1/σ2)ξdξ, and

(b) With σ = 1, we have

and

These functions are sketched in Fig. 2-23.

Fig. 2-23 Rayleigh distribution with σ = 1.

2.26. Consider a gamma r.v. X with parameter (α, λ) (α > 0 and λ > 0) defined by Eq. (2.65). (a) Show that the gamma function has the following properties: 1.

2.

3.

(b) Show that the pdf given by Eq. (2.65) satisfies Eq. (2.22).

(a) Integrating Eq. (2.66) by parts (u = xα −1, dv = e−x dx), we obtain

Replacing α by α + 1 in Eq. (2.100), we get Eq. (2.97). Next, by applying Eq. (2.97) repeatedly using an integral value of α,

say α = k, we obtain

Since

it follows that Γ(k + 1) = k!. Finally, by Eq. (2.66),

Let y = x1/2. Then , and

in view of Eq. (2.94). (b) Now

Let y = λx. Then dy = λdx and

Mean and Variance

2.27. Consider a discrete r.v. X whose pmf is given by

(a) Plot pX(x) and find the mean and variance of X. (b) Repeat (a) if the pmf is given by

(a) The pmf pX(x) is plotted in Fig. 2-24(a). By Eq. (2.26), the mean of X is

Fig. 2-24

By Eq. (2.29), the variance of X is

(b) The pmf pX(x) is plotted in Fig. 2-24(b). Again by Eqs. (2.26) and (2.29), we obtain

Note that the variance of X is a measure of the spread of a distribution about its mean.

2.28. Let a r.v. X denote the outcome of throwing a fair die. Find the mean and variance of X.

Since the die is fair, the pmf of X is

By Eqs. (2.26) and (2.29), the mean and variance of X are

Alternatively, the variance of X can be found as follows:

Hence, by Eq. (2.31),

2.29. Find the mean and variance of the geometric r.v. X defined by Eq. (2.40).

To find the mean and variance of a geometric r.v. X, we need the following results about the sum of a geometric series and its first and second derivatives. Let

Then

By Eqs. (2.26) and (2.40), and letting q = 1 − p, the mean of X is given by

where Eq. (2.102) is used with a = p and r = q.

To find the variance of X, we first find E[X(X − 1)]. Now,

where Eq. (2.103) is used with a = pq and r = q.

Since E[X(X − 1)] = E(X2 − X) = E(X2) − E(X), we have

Then by Eq. (2.31), the variance of X is

2.30. Let X be a binomial r.v. with parameters (n, p). Verify Eqs. (2.38) and (2.39).

By Eqs. (2.26) and (2.36), and letting q = 1 − p, we have

Letting i = k − 1 and using Eq. (2.86), we obtain

Next,

Similarly, letting i = k − 2 and using Eq. (2.86), we obtain

Thus,

and by Eq. (2.31),

2.31. Let X be a Poisson r.v. with parameter λ. Verify Eqs. (2.50) and (2.51).

By Eqs. (2.26) and (2.48),

Next,

Thus,

and by Eq. (2.31),

2.32. Let X be a discrete uniform r.v. defined by Eq. (2.52). Verify Eqs. (2.54) and (2.55), i.e.,

By Eqs. (2.52) and (2.26), we have

Now, using Eq. (2.31), we obtain

2.33. Find the mean and variance of the r.v. X of Prob. 2.22.

From Prob. 2.22, the pdf of X is

By Eq. (2.26), the mean of X is

By Eq. (2.27), we have

Thus, by Eq. (2.31), the variance of X is

2.34. Let X be a uniform r.v. over (a, b). Verify Eqs. (2.58) and (2.59).

By Eqs. (2.56) and (2.26), the mean of X is

By Eq. (2.27), we have

Thus, by Eq. (2.31), the variance of X is

2.35. Let X be an exponential r.v. X with parameter λ. Verify Eqs. (2.62) and (2.63).

By Eqs. (2.60) and (2.26), the mean of X is

Integrating by parts (u = x, dv = λe−λxdx) yields

Next, by Eq. (2.27),

Again integrating by parts (u = x2, dv = λe−λxdx), we obtain

Thus, by Eq. (2.31), the variance of X is

2.36. Let X = N(μ; σ2). Verify Eqs. (2.76) and (2.77).

Using Eqs. (2.71) and (2.26), we have

Writing x as (x − μ) + μ we have

Letting y = x − μ in the first integral, we obtain

The first integral is zero, since its integrand is an odd function. Thus, by the property of pdf Eq. (2.22), we get

Next, by Eq. (2.29),

From Eqs. (2.22) and (2.71), we have

Differentiating with respect to σ, we obtain

Multiplying both sides by , we have

Thus,

2.37. Find the mean and variance of a Rayleigh r.v. defined by Eq. (2.95) (Prob. 2.25).

Using Eqs. (2.95) and (2.26), we have

Now the variance of N(0; σ2) is given by

Since the integrand is an even function, we have

or

Then

Next,

Let y = x2/(2σ2). Then dy = x dx/σ2, and so

Hence, by Eq. (2.31),

2.38. Consider a continuous r.v. X with pdf ƒX (x). If ƒX(x) = 0 for x < 0, then show that, for any a > 0,

where μX = E(X). This is known as the Markov inequality.

From Eq. (2.23),

Since ƒX(x) = 0 for x < 0,

Hence,

2.39. For any a > 0, show that

where μX and are the mean and variance of X, respectively. This is known as the Chebyshev inequality.

From Eq. (2.23),

By Eq. (2.29),

Hence,

Note that by setting a = kσX in Eq. (2.116), we obtain

Equation (2.117) says that the probability that a r.v. will fall k or more standard deviations from its mean is ≤ 1/k2. Notice that nothing at all is said about the distribution function of X. The Chebyshev inequality is therefore

quite a generalized statement. However, when applied to a particular case, it may be quite weak.

Special Distributions

2.40. A binary source generates digits 1 and 0 randomly with probabilities 0.6 and 0.4, respectively. (a) What is the probability that two 1s and three 0s will occur in a five-digit

sequence? (b) What is the probability that at least three 1s will occur in a five-digit

sequence?

(a) Let X be the r.v. denoting the number of 1s generated in a five-digit sequence. Since there are only two possible outcomes (1 or 0), the probability of generating 1 is constant, and there are five digits, it is clear that X is a binomial r.v. with parameters (n, p) = (5, 0.6). Hence, by Eq. (2.36), the probability that two 1s and three 0s will occur in a five- digit sequence is

(b) The probability that at least three 1s will occur in a five-digit sequence is

where

Hence,

2.41. A fair coin is flipped 10 times. Find the probability of the occurrence of 5 or 6 heads.

Let the r.v. X denote the number of heads occurring when a fair coin is flipped 10 times. Then X is a binomial r.v. with parameters . Thus, by Eq. (2.36),

2.42. Let X be a binomial r.v. with parameters (n, p), where 0 < p < 1. Show that as k goes from 0 to n, the pmf pX(k) of X first increases monotonically and then decreases monotonically, reaching its largest value when k is the largest integer less than or equal to (n + 1)p.

By Eq. (2.36), we have

Hence, pX(k) ≥ pX (k − 1) if and only if (n − k + 1)p ≥ k(1 − p) or k ≤ (n + 1)p. Thus, we see that pX(k) increases monotonically and reaches its maximum when k is the largest integer less than or equal to (n + 1)p and then decreases monotonically.

2.43. Show that the Poisson distribution can be used as a convenient approximation to the binomial distribution for large n and small p.

From Eq. (2.36), the pmf of the binomial r.v. with parameters (n, p) is

Multiplying and dividing the right-hand side by nk, we have

If we let n → ∞ in such a way that np = λ remains constant, then

where we used the fact that

Hence, in the limit as n → ∞ with np = λ (and as p = λ/n → 0),

Thus, in the case of large n and small p,

which indicates that the binomial distribution can be approximated by the Poisson distribution.

2.44. A noisy transmission channel has a per-digit error probability p = 0.01. (a) Calculate the probability of more than one error in 10 received digits. (b) Repeat (a), using the Poisson approximation Eq. (2.119).

(a) It is clear that the number of errors in 10 received digits is a binomial r.v. X with parameters (n, p) = (10, 0.01). Then, using Eq. (2.36), we obtain

(b) Using Eq. (2.119) with λ = np = 10(0.01) = 0.1, we have

2.45. The number of telephone calls arriving at a switchboard during any 10- minute period is known to be a Poisson r.v. X with λ = 2. (a) Find the probability that more than three calls will arrive during any 10-

minute period. (b) Find the probability that no calls will arrive during any 10-minute

period.

(a) From Eq. (2.48), the pmf of X is

Thus,

(b) P(X = 0) = pX(0) = e−2 ≈ 0.135

2.46. Consider the experiment of throwing a pair of fair dice. (a) Find the probability that it will take less than six tosses to throw a 7. (b) Find the probability that it will take more than six tosses to throw a 7.

(a) From Prob. 1.37(a), we see that the probability of throwing a 7 on any toss is . Let X denote the number of tosses required for the first success of throwing a 7. Then, it is clear that X is a geometric r.v. with parameter

. Thus, using Eq. (2.90) of Prob. 2.16, we obtain

(b) Similarly, we get

2.47. Consider the experiment of rolling a fair die. Find the average number of rolls required in order to obtain a 6.

Let X denote the number of trials (rolls) required until the number 6 first appears. Then X is given by geometrical r.v. with parameter . From Eq. (2.104) of Prob. 2.29, the mean of X is given by

Thus, the average number of rolls required in order to obtain a 6 is 6.

2.48. Consider an experiment of tossing a fair coin repeatedly until the coin lands heads the sixth time. Find the probability that it takes exactly 10 tosses.

The number of tosses of a fair coin it takes to get 6 heads is a negative binomial r.v. X with parameters p = 0.5 and k = 6. Thus, by Eq. (2.45), the probability that it takes exactly 10 tosses is

2.49. Assume that the length of a phone call in minutes is an exponential r.v. X with parameter . If someone arrives at a phone booth just before you arrive, find the probability that you will have to wait (a) less than 5 minutes, and (b) between 5 and 10 minutes.

(a) From Eq. (2.60), the pdf of X is

Then

(b) Similarly,

2.50. All manufactured devices and machines fail to work sooner or later. Suppose that the failure rate is constant and the time to failure (in hours) is an exponential r.v. X with parameter λ. (a) Measurements show that the probability that the time to failure for

computer memory chips in a given class exceeds 104 hours is e−1 (≈0.368). Calculate the value of the parameter λ.

(b) Using the value of the parameter λ determined in part (a), calculate the time x0 such that the probability that the time to failure is less than x0 is 0.05.

(a) From Eq. (2.61), the cdf of X is given by

Now

from which we obtain λ = 10−4. (b) We want

from which we obtain

2.51. A production line manufactures 1000-ohm (Ω) resistors that have 10 percent tolerance. Let X denote the resistance of a resistor. Assuming that X is a normal r.v. with mean 1000 and variance 2500, find the probability that a resistor picked at random will be rejected.

Let A be the event that a resistor is rejected. Then A = {X < 900} ∪ {X > 1100}. Since {X < 900} ∩ {X > 1100} = ∅, we have

Since X is a normal r.v. with μ = 1000 and σ2 = 2500 (σ = 50), by Eq. (2.74) and Table A (Appendix A),

Thus,

2.52. The radial miss distance [in meters (m)] of the landing point of a parachuting sky diver from the center of the target area is known to be a Rayleigh r.v. X with parameter σ2 = 100. (a) Find the probability that the sky diver will land within a radius of 10 m

from the center of the target area. (b) Find the radius r such that the probability that X > r is e−1 (≈0.368).

(a) Using Eq. (2.96) of Prob. 2.25, we obtain

(b) Now

from which we obtain r2 = 200 and .

Conditional Distributions

2.53. Let X be a Poisson r.v. with parameter λ. Find the conditional pmf of X given B = (X is even).

From Eq. (2.48), the pdf of X is

Then the probability of event B is

Let A = {X is odd}. Then the probability of event A is

Now

Hence, adding Eqs. (2.120) and (2.121), we obtain

Now, by Eq. (2.81), the pmf of X given B is

If k is even, (X = k) ⊂ B and (X = k) ∩ B = (X = k). If k is odd, (X = k) ∩ B = ∅. Hence,

2.54. Show that the conditional cdf and pdf of X given the event B = (a < X ≤ b) are as follows:

Substituting B = (a < X ≤ b) in Eq. (2.78), we have

Now

Hence,

By Eq. (2.82), the conditional pdf of X given a < X ≤ b is obtained by differentiating Eq. (2.123) with respect to x. Thus,

2.55. Recall the parachuting sky diver problem (Prob. 2.48). Find the probability of the sky diver landing within a 10-m radius from the center of the target

area given that the landing is within 50 m from the center of the target area.

From Eq. (2.96) (Prob. 2.25) with σ2 = 100, we have

Setting x = 10 and b = 50 and a = − ∞ in Eq. (2.123), we obtain

2.56. Let X = N(0; σ2). Find E(X | X > 0) and Var(X | X > 0).

From Eq. (2.71), the pdf of X = N(0; σ2) is

Then by Eq. (2.124),

Let y = x2/(2σ2). Then dy = x dx/σ2, and we get

Next,

Then by Eq. (2.31), we obtain

2.57. If X is a nonnegative integer valued r.v. and it satisfies the following memoryless property [see Eq. (2.44)]

then show that X must be a geometric r.v.

Let pk = P(X = k), k = 1, 2, 3, …, then

By Eq. (1.55) and using Eqs. (2.129) and (2.130), we have

Hence,

Setting i = j = 1, we have

Now

Thus,

Finally, by Eq. (2.130), we get

Comparing with Eq. (2.40) we conclude that X is a geometric r.v. with parameter p.

2.58. If X is nonnegative continuous r.v. and it satisfies the following memoryless property (see Eq. (2.64))

then show that X must be an exponential r.v.

By Eq. (1.55), Eq. (2.132) reduces to

Hence,

Let

Then Eq. (2.133) becomes

which is satisfied only by exponential function, that is,

Since P(X > ∞) = 0 we let α = − λ(λ > 0), then

Now if X is a continuous r.v.

Comparing with Eq. (2.61), we conclude that X is an exponential r.v. with parameter λ.

SUPPLEMENTARY PROBLEMS

2.59. Consider the experiment of tossing a coin. Heads appear about once out of every three tosses. If this experiment is repeated, what is the probability of the event that heads appear exactly twice during the first five tosses?

2.60. Consider the experiment of tossing a fair coin three times (Prob. 1.1). Let X be the r.v. that counts the number of heads in each sample point. Find the following probabilities:

(a) P(X ≤ 1); (b) P (X > 1); and (c) P(0 < X < 3).

2.61. Consider the experiment of throwing two fair dice (Prob. 1.31). Let X be the r.v. indicating the sum of the numbers that appear.

(a) What is the range of X? (b) Find (i) P(X = 3); (ii) P(X ≤ 4); and (iii) P(3 < X ≤ 7).

2.62. Let X denote the number of heads obtained in the flipping of a fair coin twice. (a) Find the pmf of X. (b) Compute the mean and the variance of X.

2.63. Consider the discrete r.v. X that has the pmf

Let A = {ζ: X(ζ) = 1, 3, 5, 7, …}. Find P(A).

2.64. Consider the function given by

where k is a constant. Find the value of k such that p(x) can be the pmf of a discrete r.v. X.

2.65. It is known that the floppy disks produced by company A will be defective with probability 0.01. The company sells the disks in packages of 10 and offers a guarantee of replacement that at most 1 of the 10 disks is defective. Find the probability that a package purchased will have to be replaced.

2.66. Consider an experiment of tossing a fair coin sequentially until “head” appears. What is the probability that the number of tossing is less than 5?

2.67. Given that X is a Poisson r.v. and pX(0) = 0.0498, compute E(X) and P(X ≥ 3).

2.68. A digital transmission system has an error probability of 10−6 per digit. Find the probability of three or more errors in 106 digits by using the Poisson distribution approximation.

2.69. Show that the pmf pX(x) of a Poisson r.v. X with parameter λ satifsies the following recursion formula:

2.70. The continuous r.v. X has the pdf

where k is a constant. Find the value of k and the cdf of X.

2.71. The continuous r.v. X has the pdf

where k is a constant. Find the value of k and P(X > 1).

2.72. A r.v. X is defined by the cdf

(a) Find the value of k. (b) Find the type of X. (c) Find (i) ; (ii) ; and (iii) P(X > 2).

2.73. It is known that the time (in hours) between consecutive traffic accidents can be described by the exponential r.v. X with parameter . Find (i) P(X ≤ 60); (ii) P(X > 120); and (iii) P(10 < X ≤ 100).

2.74. Binary data are transmitted over a noisy communication channel in a block of 16 binary digits. The probability that a received digit is in error as a result of channel noise is 0.01. Assume that the errors occurring in various digit positions within a block are independent. (a) Find the mean and the variance of the number of errors per block. (b) Find the probability that the number of errors per block is greater than

or equal to 4.

2.75. Let the continuous r.v. X denote the weight (in pounds) of a package. The range of weight of packages is between 45 and 60 pounds. (a) Determine the probability that a package weighs more than 50 pounds. (b) Find the mean and the variance of the weight of packages.

2.76. In the manufacturing of computer memory chips, company A produces one defective chip for every nine good chips. Let X be time to failure (in months) of chips. It is known that X is an exponential r.v. with parameter for a defective chip and with a good chip. Find the probability that a chip purchased randomly will fail before (a) six months of use; and (b) one year of use.

2.77. The median of a continuous r.v. X is the value of x = x0 such that P(X ≥ x0) = P(X ≤ x0). The mode of X is the value of x = xm at which the pdf of X achieves its maximum value. (a) Find the median and mode of an exponential r.v. X with parameter λ. (b) Find the median and mode of a normal r.v. X 5 N(μ, σ2).

2.78. Let the r.v. X denote the number of defective components in a random sample of n components, chosen without replacement from a total of N components, r of which are defective. The r.v. X is known as the hypergeometric r.v. with parameters (N, r, n). (a) Find the pmf of X. (b) Find the mean and variance of X.

2.79. A lot consisting of 100 fuses is inspected by the following procedure: Five fuses are selected randomly, and if all five “blow” at the specified amperage, the lot is accepted. Suppose that the lot contains 10 defective fuses. Find the probability of accepting the lot.

2.80. Let X be the negative binomial r.v. with parameters p and k. Verify Eqs. (2.46) and (2.47), that is,

2.81. Suppose the probability that a bit transmitted through a digital communication channel and received in error is 0.1. Assuming that the transmissions are independent events, find the probability that the third error occurs at the 10th bit.

2.82. A r.v. X is called a Laplace r.v. if its pdf is given by

where k is a constant. (a) Find the value of k. (b) Find the cdf of X. (c) Find the mean and the variance of X.

2.83. A r.v. X is called a Cauchy r.v. if its pdf is given by

where a (> 0) and k are constants. (a) Find the value of k. (b) Find the cdf of X. (c) Find the mean and the variance of X.

ANSWERS TO SUPPLEMENTARY PROBLEMS

2.59. 0.329

2.60.

2.61.

2.62.

2.63.

2.64. k = 6/π2

2.65. 0.004

2.66. 0.9375

2.67. E(X) = 3, P(X ≥ 3) = 0.5767

2.68. 0.08

2.69. Hint: Use Eq. (2.48).

2.70.

2.71.

2.72.

2.73. (i) 0.632; (ii) 0.135; (iii) 0.658

2.74. (a) E(X) = 0.16, Var(X) = 0.158 (b) 0.165 × 10−4

2.75. Hint: Assume that X is uniformly distributed over (45, 60). (a) (b) E(X) = 52.5, Var(X) = 18.75

2.76. (a) 0.501; (b) 0.729

2.77. (a) x0 = (ln 2)/λ = 0.693/λ, xm = 0 (b) x0 = xm = μ

2.78. Hint: To find E(X), note that

To find Var(X), first find E[X(X − 1)].

2.79. Hint: Let X be a r.v. equal to the number of defective fuses in the sample of 5 and use the result of Prob. 2.78.

0.584

2.80. Hint: To find E(X), use Maclaurin’s series expansions of the negative binomial h(q) = (1 − q)−r and its derivatives h′(q) and h″(q), and note that

To find Var(X), first find E[(X − r) (X − r − 1)] using h″(q).

2.81. 0.017

2.82.

2.83. (a) k = a/π

(c) E(X) = 0, Var (X) does not exist.

CHAPTER 3

Multiple Random Variables

3.1 Introduction

In many applications it is important to study two or more r.v.’s defined on the same sample space. In this chapter, we first consider the case of two r.v.’s, their associated distribution, and some properties, such as independence of the r.v.’s. These concepts are then extended to the case of many r.v.’s defined on the same sample space.

3.2 Bivariate Random Variables

A. Definition:

Let S be the sample space of a random experiment. Let X and Y be two r.v.’s. Then the pair (X, Y) is called a bivariate r.v. (or two-dimensional random vector) if each of X and Y associates a real number with every element of S. Thus, the bivariate r.v. (X, Y) can be considered as a function that to each point ζ in S assigns a point (x, y) in the plane (Fig. 3-1). The range space of the bivariate r.v. (X, Y) is denoted by Rxy and defined by

Fig. 3-1 (X, Y) as a function from S to the plane.

If the r.v.’s X and Y are each, by themselves, discrete r.v.’s, then (X, Y) is called a discrete bivariate r.v. Similarly, if X and Y are each, by themselves, continuous r.v.’s, then (X, Y) is called a continuous bivariate r.v. If one of X and Y is discrete while the other is continuous, then (X, Y) is called a mixed bivariate r.v.

3.3 Joint Distribution Functions

A. Definition:

The joint cumulative distribution function (or joint cdf) of X and Y, denoted by FXY(x, y), is the function defined by

The event (X ≤ x, Y ≤ y) in Eq. (3.1) is equivalent to the event A ∩ B, where A and B are events of S defined by

If, for particular values of x and y, A and B were independent events of S, then by Eq. (1.62),

B. Independent Random Variables: Two r.v.’s X and Y will be called independent if

for every value of x and y.

C. Properties of FXY(x, y):

The joint cdf of two r.v.’s has many properties analogous to those of the cdf of a single r.v.

Note that the left-hand side of Eq. (3.12) is equal to P(x1, X ≤ x2, y1, Y ≤ y2) (Prob. 3.5).

D. Marginal Distribution Functions: Now

since the condition y ≤ ∞ is always satisfied. Then

Similarly,

The cdf’s FX(x) and FY(y), when obtained by Eqs. (3.13) and (3.14), are referred to as the marginal cdf ’s of X and Y, respectively.

3.4 Discrete Random Variables—Joint Probability Mass Functions

A. Joint Probability Mass Functions:

Let (X, Y) be a discrete bivariate r.v., and let (X, Y) take on the values (xi, yj) for a certain allowable set of integers i and j. Let

The function pXY(xi, yj) is called the joint probability mass function (joint pmf) of (X, Y).

B. Properties of pXY(xi, yj):

where the summation is over the points (xi, yj) in the range space RA corresponding to the event A. The joint cdf of a discrete bivariate r.v. (X, Y) is given by

C. Marginal Probability Mass Functions: Suppose that for a fixed value X = xi, the r.v. Y can take on only the possible values yj (j = 1, 2, …, n). Then

where the summation is taken over all possible pairs (xi, yj) with xi fixed. Similarly,

where the summation is taken over all possible pairs (xi, yj) with yj fixed. The pmf’s pX(xi) and pY(yj), when obtained by Eqs. (3.20) and (3.21), are referred to as the marginal pmf’s of X and Y, respectively.

D. Independent Random Variables: If X and Y are independent r.v.’s, then (Prob. 3.10)

3.5 Continuous Random Variables—Joint Probability Density Functions

A. Joint Probability Density Functions:

Let (X, Y) be a continuous bivariate r.v. with cdf FXY(x, y) and let

The function fXY(x, y) is called the joint probability density function (joint pdf) of (X, Y). By integrating Eq. (3.23), we have

B. Properties of fXY(x, y):

C. Marginal Probability Density Functions: By Eq. (3.13), Hence,

or

Similarly,

The pdf’s fX(x) and fY(y), when obtained by Eqs. (3.30) and (3.31), are referred to as the marginal pdf’s of X and Y, respectively.

D. Independent Random Variables: If X and Y are independent r.v.’s, by Eq. (3.4),

Then or

analogous with Eq. (3.22) for the discrete case. Thus, we say that the continuous r.v.’s X and Y are independent r.v.’s if and only if Eq. (3.32) is satisfied.

3.6 Conditional Distributions

A. Conditional Probability Mass Functions:

If (X, Y) is a discrete bivariate r.v. with joint pmf pXY(xi, yj), then the conditional pmf of Y, given that X = xi, is defined by

Similarly, we can define pX|Y(xi|yj) as

B. Properties of pY|X(yj|xi):

Notice that if X and Y are independent, then by Eq. (3.22),

C. Conditional Probability Density Functions: If (X, Y) is a continuous bivariate r.v. with joint pdf fXY(x, y), then the conditional pdf of Y, given that X = x, is defined by

Similarly, we can define fX|Y(x|y) as

D. Properties of fY|X(y|x):

As in the discrete case, if X and Y are independent, then by Eq. (3.32),

3.7 Covariance and Correlation Coefficient

The (k, n)th moment of a bivariate r.v. (X, Y) is defined by

If n = 0, we obtain the kth moment of X, and if k = 0, we obtain the nth moment of Y. Thus,

If (X, Y) is a discrete bivariate r.v., then using Eqs. (3.43), (3.20), and (3.21), we obtain

Similarly, we have

If (X, Y) is a continuous bivariate r.v., then using Eqs. (3.43), (3.30), and (3.31), we obtain

Similarly, we have

The variances of X and Y are easily obtained by using Eq. (2.31). The (1, 1)th joint moment of (X, Y),

is called the correlation of X and Y. If E(XY) = 0, then we say that X and Y are orthogonal. The covariance of X and Y, denoted by Cov(X, Y) or σ XY, is defined by

Expanding Eq. (3.50), we obtain

If Cov(X, Y) = 0, then we say that X and Y are uncorrelated. From Eq. (3.51), we see that X and Y are uncorrelated if

Note that if X and Y are independent, then it can be shown that they are uncorrelated (Prob. 3.32), but the converse is not true in general; that is, the fact that X and Y are uncorrelated does not, in general, imply that they are independent (Probs. 3.33, 3.34, and 3.38). The correlation coefficient, denoted by ρ (X, Y) or ρXY, is defined by

It can be shown that (Prob. 3.36)

Note that the correlation coefficient of X and Y is a measure of linear dependence between X and Y (see Prob. 4.46).

3.8 Conditional Means and Conditional Variances

If (X, Y) is a discrete bivariate r.v. with joint pmf pXY(xi, yj), then the conditional mean (or conditional expectation) of Y, given that X = xi, is defined by

The conditional variance of Y, given that X = xi, is defined by

which can be reduced to

The conditional mean of X, given that Y = yj, and the conditional variance of X, given that Y = yj, are given by similar expressions. Note that the conditional mean of Y, given that X = xi, is a function of xi alone. Similarly, the conditional mean of X, given that Y = yj, is a function of yj alone.

If (X, Y) is a continuous bivariate r.v. with joint pdf fXY(x, y), the conditional mean of Y, given that X = x, is defined by

The conditional variance of Y, given that X = x, is defined by

which can be reduced to

The conditional mean of X, given that Y = y, and the conditional variance of X, given that Y = y, are given by similar expressions. Note that the conditional mean of Y, given that X = x, is a function of x alone. Similarly, the conditional mean of X, given that Y = y, is a function of y alone (Prob. 3.40).

3.9 N-Variate Random Variables

In previous sections, the extension from one r.v. to two r.v.’s has been made. The concepts can be extended easily to any number of r.v.’s defined on the same sample space. In this section we briefly describe some of the extensions.

A. Definitions: Given an experiment, the n-tuple of r.v.’s (X1, X2, …, Xn) is called an n-variate r.v. (or n-dimensional random vector) if each Xi, i = 1, 2, …, n, associates a real number with every sample point ζ ∈ S. Thus, an n-variate r.v. is simply a rule associating an n-tuple of real numbers with every ζ ∈ S.

Let (X1, …, Xn) be an n-variate r.v. on S. Then its joint cdf is defined as

Note that

The marginal joint cdf’s are obtained by setting the appropriate Xi’s to + ∞ in Eq. (3.61). For example,

A discrete n-variate r.v. will be described by a joint pmf defined by

The probability of any n-dimensional event A is found by summing Eq. (3.65) over the points in the n-dimensional range space RA corresponding to the event A:

Properties of pX1 … Xn(x1, …, xn):

The marginal pmf’s of one or more of the r.v.’s are obtained by summing Eq. (3.65) appropriately. For example,

Conditional pmf’s are defined similarly. For example,

A continuous n-variate r.v. will be described by a joint pdf defined by

Then

and

Properties of fX1 … Xn(x1, …, xn):

The marginal pdf’s of one or more of the r.v.’s are obtained by integrating Eq. (3.72) appropriately. For example,

Conditional pdf’s are defined similarly. For example,

The r.v.’s X1, …, Xn are said to be mutually independent if

for the discrete case, and

for the continuous case. The mean (or expectation) of Xi in (X1, …, Xn) is defined as

The variance of Xi is defined as

The covariance of Xi and Xj is defined as

The correlation coefficient of Xi and Xj is defined as

3.10 Special Distributions

A. Multinomial Distribution:

The multinomial distribution is an extension of the binomial distribution. An experiment is termed a multinomial trial with parameters p1, p2, …, pk, if it has the following conditions:

1. The experiment has k possible outcomes that are mutually exclusive and exhaustive, say A1, A2, …, Ak.

2.

Consider an experiment which consists of n repeated, independent, multinomial trials with parameters p1, p2, …, pk. Let Xi be the r.v. denoting the number of trials which result in Ai. Then (X1, X2, …, Xk) is called the multinomial r.v. with parameters (n, p1, p2, …, pk) and its pmf is given by (Prob. 3.46)

for xi = 0, 1, …, n, i = 1, …, k, such that .

Note that when k = 2, the multinomial distribution reduces to the binomial distribution.

B. Bivariate Normal Distribution: A bivariate r.v. (X, Y) is said to be a bivariate normal (or Gaussian) r.v. if its joint pdf is given by

where

and are the means and variances of X and Y, respectively. It can be shown that ρ is the correlation coefficient of X and Y (Prob. 3.50) and that X and Y are independent when ρ = 0 (Prob. 3.49).

C. N-variate Normal Distribution: Let (X1, …, Xn) be an n-variate r.v. defined on a sample space S. Let X be an n- dimensional random vector expressed as an n × 1 matrix:

Let x be an n-dimensional vector (n × 1 matrix) defined by

The n-variate r.v. (X1, …, Xn) is called an n-variate normal r.v. if its joint pdf is given by

where T denotes the “transpose,” μ is the vector mean, K is the covariance matrix given by

and det K is the determinant of the matrix K. Note that fX(x) stands for fX1 … Xn(x1, …, xn).

SOLVED PROBLEMS

Bivariate Random Variables and Joint Distribution Functions

3.1. Consider an experiment of tossing a fair coin twice. Let (X, Y) be a bivariate r.v., where X is the number of heads that occurs in the two tosses and Y is the number of tails that occurs in the two tosses. (a) What is the range RX of X? (b) What is the range RY of Y? (c) Find and sketch the range RXY of (X, Y). (d) Find P(X = 2, Y = 0), P(X = 0, Y = 2), and P(X = 1, Y = 1).

The sample space S of the experiment is

(a) RX = {0, 1, 2} (b) RY = {0, 1, 2} (c) RXY = {(2, 0), (1, 1), (0, 2)} which is sketched in Fig. 3-2.

Fig. 3-2

(d) Since the coin is fair, we have

3.2. Consider a bivariate r.v. (X, Y). Find the region of the xy plane corresponding to the events

The region corresponding to event A is expressed by x + y ≤ 2, which is shown in Fig. 3-3(a), that is, the region below and including the straight line x + y = 2.

Fig. 3-3 Regions corresponding to certain joint events.

The region corresponding to event B is expressed by x2 + y2 < 22, which is shown in Fig. 3-3(b), that is, the region within the circle with its center at the origin and radius 2.

The region corresponding to event C is shown in Fig. 3-3(c), which is found by noting that

The region corresponding to event D is shown in Fig. 3-3(d), which is found by noting that

3.3. Verify Eqs. (3.7), (3.8a), and (3.8b).

Since {X ≤ ∞, Y ≤ ∞} = S and by Eq. (1.36),

Next, as we know, from Eq. (2.8),

Since and by Eq. (1.41), we have

3.4. Verify Eqs. (3.10) and (3.11). Clearly

The two events on the right-hand side are disjoint; hence by Eq. (1.37),

or

Similarly,

Again the two events on the right-hand side are disjoint, hence

or

3.5. Verify Eq. (3.12).

Clearly

The two events on the right-hand side are disjoint; hence

Then using Eq. (3.10), we obtain

Since the probability must be nonnegative, we conclude that

if x2 ≥ x1 and y2 ≥ y1.

3.6. Consider a function

Can this function be a joint cdf of a bivariate r.v. (X, Y)?

It is clear that F(x, y) satisfies properties 1 to 5 of a cdf [Eqs. (3.5) to (3.9)]. But substituting F(x, y) in Eq. (3.12) and setting x2 = y2 = 2 and x1 = y1 = 1, we get

Thus, property 7 [Eq. (3.12)] is not satisfied. Hence, F(x, y) cannot be a joint cdf.

3.7. Consider a bivariate r.v. (X, Y). Show that if X and Y are independent, then every event of the form (a < X ≤ b) is independent of every event of the form (c < Y ≤ d).

By definition (3.4), if X and Y are independent, we have

Setting x1 = a, x2 = b, y1 = c, and y2 = d in Eq. (3.95) (Prob. 3.5), we obtain

which indicates that event (a < X ≤ b) and event (c < Y ≤ d) are independent [Eq. (1.62)].

3.8. The joint cdf of a bivariate r.v. (X, Y) is given by

(a) Find the marginal cdf’s of X and Y. (b) Show that X and Y are independent. (c) Find P(X ≤ 1, Y ≤ 1), P(X ≤ 1), P (Y > 1), and P(X > x, Y > y).

(a) By Eqs. (3.13) and (3.14), the marginal cdf’s of X and Y are

(b) Since FXY(x, y) = FX(x)FY(y), X and Y are independent.

(c)

By De Morgan’s law (1.21), we have

Then by Eq. (1.44),

Finally, by Eq. (1.39), we obtain

3.9. The joint cdf of a bivariate r.v. (X, Y) is given by

(a) Find the marginal cdf’s of X and Y. (b) Find the conditions on p1, p2, and p3 for which X and Y are independent.

(a) By Eq. (3.13), the marginal cdf of X is given by

By Eq. (3.14), the marginal cdf of Y is given by

(b) For X and Y to be independent, by Eq. (3.4), we must have FXY(x, y) = FX(x)FY(y). Thus, for 0 ≤ x < a, 0 ≤ y < b, we must have p1 = p2 p3 for X and Y to be independent.

Discrete Bivariate Random Variables—Joint Probability Mass Functions

3.10. Verify Eq. (3.22).

If X and Y are independent, then by Eq. (1.62),

3.11. Two fair dice are thrown. Consider a bivariate r.v. (X, Y). Let X = 0 or 1 according to whether the first die shows an even number or an odd number of dots. Similarly, let Y = 0 or 1 according to the second die. (a) Find the range RXY of (X, Y). (b) Find the joint pmf’s of (X, Y).

(a) The range of (X, Y) is

(b) It is clear that X and Y are independent and

3.12. Consider the binary communication channel shown in Fig. 3-4 (Prob. 1.52). Let (X, Y) be a bivariate r.v., where X is the input to the channel and Y is the output of the channel. Let P(X = 0) = 0.5, P(Y = 1 | X = 0) = 0.1, and P(Y = 0 | X = 1) = 0.2.

Fig. 3-4 Binary communication channel.

(a) Find the joint pmf’s of (X, Y). (b) Find the marginal pmf’s of X and Y. (c) Are X and Y independent?

(a) From the results of Prob. 1.62, we found that

Then by Eq. (1.57), we obtain

Hence, the joint pmf’s of (X, Y) are

(b) By Eq. (3.20), the marginal pmf’s of X are

By Eq. (3.21), the marginal pmf’s of Y are

(c) Now

Hence, X and Y are not independent.

3.13. Consider an experiment of drawing randomly three balls from an urn containing two red, three white, and four blue balls. Let (X, Y) be a bivariate r.v. where X and Y denote, respectively, the number of red and white balls chosen. (a) Find the range of (X, Y). (b) Find the joint pmf’s of (X, Y). (c) Find the marginal pmf’s of X and Y. (d) Are X and Y independent?

(a) The range of (X, Y) is given by

(b) The joint pmf’s of (X, Y)

are given as follows:

which are expressed in tabular form as in Table 3-1.

TABLE 3-1 pXY(i, j)

(c) The marginal pmf’s of X are obtained from Table 3-1 by computing the row sums, and the marginal pmf’s of Y are obtained by computing the column sums. Thus,

(d) Since

X and Y are not independent.

3.14. The joint pmf of a bivariate r.v. (X, Y) is given by

where k is a constant. (a) Find the value of k. (b) Find the marginal pmf’s of X and Y. (c) Are X and Y independent?

(a) By Eq. (3.17),

Thus, . (b) By Eq. (3.20), the marginal pmf’s of X are

By Eq. (3.21), the marginal pmf’s of Y are

(c) Now pX(xi)pY (yj) ≠ pXY(xi, yj); hence X and Y are not independent.

3.15. The joint pmf of a bivariate r.v. (X, Y) is given by

where k is a constant. (a) Find the value of k. (b) Find the marginal pmf’s of X and Y. (c) Are X and Y independent?

(a) By Eq. (3.17),

Thus, . (b) By Eq. (3.20), the marginal pmf’s of X are

By Eq. (3.21), the marginal pmf’s of Y are

(c) Now

Hence, X and Y are independent.

3.16. Consider an experiment of tossing two coins three times. Coin A is fair, but coin B is not fair, with and . Consider a bivariate r.v. (X, Y), where X denotes the number of heads resulting from coin A and Y denotes the number of heads resulting from coin B. (a) Find the range of (X, Y).

(b) Find the joint pmf’s of (X, Y). (c) Find P(X = Y), P(X > Y), and P(X + Y ≤ 4).

(a) The range of (X, Y) is given by

(b) It is clear that the r.v.’s X and Y are independent, and they are both binomial r.v.’s with parameters and , respectively. Thus, by Eq. (2.36), we have

Since X and Y are independent, the joint pmf’s of (X, Y) are

which are tabulated in Table 3-2.

Table 3-2 pXY(i, j)

(c) From Table 3-2, we have

Continuous Bivariate Random Variables—Probability Density Functions

3.17. The joint pdf of a bivariate r.v. (X, Y) is given by

where k is a constant. (a) Find the value of k. (b) Find the marginal pdf’s of X and Y. (c) Are X and Y independent?

(a) By Eq. (3.26),

Thus, .

(b) By Eq. (3.30), the marginal pdf of X is

Since fXY(x, y) is symmetric with respect to x and y, the marginal pdf of Y is

(c) Since fXY(x, y) ≠ fX(x)fY(y), X and Y are not independent.

3.18. The joint pdf of a bivariate r.v. (X, Y) is given by

where k is a constant. (a) Find the value of k. (b) Are X and Y independent? (c) Find P(X + Y < 1).

(a) The range space RXY is shown in Fig. 3-5(a). By Eq. (3.26),

Fig. 3-5

Thus k = 4. (b) To determine whether X and Y are independent, we must find the marginal

pdf’s of X and Y. By Eq. (3.30),

By symmetry,

Since fXY(x, y) = fX(x)fY(y), X and Y are independent. (c) The region in the xy plane corresponding to the event (X + Y < 1) is shown

in Fig. 3-5(b) as a shaded area. Then

3.19. The joint pdf of a bivariate r.v. (X, Y) is given by

where k is a constant. (a) Find the value of k. (b) Are X and Y independent?

(a) The range space RXY is shown in Fig. 3-6. By Eq. (3.26),

Fig. 3-6

Thus k = 8.

(b) By Eq. (3.30), the marginal pdf of X is

By Eq. (3.31), the marginal pdf of Y is

Since fXY(x,y) ≠ fY(x) fY(y), X and Y are not independent. Note that if the range space RXY depends functionally on x or y, then X

and Y cannot be independent r.v.’s.

3.20. The joint pdf of a bivariate r.v. (X, Y) is given by

where k is a constant. (a) Determine the value of k. (b) Find the marginal pdf’s of X and Y. (c) Find .

(a) The range space RXY is shown in Fig. 3-7. By Eq. (3.26),

Fig. 3-7

Thus k = 2. (b) By Eq. (3.30), the marginal pdf of X is

By Eq. (3.31), the marginal pdf of Y is

(c) The region in the xy plane corresponding to the event is shown in Fig. 3-7 as the shaded area Rs. Then

Note that the bivariate r.v. (X, Y) is said to be uniformly distributed over the region RXY if its pdf is

where k is a constant. Then by Eq. (3.26), the constant k must be k = 1/(area of RXY).

3.21. Suppose we select one point at random from within the circle with radius R. If we let the center of the circle denote the origin and define X and Y to be the coordinates of the point chosen (Fig. 3-8), then (X, Y) is a uniform bivariate r.v. with joint pdf given by

Fig. 3-8

where k is a constant. (a) Determine the value of k. (b) Find the marginal pdf’s of X and Y. (c) Find the probability that the distance from the origin of the point selected is

not greater than a.

(a) By Eq. (3.26),

Thus, k = 1/ πR2. (b) By Eq. (3.30), the marginal pdf of X is

By symmetry, the marginal pdf of Y is

(c) For 0 ≤ a ≤ R,

3.22. The joint pdf of a bivariate r.v. (X, Y) is given by

where a and b are positive constants and k is a constant. (a) Determine the value of k. (b) Are X and Y independent?

(a) By Eq. (3.26),

Thus, k = ab. (b) By Eq. (3.30), the marginal pdf of X is

By Eq. (3.31), the marginal pdf of Y is

Since fXY(x, y) = fX(x)fY(y), X and Y are independent.

3.23. A manufacturer has been using two different manufacturing processes to make computer memory chips. Let (X, Y) be a bivariate r.v., where X denotes the time to failure of chips made by process A and Y denotes the time to failure of chips made by process B. Assuming that the joint pdf of (X, Y) is

where a = 10−4 and b = 1.2(10−4), determine P(X > Y).

The region in the xy plane corresponding to the event (X > Y) is shown in Fig. 3- 9 as the shaded area. Then

Fig. 3-9

3.24. A smooth-surface table is ruled with equidistant parallel lines a distance D apart. A needle of length L, where L ≤ D, is randomly dropped onto this table. What is the probability that the needle will intersect one of the lines? (This is known as Buffon’s needle problem.)

We can determine the needle’s position by specifying a bivariate r.v. (X, Θ), where X is the distance from the middle point of the needle to the nearest parallel line and Θ is the angle from the vertical to the needle (Fig. 3-10). We interpret the statement “the needle is randomly dropped" to mean that both X and Θ have uniform distributions and that X and Θ are independent. The possible values of X are between 0 and D/2, and the possible values of Θ are between 0 and π/2. Thus, the joint pdf of (X, Θ) is

Fig. 3-10 Buffon’s needle problem.

From Fig. 3-10, we see that the condition for the needle to intersect a line is X < L/2 cos θ. Thus, the probability that the needle will intersect a line is

Conditional Distributions

3.25. Verify Eqs. (3.36) and (3.41).

(a) By Eqs. (3.33) and (3.20),

(b) Similarly, by Eqs. (3.38) and (3.30),

3.26. Consider the bivariate r.v. (X, Y) of Prob. 3.14. (a) Find the conditional pmf’s pY | X(yj | xi) and pX | Y(xi | yj).

(b) Find P(Y = 2 X = 2) and P(X = 2 Y = 2). (a) From the results of Prob. 3.14, we have

Thus, by Eqs. (3.33) and (3.34),

(b) Using the results of part (a), we obtain

3.27. Find the conditional pmf’s pY | X (yj | xi) and pX | Y(xi|y j) for the bivariate r.v. (X, Y) of Prob. 3.15.

From the results of Prob. 3.15, we have

Thus, by Eqs. (3.33) and (3.34),

Note that PY | X(yj | xi) = pY(yj) and pX | Y(xi|yj) = pX(xi), as must be the case since X and Y are independent, as shown in Prob. 3.15.

3.28. Consider the bivariate r.v. (X, Y) of Prob. 3.17. (a) Find the conditional pdf’s fY | X(y |x) and fX | Y(x | y). (b) Find .

(a) From the results of Prob. 3.17, we have

Thus, by Eqs. (3.38) and (3.39),

(b) Using the results of part (a), we obtain

3.29. Find the conditional pdf’s ƒY | X(y |x) and ƒX | Y(x | y) for the bivariate r.v. (X, Y) of Prob. 3.18.

From the results of Prob. 3.18, we have

Thus, by Eqs. (3.38) and (3.39),

Again note that fY | X(y |x) = fY(y) and fX | Y (x | y) = fX(x), as must be the case since X and Y are independent, as shown in Prob. 3.18.

3.30. Find the conditional pdf’s ƒY | X(y |x) and ƒX | Y(x | y) for the bivariate r.v. (X, Y) of Prob. 3.20.

From the results of Prob. 3.20, we have

Thus, by Eqs. (3.38) and (3.39),

3.31. The joint pdf of a bivariate r.v. (X, Y) is given by

(a) Show that fXY(x, y) satisfies Eq. (3.26). (b) Find P(X > 1 | Y = y).

(a) We have

(b) First we must find the marginal pdf on Y. By Eq. (3.31),

By Eq. (3.39), the conditional pdf of X is

Then

Covariance and Correlation Coefficients

3.32. Let (X, Y) be a bivariate r.v. If X and Y are independent, show that X and Y are uncorrelated.

If (X, Y) is a discrete bivariate r.v., then by Eqs. (3.43) and (3.22),

If (X, Y) is a continuous bivariate r.v., then by Eqs. (3.43) and (3.32),

Thus, X and Y are uncorrelated by Eq. (3.52).

3.33. Suppose the joint pmf of a bivariate r.v. (X, Y) is given by

(a) Are X and Y independent? (b) Are X and Y uncorrelated ?

(a) By Eq. (3.20), the marginal pmf’s of X are

By Eq. (3.21), the marginal pmf’s of Y are

and Thus, X and Y are not independent.

(b) By Eqs. (3.45a), (3.45b), and (3.43), we have

Now by Eq. (3.51),

Thus, X and Y are uncorrelated.

3.34. Let (X, Y) be a bivariate r.v. with the joint pdf

Show that X and Y are not independent but are uncorrelated.

By Eq. (3.30), the marginal pdf of X is

Noting that the integrand of the first integral in the above expression is the pdf of N(0; 1) and the second integral in the above expression is the variance of N(0; 1), we have

Since fXY(x, y) is symmetric in x and y, we have

Now fXY(x, y) ≠ fX(x) fY(y), and hence X and Y are not independent. Next, by Eqs. (3.47a) and (3.47b),

since for each integral the integrand is an odd function. By Eq. (3.43),

The integral vanishes because the contributions of the second and the fourth quadrants cancel those of the first and the third. Thus, E(XY) = E(X)E(Y), and so X and Y are uncorrelated.

3.35. Let (X, Y) be a bivariate r.v. Show that

This is known as the Cauchy-Schwarz inequality.

Consider the expression E[(X − αY)2] for any two r.v.’s X and Y and a real variable α. This expression, when viewed as a quadratic in α, is greater than or equal to zero; that is,

for any value of α. Expanding this, we obtain

Choose a value of α for which the left-hand side of this inequality is minimum,

which results in the inequality

3.36. Verify Eq. (3.54).

From the Cauchy-Schwarz inequality [Eq. (3.97)], we have

Since ρ XY is a real number, this implies

3.37. Let (X, Y) be the bivariate r.v. of Prob. 3.12. (a) Find the mean and the variance of X. (b) Find the mean and the variance of Y. (c) Find the covariance of X and Y. (d) Find the correlation coefficient of X and Y.

(a) From the results of Prob. 3.12, the mean and the variance of X are evaluated as follows:

(b) Similarly, the mean and the variance of Y are

(c) By Eq. (3.43),

By Eq. (3.51), the covariance of X and Y is

(d) By Eq. (3.53), the correlation coefficient of X and Y is

3.38. Suppose that a bivariate r.v. (X, Y) is uniformly distributed over a unit circle (Prob. 3.21). (a) Are X and Y independent? (b) Are X and Y correlated?

(a) Setting R = 1 in the results of Prob. 3.21, we obtain

Since fXY(x, y) ≠ fX(x) fY (y), X and Y are not independent. (b) By Eqs. (3.47a) and (3.47b), the means of X and Y are

since each integrand is an odd function. Next, by Eq. (3.43),

The integral vanishes because the contributions of the second and the fourth quadrants cancel those of the first and the third. Hence, E(XY) = E(X)E(Y) = 0 and X and Y are uncorrelated.

Conditional Means and Conditional Variances

3.39. Consider the bivariate r.v. (X, Y) of Prob. 3.14 (or Prob. 3.26). Compute the conditional mean and the conditional variance of Y given xi = 2.

From Prob. 3.26, the conditional pmf pY | X(yj | xi) is

and by Eqs. (3.55) and (3.56), the conditional mean and the conditional variance of Y given xi = 2 are

3.40. Let (X, Y) be the bivariate r.v. of Prob. 3.20 (or Prob. 3.30). Compute the conditional means E(Y |x) and E(X |y).

From Prob. 3.30,

By Eq. (3.58), the conditional mean of Y, given X = x, is

Similarly, the conditional mean of X, given Y = y, is

Note that E(Y |x) is a function of x only and E(X |y) is a function of y only.

3.41. Let (X, Y) be the bivariate r.v. of Prob. 3.20 (or Prob. 3.30). Compute the conditional variances Var(Y |x) and Var(X |y).

Using the results of Prob. 3.40 and Eq. (3.59), the conditional variance of Y, given X = x, is

Similarly, the conditional variance of X, given Y = y, is

N-Dimensional Random Vectors

3.42. Let (X1, X2, X3, X4) be a four-dimensional random vector, where Xk (k = 1, 2, 3, 4) are independent Poisson r.v.’s with parameter 2. (a) Find P(X1 = 1, X2 = 3, X3 = 2, X4 = 1). (b) Find the probability that exactly one of the Xk’s equals zero.

(a) By Eq. (2.48), the pmf of Xk is

Since the Xk’s are independent, by Eq. (3.80),

(b) First, we find the probability that Xk = 0, k = 1, 2, 3, 4. From Eq. (3.98),

Next, we treat zero as “success.” If Y denotes the number of successes, then Y is a binomial r.v. with parameters (n, p) = (4, e−2). Thus, the probability that exactly one of the Xk’s equals zero is given by [Eq. (2.36)]

3.43. Let (X, Y, Z) be a trivariate r.v., where X, Y, and Z are independent uniform r.v.’s over (0, 1). Compute P(Z ≥ XY).

Since X, Y, Z are independent and uniformly distributed over (0, 1), we have

3.44. Let (X, Y, Z) be a trivariate r.v. with joint pdf

where a, b, c > 0 and k are constants. (a) Determine the value of k. (b) Find the marginal joint pdf of X and Y. (c) Find the marginal pdf of X. (d) Are X, Y, and Z independent?

(a) By Eq. (3.76),

Thus, k = abc. (b) By Eq. (3.77), the marginal joint pdf of X and Y is

(c) By Eq. (3.78), the marginal pdf of X is

(d) Similarly, we obtain

Since fXYZ(x, y, z) = fX(x) fY(y) fZ(z), X, Y, and Z are independent.

3.45. Show that

By definition (3.79),

Hence

Now, by Eq. (3.38),

Substituting this expression into Eq. (3.100), we obtain

Special Distributions

3.46. Derive Eq. (3.87).

Consider a sequence of n independent multinomial trials. Let Ai (i = 1, 2, …, k) be the outcome of a single trial. The r.v. Xi is equal to the number of times Ai occurs in the n trials. If x1, x2, …, xk are nonnegative integers such that their sum equals n, then for such a sequence the probability that Ai occurs xi times, i = 1, 2, …, k—that is, P(X1 = x1, X2 = x2, …, Xk = xk)—can be obtained by counting the number of sequences containing exactly x1 A1’s, x2 A2’s, …, xk Ak’s and multiplying by . The total number of such sequences is given by the number of ways we could lay out in a row n things, of which x1 are of one kind, x2 are of a second kind, …, xk are of a kth kind. The number of ways we could

choose x1 positions for the A1’s is ; after having put the A1’s in their position,

the number of ways we could choose positions for the A2’s is , and so

on. Thus, the total number of sequences with x1 A1’s, x2A2’s, …, xk Ak’s is given by

Thus, we obtain

3.47. Suppose that a fair die is rolled seven times. Find the probability that 1 and 2 dots appear twice each; 3, 4, and 5 dots once each; and 6 dots not at all.

Let (X1, X2, …, X6) be a six-dimensional random vector, where Xi denotes the number of times i dots appear in seven rolls of a fair die. Then (X1, X2, …, X6) is a multinomial r.v. with parameters (7, p1, p2, …, p6) where (i = 1, 2, …6). Hence, by Eq. (3.87),

3.48. Show that the pmf of a multinomial r.v. given by Eq. (3.87) satisfies the condition (3.68); that is,

where the summation is over the set of all nonnegative integers x1, x2, …, xk whose sum is n.

The multinomial theorem (which is an extension of the binomial theorem) states that

where x1 + x2 + … + xk = n and

is called the multinomial coefficient, and the summation is over the set of all nonnegative integers x1, x2, …, xk whose sum is n.

Thus, setting ai = pi in Eq. (3.102), we obtain

3.49. Let (X, Y) be a bivariate normal r.v. with its pdf given by Eq. (3.88). (a) Find the marginal pdf’s of X and Y. (b) Show that X and Y are independent when ρ = 0.

(a) By Eq. (3.30), the marginal pdf of X is

From Eqs. (3.88) and (3.89), we have

Rewriting q(x, y),

Comparing the integrand with Eq. (2.71), we see that the integrand is a normal pdf with mean μY + ρ (σY/σX) (x − μX) and variance (1 − ρ

2)σY 2. Thus,

the integral must be unity and we obtain

In a similar manner, the marginal pdf of Y is

(b) When ρ = 0, Eq. (3.88) reduces to

Hence, X and Y are independent.

3.50. Show that ρ in Eq. (3.88) is the correlation coefficient of X and Y.

By Eqs. (3.50) and (3.53), the correlation coefficient of X and Y is

where fXY(x, y) is given by Eq. (3.88). By making a change in variables v = (x − μX)/σX and w = (y − μY)/σY, we can write Eq. (3.105) as

The term in the curly braces is identified as the mean of V = N(ρw; 1 − σ2), and so

The last integral is the variance of W = N(0; 1), and so it is equal to 1 and we obtain ρXY = ρ.

3.51. Let (X, Y) be a bivariate normal r.v. with its pdf given by Eq. (3.88). Determine E(Y |x).

By Eq. (3.58),

where

Substituting Eqs. (3.88) and (3.103) into Eq. (3.107), and after some cancellation and rearranging, we obtain

which is equal to the pdf of a normal r.v. with mean μY + ρ(σY/σx)(x − μx) and variance . Thus, we get

Note that when X and Y are independent, then ρ = 0 and E(Y |x) = μY = E(Y).

3.52. The joint pdf of a bivariate r.v. (X, Y) is given by

(a) Find the means of X and Y. (b) Find the variances of X and Y. (c) Find the correlation coefficient of X and Y.

We note that the term in the bracket of the exponential is a quadratic function of x and y, and hence fXY(x, y) could be a pdf of a bivariate normal

r.v. If so, then it is simpler to solve equations for the various parameters. Now, the given joint pdf of (X, Y) can be expressed as

where

Comparing the above expressions with Eqs. (3.88) and (3.89), we see that fXY(x, y) is the pdf of a bivariate normal r.v. with μX = 0, μY = 1, and the following equations:

Solving for , and ρ, we get

Hence, (a) The mean of X is zero, and the mean of Y is 1. (b) The variance of both X and Y is 2. (c) The correlation coefficient of X and Y is .

3.53. Consider a bivariate r.v. (X, Y), where X and Y denote the horizontal and vertical miss distances, respectively, from a target when a bullet is fired. Assume that X and Y are independent and that the probability of the bullet landing on any point of the xy plane depends only on the distance of the point from the target. Show that (X, Y) is a bivariate normal r.v.

From the assumption, we have

for some function g. Differentiating Eq. (3.109) with respect to x, we have

Dividing Eq. (3.110) by Eq. (3.109) and rearranging, we get

Note that the left-hand side of Eq. (3.111) depends only on x, whereas the right- hand side depends only on x2 + y2; thus,

where c is a constant. Rewriting Eq.(3.112) as

and integrating both sides, we get

where a and k are constants. By the properties of a pdf, the constant c must be negative, and setting c = − 1/σ2, we have

Thus, by Eq. (2.71), X = N(0; σ2) and

In a similar way, we can obtain the pdf of Y as

Since X and Y are independent, the joint pdf of (X, Y) is

which indicates that (X, Y) is a bivariate normal r.v.

3.54. Let (X1, X2, …, Xn) be an n-variate normal r.v. with its joint pdf given by Eq. (3.92). Show that if the covariance of Xi and Xj is zero for i ≠ j, that is,

then X1, X2, …, Xn are independent.

From Eq. (3.94) with Eq. (3.114), the covariance matrix K becomes

It therefore follows that

and

Then we can write

Substituting Eqs. (3.116) and (3.118) into Eq. (3.92), we obtain

Now Eq. (3.119) can be rewritten as

where

Thus we conclude that X1, X2,…, Xn are independent.

SUPPLEMENTARY PROBLEMS

3.55. Consider an experiment of tossing a fair coin three times. Let (X, Y) be a bivariate r.v., where X denotes the number of heads on the first two tosses and Y denotes the number of heads on the third toss. (a) Find the range of X. (b) Find the range of Y. (c) Find the range of (X, Y). (d) Find (i) P(X ≤ 2, Y ≤ 1); (ii) P (X ≤ 1, Y ≤ 1); and (iii) P(X ≤ 0, Y ≤ 0).

3.56. Let FXY(x, y) be a joint cdf of a bivariate r.v. (X, Y). Show that

where FX(x) and FY(y) are marginal cdf’s of X and Y, respectively.

3.57. Let the joint pmf of (X, Y) be given by

where k is a constant. (a) Find the value of k. (b) Find the marginal pmf’s of X and Y.

3.58. The joint pdf of (X, Y) is given by

where k is a constant. (a) Find the value of k. (b) Find P(X > 1, Y < 1), P(X < Y), and P(X ≤ 2).

3.59. Let (X, Y) be a bivariate r.v., where X is a uniform r.v. over (0, 0.2) and Y is an exponential r.v. with parameter 5, and X and Y are independent. (a) Find the joint pdf of (X, Y). (b) Find P(Y ≤ X).

3.60. Let the joint pdf of (X, Y) be given by

(a) Show that fXY(x, y) satisfies Eq.(3.26). (b) Find the marginal pdf’s of X and Y.

3.61. The joint pdf of (X, Y) is given by

where k is a constant. (a) Find the value of k. (b) Find the marginal pdf’s of X and Y.

3.62. The joint pdf of (X, Y) is given by

(a) Find the marginal pdf’s of X and Y. (b) Are X and Y independent?

3.63. The joint pdf of (X, Y) is given by

(a) Are X and Y independent ? (b) Find the conditional pdf’s of X and Y.

3.64. The joint pdf of (X, Y) is given by

(a) Find the conditional pdf’s of Y, given that X = x. (b) Find the conditional cdf’s of Y, given that X = x.

3.65. Consider the bivariate r.v. (X, Y) of Prob. 3.14. (a) Find the mean and the variance of X. (b) Find the mean and the variance of Y. (c) Find the covariance of X and Y. (d) Find the correlation coefficient of X and Y.

3.66. Consider a bivariate r.v. (X, Y) with joint pdf

Find P[(X, Y) x2 + y2 ≤ a2].

3.67. Let (X, Y) be a bivariate normal r.v., where X and Y each have zero mean and variance σ2, and the correlation coefficient of X and Y is ρ. Find the joint pdf of (X, Y).

3.68. The joint pdf of a bivariate r.v. (X, Y) is given by

(a) Find the means and variances of X and Y. (b) Find the correlation coefficient of X and Y.

3.69. Let (X, Y, Z) be a trivariate r.v., where X, Y, and Z are independent and each has a uniform distribution over (0, 1). Compute P(X ≥ Y ≥ Z).

ANSWERS TO SUPPLEMENTARY PROBLEMS

3.55.

3.56. Hint: Set x1 = a, y1 = c, and x2 = y2 = ∞ in Eq. (3.95) and use Eqs. (3.13) and (3.14).

3.57.

3.58.

3.59.

3.60.

3.61.

3.62.

3.63.

3.64.

3.65.

3.66. 1 − e−a 2/(2σ2)

3.67.

3.68.

3.69.

CHAPTER 4

Functions of Random Variables, Expectation, Limit Theorems

4.1 Introduction

In this chapter we study a few basic concepts of functions of random variables and investigate the expected value of a certain function of a random variable. The techniques of moment generating functions and characteristic functions, which are very useful in some applications, are presented. Finally, the laws of large numbers and the central limit theorem, which is one of the most remarkable results in probability theory, are discussed.

4.2 Functions of One Random Variable

A. Random Variable g(X):

Given a r.v. X and a function g(x), the expression

defines a new r.v. Y. With y a given number, we denote DY the subset of Rx (range of X) such that g(x) ≤ y. Then

where (X ∈ DY) is the event consisting of all outcomes ζ such that the point X(ζ) ∈ DY. Hence,

If X is a continuous r.v. with pdf ƒX(x), then

B. Determination of ƒY(y) from ƒX(x):

Let X be a continuous r.v. with pdf ƒX(x). If the transformation y = g(x) is one-to- one and has the inverse transformation

then the pdf of Y is given by (Prob. 4.2)

Note that if g(x) is a continuous monotonic increasing or decreasing function, then the transformation y = g(x) is one-to-one. If the transformation y = g(x) is not one-to-one, ƒY(Y) is obtained as follows: Denoting the real roots of y = g(x) by xk, that is,

then

where g′(x) is the derivative of g(x).

4.3 Functions of Two Random Variables

A. One Function of Two Random Variables:

Given two r.v.’s X and Y and a function g(x, y), the expression

defines a new r.v. Z. With z a given number, we denote DZ the subset of RXY [range of (X, Y)] such that g(x, y) ≤ z. Then

where {(X, Y) ∈ DZ} is the event consisting of all outcomes ζ such that point {X(ζ), Y(ζ)} ∈ DZ. Hence,

If X and Y are continuous r.v.’s with joint pdf ƒXY(x, y), then

B. Two Functions of Two Random Variables: Given two r.v.’s X and Y and two functions g(x, y) and h(x, y), the expression

defines two new r.v.’s Z and W. With z and w two given numbers, we denote DZW the subset of RXY [range of (X, Y)] such that g(x, y) ≤ z and h(x, y) ≤ w. Then

where {(X, Y) ∈ DZW} is the event consisting of all outcomes ζ such that point {X(ζ), Y(ζ)} ∈ DZW. Hence,

In the continuous case, we have

Determination of ƒZW(z, w) from ƒXY(x, y): Let X and Y be two continuous r.v.’s with joint pdf ƒXY(x, y). If the transformation

is one-to-one and has the inverse transformation

then the joint pdf of Z and W is given by

where x = q(z, w), y = r(z, w), and

which is the Jacobian of the transformation (4.17). If we define

then

and Eq. (4.19) can be expressed as

4.4 Functions of n Random Variables

A. One Function of n Random Variables:

Given n r.v.’s X1, …, Xn and a function g(x1, …, xn), the expression

defines a new r.v. Y. Then

and

where DY is the subset of the range of (X1, …, Xn) such that g(x1, …, xn) ≤ y. If X1, …, Xn are continuous r.v.’s with joint pdf ƒX1, …, xn(x1, …, xn), then

B. n Functions of n Random Variables: When the joint pdf of n r.v.’s X1, …, Xn is given and we want to determine the joint pdf of n r.v.’s Y1, …, Yn, where

the approach is the same as for two r.v.’ s. We shall assume that the transformation

is one-to-one and has the inverse transformation

Then the joint pdf of Y1, …, Yn, is given by

where

which is the Jacobian of the transformation (4.29).

4.5 Expectation

A. Expectation of a Function of One Random Variable:

The expectation of Y = g(X) is given by

B. Expectation of a Function of More Than One Random Variable: Let X1, …, Xn be n r.v.’s, and let Y = g(X1, …, Xn). Then

C. Linearity Property of Expectation: Note that the expectation operation is linear (Prob. 4.45), and we have

where ai’s are constants. If r.v.’s X and Y are independent, then we have (Prob. 4.47)

The relation (4.36) can be generalized to a mutually independent set of n r.v.’s X1, …, Xn:

D. Conditional Expectation as a Random Variable: In Sec. 3.8 we defined the conditional expectation of Y given X = x, E(Y |x) [Eq. (3.58)], which is, in general, a function of x, say H(x). Now H(X) is a function of the r.v. X; that is,

Thus, E(Y | X) is a function of the r.v. X. Note that E(Y | X) has the following property (Prob. 4.38):

E. Jensen’s Inequality: A twice differentiable real-valued function g(x) is said to be convex if g″(x) ≥ 0 for all x; similarly, it is said to be concave if g″(x) ≤ 0. Examples of convex functions include x2, x < ex, x log x (x ≥ 0), and so on. Examples of concave functions include log x and . If g(x) is convex, then h(x) = − g(x) is concave and vice versa. Jensen’s Inequality: If g(x) is a convex function, then

provided that the expectations exist and are finite. Equation (4.40) is known as Jensen’s inequality (for proof see Prob. 4.50).

F. Cauchy-Schwarz Inequality:

Assume that E(X2), E(Y2) < ∞, then

Equation (4.41) is known as Cauchy-Schwarz inequality (for proof see Prob. 4.51).

4.6 Probability Generating Functions

A. Definition:

Let X be a nonnegative integer-valued discrete r.v. with pmf pX(x). The probability generating function (or z-transform) of X is defined by

where z is a variable. Note that

B. Properties of GX(z):

Differentiating Eq. (4.42) repeatedly, we have

In general,

Then, we have

One of the useful properties of the probability generating function is that it turns a sum into product.

Suppose that X1, X2, …, Xn are independent nonnegative integer-valued r.v.’s, and let Y = X1 + X2 + … + Xn. Then

Note that property (2) indicates that the probability generating function determines the distribution. Property (4) is known as the nth factorial moment.

Setting n = 2 in Eq. (4.50) and by Eq. (4.35), we have

Thus, using Eq. (4.49)

Using Eq. (2.31), we obtain

C. Lemma for Probability Generating Function: Lemma 4.1: If two nonnegative integer-valued discrete r.v.’s have the same probability generating functions, then they must have the same distribution.

4.7 Moment Generating Functions

A. Definition:

The moment generating function of a r.v. X is defined by

where t is a real variable. Note that MX(t) may not exist for all r.v.’s X. In general, MX(t) will exist only for those values of t for which the sum or integral of Eq. (4.56) converges absolutely. Suppose that MX(t) exists. If we express e

tX formally and take expectation, then

and the kth moment of X is given by

where

Note that by substituting z by et, we can obtain the moment generating function for a nonnegative integer-valued discrete r.v. from the probability generating function.

B. Joint Moment Generating Function: The joint moment generating function MXY(t1, t2) of two r.v.’s X and Y is defined by

where t1 and t2 are real variables. Proceeding as we did in Eq. (4.57), we can establish that

and the (k, n) joint moment of X and Y is given by

where

In a similar fashion, we can define the joint moment generating function of n r.v.’s X1, …, Xn by

from which the various moments can be computed. If X1, …, Xn are independent, then

C. Lemmas for Moment Generating Functions: Two important lemmas concerning moment generating functions are stated in the following:

Lemma 4.2: If two r.v.’s have the same moment generating functions, then they must have the same distribution. Lemma 4.3: Given cdf’s F(x), F1(x), F2(x), … with corresponding moment generating functions M(t), M1(t), M2(t), …, then Fn(x) → F(x) if Mn(t) → M(t).

4.8 Characteristic Functions

A. Definition:

The characteristic function of a r.v. X is defined by

where ω is a real variable and . Note that ΨX(ω) is obtained by replacing t in MX(t) by jω if MX(t) exists. Thus, the characteristic function has all the properties of the moment generating function. Now

for the discrete case and

for the continuous case. Thus, the characteristic function ΨX(ω) is always defined even if the moment generating function MX(t) is not (Prob. 4.76). Note that ΨX(ω) of Eq. (4.66) for the continuous case is the Fourier transform (with the sign of j reversed) of ƒX(x). Because of this fact, if ΨX(ω) is known, ƒX(x) can be found from the inverse Fourier transform; that is,

B. Joint Characteristic Functions: The joint characteristic function ΨXY(ω1, ω2) of two r.v.’s X and Y is defined by

where ω1 and ω2 are real variables. The expression of Eq. (4.68) for the continuous case is recognized as the two-

dimensional Fourier transform (with the sign of j reversed) of ƒXY(x, y). Thus, from the inverse Fourier transform, we have

From Eqs. (4.66) and (4.68), we see that

which are called marginal characteristic functions. Similarly, we can define the joint characteristic function of n r.v.’s X1, …, Xn by

As in the case of the moment generating function, if X1, …, Xn are independent, then

C. Lemmas for Characteristic Functions: As with the moment generating function, we have the following two lemmas:

Lemma 4.4: A distribution function is uniquely determined by its characteristic function. Lemma 4.5: Given cdf’s F(x), F1(x), F2(x), … with corresponding characteristic functions Ψ(ω), Ψ1(ω), Ψ2(ω), …, then Fn(x) → F(x) at points of continuity of F(x) if and only if Ψn(ω) → Ψ(ω) for every ω.

4.9 The Laws of Large Numbers and the Central Limit Theorem

A. The Weak Law of Large Numbers:

Let X1, …, Xn be a sequence of independent, identically distributed r.v.’s each with a finite mean E(Xi) = μ. Let

Then, for any ε > 0,

Equation (4.74) is known as the weak law of large numbers, and is known as the sample mean.

B. The Strong Law of Large Numbers: Let X1, …, Xn be a sequence of independent, identically distributed r.v.’s each with a finite mean E(Xi) = μ. Then, for any ε > 0,

where is the sample mean defined by Eq. (4.73). Equation (4.75) is known as the strong law of large numbers.

Notice the important difference between Eqs. (4.74) and (4.75). Equation (4.74) tells us how a sequence of probabilities converges, and Eq. (4.75) tells us how the sequence of r.v.’s behaves in the limit. The strong law of large numbers tells us that the sequence is converging to the constant μ.

C. The Central Limit Theorem: The central limit theorem is one of the most remarkable results in probability theory. There are many versions of this theorem. In its simplest form, the central limit theorem is stated as follows:

Let X1, …, Xn be a sequence of independent, identically distributed r.v.’s each with mean μ and variance σ2. Let

where is defined by Eq. (4.73). Then the distribution of Zn tends to the standard normal as n → ∞; that is,

or

where Φ (z) is the cdf of a standard normal r.v. [Eq. (2.73)]. Thus, the central limit theorem tells us that for large n, the distribution of the sum Sn = X1 + … + Xn is approximately normal regardless of the form of the distribution of the individual Xi’s. Notice how much stronger this theorem is than the laws of large numbers. In practice, whenever an observed r.v. is known to be a sum of a large number of r.v.’s, then the central limit theorem gives us some justification for assuming that this sum is normally distributed.

SOLVED PROBLEMS

Functions of One Random Variable

4.1. If X is N(μ; σ2), then show that Z = (X − μ)/σ is a standard normal r.v.; that is, N(0; 1).

The cdf of Z is

By the change of variable y = (x − μ)/σ (that is, x = σy + μ), we obtain

and

which indicates that Z = N(0; 1).

4.2. Verify Eq. (4.6).

Assume that y = g(x) is a continuous monotonically increasing function [Fig. 4-1(a)]. Since y = g(x) is monotonically increasing, it has an inverse that we denote by x = g−1(y) = h(y). Then

Fig. 4-1

and

Applying the chain rule of differentiation to this expression yields

which can be written as

If y = g(x) is monotonically decreasing [Fig. 4-1(b)], then

Thus,

In Eq. (4.82), since y = g(x) is monotonically decreasing, dy/dx (and dx/dy) is negative. Combining Eqs. (4.80) and (4.82), we obtain

which is valid for any continuous monotonic (increasing or decreasing) function y = g(x).

4.3. Let X be a r.v. with cdf FX(x) and pdf ƒX(x). Let Y = aX + b, where a and b are real constants and a ≠ 0. (a) Find the cdf of Y in terms of FX(x). (b) Find the pdf of Y in terms of ƒX(x).

(a) If a > 0, then [Fig. 4-2(a)]

Fig. 4-2

If a < 0, then [Fig. 4-2(b)]

Note that if X is continuous, then P[X = (y − b)/a] = 0, and

(b) From Fig. 4-2, we see that y = g(x) = ax + b is a continuous monotonically increasing (a > 0) or decreasing (a < 0) function. Its inverse is x = g−1(y) = h(y) = (y − b)/a, and dx/dy = 1/a. Thus, by Eq. (4.6),

Note that Eq. (4.86) can also be obtained by differentiating Eqs. (4.83) and (4.85) with respect to y.

4.4. Let Y = aX + b. Determine the pdf of Y, if X is a uniform r.v. over (0, 1).

The pdf of X is [Eq. (2.56)]

Then by Eq. (4.86), we get

The range RY is found as follows: From Fig. 4-3, we see that

Fig. 4-3

4.5. Let Y = aX + b. Show that if X = N(μ; σ2), then Y = N(aμ + b; a2 σ2), and find the values of a and b so that Y = N(0; 1).

Since X = N(μ; σ2), by Eq. (2.71),

Hence, by Eq. (4.86),

which is the pdf of N(aμ + b; a2σ2). Hence, Y = N(aμ + b; a2σ2). Next, let aμ + b = 0 and a2σ2 = 1, from which we get a = 1/σ and b = −μ/σ. Thus, Y = (X − μ)/σ is N(0; 1) (see Prob. 4.1).

4.6. Let X be a r.v. with pdf ƒX(x). Let Y = X 2. Find the pdf of Y.

The event A = (Y ≤ y) in RY is equivalent to the event in RX (Fig. 4-4). If y ≤ 0, then

Fig. 4-4

and ƒY(y) = 0. If y > 0, then

and

Thus,

Alternative Solution:

If y < 0, then the equation y = x2 has no real solutions; hence, ƒY(y) = 0. If y > 0, then y = x2 has two solutions, and . Now, y = g(x) = x2 and g′(x) = 2x. Hence, by Eq. (4.8),

4.7. Let Y = X2. Find the pdf of Y if X = N(0; 1).

Since X = N(0; 1)

Since ƒX(x) is an even function, by Eq. (4.90), we obtain

4.8. Let Y = X2. Find and sketch the pdf of Y if X is a uniform r.v. over (−1, 2).

The pdf of X is [Eq. (2.56)] [Fig. 4-5(a)]

Fig. 4-5

In this case, the range of Y is (0, 4), and we must be careful in applying Eq. (4.90). When 0 < y < 1, both and are in RX = (−1, 2), and by Eq. (4.90),

When is in RX = (−1, 2) but , and by Eq. (4.90),

which is sketched in Fig. 4-5(b).

4.9. Let Y = eX. Find the pdf of Y if X is a uniform r.v. over (0, 1).

The pdf of X is

The cdf of Y is

Thus,

Alternative Solution:

The function y = g(x) = ex is a continuous monotonically increasing function. Its inverse is x = g−1(y) = h(y) = ln y. Thus, by Eq. (4.6), we obtain

or

4.10. Let Y = eX. Find the pdf of Y if X = N (μ; σ2).

The pdf of X is [Eq. (2.71)]

Thus, using the technique shown in the alternative solution of Prob. 4.9, we obtain

Note that X = ln Y is the normal r.v.; hence, the r.v. Y is called the log-normal r.v.

4.11. Let X be a r.v. with pdf ƒX (x). Let Y = 1/X. Find the pdf of Y in terms of ƒX(x).

We see that the inverse of y = 1/x is x = 1/y and dx/dy = −1/y2. Thus, by Eq. (4.6)

4.12. Let Y = 1/X and X be a Cauchy r.v. with parameter a. Show that Y is also a Cauchy r.v. with parameter 1/a.

From Prob. 2.83, we have

By Eq. (4.94)

which indicates that Y is also a Cauchy r.v. with parameter 1/a.

4.13. Let Y = tan X. Find the pdf of Y if X is a uniform r.v. over (−π/2, π/2).

The cdf of X is [Eq. (2.57)]

Now

Then the pdf of Y is given by

Note that the r.v. Y is a Cauchy r.v. with parameter 1.

4.14. Let X be a continuous r.v. with the cdf FX(x). Let Y = FX(X). Show that Y is a uniform r.v. over (0, 1).

Notice from the properties of a cdf that y = FX(x) is a monotonically nondecreasing function. Since 0 ≤ FX(x) ≤ 1 for all real x, y takes on values only on the interval (0, 1). Using Eq. (4.80) (Prob. 4.2), we have

Hence, Y is a uniform r.v. over (0, 1).

4.15. Let Y be a uniform r.v. over (0, 1). Let F(x) be a function which has the properties of the cdf of a continuous r.v. with F(a) = 0, F(b) = 1, and F(x) strictly increasing for a < x < b, where a and b could be 2∞ and ∞, respectively. Let X = F−1(Y). Show that the cdf of X is F(x).

Since F(x) is strictly increasing, F−1(Y) ≤ x is equivalent to Y ≤ F(x), and hence

Now Y is a uniform r.v. over (0, 1), and by Eq. (2.57),

and accordingly,

Note that this problem is the converse of Prob. 4.14.

4.16. Let X be a continuous r.v. with the pdf

Find the transformation Y = g(X) such that the pdf of Y is

The cdf of X is

Then from the result of Prob. 4.14, the r.v. Z = 1 −e−X is uniformly distributed over (0, 1). Similarly, the cdf of Y is

and the r.v. is uniformly distributed over (0, 1). Thus, by setting Z = W, the required transformation is Y = (1 − e−X)2.

Functions of Two Random Variables

4.17. Consider Z = X + Y. Show that if X and Y are independent Poisson r.v.’s with parameters λ1 and λ2, respectively, then Z is also a Poisson r.v. with parameter λ1 + λ2.

We can write the event

where events (X = i, Y = n − i), i = 0, 1, …, n, are disjoint. Since X and Y are independent, by Eqs. (1.62) and (2.48), we have

which indicates that Z = X + Y is a Poisson r.v. with λ1 + λ2.

4.18. Consider two r.v.’s X and Y with joint pdf ƒXY (x, y). Let Z = X + Y.

(a) Determine the pdf of Z. (b) Determine the pdf of Z if X and Y are independent.

(a) The range RZ of Z corresponding to the event (Z ≤ z) = (X + Y ≤ z) is the set of points (x, y) which lie on and to the left of the line z = x + y (Fig. 4-6). Thus, we have

Fig. 4-6

Then

(b) If X and Y are independent, then Eq. (4.96) reduces to

The integral on the right-hand side of Eq. (4.97) is known as a convolution of ƒX(z) and ƒY(z). Since the convolution is commutative, Eq. (4.97) can also be written as

4.19. Using Eqs. (4.19) and (3.30), redo Prob. 4.18(a); that is, find the pdf of Z = X + Y.

Let Z = X + Y and W = X. The transformation z = x + y, w = x has the inverse transformation x = w, y = z − w, and

By Eq. (4.19), we obtain

Hence, by Eq. (3.30), we get

4.20. Suppose that X and Y are independent standard normal r.v.’s. Find the pdf of Z = X + Y.

The pdf’s of X and Y are

Then, by Eq. (4.97), we have

Now, , and we have

with the change of variables . Since the integrand is the pdf of N(0; 1), the integral is equal to unity, and we get

which is the pdf of N(0; 2). Thus, Z is a normal r.v. with zero mean and variance 2.

4.21. Let X and Y be independent uniform r.v.’s over (0, 1). Find and sketch the pdf of Z = X + Y.

Since X and Y are independent, we have

The range of Z is (0, 2), and

If 0 < z < 1 [Fig. 4-7(a)],

Fig. 4-7

and

If 1 < z < 2 [Fig. 4-7(b)],

which is sketched in Fig. 4-7(c). Note that the same result can be obtained by the convolution of ƒX(z) and ƒY(z).

4.22. Let X and Y be independent exponential r.v.’s with common parameter λ and let Z = X + Y. Find ƒZ(z).

From Eq. (2.60) we have

In order to apply Eq. (4.97) we need to rewrite ƒX(x) and ƒY(y) as

where u(ξ) is a unit step function defined as

Now by Eq. (4.97) we have

Using Eq. (4.99), we have

Thus,

Note that X and Y are gamma r.v.’s with parameter (1, λ) and Z is a gamma r.v. with parameter (2, λ) (see Prob. 4.23).

4.23. Let X and Y be independent gamma r.v.’s with respective parameters (α, λ) and (β, λ). Show that Z = X + Y is also a gamma r.v. with parameters (α + β, λ).

From Eq. (2.65),

The range of Z is (0, ∞), and using Eq. (4.97), we have

By the change of variable w = x/z, we have

where k is a constant which does not depend on z. The value of k is determined as follows: Using Eq. (2.22) and definition (2.66) of the gamma function, we have

Hence, k = λα+β/Γ(α + β) and

which indicates that Z is a gamma r.v. with parameters (α + β, λ).

4.24. Let X and Y be two r.v.’s with joint pdf ƒX Y (x, y). and let Z = X − Y.

(a) Find ƒZ (z). (b) Find ƒZ(z) if X and Y are independent.

(a) From Eq. (4.12) and Fig. 4-8 we have

Fig. 4-8

Then

(b) If X and Y are independent, by Eq. (3.32), Eq. (4.100) reduces to

which is the convolution of ƒX(−z) with ƒY (z).

4.25. Consider two r.v.’s X and Y with joint pdf ƒXY(x, y). Determine the pdf of Z = XY.

Let Z = XY and W = X. The transformation z = xy, w = x has the inverse transformation x = w, y = z/w, and

Thus, by Eq. (4.23), we obtain

and the marginal pdf of Z is

4.26. Let X and Y be independent uniform r.v.’s over (0, 1). Find the pdf of Z = XY.

We have

The range of Z is (0, 1). Then

or

By Eq. (4.103),

Thus,

4.27. Consider two r.v.’s X and Y with joint pdf ƒXY(x, y). Determine the pdf of Z = X/Y.

Let Z = X/Y and W = Y. The transformation z = x/y, w = y has the inverse transformation x = zw, y = w, and

Thus, by Eq. (4.23), we obtain

and the marginal pdf of Z is

4.28. Let X and Y be independent standard normal r.v.’s. Find the pdf of Z = X/Y.

Since X and Y are independent, using Eq. (4.105), we have

which is the pdf of a Cauchy r.v. with parameter 1.

4.29. Let X and Y be two r.v.’s with joint pdf ƒXY (x, y) and joint cdf FXY(x, y). Let Z = max(X, Y).

(a) Find the cdf of Z. (b) Find the pdf of Z if X and Y are independent.

(a) The region in the xy plane corresponding to the event {max(X, Y) ≤ z} is shown as the shaded area in Fig. 4-9. Then

Fig. 4-9

(b) If X and Y are independent, then

and differentiating with respect to z gives

4.30. Let X and Y be two r.v.’s with joint pdf ƒXY(x, y) and joint cdf FXY(x, y). Let W = min(X, Y).

(a) Find the cdf of W. (b) Find the pdf of W if X and Y are independent.

(a) The region in the xy plane corresponding to the event {min(X, Y) ≤ w} is shown as the shaded area in Fig. 4-10. Then

Fig. 4-10

Thus,

(b) If X and Y are independent, then

and differentiating with respect to w gives

4.31. Let X and Y be two r.v.’s with joint pdf ƒXY (x, y). Let Z = X 2 + Y2. Find

ƒz(z).

As shown in Fig. 4-11, DZ (X 2 + Y2 ≤ z) represents the area of a circle with

radius .

Fig. 4-11

Hence, by Eq. (4.12)

and

4.32. Let X and Y be independent normal r.v.’s with μX = μY = 0 and . Let Z = X2 + Y2. Find ƒZ(z).

Since X and Y are independent, from Eqs. (3.32) and (2.71) we have

and

Thus, using Eq. (4.110), we obtain

Let sin θ. Then

and

Hence,

which indicates that Z is an exponential r.v. with parameter 1/(2 σ2).

4.33. Let X and Y be two r.v.’s with joint pdf fXY (x, y). Let

Find fRΘ (r, θ) in terms of fXY (x, y).

We assume that r ≥ 0 and 0 ≤ θ < 2π. With this assumption, the transformation

has the inverse transformation

Since

by Eq. (4.23) we obtain

4.34. A voltage V is a function of time t and is given by

in which ω is a constant angular frequency and X = Y = N (0; σ2) and they are independent.

(a) Show that V(t) may be written as

(b) Find the pdf’s of r.v.’s R and Θ and show that R and Θ are independent.

(a) We have

Where which is the transformation (4.113).

(b) Since X = Y = N (0; σ2) and they are independent, we have

Thus, using Eq. (4.114), we get

Now

and ƒRΘ(r, θ) = ƒR(r) ƒΘ(θ); hence, R and Θ are independent. Note that R is a Rayleigh r.v. (Prob. 2.26), and Θ is a uniform r.v. over

(0, 2π).

Functions of N Random Variables

4.35. Let X, Y, and Z be independent standard normal r.v.’s. Let W = (X2 + Y2 + Z2)1/2. Find the pdf of W.

We have

where RW = {(x, y, z): x 2 + y2 + z2 ≤ w2). Using spherical coordinates (Fig. 4-

12), we have

Fig. 4-12 Spherical coordinates.

and

Thus, the pdf of W is

4.36. Let X1, …, Xn be n independent r.v.’s each with the identical pdf ƒ(x). Let Z = max(X1, …, Xn). Find the pdf of Z.

The probability P(z < Z < z + dz) is equal to the probability that one of the r.v.’s falls in (z, z + dz) and all others are less than z. The probability that one of Xi (i = 1, …, n) falls in (z, z + dz) and all others are all less than z is

Since there are n ways of choosing the variables to be maximum, we have

When n = 2, Eq. (4.122) reduces to

which is the same as Eq. (4.107) (Prob. 4.29) with ƒX(z) = ƒY(z) = ƒ(z) and FX(z) = FY(z) = F(z).

4.37. Let X1, …, Xn be n independent r.v.’s each with the identical pdf ƒ(x). Let W 5 min(X1, …, Xn). Find the pdf of W.

The probability P(w < W < w + dw) is equal to the probability that one of the r.v.’s falls in (w, w + dw) and all others are greater than w. The probability that one of Xi (i = 1, …, n) falls in (w, w + dw) and all others are greater than w is

Since there are n ways of choosing the variables to be minimum, we have

When n = 2, Eq. (4.124) reduces to

which is the same as Eq. (4.109) (Prob. 4.30) with ƒX(w) = ƒY(w) = ƒ(w) and FX(w) = FY(w) = F(w).

4.38. Let Xi, i = 1, …, n, be n independent gamma r.v.’s with respective parameters (αi, λ), i = 1, …, n. Let

Show that Y is also a gamma r.v. with parameters .

We prove this proposition by induction. Let us assume that the proposition is true for n = k; that is,

is a gamma r.v. with parameters

Let

Then, by the result of Prob. 4.23, we see that W is a gamma r.v. with parameters . Hence, the proposition is true for n = k + 1. Next, by the result of Prob. 4.23, the proposition is true for n = 2. Thus, we conclude that the proposition is true for any n ≥ 2.

4.39. Let X1, …, Xn be n independent exponential r.v.’s each with parameter λ. Let

Show that Y is a gamma r.v. with parameters (n, λ).

We note that an exponential r.v. with parameter λ is a gamma r.v. with parameters (1, λ). Thus, from the result of Prob. 4.38 and setting αi = 1, we conclude that Y is a gamma r.v. with parameters (n, λ).

4.40. Let Z1, …, Zn be n independent standard normal r.v.’s. Let

Find the pdf of Y.

Let . Then by Eq. (4.91) (Prob. 4.7), the pdf of Yi is

Now, using Eq. (2.96), we can rewrite

and we recognize the above as the pdf of a gamma r.v. with parameters [Eq. (2.65)]. Thus, by the result of Prob. 4.38, we conclude that Y is the gamma r.v. with parameters and

When n is an even integer, Γ(n/2) = [(n/2) − 1]!, whereas when n is odd, Γ(n/2) can be obtained from Γ(α) = (α − 1) Γ(α − 1) [Eq. (2.97)] and

[Eq. (2.99)].

Note that Equation (4.126) is referred to as the chi-square (χ2) density function with n degrees of freedom, and Y is known as the chi-square (χ2) r.v. with n degrees of freedom. It is important to recognize that the sum of the squares of n independent standard normal r.v.’s is a chi-square r.v. with n degrees of freedom. The chi-square distribution plays an important role in statistical analysis.

4.41. Let X1, X2, and X3 be independent standard normal r.v.’s. Let

Determine the joint pdf of Y1, Y2, and Y3.

Let

By Eq. (4.32), the Jacobian of transformation (4.127) is

Thus, solving the system (4.127), we get

Then by Eq. (4.31), we obtain

Since X1, X2, and X3 are independent,

Hence,

where

Expectation

4.42. Let X be a uniform r.v. over (0, 1) and Y = eX. (a) Find E(Y) by using ƒY(y). (b) Find E(Y) by using ƒX(x).

(a) From Eq. (4.92) (Prob. 4.9),

Hence, (b) The pdf of X is

Then, by Eq. (4.33),

4.43. Let Y = aX + b, where a and b are constants. Show that

(a)

(b)

We verify for the continuous case. The proof for the discrete case is similar. (a) By Eq. (4.33),

(b) Using Eq. (4.129), we have

4.44. Verify Eq. (4.39).

Using Eqs. (3.58) and (3.38), we have

4.45. Let Z = aX + bY, where a and b are constants. Show that

We verify for the continuous case. The proof for the discrete case is similar.

Note that Eq. (4.131) (the linearity of E) can be easily extended to n r.v.’s:

4.46. Let Y = aX + b. (a) Find the covariance of X and Y. (b) Find the correlation coefficient of X and Y.

(a) By Eq. (4.131), we have

Thus, the covariance of X and Y is [Eq. (3.51)]

(b) By Eq. (4.130), we have σY = |a| σX. Thus, the correlation coefficient of X and Y is [Eq. (3.53)]

4.47. Verify Eq. (4.36).

Since X and Y are independent, we have

The proof for the discrete case is similar.

4.48. Let X and Y be defined by

where Θ is a random variable uniformly distributed over (0, 2π). (a) Show that X and Y are uncorrelated. (b) Show that X and Y are not independent.

(a) We have

Similarly,

Thus, by Eq. (3.52), X and Y are uncorrelated. (b)

Hence,

If X and Y were independent, then by Eq. (4.36), we would have E(X2Y2) = E(X2)E(Y2). Therefore, X and Y are not independent.

4.49. Let X1, …, Xn be n r.v.’s. Show that

If X1, …, Xn are pairwise independent, then

Let

Then by Eq. (4.132), we have

If X1, …, Xn are pairwise independent, then (Prob. 3.22)

and Eq. (4.135) reduces to

4.50. Verify Jensen’s inequality (4.40),

Expanding g(x) in a Taylor’s series expansion around μ = E(x), we have

where ξ is some value between x and μ. If g(x) is convex, then g″(ζ) ≥ 0 and we obtain

Hence,

Taking expectation, we get

4.51. Verify Cauchy-Schwarz inequality (4.41),

We have

The discriminant of the quadratic in α appearing in Eq. (4.138) must be nonpositive because the quadratic cannot have two distinct real roots. Therefore,

and we obtain

or Probability Generating Functions

4.52. Let X be a Bernoulli r.v. with parameter p. (a) Find the probability generating function GX(z) of X. (b) Find the mean and variance of X.

(a) From Eq. (2.32)

By Eq. (4.42)

(b) Differentiating Eq. (4.139), we have

Using Eqs. (4.49) and (4.55), we obtain

4.53. Let X be a binomial r.v. with parameters (n, p). (a) Find the probability generating function GX(z) of X. (b) Find P(X = 0) and P(X = 1). (c) Find the mean and variance of X.

(a) From Eq. (2.36)

By Eq. (4.41)

(b) From Eqs. (4.47) and (4.140)

Differentiating Eq. (4.140), we have

Then from Eq. (4.48)

(c) Differentiating Eq. (4.141) again, we have

Thus, using Eqs. (4.49) and (4.53), we obtain

4.54. Let X1, X2, …, Xn be independent Bernoulli r.v.’s with the same parameter p, and let Y = X1 + X2 + … + Xn. Show that Y is a binomial r.v. with parameters (n, p).

By Eq. (4.139)

Now applying property 5 Eq. (4.52), we have

Comparing with Eq. (4.140), we conclude that Y is a binomial r.v. with parameters (n, p).

4.55. Let X be a geometric r.v. with parameter p. (a) Find the probability generating function GX(z) of X. (b) Find the mean and variance of X.

(a) From Eq. (2.40) we have

Then by Eq. (4.42)

Thus,

(b) Differentiating Eq. (4.143), we have

Thus, using Eqs. (4.49) and (4.55), we obtain

4.56. Let X be a negative binomial r.v. with parameters p and k. (a) Find the probability generating function GX(z) of X. (b) Find the mean and variance of X.

(a) From Eq. (2.45)

By Eq. (4.42)

(b) Differentiating Eq. (4.146), we have

Then,

Thus, by Eqs. (4.49) and (4.55), we obtain

4.57. Let X1, X2, …, Xn be independent geometric r.v.’s with the same parameter p, and let Y = X1+X2+…+ Xk. Show that Y is a negative binomial r.v. with parameters p and k.

By Eq. (4.143) (Prob.4.55)

Now applying Eq. (4.52), we have

Comparing with Eq. (4.146) (Prob. 4.56), we conclude that Y is a negative binomial r.v. with parameters p and k.

4.58. Let X be a Poisson r.v. with parameter λ. (a) Find the probability generating function of GX (z) of X. (b) Find the mean and variance of X.

(a) From Eq. (2.48)

By Eq. (4.42)

(b) Differentiating Eq. (4.147), we have

Then,

Thus, by Eq. (4.49) and Eq. (4.55), we obtain

4.59. Let X1, X2, …, Xn be independent Poisson r.v.’s with the same parameter ρ and let Y = X1 + X2 + … + Xn. Show that Y is a Poisson r.v. with parameter nλ.

Using Eq. (4.147) and Eq. (4.52), we have

which indicates that Y is a Poisson r.v. with parameter nλ.

Moment Generating Functions

4.60. Let the moment of a discrete r.v. X be given by

(a) Find the moment generating function of X. (b) Find P(X = 0) and P(X = 1).

(a) By Eq. (4.57), the moment generating function of X is

(b) By definition (4.56),

Thus, equating Eqs. (4.149) and (4.150), we obtain

4.61. Let X be a Bernoulli r.v. (a) Find the moment generating function of X. (b) Find the mean and variance of X.

(a) By definition (4.56) and Eq. (2.32),

which can also be obtained by substituting z by et in Eq. (4.139). (b) By Eq. (4.58),

4.62. Let X be a binomial r.v. with parameters (n, p). (a) Find the moment generating function of X. (b) Find the mean and variance of X.

(a) By definition (4.56) and Eq. (2.36), and letting q = 1 − p, we get

which can also be obtained by substituting z by et in Eq. (4.140). (b) The first two derivatives of MX(t) are

Thus, by Eq. (4.58),

Hence,

4.63. Let X be a Poisson r.v. with parameter λ. (a) Find the moment generating function of X. (b) Find the mean and variance of X.

(a) By definition (4.56) and Eq. (2.48),

which can also be obtained by substituting z by et in Eq. (4.147). (b) The first two derivatives of MX (t) are

Thus, by Eq. (4.58),

Hence,

4.64. Let X be an exponential r.v. with parameter λ. (a) Find the moment generating function of X. (b) Find the mean and variance of X.

(a) By definition (4.56) and Eq. (2.60),

(b) The first two derivatives of MX(t) are

Thus, by Eq. (4.58),

Hence,

4.65. Let X be a gamma r.v. with parameters (α, λ). (a) Find the moment generating function of X. (b) Find the mean and variance of X.

(a) By definition (4.56) and Eq. (2.65)

Let y = (λ − t)x, dy = (λ − t)dx. Then

Since (Eq. (2.66)) we obtain

(b) The first two derivatives of MX(t) are

Thus, by Eq. (4.58)

Hence,

4.66. Find the moment generating function of the standard normal r.v. X = N(0; 1), and calculate the first three moments of X.

By definition (4.56) and Eq. (2.71),

Combining the exponents and completing the square, that is,

we obtain

since the integrand is the pdf of N(t; 1).

Differentiating MX(t) with respect to t three times, we have

Thus, by Eq. (4.58),

4.67. Let Y = aX + b. Let MX(t) be the moment generating function of X. Show that the moment generating function of Y is given by

By Eqs. (4.56) and (4.129),

4.68. Find the moment generating function of a normal r.v. N(μ; σ2). If X is N(0; 1), then from Prob. 4.1 (or Prob. 4.43), we see that Y = σX + μ is N(μ; σ2). Then by setting a = σ and b = μ in Eq. (4.157) (Prob. 4.67) and using Eq. (4.156), we get

4.69. Suppose that r.v. X = N (0;1). Find the moment generating function of Y = X2.

By definition (4.56) and Eq. (2.71)

Using we obtain

4.70. Let X1, …, Xn be n independent r.v.’s and let the moment generating function of Xi be MXi(t). Let Y = X1 + … + Xn. Find the moment generating function

of Y.

By definition (4.56),

4.71. Show that if X1, …, Xn are independent Poisson r.v.’s Xi having parameter λi, then Y = X1 + … + Xn is also a Poisson r.v. with parameter λ = λ1 + … + λn.

Using Eqs. (4.160) and (4.153), the moment generating function of Y is

which is the moment generating function of a Poisson r.v. with parameter λ. Hence, Y is a Poisson r.v. with parameter λ = Σλi = λ1 + … + λn.

Note that Prob. 4.17 is a special case for n = 2.

4.72. Show that if X1, …, Xn are independent normal r.v.’s and , then Y = X1 + … + Xn is also a normal r.v. with mean μ = μ1 + … + μn and variance .

Using Eqs. (4.160) and (4.158), the moment generating function of Y is

which is the moment generating function of a normal r.v. with mean μ and variance σ2. Hence, Y is a normal r.v. with mean μ + μ1 + … + μn and variance .

Note that Prob. 4.20 is a special case for n = 2 with μi = 0 and .

4.73. Find the moment generating function of a gamma r.v. Y with parameters (n, λ).

From Prob. 4.39, we see that if X1, …, Xn are independent exponential r.v.’s, each with parameter λ, then Y = X1 + … + Xn is a gamma r.v. with parameters (n, λ). Thus, by Eqs. (4.160) and (4.154), the moment generating function of Y is

4.74. Suppose that X1, X2, …, Xn be independent standard normal r.v.’s and Xi = N(0;1). Let Y = X1

2 + X2 2 + … + Xn

2

(a) Find the moment generating function of Y. (b) Find the mean and variance of Y.

(a) By Eqs. (4.159) and (4.160)

(b) Differentiating Eq. (4.162), we obtain

Thus, by Eq. (4.58)

Hence,

Characteristic Functions

4.75. The r.v. X can take on the values x1 = −1 and x2 = +1 with pmf’s pX(x1) = pX(x2) = 0.5. Determine the characteristic function of X.

By definition (4.66), the characteristic function of X is

4.76. Find the characteristic function of a Cauchy r.v. X with parameter α and pdf given by

By direct integration (or from the Table of Fourier transforms in Appendix B), we have the following Fourier transform pair:

Now, by the duality property of the Fourier transform, we have the following Fourier transform pair:

or (by the linearity property of the Fourier transform)

Thus, the characteristic function of X is

Note that the moment generating function of the Cauchy r.v. X does not exist, since E(Xn) → ∞ for n ≥ 2.

4.77. The characteristic function of a r.v. X is given by

Find the pdf of X.

From formula (4.67), we obtain the pdf of X as

4.78. Find the characteristic function of a normal r.v. X = N(μ; σ2).

The moment generating function of N(μ; σ2) is [Eq. (4.158)]

Thus, the characteristic function of N(μ; σ2) is obtained by setting t = jω in MX(t); that is,

4.79. Let Y = aX + b. Show that if ΨX(ω) is the characteristic function of X, then the characteristic function of Y is given by

By definition (4.66),

4.80. Using the characteristic equation technique, redo part (b) of Prob. 4.18.

Let Z = X + Y, where X and Y are independent. Then

Applying the convolution theorem of the Fourier transform (Appendix B), we obtain

The Laws of Large Numbers and the Central Limit Theorem

4.81. Verify the weak law of large numbers (4.74); that is,

where and E(Xi) = μ, Var(Xi) = σ 2.

Using Eqs. (4.132) and (4.136), we have

Then it follows from Chebyshev’s inequality [Eq. (2.116)] (Prob. 2.39) that

Since limn → ∞σ 2/(ne2) = 0, we get

4.82. Let X be a r.v. with pdf ƒX(x) and let X1, …, Xn be a set of independent r.v.’s each with pdf ƒX(x). Then the set of r.v.’s X1, …, Xn is called a random sample of size n of X. The sample mean is defined by

Let X1, …, Xn be a random sample of X with mean μ and variance σ 2. How

many samples of X should be taken if the probability that the sample mean will not deviate from the true mean μ by more than σ/10 is at least 0.95?

Setting ε = σ/10 in Eq. (4.171), we have

or

Thus, if we want this probability to be at least 0.95, we must have 100/n ≤ 0.05 or n ≥ 100/0.05 = 2000.

4.83. Verify the central limit theorem (4.77).

Let X1, …, Xn be a sequence of independent, identically distributed r.v.’s with E(Xi) = μ and Var(Xi) = σ

2. Consider the sum Sn = X1 + … + Xn. Then by Eqs. (4.132) and (4.136), we have E(Sn) = nμ and Var(Sn) = nσ

2. Let

Then by Eqs. (4.129) and (4.130), we have E(Zn) = 0 and Var(Zn) = 1. Let M(t) be the moment generating function of the standardized r.v. Yi = (Xi − μ)/ σ. Since E(Yi) = 0 and , by Eq. (4.58), we have

Given that M′(t) and M″(t) are continuous functions of t, a Taylor (or Maclaurin) expansion of M(t) about t = 0 can be expressed as

By adding and subtracting t2/2, we have

Now, by Eqs. (4.157) and (4.160), the moment generating function of Zn is

Using Eqs. (4.174), Eq. (4.175) can be written as

where now t1 is between 0 and . Since M″(t) is continuous at t = 0 and t1 → 0 as n → ∞, we have

Thus, from elementary calculus, limn → ∞ (1 + x/n) n = ex, and we obtain

The right-hand side is the moment generating function of the standard normal r.v. Z = N(0; 1) [Eq. (4.156)]. Hence, by Lemma 4.3 of the moment generating function,

4.84. Let X1, …, Xn be n independent Cauchy r.v.’s with identical pdf shown in Prob. 4.76. Let

(a) Find the characteristic function of Yn. (b) Find the pdf of Yn. (c) Does the central limit theorem hold?

(a) From Eq. (4.166), the characteristic function of Xi is

Let Y = X1 + … + Xn. Then the characteristic function of Y is

Now Yn = (1/n)Y. Thus, by Eq. (4.168), the characteristic function of Yn is

(b) Equation (4.177) indicates that Yn is also a Cauchy r.v. with parameter a, and its pdf is the same as that of Xi.

(c) Since the characteristic function of Yn is independent of n and so is its pdf, Yn does not tend to a normal r.v. as n → ∞. Random variables in the given sequence all have finite mean but infinite variance. The central limit theorem does hold but for infinite variances for which Cauchy distribution is the stable (or convergent) distribution.

4.85. Let Y be a binomial r.v. with parameters (n, p). Using the central limit theorem, derive the approximation formula

where Φ(z) is the cdf of a standard normal r.v. [Eq. (2.73)].

We saw in Prob. 4.54 that if X1, …, Xn are independent Bernoulli r.v.’s, each with parameter p, then Y = X1 + … + Xn is a binomial r.v. with parameters (n, p). Since Xi’s are independent, we can apply the central limit theorem to the r.v. Zn defined by

Thus, for large n, Zn is normally distributed and

Substituting Eq. (4.179) into Eq. (4.180) gives

or

Because we are approximating a discrete distribution by a continuous one, a slightly better approximation is given by

Formula (4.181) is referred to as a continuity correction of Eq. (4.178).

4.86. Let Y be a Poisson r.v. with parameter λ. Using the central limit theorem, derive approximation formula:

We saw in Prob. 4.71 that if X1, …, Xn are independent Poisson r.v.’s Xi having parameter λi, then Y = X1 + … + Xn is also a Poisson r.v. with parameter λ = λ1 + … + λn. Using this fact, we can view a Poisson r.v. Y with parameter λ as a sum of independent Poisson r.v.’s Xi, i = 1, …, n, each with parameter λ /n; that is,

The central limit theorem then implies that the r.v. Z is defined by

is approximately normal and

Substituting Eq. (4.183) into Eq. (4.184) gives

or

Again, using a continuity correction, a slightly better approximation is given by

SUPPLEMENTARY PROBLEMS

4.87. Let Y = 2X + 3. Find the pdf of Y if X is a uniform r.v. over (−1, 2).

4.88. Let X be a r.v. with pdf ƒX(x). Let Y = |X| > Find the pdf of Y in terms of ƒX(x).

4.89. Let Y = sin X, where X is uniformly distributed over (0, 2π). Find the pdf of Y.

4.90. Let X and Y be independent r.v.’s, each uniformly distributed over (0, 1). Let Z = X + Y, W = X − Y. Find the marginal pdf’s of Z and W.

4.91. Let X and Y be independent exponential r.v.’s with parameters α and β, respectively. Find the pdf of (a) Z = X − Y; (b) Z = X/Y; (c) Z = max(X, Y); (d) Z = min(X, Y).

4.92. Let X denote the number of heads obtained when three independent tossings of a fair coin are made. Let Y = X2. Find E(Y).

4.93. Let X be a uniform r.v. over (−1, 1). Let Y = Xn.

(a) Calculate the covariance of X and Y.

(b) Calculate the correlation coefficient of X and Y.

4.94. What is the pmf of r.v. X whose probability generating function is

4.95. Let Y = a X + b. Express the probability generating function of Y, GY(z), in terms of the probability generating function of X, GX(z).

4.96. Let the moment generating function of a discrete r.v. X be given by

Find P(X = 3).

4.97. Let X be a geometric r.v. with parameter p.

(a) Determine the moment generating function of X.

(b) Find the mean of X for .

4.98. Let X be a uniform r.v. over (a, b).

(a) Determine the moment generating function of X.

(b) Using the result of (a), find E(X), E(X2), and E(X3).

4.99. Consider a r.v. X with pdf

Find the moment generating function of X.

4.100. Let X = N(0; 1). Using the moment generating function of X, determine E(X″).

4.101. Let X and Y be independent binomial r.v.’s with parameters (n, p) and (m, p), respectively. Let Z = X + Y. What is the distribution of Z?

4.102. Let (X, Y) be a continuous bivariate r.v. with joint pdf

(a) Find the joint moment generating function of X and Y.

(b) Find the joint moments m10, m01, and m11.

4.103. Let (X, Y) be a bivariate normal r.v. defined by Eq. (3.88). Find the joint moment generating function of X and Y.

4.104. Let X1, …, Xn be n independent r.v.’s and Xi > 0. Let

Show that for large n, the pdf of Y is approximately log-normal.

4.105. Let , where X is a Poisson r.v. with parameter λ. Show that Y ≈ N(0; 1) when λ is sufficiently large.

4.106. Consider an experiment of tossing a fair coin 1000 times. Find the probability of obtaining more that 520 heads (a) by using formula (4.178), and (b) by formula (4.181).

4.107. The number of cars entering a parking lot is Poisson distributed with a rate of 100 cars per hour. Find the time required for more than 200 cars to have entered the parking lot with probability 0.90 (a) by using formula (4.182), and (b) by formula (4.185).

ANSWERS TO SUPPLEMENTARY PROBLEMS

4.87.

4.88.

4.89.

4.90.

4.91.

4.92. 3

4.93.

4.94.

4.95. GY (z) = z aGX (z

b).

4.96. 0.35

4.97.

4.98.

4.99. MX(t) = e −7t + 8 t2

4.100.

4.101. Hint: Use the moment generating functions.

Z is a binomial r.v. with parameters (n + m, p).

4.102.

4.103.

4.104. Hint: Take the natural logarithm of Y and use the central limit theorem and the result of Prob. 4.10.

4.105. Hint: Find the moment generating function of Y and let λ → ∞.

4.106. (a) 0.1038 (b) 0.0974

4.107. (a) 2.189 h (b) 2.1946 h

CHAPTER 5

Random Processes

5.1 Introduction

In this chapter, we introduce the concept of a random (or stochastic) process. The theory of random processes was first developed in connection with the study of fluctuations and noise in physical systems. A random process is the mathematical model of an empirical process whose development is governed by probability laws. Random processes provides useful models for the studies of such diverse fields as statistical physics, communication and control, time series analysis, population growth, and management sciences.

5.2 Random Processes

A. Definition:

A random process is a family of r.v.’s {X(t), t ∈ T} defined on a given probability space, indexed by the parameter t, where t varies over an index set T.

Recall that a random variable is a function defined on the sample space S (Sec. 2.2). Thus, a random process {X(t, ζ), t ∈ T, ζ ∈ S} is really a function of two arguments {X(t, ζ), t ∈ T, ζ ∈ S}. For a fixed t(= tk), X(tk, ζ) = Xk(ζ) is a r.v. denoted by X(tk), as ζ varies over the sample space S. On the other hand, for a fixed sample point ζi ∈ S, X(t, ζi) = Xi(t) is a single function of time t, called a sample function or a realization of the process. (See Fig. 5- 1.) The totality of all sample functions is called an ensemble.

Of course if both ζ and t are fixed, X(tk, ζi) is simply a real number. In the following we use the notation X(t) to represent X(t, ζ).

Fig. 5-1 Random process.

B. Description of Random Process: In a random process {X(t), t ∈ T}, the index set T is called the parameter set of the random process. The values assumed by X(t) are called states, and the set of all possible values forms the state space E of the random process. If the index set T of a random process is discrete, then the process is called a discrete-parameter (or discrete-time) process. A discrete-parameter process is also called a random sequence and is denoted by {Xn, n = 1, 2, …}. If T is continuous, then we have a continuous-parameter (or continuous-time) process. If the state space E of a random process is discrete, then the process is called a discrete-state process, often referred to as a chain. In this case, the state space E is often assumed to be {0, 1, 2, …}. If the state space E is continuous, then we have a continuous-state process.

A complex random process X(t) is defined by

where X1(t) and X2(t) are (real) random processes and . Throughout this book, all random processes are real random processes unless specified otherwise.

5.3 Characterization of Random Processes

A. Probabilistic Descriptions:

Consider a random process X(t). For a fixed time t1, X(t1) = X1 is a r.v., and its cdf FX(x1; t1) is defined as

FX(x1; t1) is known as the first-order distribution of X(t). Similarly, given t1 and t2, X(t1) = X1 and X(t2) = X2 represent two r.v.’s. Their joint distribution is known as the second-order distribution of X(t) and is given by

In general, we define the nth-order distribution of X(t) by

If X(t) is a discrete-state process, then X(t) is specified by a collection of pmf’s:

If X(t) is a continuous-time process, then X(t) is specified by a collection of pdf’s:

The complete characterization of X(t) requires knowledge of all the distributions as n → ∞. Fortunately, often much less is sufficient.

B. Mean, Correlation, and Covariance Functions: As in the case of r.v.’s, random processes are often described by using statistical averages. The mean of X(t) is defined by

where X(t) is treated as a random variable for a fixed value of t. In general, μX(t) is a function of time, and it is often called the ensemble average of X(t). A measure of dependence among the r.v.’s of X(t) is provided by its autocorrelation function, defined by

Note that

and

The autocovariance function of X(t) is defined by

It is clear that if the mean of X(t) is zero, then KX(t, s) = RX(t, s). Note that the variance of X(t) is given by

If X(t) is a complex random process, then its autocorrelation function RX(t, s) and autocovariance function KX(t, s) are defined, respectively, by

and

where * denotes the complex conjugate.

5.4 Classification of Random Processes

If a random process X(t) possesses some special probabilistic structure, we can specify less to characterize X(t) completely. Some simple random processes are characterized completely by only the first-and second-order distributions.

A. Stationary Processes: A random process {X(t), t ∈ T} is said to be stationary or strict-sense stationary if, for all n and for every set of time instants (ti ∈ T, i = 1, 2, …, n},

for any τ. Hence, the distribution of a stationary process will be unaffected by a shift in the time origin, and X(t) and X(t + τ) will have the same distributions for any τ. Thus, for the first-order distribution,

and

Then

where μ and σ2 are constants. Similarly, for the second-order distribution,

and

Nonstationary processes are characterized by distributions depending on the points t1, t2, …, tn.

B. Wide-Sense Stationary Processes: If stationary condition (5.14) of a random process X(t) does not hold for all n but holds for n ≤ k, then we say that the process X(t)is stationary to order k. If X(t) is stationary to order 2, then X(t) is said to be wide-sense stationary (WSS) or weak stationary. If X(t) is a WSS random process, then we have

Note that a strict-sense stationary process is also a WSS process, but, in general, the converse is not true.

C. Independent Processes: In a random process X(t), if X(ti) for i = 1, 2, …, n are independent r.v.’s, so that for n = 2, 3, …,

then we call X(t) an independent random process. Thus, a first-order distribution is sufficient to characterize an independent random process X(t).

D. Processes with Stationary Independent Increments: A random process {X(t), t ≥ 0)} is said to have independent increments if whenever 0 < t1 < t2 < … < tn,

are independent. If {X(t), t ≥ 0)} has independent increments and X(t) − X(s) has the same distribution as X(t + h) − X(s + h) for all s, t, h ≥ 0, s < t, then the process X(t) is said to have stationary independent increments.

Let {X(t), t ≥ 0} be a random process with stationary independent increments and assume that X(0) = 0. Then (Probs. 5.21 and 5.22)

where μ1 = E[X(1)] and

where σ1 2 = Var[X(1)].

From Eq. (5.24), we see that processes with stationary independent increments are nonstationary. Examples of processes with stationary independent increments are Poisson processes and Wiener processes, which are discussed in later sections.

E. Markov Processes: A random process {X(t), t ∈ T} is said to be a Markov process if

whenever t1 < t2 < … < tn < tn+1. A discrete-state Markov process is called a Markov chain. For a discrete-

parameter Markov chain {Xn, n ≥ 0} (see Sec. 5.5), we have for every n

Equation (5.26) or Eq. (5.27) is referred to as the Markov property (which is also known as the memoryless property). This property of a Markov process states that the future state of the process depends only on the present state and not on the past history. Clearly, any process with independent increments is a Markov process.

Using the Markov property, the nth-order distribution of a Markov process X(t) can be expressed as (Prob. 5.25)

Thus, all finite-order distributions of a Markov process can be expressed in terms of the second-order distributions.

F. Normal Processes: A random process {X(t), t ∈ T} is said to be a normal (or Gaussian) process if for any integer n and any subset {t1, …, tn} of T, the n r.v.’s X(t1), …, X(tn) are jointly normally distributed in the sense that their joint characteristic function is given by

where ω1, …, ωn are any real numbers (see Probs. 5.59 and 5.60). Equation (5.29) shows that a normal process is completely characterized by the second-order distributions. Thus, if a normal process is wide-sense stationary, then it is also strictly stationary.

G. Ergodic Processes: Consider a random process {X(t), − ∞ < t < ∞} with a typical sample function x(t). The time average of x(t) is defined as

Similarly, the time autocorrelation function of x(t) is defined as

A random process is said to be ergodic if it has the property that the time averages of sample functions of the process are equal to the corresponding statistical or ensemble averages. The subject of ergodicity is extremely complicated. However, in most physical applications, it is assumed that stationary processes are ergodic.

5.5 Discrete-Parameter Markov Chains

In this section we treat a discrete-parameter Markov chain {Xn, n ≥ 0} with a discrete state space E = {0, 1, 2, … }, where this set may be finite or infinite. If Xn = i, then the Markov chain is said to be in state i at time n (or the nth step). A discrete-parameter Markov chain {Xn, n ≥ 0} is characterized by [Eq. (5.27)]

where P{xn+1 = j|Xn = i} are known as one-step transition probabilities. If P{xn+1 = j|Xn = i} is independent of n, then the Markov chain is said to possess stationary transition probabilities and the process is referred to as a homogeneous Markov chain. Otherwise the process is known as a nonhomogeneous Markov chain. Note that the concepts of a Markov chain’s having stationary transition probabilities and being a stationary random process should not be confused. The Markov process, in general, is not stationary. We shall consider only homogeneous Markov chains in this section.

A. Transition Probability Matrix:

Let {Xn, n ≥ 0} be a homogeneous Markov chain with a discrete infinite state space E = {0, 1, 2, … }. Then

regardless of the value of n. A transition probability matrix of {Xn, n ≥ 0} is defined by

where the elements satisfy

In the case where the state space E is finite and equal to {1, 2, …, m}, P is m × m dimensional; that is,

where

A square matrix whose elements satisfy Eq. (5.34) or (5.35) is called a Markov matrix or stochastic matrix.

B. Higher-Order Transition Probabilities—Chapman- Kolmogorov Equation: Tractability of Markov chain models is based on the fact that the probability distribution of {Xn, n ≥ 0} can be computed by matrix manipulations.

Let P = [pij] be the transition probability matrix of a Markov chain {Xn, n ≥ 0}. Matrix powers of P are defined by

with the (i, j)th element given by

Note that when the state space E is infinite, the series above converges, since by Eq. (5.34),

Similarly, P3 = PP2 has the (i, j)th element

and in general, Pn+1 = PPn has the (i, j)th element

Finally, we define P0 = I, where I is the identity matrix. The n-step transition probabilities for the homogeneous Markov chain

{Xn, n ≥ 0} are defined by

Then we can show that (Prob. 5.90)

We compute pij (n) by taking matrix powers.

The matrix identity

when written in terms of elements

is known as the Chapman-Kolmogorov equation. It expresses the fact that a transition from i to j in n + m steps can be achieved by moving from i to an intermediate k in n steps (with probability pik

(n), and then proceeding to j from k in m steps (with probability pkj

(m)). Furthermore, the events “go from i to k in n steps” and “go from k to j in m steps” are independent. Hence, the probability of the transition from i to j in n + m steps via i, k, j is pik

(n)pkj (m).

Finally, the probability of the transition from i to j is obtained by summing over the intermediate state k.

C. The Probability Distribution of {Xn, n ≥ 0}:

Let pi(n) = P(Xn = i) and

Then pi(0) = P(X0 = i) are the initial-state probabilities,

is called the initial-state probability vector, and p(n) is called the state probability vector after n transitions or the probability distribution of Xn. Now it can be shown that (Prob. 5.29)

which indicates that the probability distribution of a homogeneous Markov chain is completely determined by the one-step transition probability matrix P and the initial-state probability vector p(0).

D. Classification of States:

1. Accessible States: State j is said to be accessible from state i if for some n ≥ 0, pij

(n) > 0, and we write i → j. Two states i and j accessible to each other are said to communicate, and we write i ↔ j. If all states communicate with each other, then we say that the Markov chain is irreducible.

2. Recurrent States: Let Tj be the time (or the number of steps) of the first visit to state j after time zero, unless state j is never visited, in which case we set Tj = ∞. Then Tj is a discrete r.v. taking values in {1, 2, …, ∞}.

Let

and fij (0) = 0 since Tj ≥ 1. Then

and

The probability of visiting j in finite time, starting from i, is given by

Now state j is said to be recurrent if

That is, starting from j, the probability of eventual return to j is one. A recurrent state j is said to be positive recurrent if

and state j is said to be null recurrent if

Note that

3. Transient States: State j is said to be transient (or nonrecurrent) if

In this case there is positive probability of never returning to state j.

4. Periodic and Aperiodic States: We define the period of state j to be

where gcd stands for greatest common divisor. If d(j) > 1, then state j is called periodic with period d(j). If d(j) = 1, then

state j is called aperiodic. Note that whenever pjj > 0, j is aperiodic.

5. Absorbing States: State j is said to be an absorbing state if pjj = 1; that is, once state j is reached, it is never left.

E. Absorption Probabilities:

Consider a Markov chain X(n) = {Xn, n ≥ 0} with finite state space E = {1, 2, …, N} and transition probability matrix P. Let A = {1, …, m} be the set of absorbing states and B = {m + 1, …, N} be a set of nonabsorbing states. Then the transition probability matrix P can be expressed as

where I is an m × m identity matrix, 0 is an m × (N − m) zero matrix, and

Note that the elements of R are the one-step transition probabilities from nonabsorbing to absorbing states, and the elements of Q are the one-step transition probabilities among the nonabsorbing states.

Let U = [ukj], where

It is seen that U is an (N − m) × m matrix and its elements are the absorption probabilities for the various absorbing states. Then it can be shown that (Prob. 5.40)

The matrix Φ = (I − Q)−1 is known as the fundamental matrix of the Markov chain X(n). Let Tk denote the total time units (or steps) to absorption from state k. Let

Then it can be shown that (Prob. 5.74)

where ωki is the (k, i)th element of the fundamental matrix Φ.

F. Stationary Distributions: Let P be the transition probability matrix of a homogeneous Markov chain {Xn, n ≥ 0}. If there exists a probability vector such that

then is called a stationary distribution for the Markov chain. Equation (5.52) indicates that a stationary distribution is a (left) eigenvector of P with eigenvalue 1. Note that any nonzero multiple of is also an eigenvector of P. But the stationary distribution is fixed by being a probability vector; that is, its components sum to unity.

G. Limiting Distributions: A Markov chain is called regular if there is a finite positive integer m such that after m time-steps, every state has a nonzero chance of being occupied, no matter what the initial state. Let A > O denote that every element aij of A satisfies the condition aij > 0. Then, for a regular Markov chain with transition probability matrix P, there exists an m > 0 such that Pm > O. For a regular homogeneous Markov chain we have the following theorem:

THEOREM 5.5.1 Let {Xn, n ≥ 0} be a regular homogeneous finite-state Markov chain with transition matrix P. Then

where is a matrix whose rows are identical and equal to the stationary distribution for the Markov chain defined by Eq. (5.52).

5.6 Poisson Processes

A. Definitions:

Let t represent a time variable. Suppose an experiment begins at t = 0. Events of a particular kind occur randomly, the first at T1, the second at T2, and so on. The r.v. Ti denotes the time at which the ith event occurs, and the values ti of Ti (i = 1, 2, …) are called points of occurrence (Fig. 5-2).

Fig. 5-2

Let

and T0 = 0. Then Zn denotes the time between the (n − 1)st and the nth events (Fig. 5-2). The sequence of ordered r.v.’s {Zn, n ≥ 1} is sometimes called an interarrival process. If all r.v.’s Zn are independent and identically distributed, then {Zn, n ≥ 1} is called a renewal process or a recurrent process. From Eq. (5.54), we see that

where Tn denotes the time from the beginning until the occurrence of the nth event. Thus, {Tn, n ≥ 0} is sometimes called an arrival process.

B. Counting Processes: A random process {X(t), t ≥ 0} is said to be a counting process if X(t) represents the total number of “events” that have occurred in the interval (0, t). From its definition, we see that for a counting process, X(t) must satisfy the following conditions:

1. X(t) ≥ 0 and X(0) = 0. 2. X(t) is integer valued. 3. X(s) ≤ X(t) if s < t. 4. X(t) − X(s) equals the number of events that have occurred on the

interval (s, t).

A typical sample function (or realization) of X(t) is shown in Fig. 5-3. A counting process X(t) is said to possess independent increments if the

numbers of events which occur in disjoint time intervals are independent. A counting process X(t) is said to possess stationary increments if the number of events in the interval (s + h, t + h)—that is, X(t + h) − X(s + h)—has the same distribution as the number of events in the interval (s, t)—that is, X(t) − X(s)—for all s < t and h > 0.

Fig. 5-3 A sample function of a counting process.

C. Poisson Processes: One of the most important types of counting processes is the Poisson process (or Poisson counting process), which is defined as follows:

DEFINITION 5.6.1

A counting process X(t) is said to be a Poisson process with rate (or intensity) λ(> 0) if

1. X(0) = 0. 2. X(t) has independent increments. 3. The number of events in any interval of length t is Poisson distributed

with mean λt; that is, for all s, t > 0,

It follows from condition 3 of Def. 5.6.1 that a Poisson process has stationary increments and that

Then by Eq. (2.51) (Sec. 2.7E), we have

Thus, the expected number of events in the unit interval (0, 1), or any other interval of unit length, is just λ (hence the name of the rate or intensity).

An alternative definition of a Poisson process is given as follows:

DEFINITION 5.6.2 A counting process X(t) is said to be a Poisson process with rate (or intensity) λ(> 0) if

1. X(0) = 0.

2. X(t) has independent and stationary increments. 3. P[X(t + Δt) − X(t) = 1] = λ Δt + o(Δt) 4. P[X(t + Δt) − X(t) ≥ 2] = o(Δt)

where o(Δt) is a function of)t which goes to zero faster than does Δt; that is,

Note: Since addition or multiplication by a scalar does not change the property of approaching zero, even when divided by Δt, o((Δt) satisfies useful identities such as o(Δt) + o(Δt) = o(Δt) and ao(Δt) = o(Δt) for all constant a.

It can be shown that Def. 5.6.1 and Def. 5.6.2 are equivalent (Prob. 5.49). Note that from conditions 3 and 4 of Def. 5.6.2, we have (Prob. 5.50)

Equation (5.59) states that the probability that no event occurs in any short interval approaches unity as the duration of the interval approaches zero. It can be shown that in the Poisson process, the intervals between successive events are independent and identically distributed exponential r.v.’s (Prob. 5.53). Thus, we also identify the Poisson process as a renewal process with exponentially distributed intervals.

The autocorrelation function RX(t, s) and the autocovariance function KX(t, s) of a Poisson process X(t) with rate λ are given by (Prob. 5.52)

5.7 Wiener Processes

Another example of random processes with independent stationary increments is a Wiener process.

DEFINITION 5.7.1 A random process {X(t), t ≥ 0} is called a Wiener process if

1. X(t) has stationary independent increments. 2. The increment X(t) − X(s)(t > s) is normally distributed. 3. E[X(t)] = 0. 4. X(0) = 0.

The Wiener process is also known as the Brownian motion process, since it originates as a model for Brownian motion, the motion of particles suspended in a fluid. From Def. 5.7.1, we can verify that a Wiener process is a normal process (Prob. 5.61) and

where σ2 is a parameter of the Wiener process which must be determined from observations. When σ2 = 1, X(t) is called a standard Wiener (or standard Brownian motion) process.

The autocorrelation function RX(t, s) and the autocovariance function KX(t, s) of a Wiener process X(t) are given by (see Prob. 5.23)

DEFINITION 5.7.2 A random process {X(t), t ≥ 0} is called a Wiener process with drift coefficient μ if

1. X(t) has stationary independent increments. 2. X(t) is normally distributed with mean μt. 3. X(0) = 0.

From condition 2, the pdf of a standard Wiener process with drift coefficient μ is given by

5.8 Martingales

Martingales have their roots in gaming theory. A martingale is a random process that models a fair game. It is a powerful tool with many applications, especially in the field of mathematical finance.

A. Conditional Expectation and Filtrations: The conditional expectation E(Y|X1, …, Xn) is a r.v. (see Sec. 4.5 D) characterized by two properties:

1. The value of E(Y|X1, …, Xn) depends only on the values of X1, …, Xn, that is,

2.

If X1, …, Xn is a sequence of r.v.’s, we will use Fn to denote the information contained in X1, …, Xn and we write E(Y|Fn) for E(Y|X1, …, Xn), that is,

We also define information carried by r.v.’s X1, …, Xn in terms of the associated event space (σ-field), σ(X1, …, Xn). Thus,

and we say that Fn is an event space generated by X1, …, Xn. We have

A collection {Fn, n = 1, 2, …} satisfying Eq. (5.70) is called a filtration. Note that if a r.v. Z can be written as a function of X1, …, Xn, it is called

measurable with respect to X1, …, Xn, or Fn-measurable.

Properties of Conditional Expectations:

1. Linearity:

where a and b are constants.

2. Positivity:

If Y ≥ 0, then

3. Measurablity:

If Y is Fn-measurable, then

4. Stability:

If Z is Fn-measurable, then

5. Independence Law: If Y is independent of Fn, then

6. Tower Property:

7. Projection Law:

8. Jensen’s Inequality: If g is a convex function and E(|Y|) < ∞, then

B. Martingale in Discrete Time: Definition: A discrete-time random process {Mn, n ≥ 0} is a martingale with respect to Fn if

It immediately follows from Eq. (5.79) that for a martingale

A discrete-time random process {Mn, n ≥ 0} is a submartingale (supermartingale) with respect to Fn if

While a martingale models a fair game, the submartingale and supermartingale model favorable and unfavorable games, respectively.

Theorem 5.8.1 Let {Mn, n ≥ 0} be a martingale. Then for any given n

Equation (5.82) indicates that in a martingale all the r.v.’s have the same expectation (Prob. 5.67).

Theorem 5.8.2 (Doob decomposition) Let X = {Xn, n ≥ 0} be a submartingale with respect to Fn. Then there exists a martingale M = {Mn, n ≥ 0} and a process A = {An, n ≥ 0} such that

(1) M is a martingale with respect to Fn; (2) A is an increasing process An+1 ≥ An; (3) An is Fn−1-measurable for all n; (4) Xn = Mn + An.

(For the proof of this theorem see Prob. 5.78.)

C. Stopping Time and the Optional Stopping Theorem: Definition: A r.v. T is called a stopping time with respect to Fn if

1. T takes values from the set {0, 1, 2, …, ∞} 2. The event {T = n} is Fn-measurable.

EXAMPLE 5.1: A gambler has $100 and plays the slot machine at $1 per play.

1. The gambler stops playing when his capital is depleted. The number T = n1 of plays that it takes the gambler to stop play is a stopping time.

2. The gambler stops playing when his capital reaches $200. The number T = n2 of plays that it takes the gambler to stop play is a stopping time.

3. The gambler stops playing when his capital reaches $200, or is depleted, whichever comes first. The number T = min(n1, n2) of plays that it takes the gambler to stop play is a stopping time.

EXAMPLE 5.2 A typical example of the event T is not a stopping time; it is the moment the stock price attains its maximum over a certain period. To determine whether T is a point of maximum, we have to know the future values of the stock price and event {T = n} ∉ Fn.

Lemma 5.8.1

1. If T1 and T2 are stopping times, then so is T1 + T2. 2. If T1 and T2 are stopping times, then T = min(n1, n2) and T = max(n1,

n2) are also stopping times. 3. min (T, n) is a stopping time for any fixed n.

Let IA denote the indicator function of A, that is, the r.v. which equals 1 if A occurs and 0 otherwise. Note that I{T.n}, the indicator function of the event {T > n}, is Fn-measurable (since we need only the information up through time n to determine if we have stopped by time n).

Optional Stopping Theorem: Suppose {Mn, n ≥ 0} is a martingale and T is a stopping time. If

Then

Note that Eqs. (5.84) and (5.85) are always satisfied if the martingale is bounded and P(T < ∞) = 1.

D. Martingale in Continuous Time

A continuous-time filtration is a family {Ft, t ≥ 0} contained in the event space F such that Fs ⊂ Ft for s < t.

The continuous random process X(t) is a martingale with respect to Ft if

Similarly, continuous-time submartingales and supermartingales can be defined by replacing equal (=) sign by ≥ and ≤, respectively, in Eq. (5.88).

SOLVED PROBLEMS

Random Processes

5.1. Let X1, X2, … be independent Bernoulli r.v.’s (Sec. 2.7A) with P(Xn = 1) = p and P(Xn = 0) = q = 1 − p for all n. The collection of r.v.’s {Xn, n ≥ 1} is a random process, and it is called a Bernoulli process. (a) Describe the Bernoulli process. (b) Construct a typical sample sequence of the Bernoulli process.

(a) The Bernoulli process {Xn, n ≥ 1} is a discrete-parameter, discrete-state process. The state space is E = {0, 1}, and the index set is T = {1, 2, …}.

(b) A sample sequence of the Bernoulli process can be obtained by tossing a coin consecutively. If a head appears, we assign 1, and if a tail appears, we assign 0. Thus, for instance,

The sample sequence {xn} obtained above is plotted in Fig. 5-4.

Fig. 5-4 A sample function of a Bernoulli process.

5.2. Let Z1, Z2, … be independent identically distributed r.v.’s with P(Zn = 1) = p and P(Zn = − 1) = q = 1 − p for all n. Let

and X0 = 0. The collection of r.v.’s {Xn, n ≥ 0} is a random process, and it is called the simple random walk X(n) in one dimension. (a) Describe the simple random walk X(n). (b) Construct a typical sample sequence (or realization) of X(n).

(a) The simple random walk X(n) is a discrete-parameter (or time), discrete-state random process. The state space is E = {…, −2, −1, 0, 1, 2, …}, and the index parameter set is T = {0, 1, 2, …}.

(b) A sample sequence x(n) of a simple random walk X(n) can be produced by tossing a coin every second and letting x(n) increase by unity if a head appears and decrease by unity if a tail appears. Thus, for instance,

The sample sequence x(n) obtained above is plotted in Fig. 5-5. The simple random walk X(n) specified in this problem is said to be unrestricted because there are no bounds on the possible values of Xn.

The simple random walk process is often used in the following primitive gambling model: Toss a coin. If a head appears, you win one dollar; if a tail appears, you lose one dollar (see Prob. 5.38).

5.3. Let {Xn, n ≥ 0} be a simple random walk of Prob. 5.2. Now let the random process X(t) be defined by

(a) Describe X(t). (b) Construct a typical sample function of X(t).

(a) The random process X(t) is a continuous-parameter (or time), discrete-state random process. The state space is E = {…, −2, −1, 0, 1, 2, …}, and the index parameter set is T = {t, t ≥ 0}.

(b) A sample function x(t) of X(t) corresponding to Fig. 5-5 is shown in Fig. 5-6.

Fig. 5-5 A sample function of a random walk.

Fig. 5-6

5.4. Consider a random process X(t) defined by

where ω is a constant and Y is a uniform r.v. over (0, 1). (a) Describe X(t). (b) Sketch a few typical sample functions of X(t).

(a) The random process X(t) is a continuous-parameter (or time), continuous-state random process. The state space is E = {x: −1 < x < 1} and the index parameter set is T = {t: t ≥ 0}.

(b) Three sample functions of X(t) are sketched in Fig. 5-7.

Fig. 5-7

5.5. Consider patients coming to a doctor’s office at random points in time. Let Xn denote the time (in hours) that the nth patient has to wait in the office before being admitted to see the doctor. (a) Describe the random process X(n) = {Xn, n ≥ 1}. (b) Construct a typical sample function of X(n).

(a) The random process X(n) is a discrete-parameter, continuous-state random process. The state space is E = {x: x ≥ 0), and the index

parameter set is T = {1, 2, …}. (b) A sample function x(n) of X(n) is shown in Fig. 5-8.

Characterization of Random Processes

5.6. Consider the Bernoulli process of Prob. 5.1. Determine the probability of occurrence of the sample sequence obtained in part (b) of Prob. 5.1.

Since Xn’s are independent, we have

Thus, for the sample sequence of Fig. 5-4,

Fig. 5-8

5.7. Consider the random process X(t) of Prob. 5.4. Determine the pdf’s of X(t) at t = 0, π/4ω,π/2ω,π/ω.

For t = 0, X(0) = Y cos 0 = Y. Thus,

For . Thus,

For t = π/2ω, X(π/2ω) = Y cos π/2ω = 0; that is, X(π/2ω) = 0 irrespective of the value of Y. Thus, the pmf of X(ω/2ω) is

For t = π/ω, X(π/ω) = Y cos π = − Y. Thus,

5.8. Derive the first-order probability distribution of the simple random walk X(n) of Prob. 5.2.

The first-order probability distribution of the simple random walk X(n) is given by

where k is an integer. Note that P(X0 = 0) = 1. We note that pn(k) = 0 if n < |k| because the simple random walk cannot get to level k in less than |k| steps. Thus, n ≥ |k|.

Let Nn + and Nn

− be the r.v.’s denoting the numbers of +1s and −1s, respectively, in the first n steps.

Then

Adding Eqs. (5.91) and (5.92), we get

Thus, Xn = k if and only if . From Eq. (5.93), we note that 2Nn

+ = n + Xn must be even. Thus, Xn must be even if n is even, and Xn must be odd if n is odd. We note that Nn

+ is a binomial r.v. with parameters (n, p). Thus, by Eq. (2.36), we obtain

where n ≥ |k|, and n and k are either both even or both odd.

5.9. Consider the simple random walk X(n) of Prob. 5.2. (a) Find the probability that X(n) = −2 after four steps. (b) Verify the result of part (a) by enumerating all possible sample

sequences that lead to the value X(n) = −2 after four steps.

(a) Setting k = −2 and n = 4 in Eq. (5.94), we obtain

Fig. 5-9

(b) All possible sample functions that lead to the value X4 = −2 after four steps are shown in Fig. 5-9. For each sample sequence, P(X4 = −2) = pq3. There are only four sample functions that lead to the value X4 = −2 after four steps. Thus, P(X4 = −2) = 4pq

3.

5.10. Find the mean and variance of the simple random walk X(n) of Prob. 5.2.

From Eq. (5.89), we have

and X0 = 0 and Zn (n = 1, 2, …) are independent and identically distributed (iid) r.v.’s with

From Eq. (5.95), we observe that

Then, because the Zn are iid r.v.’s and X0 = 0, by Eqs. (4.132) and (4.136), we have

Now

Thus,

Hence,

Note that if , then

5.11. Find the autocorrelation function Rx(n, m) of the simple random walk X(n) of Prob. 5.2.

From Eq. (5.96), we can express Xn as

where Z0 = X0 = 0 and Zi (i ≥ 1) are iid r.v.’s with

By Eq. (5.7),

Then by Eq. (5.104),

Using Eqs. (5.97) and (5.98), we obtain

or

Note that if , then

5.12. Consider the random process X(t) of Prob. 5.4; that is,

where ω is a constant and Y is a uniform r.v. over (0, 1). (a) Find E[X(t)]. (b) Find the autocorrelation function Rx(t, s) of X(t). (c) Find the autocovariance function Kx(t, s) of X(t).

(a) From Eqs. (2.58) and (2.110), we have and . Thus,

(b) By Eq. (5.7), we have

(c) By Eq. (5.10), we have

5.13. Consider a discrete-parameter random process X(n) = {Xn, n ≥ 1} where the Xn’s are iid r.v.’s with common cdf Fx(x), mean μ, and variance σ2. (a) Find the joint cdf of X(n). (b) Find the mean of X(n). (c) Find the autocorrelation function Rx(n, m) of X(n). (d) Find the autocovariance function Kx(n, m) of X(n).

(a) Since the Xn’s are iid r.v.’s with common cdf Fx(x), the joint cdf of X(n) is given by

(b) The mean of X(n) is

(c) If n ≠ m, by Eqs. (5.7) and (5.113),

If n = m, then by Eq. (2.31),

(d) By Eq. (5.10),

Classification of Random Processes

5.14. Show that a random process which is stationary to order n is also stationary to all orders lower than n.

Assume that Eq. (5.14) holds for some particular n; that is,

for any τ. Letting xn → ∞, we have [see Eq. (3.63)]

and the process is stationary to order n − 1. Continuing the same procedure, we see that the process is stationary to all orders lower than n.

5.15. Show that if {X(t), t ∈ T} is a strict-sense stationary random process, then it is also WSS.

Since X(t) is strict-sense stationary, the first- and second-order distributions are invariant through time translation for all τ ∈ T. Then we have

And, hence, the mean function μx(t) must be constant; that is,

Similarly, we have

so that the autocorrelation function would depend on the time points s and t only through the difference |t − s|. Thus, X(t) is WSS.

5.16. Let {Xn, n ≥ 0} be a sequence of iid r.v.’s with mean 0 and variance 1. Show that {Xn, n ≥ 0} is a WSS process.

By Eq. (5.113),

and by Eq. (5.114),

which depends only on k. Thus, {Xn} is a WSS process.

5.17. Show that if a random process X(t) is WSS, then it must also be covariance stationary.

If X(t) is WSS, then

which indicates that KX(t, t + τ) depends only on τ; thus, X(t) is covariance stationary.

5.18. Consider a random process X(t) defined by

where ω is constant and U and V are r.v.’s. (a) Show that the condition

is necessary for X(t) to be stationary.

(b) Show that X(t) is WSS if and only if U and V are uncorrelated with equal variance; that is,

(a) Now

must be independent of t for X(t) to be stationary. This is possible only if μx(t) = 0, that is, E(U) = E(V) = 0.

(b) If X(t) is WSS, then

But X(0) = U and X(μ/2ω) = V; thus,

Using the above result, we obtain

which will be a function of τ only if E(UV) = 0. Conversely, if E(UV) = 0 and E(U2) = E(V2) = σ2, then from the result of part (a) and Eq. (5.119), we have

Hence, X(t) is WSS.

5.19. Consider a random process X(t) defined by

where U and V are independent r.v.’s, each of which assumes the values −2 and 1 with the probabilities and , respectively. Show that

X(t) is WSS but not strict-sense stationary.

We have

Since U and V are independent,

Thus, by the results of Prob. 5.18, X(t) is WSS. To see if X(t) is strict- sense stationary, we consider E[X3(t)].

which is a function of t. From Eq. (5.16), we see that all the moments of a strict-sense stationary process must be independent of time. Thus, X(t) is not strict-sense stationary.

5.20. Consider a random process X(t) defined by

where A and ω are constants and Θ is a uniform r.v. over (−π, π). Show that X(t) is WSS.

From Eq. (2.56), we have

Then

Setting s = t + τ in Eq. (5.7), we have

Since the mean of X(t) is a constant and the autocorrelation of X(t) is a function of time difference only, we conclude that X(t) is WSS.

5.21. Let {X(t), t ≥ 0} be a random process with stationary independent increments, and assume that X(0) = 0. Show that

where μ1 = E[X(1)].

Let

Then, for any t and s and using Eq. (4.132) and the property of the stationary independent increments, we have

The only solution to the above functional equation is f(t) = ct, where c is a constant. Since c = f(1) = E[X(1)], we obtain

5.22. Let {X(t), t ≥ 0} be a random process with stationary independent increments, and assume that X(0) = 0. Show that (a)

(b)

where σ1 2 = Var[X(1)].

(a) Let g(t) = Var[X(t)] = Var[X(t) − X(0)] Then, for any t and s and using Eq. (4.136) and the property of the stationary independent increments, we get

which is the same functional equation as Eq. (5.123). Thus, g(t) = kt, where k is a constant. Since k = g(1) = Var[X(1)], we obtain

(b) Let t > s. Then

Thus, using Eq. (5.124), we obtain

5.23. Let {X(t), t ≥ 0} be a random process with stationary independent increments, and assume that X(0) = 0. Show that

where σ1 2 = Var[X(1)].

By definition (2.28),

Thus,

Using Eqs. (5.124) and (5.125), we obtain

where σ1 2 = Var[X(1)].

5.24. (a) Show that a simple random walk X(n) of Prob. 5.2 is a Markov chain. (b) Find its one-step transition probabilities.

(a) From Eq. (5.96) (Prob. 5.10), X(n) = {Xn, n ≥ 0} can be expressed as

where Zn (n = 1, 2, …) are iid r.v.’s with

Then X(n) = {Xn, n ≥ 0} is a Markov chain, since

since Zn+1 is independent of X0, X1, …, Xn. (b) The one-step transition probabilities are given by

which do not depend on n. Thus, a simple random walk X(n) is a homogeneous Markov chain.

5.25. Show that for a Markov process X(t), the second-order distribution is sufficient to characterize X(t).

Let X(t) be a Markov process with the nth-order distribution

Then, using the Markov property (5.26), we have

Applying the above relation repeatedly for lower-order distribution, we can write

Hence, all finite-order distributions of a Markov process can be completely determined by the second-order distribution.

5.26. Show that if a normal process is WSS, then it is also strict-sense stationary.

By Eq. (5.29), a normal random process X(t) is completely characterized by the specification of the mean E[X(t)] and the

covariance function KX(t, s) of the process. Suppose that X(t) is WSS. Then, by Eqs. (5.21) and (5.22), Eq. (5.29) becomes

Now we translate all of the time instants t1, t2, …, tn by the same amount τ. The joint characteristic function of the new r.v.’s X(ti + τ), i = 1, 2, …, n, is then

which indicates that the joint characteristic function (and hence the corresponding joint pdf) is unaffected by a shift in the time origin. Since this result holds for any n and any set of time instants (ti ∈ T, i = 1, 2, …, n), it follows that if a normal process is WSS, then it is also strict-sense stationary.

5.27. Let {X(t), −∞ < t < ∞} be a zero-mean, stationary, normal process with the autocorrelation function

Let {X(ti), i = 1, 2, …, n} be a sequence of n samples of the process taken at the time instants

Find the mean and the variance of the sample mean

Since X(t) is zero-mean and stationary, we have

and

Thus,

and

By Eq. (5.130),

Thus,

Discrete-Parameter Markov Chains

5.28. Show that if P is a Markov matrix, then Pn is also a Markov matrix for any positive integer n.

Let

Then by the property of a Markov matrix [Eq. (5.35)], we can write

or

where

Premultiplying both sides of Eq. (5.134) by P, we obtain

which indicates that P2 is also a Markov matrix. Repeated premultiplication by P yields

which shows that Pn is also a Markov matrix.

5.29. Verify Eq. (5.39); that is,

We verify Eq. (5.39) by induction. If the state of X0 is i, state X1 will be j only if a transition is made from i to j. The events {X0 = i, i = 1, 2, …} are mutually exclusive, and one of them must occur. Hence, by the law of total probability [Eq. (1.44)],

or

In terms of vectors and matrices, Eq. (5.135) can be expressed as

Thus, Eq. (5.39) is true for n = 1. Assume now that Eq. (5.39) is true for n = k; that is,

Again, by the law of total probability,

or

In terms of vectors and matrices, Eq. (5.137) can be expressed as

which indicates that Eq. (5.39) is true for k + 1. Hence, we conclude that Eq. (5.39) is true for all n ≥ 1.

5.30. Consider a two-state Markov chain with the transition probability matrix

(a) Show that the n-step transition probability matrix Pn is given by

(b) Find Pn when n → ∞.

(a) From matrix analysis, the characteristic equation of P is

Thus, the eigenvalues of P are λ1 = 1 and λ2 = 1 − a − b. Then, using the spectral decomposition method, Pn can be expressed as

where E1 and E2 are constituent matrices of P, given by

Substituting λ1 = 1 and λ2 = 1 − a − b in the above expressions, we obtain

Thus, by Eq. (5.141), we obtain

(b) If 0 < a < 1, 0 < b < 1, then 0 < 1 − a < 1 and |1 − a − b| < 1. So limn→∞ (1 − a − b)

n = 0 and

Note that a limiting matrix exists and has the same rows (see Prob. 5.47).

5.31. An example of a two-state Markov chain is provided by a communication network consisting of the sequence (or cascade) of stages of binary communication channels shown in Fig. 5-10. Here Xn

Fig. 5-10 Binary communication network.

denotes the digit leaving the nth stage of the channel and X0 denotes the digit entering the first stage. The transition probability matrix of this communication network is often called the channel matrix and is given by Eq. (5.139); that is,

Assume that a = 0.1 and b = 0.2, and the initial distribution is P(X0 = 0) = P(X0 = 1) = 0.5. (a) Find the distribution of Xn. (b) Find the distribution of Xn when n → ∞.

(a) The channel matrix of the communication network is

and the initial distribution is

By Eq. (5.39), the distribution of Xn is given by

Letting a = 0.1 and b = 0.2 in Eq. (5.140), we get

Thus, the distribution of Xn is

that is,

(b) Since limn→∞ (0.7) n = 0, the distribution of Xn when n → ∞ is

5.32. Verify the transitivity property of the Markov chain; that is, if i → j and j → k, then i → k.

By definition, the relations i → j and j → k imply that there exist integers n and m such that pij

(n) > 0 and pjk (m) > 0. Then, by the

Chapman-Kolmogorov equation (5.38), we have

Therefore, i → k.

5.33. Verify Eq. (5.42).

If the Markov chain {Xn} goes from state i to state j in m steps, the first step must take the chain from i to some state k, where k ≠ j. Now after that first step to k, we have m − 1 steps left, and the chain must get to state j, from state k, on the last of those steps. That is, the first visit to state j must occur on the (m − 1)st step, starting now in state k. Thus, we must have

5.34. Show that in a finite-state Markov chain, not all states can be transient.

Suppose that the states are 0, 1, …, m, and suppose that they are all transient. Then by definition, after a finite amount of time (say T0), state 0 will never be visited; after a finite amount of time (say T1), state 1 will never be visited; and so on. Thus, after a finite time T = max{T0, T1, …, Tm}, no state will be visited. But as the process must be in some state after time T, we have a contradiction. Thus, we conclude that not all states can be transient and at least one of the states must be recurrent.

5.35. A state transition diagram of a finite-state Markov chain is a line diagram with a vertex corresponding to each state and a directed line between two vertices i and j if pij > 0. In such a diagram, if one can

move from i and j by a path following the arrows, then i → j. The diagram is useful to determine whether a finite-state Markov chain is irreducible or not, or to check for periodicities. Draw the state transition diagrams and classify the states of the Markov chains with the following transition probability matrices:

Fig. 5-11 State transition diagram.

(a) The state transition diagram of the Markov chain with P of part (a) is shown in Fig. 5-11(a). From Fig. 5-11(a), it is seen that the Markov chain is irreducible and aperiodic. For instance, one can get back to state 0 in two steps by going from 0 to 1 to 0.

However, one can also get back to state 0 in three steps by going from 0 to 1 to 2 to 0. Hence 0 is aperiodic. Similarly, we can see that states 1 and 2 are also aperiodic.

(b) The state transition diagram of the Markov chain with P of part (b) is shown in Fig. 5-11(b). From Fig. 5-11(b), it is seen that the Markov chain is irreducible and periodic with period 3.

(c) The state transition diagram of the Markov chain with P of part (c) is shown in Fig. 5-11(c). From Fig. 5-11(c), it is seen that the Markov chain is not irreducible, since states 0 and 4 do not communicate, and state 1 is absorbing.

5.36. Consider a Markov chain with state space {0, 1} and transition probability matrix

(a) Show that state 0 is recurrent. (b) Show that state 1 is transient.

(a) By Eqs. (5.41) and (5.42), we have

Then, by Eqs. (5.43),

Thus, by definition (5.44), state 0 is recurrent. (b) Similarly, we have

and

Thus, by definition (5.48), state 1 is transient.

5.37. Consider a Markov chain with state space {0, 1, 2} and transition probability matrix

Show that state 0 is periodic with period 2.

The characteristic equation of P is given by

Thus, by the Cayley-Hamilton theorem (in matrix analysis), we have P3 = P. Thus, for n ≥ 1,

Therefore,

Thus, state 0 is periodic with period 2. Note that the state transition diagram corresponding to the given P

is shown in Fig. 5-12. From Fig. 5-12, it is clear that state 0 is periodic with period 2.

Fig. 5-12

5.38. Let two gamblers, A and B, initially have k dollars and m dollars, respectively. Suppose that at each round of their game, A wins one dollar from B with probability p and loses one dollar to B with

probability q = 1 − p. Assume that A and B play until one of them has no money left. (This is known as the Gambler’s Ruin problem.) Let Xn be A’s capital after round n, where n = 0, 1, 2, … and X0 = k. (a) Show that X(n) = {Xn, n ≥ 0} is a Markov chain with absorbing

states. (b) Find its transition probability matrix P.

(a) The total capital of the two players at all times is

Let Zn (n ≥ 1) be independent r.v.’s with P(Zn = 1) = p and P(Zn = −1) = q = 1 − p for all n. Then

and X0 = k. The game ends when Xn = 0 or Xn = N. Thus, by Probs. 5.2 and 5.24, X(n) = {Xn, n ≥ 0} is a Markov chain with state space E = {0, 1, 2, …, N}, where states 0 and N are absorbing states. The Markov chain X(n) is also known as a simple random walk with absorbing barriers.

(b) Since

the transition probability matrix P is

For example, when and N = 4,

5.39. Consider a homogeneous Markov chain X(n) = {Xn, n ≥ 0} with a finite state space E = {0, 1, …, N}, of which A = {0, 1, …, m}, m ≥ 1, is a set of absorbing states and B = {m + 1, …, N} is a set of nonabsorbing states. It is assumed that at least one of the absorbing states in A is accessible from any nonabsorbing states in B. Show that absorption of X(n) in one or another of the absorbing states is certain.

If X0 ∈ A, then there is nothing to prove, since X(n) is already absorbed. Let X0 ∈ B. By assumption, there is at least one state in A which is accessible from any state in B. Now assume that state k ∈ A is accessible from j ∈ B. Let njk (< ∞) be the smallest number n such that pjk

(n) > 0. For a given state j, let nj be the largest of njk as k varies and n′ be the largest of nj as j varies. After n′ steps, no matter what the

initial state of X(n), there is a probabilityp > 0 that X(n) is in an absorbing state. Therefore,

and 0 < 1 − p < 1. It follows by homogeneity and the Markov property that

Now since limk→∞(1 − p) k = 0, we have

which shows that absorption of X(n) in one or another of the absorption states is certain.

5.40. Verify Eq. (5.50).

Let X(n) = {Xn, n ≥ 0} be a homogeneous Markov chain with a finite state space E = {0, 1, …, N}, of which A = {0, 1, …, m}, m ≥ 1, is a set of absorbing states and B = {m + 1, …, N} is a set of nonabsorbing states. Let state k ∈ B at the first step go to i → E with probability pki. Then

Now

Then Eq. (5.147) becomes

But pkj, k = m + 1, …, N; j = 1, …, m, are the elements of R, whereas pki, k = m + 1, …, N; i = m + 1, …, N are the elements of Q [see Eq. (5.49a)]. Hence, in matrix notation, Eq. (5.148) can be expressed as

Premultiplying both sides of the second equation of Eq. (5.149) with (I − Q)−1, we obtain

5.41. Consider a simple random walk X(n) with absorbing barriers at state 0 and state N = 3 (see Prob. 5.38). (a) Find the transition probability matrix P. (b) Find the probabilities of absorption into states 0 and 3.

(a) The transition probability matrix P is [Eq. (5.146)]

(b) Rearranging the transition probability matrix P as [Eq. (5.49a)],

and by Eq. (5.49b), the matrices Q and R are given by

Then

and

By Eq. (5.50),

Thus, the probabilities of absorption into state 0 from states 1 and 2 are given, respectively, by

and the probabilities of absorption into state 3 from states 1 and 2 are given, respectively, by

Note that

which confirm the proposition of Prob. 5.39.

5.42. Consider the simple random walk X(n) with absorbing barriers at 0 and 3 (Prob. 5.41). Find the expected time (or steps) to absorption when X0 = 1 and when X0 = 2.

The fundamental matrix Φ of X(n) is [Eq. (5.150)]

Let Ti be the time to absorption when X0 = i. Then by Eq. (5.51), we get

5.43. Consider the gambler’s game described in Prob. 5.38. What is the probability of A’s losing all his money?

Let P(k), k = 0, 1, 2, …, N, denote the probability that A loses all his money when his initial capital is k dollars. Equivalently, P(k) is the probability of absorption at state 0 when X0 = k in the simple random walk X(n) with absorbing barriers at states 0 and N. Now if 0 < k < N, then

where pP(k + 1) is the probability that A wins the first round and subsequently loses all his money and qP(k − 1) is the probability that A loses the first round and subsequently loses all his money. Rewriting Eq. (5.153), we have

which is a second-order homogeneous linear constant-coefficient difference equation. Next, we have

since if k = 0, absorption at 0 is a sure event, and if k = N, absorption at N has occurred and absorption at 0 is impossible. Thus, finding P(k) reduces to solving Eq. (5.154) subject to the boundary conditions given by Eq. (5.155). Let P(k) = rk. Then Eq. (5.154) becomes

Setting k = 1 (and noting that p + q = 1), we get

from which we get r = 1 and r = q/p. Thus,

where c1 and c2 are arbitrary constants. Now, by Eq. (5.155),

Solving for c1 and c2, we obtain

Hence,

Note that if N ≫ k,

Setting r = q/p in Eq. (5.157), we have

Thus, when ,

5.44. Show that Eq. (5.157) is consistent with Eq. (5.151).

Substituting k = 1 and N = 3 in Eq. (5.134), and noting that p + q = 1, we have

Now from Eq. (5.151), we have

5.45. Consider the simple random walk X(n) with state space E = {0, 1, 2, …, N}, where 0 and N are absorbing states (Prob. 5.38). Let r.v. Tk denote the time (or number of steps) to absorption of X(n) when X0 = k, k = 0, 1, …, N. Find E(Tk).

Let Y(k) = E(Tk). Clearly, if k = 0 or k = N, then absorption is immediate, and we have

Let the probability that absorption takes m steps when X0 = k be defined by

Then, we have (Fig. 5-13)

and

Setting m − 1 = i, we get

Fig. 5-13 Simple random walk with absorbing barriers.

Now by the result of Prob. 5.39, we see that absorption is certain; therefore,

Thus,

or

Rewriting Eq. (5.163), we have

Thus, finding P(k) reduces to solving Eq. (5.164) subject to the boundary conditions given by Eq. (5.160). Let the general solution of Eq. (5.164) be

where Yh(k) is the homogeneous solution satisfying

and Yp(k) is the particular solution satisfying

Let Yp(k) = αk, where α is a constant. Then Eq. (5.166) becomes

from which we get α = 1/(q − p) and

Since Eq. (5.165) is the same as Eq. (5.154), by Eq. (5.156), we obtain

where c1 and c2 are arbitrary constants. Hence, the general solution of Eq. (5.164) is

Now, by Eq. (5.160),

Solving for c1 and c2, we obtain

Substituting these values in Eq. (5.169), we obtain (for p ≠ q)

When , we have

5.46. Consider a Markov chain with two states and transition probability matrix

(a) Find the stationary distribution of the chain.

(b) Find limn→∞ P n.

(a) By definition (5.52),

or

which yields p1 = p2. Since p1 + p2 = 1, we obtain

(b) Now

and limn→∞ P n does not exist.

5.47. Consider a Markov chain with two states and transition probability matrix

(a) Find the stationary distribution of the chain.

(b) Find limn→∞ P n.

(c) Find limn→∞ P n by first evaluating Pn.

(a) By definition (5.52), we have

which yields

Each of these equations is equivalent to p1 = 2p2. Since p1 + p2 = 1, we obtain

(b) Since the Markov chain is regular, by Eq. (5.53), we obtain

(c) Setting and in Eq. (5.143) (Prob. 5.30), we get

Since , we obtain

Poisson Processes

5.48. Let Tn denote the arrival time of the nth customer at a service station. Let Zn denote the time interval between the arrival of the nth customer and the (n − 1)st customer; that is,

and T0 = 0. Let {X(t), t ≥ 0} be the counting process associated with {Tn, n ≥ 0}. Show that if X(t) has stationary increments, then Zn, n = 1, 2, …, are identically distributed r.v.’s.

We have

By Eq. (5.172),

Suppose that the observed value of Tn − 1 is tn − 1. The event (Tn > Tn − 1 + z|Tn − 1 = tn − 1) occurs if and only if X(t) does not change count during the time interval (tn − 1, tn − 1 + z) (Fig. 5-14). Thus,

or

Since X(t) has stationary increments, the probability on the right-hand side of Eq. (5.173) is a function only of the time difference z. Thus,

which shows that the conditional distribution function on the left-hand side of Eq. (5.174) is independent of the particular value of n in this case, and hence we have

which shows that the cdf of Zn is independent of n. Thus, we conclude that the Zn’s are identically distributed r.v.’s.

Fig. 5-14

5.49. Show that Definition 5.6.2 implies Definition 5.6.1.

Let pn(t) = P[X(t) = n]. Then, by condition 2 of Definition 5.6.2, we have

Now, by Eq. (5.59), we have

Letting Δt→0, and by Eq. (5.58), we obtain

Solving the above differential equation, we get

where k is an integration constant. Since p0(0) = P[X(0) = 0] = 1, we obtain

Similarly, for n > 0,

Now, by condition 4 of Definition 5.6.2, the last term in the above expression is o(Δt). Thus, by conditions 2 and 3 of Definition 5.6.2, we have

and letting Δt → 0 yields

Multiplying both sides by eλt, we get

Hence,

Then by Eq. (5.177), we have

or

where c is an integration constant. Since p1(0) = P[X(0) = 1] = 0, we obtain

To show that

we use mathematical induction. Assume that it is true for n − 1; that is,

Substituting the above expression into Eq. (5.179), we have

Integrating, we get

Since pn(0) = 0, c1 = 0, and we obtain

which is Eq. (5.55) of Definition 5.6.1. Thus we conclude that Definition 5.6.2 implies Definition 5.6.1.

5.50. Verify Eq. (5.59).

We note first that X(t) can assume only nonnegative integer values; therefore, the same is true for the counting increment X(t + Δt) − X(t). Thus, summing over all possible values of the increment, we get

Substituting conditions 3 and 4 of Definition 5.6.2 into the above equation, we obtain

5.51. (a) Using the Poisson probability distribution in Eq. (5.181), obtain an analytical expression for the correction term oΔ(t) in the

expression (condition 3 of Definition 5.6.2)

(b) Show that this correction term does have the property of Eq. (5.58); that is,

(a) Since the Poisson process X(t) has stationary increments, Eq. (5.182) can be rewritten as

Using Eq. (5.181) [or Eq. (5.180)], we have

Equating the above expression with Eq. (5.183), we get

from which we obtain

(b) From Eq. (5.184), we have

5.52. Find the autocorrelation function RX(t, s) and the autocovariance function KX(t, s) of a Poisson process X(t) with rate λ.

From Eqs. (5.56) and (5.57),

Now, the Poisson process X(t) is a random process with stationary independent increments and X(0) = 0. Thus, by Eq. (5.126) (Prob. 5.23), we obtain

since σ1 2 = Var[X(1)] = λ. Next, since E[X(t)] E[X(s)] = λ2ts, by Eq.

(5.10), we obtain

5.53. Show that the time intervals between successive events (or interarrival times) in a Poisson process X(t) with rate λ are independent and identically distributed exponential r.v.’s with parameter λ.

Let Z1, Z2, … be the r.v.’s representing the lengths of interarrival times in the Poisson process X(t). First, notice that {Z1 > t} takes place if and only if no event of the Poisson process occurs in the interval (0, t), and thus by Eq. (5.177),

or

Hence, Z1 is an exponential r.v. with parameter λ [Eq. (2.61)]. Let f1(t) be the pdf of Z1. Then we have

which indicates that Z2 is also an exponential r.v. with parameter λ and is independent of Z1. Repeating the same argument, we conclude that Z1, Z2, … are iid exponential r.v.’s with parameter λ.

5.54. Let Tn denote the time of the nth event of a Poisson process X(t) with rate λ. Show that Tn is a gamma r.v. with parameters (n, λ).

Clearly,

where Zn, n = 1, 2, …, are the interarrival times defined by Eq. (5.172). From Prob. 5.53, we know that Zn are iid exponential r.v.’s with parameter λ. Now, using the result of Prob. 4.39, we see that Tn is a gamma r.v. with parameters (n, λ), and its pdf is given by [Eq. (2.65)]:

The random process {Tn, n ≥ 1} is often called an arrival process.

5.55. Suppose t is not a point at which an event occurs in a Poisson process X(t) with rate λ. Let W(t) be the r.v. representing the time until the next occurrence of an event. Show that the distribution of W(t) is independent of t and W(t) is an exponential r.v. with parameter λ.

Let s (0 ≤ s < t) be the point at which the last event [say the (n − 1)st event] occurred (Fig. 5-15). The event {W(t) > τ} is equivalent to the event

Fig. 5-15

Thus, using Eq. (5.187), we have

and

which indicates that W(t) is an exponential r.v. with parameter λ and is independent of t. Note that W(t) is often called a waiting time.

5.56. Patients arrive at the doctor’s office according to a Poisson process with rate minute. The doctor will not see a patient until at least three patients are in the waiting room. (a) Find the expected waiting time until the first patient is admitted to

see the doctor. (b) What is the probability that nobody is admitted to see the doctor in

the first hour?

(a) Let Tn denote the arrival time of the nth patient at the doctor’s office. Then

where Zn, n = 1, 2, …, are iid exponential r.v.’s with parameter . By Eqs. (4.132) and (2.62),

The expected waiting time until the first patient is admitted to see the doctor is

(b) Let X(t) be the Poisson process with parameter . The probability that nobody is admitted to see the doctor in the first hour is the same as the probability that at most two patients arrive in the first 60 minutes. Thus, by Eq. (5.55),

5.57. Let Tn denote the time of the nth event of a Poisson process X(t) with rate λ. Suppose that one event has occurred in the interval (0, t). Show that the conditional distribution of arrival time T1 is uniform over (0, t).

For τ ≤ t,

which indicates that T1 is uniform over (0, t) [see Eq. (2.57)].

5.58. Consider a Poisson process X(t) with rate λ, and suppose that each time an event occurs, it is classified as either a type 1 or a type 2 event. Suppose further that the event is classified as a type 1 event with probability p and a type 2 event with probability 1 − p. Let X1(t) and X2(t) denote the number of type 1 and type 2 events, respectively, occurring in (0, t). Show that {X1(t), t ≥ 0} and {X2(t), t ≥ 0} are both Poisson processes with rates λp and λ(1 − p), respectively. Furthermore, the two processes are independent.

We have

First we calculate the joint probability P[X1(t) = k, X2(t) = m].

Note that

Thus, using Eq. (5.181), we obtain

Now, given that k + m events occurred, since each event has probability p of being a type 1 event and probability 1 − p of being a type 2 event, it follows that

Thus,

Then,

which indicates that X1(t) is a Poisson process with rate λp. Similarly, we can obtain

and so X2(t) is a Poisson process with rate λ(1 − p). Finally, from Eqs. (5.193), (5.194), and (5.192), we see that

Hence, X1(t) and X2(t) are independent.

Wiener Processes

5.59. Let X1, …, Xn be jointly normal r.v.’s. Show that the joint characteristic function of X1, …, Xn is given by

where ωi = E(Xi) and σik = Cov(Xi, Xk).

Let

By definition (4.66), the characteristic function of Y is

Now, by the results of Prob. 4.72, we see that Y is a normal r.v. with mean and variance given by [Eqs. (4.132) and (4.135)]

Thus, by Eq. (4.167),

Equating Eqs. (5.199) and (5.196) and setting ω = 1, we get

By replacing ai’s with ωi’s, we obtain Eq. (5.195); that is,

Let

Then we can write

and Eq. (5.195) can be expressed more compactly as

5.60. Let X1, …, Xn be jointly normal r.v.’s Let

where aik (i = 1, …, m; j = 1, …, n) are constants. Show that Y1, …, Ym are also jointly normal r.v.’s.

Let

Then Eq. (5.201) can be expressed as

Then the characteristic function for Y can be written as

Since X is a normal random vector, by Eq. (5.177) we can write

Thus,

where

Comparing Eqs. (5.200) and (5.203), we see that Eq. (5.203) is the characteristic function of a random vector Y. Hence, we conclude that Y1, …, Ym are also jointly normal r.v.’s

Note that on the basis of the above result, we can say that a random process {X(t), t ∈ T} is a normal process if every finite linear combination of the r.v.’s X(ti), ti ∈ T is normally distributed.

5.61. Show that a Wiener process X(t) is a normal process.

Consider an arbitrary linear combination

where 0 ≤ t1 < … < tn and ai are real constants. Now we write

Now from conditions 1 and 2 of Definition 5.7.1, the right-hand side of Eq. (5.206) is a linear combination of independent normal r.v.’s. Thus, based on the result of Prob. 5.60, the left-hand side of Eq. (5.206) is also a normal r.v.; that is, every finite linear combination of the r.v.’s X(ti) is a normal r.v. Thus, we conclude that the Wiener process X(t) is a normal process.

5.62. A random process {X(t), t ∈ T} is said to be continuous in probability if for every ε > 0 and t ∈ T,

Show that a Wiener process X(t) is continuous in probability.

From Chebyshev inequality (2.116), we have

Since X(t) has stationary increments, we have

in view of Eq. (5.63). Hence,

Thus, the Wiener process X(t) is continuous in probability.

Martingales

5.63. Let Y = X1 + X2 + X3 where Xi is the outcome of the ith toss of a fair coin. Verify the tower property Eq. (5.76).

Let Xi = 1 when it is a head and Xi = 0 when it is a tail. Since the coin is fair, we have

and Xi’s are independent. Now

and

Thus,

5.64. Let X1, X2, … be i.i.d. r.v.’s with mean μ. Let

Let Fn denote the information contained in X1, …, Xn. Show that

Let m < n, then by Eq. (5.71)

Since X1 + X2 + … + Xm is measurable with respect to Fm, by Eq. (5.73)

Since Xm+1 + … + Xn is independent of X1, …, Xm, by Eq. (5.75)

Thus, we obtain

5.65. Let X1, X2, … be i.i.d. r.v.’s with E(Xi) = 0 and E(Xi 2) = σ2 for all i. Let

. Let Fn denote the information contained in X1, …, Xn.

Show that

Let m < n, then by Eq. (5.71)

Since Sm is dependent only on X1, …, Xm, by Eqs. (5.73) and (5.75)

since E(Xi) = μ = 0, Var(Xi) = E(Xi 2) = σ2 and Var(Sn − Sm) = Var(Xm+1

+ … + Xn) = (n − m)σ 2. Next, by Eq. (5.74)

Thus, we obtain

5.66. Verify Eq. (5.80), that is

By condition (2) of martingale, Eq. (5.79), we have

Then by tower property Eq. (5.76)

and so on, and we obtain Eq. (5.80), that is

5.67. Verify Eq. (5.82), that is

Since {Mn, n ≥ 0} is a martingale, we have

Applying Eq. (5.77), we have

Thus, by induction we obtain

5.68. Let X1, X2, … be a sequence of independent r.v.’s with E[|Xn|] = < ∞ and E(Xn) = 0 for all n. Set S0 = 0, . Show that {Sn, n ≥ 0} is a martingale.

since E(Xn) = 0 for all n. Thus, {Sn, n ≥ 0} is a martingale.

5.69. Consider the same problem as Prob. 5.68 except E(Xn) ≥ 0 for all n. Show that {Sn, n ≥ 0} is a submartingale.

Assume max E(|Xn|) = k < ∞, then

since E(Xn) ≥ 0 for all n. Thus, {Sn, n ≥ 0} is a submartingale.

5.70. Let X1, X2, … be a sequence of Bernoulli r.v.’s with

Let . Show that (1) if then {Sn} is a martingale. (2) if then {Sn} is a submartingale, and (3) if then {Sn} is a supermartingale.

(1) If , E(Xi)= 0, and

Thus, {Sn} is a martingale.

(2) If , 0 < E(Xi) ≤ 1, and

Thus, {Sn} is a submartingale.

(3) If , E(Xi) < 0, and

Thus, {Sn} is a supermartingale.

Note that this problem represents a tossing a coin game, “heads” you win $1 and “tails” you lose $1. Thus, if p = , it is a fair coin and if p > , the game is favorable, and if p < , the game is unfavorable.

5.71. Let X1, X2, … be a sequence of i.i.d. r.v.’s with E(Xi) = μ > 0. Set

Show that {Mn, n ≥ 0} is a martingale.

Next, using Eq. (5.208) of Prob. 5.64, we have

Thus, {Mn, n ≥ 0} is a martingale.

5.72. Let X1, X2, … be i.i.d. r.v.’s with E(Xi) = 0 and E(Xi 2) = σ2 for all i. Let

S0 = 0, , and

Show that {Mn, n ≥ 0} is a martingale.

Using the triangle inequality, we have

Using Cauchy-Schwarz inequality (Eq. (4.41)), we have

Thus,

Next,

Thus, {Mn, n ≥ 0} is a martingale.

5.73. Let X1, X2, … be a sequence of i.i.d. r.v.’s with E(Xi) = μ and E(|Xi|) < ∞ for all i. Show that

is a martingale.

Thus, {Mn} is a martingale.

5.74. An urn contains initially a red and black ball. At each time n ≥ 1, a ball is taken randomly, its color noted, and both this ball and another ball of the same color are put back into the urn. Continue similarly after n draws, the urn contains n + 2 balls. Let Xn denote the number of black balls after n draws. Let Mn = Xn / (n + 2) be the fraction of black balls after n draws. Show that {Mn, n ≥ 0} is a martingale. (This is known as Polya’s Urn.)

and Xn takes values in {1, 2, …, n + 1} and .

Now,

and

Thus, {Mn, n ≥ 0} is a martingale.

5.75. Let X1, X2, … be a sequence of independent r.v.’s with

We can think of Xi as the result of a tossing a fair coin game where one wins $1 if heads come up and loses $1 if tails come up. The one way of betting strategy is to keep doubling the bet until one eventually wins. At this point one stops. (This strategy is the original martingale game.) Let Sn denote the winnings (or losses) up through n tosses. S0 = 0. Whenever one wins, one stops playing, so P(Sn+1 = 1 | Sn = 1) = 1. Show that {Sn, n ≥ 0} is a martingale—that is, the game is fair.

Suppose the first n tosses of the coin have turned up tails. So the loss Sn is given by

At this time, one double the bet again and bet 2n on the next toss. This gives

and

Thus, {Sn, n ≥ 0} is a martingale.

5.76. Let {Xn, n ≥ 0} be a martingale with respect to the filtration Fn and let g be a convex function such that E[g(Xn)] < ∞ for all n ≥ 0. Then show that the sequence {Zn, n ≥ 0} defined by

is a submartingale with respect to Fn.

By Jensen’s inequality Eq. (4.40) and the martingale property of Xn, we have

Thus, {Zn, n ≥ 0} is a submartingale.

5.77. Let Fn be a filtration and E(X) < ∞. Define

Show that {Xn, n ≥ 0} is a martingale with respect to Fn.

Thus, {Xn, n ≥ 0} is a martingale with respect to Fn.

5.78. Prove Theorem 5.8.2 (Doob decomposition).

Since X is a submartingale, we have

Let

and dn is Fn-measurable.

Set , and Mn = Xn − An. Then it is easily seen that (2), (3), and (4) of Theorem 5.82 are satisfied. Next,

Thus, (1) of Theorem 5.82 is also verified.

5.79. Let {Mn, n ≥ 0} be a martingale. Suppose that the stopping time T is bounded, that is T ≤ k. Then show that

Note that I{T=j}, the indicator function of the event {T = j}, is Fn- measurable (since we need only the information up to time n to determine if we have stopped by time n). Then we can write

and

For j ≤ k − 1, MjI{T=j} is Fk−1-measurable, thus,

Since T is known to be no more than k, the event {T = k} is the same as the event {T > k − 1 which is Fk−1-measurable. Thus,

since {Mn} be a martingale. Hence,

In a similar way, we can derive

And continue this process until we get

and finally

5.80. Verify the Optional Stopping Theorem.

Consider the stopping times Tn = min{T, n}. Note that

Hence,

Since Tn is a bounded stopping time, by Eq. (5.217), we have

and , then if E(|MT|) < ∞, (condition (1), Eq. (5.83)) we have. . Thus, by condition (3), Eq. (5.85), we get. . Hence, by Eqs. (5.219) and (5.220), we obtain

5.81. Let two gamblers, A and B, initially have a dollars and b dollars, respectively. Suppose that at each round of tossing a fair coin A wins one dollar from B if “heads” comes up, and gives one dollar to B if “tails” comes up. The game continues until either A or B runs out of money. (a) What is the probability that when the game ends, A has all the

cash? (b) What is the expected duration of the game?

(a) Let X1, X2, … be the sequence of play-by-play increments in A’s fortune; thus, Xi = ≠1 according to whether ith toss is “heads” or “tails.” The total change in A’s fortune after n plays is . The game continues until time T where T = min{n : sn = −a or +b}. It is easily seen that T is a stopping time with respect to Fn = σ(X1, X2, …, Xn) and {Sn} is a martingale with respect to Fn. (See Prob. 5.68.) Thus, by the Optional Stopping Theorem, for each n < ∞

As n → ∞, the probability of the event {T > n} converges to zero. Since Sn must be between −a and b on the event {T > n}, it follows

that E(Sn I{T > n}) converges to zero as n → ∞. Thus, letting n → ∞, we obtain

Since ST must be −a or b, we have

Solving Eqs. (5.221) and (5.222) for P(ST = −a) and P(ST = b), we obtain (cf. Prob. 5.43)

Thus, the probability that when the game ends, A has all the cash is a/(a + b).

(b) It is seen that {Sn 2 − n} is a martingale (see Prob. 5.72, σ2 = 1).

Then the Optional Stopping Theorem implies that, for each n = 1, 2, …,

Thus,

Now, as n → ∞, min(T, n) → T and ST 2 I{T≤n} → ST

2, and

Since Ssn is bounded on the event {T > n}, and since the probability of this event converges to zero as n → ∞, E(S2nI{T > n}) → 0 as n → ∞. Thus, as n → ∞, Eq. (5.225) reduces to

5.82. Let X(t) be a Poisson process with rate λ > 0. Show that x(t) − λt is a martingale.

We have

since X(t) ≥ 0 and by Eq. (5.56), E[X(t)] = λt.

Thus, x(t) − λt is a martingale.

SUPPLEMENTARY PROBLEMS

5.83. Consider a random process X(n) = {Xn, n ≥ 1}, where

and Zn are iid r.v.’s with zero mean and variance σ 2. Is X(n) stationary?

5.84. Consider a random process X(t) defined by

where Y and Θ are independent r.v.’s and are uniformly distributed over (−A, A) and (−π, π), respectively. (a) Find the mean of X(t).

(b) Find the autocorrelation function RX(t, s) of X(t).

5.85. Suppose that a random process X(t) is wide-sense stationary with autocorrelation

(a) Find the second moment of the r.v. X(5). (b) Find the second moment of the r.v. X(5) − X(3).

5.86. Consider a random process X(t) defined by

where U and V are independent r.v.’s for which

(a) Find the autocovariance function KX(t, s) of X(t). (b) Is X(t) WSS?

5.87. Consider the random processes

where A0, A1, ω0, and ω1 are constants, and r.v.’s Θ and Φ are independent and uniformly distributed over (−π, π). (a) Find the cross-correlation function of RXY(t, t + τ) of X(t) and Y(t). (b) Repeat (a) if Θ = Φ.

5.88. Given a Markov chain {Xn, n ≥ 0}, find the joint pmf

5.89. Let {Xn, n ≥ 0} be a homogeneous Markov chain. Show that

5.90. Verify Eq. (5.37).

5.91. Find Pn for the following transition probability matrices:

5.92. A certain product is made by two companies, A and B, that control the entire market. Currently, A and B have 60 percent and 40 percent, respectively, of the total market. Each year, A loses of its market share to B, while B loses of its share to A. Find the relative proportion of the market that each hold after 2 years.

5.93. Consider a Markov chain with state {0, 1, 2} and transition probability matrix

Is state 0 periodic?

5.94. Verify Eq. (5.51).

5.95. Consider a Markov chain with transition probability matrix

Find the steady-state probabilities.

5.96. Let X(t) be a Poisson process with rate λ. Find E[X2(t)].

5.97. Let X(t) be a Poisson process with rate λ. Find E{[X(t) − X(s)]2} for t > s.

5.98. Let X(t) be a Poisson process with rate λ. Find

5.99. Let Tn denote the time of the nth event of a Poisson process with rate λ. Find the variance of Tn.

5.100. Assume that customers arrive at a bank in accordance with a Poisson process with rate λ = 6 per hour, and suppose that each customer is a man with probability and a woman with probability . Now suppose that 10 men arrived in the first 2 hours. How many woman would you expect to have arrived in the first 2 hours?

5.101. Let X1,…, Xn be jointly normal r.v.’s. Let

where ci are constants. Show that Y1, …, Yn are also jointly normal r.v.’s.

5.102. Derive Eq. (5.63).

5.103. Let X1, X2, … be a sequence of Bernoulli r.v.’s in Prob. 5.70. Let Mn = Sn − n(2p − 1). Show that {Mn} is a martingale.

5.104. Let X1, X2, … be i.i.d. r.v.’s where Xi can take only two values and with equal probability. Let M0 = 1 and . Show that {Mn, n ≥ 0} is a martingale.

5.105. Consider {Xn} of Prob. 5.70 and . Let . Show

that {Yn} is a martingale.

5.106. Let X(t) be a Wiener’s process (or Brownian motion). Show that {X(t)} is a martingale.

ANSWERS TO SUPPLEMENTARY PROBLEMS

5.83. No.

5.84. (a) E[X(t)] = 0; (b)

5.85. (a) E[X2(5)] = 1; (b) E{[X(5) − X(3)]2} = 2(1 − e−1)

5.86. (a) KX(t, s) = cos(s − t); (b) No.

5.87. (a) RXY(t, t + τ)] = 0

(b)

5.88. Hint: Use Eq. (5.32).

pi0(0) pi0i1pi1i2 … pin−1in

5.89. Use the Markov property (5.27) and the homogeneity property.

5.90. Hint: Write Eq. (5.39) in terms of components.

5.91. (a)

(b)

(c)

5.92. A has 43.3 percent and B has 56.7 percent.

5.93. Hint: Draw the state transition diagram.

No.

5.94. Hint: Let , where Njk is the number of times the state k(∈ B) is occupied until absorption takes place when X(n) starts in state j(∈ B). Then ; calculate E(NJk).

5.95.

5.96. λt + λ2t2

5.97. Hint: Use the independent stationary increments condition and the result of Prob. 5.76. λ(t − s) + λ2(t − s)2

5.98.

5.99. n/λ2

5.100. 4

5.101. Hint: See Prob. 5.60.

5.102. Hint: Use condition (1) of a Wiener process and Eq. (5.102) of Prob. 5.22.

5.103. Hint: Note that Mn is the random number Sn minus its expected value.

5.104. Hint: Use definition 5.7.1.

CHAPTER 6

Analysis and Processing of Random Processes

6.1 Introduction

In this chapter, we introduce the methods for analysis and processing of random processes. First, we introduce the definitions of stochastic continuity, stochastic derivatives, and stochastic integrals of random processes. Next, the notion of power spectral density is introduced. This concept enables us to study wide-sense stationary processes in the frequency domain and define a white noise process. The response of linear systems to random processes is then studied. Finally, orthogonal and spectral representations of random processes are presented.

6.2 Continuity, Differentiation, Integration

In this section, we shall consider only the continuous-time random processes.

A. Stochastic Continuity: A random process X(t) is said to be continuous in mean square or mean square (m.s.) continuous if

The random process X(t) is m.s. continuous if and only if its autocorrelation function is continuous at t = s (Prob. 6.1). If X(t) is WSS, then it is m.s. continuous if and only if its autocorrelation function RX(τ) is continuous at τ = 0. If X(t) is m.s. continuous, then its mean is continuous; that is,

which can be written as

Hence, if X(t) is m.s. continuous, then we may interchange the ordering of the operations of expectation and limiting. Note that m.s. continuity of X(t) does not imply that the sample functions of X(t) are continuous. For instance, the Poisson process is m.s. continuous (Prob. 6.46), but sample functions of the Poisson process have a countably infinite number of discontinuities (see Fig. 5-3).

B. Stochastic Derivatives: A random process X(t) is said to have a m.s. derivative X′(t) if

where l.i.m. denotes limit in the mean (square); that is,

The m.s. derivative of X(t) exists if ∂2RX (t, s)/∂t ∂s exists at t = s (Prob. 6.6). If X(t) has the m.s. derivative X′(t), then its mean and autocorrelation function are given by

Equation (6.6) indicates that the operations of differentiation and expectation may be interchanged. If X(t) is a normal random process for which the m.s. derivative X′(t) exists, then X′(t) is also a normal random process (Prob. 6.10).

C. Stochastic Integrals: A m.s. integral of a random process X(t) is defined by

where t0 < t1 < … < t and Δti = ti+1 − ti. The m.s. integral of X(t) exists if the following integral exists (Prob. 6.11):

This implies that if X(t) is m.s. continuous, then its m.s. integral Y(t) exists (see Prob. 6.1). The mean and the autocorrelation function of Y(t) are given by

Equation (6.10) indicates that the operations of integration and expectation may be interchanged. If X(t) is a normal random process, then its integral Y(t) is also a normal random process. This follows from the fact that Σ1 X(ti) Δti is a linear combination of the jointly normal r.v.’s. (see Prob. 5.60).

6.3 Power Spectral Densities

In this section we assume that all random processes are WSS.

A. Autocorrelation Functions: The autocorrelation function of a continuous-time random process X(t) is defined as [Eq. (5.7)]

Properties of RX(τ):

Property 3 [Eq. (6.15)] is easily obtained by setting τ = 0 in Eq. (6.12). If we assume that X(t) is a voltage waveform across a 1-Ω resistor, then E[X2(t)] is the average value of power delivered to the 1-Ω resistor by X(t). Thus, E[X2(t)] is often called the average power of X(t). Properties 1 and 2 are verified in Prob. 6.13.

In case of a discrete-time random process X(n), the autocorrelation function of X(n) is defined by

Various properties of RX(k) similar to those of RX(τ) can be obtained by replacing τ by k in Eqs. (6.13) to (6.15).

B. Cross-Correlation Functions: The cross-correlation function of two continuous-time jointly WSS random processes X(t) and Y(t) is defined by

Properties of RXY(τ):

These properties are verified in Prob. 6.14. Two processes X(t) and Y(t) are called (mutually) orthogonal if

Similarly, the cross-correlation function of two discrete-time jointly WSS random processes X(n) and Y(n) is defined by

and various properties of RXY(k) similar to those of RXY(τ) can be obtained by replacing τ by k in Eqs. (6.18) to (6.20).

C. Power Spectral Density: The power spectral density (or power spectrum) SX(ω) of a continuous-time random process X(t) is defined as the Fourier transform of RX(τ):

Thus, taking the inverse Fourier transform of SX(ω), we obtain

Equations (6.23) and (6.24) are known as the Wiener-Khinchin relations.

Properties of SX(ω):

Similarly, the power spectral density SX(Ω) of a discrete-time random process X(n) is defined as the Fourier transform of RX(k):

Thus, taking the inverse Fourier transform of SX(Ω), we obtain

Properties of SX(Ω):

Note that property 1 [Eq. (6.30)] follows from the fact that e−jΩk is periodic with period 2π. Hence, it is sufficient to define SX(Ω) only in the range (−π, π).

D. Cross Power Spectral Densities: The cross power spectral density (or cross power spectrum) SXY(ω) of two continuous-time random processes X(t) and Y(t) is defined as the Fourier transform of RXY(τ):

Thus, taking the inverse Fourier transform of SXY(ω), we get

Properties of SXY(ω): Unlike SX(ω), which is a real-valued function of ω, SXY(ω), in general, is a complex-valued function.

Similarly, the cross power spectral density SXY(Ω) of two discrete-time random processes X(n) and Y(n) is defined as the Fourier transform of RXY(k):

Thus, taking the inverse Fourier transform of SXY(Ω), we get

Properties of SXY(Ω): Unlike SX(Ω), which is a real-valued function of ω, SXY(Ω), in general, is a complex-valued function.

6.4 White Noise

A continuous-time white noise process, W(t), is a WSS zero-mean continuous-time random process whose autocorrelation function is given by

where δ(τ) is a unit impulse function (or Dirac δ function) defined by

where φ(τ) is any function continuous at τ = 0. Taking the Fourier transform of Eq. (6.43), we obtain

which indicates that X(t) has a constant power spectral density (hence the name white noise). Note that the average power of W(t) is not finite.

Similarly, a WSS zero-mean discrete-time random process W(n) is called a discrete-time white noise if its autocorrelation function is given by

where δ(k) is a unit impulse sequence (or unit sample sequence) defined by

Taking the Fourier transform of Eq. (6.46), we obtain

Again the power spectral density of W(n) is a constant. Note that SW(Ω + 2π) = SW(Ω) and the average power of W(n) is σ

2 = Var[W(n)], which is finite.

6.5 Response of Linear Systems to Random Inputs

A. Linear Systems: A system is a mathematical model of a physical process that relates the input (or excitation) signal x to the output (or response) signal y. Then the system is viewed as a transformation (or mapping) of x into y. This transformation is represented by the operator T as (Fig. 6-1)

Fig. 6-1

If x and y are continuous-time signals, then the system is called a continuous-time system, and if x and y are discrete-time signals, then the system is called a discrete- time system. If the operator T is a linear operator satisfying

where α is a scalar number, then the system represented by T is called a linear system. A system is called time-invariant if a time shift in the input signal causes the same time shift in the output signal. Thus, for a continuous-time system,

for any value of t0, and for a discrete-time system,

for any integer n0. For a continuous-time linear time-invariant (LTI) system, Eq. (6.49) can be expressed as

where

is known as the impulse response of a continuous-time LTI system. The right-hand side of Eq. (6.50) is commonly called the convolution integral of h(t) and x(t), denoted by h(t) * x(t). For a discrete-time LTI system, Eq. (6.49) can be expressed as

where

is known as the impulse response (or unit sample response) of a discrete-time LTI system. The right-hand side of Eq. (6.52) is commonly called the convolution sum of h(n) and x(n), denoted by h(n) * x(n).

B. Response of a Continuous-Time Linear System to Random Input: When the input to a continuous-time linear system represented by Eq. (6.49) is a random process {X(t), t ∈ Tx}, then the output will also be a random process {Y(t), t ∈ Ty}; that is,

For any input sample function xi(t), the corresponding output sample function is

If the system is LTI, then by Eq. (6.50), we can write

Note that Eq. (6.56) is a stochastic integral. Then

The autocorrelation function of Y(t) is given by (Prob. 6.24)

If the input X(t) is WSS, then from Eq. (6.57),

where H(0) = H(ω)|ω = 0 and H(ω) is the frequency response of the system defined by the Fourier transform of h(t); that is,

The autocorrelation function of Y(t) is, from Eq. (6.58),

Setting s = t + τ, we get

From Eqs. (6.59) and (6.62), we see that the output Y(t) is also WSS. Taking the Fourier transform of Eq. (6.62), the power spectral density of Y(t) is given by (Prob. 6.25)

Thus, we obtain the important result that the power spectral density of the output is the product of the power spectral density of the input and the magnitude squared of the frequency response of the system.

When the autocorrelation function of the output RY(τ) is desired, it is easier to determine the power spectral density SY(ω) and then evaluate the inverse Fourier transform (Prob. 6.26). Thus,

By Eq. (6.15), the average power in the output Y(t) is

C. Response of a Discrete-Time Linear System to Random Input: When the input to a discrete-time LTI system is a discrete-time random process X(n), then by Eq. (6.52), the output Y(n) is

The autocorrelation function of Y(n) is given by

When X(n) is WSS, then from Eq. (6.66),

where H(0) = H(Ω)|Ω = 0 and H(Ω) is the frequency response of the system defined by the Fourier transform of h(n):

The autocorrelation function of Y(n) is, from Eq. (6.67),

Setting m = n + k, we get

From Eqs. (6.68) and (6.71), we see that the output Y(n) is also WSS. Taking the Fourier transform of Eq. (6.71), the power spectral density of Y(n) is given by (Prob. 6.28)

which is the same as Eq. (6.63).

6.6 Fourier Series and Karhunen-Loéve Expansions

A. Stochastic Periodicity: A continuous-time random process X(t) is said to be m.s. periodic with period T if

If X(t) is WSS, then X(t) is m.s. periodic if and only if its autocorrelation function is periodic with period T; that is,

B. Fourier Series: Let X(t) be a WSS random process with periodic RX(τ) having period T. Expanding RX(τ) into a Fourier series, we obtain

where

Let be expressed as

where Xn are r.v.’s given by

Note that, in general, Xn are complex-valued r.v.’s. For complex-valued r.v.’s, the correlation between two r.v.’s X and Y is defined by E(XY*). Then is called the m.s. Fourier series of X(t) such that (Prob. 6.34)

Furthermore, we have (Prob. 6.33)

C. Karhunen-Loéve Expansion Consider a random process X(t) which is not periodic. Let be expressed as

where a set of functions {φn(t)} is orthonormal on an interval (0, T) such that

and Xn are r.v.’s given by

Then is called the Karhunen-Loéve expansion of X(t) such that (Prob. 6.38)

Let RX(t, s) be the autocorrelation function of X(t), and consider the following integral equation:

where λn and φn(t) are called the eigenvalues and the corresponding eigenfunctions of the integral equation (6.86). It is known from the theory of integral equations that if RX(t, s) is continuous, then φn(t) of Eq. (6.86) are orthonormal as in Eq. (6.83), and they satisfy the following identity:

which is known as Mercer’s theorem. With the above results, we can show that Eq. (6.85) is satisfied and the

coefficient Xn are orthogonal r.v.’s (Prob. 6.37); that is,

6.7 Fourier Transform of Random Processes

A. Continuous-Time Random Processes: The Fourier transform of a continuous-time random process X(t) is a random process given by

which is the stochastic integral, and the integral is interpreted as a m.s. limit; that is,

Note that is a complex random process. Similarly, the inverse Fourier transform

is also a stochastic integral and should also be interpreted in the m.s. sense. The properties of continuous-time Fourier transforms (Appendix B) also hold for random processes (or random signals). For instance, if Y(t) is the output of a continuous-time LTI system with input X(t), then

where H(ω) is the frequency response of the system. Let be the two-dimensional Fourier transform of RX(t, s); that is,

Then the autocorrelation function of is given by (Prob. 6.41)

If X(t) is real, then

If X(t) is a WSS random process with autocorrelation function RX(t, s) = RX(t − s) = RX(τ) and power spectral density SX(ω), then (Prob. 6.42)

Equation (6.99) shows that the Fourier transform of a WSS random process is nonstationary white noise.

B. Discrete-Time Random Processes: The Fourier transform of a discrete-time random process X(n) is a random process

given by (in m.s. sense)

Similarly, the inverse Fourier transform

should also be interpreted in the m.s. sense. Note that and the properties of discrete-time Fourier transforms (Appendix B) also hold for discrete-time random signals. For instance, if Y(n) is the output of a discrete-time LTI system with input X(n), then

where H(Ω) is the frequency response of the system. Let be the two-dimensional Fourier transform of RX(n, m):

Then the autocorrelation function of is given by (Prob. 6.44)

If X(n) is a WSS random process with autocorrelation function RX(n, m) = RX(n − m) = RX(k) and power spectral density SX(Ω), then

Equation (6.106) shows that the Fourier transform of a discrete-time WSS random process is nonstationary white noise.

SOLVED PROBLEMS

Continuity, Differentiation, Integration

6.1. Show that the random process X(t) is m.s. continuous if and only if its autocorrelation function RX(t, s) is continuous.

We can write

Thus, if RX(t, s) is continuous, then

and X(t) is m.s. continuous. Next, consider

Applying Cauchy-Schwarz inequality (3.97) (Prob. 3.35), we obtain

Thus, if X(t) is m.s. continuous, then by Eq. (6.1) we have

that is, RX(t, s) is continuous. This completes the proof.

6.2. Show that a WSS random process X(t) is m.s. continuous if and only if its autocorrelation function RX(τ) is continuous at τ = 0.

If X(t) is WSS, then Eq. (6.107) becomes

Thus if RX(τ) is continuous at τ = 0, that is,

that is, X(t) is m.s. continuous. Similarly, we can show that if X(t) is m.s. continuous, then by Eq. (6.108), RX(τ) is continuous at τ = 0.

6.3. Show that if X(t) is m.s. continuous, then its mean is continuous; that is,

We have

Thus,

If X(t) is m.s. continuous, then as ε → 0, the left-hand side of the above expression approaches zero. Thus,

6.4. Show that the Wiener process X(t) is m.s. continuous.

From Eq. (5.64), the autocorrelation function of the Wiener process X(t) is given by

Thus, we have

Since

RX(t, s) is continuous. Hence, the Wiener process X(t) is m.s. continuous.

6.5. Show that every m.s. continuous random process is continuous in probability.

A random process X(t) is continuous in probability if, for every t and a > 0 (see Prob. 5.62),

Applying Chebyshev inequality (2.116) (Prob. 2.39), we have

Now, if X(t) is m.s. continuous, then the right-hand side goes to 0 as ε → 0, which implies that the left-hand side must also go to 0 as ε → 0. Thus, we have proved that if X(t) is m.s. continuous, then it is also continuous in probability.

6.6. Show that a random process X(t) has a m.s. derivative X′(t) if ∂2RX(t, s)/∂t ∂s exists at s = t.

Let

By the Cauchy criterion (see the note at the end of this solution), the m.s. derivative X′(t) exists if

and

Thus,

provided ∂2RX(t, s)/∂t ∂s exists at s = t. Setting ε1 = ε2 in Eq. (6.112), we get

and by Eq. (6.111), we obtain

Thus, we conclude that X(t) has a m.s. derivative X′(t) if ∂2RX(t, s)/∂t ∂s exists at s = t. If X(t) is WSS, then the above conclusion is equivalent to the existence of ∂2RX(τ)/ ∂

2τ at τ = 0.

Note: In real analysis, a function g(ε) of some parameter ε converges to a finite value if

This is known as the Cauchy criterion.

6.7. Suppose a random process X(t) has a m.s. derivative X′(t). (a) Find E[X′(t)]. (b) Find the cross-correlation function of X(t) and X′(t). (c) Find the autocorrelation function of X′(t).

(a) We have

(b) From Eq. (6.17), the cross-correlation function of X(t) and X′(t) is

(c) Using Eq. (6.115), the autocorrelation function of X′(t) is

6.8. If X(t) is a WSS random process and has a m.s. derivative X′(t), then show that

(a) For a WSS process X(t), RX(t, s) = RX(s − t). Thus, setting s − t = τ in Eq. (6.115) of Prob. 6.7, we obtain ∂RX(s − t)/∂s = dRX(τ)/dτ and

(b) Now ∂RX(s − t)/∂t = −dRX(τ)/dτ. Thus, ∂ 2RX(s − t)/ ∂t ∂s = d

2RX(τ)/d τ 2,

and by Eq. (6.116) of Prob. 6.7, we have

6.9. Show that the Wiener process X(t) does not have a m.s. derivative.

From Eq. (5.64), the autocorrelation function of the Wiener process X(t) is given by

Thus,

where u(t − s) is a unit step function defined by

and it is not continuous at s = t (Fig. 6-2). Thus, ∂2 RX(t, s)/∂t ∂s does not exist at s = t, and the Wiener process X(t) does not have a m.s. derivative.

Fig. 6-2 Shifted unit step function.

Note that although a m.s. derivative does not exist for the Wiener process, we can define a generalized derivative of the Wiener process (see Prob. 6.20).

6.10. Show that if X(t) is a normal random process for which the m.s. derivative X′ (t) exists, then X′(t) is also a normal random process.

Let X(t) be a normal random process. Now consider

Then, n r.v.’s Yε(t1), Yε(t2), …, Yε(tn) are given by a linear transformation of the jointly normal r.v.’s X(t1), X(t1 + ε), X(t2), X(t2 + ε), …, X(tn), X(tn + ε). It then follows by the result of Prob. 5.60 that Yε(t1), Yε(t2), …, Yε(tn) are

jointly normal r.v.’s, and hence Yε(t) is a normal random process. Thus, we conclude that the m.s. derivative X′(t), which is the limit of Yε(t) as ε → 0, is also a normal random process, since m.s. convergence implies convergence in probability (see Prob. 6.5).

6.11. Show that the m.s. integral of a random process X(t) exists if the following integral exists:

A m.s. integral of X(t) is defined by [Eq. (6.8)]

Again using the Cauchy criterion, the m.s. integral Y(t) of X(t) exists if

As in the case of the m.s. derivative [Eq. (6.111)], expanding the square, we obtain

and Eq. (6.120) holds if

exists, or, equivalently,

exists.

6.12. Let X(t) be the Wiener process with parameter σ2. Let

(a) Find the mean and the variance of Y(t). (b) Find the autocorrelation function of Y(t).

(a) By assumption 3 of the Wiener process (Sec. 5.7), that is, E[X(t)] = 0, we have

Then

By Eq. (5.64), RX(α, β) = σ 2 min(α, β); thus, referring to Fig. 6-3, we

obtain

Fig. 6-3

(b) Let t > s ≥ 0 and write

Then, for t > s ≥ 0,

Now by Eq. (6.122),

Using assumptions 1, 3, and 4 of the Wiener process (Sec. 5.7), and since s ≤ α ≤ t, we have

Finally, for 0 ≤ β ≤ s,

Substituting these results into Eq. (6.123), we get

Since RY(t, s) = RY(s, t), we obtain

Power Spectral Density

6.13. Verify Eqs. (6.13) and (6.14).

From Eq. (6.12),

Setting t + τ = s, we get

Next, we have

Expanding the square, we have

from which we obtain Eq. (6.14); that is,

6.14. Verify Eqs. (6.18) to (6.20).

By Eq. (6.17),

Setting t − τ = s, we get

Next, from the Cauchy-Schwarz inequality, Eq. (3.97) (Prob. 3.35), it follows that

from which we obtain Eq. (6.19); that is,

Expanding the square, we have

from which we obtain Eq. (6.20); that is,

6.15. Two random processes X(t) and Y(t) are given by

where A and ω are constants and Θ is a uniform r.v. over (0, 2π). Find the cross-correlation function of X(t) and Y(t) and verify Eq. (6.18).

From Eq. (6.17), the cross-correlation function of X(t) and Y(t) is

Similarly,

From Eqs. (6.125) and (6.126), we see that

which verifies Eq. (6.18).

6.16. Show that the power spectrum of a (real) random process X(t) is real and verify Eq. (6.26).

From Eq. (6.23) and expanding the exponential, we have

Since RX(− τ) = RX(τ), RX(τ) cos ωτ is an even function of τ and RX(τ) sin ωτ is an odd function of τ, and hence the imaginary term in Eq. (6.127) vanishes and we obtain

which indicates that SX(ω) is real. Since cos(−ωτ) = cos(ωτ), it follows that

which indicates that the power spectrum of a real random process X(t) is an even function of frequency.

6.17. Consider the random process

where X(t) is a Poisson process with rate λ. Thus, Y(t) starts at Y(0) = 1 and switches back and forth from +1 to −1 at random Poisson times Ti, as shown in Fig. 6-4. The process Y(t) is known as the semirandom telegraph signal because its initial value Y(0) = 1 is not random.

Fig. 6-4 Semirandom telegraph signal.

(a) Find the mean of Y(t). (b) Find the autocorrelation function of Y(t).

(a) We have

Thus, using Eq. (5.55), we have

Hence,

(b) Similarly, since Y(t)Y(t + τ) = 1 if there are an even number of events in (t, t + τ) for τ > 0 and Y(t)Y(t + τ) = −1 if there are an odd number of events, then for t > 0 and t + τ > 0,

which indicates that RY(t, t + τ) = RY(τ), and by Eq. (6.13),

Note that since E[Y(t)] is not a constant, Y(t) is not WSS.

6.18. Consider the random process

where Y(t) is the semirandom telegraph signal of Prob. 6.17 and A is a r.v. independent of Y(t) and takes on the values ≠1 with equal probability. The process Z(t) is known as the random telegraph signal. (a) Show that Z(t) is WSS. (b) Find the power spectral density of Z(t).

(a) Since E(A) = 0 and E(A2) = 1, the mean of Z(t) is

and the autocorrelation of Z(t) is

Thus, using Eq. (6.130), we obtain

Thus, we see that Z(t) is WSS. (b) Taking the Fourier transform of Eq. (6.132) (see Appendix B), we see

that the power spectrum of Z(t) is given by

6.19. Let X(t) and Y(t) be both zero-mean and WSS random processes. Consider the random process Z(t) defined by

(a) Determine the autocorrelation function and the power spectral density of Z(t), (i) if X(t) and Y(t) are jointly WSS; (ii) if X(t) and Y(t) are orthogonal.

(b) Show that if X(t) and Y(t) are orthogonal, then the mean square of Z(t) is equal to the sum of the mean squares of X(t) and Y(t).

(a) The autocorrelation of Z(t) is given by

(i) If X(t) and Y(t) are jointly WSS, then we have

where τ = s − t. Taking the Fourier transform of the above expression, we obtain

(ii) If X(t) and Y(t) are orthogonal [Eq. (6.21)],

Then

(b) Setting τ = 0 in Eq. (6.134a), and using Eq. (6.15), we get

which indicates that the mean square of Z(t) is equal to the sum of the mean squares of X(t) and Y(t).

White Noise

6.20. Using the notion of generalized derivative, show that the generalized derivative X′(t) of the Wiener process X(t) is a white noise.

From Eq. (5.64),

and from Eq. (6.119) (Prob. 6.9), we have

Now, using the δ function, the generalized derivative of a unit step function u(t) is given by

Applying the above relation to Eq. (6.135), we obtain

which is, by Eq. (6.116) (Prob. 6.7), the autocorrelation function of the generalized derivative X′(t) of the Wiener process X(t); that is,

where τ = t − s. Thus, by definition (6.43), we see that the generalized derivative X′(t) of the Wiener process X(t) is a white noise.

Recall that the Wiener process is a normal process and its derivative is also normal (see Prob. 6.10). Hence, the generalized derivative X′(t) of the Wiener process is called white normal (or white Gaussian) noise.

6.21. Let X(t) be a Poisson process with rate λ. Let

Show that the generalized derivative Y′(t) of Y(t) is a white noise.

Since Y(t) = X(t) − λt, we have formally

Then

Now, from Eqs. (5.56) and (5.60), we have

Thus,

and from Eqs. (6.7) and (6.137),

Substituting Eq. (6.141) into Eq. (6.139), we obtain

Substituting Eqs. (6.141) and (6.142) into Eq. (6.140), we get

Hence, we see that Y′(t) is a zero-mean WSS random process, and by definition (6.43), Y′(t) is a white noise with σ2 = λ. The process Y′(t) is known as the Poisson white noise.

6.22. Let X(t) be a white normal noise. Let

(a) Find the autocorrelation function of Y(t). (b) Show that Y(t) is the Wiener process.

(a) From Eq. (6.137) of Prob. 6.20,

Thus, by Eq. (6.11), the autocorrelation function of Y(t) is

(b) Comparing Eq. (6.145) and Eq. (5.64), we see that Y(t) has the same autocorrelation function as the Wiener process. In addition, Y(t) is normal, since X(t) is a normal process and Y(0) = 0. Thus, we conclude that Y(t) is the Wiener process.

6.23. Let Y(n) = X(n) + W(n), where X(n) = A (for all n) and A is a r.v. with zero mean and variance , and W(n) is a discrete-time white noise with average power σ2. It is also assumed that X(n) and W(n) are independent. (a) Show that Y(n) is WSS. (b) Find the power spectral density SY(Ω) of Y(n).

(a) The mean of Y(n) is

The autocorrelation function of Y(n) is

Thus, Y(n) is WSS. (b) Taking the Fourier transform of Eq. (6.146), we obtain

Response of Linear Systems to Random Inputs

6.24. Derive Eq. (6.58).

Using Eq. (6.56), we have

6.25. Derive Eq. (6.63).

From Eq. (6.62), we have

Taking the Fourier transform of RY(τ), we obtain

Letting τ + α − β = λ, we get

6.26. A WSS random process X(t) with autocorrelation function

where a is a real positive constant, is applied to the input of an LTI system with impulse response

where b is a real positive constant. Find the autocorrelation function of the output Y(t) of the system.

The frequency response H(ω) of the system is

The power spectral density of X(t) is

By Eq. (6.63), the power spectral density of Y(t) is

Taking the inverse Fourier transform of both sides of the above equation, we obtain

6.27. Verify Eq. (6.25), that is, the power spectral density of any WSS process X(t) is real and SX(ω) ≥ 0.

The realness of SX(ω) was shown in Prob. 6.16. Consider an ideal bandpass filter with frequency response (Fig. 6-5)

Fig. 6-5

with a random process X(t) as its input.

From Eq. (6.63), it follows that the power spectral density SY(ω) of the output Y(t) equals

Hence, from Eq. (6.27), we have

which indicates that the area of SX(ω) in any interval of ω is nonnegative. This is possible only if SX(ω) ≥ 0 for every ω.

6.28. Verify Eq. (6.72).

From Eq. (6.71), we have

Taking the Fourier transform of RY(k), we obtain

Letting k + i − l = n, we get

6.29. The discrete-time system shown in Fig. 6-6 consists of one unit delay element and one scalar multiplier (a < 1). The input X(n) is discrete-time white noise with average power σ2. Find the spectral density and average power of the output Y(n).

Fig. 6-6

From Fig. 6-6, Y(n) and X(n) are related by

The impulse response h(n) of the system is defined by

Solving Eq. (6.149), we obtain

where u(n) is the unit step sequence defined by

Taking the Fourier transform of Eq. (6.150), we obtain

Now, by Eq. (6.48),

and by Eq. (6.72), the power spectral density of Y(n) is

Taking the inverse Fourier transform of Eq. (6.151), we obtain

Thus, by Eq. (6.33), the average power of Y(n) is

6.30. Let Y(t) be the output of an LTI system with impulse response h(t), when X(t) is applied as input. Show that

(a) Using Eq. (6.56), we have

(b) Similarly,

6.31. Let Y(t) be the output of an LTI system with impulse response h(t) when a WSS random process X(t) is applied as input. Show that

(a) If X(t) is WSS, then Eq. (6.152) of Prob. 6.30 becomes

which indicates that RXY(t, s) is a function of the time difference τ = s − t only. Hence,

Taking the Fourier transform of Eq. (6.157), we obtain

(b) Similarly, if X(t) is WSS, then by Eq. (6.156), Eq. (6.153) becomes

which indicates that RY(t, s) is a function of the time difference τ = s − t only. Hence,

Taking the Fourier transform of RY(τ), we obtain

Note that from Eqs. (6.154) and (6.155), we obtain Eq. (6.63); that is,

6.32. Consider a WSS process X(t) with autocorrelation function RX(τ) and power spectral density SX(ω). Let X′(t) = dX(t)/dt. Show that

(a) If X(t) is the input to a differentiator, then its output is Y(t) = X′(t). The frequency response of a differentiator is known as H(ω) = jω. Then from Eq. (6.154),

Taking the inverse Fourier transform of both sides, we obtain

(b) From Eq. (6.155),

Again taking the inverse Fourier transform of both sides and using the result of part (a), we have

(c) From Eq. (6.63),

Note that Eqs. (6.159) and (6.160) were proved in Prob. 6.8 by a different method.

Fourier Series and Karhunen-Loéve Expansions

6.33. Verify Eqs. (6.80) and (6.81).

From Eq. (6.78),

Since X(t) is WSS, E[X(t)] = μX, and we have

Again using Eq. (6.78), we have

Now

Letting t − s = τ, and using Eq. (6.76), we obtain

Thus,

6.34. Let be the Fourier series representation of X(t) shown in Eq. (6.77). Verify Eq. (6.79).

From Eq. (6.77), we have

Now, by Eqs. (6.81) and (6.162), we have

Using these results, finally we obtain

since each sum above equals RX(0) [see Eq. (6.75)].

6.35. Let X(t) be m.s. periodic and represented by the Fourier series [Eq. (6.77)]

Show that

From Eq. (6.81), we have

Setting τ = 0 in Eq. (6.75), we obtain

Equation (6.163) is known as Parseval’s theorem for the Fourier series.

6.36. If a random process X(t) is represented by a Karhunen-Loéve expansion [Eq. (6.82)]

and Xn’s are orthogonal, show that φn(t) must satisfy integral equation (6.86); that is,

Consider

Then

since Xn’s are orthogonal; that is, if m ≠ n. But by Eq. (6.84),

Thus, equating Eqs. (6.165) and (6.166), we obtain

where λn = E(|Xn| 2).

6.37. Let be the Karhunen-Loéve expansion of X(t) shown in Eq. (6.82). Verify Eq. (6.88).

From Eqs. (6.166) and (6.86), we have

Now by Eqs. (6.83), (6.84), and (6.167) we obtain

6.38. Let be the Karhunen-Loéve expansion of X(t) shown in Eq. (6.82). Verify Eq. (6.85).

From Eq. (6.82), we have

Using Eqs. (6.167) and (6.168), we have

since by Mercer’s theorem [Eq. (6.87)]

and

6.39. Find the Karhunen-Loéve expansion of the Wiener process X(t).

From Eq. (5.64),

Substituting the above expression into Eq. (6.86), we obtain

or

Differentiating Eq. (6.170) with respect to t, we get

Differentiating Eq. (6.171) with respect to t again, we obtain

A general solution of Eq. (6.172) is

In order to determine the values of an, bn, and λn (or ωn), we need appropriate boundary conditions. From Eq. (6.170), we see that φn(0) = 0. This implies that bn = 0. From Eq. (6.171), we see that φn(T) = 0. This implies that

Therefore, the eigenvalues are given by

The normalization requirement [Eq. (6.83)] implies that

Thus, the eigenfunctions are given by

and the Karhunen-Loéve expansion of the Wiener process X(t) is

where Xn are given by

and they are uncorrelated with variance λn.

6.40. Find the Karhunen-Loéve expansion of the white normal (or white Gaussian) noise W(t).

From Eq. (6.43),

Substituting the above expression into Eq. (6.86), we obtain

or [by Eq. (6.44)]

which indicates that all λn = σ 2 and φn(t) are arbitrary. Thus, any complete

orthogonal set {φn(t)} with corresponding eigenvalues λn = σ 2 can be used in

the Karhunen-Loéve expansion of the white Gaussian noise.

Fourier Transform of Random Processes

6.41. Derive Eq. (6.94).

From Eq. (6.89),

Then

in view of Eq. (6.93).

6.42. Derive Eqs. (6.98) and (6.99).

Since X(t) is WSS, by Eq. (6.93), and letting t − s = τ, we have

From the Fourier transform pair (Appendix B) 1 ↔ 2πδ(ω), we have

Hence,

Next, from Eq. (6.94) and the above result, we obtain

6.43. Let be the Fourier transform of a random process X(t). If is a white noise with zero mean and autocorrelation function q(ω1) δ (ω1 − ω2), then show that X(t) is WSS with power spectral density q(ω)/2π.

By Eq. (6.91),

Then

Assuming that X(t) is a complex random process, we have

which depends only on t − s = τ. Hence, we conclude that X(t) is WSS. Setting t − s = τ and ω1 = ω in Eq. (6.178), we have

in view of Eq. (6.24). Thus, we obtain SX(ω) = q(ω)/2π.

6.44. Verify Eq. (6.104).

By Eq. (6.100),

Then

in view of Eq. (6.103).

6.45. Derive Eqs. (6.105) and (6.106).

If X(n) is WSS, then RX(n, m) = RX(n − m). By Eq. (6.103), and letting n − m = k, we have

From the Fourier transform pair (Appendix B) x(n) = 1 ↔ 2πδ(Ω), we have

Hence,

Next, from Eq. (6.104) and the above result, we obtain

SUPPLEMENTARY PROBLEMS

6.46. Is the Poisson process X(t) m.s. continuous?

6.47. Let X(t) be defined by (Prob. 5.4)

where Y is a uniform r.v. over (0, 1) and ω is a constant. (a) Is X(t) m.s. continuous? (b) Does X(t) have a m.s. derivative?

6.48. Let Z(t) be the random telegraph signal of Prob. 6.18. (a) Is Z(t) m.s. continuous? (b) Does Z(t) have a m.s. derivative?

6.49. Let X(t) be a WSS random process, and let X′(t) be its m.s. derivative. Show that E[X(t)X′(t)] = 0.

6.50. Let

where X(t) is given by Prob. 6.47 with ω = 2π/T. (a) Find the mean of Z(t). (b) Find the autocorrelation function of Z(t).

6.51. Consider a WSS random process X(t) with E[X(t)] = μX. Let

The process X(t) is said to be ergodic in the mean if

Find E [〈X(t)〉T].

6.52. Let X(t) = A cos(ω0 t + Θ), where A and ω0 are constants, Θ is a uniform r.v. over (−π, π) (Prob. 5.20). Find the power spectral density of X(t).

6.53. A random process Y(t) is defined by

where A and ωc are constants, Θ is a uniform r.v. over (−π, π), and X(t) is a zero-mean WSS random process with the autocorrelation function RX(τ) and the power spectral density SX(ω). Furthermore, X(t) and Θ are

independent. Show that Y(t) is WSS, and find the power spectral density of Y(t).

6.54. Consider a discrete-time random process defined by

where ai and Ωi are real constants and Θi are independent uniform r.v.’s over (−π, π).

(a) Find the mean of X(n). (b) Find the autocorrelation function of X(n).

6.55. Consider a discrete-time WSS random process X(n) with the autocorrelation function

Find the power spectral density of X(n).

6.56. Let X(t) and Y(t) be defined by

where ω0 is constant and U and V are independent r.v.’s both having zero mean and variance σ2.

(a) Find the cross-correlation function of X(t) and Y(t). (b) Find the cross power spectral density of X(t) and Y(t).

6.57. Verify Eqs. (6.36) and (6.37).

6.58. Let Y(t) = X(t) + W(t), where X(t) and W(t) are orthogonal and W(t) is a white noise specified by Eq. (6.43) or (6.45). Find the autocorrelation function of Y(t).

6.59. A zero-mean WSS random process X(t) is called band-limited white noise if its spectral density is given by Find the autocorrelation function of X(t).

6.60. A WSS random process X(t) is applied to the input of an LTI system with impulse response h(t) = 3e−2tu(t). Find the mean value of Y(t) of the system if E[X(t)] = 2.

6.61. The input X(t) to the RC filter shown in Fig. 6-7 is a white noise specified by Eq. (6.45). Find the mean-square value of Y(t).

Fig. 6-7 RC filter.

6.62. The input X(t) to a differentiator is the random telegraph signal of Prob. 6.18. (a) Determine the power spectral density of the differentiator output. (b) Find the mean-square value of the differentiator output.

6.63. Suppose that the input to the filter shown in Fig. 6-8 is a white noise specified by Eq. (6.45). Find the power spectral density of Y(t).

Fig. 6-8

6.64. Verify Eq. (6.67).

6.65. Suppose that the input to the discrete-time filter shown in Fig. 6-9 is a discrete-time white noise with average power σ2. Find the power spectral density of Y(n).

Fig. 6-9

6.66. Using the Karhunen-Loéve expansion of the Wiener process, obtain the Karhunen-Loéve expansion of the white normal noise.

6.67. Let Y(t) = X(t) + W(t), where X(t) and W(t) are orthogonal and W(t) is a white noise specified by Eq. (6.43) or (6.45). Let φn(t) be the eigenfunctions of the integral equation (6.86) and λn the corresponding eigenvalues. (a) Show that φn(t) are also the eigenfunctions of the integral equation for

the Karhunen-Loéve expansion of Y(t) with RY(t, s). (b) Find the corresponding eigenvalues.

6.68. Suppose that

where Xn are r.v.’s and ω0 is a constant. Find the Fourier transform of X(t).

6.69. Let be the Fourier transform of a continuous-time random process X(t). Find the mean of .

6.70. Let

where E[X(n)] = 0 and . Find the mean and the autocorrelation function of .

ANSWERS TO SUPPLEMENTARY PROBLEMS

6.46. Hint: Use Eq. (5.60) and proceed as in Prob. 6.4.

Yes.

6.47. Hint: Use Eq. (5.87) of Prob. 5.12. (a) Yes; (b) yes.

6.48. Hint: Use Eq. (6.132) of Prob. 6.18. (a) Yes; (b) no.

6.49. Hint: Use Eqs. (6.13) [or (6.14)] and (6.117).

6.50.

6.51. μX

6.52.

6.53.

6.54.

6.55.

6.56. (a) RXY(t, t + τ) = −σ 2 sin ω0τ

(b) SXY(ω) = jσ 2π[δ (ω − ω0) − δ(ω + ω0)]

6.57. Hint: Substitute Eq. (6.18) into Eq. (6.34).

6.58. RY(t, s) = RX(t, s) + σ 2δ(t − s)

6.59.

6.60. Hint: Use Eq. (6.59).

3

6.61. Hint: Use Eqs. (6.64) and (6.65). σ2/(2RC)

6.62.

6.63. SY(ω) = σ 2(1 + a2 + 2a cos ωT)

6.64. Hint: Proceed as in Prob. 6.24.

6.65. SY(Ω) = σ 2(1 + a2 + 2a cos Ω)

6.66. Hint: Take the derivative of Eq. (6.175) of Prob. 6.39.

where Wn are independent normal r.v.’s with the same variance σ 2.

6.67. Hint: Use the result of Prob. 6.58. (b) λn + σ

2

6.68.

6.69.

6.70.

CHAPTER 7

Estimation Theory

7.1 Introduction

In this chapter, we present a classical estimation theory. There are two basic types of estimation problems. In the first type, we are interested in estimating the parameters of one or more r.v.’s, and in the second type, we are interested in estimating the value of an inaccessible r.v. Y in terms of the observation of an accessible r.v. X.

7.2 Parameter Estimation

Let X be a r.v. with pdf ƒ(x) and X1, …, Xn a set of n independent r.v.’s each with pdf ƒ(x). The set of r.v.’s(X1, …, Xn) is called a random sample (or sample vector) of size n of X. Any real-valued function of a random sample s(X1, …, Xn) is called a statistic.

Let X be a r.v. with pdf ƒ(x; θ) which depends on an unknown parameter θ. Let (X1, …, Xn) be a random sample of X. In this case, the joint pdf of X1, …, Xn is given by

where x1, …, xn are the values of the observed data taken from the random sample. An estimator of θ is any statistic s(X1, …, Xn), denoted as

For a particular set of observations X1 = x1, …, Xn = xn, the value of the estimator s(x1, …, xn) will be called an estimate of θ and denoted by . Thus, an estimator is a r.v. and an estimate is a particular realization of it. It is not necessary that an estimate of a parameter be one single value; instead, the estimate could be a range of values. Estimates which specify a single value are called point estimates, and estimates which specify a range of values are called interval estimates.

7.3 Properties of Point Estimators

A. Unbiased Estimators:

An estimator Θ = s(X1, …, Xn) is said to be an unbiased estimator of the parameter θ if

for all possible values of θ. If Θ is an unbiased estimator, then its mean square error is given by

That is, its mean square error equals its variance.

B. Efficient Estimators: An estimator Θ1 is said to be a more efficient estimator of the parameter θ than the estimator Θ2 if

1. Θ1 and Θ2 are both unbiased estimators of θ. 2. Var(Θ1) < Var(Θ2).

The estimator ΘMV = s(X1, …, Xn) is said to be a most efficient (or minimum variance) unbiased estimator of the parameter θ if

1. It is an unbiased estimator of θ. 2. Var(ΘMV) ≤ Var(Θ) for all Θ.

C. Consistent Estimators: The estimator Θn of θ based on a random sample of size n is said to be consistent if for any small ε > 0,

or equivalently,

The following two conditions are sufficient to define consistency (Prob. 7.5):

7.4 Maximum-Likelihood Estimation

Let ƒ(x; θ) = ƒ(x1, …, xn; θ) denote the joint pmf of the r.v.’s X1, …, Xn when they are discrete, and let it be their joint pdf when they are continuous. Let

Now L(θ) represents the likelihood that the values x1, …, xn will be observed when θ is the true value of the parameter. Thus, L(θ) is often referred to as the likelihood function of the random sample. Let θML = s(x1, …, xn) be the maximizing value of L(θ); that is,

Then the maximum-likelihood estimator of θ is

and θML is the maximum-likelihood estimate of θ. Since L(θ) is a product of either pmf’s or pdf’s, it will always be positive (for

the range of possible values of θ). Thus, ln L(θ) can always be defined, and in determining the maximizing value of θ, it is often useful to use the fact that L(θ)

and ln L(θ) have their maximum at the same value of θ. Hence, we may also obtain θML by maximizing ln L(θ).

7.5 Bayes’ Estimation

Suppose that the unknown parameter θ is considered to be a r.v. having some fixed distribution or prior pdf ƒ(θ). Then ƒ(x; θ) is now viewed as a conditional pdf and written as ƒ(x | θ), and we can express the joint pdf of the random sample (X1, …, Xn) and θ as

and the marginal pdf of the sample is given by

where Rθ is the range of the possible value of θ. The other conditional pdf,

is referred to as the posterior pdf of θ. Thus, the prior pdf ƒ(θ) represents our information about θ prior to the observation of the outcomes of X1, …, Xn, and the posterior pdf ƒ(θ |x1, …, xn) represents our information about θ after having observed the sample. The conditional mean of θ, defined by

is called the Bayes’ estimate of θ, and

is called the Bayes’ estimator of θ.

7.6 Mean Square Estimation

In this section, we deal with the second type of estimation problem—that is, estimating the value of an inaccessible r.v. Y in terms of the observation of an accessible r.v. X. In general, the estimator of Y is given by a function of X, g(X). Then is called the estimation error, and there is a cost associated with this error, C[Y − g(X)]. We are interested in finding the function g(X) that minimizes this cost. When X and Y are continuous r.v.’s, the mean square (m.s.) error is often used as the cost function,

It can be shown that the estimator of Y given by (Prob. 7.17),

is the best estimator in the sense that the m.s. error defined by Eq. (7.17) is a minimum.

7.7 Linear Mean Square Estimation

Now consider the estimator of Y given by

We would like to find the values of a and b such that the m.s. error defined by

is minimum. We maintain that a and b must be such that (Prob. 7.20)

and a and b are given by

and the minimum m.s. error em is (Prob. 7.22)

where σXY = Cov(X, Y) and ρXY is the correlation coefficient of X and Y. Note that Eq. (7.21) states that the optimum linear m.s. estimator of Y is such that the estimation error is orthogonal to the observation X. This is known as the orthogonality principle. The line y = ax + b is often called a regression line.

Next, we consider the estimator of Y with a linear combination of the random sample (X1, …, Xn) by

Again, we maintain that in order to produce the linear estimator with the minimum m.s. error, the coefficients ai must be such that the following orthogonality conditions are satisfied (Prob. 7.35):

Solving Eq. (7.25) for ai, we obtain

where

and R−1 is the inverse of R.

SOLVED PROBLEMS

Properties of Point Estimators

7.1. Let (X1, …, Xn) be a random sample of X having unknown mean μ. Show that the estimator of μ defined by

is an unbiased estimator of μ. Note that is known as the sample mean (Prob. 4.64).

By Eq. (4.108),

Thus, M is an unbiased estimator of μ.

7.2. Let (X1, …, Xn) be a random sample of X having unknown mean μ and variance σ2. Show that the estimator of σ2 defined by

where is the sample mean, is a biased estimator of σ2.

By definition, we have

Now

By Eqs. (4.112) and (7.27), we have

Thus,

which shows that S2 is a biased estimator of σ2.

7.3. Let (X1, …, Xn) be a random sample of a Poisson r.v. X with unknown parameter λ. (a) Show that

are both unbiased estimators of λ. (b) Which estimator is more efficient?

(a) By Eqs. (2.50) and (4.132), we have

Thus, both estimators are unbiased estimators of λ. (b) By Eqs. (2.51) and (4.136),

Thus, if n > 2, Λ1 is a more efficient estimator of λ than Λ2, since λ/n < λ/2.

7.4. Let (X1, …, Xn) be a random sample of X with mean μ and variance σ 2. A

linear estimator of μ is defined to be a linear function of X1, …, Xn, l(X1, …, Xn). Show that the linear estimator defined by [Eq. (7.27)],

is the most efficient linear unbiased estimator of μ.

Assume that

is a linear unbiased estimator of μ with lower variance than M. Since M1 is unbiased, we must have

which implies that . By Eq. (4.136),

By assumption,

Consider the sum

which, by assumption (7.30), is less than 0. This is impossible unless ai = 1/n, implying that M is the most efficient linear unbiased estimator of μ.

7.5. Show that if

then the estimator Θn is consistent.

Using Chebyshev’s inequality (2.116), we can write

Thus, if

then

that is, Θn is consistent [see Eq. (7.6)].

7.6. Let (X1, …, Xn) be a random sample of a uniform r.v. X over (0, a), where a is unknown. Show that

is a consistent estimator of the parameter a.

If X is uniformly distributed over (0, a), then from Eqs. (2.56), (2.57), and (4.122) of Prob. 4.36, the pdf of Z = max(X1, …, Xn) is

Thus,

and

Next,

and

Thus, by Eqs. (7.7) and (7.8), A is a consistent estimator of parameter a.

Maximum-Likelihood Estimation

7.7. Let (X1, …, Xn) be a random sample of a binomial r.v. X with parameters (m, p), where m is assumed to be known and p unknown. Determine the maximum-likelihood estimator of p.

The likelihood function is given by [Eq. (2.36)]

Taking the natural logarithm of the above expression, we get

Setting d[ln L(p)]/dp = 0, the maximum-likelihood estimate of p is obtained as

Hence, the maximum-likelihood estimator of p is given by

7.8. Let (X1, …, Xn) be a random sample of a Poisson r.v. with unknown parameter λ. Determine the maximum-likelihood estimator of λ.

The likelihood function is given by [Eq. (2.48)]

Thus,

where

and

Setting d[ln L(λ)]/dλ = 0, the maximum-likelihood estimate of λ is obtained as

Hence, the maximum-likelihood estimator of λ is given by

7.9. Let (X1, …, Xn) be a random sample of an exponential r.v. X with unknown parameter λ. Determine the maximum-likelihood estimator of λ.

The likelihood function is given by [Eq. (2.60)]

Thus,

and

Setting d[ln L(λ)]/dλ = 0, the maximum-likelihood estimate of λ is obtained as

Hence, the maximum-likelihood estimator of λ is given by

7.10. Let (X1, …, Xn) be a random sample of a normal random r.v. X with unknown mean μ and unknown variance σ2. Determine the maximum- likelihood estimators of μ and σ2.

The likelihood function is given by [Eq. (2.71)]

Thus,

In order to find the values of μ and σ maximizing the above, we compute

Equating these equations to zero, we get

and

Solving for and , the maximum-likelihood estimates of μ and σ2 are given, respectively, by

Hence, the maximum-likelihood estimators of μ and σ2 are given, respectively, by

Bayes’ Estimation

7.11. Let (X1, …, Xn) be the random sample of a Bernoulli r.v. X with pmf given by [Eq. (2.32)]

where p, 0 ≤ p ≤ 1, is unknown. Assume that p is a uniform r.v. over (0, 1). Find the Bayes’ estimator of p.

The prior pdf of p is the uniform pdf; that is,

The posterior pdf of p is given by

Then, by Eq. (7.12),

where and by Eq. (7.13),

Now, from calculus, for integers m and k, we have

Thus, by Eq. (7.14), the posterior pdf of p is

and by Eqs. (7.15) and (7.44),

Hence, by Eq. (7.16), the Bayes’ estimator of p is

7.12. Let (X1, …, Xn) be a random sample of an exponential r.v. X with unknown parameter λ. Assume that λ is itself to be an exponential r.v. with parameter

α. Find the Bayes’ estimator of λ.

The assumed prior pdf of λ is [Eq. (2.48)]

where . Then, by Eqs. (7.12) and (7.13),

By Eq. (7.14), the posterior pdf of λ is given by

Thus, by Eq. (7.15), the Bayes’ estimate of λ is

and the Bayes’ estimator of λ is

7.13. Let (X1, …, Xn) be a random sample of a normal r.v. X with unknown mean μ and variance 1. Assume that μ is itself to be a normal r.v. with mean 0 and variance 1. Find the Bayes’ estimator of μ.

The assumed prior pdf of μ is

Then by Eq. (7.12),

Then, by Eq. (7.14), the posterior pdf of μ is given by

where C = C(x1, …, xn) is independent of μ. However, Eq. (7.48) is just the pdf of a normal r.v. with mean

and variance

Hence, the conditional distribution of μ given x1, …, xn is the normal distribution with mean

and variance

Thus, the Bayes’ estimate of μ is given by

and the Bayes’ estimator of μ is

7.14. Let (X1, …, Xn) be a random sample of a r.v. X with pdf ƒ(x; θ), where θ is an unknown parameter. The statistics L and U determine a 100(1 − α) percent confidence interval (L, U) for the parameter θ if

and 1 − α is called the confidence coefficient. Find L and U if X is a normal r.v. with known variance σ2 and mean μ is an unknown parameter.

If X = N(μ; σ2), then

is a standard normal r.v., and hence for a given α we can find a number zα/2 from Table A (Appendix A) such that

For example, if 1 − α = 0.95, then zα/2 = z0.025 = 1.96, and if 1 − α = 0.9, then zα/2 = z0.05 = 1.645. Now, recalling that σ > 0, we have the following equivalent inequality relationships;

Thus, we have

and so

7.15. Consider a normal r.v. with variance 1.66 and unknown mean μ. Find the 95 percent confidence interval for the mean based on a random sample of size 10.

As shown in Prob. 7.14, for 1 − α = 0.95, we have zα/2 = z0.025 = 1.96 and

Thus, by Eq. (7.54), the 95 percent confidence interval for μ is

Mean Square Estimation

7.16. Find the m.s. estimate of a r.v. Y by a constant c.

By Eq. (7.17), the m.s. error is

Clearly the m.s. error e depends on c, and it is minimum if

or

Thus, we conclude that the m.s. estimate c of Y is given by

7.17. Find the m.s. estimator of a r.v. Y by a function g(X) of the r.v. X.

By Eq. (7.17), the m.s. error is

Since ƒ(x, y) = ƒ(y |x) ƒ(x), we can write

Since the integrands above are positive, the m.s. error e is minimum if the inner integrand,

is minimum for every x. Comparing Eq. (7.58) with Eq. (7.55) (Prob. 7.16), we see that they are the same form if c is changed to g(x) and ƒ(y) is changed to ƒ(y |x). Thus, by the result of Prob. 7.16 [Eq. (7.56)], we conclude that the m.s. estimate of Y is given by

Hence, the m.s. estimator of Y is

7.18. Find the m.s. error if g(x) = E(Y|x) is the m.s. estimate of Y.

As we see from Eq. (3.58), the conditional mean E(Y|x) of Y, given that X = x, is a function of x, and by Eq. (4.39),

Similarly, the conditional mean E[g(X, Y)|x] of g(X, Y), given that X = x, is a function of x. It defines, therefore, the function E[g(X, Y)|X] of the r.v. X. Then

Note that Eq. (7.62) is the generalization of Eq. (7.61). Next, we note that

Then by Eqs. (7.62) and (7.63), we have

Now, setting g1(X) = g(X) and g2(Y) = Y in Eq. (7.64), and using Eq. (7.18), we obtain

Thus, the m.s. error is given by

7.19. Let Y = X2 and X be a uniform r.v. over (−1, 1). Find the m.s. estimator of Y in terms of X and its m.s. error.

By Eq. (7.18), the m.s. estimate of Y is given by

Hence, the m.s. estimator of Y is

The m.s. error is

Linear Mean Square Estimation

7.20. Derive the orthogonality principle (7.21) and Eq. (7.22).

By Eq. (7.20), the m.s. error is

Clearly, the m.s. error e is a function of a and b, and it is minimum if ∂e/∂a = 0 and ∂e/∂ b = 0. Now

Setting ∂e/∂a = 0 and ∂e/∂b = 0, we obtain

Note that Eq. (7.68) is the orthogonality principle (7.21).

Rearranging Eqs. (7.68) and (7.69), we get

Solving for a and b, we obtain Eq. (7.22); that is,

where we have used Eqs. (2.31), (3.51), and (3.53).

7.21. Show that m.s. error defined by Eq. (7.20) is minimum when Eqs. (7.68) and (7.69) are satisfied.

Assume that , where c and d are arbitrary constants. Then

The last two terms on the right-hand side are zero when Eqs. (7.68) and (7.69) are satisfied, and the second term on the right-hand side is positive if

a ≠ c and b ≠ d. Thus, e(c, d) ≥ e(a, b) for any c and d. Hence, e(a, b) is minimum.

7.22. Derive Eq. (7.23).

By Eqs. (7.68) and (7.69), we have

Using Eqs. (2.31), (3.51), and (3.53), and substituting the values of a and b [Eq. (7.22)] in the above expression, the minimum m.s. error is

which is Eq. (7.23).

7.23. Let Y = X2, and let X be a uniform r.v. over (−1, 1) (see Prob. 7.19). Find the linear m.s. estimator of Y in terms of X and its m.s. error.

The linear m.s. estimator of Y in terms of X is

where a and b are given by [Eq. (7.22)]

Now, by Eqs. (2.58) and (2.56),

By Eq. (3.51),

Thus, a = 0 and b = E(Y), and the linear m.s. estimator of Y is

and the m.s. error is

7.24. Find the minimum m.s. error estimator of Y in terms of X when X and Y are jointly normal r.v.’s.

By Eq. (7.18), the minimum m.s. error estimator of Y in terms of X is

Now, when X and Y are jointly normal, by Eq. (3.108) (Prob. 3.51), we have

Hence, the minimum m.s. error estimator of Y is

Comparing Eq. (7.72) with Eqs. (7.19) and (7.22), we see that for jointly normal r.v.’s the linear m.s. estimator is the minimum m.s. error estimator.

SUPPLEMENTARY PROBLEMS

7.25. Let (X1, …, Xn) be a random sample of X having unknown mean μ and variance σ2. Show that the estimator of σ2 defined by

where is the sample mean, is an unbiased estimator of σ2. Note that is often called the sample variance.

7.26. Let (X1, …, Xn) be a random sample of X having known mean μ and unknown variance σ2. Show that the estimator of σ2 defined by

is an unbiased estimator of σ2.

7.27. Let (X1, …, Xn) be a random sample of a binomial r.v. X with parameter (m, p), where p is unknown. Show that the maximum-likelihood estimator of p given by Eq. (7.34) is unbiased.

7.28. Let (X1, …, Xn) be a random sample of a Bernoulli r.v. X with pmf ƒ(x; p) = px(1 − p)1−x, x = 0, 1, where p, 0 ≤ p ≤ 1, is unknown. Find the maximum- likelihood estimator of p.

7.29. The values of a random sample, 2.9, 0.5, 1.7, 4.3, and 3.2, are obtained from a r.v. X that is uniformly distributed over the unknown interval (a, b). Find the maximum-likelihood estimates of a and b.

7.30. In analyzing the flow of traffic through a drive-in bank, the times (in minutes) between arrivals of 10 customers are recorded as 3.2, 2.1, 5.3, 4.2, 1.2, 2.8, 6.4, 1.5, 1.9, and 3.0. Assuming that the interarrival time is an exponential r.v. with parameter λ, find the maximum likelihood estimate of λ.

7.31. Let (X1, …, Xn) be a random sample of a normal r.v. X with known mean μ and unknown variance σ2. Find the maximum likelihood estimator of σ2.

7.32. Let (X1, …, Xn) be the random sample of a normal r.v. X with mean μ and variance σ2, where μ is unknown. Assume that μ is itself to be a normal r.v.

with mean μ1 and variance . Find the Bayes’ estimate of μ.

7.33. Let (X1, …, Xn) be the random sample of a normal r.v. X with variance 100 and unknown μ. What sample size n is required such that the width of 95 percent confidence interval is 5?

7.34. Find a constant a such that if Y is estimated by aX, the m.s. error is minimum, and also find the minimum m.s. error em.

7.35. Derive Eqs. (7.25) and (7.26).

ANSWERS TO SUPPLEMENTARY PROBLEMS

7.25. Hint: Show that , and use Eq.(7.29)

7.26. Hint: Proceed as in Prob. 7.2.

7.27. Hint: Use Eq. (2.38).

7.28.

7.29.

7.30.

7.31.

7.32.

7.33. n = 62

7.34.

7.35. Hint: Proceed as in Prob. 7.20.

CHAPTER 8

Decision Theory

8.1 Introduction

There are many situations in which we have to make decisions based on observations or data that are random variables. The theory behind the solutions for these situations is known as decision theory or hypothesis testing. In communication or radar technology, decision theory or hypothesis testing is known as (signal) detection theory. In this chapter we present a brief review of the binary decision theory and various decision tests.

8.2 Hypothesis Testing

A. Definitions: A statistical hypothesis is an assumption about the probability law of r.v.’s. Suppose we observe a random sample (X1, …, Xn) of a r.v. X whose pdf ƒ(x; θ) = ƒ(x1, …, xn; θ) depends on a parameter θ. We wish to test the assumption θ = θ0 against the assumption θ = θ1. The assumption θ = θ0 is denoted by H0 and is called the null hypothesis. The assumption θ = θ1 is denoted by H1 and is called the alternative hypothesis.

A hypothesis is called simple if all parameters are specified exactly. Otherwise it is called composite. Thus, suppose H0: θ = θ0 and H1: θ ≠ θ0; then H0 is simple and H1 is composite.

B. Hypothesis Testing and Types of Errors: Hypothesis testing is a decision process establishing the validity of a hypothesis. We can think of the decision process as dividing the observation space Rn (Euclidean n-space) into two regions R0 and R1. Let x = (x1, …, xn) be the observed vector. Then if x ∈ R0, we will decide on H0; if x ∈ R1, we decide on H1. The region R0 is known as the acceptance region and the region R1 as the rejection (or critical) region (since the null hypothesis is rejected). Thus, with the observation vector (or data), one of the following four actions can happen:

1. H0 true; accept H0 2. H0 true; reject H0 (or accept H1) 3. H1 true; accept H1 4. H1 true; reject H1 (or accept H0)

The first and third actions correspond to correct decisions, and the second and fourth actions correspond to errors. The errors are classified as

1. Type I error: Reject H0 (or accept H1) when H0 is true. 2. Type II error: Reject H1 (or accept H0) when H1 is true.

Let PI and PII denote, respectively, the probabilities of Type I and Type II errors:

where Di (i = 0, 1) denotes the event that the decision is made to accept Hi. P1 is often denoted by α and is known as the level of significance, and PII is denoted by β and (1 − β) is known as the power of the test. Note that since α

and β represent probabilities of events from the same decision problem, they are not independent of each other or of the sample size n. It would be desirable to have a decision process such that both α and β will be small. However, in general, a decrease in one type of error leads to an increase in the other type for a fixed sample size (Prob. 8.4). The only way to simultaneously reduce both types of errors is to increase the sample size (Prob. 8.5). One might also attach some relative importance (or cost) to the four possible courses of action and minimize the total cost of the decision (see Sec. 8.3D).

The probabilities of correct decisions (actions 1 and 3) may be expressed as

In radar signal detection, the two hypotheses are

In this case, the probability of a Type I error PI = P(D1 |H0) is often referred to as the false-alarm probability (denoted by PF), the probability of a Type II error PII = P(D0 |H1) as the miss probability (denoted by PM), and P(D1 |H1) as the detection probability (denoted by PD). The cost of failing to detect a target cannot be easily determined. In general we set a value of PF which is acceptable and seek a decision test that constrains PF to this value while maximizing PD (or equivalently minimizing PM). This test is known as the Neyman-Pearson test (see Sec. 8.3C).

8.3 Decision Tests

A. Maximum-Likelihood Test:

Let x be the observation vector and P(x |Hi), i = 0.1, denote the probability of observing x given that Hi was true. In the maximum-likelihood test, the decision regions R0 and R1 are selected as

Thus, the maximum-likelihood test can be expressed as

The above decision test can be rewritten as

If we define the likelihood ratio Λ(x) as

then the maximum-likelihood test (8.7) can be expressed as

which is called the likelihood ratio test, and 1 is called the threshold value of the test.

Note that the likelihood ratio Λ(x) is also often expressed as

B. MAP Test: Let P(Hi |x), i = 0, 1, denote the probability that Hi was true given a particular value of x. The conditional probability P(Hi |x) is called a posteriori (or posterior) probability, that is, a probability that is computed after an observation has been made. The probability P(Hi), i = 0, 1, is called a priori (or prior) probability. In the maximum a posteriori (MAP) test, the decision regions R0 and R1 are selected as

Thus, the MAP test is given by

which can be rewritten as

Using Bayes’ rule [Eq. (1.58)], Eq. (8.13) reduces to

Using the likelihood ratio Λ(x) defined in Eq. (8.8), the MAP test can be expressed in the following likelihood ratio test as

where η = P(H0)/P(H1) is the threshold value for the MAP test. Note that when P(H0) = P(H1), the maximum-likelihood test is also the MAP test.

C. Neyman-Pearson Test: As mentioned, it is not possible to simultaneously minimize both α (= PI) and β(= PII). The Neyman-Pearson test provides a workable solution to this problem in that the test minimizes β for a given level of α. Hence, the Neyman-Pearson test is the test which maximizes the power of the test 1 − β for a given level of significance α. In the Neyman-Pearson test, the critical (or rejection) region R1 is selected such that 1 − β = 1 − P(D0 |H1) = P(D1 |H1) is maximum subject to the constraint α = P(D1 |H0) = α0. This is a classical problem in optimization: maximizing a function subject to a constraint, which can be solved by the use of Lagrange multiplier method. We thus construct the objective function

where λ ≥ 0 is a Lagrange multiplier. Then the critical region R1 is chosen to maximize J. It can be shown that the Neyman-Pearson test can be expressed in terms of the likelihood ratio test as (Prob. 8.8)

where the threshold value η of the test is equal to the Lagrange multiplier λ, which is chosen to satisfy the constraint α = α0.

D. Bayes’ Test: Let Cij be the cost associated with (Di, Hj), which denotes the event that we accept Hi when Hj is true. Then the average cost, which is known as the Bayes’ risk, can be written as

where P(Di, Hj) denotes the probability that we accept Hi when Hj is true. By Bayes’ rule (1.42), we have

In general, we assume that

since it is reasonable to assume that the cost of making an incorrect decision is higher than the cost of making a correct decision. The test that minimizes the average cost is called the Bayes’ test, and it can be expressed in terms of the likelihood ratio test as (Prob. 8.10)

Note that when C10 − C00 = C01 − C11, the Bayes’ test (8.21) and the MAP test (8.15) are identical.

E. Minimum Probability of Error Test: If we set C00 = C11 = 0 and C01 = C10 = 1 in Eq. (8.18), we have

which is just the probability of making an incorrect decision. Thus, in this case, the Bayes’ test yields the minimum probability of error, and Eq. (8.21) becomes

We see that the minimum probability of error test is the same as the MAP test.

F. Minimax Test:

We have seen that the Bayes’ test requires the a priori probabilities P(H0) and P(H1). Frequently, these probabilities are not known. In such a case, the Bayes’ test cannot be applied, and the following minimax (min-max) test may be used. In the minimax test, we use the Bayes’ test which corresponds to the least favorable P(H0) (Prob. 8.12). In the minimax test, the critical region is defined by

for all . In other words, is the critical region which yields the minimum Bayes’ risk for the least favorable P(H0). Assuming that the minimization and maximization operations are interchangeable, then we have

The minimization of [P(H0), R1] with respect to R1 is simply the Bayes’ test, so that

where C*[P(H0)] is the minimum Bayes’ risk associated with the a priori probability P(H0). Thus, Eq. (8.25) states that we may find the minimax test by finding the Bayes’ test for the least favorable P(H0), that is, the P(H0) which maximizes [P(H0)].

SOLVED PROBLEMS

Hypothesis Testing

8.1. Suppose a manufacturer of memory chips observes that the probability of chip failure is p = 0.05. A new procedure is introduced to improve

the design of chips. To test this new procedure, 200 chips could be produced using this new procedure and tested. Let r.v. X denote the number of these 200 chips that fail. We set the test rule that we would accept the new procedure if X ≤ 5. Let

Find the probability of a Type I error.

If we assume that these tests using the new procedure are independent and have the same probability of failure on each test, then X is a binomial r.v. with parameters (n, p) = (200, p). We make a Type I error if X ≤ 5 when in fact p = 0.05. Thus, using Eq. (2.37), we have

Since n is rather large and p is small, these binomial probabilities can be approximated by Poisson probabilities with λ = np = 200(0.05) = 10 (see Prob. 2.43). Thus, using Eq. (2.119), we obtain

Note that H0 is a simple hypothesis but H1 is a composite hypothesis.

8.2. Consider again the memory chip manufacturing problem of Prob. 8.1. Now let

Again our rule is, we would reject the new procedure if X > 5. Find the probability of a Type II error.

Now both hypotheses are simple. We make a Type II error if X > 5 when in fact p = 0.02. Hence, by Eq. (2.37),

Again using the Poisson approximation with λ = np = 200(0.02) = 4, we obtain

8.3. Let (X1, …, Xn) be a random sample of a normal r.v. X with mean μ and variance 100. Let

and sample size n = 25. As a decision procedure, we use the rule to reject H0 if ≥ 52, where is the value of the sample mean defined by Eq. (7.27). (a) Find the probability of rejecting H0: μ = 50 as a function of μ (>

50). (b) Find the probability of a Type I error α. (c) Find the probability of a Type II error β (i) when μ1 = 53 and (ii)

when μ1 = 55.

(a) Since the test calls for the rejection of H0: μ = 50 when ≥ 52, the probability of rejecting H0 is given by

Now, by Eqs. (4.136) and (7.27), we have

Thus, is N(μ; 4), and using Eq. (2.74), we obtain

The function g(μ) is known as the power function of the test, and the value of g(μ) at μ = μ1, g(μ1), is called the power at μ1.

(b) Note that the power at μ = 50, g(50), is the probability of rejecting H0: μ = 50 when H0 is true—that is, a Type I error. Thus, using Table A (Appendix A), we obtain

(c) Note that the power at μ = μ1, g(μ1), is the probability of rejecting H0: μ = 50 when μ = μ1. Thus, 1 − g(μ1) is the probability of accepting H0 when μ = μ1—that is, the probability of a Type II error β.

(i) Setting μ = μ1 = 53 in Eq. (8.28) and using Table A (Appendix A), we obtain

(ii) Similarly, for μ = μ1 = 55 we obtain

Notice that clearly, the probability of a Type II error depends on the value of μ1.

8.4. Consider the binary decision problem of Prob. 8.3. We modify the decision rule such that we reject H0 if ≥ c. (a) Find the value of c such that the probability of a Type I error α =

0.05. (b) Find the probability of a Type II error β when μ1 = 55 with the

modified decision rule.

(a) Using the result of part (b) in Prob. 8.3, c is selected such that [see Eq. (8.27)]

However, when μ = 50, = N(50; 4), and [see Eq. (8.28)]

From Table A (Appendix A), we have Φ (1.645) = 0.95. Thus,

(b) The power function g(μ) with the modified decision rule is

Setting μ = μ1 = 55 and using Table A (Appendix A), we obtain

Comparing with the results of Prob. 8.3, we notice that with the change of the decision rule, α is reduced from 0.1587 to 0.05, but β is increased from 0.0668 to 0.1963.

8.5. Redo Prob. 8.4 for the case where the sample size n = 100. (a) With n = 100, we have

As in part (a) of Prob. 8.4, c is selected so that

Since = N(50; 1), we have

Thus, (b) The power function is

Setting μ = μ1 = 55 and using Table A (Appendix A), we obtain

Notice that with sample size n = 100, both α and β have decreased from their respective original values of 0.1587 and 0.0668 when n = 25.

Decision Tests

8.6. In a simple binary communication system, during every T seconds, one of two possible signals s0(t) and s1(t) is transmitted. Our two hypotheses are

We assume that

The communication channel adds noise n(t), which is a zero-mean normal random process with variance 1. Let x(t) represent the received signal:

We observe the received signal x(t) at some instant during each signaling interval. Suppose that we received an observation x = 0.6. (a) Using the maximum likelihood test, determine which signal is

transmitted. (b) Find PI and PII.

(a) The received signal under each hypothesis can be written as

Then the pdf of x under each hypothesis is given by

The likelihood ratio is then given by

By Eq. (8.9), the maximum likelihood test is

Taking the natural logarithm of the above expression, we get

Since , we determine that signal s1(t) was transmitted. (b) The decision regions are given by

Then by Eqs. (8.1) and (8.2) and using Table A (Appendix A), we obtain

8.7. In the binary communication system of Prob. 8.6, suppose that and .

(a) Using the MAP test, determine which signal is transmitted when x = 0.6.

(b) Find PI and PII.

(a) Using the result of Prob. 8.6 and Eq. (8.15), the MAP test is given by

Taking the natural logarithm of the above expression, we get

Since x = 0.6 < 1.193, we determine that signal s0(t) was transmitted.

(b) The decision regions are given by

Thus, by Eqs. (8.1) and (8.2) and using Table A (Appendix A), we obtain

8.8. Derive the Neyman-Pearson test, Eq. (8.17).

From Eq. (8.16), the objective function is

where λ is an undetermined Lagrange multiplier which is chosen to satisfy the constraint α = α0. Now, we wish to choose the critical region R1 to maximize J. Using Eqs. (8.1) and (8.2), we have

To maximize J by selecting the critical region R1, we select x ∈ R1 such that the integrand in Eq. (8.30) is positive. Thus, R1 is given by

and the Neyman-Pearson test is given by

and λ is determined such that the constraint

is satisfied.

8.9. Consider the binary communication system of Prob. 8.6 and suppose that we require that α = PI = 0.25. (a) Using the Neyman-Pearson test, determine which signal is

transmitted when x = 0.6. (b) Find PII.

(a) Using the result of Prob. 8.6 and Eq. (8.17), the Neyman-Pearson test is given by

Taking the natural logarithm of the above expression, we get

The critical region R1 is thus

Now we must determine λ such that α = PI = P(D1 |H0) = 0.25. By Eq. (8.1), we have

From Table A (Appendix A), we find that Φ (0.674) = 0.75. Thus,

Then the Neyman-Pearson test is

Since x = 0.6 < 0.674, we determine that signal s0(t) was transmitted.

(b) By Eq. (8.2), we have

8.10. Derive Eq. (8.21).

By Eq. (8.19), the Bayes’ risk is

Now we can express

Then can be expressed as

Since R0 ∪ R1 = S and R0 ∩ R1 = ϕ, we can write

Then Eq. (8.32) becomes

The only variable in the above expression is the critical region R1. By the assumptions [Eq. (8.20)] C10 > C00 and C01 > C11, the two terms inside the brackets in the integral are both positive. Thus, is minimized if R1 is chosen such that

for all x ∈ R1. That is, we decide to accept H1 if

In terms of the likelihood ratio, we obtain

which is Eq. (8.21).

8.11. Consider a binary decision problem with the following conditional pdf’s:

The Bayes’ costs are given by

(a) Determine the Bayes’ test if and the associated Bayes’ risk.

(b) Repeat (a) with .

(a) The likelihood ratio is

By Eq. (8.21), the Bayes’ test is given by

Taking the natural logarithm of both sides of the last expression yields

Thus, the decision regions are given by

and by Eq. (8.19), the Bayes’ risk is

(b) The Bayes’ test is

Again, taking the natural logarithm of both sides of the last expression yields

Thus, the decision regions are given by

Then

and by Eq. (8.19), the Bayes’ risk is

8.12. Consider the binary decision problem of Prob. 8.11 with the same Bayes’ costs. Determine the minimax test.

From Eq. (8.33), the likelihood ratio is

In terms of P(H0), the Bayes’ test [Eq. (8.21)] becomes

Taking the natural logarithm of both sides of the last expression yields

For P(H0) > 0.8, δ becomes negative, and we always decide H0. For P(H0) ≤ 0.8, the decision regions are

Then, by setting C00 = C11 = 0, C01 = 2, and C10 = 1 in Eq. (8.19), the minimum Bayes’ risk can be expressed as a function of P(H0) as

From the definition of δ [Eq. (8.34)], we have

Thus,

Substituting these values into Eq. (8.35), we obtain

Now the value of P(H0) which maximizes can be obtained by setting d [P(H0)]/dP(H0) equal to zero and solving for P(H0). The result yields . Substituting this value into Eq. (8.34), we obtain the following minimax test:

8.13. Suppose that we have n observations Xi, i = 1, …, n, of radar signals, and Xi are normal iid r.v.’s under each hypothesis. Under H0, Xi have mean μ0 and variance σ

2, while under H1, Xi have mean μ1 and variance σ2, and μ1 > μ0. Determine the maximum likelihood test.

By Eq. (2.71) for each Xi, we have

Since the Xi are independent, we have

With μ1 − μ0 > 0, the likelihood ratio is then given by

Hence, the maximum likelihood test is given by

Taking the natural logarithm of both sides of the above expression yields

or

Equation (8.36) indicates that the statistic

provides enough information about the observations to enable us to make a decision. Thus, it is called the sufficient statistic for the

maximum likelihood test.

8.14. Consider the same observations Xi, i = 1, …, n, of radar signals as in Prob. 8.13, but now, under H0, Xi have zero mean and variance , while under H1, Xi have zero mean and variance , and . Determine the maximum likelihood test.

In a similar manner as in Prob. 8.13, we obtain

With , the likelihood ratio is

and the maximum likelihood test is

Taking the natural logarithm of both sides of the above expression yields

Note that in this case,

is the sufficient statistic for the maximum likelihood test.

8.15. In the binary communication system of Prob. 8.6, suppose that we have n independent observations Xi = X(ti), i = 1, …, n, where 0 < t1 < … < tn ≤ T. (a) Determine the maximum likelihood test. (b) Find PI and PII for n = 5 and n = 10.

(a) Setting μ0 = 0 and μ1 = 1 in Eq. (8.36), the maximum likelihood test is

(b) Let

Then by Eqs. (4.132) and (4.136), and the result of Prob. 5.60, we see that Y is a normal r.v. with zero mean and variance 1/n under H0, and is a normal r.v. with mean 1 and variance 1/n under H1. Thus,

Note that PI = PII. Using Table A (Appendix A), we have

8.16. In the binary communication system of Prob. 8.6, suppose that s0(t) and s1(t) are arbitrary signals and that n observations of the received signal x(t) are made. Let n samples of s0(t) and s1(t) be represented, respectively, by

where T denotes "transpose of." Determine the MAP test.

For each Xi, we can write

Since the noise components are independent, we have

and the likelihood ratio is given by

Thus, by Eq. (8.15), the MAP test is given by

Taking the natural logarithm of both sides of the above expression yields

SUPPLEMENTARY PROBLEMS

8.17. Let (X1, …, Xn) be a random sample of a Bernoulli r.v. X with pmf

where it is known that . Let

and n = 20. As a decision test, we use the rule to reject H0 if .

(a) Find the power function g(p) of the test. (b) Find the probability of a Type I error α. (c) Find the probability of a Type II error β (i) when and (ii)

when .

8.18. Let (X1, …, Xn) be a random sample of a normal r.v. X with mean μ and variance 36. Let

As a decision test, we use the rule to accept H0 if < 53, where is the value of the sample mean. (a) Find the expression for the critical region R1. (b) Find α and β for n = 16.

8.19. Let (X1, …, Xn) be a random sample of a normal r.v. X with mean μ and variance 100. Let

As a decision test, we use the rule that we reject H0 if ≥ c. Find the value of c and sample size n such that α = 0.025 and β = 0.05.

8.20. Let X be a normal r.v. with zero mean and variance σ2. Let

Determine the maximum likelihood test.

8.21. Consider the binary decision problem of Prob. 8.20. Let and . Determine the MAP test.

8.22. Consider the binary communication system of Prob. 8.6. (a) Construct a Neyman-Pearson test for the case where α = 0.1. (b) Find β.

8.23. Consider the binary decision problem of Prob. 8.11. Determine the Bayes’ test if P(H0) = 0.25 and the Bayes’ costs are

ANSWERS TO SUPPLEMENTARY PROBLEMS

8.17.

8.18.

8.19. c = 52.718, n = 52

8.20.

8.21.

8.22.

8.23.

CHAPTER 9

Queueing Theory

9.1 Introduction

Queueing theory deals with the study of queues (waiting lines). Queues abound in practical situations. The earliest use of queueing theory was in the design of a telephone system. Applications of queueing theory are found in fields as seemingly diverse as traffic control, hospital management, and time-shared computer system design. In this chapter, we present an elementary queueing theory.

9.2 Queueing Systems

A. Description:

A simple queueing system is shown in Fig. 9-1. Customers arrive randomly at an average rate of λa. Upon arrival, they are served without delay if there are available servers; otherwise, they are made to wait in the queue until it is their turn to be served. Once served, they are assumed to leave the system. We will be interested in determining such quantities as the average number of customers in the system, the average time a customer spends in the system, the average time spent waiting in the queue, and so on.

Fig. 9-1 A simple queueing system.

The description of any queueing system requires the specification of three parts:

1. The arrival process 2. The service mechanism, such as the number of servers and service-time

distribution 3. The queue discipline (for example, first-come, first-served)

B. Classification: The notation A /B /s /K is used to classify a queueing system, where A specifies the type of arrival process, B denotes the service-time distribution, s specifies the number of servers, and K denotes the capacity of the system, that is, the maximum number of customers that can be accommodated. If K is not specified, it is assumed that the capacity of the system is unlimited. For example, an M/M /2 queueing system (M stands for Markov) is one with Poisson arrivals (or exponential interarrival time distribution), exponential service-time distribution, and 2 servers. An M /G/1 queueing system has Poisson arrivals, general service-time distribution, and a single server. A special case is the M/D/1 queueing system, where D stands for constant (deterministic) service time. Examples of queueing systems with limited capacity are telephone systems with limited trunks, hospital emergency rooms with limited beds, and airline terminals with limited space in which to park aircraft for loading and unloading. In each case, customers who arrive when the system is saturated are denied entrance and are lost.

C. Useful Formulas: Some basic quantities of queueing systems are

L: the average number of customers in the system Lq: the average number of customers waiting in the queue Ls: the average number of customers in service W: the average amount of time that a customer spends in the system Wq: the average amount of time that a customer spends waiting in the

queue

Ws: the average amount of time that a customer spends in service

Many useful relationships between the above and other quantities of interest can be obtained by using the following basic cost identity:

Assume that entering customers are required to pay an entrance fee (according to some rule) to the system. Then we have

where λa is the average arrival rate of entering customers

and X(t) denotes the number of customer arrivals by time t. If we assume that each customer pays $1 per unit time while in the

system, Eq. (9.1) yields

Equation (9.2) is sometimes known as Little’s formula. Similarly, if we assume that each customer pays $1 per unit time while in

the queue, Eq. (9.1) yields

If we assume that each customer pays $1 per unit time while in service, Eq. (9.1) yields

Note that Eqs. (9.2) to (9.4) are valid for almost all queueing systems, regardless of the arrival process, the number of servers, or queueing discipline.

9.3 Birth-Death Process

We say that the queueing system is in state Sn if there are n customers in the system, including those being served. Let N(t) be the Markov process that takes on the value n when the queueing system is in state Sn with the following assumptions:

1. If the system is in state Sn, it can make transitions only to Sn − 1 or Sn − 1, n ≥ 1; that is, either a customer completes service and leaves the system or, while the present customer is still being serviced, another customer arrives at the system; from S0, the next state can only be S1.

2. If the system is in state Sn at time t, the probability of a transition to Sn + 1 in the time interval (t, t + Δt) is an Δt. We refer to an as the arrival parameter (or the birth parameter).

3. If the system is in state Sn at time t, the probability of a transition to Sn − 1 in the time interval (t, t + Δt) is dn Δt. We refer to dn as the departure parameter (or the death parameter).

The process N(t) is sometimes referred to as the birth-death process. Let pn(t) be the probability that the queueing system is in state Sn at time

t; that is,

Then we have the following fundamental recursive equations for N(t) (Prob. 9.2):

Assume that in the steady state we have

and setting and in Eqs. (9.6), we obtain the following steady- state recursive equation:

and for the special case with d0 = 0,

Equations (9.8) and (9.9) are also known as the steady-state equilibrium equations. The state transition diagram for the birth-death process is shown in Fig. 9-2.

Fig. 9-2 State transition diagram for the birth-death process.

Solving Eqs. (9.8) and (9.9) in terms of p0, we obtain

where p0 can be determined from the fact that

provided that the summation in parentheses converges to a finite value.

9.4 The M/M/1 Queueing System

In the M/M/1 queueing system, the arrival process is the Poisson process with rate λ (the mean arrival rate) and the service time is exponentially distributed with parameter μ (the mean service rate). Then the process N(t) describing the state of the M/M/1 queueing system at time t is a birth-death process with the following state independent parameters:

Then from Eqs. (9.10) and (9.11), we obtain (Prob. 9.3)

where ρ = λ/μ < 1, which implies that the server, on the average, must process the customers faster than their average arrival rate; otherwise, the queue length (the number of customers waiting in the queue) tends to infinity. The ratio ρ = λ/μ is sometimes referred to as the traffic intensity of the system. The traffic intensity of the system is defined as

The average number of customers in the system is given by (Prob. 9.4)

Then, setting λa = λ in Eqs. (9.2) to (9.4), we obtain (Prob. 9.5)

9.5 The M/M/s Queueing System

In the M/M/s queueing system, the arrival process is the Poisson process with rate λ and each of the s servers has an exponential service time with parameter μ. In this case, the process N(t) describing the state of the M/M/s queueing system at time t is a birth-death process with the following parameters:

Note that the departure parameter dn is state dependent. Then, from Eqs. (9.10) and (9.11) we obtain (Prob. 9.10)

where ρ = λ /(sμ) < 1. Note that the ratio ρ = λ /(sμ) is the traffic intensity of the M/M/s queueing system. The average number of customers in the system and the average number of customers in the queue are given, respectively, by (Prob. 9.12)

By Eqs. (9.2) and (9.3), the quantities W and Wq are given by

9.6 The M/M/1/K Queueing System

In the M/M/1/K queueing system, the capacity of the system is limited to K customers. When the system reaches its capacity, the effect is to reduce the arrival rate to zero until such time as a customer is served to again make queue space available. Thus, the M/M/1/K queueing system can be modeled as a birth-death process with the following parameters:

Then, from Eqs. (9.10) and (9.11) we obtain (Prob. 9.14)

where ρ = λ/μ. It is important to note that it is no longer necessary that the traffic intensity ρ = λ/μ be less than 1. Customers are denied service when

the system is in state K. Since the fraction of arrivals that actually enter the system is 1 − pK, the effective arrival rate is given by

The average number of customers in the system is given by (Prob. 9.15)

Then, setting λa = λe in Eqs. (9.2) to (9.4), we obtain

9.7 The M/M/s/K Queueing System

Similarly, the M/M/s/K queueing system can be modeled as a birth-death process with the following parameters:

Then, from Eqs. (9.10) and (9.11), we obtain (Prob. 9.17)

where ρ = λ /(sμ). Note that the expression for pn is identical in form to that for the M/M/s system, Eq. (9.21). They differ only in the p0 term. Again, it is not necessary that ρ = λ /(sμ) be less than 1. The average number of customers in the queue is given by (Prob. 9.18)

The average number of customers in the system is

The quantities W and Wq are given by

SOLVED PROBLEMS

9.1. Deduce the basic cost identity (9.1).

Let T be a fixed large number. The amount of money earned by the system by time T can be computed by multiplying the average rate at which the system earns by the length of time T. On the other hand, it can also be computed by multiplying the average amount paid by an

entering customer by the average number of customers entering by time T, which is equal to λaT, where λa is the average arrival rate of entering customers. Thus, we have

Average rate at which the system earns × T = average amount paid by an entering customer × (λaT) Dividing both sides by T (and letting T → ∞), we obtain Eq. (9.1).

9.2. Derive Eq. (9.6).

From properties 1 to 3 of the birth-death process N(t), we see that at time t + Δt the system can be in state Sn in three ways: 1. By being in state Sn at time t and no transition occurring in the time

interval (t, t + Δ t). This happens with probability (1 − an Δ t)(1 − dn Δ t) ≈ 1 − (an + dn) Δ t [by neglecting the second-order effect andn(Δ t)

2]. 2. By being in state Sn − 1 at time t and a transition to Sn occurring in

the time interval (t, t + Δ t). This happens with probability an − 1 Δ t. 3. By being in state Sn − 1 at time t and a transition to Sn occurring in

the time interval (t, t + Δ t). This happens with probability dn − 1 Δ t. Let

Then, using the Markov property of N(t), we obtain

Rearranging the above relations

Letting Δt → 0, we obtain

9.3. Derive Eqs. (9.13) and (9.14).

Setting an = λ, d0 = 0, and dn = μ in Eq. (9.10), we get

where p0 is determined by equating

from which we obtain

9.4. Derive Eq. (9.15).

Since pn is the steady-state probability that the system contains exactly n customers, using Eq. (9.14), the average number of customers in the M/M/1 queueing system is given by

where ρ = λ/μ < 1. Using the algebraic identity

we obtain

9.5. Derive Eqs. (9.16) to (9.18).

Since λa = λ, by Eqs. (9.2) and (9.15), we get

which is Eq. (9.16). Next, by definition,

where Ws = 1/μ, that is, the average service time. Thus,

which is Eq. (9.17). Finally, by Eq. (9.3),

which is Eq. (9.18).

9.6. Let Wa denote the amount of time an arbitrary customer spends in the M/M/1 queueing system. Find the distribution of Wa.

We have

where n is the number of customers in the system. Now consider the amount of time Wa that this customer will spend in the system when there are already n customers when he or she arrives. When n = 0, then Wa = Ws(a), that is, the service time. When n ≥ 1, there will be one customer in service and n − 1 customers waiting in line ahead of this customer’s arrival. The customer in service might have been in service for some time, but because of the memoryless property of the exponential distribution of the service time, it follows that (see Prob. 2.58) the arriving customer would have to wait an exponential amount of time with parameter μ for this customer to complete service. In addition, the customer also would have to wait an exponential amount of time for each of the other n − 1 customers in line. Thus, adding his or her own service time, the amount of time Wa that the customer will spend in the system when there are already n customers when he or

she arrives is the sum of n + 1 iid exponential r. v.’s with parameter μ. Then by Prob. 4.39, we see that this r. v. is a gamma r. v. with parameters (n + 1, μ). Thus, by Eq. (2.65),

From Eq. (9.14),

Hence, by Eq. (9.44),

Thus, by Eq. (2.61), Wa is an exponential r. v. with parameter μ = λ. Note that from Eq. (2.62), E(Wa) = 1/(μ − λ), which agrees with Eq. (9.16), since W = E(Wa).

9.7. Customers arrive at a watch repair shop according to a Poisson process at a rate of one per every 10 minutes, and the service time is an exponential r.v. with mean 8 minutes. (a) Find the average number of customers L, the average time a

customer spends in the shop W, and the average time a customer spends in waiting for service Wq.

(b) Suppose that the arrival rate of the customers increases 10 percent. Find the corresponding changes in L, W, and Wq.

(a) The watch repair shop service can be modeled as an M/M/1 queueing system with . Thus, from Eqs. (9.15), (9.16), and (9.43), we have

(b) Now . Then

It can be seen that an increase of 10 percent in the customer arrival rate doubles the average number of customers in the system. The average time a customer spends in queue is also doubled.

9.8. A drive-in banking service is modeled as an M/M/1 queueing system with customer arrival rate of 2 per minute. It is desired to have fewer than 5 customers line up 99 percent of the time. How fast should the service rate be?

From Eq. (9.14),

In order to have fewer than 5 customers line up 99 percent of the time, we require that this probability be less than 0.01. Thus,

from which we obtain

Thus, to meet the requirements, the average service rate must be at least 5.024 customers per minute.

9.9. People arrive at a telephone booth according to a Poisson process at an average rate of 12 per hour, and the average time for each call is an exponential r.v. with mean 2 minutes. (a) What is the probability that an arriving customer will find the

telephone booth occupied? (b) It is the policy of the telephone company to install additional

booths if customers wait an average of 3 or more minutes for the phone. Find the average arrival rate needed to justify a second booth.

(a) The telephone service can be modeled as an M/ M/ 1 queueing system with , and . The probability that an arriving customer will find the telephone occupied is P(L > 0), where L is the average number of customers in the system. Thus, from Eq. (9.13),

(b) From Eq. (9.17),

from which we obtain λ ≥ 0.3 per minute. Thus, the required average arrival rate to justify the second booth is 18 per hour.

9.10. Derive Eqs. (9.20) and (9.21).

From Eqs. (9.19) and (9.10), we have

Let ρ = λ/(sμ). Then Eqs. (9.46) and (9.47) can be rewritten as

which is Eq. (9.21). From Eq. (9.11), p0 is obtained by equating

Using the summation formula

we obtain Eq. (9.20); that is,

provided ρ = λ/(sμ) < 1.

9.11. Consider an M / M /s queueing system. Find the probability that an arriving customer is forced to join the queue.

An arriving customer is forced to join the queue when all servers are busy—that is, when the number of customers in the system is equal to or greater than s. Thus, using Eqs. (9.20) and (9.21), we get

Equation (9.49) is sometimes referred to as Erlang’s delay (or C) formula and denoted by C(s, λ /μ). Equation (9.49) is widely used in telephone systems and gives the probability that no trunk (server) is available for an incoming call (arriving customer) in a system of s trunks.

9.12. Derive Eqs. (9.22) and (9.23).

Equation (9.21) can be rewritten as

Then the average number of customers in the system is

Using the summation formulas,

and Eq. (9.20), we obtain

Next, using Eqs. (9.21) and (9.50), the average number of customers in the queue is

9.13. A corporate computing center has two computers of the same capacity. The jobs arriving at the center are of two types, internal jobs and external jobs. These jobs have Poisson arrival times with rates 18 and 15 per hour, respectively. The service time for a job is an exponential r.v. with mean 3 minutes. (a) Find the average waiting time per job when one computer is used

exclusively for internal jobs and the other for external jobs. (b) Find the average waiting time per job when two computers handle

both types of jobs.

(a) When the computers are used separately, we treat them as two M/M/1 queueing systems. Let Wq1 and Wq2 be the average waiting

time per internal job and per external job, respectively. For internal jobs, and . Then, from Eq. (9.16),

(b) When two computers handle both types of jobs, we model the computing service as an M/M/2 queueing system with

Now, substituting s = 2 in Eqs. (9.20), (9.22), (9.24), and (9.25), we get

Thus, from Eq. (9.54), the average waiting time per job when both computers handle both types of jobs is given by

From these results, we see that it is more efficient for both computers to handle both types of jobs.

9.14. Derive Eqs. (9.27) and (9.28).

From Eqs. (9.26) and (9.10), we have

From Eq. (9.11), p0 is obtained by equating

Using the summation formula

we obtain

Note that in this case, there is no need to impose the condition that ρ = λ /μ < 1.

9.15. Derive Eq. (9.30).

Using Eqs. (9.28) and (9.51), the average number of customers in the system is given by

9.16. Consider the M/M/1/K queueing system. Show that

In the M/M/1/K queueing system, the average number of customers in the system is

The average number of customers in the queue is

Acustomer arriving with the queue in state Sn has a wait time Tq that is the sum of n independent exponential r. v.’s, each with parameter μ.

The expected value of this sum is n/μ [Eq. (4.132)]. Thus, the average amount of time that a customer spends waiting in the queue is

Similarly, the amount of time that a customer spends in the system is

Note that Eqs. (9.57) to (9.59) are equivalent to Eqs. (9.31) to (9.33) (Prob. 9.27).

9.17. Derive Eqs. (9.35) and (9.36).

As in Prob. 9.10, from Eqs. (9.34) and (9.10), we have

Let ρ = λ /(sμ). Then Eqs. (9.60) and (9.61) can be rewritten as

which is Eq. (9.36). From Eq. (9.11), p0 is obtained by equating

Using the summation formula (9.56), we obtain

which is Eq. (9.35).

9.18. Derive Eq. (9.37).

Using Eq. (9.36) and (9.51), the average number of customers in the queue is given by

9.19. Consider an M/M/s/s queueing system. Find the probability that all servers are busy.

Setting K = s in Eqs. (9.60) and (9.61), we get

and p0 is obtained by equating

Thus,

The probability that all servers are busy is given by

Note that in an M/M/s/s queueing system, if an arriving customer finds that all servers are busy, the customer will turn away and is lost. In a telephone system with s trunks, ps is the portion of incoming calls which will receive a busy signal. Equation (9.64) is often referred to as Erlang’s loss (or B) formula and is commonly denoted as B(s, λ /μ).

9.20. An air freight terminal has four loading docks on the main concourse. Any aircraft which arrive when all docks are full are diverted to docks on the back concourse. The average aircraft arrival rate is 3 aircraft per hour. The average service time per aircraft is 2 hours on the main concourse and 3 hours on the back concourse. (a) Find the percentage of the arriving aircraft that are diverted to the

back concourse.

(b) If a holding area which can accommodate up to 8 aircraft is added to the main concourse, find the percentage of the arriving aircraft that are diverted to the back concourse and the expected delay time awaiting service.

(a) The service system at the main concourse can be modeled as an M/M/s/s queueing system with s = 4, λ = 3, , and λ /μ = 6. The percentage of the arriving aircraft that are diverted to the back concourse is

100 × P(all docks on the main concourse are full)

From Eq. (9.64).

Thus, the percentage of the arriving aircraft that are diverted to the back concourse is about 47 percent.

(b) With the addition of a holding area for 8 aircraft, the service system at the main concourse can now be modeled as an M/M/s/K queueing system with s = 4, K = 12, and ρ = λ /(sμ) = 1.5. Now, from Eqs. (9.35) and (9.36),

Thus, about 33.2 percent of the arriving aircraft will still be diverted to the back concourse.

Next, from Eq. (9.37), the average number of aircraft in the queue is

Then, from Eq. (9.40), the expected delay time waiting for service is

Note that when the 2-hour service time is added, the total expected processing time at the main concourse will be 5.022 hours compared to the 3-hour service time at the back concourse.

SUPPLEMENTARY PROBLEMS

9.21. Customers arrive at the express checkout lane in a supermarket in a Poisson process with a rate of 15 per hour. The time to check out a customer is an exponential r. v. with mean of 2 minutes. (a) Find the average number of customers present. (b) What is the expected idle delay time experienced by a customer? (c) What is the expected time for a customer to clear a system?

9.22. Consider an M/M/1 queueing system. Find the probability of finding at least k customers in the system.

9.23. In a university computer center, 80 jobs an hour are submitted on the average. Assuming that the computer service is modeled as an M/M/1 queueing system, what should the service rate be if the average turnaround time (time at submission to time of getting job back) is to be less than 10 minutes?

9.24. The capacity of a communication line is 2000 bits per second. The line is used to transmit 8-bit characters, and the total volume of expected calls for transmission from many devices to be sent on the line is 12,000 characters per minute. Find (a) the traffic intensity, (b) the average number of characters waiting to be transmitted, and (c) the average transmission (including queueing delay) time per character.

9.25. Abank counter is currently served by two tellers. Customers entering the bank join a single queue and go to the next available teller when they reach the head of the line. On the average, the service time for a customer is 3 minutes, and 15 customers enter the bank per hour. Assuming that the arrivals process is Poisson and the service time is an exponential r. v., find the probability that a customer entering the bank will have to wait for service.

9.26. Apost office has three clerks serving at the counter. Customers arrive on the average at the rate of 30 per hour, and arriving customers are asked to form a single queue. The average service time for each customer is 3 minutes. Assuming that the arrivals process is Poisson and the service time is an exponential r. v., find (a) the probability that all the clerks will be busy, (b) the average number of customers in the queue, and (c) the average length of time customers have to spend in the post office.

9.27. Show that Eqs. (9.57) to (9.59) and Eqs. (9.31) to (9.33) are equivalent.

9.28. Find the average number of customers L in the M/M/1/K queueing system when λ = μ.

9.29. Agas station has one diesel fuel pump for trucks only and has room for three trucks (including one at the pump). On the average, trucks arrive

at the rate of 4 per hour, and each truck takes 10 minutes to service. Assume that the arrivals process is Poisson and the service time is an exponential r. v. (a) What is the average time for a truck from entering to leaving the

station? (b) What is the average time for a truck to wait for service? (c) What percentage of the truck traffic is being turned away?

9.30. Consider the air freight terminal service of Prob. 9.20. How many additional docks are needed so that at least 80 percent of the arriving aircraft can be served in the main concourse with the addition of holding area?

ANSWERS TO SUPPLEMENTARY PROBLEMS

9.21. (a) 1; (b) 2 min; (c) 4 min

9.22. ρk = (λ /μ)k

9.23. 1.43 jobs per minute

9.24. (a) 0.8; (b) 3.2; (c) 20 ms

9.25. 0.205

9.26. (a) 0.237; (b) 0.237; (c) 3.947 min

9.27. Hint: Use Eq. (9.29).

9.28. K/2

9.29. (a) 20.15 min; (b) 10.14 min; (c) 12.3 percent

9.30. 4

CHAPTER 10

Information Theory

10.1 Introduction

Information theory provides a quantitative measure of the information contained in message signals and allows us to determine the capacity of a communication system to transfer this information from source to destination. In this chapter we briefly explore some basic ideas involved in information theory.

10.2 Measure of Information

A. Information Sources:

An information source is an object that produces an event, the outcome of which is selected at random according to a probability distribution. A practical source in a communication system is a device that produces messages, and it can be either analog or discrete. In this chapter we deal mainly with the discrete sources, since analog sources can be transformed to discrete sources through the use of sampling and quantization techniques. A discrete information source is a source that has only a finite set of symbols as possible outputs. The set of source symbols is called the source alphabet, and the elements of the set are called symbols or letters.

Information sources can be classified as having memory or being memoryless. A source with memory is one for which a current symbol depends on the previous symbols. A memoryless source is one for which each symbol produced is independent of the previous symbols.

A discrete memoryless source (DMS) can be characterized by the list of the symbols, the probability assignment to these symbols, and the specification of the rate of generating these symbols by the source.

B. Information Content of a Discrete Memoryless Source:

The amount of information contained in an event is closely related to its uncertainty. Messages containing knowledge of high probability of occurrence convey relatively little information. We note that if an event is certain (that is, the event occurs with probability 1), it conveys zero information.

Thus, a mathematical measure of information should be a function of the probability of the outcome and should satisfy the following axioms:

1. Information should be proportional to the uncertainty of an outcome. 2. Information contained in independent outcomes should add.

1. Information Content of a Symbol: Consider a DMS, denoted by X, with alphabet {x1, x2, …, xm}. The information content of a symbol xi, denoted by I(xi), is defined by

where P(xi) is the probability of occurrence of symbol xi. Note that I(xi) satisfies the following properties:

The unit of I(xi) is the bit (binary unit) if b = 2, Hartley or decit if b = 10, and nat (natural unit) if b = e. It is standard to use b = 2. Here the unit bit (abbreviated “b”) is a measure of information content and is not to be confused with the term bit meaning “binary digit.” The conversion of these units to other units can be achieved by the following relationships.

2. Average Information or Entropy: In a practical communication system, we usually transmit long sequences of symbols from an information source. Thus, we are more interested in the average information that a source produces than the information content of a single symbol.

The mean value of I(xi) over the alphabet of source X with m different symbols is given by

The quantity H(X) is called the entropy of source X. It is a measure of the average information content per source symbol. The source entropy H(X) can be considered as the average amount of uncertainty within source X that is resolved by use of the alphabet.

Note that for a binary source X that generates independent symbols 0 and 1 with equal probability, the source entropy H(X) is

The source entropy H(X) satisfies the following relation:

where m is the size (number of symbols) of the alphabet of source X (Prob. 10.5). The lower bound corresponds to no uncertainty, which occurs when one symbol has probability P(xi) = 1 while P(xj) = 0 for j ≠ i, so X emits the same symbol xi all the time. The upper bound corresponds to the maximum uncertainty which occurs when P(xi) = 1/m for all i—that is, when all symbols are equally likely to be emitted by X.

3. Information Rate: If the time rate at which source X emits symbols is r (symbols/s), the information rate R of the source is given by

10.3 Discrete Memoryless Channels

A. Channel Representation:

A communication channel is the path or medium through which the symbols flow to the receiver. A discrete memoryless channel (DMC) is a statistical model with an input X and an output Y (Fig. 10-1). During each unit of the time (signaling interval), the channel accepts an input symbol from X, and in response it generates an output symbol from Y. The channel is “discrete” when the alphabets of X and Y are both finite. It is “memoryless” when the current output depends on only the current input and not on any of the previous inputs.

Fig. 10-1 Discrete memoryless channel.

A diagram of a DMC with m inputs and n outputs is illustrated in Fig. 10-1. The input X consists of input symbols x1, x2, …, xm. The a priori probabilities of these source symbols P(xi) are assumed to be known. The output Y consists of output symbols y1, y2, …, yn. Each possible input-to-output path is indicated along with a conditional probability P(yj|xi), where P(yj|xi) is the conditional probability of obtaining output yj given that the input is xi, and is called a channel transition probability.

B. Channel Matrix: A channel is completely specified by the complete set of transition probabilities. Accordingly, the channel of Fig. 10-1 is often specified by the matrix of transition probabilities [P(Y| X)], given by

The matrix [P(Y| X)] is called the channel matrix. Since each input to the channel results in some output, each row of the channel matrix must sum to unity; that is,

Now, if the input probabilities P(X) are represented by the row matrix

and the output probabilities P(Y) are represented by the row matrix

then

If P(X) is represented as a diagonal matrix

then

where the (i, j) element of matrix [P(X, Y)] has the form P(xi, yj). The matrix [P(X, Y)] is known as the joint probability matrix, and the element P(xi, yj) is the joint probability of transmitting xi and receiving yj.

C. Special Channels: 1. Lossless Channel: A channel described by a channel matrix with only one nonzero element in each column is called a lossless channel. An example of a lossless channel is shown in Fig. 10-2, and the corresponding channel matrix is shown in Eq. (10.18).

Fig. 10-2 Lossless channel.

It can be shown that in the lossless channel no source information is lost in transmission. [See Eq. (10.35) and Prob. 10.10.]

2. Deterministic Channel: A channel described by a channel matrix with only one nonzero element in each row is called a deterministic channel. An example of a deterministic channel is shown in Fig. 10-3, and the corresponding channel matrix is shown in Eq. (10.19).

Fig. 10-3 Deterministic channel.

Note that since each row has only one nonzero element, this element must be unity by Eq. (10.12). Thus, when a given source symbol is sent in the deterministic channel, it is clear which output symbol will be received.

3. Noiseless Channel: A channel is called noiseless if it is both lossless and deterministic. A noiseless channel is shown in Fig. 10-4. The channel matrix has only one element in each row and in each column, and this element is unity. Note that the input and output alphabets are of the same size; that is, m = n for the noiseless channel.

Fig. 10-4 Noiseless channel.

4. Binary Symmetric Channel:

The binary symmetric channel (BSC) is defined by the channel diagram shown in Fig. 10-5, and its channel matrix is given by

The channel has two inputs (x1 = 0, x2 = 1) and two outputs (y1 = 0, y2 = 1). The channel is symmetric because the probability of receiving a 1 if a 0 is sent is the same as the probability of receiving a 0 if a 1 is sent. This common transition probability is denoted by p.

Fig. 10-5 Binary symmetrical channel.

10.4 Mutual Information

A. Conditional and Joint Entropies:

Using the input probabilities P(xi), output probabilities P(yj), transition probabilities P(yj | xi), and joint probabilities P(xi, yj), we can define the following various entropy functions for a channel with m inputs and n outputs:

These entropies can be interpreted as follows: H(X) is the average uncertainty of the channel input, and H(Y) is the average uncertainty of the channel output. The conditional entropy H(X| Y) is a measure of the average uncertainty remaining about the channel input after the channel output has been observed. And H(X|Y) is sometimes called the equivocation of X with respect to Y. The conditional entropy H(Y|X) is the average uncertainty of the channel output given that X was transmitted. The joint entropy H(X, Y) is the average uncertainty of the communication channel as a whole.

Two useful relationships among the above various entropies are

Note that if X and Y are independent, then

B. Mutual Information: The mutual information I(X; Y) of a channel is defined by

Since H(X) represents the uncertainty about the channel input before the channel output is observed and H(X|Y) represents the uncertainty about the channel input after the channel output is observed, the mutual information I(X; Y) represents the uncertainty about the channel input that is resolved by observing the channel output.

Properties of I(X; Y):

Note that from Eqs. (10.31), (10.33), and (10.34) we have

with equality if X and Y are independent.

C. Relative Entropy: The relative entropy between two pmf’s p(xi), and q(xi) on X is defined as

The relative entropy, also known as Kullback-Leibler divergence, measures the “closeness” of one distribution from another. It can be shown that (Prob. 10.15)

and equal zero if p(xi) = q(xi). Note that, in general, it is not symmetric, that is, D(p // q) ≠ D(q // p).

The mutual information I(X; Y) can be expressed as the relative entropy between the joint distribution pXY (x, y) and the product of distribution pX(x) pY (y); that is,

10.5 Channel Capacity

A. Channel Capacity per Symbol Cs:

The channel capacity per symbol of a DMC is defined as

where the maximization is over all possible input probability distributions {P(xi)} on X. Note that the channel capacity Cs is a function of only the channel transition probabilities that define the channel.

B. Channel Capacity per Second C: If r symbols are being transmitted per second, then the maximum rate of transmission of information per second is rCs. This is the channel capacity per second and is denoted by C (b/ s):

C. Capacities of Special Channels: 1. Lossless Channel: For a lossless channel, H(X|Y) = 0 (Prob. 10.12) and

Thus, the mutual information (information transfer) is equal to the input (source) entropy, and no source information is lost in transmission. Consequently, the channel capacity per symbol is

where m is the number of symbols in X.

2. Deterministic Channel: For a deterministic channel, H(Y|X) = 0 for all input distributions P(xi), and

Thus, the information transfer is equal to the output entropy. The channel capacity per symbol is

where n is the number of symbols in Y.

3. Noiseless Channel: Since a noiseless channel is both lossless and deterministic, we have

and the channel capacity per symbol is

4. Binary Symmetric Channel: For the BSC of Fig. 10-5, the mutual information is (Prob. 10.20)

and the channel capacity per symbol is

10.6 Continuous Channel

In a continuous channel an information source produces a continuous signal x(t). The set of possible signals is considered as an ensemble of waveforms generated by some ergodic random process. It is further assumed that x(t) has a finite bandwidth so that x(t) is completely characterized by its periodic sample values. Thus, at any sampling instant, the collection of possible sample values constitutes a continuous random variable X described by its probability density function ƒX(x).

A. Differential Entropy: The average amount of information per sample value of x(t) is measured by

The entropy defined by Eq. (10.53) is known as the differential entropy of a continuous r.v. X with pdf ƒX(x). Note that as opposed to the discrete case (see Eq. (10.9)), the differential entropy can be negative (see Prob. 10.24).

Properties of Differential Entropy:

Translation does not change the differential entropy. Equation (10.54) follows directly from the definition Eq. (10.53).

Equation (10.55) is proved in Prob. 10.29.

B. Mutual Information: The average mutual information in a continuous channel is defined (by analogy with the discrete case) as

or

where

C. Relative Entropy: Similar to the discrete case, the relative entropy (or Kullback-Leibler divergence) between two pdf’s ƒX (x) and gX(x) on continuous r.v. X is defined by

From this definition, we can express the average mutual information I(X; Y) as

D. Properties of Differential Entropy, Relative Entropy, and Mutual Information:

with equality iff ƒ = g. (See Prob. 10.30.)

with equality iff X and Y are independent.

10.7 Additive White Gaussian Noise Channel

A. Additive White Gaussian Noise Channel:

An additive white Gaussian noise (AWGN) channel is depicted in Fig. 10-6.

Fig. 10-6 Gaussian channel.

This is a time discrete channel with output Yi at time i, where Yi is the sum of the input Xi and the noise Zi.

The noise Zi is drawn i.i.d. from a Gaussian distribution with zero mean and variance N. The noise Zi is assumed to be independent of the input signal Xi. We also assume that the average power of input signal is finite. Thus,

The capacity Cs of an AWGN channel is given by (Prob. 10.31)

where S/N is the signal-to-noise ratio at the channel output.

B. Band-Limited Channel: A common model for a communication channel is a band-limited channel with additive white Gaussian noise. This is a continuous-time channel.

Nyquist Sampling Theorem Suppose a signal x(t) is band-limited to B; namely, the Fourier transform of the signal x(t) is 0 for all frequencies greater than B(Hz). Then the signal x(t) is completely determined by its periodic sample values taken at the Nyquist rate 2B samples/sec. (For the proof of this sampling theorem, see any Fourier transform text.)

C. Capacity of the Continuous-Time Gaussian Channel: If the channel bandwidth B(Hz) is fixed, then the output y(t) is also a band-limited signal. Then the capacity C(b/s) of the continuous-time AWGN channel is given by (see Prob. 10.32)

Equation (10.72) is known as the Shannon-Hartley law. The Shannon-Hartley law underscores the fundamental role of bandwidth and

signal-to-noise ratio in communication. It also shows that we can exchange increased bandwidth for decreased signal power (Prob. 10.35) for a system with given capacity C.

10.8 Source Coding

A conversion of the output of a DMS into a sequence of binary symbols (binary code word) is called source coding. The device that performs this conversion is called the source encoder (Fig. 10-7).

Fig. 10-7 Source coding.

An objective of source coding is to minimize the average bit rate required for representation of the source by reducing the redundancy of the information source.

A. Code Length and Code Efficiency: Let X be a DMS with finite entropy H(X) and an alphabet {x1, …, xm} with corresponding probabilities of occurrence P(xi)(i = 1, …, m). Let the binary code word assigned to symbol xi by the encoder have length ni, measured in bits. The length of a code word is the number of binary digits in the code word. The average code word length L, per source symbol, is given by

The parameter L represents the average number of bits per source symbol used in the source coding process.

The code efficiency η is defined as

where Lmin is the minimum possible value of L. When η approaches unity, the code is said to be efficient.

The code redundancy γ is defined as

B. Source Coding Theorem:

The source coding theorem states that for a DMS X with entropy H(X), the average code word length L per symbol is bounded as (Prob. 10.39)

and further, L can be made as close to H(X) as desired for some suitably chosen code.

Thus, with Lmin = H(X), the code efficiency can be rewritten as

C. Classification of Codes: Classification of codes is best illustrated by an example. Consider Table 10-1 where a source of size 4 has been encoded in binary codes with symbol 0 and 1.

TABLE 10-1 Binary Codes

1. Fixed-Length Codes: A fixed-length code is one whose code word length is fixed. Code 1 and code 2 of Table 10-1 are fixed-length codes with length 2.

2. Variable-Length Codes: A variable-length code is one whose code word length is not fixed. All codes of Table 10-1 except codes 1 and 2 are variable-length codes.

3. Distinct Codes: A code is distinct if each code word is distinguishable from other code words. All codes of Table 10-1 except code 1 are distinct codes—notice the codes for x1 and x3.

4. Prefix-Free Codes:

A code in which no code word can be formed by adding code symbols to another code word is called a prefix-free code. Thus, in a prefix-free code no code word is a prefix of another. Codes 2, 4, and 6 of Table 10-1 are prefix-free codes.

5. Uniquely Decodable Codes: A distinct code is uniquely decodable if the original source sequence can be reconstructed perfectly from the encoded binary sequence. Note that code 3 of Table 10-1 is not a uniquely decodable code. For example, the binary sequence 1001 may correspond to the source sequences x2x3x2 or x2x1x1x2. A sufficient condition to ensure that a code is uniquely decodable is that no code word is a prefix of another. Thus, the prefix-free codes 2, 4, and 6 are uniquely decodable codes. Note that the prefix-free condition is not a necessary condition for unique decodability. For example, code 5 of Table 10-1 does not satisfy the prefix-free condition, and yet it is uniquely decodable since the bit 0 indicates the beginning of each code word of the code.

6. Instantaneous Codes: A uniquely decodable code is called an instantaneous code if the end of any code word is recognizable without examining subsequent code symbols. The instantaneous codes have the property previously mentioned that no code word is a prefix of another code word. For this reason, prefix-free codes are sometimes called instantaneous codes.

7. Optimal Codes: A code is said to be optimal if it is instantaneous and has minimum average length L for a given source with a given probability assignment for the source symbols.

D. Kraft Inequality: Let X be a DMS with alphabet {xi} (i = 1, 2, …, m). Assume that the length of the assigned binary code word corresponding to xi is ni.

A necessary and sufficient condition for the existence of an instantaneous binary code is

which is known as the Kraft inequality. (See Prob. 10.43.)

Note that the Kraft inequality assures us of the existence of an instantaneously decodable code with code word lengths that satisfy the inequality. But it does not show us how to obtain these code words, nor does it say that any code that satisfies the inequality is automatically uniquely decodable (Prob. 10.38).

10.9 Entropy Coding

The design of a variable-length code such that its average code word length approaches the entropy of the DMS is often referred to as entropy coding. This section presents two examples of entropy coding.

A. Shannon-Fano Coding: An efficient code can be obtained by the following simple procedure, known as Shannon-Fano algorithm:

1. List the source symbols in order of decreasing probability. 2. Partition the set into two sets that are as close to equiprobable as possible, and

assign 0 to the upper set and 1 to the lower set. 3. Continue this process, each time partitioning the sets with as nearly equal

probabilities as possible until further partitioning is not possible.

An example of Shannon-Fano encoding is shown in Table 10-2. Note that in Shannon-Fano encoding the ambiguity may arise in the choice of approximately equiprobable sets. (See Prob. 10.46.) Note also that Shannon-Fano coding results in suboptimal code.

TABLE 10-2 Shannon-Fano Encoding

B. Huffman Encoding: In general, Huffman encoding results in an optimum code. Thus, it is the code that has the highest efficiency (Prob. 10.47). The Huffman encoding procedure is as follows:

1. List the source symbols in order of decreasing probability. 2. Combine the probabilities of the two symbols having the lowest probabilities,

and reorder the resultant probabilities; this step is called reduction 1. The same procedure is repeated until there are two ordered probabilities remaining.

3. Start encoding with the last reduction, which consists of exactly two ordered probabilities. Assign 0 as the first digit in the code words for all the source symbols associated with the first probability; assign 1 to the second probability.

4. Now go back and assign 0 and 1 to the second digit for the two probabilities that were combined in the previous reduction step, retaining all assignments made in Step 3.

5. Keep regressing this way until the first column is reached.

An example of Huffman encoding is shown in Table 10-3.

TABLE 10-3 Huffman Encoding

Note that the Huffman code is not unique depending on Huffman tree and the labeling (see Prob. 10.33).

SOLVED PROBLEMS

Measure of Information

10.1. Consider event E occurred when a random experiment is performed with probability p. Let I(p) be the information content (or surprise measure) of event E, and assume that it satisfies the following axioms:

Then show that I(p) can be expressed as

where C is an arbitrary positive integer.

From Axiom 5, we have

and by induction, we have

Also, for any integer n, we have

and

Thus, from Eqs. (10.80) and (10.81), we have

or

where r is any positive rational number. Then by Axiom 4

where α is any nonnegative number.

Let α = − log2 p (0 < p ≤ 1). Then p = (1/ 2) α and from Eq. (10.82), we have

where C = I (1/ 2) > I(1) = 0 by Axioms 2 and 3. Setting C = 1, we have I(p) =−log2 p bits.

10.2. Verify Eq. (10.5); that is,

If xi and xj are independent, then by Eq. (3.22)

By Eq. (10.1)

10.3. A DMS X has four symbols x1, x2, x3, x4 with probabilities P(x1) = 0.4, P(x2) = 0.3, P(x3) = 0.2, P(x4) = 0.1. (a) Calculate H(X). (b) Find the amount of information contained in the messages x1x2x1x3 and

x4x3x3x2, and compare with the H(X) obtained in part (a).

(a)

(b)

Thus,

Thus,

10.4. Consider a binary memoryless source X with two symbols x1 and x2. Show that H(X) is maximum when both x1 and x2 are equiprobable.

Let P(x1) = α. P(x2) = 1 − α.

Using the relation

we obtain

The maximum value of H(X) requires that

that is,

Note that H(X) = 0 when α = 0 or 1. When is maximum and is given by

10.5. Verify Eq. (10.9); that is,

where m is the size of the alphabet of X.

Proof of the lower bound: Since 0 ≤ P(xi) ≤ 1,

Then it follows that

Thus,

Next, we note that

if and only if P(xi) = 0 or 1. Since

when P(xi) = 1, then P(xj) = 0 for j ≠ i. Thus, only in this case, H(X) = 0.

Proof of the upper bound: Consider two probability distributions {P(xi) = Pi} and {Q(xi) = Qi} on the alphabet {xi}, i = 1, 2, …, m, such that

Using Eq. (10.6), we have

Next, using the inequality

and noting that the equality holds only if α = 1, we get

by using Eq. (10.86). Thus,

where the equality holds only if Qi = Pi for all i. Setting

we obtain

Hence,

and the equality holds only if the symbols in X are equiprobable, as in Eq. (10.90).

10.6. Find the discrete probability distribution of X, pX (xi), which maximizes information entropy H(X).

From Eqs. (10.7) and (2.17), we have

Thus, the problem is the maximization of Eq. (10.92) with constraint Eq. (10.93). Then using the method of Lagrange multipliers, we set Lagrangian J as

where λ is the Lagrangian multiplier. Taking the derivative of J with respect to pX(xi) and setting equal to zero, we obtain

and

This shows that all px (xi) are equal (because they depend on λ only) and using the constraint Eq. (10.93), we obtain px(xi) = 1/m. Hence, the uniform distribution is the distribution with the maximum entropy. (cf. Prob. 10.5.)

10.7. A high-resolution black-and-white TV picture consists of about 2 × 106 picture elements and 16 different brightness levels. Pictures are repeated at the rate of 32 per second. All picture elements are assumed to be independent, and all levels have equal likelihood of occurrence. Calculate the average rate of information conveyed by this TV picture source.

Hence, by Eq. (10.10)

10.8. Consider a telegraph source having two symbols, dot and dash. The dot duration is 0.2 s. The dash duration is 3 times the dot duration. The probability of the dot’s occurring is twice that of the dash, and the time between symbols is 0.2 s. Calculate the information rate of the telegraph source.

Thus,

By Eq. (10.7)

Thus, the average time per symbol is

and the average symbol rate is

Thus, the average information rate of the telegraph source is

Discrete Memoryless Channels

10.9. Consider a binary channel shown in Fig. 10-8.

Fig. 10-8

(a) Find the channel matrix of the channel. (b) Find P(y1) and P(y2) when P(x1) = P(x2) = 0.5. (c) Find the joint

probabilities P(x1, y2) and P(x2, y1) when P(x1) = P(x2) = 0.5.

(a) Using Eq. (10.11), we see the channel matrix is given by

(b) Using Eqs. (10.13), (10.14), and (10.15), we obtain

Hence, P(y1) = 0.55 and P(y2) = 0.45. (c) Using Eqs. (10.16) and (10.17), we obtain

Hence, P(x1, y2) = 0.05 and P(x2, y1) = 0.1.

10.10. Two binary channels of Prob. 10.9 are connected in a cascade, as shown in Fig. 10-9.

Fig. 10-9

(a) Find the overall channel matrix of the resultant channel, and draw the resultant equivalent channel diagram.

(b) Find P(z1) and P(z2) when P(x1) = P(x2) = 0.5.

(a) By Eq. (10.15)

Thus, from Fig. 10-9

The resultant equivalent channel diagram is shown in Fig. 10-10.

Fig. 10-10

(b)

Hence, P(z1) = 0.585 and P(z2) = 0.415.

10.11. A channel has the following channel matrix:

(a) Draw the channel diagram. (b) If the source has equally likely outputs, compute the probabilities

associated with the channel outputs for p = 0.2.

(a) The channel diagram is shown in Fig. 10-11. Note that the channel represented by Eq. (10.96) (see Fig. 10-11) is known as the binary erasure channel. The binary erasure channel has two inputs x1 = 0 and x2 = 1 and three outputs y1 = 0, y2 = e, and y3 = 1, where e indicates an erasure; that is, the output is in doubt, and it should be erased.

Fig. 10-11 Binary erasure channel.

(b) By Eq. (10.15)

Thus, P(y1) = 0.4, P(y2) = 0.2, and P(y3) = 0.4.

Mutual Information

10.12. For a lossless channel show that

When we observe the output yj in a lossless channel (Fig. 10-2), it is clear which xi was transmitted; that is,

Now by Eq. (10.23)

Note that all the terms in the inner summation are zero because they are in the form of 1 × log2 1 or 0 × log2 0. Hence, we conclude that for a lossless channel

10.13. Consider a noiseless channel with m input symbols and m output symbols (Fig. 10-4). Show that

and

For a noiseless channel the transition probabilities are

Hence,

and

Thus, by Eqs. (10.7) and (10.104)

Next, by Eqs. (10.24), (10.102), and (10.103)

10.14. Verify Eq. (10.26); that is,

From Eqs. (3.34) and (3.37)

and

So by Eq. (10.25) and using Eqs. (10.22) and (10.23), we have

10.15. Verify Eq. (10.40); that is,

By definition of D(p // q), Eq. (10.39), we have

Since minus the logarithm is convex, then by Jensen’s inequality (Eq. 4.40), we obtain

Now

Thus, D(p // q) ≥ 0. Next, when p(xi) = q(xi), then log2 (q / p) = log2 1 = 0, and we have D(p // q) = 0. On the other hand, if D(p // q) = 0, then log2(q / p) = 0, which implies that q / p = 1; that is, p(xi) = q(x i). Thus, we have shown that D(p // q) ≥ 0 and equality holds iff p(xi) = q(xi).

10.16. Verify Eq. (10.41); that is,

From Eqs. (10.31) and using Eqs. (10.21) and (10.23), we obtain

10.17. Verify Eq. (10.32); that is,

Since pXY(x, y) = pYX(y, x), pX (x) pY (y) = pY(y) pX(x), by Eq. (10.41), we have

10.18. Verify Eq. (10.33); that is

From Eqs. (10.41) and (10.40), we have

10.19. Using Eq. (10.42) verify Eq. (10.35); that is

From Eq. (10.42) we have

10.20. Consider a BSC (Fig. 10-5) with P(x1) = α. (a) Show that the mutual information I(X; Y) is given by

(b) Calculate I(X; Y) for α = 0.5 and p = 0.1. (c) Repeat (b) for α = 0.5 and p = 0.5, and comment on the result.

Figure 10-12 shows the diagram of the BSC with associated input probabilities.

Fig. 10-12

(a) Using Eqs. (10.16), (10.17), and (10.20), we have

Then by Eq. (10.24)

Hence, by Eq. (10.31)

(b) When α = 0.5 and p = 0.1, by Eq. (10.15)

Thus, P(y1) = P(y2) = 0.5. By Eq. (10.22)

Thus,

(c) When α = 0.5 and p = 0.5,

Thus,

Note that in this case (p = 0.5) no information is being transmitted at all. An equally acceptable decision could be made by dispensing with the channel entirely and “flipping a coin” at the receiver. When I(X; Y) = 0, the channel is said to be useless.

Channel Capacity

10.21. Verify Eq. (10.46); that is,

where Cs is the channel capacity of a lossless channel and m is the number of symbols in X.

For a lossless channel [Eq. (10.97), Prob. 10.12]

Then by Eq. (10.31)

Hence, by Eqs. (10.43) and (10.9)

10.22. Verify Eq. (10.52); that is,

where Cs is the channel capacity of a BSC (Fig. 10-5).

By Eq. (10.106) (Prob. 10.20) the mutual information I(X; Y) of a BSC is given by

which is maximum when H(Y) is maximum. Since the channel output is binary, H(Y) is maximum when each output has a probability of 0.5 and is achieved for equally likely inputs [Eq. (10.9)]. For this case H(Y) = 1, and the channel capacity is

10.23. Find the channel capacity of the binary erasure channel of Fig. 10-13 (Prob. 10.11).

Let P(x1) = α. Then P(x2) = 1 − α. By Eq. (10.96)

Fig. 10-13

By Eq. (10.15)

By Eq. (10.17)

In addition, from Eqs. (10.22) and (10.24) we can calculate

Thus, by Eqs. (10.34) and (10.83)

And by Eqs. (10.43) and (10.84)

Continuous Channel

10.24. Find the differential entropy H(X) of the uniformly distributed random variable X with probability density function

for

By Eq. (10.53)

Note that the differential entropy H(X) is not an absolute measure of information, and unlike discrete entropy, differential entropy can be negative.

10.25. Find the probability density function ƒX(x) of X for which differential entropy H(X) is maximum.

Let the support of ƒX(x) be (a, b). [Note that the support of ƒX(x) is the region where ƒX(x) > 0.] Now

Using Lagrangian multipliers technique, we let

where λ is the Lagrangian multiplier. Taking the functional derivative of J with respect to ƒX(x) and setting equal to zero, we obtain

or

so ƒX(x) is a constant. Then, by constraint Eq. (10.115), we obtain

Thus, the uniform distribution results in the maximum differential entropy.

10.26. Let X be N (0; σ2). Find its differential entropy.

The differential entropy in nats is expressed as

By Eq. (2.71), the pdf of N(0, σ2) is

Then

Thus, we obtain

Changing the base of the logarithm, we have

10.27. Let (X1, …, Xn) be an n-variate r.v. defined by Eq. (3.92). Find the differential entropy of (X1,…, Xn).

From Eq. (3.92) the joint pdf of (X1, …, Xn) is given by

where

Then, we obtain

10.28. Find the probability density function ƒX(x) of X with zero mean and variance σ2 for which differential entropy H(X) is maximum.

The differential entropy in nats is expressed as

The constraints are

Using the method of Lagrangian multipliers, and setting Lagrangian as

and taking the functional derivative of J with respect to ƒX(x) and setting equal to zero, we obtain

or

It is obvious from the generic form of the exponential family restricted to the second order polynomials that they cover Gaussian distribution only. Thus, we obtain N(0; σ2) and

10.29. Verify Eq. (10.55); that is,

Let Y = a X. Then by Eq. (4.86)

and

After a change of variables y / a = x, we have

10.30. Verify Eq. (10.64); that is,

By definition (10.61), we have

Thus, D (ƒ // g) ≥ 0. We have equality in Jensen’s inequality which occurs iff ƒ = g.

Additive White Gaussian Noise Channel

10.31. Verify Eq. (10.71); that is,

From Eq. (10.68), Y = X + Z. Then by Eq. (10.57) we have

since Z is independent of X.

Now, from Eq. (10.124) (Prob. 10.26) and setting σ2 = N, we have

Since X and Z are independent and E(Z) = 0, we have

Given E(Y2) = S + N, the differential entropy of Y is bounded by , since the Gaussian distribution maximizes the differential

entropy for a given variance (see Prob. 10.28). Applying this result to bound the mutual information, we obtain

Hence, the capacity Cs of the AWGN channel is

10.32. Verify Eq. (10.72); that is,

Assuming that the channel bandwidth is B Hz, then by Nyquist sampling theorem we can represent both the input and output by samples taken 1/(2B)

seconds apart. Each of the input samples is corrupted by the noise to produce the corresponding output sample. Since the noise is white and Gaussian, each of the noise samples is an i.i.d. Gaussian r. v.. If the noise has power spectral density N0 / 2 and bandwidth B, then the noise has power (N0 / 2) 2B = N0 B and each of the 2BT noise samples in time t has variance N0BT / (2BT) = N0 / 2. Now the capacity of the discrete time Gaussian channel is given by (Eq. (10.71))

Let the channel be used over the time interval [0, T]. Then the power per sample is ST / (2 BT) = S / (2 B), the noise variance per sample is N0 / 2, and hence the capacity per sample is

Since there are 2B samples each second, the capacity of the channel can be rewritten as

where N = N0 B is the total noise power.

10.33. Show that the channel capacity of an ideal AWGN channel with infinite bandwidth is given by

where S is the average signal power and η/2 is the power spectral density of white Gaussian noise.

The noise power N is given by N = ηB. Thus, by Eq. (10.72)

Let S/(ηB) = λ. Then

Now

Since , we obtain

Note that Eq. (10.140) can be used to estimate upper limits on the performance of any practical communication system whose transmission channel can be approximated by the AWGN channel.

10.34. Consider an AWGN channel with 4-kHz bandwidth and the noise power spectral density η /2 = 10−12 W / Hz. The signal power required at the receiver is 0.1 mW. Calculate the capacity of this channel.

Thus,

And by Eq. (10.72)

10.35. An analog signal having 4-kHz bandwidth is sampled at 1.25 times the Nyquist rate, and each sample is quantized into one of 256 equally likely levels. Assume that the successive samples are statistically independent. (a) What is the information rate of this source? (b) Can the output of this source be transmitted without error over an AWGN

channel with a bandwidth of 10 kHz and an S / N ratio of 20 dB? (c) Find the S / N ratio required for error-free transmission for part (b). (d) Find the bandwidth required for an AWGN channel for error-free

transmission of the output of this source if the S / N ratio is 20 dB.

(a)

By Eq (10.10) the information rate R of the source is

(b) By Eq. (10.72)

Since R > C, error-free transmission is not possible. (c) The required S/N ratio can be found by

or

or

Thus, the required S/N ratio must be greater than or equal to 24.1 dB for error-free transmission.

(d) The required bandwidth B can be found by

or

and the required bandwidth of the channel must be greater than or equal to 12 kHz.

Source Coding

10.36. Consider a DMS X with two symbols x1 and x2 and P(x1) = 0.9, P(x2) = 0.1. Symbols x1 and x2 are encoded as follows (Table 10-4):

TABLE 10-4

Find the efficiency η and the redundancy γ of this code.

By Eq. (10.73) the average code length L per symbol is

By Eq. (10.7)

Thus, by Eq. (10.77) the code efficiency η is

By Eq. (10.75) the code redundancy γ is

10.37. The second-order extension of the DMS X of Prob. 10.36, denoted by X2, is formed by taking the source symbols two at a time. The coding of this extension is shown in Table 10-5. Find the efficiency η and the redundancy γ of this extension code.

TABLE 10-5

The entropy of the second-order extension of X, H(X2), is given by

Therefore, the code efficiency η is

and the code redundancy γ is

Note that H(X2) = 2H(X).

10.38. Consider a DMS X with symbols xi, i = 1, 2, 3, 4. Table 10-6 lists four possible binary codes.

TABLE 10-6

(a) Show that all codes except code B satisfy the Kraft inequality. (b) Show that codes A and D are uniquely decodable but codes B and C are

not uniquely decodable.

(a) From Eq. (10.78) we obtain the following:

All codes except code B satisfy the Kraft inequality. (b) Codes A and D are prefix-free codes. They are therefore uniquely

decodable. Code B does not satisfy the Kraft inequality, and it is not uniquely decodable. Although code C does satisfy the Kraft inequality, it is not uniquely decodable. This can be seen by the following example: Given the binary sequence 0110110. This sequence may correspond to the source sequences x1 x2 x1 x4 or x1 x4 x4.

10.39. Verify Eq. (10.76); that is,

where L is the average code word length per symbol and H(X) is the source entropy.

From Eq. (10.89) (Prob. 10.5), we have

where the equality holds only if Qi = Pi. Let

where

which is defined in Eq. (10.78). Then

and

From the Kraft inequality (10.78) we have

Thus,

or

The equality holds when K = 1 and Pi = Qi.

10.40. Let X be a DMS with symbols xi and corresponding probabilities P(xi) = Pi, i = 1, 2, …, m. Show that for the optimum source encoding we require that

and

where ni is the length of the code word corresponding to xi and Ii is the information content of xi.

From the result of Prob. 10.39, the optimum source encoding with L = H(X) requires K = 1 and Pi = Qi. Thus, by Eqs. (10.143) and (10.142)

and

Hence,

Note that Eq. (10.149) implies the following commonsense principle: Symbols that occur with high probability should be assigned shorter code words than symbols that occur with low probability.

10.41. Consider a DMS X with symbols xi and corresponding probabilities P(xi) = Pi, i = 1, 2, …, m. Let ni be the length of the code word for xi such that

Show that this relationship satisfies the Kraft inequality (10.78), and find the bound on K in Eq. (10.78).

Equation (10.152) can be rewritten as

or

Then

or

Thus,

or

which indicates that the Kraft inequality (10.78) is satisfied, and the bound on K is

10.42. Consider a DMS X with symbols xi and corresponding probabilities P(xi) = Pi, i = 1, 2, …, m. Show that a code constructed in agreement with Eq. (10.152) will satisfy the following relation:

where H(X) is the source entropy and L is the average code word length.

Multiplying Eq. (10.153) by Pi and summing over i yields

Now

Thus, Eq. (10.159) reduces to

10.43. Verify Kraft inequality Eq. (10.78); that is,

Consider a binary tree representing the code words; (Fig. 10-14). This tree extends downward toward infinity. The path down the tree is the sequence of symbols (0, 1), and each leaf of the tree with its unique path corresponds to a code word. Since an instantaneous code is a prefix-free code, each code word eliminates its descendants as possible code words.

Fig. 10-14 Binary tree.

Let nmax be the length of the longest code word. Of all the possible nodes at a level of nmax, some may be code words, some may be descendants of code words, and some may be neither. Acode word of level ni has 2

nmax − ni

descendants at level nmax. The total number of possible leaf nodes at level nmax is 2

nmax. Hence, summing over all code words, we have

Dividing through by 2nmax, we obtain

Conversely, given any set of code words with length ni (i = 1, …, m) which satisfy the inequality, we can always construct a tree. First, order the code word lengths according to increasing length, then construct code words in terms of the binary tree introduced in Fig. 10-14.

10.44. Show that every source alphabet X = {x1, …, xm} has a binary prefix code.

Given source symbols x1, …, xm, choose the code length ni such that 2 ni ≥ m;

that is, ni ≥ log2 m. Then

Thus, Kraft’s inequality is satisfied and there is a prefix code.

Entropy Coding

10.45. A DMS X has four symbols x1, x2, x3, and x4 with , and . Construct a Shannon-Fano code for X; show that this

code has the optimum property that ni = I(xi) and that the code efficiency is 100 percent.

The Shannon-Fano code is constructed as follows (see Table 10-7):

TABLE 10-7

10.46. A DMS X has five equally likely symbols. (a) Construct a Shannon-Fano code for X, and calculate the efficiency of the

code. (b) Construct another Shannon-Fano code and compare the results. (c) Repeat for the Huffman code and compare the results.

(a) AShannon-Fano code [by choosing two approximately equiprobable (0.4 versus 0.6) sets] is constructed as follows (see Table 10-8):

TABLE 10-8

The efficiency η is

(b) Another Shannon-Fano code [by choosing another two approximately equiprobable (0.6 versus 0.4) sets] is constructed as follows (see Table 10-9):

TABLE 10-9

Since the average code word length is the same as that for the code of part (a), the efficiency is the same. (c) The Huffman code is constructed as follows (see Table 10-10):

Since the average code word length is the same as that for the Shannon-Fano code, the efficiency is also the same.

TABLE 10-10

10.47. A DMS X has five symbols x1, x2, x3, x4, and x5 with P(x1) = 0.4, P(x2) = 0.19, P(x3) = 0.16, P(x4) = 0.15, and P(x5) = 0.1. (a) Construct a Shannon-Fano code for X, and calculate the efficiency of the

code. (b) Repeat for the Huffman code and compare the results.

(a) The Shannon-Fano code is constructed as follows (see Table 10-11):

TABLE 10-11

(b) The Huffman code is constructed as follows (see Table 10-12):

TABLE 10-12

The average code word length of the Huffman code is shorter than that of the Shannon-Fano code, and thus the efficiency is higher than that of the Shannon-Fano code.

SUPPLEMENTARY PROBLEMS

10.48. Consider a source X that produces five symbols with probabilities

. Determine the source entropy H(X).

10.49. Calculate the average information content in the English language, assuming that each of the 26 characters in the alphabet occurs with equal probability.

10.50. Two BSCs are connected in cascade, as shown in Fig. 10-15.

Fig. 10-15

(a) Find the channel matrix of the resultant channel. (b) Find P(z1) and P(z2) if P(x1) = 0.6 and P(x2) = 0.4.

10.51. Consider the DMC shown in Fig. 10-16.

Fig. 10-16

(a) Find the output probabilities if

.

(b) Find the output entropy H(Y).

10.52. Verify Eq. (10.35), that is,

10.53. Show that H(X, Y) ≤ H(X) + H(Y) with equality if and only if X and Y are independent.

10.54. Show that for a deterministic channel

10.55. Consider a channel with an input X and an output Y. Show that if X and Y are statistically independent, then H(X|Y) = H(X) and I(X; Y) = 0.

10.56. Achannel is described by the following channel matrix. (a) Draw the channel diagram. (b) Find the channel capacity.

10.57. Let X be a random variable with probability density function ƒX(x), and let Y = aX + b, where a and b are constants. Find H(Y) in terms of H(X).

10.58. Show that H(X + c) = H(X), where c is a constant.

10.59. Show that H(X) ≥ H(X|Y), and H (Y) ≥ H(Y| X).

10.60. Verify Eq. (10.30), that is, H(X, Y) = H(X) + H(Y), if X and Y are independent.

10.61. Find the pdf ƒX (x) of a continuous r.v X with E(X) = μ which maximizes the differential entropy H(X).

10.62. Calculate the capacity of AWGN channel with a bandwidth of 1 MHz and an S/N ratio of 40 dB.

10.63. Consider a DMS X with m equiprobable symbols xi, i = 1, 2, …, m.

(a) Show that the use of a fixed-length code for the representation of xi is most efficient.

(b) Let n0 be the fixed code word length. Show that if n0 = log2m, then the code efficiency is 100 percent.

10.64. Construct a Huffman code for the DMS X of Prob. 10.45, and show that the code is an optimum code.

10.65. ADMS X has five symbols x1, x2, x3, x4, and x5 with respective probabilities 0.2, 0.15, 0.05, 0.1, and 0.5. (a) Construct a Shannon-Fano code for X, and calculate the code efficiency. (b) Repeat (a) for the Huffman code.

10.66. Show that the Kraft inequality is satisfied by the codes of Prob. 10.46.

ANSWERS TO SUPPLEMENTARY PROBLEMS

10.48. 1.875 b / symbol

10.49. 4.7 b / character

10.50.

10.51. (a) P(y1) = 7/24, P(y2) = 17/48, and P(y3) = 17/48 (b) 1.58 b / symbol

10.52. Hint: Use Eqs. (10.31) and (10.26).

10.53. Hint: Use Eqs. (10.33) and (10.35).

10.54. Hint: Use Eq. (10.24), and note that for a deterministic channel P(yj| xi) are either 0 or 1.

10.55. Hint: Use Eqs. (3.32) and (3.37) in Eqs. (10.23) and (10.31).

10.56. (a) See Fig. 10-17. (b) 1 b/symbol

Fig. 10-17

10.57. H(Y) = H(X) + log2 a

10.58. Hint: Let Y = X + c and follow Prob. 10.29.

10.59. Hint: Use Eqs. (10.31), (10.33), and (10.34).

10.60. Hint: Use Eqs. (10.28), (10.29), and (10.26).

10.61.

10.62. 13.29 Mb/s

10.63. Hint: Use Eqs. (10.73) and (10.76).

10.64.

10.65.

10.66. Hint: Use Eq. (10.78)

APPENDIX A

Normal Distribution

Fig. A

TABLE A Normal Distribution Φ(z)

The material below refers to Fig. A.

APPENDIX B

Fourier Transform

B.1 Continuous-Time Fourier Transform

Definition:

TABLE B-1 Properties of the Continuous-Time Fourier Transform

TABLE B-2 Common Continuous-Time Fourier Transform Pairs

B.2 Discrete-Time Fourier Transform

Definition:

TABLE B-3 Properties of the Discrete-Time Fourier Transform

TABLE B-4 Common Discrete-Time Fourier Transform Pairs

INDEX

Please note that index links point to page beginnings from the print edition. Locations are approximate in e-readers, and you may need to page down one or more times after clicking a link to get to the indexed material.

a priori probability, 333 a posteriori probability, 333 Absorbing barrier, 241

states, 215 Absorption, 215

probability, 215 Acceptance region, 331 Accessible states, 214 Additive white gaussian noise channel, 397 Algebra of sets, 2–6, 15 Alternative hypothesis, 331 Aperiodic states, 215 Arrival (or birth) parameter, 350 Arrival process, 216, 253 Assemble average, 208 Autocorrelation function, 208, 273 Autocovariance function, 209 Average information, 368 Axioms of probability, 8

Band-limited, channel, 376 white noise, 308

Bayes’, estimate, 314

estimator, 314 estimation, 314, 321 risk, 334 rule, 10 test, 334 theorem, 10

Bernoulli, distribution, 55 experiment, 44 process, 222 r.v., 55 trials, 44, 56

Best estimator, 314 Biased estimator, 316 Binary, communication channel, 117–118

erasure channel, 386, 392 symmetrical channel (BSC), 371

Binomial, distribution, 56 coefficient, 56 r.v., 56

Birth-death process, 350 Birth parameter, 350 Bivariate, normal distribution, 111

r.v., 101, 122 Bonferroni’s inequality, 23 Boole’s inequality, 24 Brownian motion process (see Wiener process) Buffon’s needle, 128

Cardinality, 4 Cauchy, criterion, 284

r.v., 98 Cauchy-Schwarz inequality, 133, 154, 186 Central limit theorem, 64, 158, 198–199

Chain, 208 Markov, 210

Channel, band-limited, 376 binary symmetrical, 371 capacity, 391, 393 continuous, 374, 393 deterministic, 370 discrete memoryless (DMC), 369 lossless, 370 matrix, 238, 369 noiseless, 371 representation, 369 transition probability, 369

Chapman-Kolomogorov equation, 212 Characteristic function, 156–157, 196 Chebyshev inequality, 86 Chi-square (χ2) r.v., 181 Code, classification, 377

distinct, 378 efficiency, 377 Huffman, 380 instantaneous, 378 length, 377 optimal, 378 prefix-free, 378 redundancy, 377 Shannon-Fano, 405 uniquely decidable, 378

Coding, entropy, 378, 405 source, 400

Complement of set, 2 Complex random process, 208, 280 Composite hypothesis, 331

Concave function, 153 Conditional, distribution, 64, 92, 105, 128

expectation, 107, 219 mean, 107, 135 probability, 10, 31 probability density function (pdf), 105 probability mass function (pmf), 105 variance, 107, 135

Confidence, coefficient, 324 interval, 324

Consistent estimator, 313 Continuity theorem of probability, 26 Continuity correction, 201 Convex function, 153 Convolution, 168, 276

integral, 276 sum, 277

Correlation, 107 coefficient, 106–107, 131, 208

Counting process, 217 Poisson, 217

Covariance, 106–107, 131, 208 matrix, 111 stationary, 230

Craps, 45 Critical region, 331 Cross-correlation function, 273, Cross power spectral density (or spectrum), 274 Cumulative distribution function (cdf), 52

Death parameter, 351 Decision test, 332, 338

Bayes’, 334

likelihood ratio, 333 MAP (maximum a posteriori), 333 maximum-likelihood, 332 minimax (min-max), 335 minimum probablity of error, 334 Neyman-Pearson, 333

Decision theory, 331 De Morgan’s laws, 6, 18 Departure (or death) parameter, 351 Detection probability, 332 Differential entropy, 394 Dirac δ function, 275 Discrete, memoryless channel (DMC), 369 memoryless source (DMS), 369

r.v., 71, 116 Discrete-parameter Markov chain, 235 Disjoint sets, 3 Distribution:

Bernoulli, 55 binomial, 56 conditional,64, 92, 105, 128 continuous uniform, 61 discrete uniform, 60 exponential, 61 first-order, 208 gamma, 62 geometric, 57 limiting, 216 multinomial, 110 negative binomial, 58 normal (or gaussian), 63 nth-order, 208 Poisson, 59

second-order, 208 stationary, 216 uniform, continuous, 61 discrete, 60

Distribution function, 52, 66 cumulative (cdf), 52

Domain, 50, 101 Doob decomposition, 221

Efficient estimator, 313 Eigenvalue, 216 Eigenvector, 216 Ensemble, 207

average, 208 Entropy, 368

coding, 378, 405 conditional, 371 differential, 374 joint, 371 relative, 373, 375

Equally likely events, 9, 27 Equivocation, 372 Ergodic, in the mean, 307

process, 211 Erlang’s, delay (or C) formula, 360

loss (or B) formula, 364 Estimates, Bayes’, 314

point, 312 interval, 312 maximum likelyhood, 313

Estimation, 312 Bayes’, 314 error, 314

mean square, 314 maximum likelihood, 313, 319 mean square, 314, 325

linear, 315 parameter, 312 theory, 312

Estimator, Bayes’ 314 best, 314 biased, 316 consistent, 313 efficient, 313

most, 313 maximum-likelihood, 313 minimum mean square error, 314 minimum variance,313 point, 312, 316 unbiased, 312

Event, 2, 12, 51 certain, 2 elementary, 2 equally likely, 9, 27 impossible, 4 independent, 11, 41 mutually exclusive, 8, 9

and exhaustive, 10 space (or σ-field), 6

Expectation, 54, 152–153, 182 conditional, 107, 153, 219

properties of, 220 Expected value (see Mean) Experiment, Bernoulli, 44

random, 1 Exponential, distribution, 61

r.v., 61

Factorial moment, 155 False-alarm probability, 332 Filtration, 219, 222 Fn-measurable, 219 Fourier series, 279, 300

Perseval’s theorem for, 301 Fourier transform, 280, 304 Functions of r.v.’s, 149–152, 159, 167, 178 Fundamental matrix, 215

Gambler’s ruin, 241 Gamma, distribution, 62

function, 62, 79 r.v.,62, 180

Gaussian distribution (see Normal distribution) Geometric, distribution, 57

memoyless property, 58, 74 r.v.,57

Huffman, code, 380 Encoding, 379

Hypergeometric r.v., 97 Hypothesis, alternative, 331

composite, 331 null, 331 simple, 331

Hypothesis testing, 331, 335 level of significance, 332 power of, 332

Impulse response, 276 Independence law, 220

Independent (statistically), events, 11 increments, 210 process, 210 r.v.’s, 102, 105

Information, content, 368 measure of, 367 mutual, 371–372 rate, 368 source, 367 theory, 367

Initial-state probability vector, 213 Interarrival process, 216 Intersection of sets, 3 Interval estimate, 312

Jacobian, 151 Jenson’s inequlity, 153, 186, 220 Joint, characteristic function, 157

distribution function,102, 112 moment generating function, 156 probability density function (pdf), 104, 122 probability mass function (pmf), 103, 116 probability matrix, 370

Karhunen-Loéve expansion, 280, 300 Kraft inequality, 378, 404 Kullback-Leibler divergence, 373

Lagrange multiplier, 334, 384, 394, 396 Laplace r.v., 98 Law of large numbers, 158, 198 Level of significance, 332 Likelihood, function, 313

ratio, 333

test, 333 Limiting distribution, 216 Linearity, 153, 220 Linear mean-square estimation, 314, 327 Linear system, 276, 294

continuous-time, 276 impulse response of, 276 response to random inputs, 277, 294

discrete-time, 276 impulse (or unit sample) response, 277 response to random inputs, 278, 294

Little’s formula, 350 Log-normal r.v., 165

MAP (maximum a posteriori) test, 333 Marginal, distribution function, 103

cumulative distribution function (cdf), 103 probability density function (pdf), 104 probability mass function (pmf), 103

Markov, chains, 210 discrete-parameter, 211, 235 fundamental matrix, 215 homogeneous, 212 irreducible, 214 nonhomogeneous, 212 regular, 216

inequality, 86 matrix, 212 process, 210 property, 211

Maximum likelihood estimator, 313, 319 Mean, 54, 86, 208

conditional, 107

Mean square, continuity, 271 derivative, 272 error, 314

minimum, 314 estimation, 314, 325

linear, 315 integral, 272

Median, 97 Memoryless property (or Markov property), 58, 62, 74, 94, 211 Mercer’s theorem, 280 Measurability, 220 Measure of information, 380 Minimax (min-max) test, 335 Minimum probability of error test, 334 Minimum variance estimator, 313 Mixed r.v., 54 Mode, 97 Moment, 55, 155 Moment generating function, 155, 191

joint, 156 Most efficient estimator, 313 Multinomial, coefficient, 140

distribution, 110 r.v., 110 theorem, 140 trial, 110

Multiple r.v., 101 Mutual information, 367, 371, 387 Mutually exclusive, events, 3, 9

and exhaustive events, 10 sets, 3

Negative binomial r.v., 58

Neyman-Pearson test, 332, 340 Nonstationary process, 209 Normal, distribution, 63, 411

bivariate, 111 n-variate, 111

process, 211, 234 r.v., 63

standard, 64 Null, event (set), 3

hypothesis, 331 recurrent state, 214

Optimal stopping theorem, 221, 265 Orthogonal, processes, 273

r.v., 107 Orthogonality principle, 315 Outcomes, 1

Parameter estimation, 312 Parameter set, 208 Parseval’s theorem, 301 Periodic states, 215 Point estimators, 312 Point of occurrence, 216 Poisson, distribution, 59

process, 216, 249 r.v., 59 white noise, 293

Polya’s urn, 263 Positive recurrent states, 214 Positivity, 220 Posterior probability, 333 Power, function, 336

of test, 332 Power spectral density (or spectrum), 273, 288

cross, 274 Prior probability, 333 Probability, 1

conditional, 10, 31 continuity theorem of, 26 density function (pdf), 54 distribution, 213 generating function, 154, 187 initial state, 213 mass function (pmf), 53 measure, 7–8 space, 6–7, 21 total, 10, 38

Projection law, 220

Queueing, system, 349 M / M /1, 352 M / M /1/ K, 353 M/M/s, 352 M/M/s/K, 354 theory, 349

Random, experiment, 1 process, 207

complex, 208 Fourier transform of, 280 independent, 210 real, 208 realization of, 207 sample function of, 207

sample, 199, 312

sequence, 208 telegraph signal, 291

semi, 290 variable (r.v.), 50

continuous, 54 discrete, 53 function of, 149 mixed, 54 uncorrelated, 107

vector, 101, 108, 111, 137 walk, 222

simple, 222 Range, 50, 101 Rayleigh r.v., 78, 178 Real random process, 208 Recurrent process, 216 Recurrent states, 214

null, 214 positive, 214

Regression line, 315 Rejection region, 331 Relative frequency, 8 Renewal process, 216

Sample, function, 207 mean, 158, 199, 316 point, 1 random, 312 space, 1, 12 variance, 329 vector (see Random sample)

Sequence, decreasing, 25 increasing, 25

Sets, 1 algebra of, 2–6, 15 cardinality of, 4 countable, 2 difference of, 3

symmetrical, 3 disjoint, 3 intersection of, 3 mutually exclusive, 3 product of, 4 size of, 4 union of, 3

Shannon-Fano coding, 379 Shannon-Hartley law, 376 Sigma field (see event space) Signal-to-noise (S/N) ratio, 370 Simple, hypothesis, 331

random walk, 222 Source, alphabet, 367

coding, 376, 400 theorem, 377

encoder, 376 Stability, 220 Standard, deviation, 55

normal r.v., 63 State probability vector, 213 State space, 208 States, absorbing, 215

accessible, 214 aperiodic, 215 periodic, 215 recurrent, 214

null, 214

positive, 214 transient, 214

Stationary, distributions, 216 independent increments, 210 processes, 209

strict sense, 209 weak, 210 wide sense (WSS), 210

transition probability, 212 Statistic, 312

sufficient, 345 Statistical hypothesis, 331 Stochastic, continuity, 271

derivative, 272 integral, 272 matrix (see Markov matrix) periodicity, 279 process (see Random process)

Stopping time, 221 optimal, 221

System, linear, 276 linear time invariance (LTI), 276

response to random inputs, 276, 294 parallel, 43 series, 42

Threshold value, 333 Time-average, 211 Time autocorrelation function, 211 Total probability, 10, 38 Tower property, 220 Traffic intensity, 352 Transient states, 214

Transition probability, 212 matrix, 212 stationary, 212

Type I error, 332 Type II error, 332

Unbiased estimator, 312 Uncorrelated r.v.’s, 107 Uniform, distribution, 60–61

continuous, 60 r.v., 60

discontinuous, 61 r.v., 61

Union of sets, 3 Unit, impulse function (see Dirac δ function)

impulse sequence, 275 sample response, 277 sample sequence, 275 step function, 171

Universal set, 1

Variance, 55 conditional, 107

Vector mean, 111 Venn diagram, 4

Waiting time, 253 White noise, 275, 292

normal (or gaussian), 293 Poisson, 293

Wiener-Khinchin relations, 274 Wiener process, 218, 256, 303

standard, 218 with drift coefficient, 219

Z-transform, 154

  • Title Page
  • Copyright Page
  • Preface to The Second Edition
  • Preface to The First Edition
  • Contents
  • CHAPTER 1 Probability
    • 1.1 Introduction
    • 1.2 Sample Space and Events
    • 1.3 Algebra of Sets
    • 1.4 Probability Space
    • 1.5 Equally Likely Events
    • 1.6 Conditional Probability
    • 1.7 Total Probability
    • 1.8 Independent Events
    • Solved Problems
  • CHAPTER 2 Random Variables
    • 2.1 Introduction
    • 2.2 Random Variables
    • 2.3 Distribution Functions
    • 2.4 Discrete Random Variables and Probability Mass Functions
    • 2.5 Continuous Random Variables and Probability Density Functions
    • 2.6 Mean and Variance
    • 2.7 Some Special Distributions
    • 2.8 Conditional Distributions
    • Solved Problems
  • CHAPTER 3 Multiple Random Variables
    • 3.1 Introduction
    • 3.2 Bivariate Random Variables
    • 3.3 Joint Distribution Functions
    • 3.4 Discrete Random Variables—Joint Probability Mass Functions
    • 3.5 Continuous Random Variables—Joint Probability Density Functions
    • 3.6 Conditional Distributions
    • 3.7 Covariance and Correlation Coefficient
    • 3.8 Conditional Means and Conditional Variances
    • 3.9 N-Variate Random Variables
    • 3.10 Special Distributions
    • Solved Problems
  • CHAPTER 4 Functions of Random Variables, Expectation, Limit Theorems
    • 4.1 Introduction
    • 4.2 Functions of One Random Variable
    • 4.3 Functions of Two Random Variables
    • 4.4 Functions of n Random Variables
    • 4.5 Expectation
    • 4.6 Probability Generating Functions
    • 4.7 Moment Generating Functions
    • 4.8 Characteristic Functions
    • 4.9 The Laws of Large Numbers and the Central Limit Theorem
    • Solved Problems
  • CHAPTER 5 Random Processes
    • 5.1 Introduction
    • 5.2 Random Processes
    • 5.3 Characterization of Random Processes
    • 5.4 Classification of Random Processes
    • 5.5 Discrete-Parameter Markov Chains
    • 5.6 Poisson Processes
    • 5.7 Wiener Processes
    • 5.8 Martingales
    • Solved Problems
  • CHAPTER 6 Analysis and Processing of Random Processes
    • 6.1 Introduction
    • 6.2 Continuity, Differentiation, Integration
    • 6.3 Power Spectral Densities
    • 6.4 White Noise
    • 6.5 Response of Linear Systems to Random Inputs
    • 6.6 Fourier Series and Karhunen-Loéve Expansions
    • 6.7 Fourier Transform of Random Processes
    • Solved Problems
  • CHAPTER 7 Estimation Theory
    • 7.1 Introduction
    • 7.2 Parameter Estimation
    • 7.3 Properties of Point Estimators
    • 7.4 Maximum-Likelihood Estimation
    • 7.5 Bayes’ Estimation
    • 7.6 Mean Square Estimation
    • 7.7 Linear Mean Square Estimation
    • Solved Problems
  • CHAPTER 8 Decision Theory
    • 8.1 Introduction
    • 8.2 Hypothesis Testing
    • 8.3 Decision Tests
    • Solved Problems
  • CHAPTER 9 Queueing Theory
    • 9.1 Introduction
    • 9.2 Queueing Systems
    • 9.3 Birth-Death Process
    • 9.4 The M/M/1 Queueing System
    • 9.5 The M/M/s Queueing System
    • 9.6 The M/M/1/K Queueing System
    • 9.7 The M/M/s/K Queueing System
    • Solved Problems
  • CHAPTER 10 Information Theory
    • 10.1 Introduction
    • 10.2 Measure of Information
    • 10.3 Discrete Memoryless Channels
    • 10.4 Mutual Information
    • 10.5 Channel Capacity
    • 10.6 Continuous Channel
    • 10.7 Additive White Gaussian Noise Channel
    • 10.8 Source Coding
    • 10.9 Entropy Coding
    • Solved Problems
  • APPENDIX A Normal Distribution
  • APPENDIX B Fourier Transform
    • B.1 Continuous-Time Fourier Transform
    • B.2 Discrete-Time Fourier Transform
  • INDEX