Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Math / Probability and Statistics

Intro to Probability, Grinstead and Snell (2012)

PDF · 510 pages · 2.2 MB
Open PDF file

A probability textbook running through twelve chapters. The opening section covers discrete probability distributions, random variables, and simulating dice and coin experiments with random number generators. Examples include the de Méré dice problem from the Pascal-Fermat correspondence and the heads-or-tails game, and the text refers ahead to Bernoulli trials and the Law of Large Numbers. It appears to be a published text rather than Phil's own writing.

AI-written summary; may contain errors. This description is approximate.

Extracted text (machine-read; may contain errors)
Chapter 1 Discrete Probability Distributions 1.1 Simulation of Discrete Probabilities Probability In this chapter, we shall flrst consider chance experiments with a flnite number of possible outcomes !1,!2, ...,!n. For example, we roll a die and the possible outcomes are 1, 2, 3, 4, 5, 6 corresponding to the side that turns up. We toss a coinwith possible outcomes H (heads) and T (tails). It is frequently useful to be able to refer to an outcome of an experiment. For example, we might want to write the mathematical expression which gives the sumof four rolls of a die. To do this, we could let X i,i=1;2;3;4;represent the values of the outcomes of the four rolls, and then we could write the expression X1+X2+X3+X4 for the sum of the four rolls. The Xi’s are called random variables . A random vari- able is simply an expression whose value is the outcome of a particular experiment.Just as in the case of other types of variables in mathematics, random variables cantake on difierent values. LetXbe the random variable which represents the roll of one die. We shall assign probabilities to the possible outcomes of this experiment. We do this byassigning to each outcome ! ja nonnegative number m(!j) in such a way that m(!1)+m(!2)+¢¢¢+m(!6)=1: The function m(!j) is called the distribution function of the random variable X. For the case of the roll of the die we would assign equal probabilities or probabilities1/6 to each of the outcomes. With this assignment of probabilities, one could write P(X•4) =2 3 1 2 CHAPTER 1. DISCRETE PROBABILITY DISTRIBUTIONS to mean that the probability is 2 =3 that a roll of a die will have a value which does not exceed 4. LetYbe the random variable which represents the toss of a coin. In this case, there are two possible outcomes, which we can label as H and T. Unless we havereason to suspect that the coin comes up one way more often than the other way,it is natural to assign the probability of 1/2 to each of the two outcomes. In both of the above experiments, each outcome is assigned an equal probability. This would certainly not be the case in general. For example, if a drug is found tobe efiective 30 percent of the time it is used, we might assign a probability .3 thatthe drug is efiective the next time it is used and .7 that it is not efiective. This lastexample illustrates the intuitive frequency concept of probability. That is, if we have a probability pthat an experiment will result in outcome A, then if we repeat this experiment a large number of times we should expect that the fraction of times thatAwill occur is about p. To check intuitive ideas like this, we shall flnd it helpful to look at some of these problems experimentally. We could, for example, toss a coina large number of times and see if the fraction of times heads turns up is about 1/2.We could also simulate this experiment on a computer. Simulation We want to be able to perform an experiment that corresponds to a given set ofprobabilities; for example, m(! 1)=1=2,m(!2)=1=3, andm(!3)=1=6. In this case, one could mark three faces of a six-sided die with an !1, two faces with an !2, and one face with an !3. In the general case we assume that m(!1),m(!2) , ...,m(!n) are all rational numbers, with least common denominator n.I fn> 2, we can imagine a long cylindrical die with a cross-section that is a regular n-gon. Ifm(!j)=nj=n, then we can label njof the long faces of the cylinder with an !j, and if one of the end faces comes up, we can just roll the die again. If n= 2, a coin could be used to perform the experiment. We will be particularly interested in repeating a chance experiment a large num- ber of times. Although the cylindrical die would be a convenient way to carry outa few repetitions, it would be di–cult to carry out a large number of experiments.Since the modern computer can do a large number of operations in a very shorttime, it is natural to turn to the computer for this task. Random Numbers We must flrst flnd a computer analog of rolling a die. This is done on the computerby means of a random number generator. Depending upon the particular software package, the computer can be asked for a real number between 0 and 1, or an integerin a given set of consecutive integers. In the flrst case, the real numbers are chosenin such a way that the probability that the number lies in any particular subintervalof this unit interval is equal to the length of the subinterval. In the second case,each integer has the same probability of being chosen. 1.1. SIMULATION OF DISCRETE PROBABILITIES 3 .203309 .762057 .151121 .623868 .932052 .415178 .716719 .967412.069664 .670982 .352320 .049723.750216 .784810 .089734 .966730.946708 .380365 .027381 .900794 Table 1.1: Sample output of the program RandomNumbers . LetXbe a random variable with distribution function m(!), where!is in the setf! 1;!2;!3g, andm(!1)=1=2,m(!2)=1=3, andm(!3)=1=6. If our computer package can return a random integer in the set f1;2;:::;6g, then we simply ask it to do so, and make 1, 2, and 3 correspond to !1, 4 and 5 correspond to !2, and 6 correspond to !3. If our computer package returns a random real number rin the interval (0;1), then the expression b6rc+1 will be a random integer between 1 and 6. (The notation bxcmeans the greatest integer not exceeding x, and is read \°oor of x.") The method by which random real numbers are generated on a computer is described in the historical discussion at the end of this section. The followingexample gives sample output of the program RandomNumbers . Example 1.1 (Random Number Generation) The program RandomNumbers generatesnrandom real numbers in the interval [0 ;1], wherenis chosen by the user. When we ran the program with n= 20, we obtained the data shown in Table 1.1. 2 Example 1.2 (Coin Tossing) As we have noted, our intuition suggests that the probability of obtaining a head on a single toss of a coin is 1/2. To have thecomputer toss a coin, we can ask it to pick a random real number in the interval[0;1] and test to see if this number is less than 1/2. If so, we shall call the outcome heads ; if not we call it tails. Another way to proceed would be to ask the computer to pick a random integer from the set f0;1g. The program CoinTosses carries out the experiment of tossing a coin ntimes. Running this program, with n= 20, resulted in: THTTTHTTTTHTTTTTHHTT. Note that in 20 tosses, we obtained 5 heads and 15 tails. Let us toss a coin n times, where nis much larger than 20, and see if we obtain a proportion of heads closer to our intuitive guess of 1/2. The program CoinTosses keeps track of the number of heads. When we ran this program with n= 1000, we obtained 494 heads. When we ran it with n= 10000, we obtained 5039 heads. 4 CHAPTER 1. DISCRETE PROBABILITY DISTRIBUTIONS We notice that when we tossed the coin 10,000 times, the proportion of heads was close to the \true value" .5 for obtaining a head when a coin is tossed. A math-ematical model for this experiment is called Bernoulli Trials (see Chapter 3). TheLaw of Large Numbers, which we shall study later (see Chapter 8), will show that in the Bernoulli Trials model, the proportion of heads should be near .5, consistentwith our intuitive idea of the frequency interpretation of probability. Of course, our program could be easily modifled to simulate coins for which the probability of a head is p, wherepis a real number between 0 and 1. 2 In the case of coin tossing, we already knew the probability of the event occurring on each experiment. The real power of simulation comes from the ability to estimateprobabilities when they are not known ahead of time. This method has been used inthe recent discoveries of strategies that make the casino game of blackjack favorableto the player. We illustrate this idea in a simple situation in which we can computethe true probability and see how efiective the simulation is. Example 1.3 (Dice Rolling) We consider a dice game that played an important role in the historical development of probability. The famous letters between Pas-cal and Fermat, which many believe started a serious study of probability, wereinstigated by a request for help from a French nobleman and gambler, Chevalierde M¶ er¶e. It is said that de M¶ er¶e had been betting that, in four rolls of a die, at least one six would turn up. He was winning consistently and, to get more peopleto play, he changed the game to bet that, in 24 rolls of two dice, a pair of sixeswould turn up. It is claimed that de M¶ er¶e lost with 24 and felt that 25 rolls were necessary to make the game favorable. It was un grand scandale that mathematics was wrong. We shall try to see if de M¶ er¶e is correct by simulating his various bets. The program DeMere1 simulates a large number of experiments, seeing, in each one, if a six turns up in four rolls of a die. When we ran this program for 1000 plays,a six came up in the flrst four rolls 48.6 percent of the time. When we ran it for10,000 plays this happened 51.98 percent of the time. We note that the result of the second run suggests that de M¶ er¶e was correct in believing that his bet with one die was favorable; however, if we had based ourconclusion on the flrst run, we would have decided that he was wrong. Accurate results by simulation require a large number of experiments. 2 The program DeMere2 simulates de M¶ er¶e’s second bet that a pair of sixes will occur in nrolls of a pair of dice. The previous simulation shows that it is important to know how many trials we should simulate in order to expect a certaindegree of accuracy in our approximation. We shall see later that in these types ofexperiments, a rough rule of thumb is that, at least 95% of the time, the error doesnot exceed the reciprocal of the square root of the number of trials. Fortunately,for this dice game, it will be easy to compute the exact probabilities. We shallshow in the next section that for the flrst bet the probability that de M¶ er¶e wins is 1¡(5=6) 4=:518. 1.1. SIMULATION OF DISCRETE PROBABILITIES 5 510 15 20 25 30 35 40 -10-8-6-4-2246810 Figure 1.1: Peter’s winnings in 40 plays of heads or tails. One can understand this calculation as follows: The probability that no 6 turns up on the flrst toss is (5 =6). The probability that no 6 turns up on either of the flrst two tosses is (5 =6)2. Reasoning in the same way, the probability that no 6 turns up on any of the flrst four tosses is (5 =6)4. Thus, the probability of at least one 6 in the flrst four tosses is 1 ¡(5=6)4. Similarly, for the second bet, with 24 rolls, the probability that de M¶ er¶e wins is 1¡(35=36)24=:491, and for 25 rolls it is 1¡(35=36)25=:506. Using the rule of thumb mentioned above, it would require 27,000 rolls to have a reasonable chance to determine these probabilities with su–cient accuracy to assertthat they lie on opposite sides of .5. It is interesting to ponder whether a gamblercan detect such probabilities with the required accuracy from gambling experience.Some writers on the history of probability suggest that de M¶ er¶e was, in fact, just interested in these problems as intriguing probability problems. Example 1.4 (Heads or Tails) For our next example, we consider a problem where the exact answer is di–cult to obtain but for which simulation easily gives thequalitative results. Peter and Paul play a game called heads or tails. In this game, a fair coin is tossed a sequence of times|we choose 40. Each time a head comes upPeter wins 1 penny from Paul, and each time a tail comes up Peter loses 1 pennyto Paul. For example, if the results of the 40 tosses are THTHHHHTTHTHHTTHHTTTTHHHTHHTHHHTHHHTTTHH. Peter’s winnings may be graphed as in Figure 1.1. Peter has won 6 pennies in this particular game. It is natural to ask for the probability that he will win jpennies; here jcould be any even number from ¡40 to 40. It is reasonable to guess that the value of jwith the highest probability isj= 0, since this occurs when the number of heads equals the number of tails. Similarly, we would guess that the values of jwith the lowest probabilities are j=§40. 6 CHAPTER 1. DISCRETE PROBABILITY DISTRIBUTIONS A second interesting question about this game is the following: How many times in the 40 tosses will Peter be in the lead? Looking at the graph of his winnings(Figure 1.1), we see that Peter is in the lead when his winnings are positive, butwe have to make some convention when his winnings are 0 if we want all tosses tocontribute to the number of times in the lead. We adopt the convention that, whenPeter’s winnings are 0, he is in the lead if he was ahead at the previous toss andnot if he was behind at the previous toss. With this convention, Peter is in the lead34 times in our example. Again, our intuition might suggest that the most likelynumber of times to be in the lead is 1/2 of 40, or 20, and the least likely numbersare the extreme cases of 40 or 0. It is easy to settle this by simulating the game a large number of times and keeping track of the number of times that Peter’s flnal winnings are j, and the number of times that Peter ends up being in the lead by k. The proportions over all games then give estimates for the corresponding probabilities. The programHTSimulation carries out this simulation. Note that when there are an even number of tosses in the game, it is possible to be in the lead only an even numberof times. We have simulated this game 10,000 times. The results are shown inFigures 1.2 and 1.3. These graphs, which we call spike graphs, were generatedusing the program Spikegraph . The vertical line, or spike, at position xon the horizontal axis, has a height equal to the proportion of outcomes which equal x. Our intuition about Peter’s flnal winnings was quite correct, but our intuition aboutthe number of times Peter was in the lead was completely wrong. The simulationsuggests that the least likely number of times in the lead is 20 and the most likelyis 0 or 40. This is indeed correct, and the explanation for it is suggested by playingthe game of heads or tails with a large number of tosses and looking at a graph ofPeter’s winnings. In Figure 1.4 we show the results of a simulation of the game, for1000 tosses and in Figure 1.5 for 10,000 tosses. In the second example Peter was ahead most of the time. It is a remarkable fact, however, that, if play is continued long enough, Peter’s winnings will continueto come back to 0, but there will be very long times between the times that thishappens. These and related results will be discussed in Chapter 12. 2 In all of our examples so far, we have simulated equiprobable outcomes. We illustrate next an example where the outcomes are not equiprobable. Example 1.5 (Horse Races) Four horses (Acorn, Balky, Chestnut, and Dolby) have raced many times. It is estimated that Acorn wins 30 percent of the time,Balky 40 percent of the time, Chestnut 20 percent of the time, and Dolby 10 percentof the time. We can have our computer carry out one race as follows: Choose a random numberx.I fx<: 3 then we say that Acorn won. If :3•x<: 7 then Balky wins. If:7•x<: 9 then Chestnut wins. Finally, if :9•xthen Dolby wins. The program HorseRace uses this method to simulate the outcomes of nraces. Running this program for n= 10 we found that Acorn won 40 percent of the time, Balky 20 percent of the time, Chestnut 10 percent of the time, and Dolby 30 percent 1.1. SIMULATION OF DISCRETE PROBABILITIES 7 Figure 1.2: Distribution of winnings. Figure 1.3: Distribution of number of times in the lead. 8 CHAPTER 1. DISCRETE PROBABILITY DISTRIBUTIONS 200 400 600 800 10001000 plays -50-40-30-20-1001020 Figure 1.4: Peter’s winnings in 1000 plays of heads or tails. 2000 4000 6000 8000 1000010000 plays 050100150200 Figure 1.5: Peter’s winnings in 10,000 plays of heads or tails. 1.1. SIMULATION OF DISCRETE PROBABILITIES 9 of the time. A larger number of races would be necessary to have better agreement with the past experience. Therefore we ran the program to simulate 1000 raceswith our four horses. Although very tired after all these races, they performed ina manner quite consistent with our estimates of their abilities. Acorn won 29.8percent of the time, Balky 39.4 percent, Chestnut 19.5 percent, and Dolby 11.3percent of the time. The program GeneralSimulation uses this method to simulate repetitions of an arbitrary experiment with a flnite number of outcomes occurring with knownprobabilities. 2 Historical Remarks Anyone who plays the same chance game over and over is really carrying out a sim- ulation, and in this sense the process of simulation has been going on for centuries.As we have remarked, many of the early problems of probability might well havebeen suggested by gamblers’ experiences. It is natural for anyone trying to understand probability theory to try simple experiments by tossing coins, rolling dice, and so forth. The naturalist Bufion tosseda coin 4040 times, resulting in 2048 heads and 1992 tails. He also estimated thenumber…by throwing needles on a ruled surface and recording how many times the needles crossed a line (see Section 2.1). The English biologist W. F. R. Weldon 1 recorded 26,306 throws of 12 dice, and the Swiss scientist Rudolf Wolf2recorded 100,000 throws of a single die without a computer. Such experiments are very time-consuming and may not accurately represent the chance phenomena being studied.For example, for the dice experiments of Weldon and Wolf, further analysis of therecorded data showed a suspected bias in the dice. The statistician Karl Pearsonanalyzed a large number of outcomes at certain roulette tables and suggested thatthe wheels were biased. He wrote in 1894: Clearly, since the Casino does not serve the valuable end of huge lab- oratory for the preparation of probability statistics, it has no scientiflcraison d’^ etre. Men of science cannot have their most reflned theories disregarded in this shameless manner! The French Government must beurged by the hierarchy of science to close the gaming-saloons; it wouldbe, of course, a graceful act to hand over the remaining resources of theCasino to the Acad¶ emie des Sciences for the endowment of a laboratory of orthodox probability; in particular, of the new branch of that study,the application of the theory of chance to the biological problems ofevolution, which is likely to occupy so much of men’s thoughts in thenear future. 3 However, these early experiments were suggestive and led to important discov- eries in probability and statistics. They led Pearson to the chi-squared test, which 1T. C. Fry, Probability and Its Engineering Uses, 2nd ed. (Princeton: Van Nostrand, 1965). 2E. Czuber, Wahrscheinlichkeitsrechnung, 3rd ed. (Berlin: Teubner, 1914). 3K. Pearson, \Science and Monte Carlo," Fortnightly Review , vol. 55 (1894), p. 193; cited in S. M. Stigler, The History of Statistics (Cambridge: Harvard University Press, 1986). 10 CHAPTER 1. DISCRETE PROBABILITY DISTRIBUTIONS is of great importance in testing whether observed data flt a given probability dis- tribution. By the early 1900s it was clear that a better way to generate random numbers was needed. In 1927, L. H. C. Tippett published a list of 41,600 digits obtained byselecting numbers haphazardly from census reports. In 1955, RAND Corporationprinted a table of 1,000,000 random numbers generated from electronic noise. Theadvent of the high-speed computer raised the possibility of generating random num-bers directly on the computer, and in the late 1940s John von Neumann suggestedthat this be done as follows: Suppose that you want a random sequence of four-digitnumbers. Choose any four-digit number, say 6235, to start. Square this numberto obtain 38,875,225. For the second number choose the middle four digits of thissquare (i.e., 8752). Do the same process starting with 8752 to get the third number,and so forth. More modern methods involve the concept of modular arithmetic. If ais an integer and mis a positive integer, then by a(modm) we mean the remainder whenais divided by m. For example, 10 (mod 4) = 2, 8 (mod 2) = 0, and so forth. To generate a random sequence X 0;X1;X2;:::of numbers choose a starting numberX0and then obtain the numbers Xn+1fromXnby the formula Xn+1=(aXn+c) (modm); wherea,c, andmare carefully chosen constants. The sequence X0;X1;X2;::: is then a sequence of integers between 0 and m¡1. To obtain a sequence of real numbers in [0 ;1), we divide each Xjbym. The resulting sequence consists of rational numbers of the form j=m, where 0•j•m¡1. Sincemis usually a very large integer, we think of the numbers in the sequence as being random realnumbers in [0 ;1). For both von Neumann’s squaring method and the modular arithmetic technique the sequence of numbers is actually completely determined by the flrst number.Thus, there is nothing really random about these sequences. However, they producenumbers that behave very much as theory would predict for random experiments.To obtain difierent sequences for difierent experiments the initial number X 0is chosen by some other procedure that might involve, for example, the time of day.4 During the Second World War, physicists at the Los Alamos Scientiflc Labo- ratory needed to know, for purposes of shielding, how far neutrons travel throughvarious materials. This question was beyond the reach of theoretical calculations.Daniel McCracken, writing in the Scientiflc American , states: The physicists had most of the necessary data: they knew the average distance a neutron of a given speed would travel in a given substancebefore it collided with an atomic nucleus, what the probabilities werethat the neutron would bounce ofi instead of being absorbed by thenucleus, how much energy the neutron was likely to lose after a given 4For a detailed discussion of random numbers, see D. E. Knuth, The Art of Computer Pro- gramming, vol. II (Reading: Addison-Wesley, 1969). 1.1. SIMULATION OF DISCRETE PROBABILITIES 11 collision and so on.5 John von Neumann and Stanislas Ulam suggested that the problem be solved by modeling the experiment by chance devices on a computer. Their work beingsecret, it was necessary to give it a code name. Von Neumann chose the name\Monte Carlo." Since that time, this method of simulation has been called theMonte Carlo Method. William Feller indicated the possibilities of using computer simulations to illus- trate basic concepts in probability in his book An Introduction to Probability Theory and Its Applications. In discussing the problem about the number of times in the lead in the game of \heads or tails" Feller writes: The results concerning °uctuations in coin tossing show that widely held beliefs about the law of large numbers are fallacious. These resultsare so amazing and so at variance with common intuition that evensophisticated colleagues doubted that coins actually misbehave as theorypredicts. The record of a simulated experiment is therefore included. 6 Feller provides a plot showing the result of 10,000 plays of heads or tails similar to that in Figure 1.5. The martingale betting system described in Exercise 10 has a long and interest- ing history. Russell Barnhart pointed out to the authors that its use can be tracedback at least to 1754, when Casanova, writing in his memoirs, History of My Life, writes She [Casanova’s mistress] made me promise to go to the casino [the Ridotto in Venice] for money to play in partnership with her. I wentthere and took all the gold I found, and, determinedly doubling mystakes according to the system known as the martingale, I won three orfour times a day during the rest of the Carnival. I never lost the sixthcard. If I had lost it, I should have been out of funds, which amountedto two thousand zecchini. 7 Even if there were no zeros on the roulette wheel so the game was perfectly fair, the martingale system, or any other system for that matter, cannot make the gameinto a favorable game. The idea that a fair game remains fair and unfair gamesremain unfair under gambling systems has been exploited by mathematicians toobtain important results in the study of probability. We will introduce the generalconcept of a martingale in Chapter 6. The word martingale itself also has an interesting history. The origin of the word is obscure. The Oxford English Dictionary gives examples of its use in the 5D. D. McCracken, \The Monte Carlo Method," Scientiflc American, vol. 192 (May 1955), p. 90. 6W. Feller, Introduction to Probability Theory and its Applications, vol. 1, 3rd ed. (New York: John Wiley & Sons, 1968), p. xi. 7G. Casanova, History of My Life, vol. IV, Chap. 7, trans. W. R. Trask (New York: Harcourt- Brace, 1968), p. 124. 12 CHAPTER 1. DISCRETE PROBABILITY DISTRIBUTIONS early 1600s and says that its probable origin is the reference in Rabelais’s Book One, Chapter 19: Everything was done as planned, the only thing being that Gargantua doubted if they would be able to flnd, right away, breeches suitable to the old fellow’s legs; he was doubtful, also, as to what cut would be mostbecoming to the orator|the martingale, which has a draw-bridge efiectin the seat, to permit doing one’s business more easily; the sailor-style,which afiords more comfort for the kidneys; the Swiss, which is warmeron the belly; or the codflsh-tail, which is cooler on the loins. 8 In modern uses martingale has several difierent meanings, all related to holding down, in addition to the gambling use. For example, it is a strap on a horse’s harness used to hold down the horse’s head, and also part of a sailing rig used tohold down the bowsprit. The Labouchere system described in Exercise 9 is named after Henry du Pre Labouchere (1831{1912), an English journalist and member of Parliament. Labou-chere attributed the system to Condorcet. Condorcet (1743{1794) was a politicalleader during the time of the French revolution who was interested in applying prob-ability theory to economics and politics. For example, he calculated the probabilitythat a jury using majority vote will give a correct decision if each juror has thesame probability of deciding correctly. His writings provided a wealth of ideas onhow probability might be applied to human afiairs. 9 Exercises 1Modify the program CoinTosses to toss a coin ntimes and print out after every 100 tosses the proportion of heads minus 1/2. Do these numbers appearto approach 0 as nincreases? Modify the program again to print out, every 100 times, both of the following quantities: the proportion of heads minus 1/2,and the number of heads minus half the number of tosses. Do these numbersappear to approach 0 as nincreases? 2Modify the program CoinTosses so that it tosses a coin ntimes and records whether or not the proportion of heads is within .1 of .5 (i.e., between .4and .6). Have your program repeat this experiment 100 times. About howlarge must nbe so that approximately 95 out of 100 times the proportion of heads is between .4 and .6? 3In the early 1600s, Galileo was asked to explain the fact that, although the number of triples of integers from 1 to 6 with sum 9 is the same as the numberof such triples with sum 10, when three dice are rolled, a 9 seemed to comeup less often than a 10|supposedly in the experience of gamblers. 8Quoted in the Portable Rabelais, ed. S. Putnam (New York: Viking, 1946), p. 113. 9Le Marquise de Condorcet, Essai sur l’Application de l’Analyse µ a la Probabilit¶ ed µes D¶ ecisions Rendues a la Pluralit¶ e des Voix (Paris: Imprimerie Royale, 1785). 1.1. SIMULATION OF DISCRETE PROBABILITIES 13 (a) Write a program to simulate the roll of three dice a large number of times and keep track of the proportion of times that the sum is 9 andthe proportion of times it is 10. (b) Can you conclude from your simulations that the gamblers were correct? 4In raquetball, a player continues to serve as long as she is winning; a point is scored only when a player is serving and wins the volley. The flrst playerto win 21 points wins the game. Assume that you serve flrst and have aprobability .6 of winning a volley when you serve and probability .5 whenyour opponent serves. Estimate, by simulation, the probability that you willwin a game. 5Consider the bet that all three dice will turn up sixes at least once in nrolls of three dice. Calculate f(n), the probability of at least one triple-six when three dice are rolled ntimes. Determine the smallest value of nnecessary for a favorable bet that a triple-six will occur when three dice are rolled ntimes. (DeMoivre would say it should be about 216 log 2 = 149 :7 and so would answer 150|see Exercise 1.2.17. Do you agree with him?) 6In Las Vegas, a roulette wheel has 38 slots numbered 0, 00, 1, 2, ..., 3 6 . T h e 0 and 00 slots are green and half of the remaining 36 slots are red and halfare black. A croupier spins the wheel and throws in an ivory ball. If you bet1 dollar on red, you win 1 dollar if the ball stops in a red slot and otherwiseyou lose 1 dollar. Write a program to flnd the total winnings for a player whomakes 1000 bets on red. 7Another form of bet for roulette is to bet that a speciflc number (say 17) will turn up. If the ball stops on your number, you get your dollar back plus 35dollars. If not, you lose your dollar. Write a program that will plot yourwinnings when you make 500 plays of roulette at Las Vegas, flrst when youbet each time on red (see Exercise 6), and then for a second visit to LasVegas when you make 500 plays betting each time on the number 17. Whatdifierences do you see in the graphs of your winnings on these two occasions? 8An astute student noticed that, in our simulation of the game of heads or tails (see Example 1.4), the proportion of times the player is always in the lead isvery close to the proportion of times that the player’s total winnings end up 0.Work out these probabilities by enumeration of all cases for two tosses andfor four tosses, and see if you think that these probabilities are, in fact, thesame. 9The Labouchere system for roulette is played as follows. Write down a list of numbers, usually 1, 2, 3, 4. Bet the sum of the flrst and last, 1 + 4 = 5, onred. If you win, delete the flrst and last numbers from your list. If you lose,add the amount that you last bet to the end of your list. Then use the newlist and bet the sum of the flrst and last numbers (if there is only one number,bet that amount). Continue until your list becomes empty. Show that, if this 14 CHAPTER 1. DISCRETE PROBABILITY DISTRIBUTIONS happens, you win the sum, 1+2+3+4=1 0 ,o fy o u r original list. Simulate this system and see if you do always stop and, hence, always win. If so, whyis this not a foolproof gambling system? 10Another well-known gambling system is the martingale doubling system . Sup- pose that you are betting on red to turn up in roulette. Every time you win,bet 1 dollar next time. Every time you lose, double your previous bet. Con-tinue to play until you have won at least 5 dollars or you have lost more than100 dollars. Write a program to simulate this system and play it a numberof times and see how you do. In his book The Newcomes, W. M. Thack- eray remarks \You have not played as yet? Do not do so; above all avoid amartingale if you do." 10Was this good advice? 11Modify the program HTSimulation so that it keeps track of the maximum of Peter’s winnings in each game of 40 tosses. Have your program print out theproportion of times that your total winnings take on values 0 ;2;4; :::; 40. Calculate the corresponding exact probabilities for games of two tosses andfour tosses. 12In an upcoming national election for the President of the United States, a pollster plans to predict the winner of the popular vote by taking a randomsample of 1000 voters and declaring that the winner will be the one obtainingthe most votes in his sample. Suppose that 48 percent of the voters planto vote for the Republican candidate and 52 percent plan to vote for theDemocratic candidate. To get some idea of how reasonable the pollster’splan is, write a program to make this prediction by simulation. Repeat thesimulation 100 times and see how many times the pollster’s prediction wouldcome true. Repeat your experiment, assuming now that 49 percent of thepopulation plan to vote for the Republican candidate; flrst with a sample of1000 and then with a sample of 3000. (The Gallup Poll uses about 3000.)(This idea is discussed further in Chapter 9, Section 9.1.) 13The psychologist Tversky and his colleagues 11say that about four out of flve people will answer (a) to the following question: A certain town is served by two hospitals. In the larger hospital about 45 babies are born each day, and in the smaller hospital 15 babies are born eachday. Although the overall proportion of boys is about 50 percent, the actualproportion at either hospital may be more or less than 50 percent on any day.At the end of a year, which hospital will have the greater number of days onwhich more than 60 percent of the babies born were boys? (a) the large hospital 10W. M. Thackerey, The Newcomes (London: Bradbury and Evans, 1854{55). 11See K. McKean, \Decisions, Decisions," Discover, June 1985, pp. 22{31. Kevin McKean, Discover Magazine, c°1987 Family Media, Inc. Reprinted with permission. This popular article reports on the work of Tverksy et. al. in Judgement Under Uncertainty: Heuristics and Biases (Cambridge: Cambridge University Press, 1982). 1.1. SIMULATION OF DISCRETE PROBABILITIES 15 (b) the small hospital (c) neither|the number of days will be about the same. Assume that the probability that a baby is a boy is .5 (actual estimates make this more like .513). Decide, by simulation, what the right answer is to thequestion. Can you suggest why so many people go wrong? 14You are ofiered the following game. A fair coin will be tossed until the flrst time it comes up heads. If this occurs on the jth toss you are paid 2 jdollars. You are sure to win at least 2 dollars so you should be willing to pay to playthis game|but how much? Few people would pay as much as 10 dollars toplay this game. See if you can decide, by simulation, a reasonable amountthat you would be willing to pay, per game, if you will be allowed to makea large number of plays of the game. Does the amount that you would bewilling to pay per game depend upon the number of plays that you will beallowed? 15Tversky and his colleagues 12studied the records of 48 of the Philadelphia 76ers basketball games in the 1980{81 season to see if a player had timeswhen he was hot and every shot went in, and other times when he was coldand barely able to hit the backboard. The players estimated that they wereabout 25 percent more likely to make a shot after a hit than after a miss.In fact, the opposite was true|the 76ers were 6 percent more likely to scoreafter a miss than after a hit. Tversky reports that the number of hot and coldstreaks was about what one would expect by purely random efiects. Assumingthat a player has a flfty-flfty chance of making a shot and makes 20 shots agame, estimate by simulation the proportion of the games in which the playerwill have a streak of 5 or more hits. 16Estimate, by simulation, the average number of children there would be in a family if all people had children until they had a boy. Do the same if allpeople had children until they had at least one boy and at least one girl. Howmany more children would you expect to flnd under the second scheme thanunder the flrst in 100,000 families? (Assume that boys and girls are equallylikely.) 17Mathematicians have been known to get some of the best ideas while sitting in a cafe, riding on a bus, or strolling in the park. In the early 1900s the famousmathematician George P¶ olya lived in a hotel near the woods in Zurich. He liked to walk in the woods and think about mathematics. P¶ olya describes the following incident: At the hotel there lived also some students with whom I usually took my meals and had friendly relations. On a certain day oneof them expected the visit of his flanc¶ ee, what (sic) I knew, but I did not foresee that he and his flanc¶ ee would also set out for a 12ibid. 16 CHAPTER 1. DISCRETE PROBABILITY DISTRIBUTIONS 01 23 -1-2-3 c. Random walk in three dimensions. b. Random walk in two dimensions.a. Random walk in one dimension. Figure 1.6: Random walk. 1.1. SIMULATION OF DISCRETE PROBABILITIES 17 stroll in the woods, and then suddenly I met them there. And then I met them the same morning repeatedly, I don’t remember howmany times, but certainly much too often and I felt embarrassed:It looked as if I was snooping around which was, I assure you, notthe case. 13 This set him to thinking about whether random walkers were destined tomeet. P¶olya considered random walkers in one, two, and three dimensions. In one dimension, he envisioned the walker on a very long street. At each intersec-tion the walker °ips a fair coin to decide which direction to walk next (seeFigure 1.6a). In two dimensions, the walker is walking on a grid of streets, andat each intersection he chooses one of the four possible directions with equalprobability (see Figure 1.6b). In three dimensions (we might better speak ofa random climber), the walker moves on a three-dimensional grid, and at eachintersection there are now six difierent directions that the walker may choose,each with equal probability (see Figure 1.6c). The reader is referred to Section 12.1, where this and related problems are discussed. (a) Write a program to simulate a random walk in one dimension starting at 0. Have your program print out the lengths of the times betweenreturns to the starting point (returns to 0). See if you can guess fromthis simulation the answer to the following question: Will the walkeralways return to his starting point eventually or might he drift away forever? (b) The paths of two walkers in two dimensions who meet after nsteps can be considered to be a single path that starts at (0 ;0) and returns to (0 ;0) after 2nsteps. This means that the probability that two random walkers in two dimensions meet is the same as the probability that a single walkerin two dimensions ever returns to the starting point. Thus the questionof whether two walkers are sure to meet is the same as the question ofwhether a single walker is sure to return to the starting point. Write a program to simulate a random walk in two dimensions and see if you think that the walker is sure to return to (0 ;0). If so, P¶ olya would be sure to keep meeting his friends in the park. Perhaps by now youhave conjectured the answer to the question: Is a random walker in oneor two dimensions sure to return to the starting point? P¶ olya answered this question for dimensions one, two, and three. He established theremarkable result that the answer is yesin one and two dimensions and noin three dimensions. 13G. P¶ olya, \Two Incidents," Scientists at Work: Festschrift in Honour of Herman Wold, ed. T. Dalenius, G. Karlsson, and S. Malmquist (Uppsala: Almquist & Wiksells Boktryckeri AB,1970). 18 CHAPTER 1. DISCRETE PROBABILITY DISTRIBUTIONS (c) Write a program to simulate a random walk in three dimensions and see whether, from this simulation and the results of (a) and (b), you couldhave guessed P¶ olya’s result. 1.2 Discrete Probability Distributions In this book we shall study many difierent experiments from a probabilistic point of view. What is involved in this study will become evident as the theory is developedand examples are analyzed. However, the overall idea can be described and illus-trated as follows: to each experiment that we consider there will be associated arandom variable, which represents the outcome of any particular experiment. Theset of possible outcomes is called the sample space . In the flrst part of this section, we will consider the case where the experiment has only flnitely many possible out-comes, i.e., the sample space is flnite. We will then generalize to the case that thesample space is either flnite or countably inflnite. This leads us to the followingdeflnition. Random Variables and Sample Spaces Deflnition 1.1 Suppose we have an experiment whose outcome depends on chance. We represent the outcome of the experiment by a capital Roman letter, such as X, called a random variable . The sample space of the experiment is the set of all possible outcomes. If the sample space is either flnite or countably inflnite, therandom variable is said to be discrete . 2 We generally denote a sample space by the capital Greek letter ›. As stated above, in the correspondence between an experiment and the mathematical theory by whichit is studied, the sample space › corresponds to the set of possible outcomes of theexperiment. We now make two additional deflnitions. These are subsidiary to the deflnition of sample space and serve to make precise some of the common terminology usedin conjunction with sample spaces. First of all, we deflne the elements of a samplespace to be outcomes . Second, each subset of a sample space is deflned to be an event . Normally, we shall denote outcomes by lower case letters and events by capital letters. Example 1.6 A die is rolled once. We let Xdenote the outcome of this experiment. Then the sample space for this experiment is the 6-element set ›=f1;2;3;4;5;6g; where each outcome i, fori= 1 , ..., 6 , corresponds to the number of dots on the face which turns up. The event E=f2;4;6g 1.2. DISCRETE PROBABILITY DISTRIBUTIONS 19 corresponds to the statement that the result of the roll is an even number. The eventEcan also be described by saying that Xis even. Unless there is reason to believe the die is loaded, the natural assumption is that every outcome is equallylikely. Adopting this convention means that we assign a probability of 1/6 to eachof the six outcomes, i.e., m(i)=1=6, for 1•i•6. 2 Distribution Functions We next describe the assignment of probabilities. The deflnitions are motivated by the example above, in which we assigned to each outcome of the sample space anonnegative number such that the sum of the numbers assigned is equal to 1. Deflnition 1.2 LetXbe a random variable which denotes the value of the out- come of a certain experiment, and assume that this experiment has only flnitelymany possible outcomes. Let › be the sample space of the experiment (i.e., theset of all possible values of X, or equivalently, the set of all possible outcomes of the experiment.) A distribution function forXis a real-valued function mwhose domain is › and which satisfles: 1.m(!)‚0; for all!2› , and 2.P !2›m(!)=1. For any subset Eof ›, we deflne the probability ofEto be the number P(E) given by P(E)=X !2Em(!): 2 Example 1.7 Consider an experiment in which a coin is tossed twice. Let Xbe the random variable which corresponds to this experiment. We note that there areseveral ways to record the outcomes of this experiment. We could, for example,record the two tosses, in the order in which they occurred. In this case, we have›=fHH,HT,TH,TTg. We could also record the outcomes by simply noting the number of heads that appeared. In this case, we have › = f0,1,2g. Finally, we could record the two outcomes, without regard to the order in which they occurred. Inthis case, we have › = fHH,HT,TTg. We will use, for the moment, the flrst of the sample spaces given above. We will assume that all four outcomes are equally likely, and deflne the distributionfunctionm(!)b y m(HH) =m(HT) =m(TH) =m(TT) =1 4: 20 CHAPTER 1. DISCRETE PROBABILITY DISTRIBUTIONS LetE=fHH,HT,THgbe the event that at least one head comes up. Then, the probability of Ecan be calculated as follows: P(E)=m(HH) +m(HT) +m(TH) =1 4+1 4+1 4=3 4: Similarly, if F=fHH,HTgis the event that heads comes up on the flrst toss, then we have P(F)=m(HH) +m(HT) =1 4+1 4=1 2: 2 Example 1.8 (Example 1.6 continued) The sample space for the experiment in which the die is rolled is the 6-element set › = f1;2;3;4;5;6g. We assumed that the die was fair, and we chose the distribution function deflned by m(i)=1 6; fori=1;:::; 6: IfEis the event that the result of the roll is an even number, then E=f2;4;6g and P(E)=m(2) +m(4) +m(6) =1 6+1 6+1 6=1 2: 2 Notice that it is an immediate consequence of the above deflnitions that, for every!2›, P(f!g)=m(!): That is, the probability of the elementary event f!g, consisting of a single outcome !, is equal to the value m(!) assigned to the outcome !by the distribution function. Example 1.9 Three people, A, B, and C, are running for the same o–ce, and we assume that one and only one of them wins. The sample space may be taken as the3-element set › = fA,B,Cgwhere each element corresponds to the outcome of that candidate’s winning. Suppose that A and B have the same chance of winning, butthat C has only 1/2 the chance of A or B. Then we assign m(A) =m( B )=2m(C): Since m(A) +m(B) +m( C )=1; 1.2. DISCRETE PROBABILITY DISTRIBUTIONS 21 we see that 2m( C )+2m(C) +m( C )=1; which implies that 5 m(C) = 1. Hence, m(A) =2 5;m (B) =2 5;m (C) =1 5: LetEbe the event that either A or C wins. Then E=fA,Cg, and P(E)=m(A) +m(C) =2 5+1 5=3 5: 2 In many cases, events can be described in terms of other events through the use of the standard constructions of set theory. We will brie°y review the deflnitions ofthese constructions. The reader is referred to Figure 1.7 for Venn diagrams whichillustrate these constructions. LetAandBbe two sets. Then the union of AandBis the set A[B=fxjx2Aorx2Bg: The intersection of AandBis the set A\B=fxjx2Aandx2Bg: The difierence of AandBis the set A¡B=fxjx2Aandx62Bg: The setAis a subset of B, writtenA‰B, if every element of Ais also an element ofB. Finally, the complement of Ais the set ~A=fxjx2› andx62Ag: The reason that these constructions are important is that it is typically the case that complicated events described in English can be broken down into simplerevents using these constructions. For example, if Ais the event that \it will snow tomorrow and it will rain the next day," Bis the event that \it will snow tomorrow," andCis the event that \it will rain two days from now," then Ais the intersection of the events BandC. Similarly, if Dis the event that \it will snow tomorrow or it will rain the next day," then D=B[C. (Note that care must be taken here, because sometimes the word \or" in English means that exactly one of the twoalternatives will occur. The meaning is usually clear from context. In this book,we will always use the word \or" in the inclusive sense, i.e., AorBmeans that at least one of the two events A,Bis true.) The event ~Bis the event that \it will not snow tomorrow." Finally, if Eis the event that \it will snow tomorrow but it will not rain the next day," then E=B¡C. 22 CHAPTER 1. DISCRETE PROBABILITY DISTRIBUTIONS A B A B ABA A A B ∼ ⊃ ⊃AB A B Figure 1.7: Basic set operations. Properties Theorem 1.1 The probabilities assigned to events by a distribution function on a sample space › satisfy the following properties: 1.P(E)‚0 for every E‰›. 2.P( › )=1. 3. IfE‰F‰›, thenP(E)•P(F). 4. IfAandBaredisjoint subsets of ›, then P(A[B)=P(A)+P(B). 5.P(~A)=1¡P(A) for every A‰›. Proof. For any event Ethe probability P(E) is determined from the distribution mby P(E)=X !2Em(!); for everyE‰›. Since the function mis nonnegative, it follows that P(E) is also nonnegative. Thus, Property 1 is true. Property 2 is proved by the equations P(›) =X !2›m(!)=1: Suppose that E‰F‰›. Then every element !that belongs to Ealso belongs toF. Therefore,X !2Em(!)•X !2Fm(!); since each term in the left-hand sum is in the right-hand sum, and all the terms in both sums are non-negative. This implies that P(E)•P(F); and Property 3 is proved. 1.2. DISCRETE PROBABILITY DISTRIBUTIONS 23 Suppose next that AandBare disjoint subsets of ›. Then every element !of A[Blies either in Aand not in Bor inBand not in A. It follows that P(A[B)=P !2A[Bm(!)=P !2Am(!)+P !2Bm(!) =P(A)+P(B); and Property 4 is proved. Finally, to prove Property 5, consider the disjoint union ›=A[~A: SinceP(›) = 1, the property of disjoint additivity (Property 4) implies that 1=P(A)+P(~A); whenceP(~A)=1¡P(A). 2 It is important to realize that Property 4 in Theorem 1.1 can be extended to more than two sets. The general flnite additivity property is given by the followingtheorem. Theorem 1.2 IfA 1, ...,Anare pairwise disjoint subsets of › (i.e., no two of the Ai’s have an element in common), then P(A1[¢¢¢[An)=nX i=1P(Ai): Proof. Let!be any element in the union A1[¢¢¢[An: Thenm(!) occurs exactly once on each side of the equality in the statement of the theorem. 2 We shall often use the following consequence of the above theorem. Theorem 1.3 LetA1, ...,Anbe pairwise disjoint events with › = A1[¢¢¢[An, and letEbe any event. Then P(E)=nX i=1P(E\Ai): Proof. The setsE\A1, ...,E\Anare pairwise disjoint, and their union is the setE. The result now follows from Theorem 1.2. 2 24 CHAPTER 1. DISCRETE PROBABILITY DISTRIBUTIONS Corollary 1.1 For any two events AandB, P(A)=P(A\B)+P(A\~B): 2 Property 4 can be generalized in another way. Suppose that AandBare subsets of › which are not necessarily disjoint. Then: Theorem 1.4 IfAandBare subsets of ›, then P(A[B)=P(A)+P(B)¡P(A\B): (1.1) Proof. The left side of Equation 1.1 is the sum of m(!) for!in eitherAorB.W e must show that the right side of Equation 1.1 also adds m(!) for!inAorB.I f! is in exactly one of the two sets, then it is counted in only one of the three termson the right side of Equation 1.1. If it is in both AandB, it is added twice from the calculations of P(A) andP(B) and subtracted once for P(A\B). Thus it is counted exactly once by the right side. Of course, if A\B=;, then Equation 1.1 reduces to Property 4. (Equation 1.1 can also be generalized; see Theorem 3.8.) 2 Tree Diagrams Example 1.10 Let us illustrate the properties of probabilities of events in terms of three tosses of a coin. When we have an experiment which takes place in stagessuch as this, we often flnd it convenient to represent the outcomes by a tree diagram as shown in Figure 1.8. Apath through the tree corresponds to a possible outcome of the experiment. For the case of three tosses of a coin, we have eight paths ! 1,!2, ...,!8and, assuming each outcome to be equally likely, we assign equal weight, 1/8, to eachpath. LetEbe the event \at least one head turns up." Then ~Eis the event \no heads turn up." This event occurs for only one outcome, namely, ! 8= TTT. Thus, ~E=fTTTgand we have P(~E)=P(fTTTg)=m(TTT) =1 8: By Property 5 of Theorem 1.1, P(E)=1¡P(~E)=1¡1 8=7 8: Note that we shall often flnd it is easier to compute the probability that an event does not happen rather than the probability that it does. We then use Property 5to obtain the desired probability. 1.2. DISCRETE PROBABILITY DISTRIBUTIONS 25 First toss Second toss Third toss Outcome H HH HH H TT TT T T(Start)ω ω ω ω ωω ω ω1 2 3 4 5 6 7 8H T Figure 1.8: Tree diagram for three tosses of a coin. LetAbe the event \the flrst outcome is a head," and Bthe event \the second outcome is a tail." By looking at the paths in Figure 1.8, we see that P(A)=P(B)=1 2: Moreover,A\B=f!3;!4g, and soP(A\B)=1=4:Using Theorem 1.4, we obtain P(A[B)=P(A)+P(B)¡P(A\B) =1 2+1 2¡1 4=3 4: SinceA[Bis the 6-element set, A[B=fHHH,HHT,HTH,HTT,TTH,TTT g; we see that we obtain the same result by direct enumeration. 2 In our coin tossing examples and in the die rolling example, we have assigned an equal probability to each possible outcome of the experiment. Corresponding tothis method of assigning probabilities, we have the following deflnitions. Uniform Distribution Deflnition 1.3 The uniform distribution on a sample space › containing nele- ments is the function mdeflned by m(!)=1 n; for every!2›. 2 26 CHAPTER 1. DISCRETE PROBABILITY DISTRIBUTIONS It is important to realize that when an experiment is analyzed to describe its possible outcomes, there is no single correct choice of sample space. For the ex-periment of tossing a coin twice in Example 1.2, we selected the 4-element set›=fHH,HT,TH,TTgas a sample space and assigned the uniform distribution func- tion. These choices are certainly intuitively natural. On the other hand, for somepurposes it may be more useful to consider the 3-element sample space „›=f0;1;2g in which 0 is the outcome \no heads turn up," 1 is the outcome \exactly one headturns up," and 2 is the outcome \two heads turn up." The distribution function „ m on„› deflned by the equations „m(0) =1 4; „m(1) =1 2; „m(2) =1 4 is the one corresponding to the uniform probability density on the original sample space ›. Notice that it is perfectly possible to choose a difierent distribution func-tion. For example, we may consider the uniform distribution function on „›, which is the function „ qdeflned by „q(0) = „q(1) = „q(2) =1 3: Although „qis a perfectly good distribution function, it is not consistent with ob- served data on coin tossing. Example 1.11 Consider the experiment that consists of rolling a pair of dice. We take as the sample space › the set of all ordered pairs ( i;j) of integers with 1 •i•6 and 1•j•6. Thus, ›=f(i;j):1•i;j•6g: (There is at least one other \reasonable" choice for a sample space, namely the set of all unordered pairs of integers, each between 1 and 6. For a discussion of whywe do not use this set, see Example 3.14.) To determine the size of ›, we notethat there are six choices for i, and for each choice of ithere are six choices for j, leading to 36 difierent outcomes. Let us assume that the dice are not loaded. Inmathematical terms, this means that we assume that each of the 36 outcomes isequally likely, or equivalently, that we adopt the uniform distribution function on› by setting m((i;j)) =1 36; 1•i;j•6: What is the probability of getting a sum of 7 on the roll of two dice|or getting a sum of 11? The flrst event, denoted by E, is the subset E=f(1;6);(6;1);(2;5);(5;2);(3;4);(4;3)g: A sum of 11 is the subset Fgiven by F=f(5;6);(6;5)g: Consequently, P(E)=P !2Em(!)=6¢1 36=1 6; P(F)=P !2Fm(!)=2¢1 36=1 18: 1.2. DISCRETE PROBABILITY DISTRIBUTIONS 27 What is the probability of getting neither snakeeyes (double ones) nor boxcars (double sixes)? The event of getting either one of these two outcomes is the set E=f(1;1);(6;6)g: Hence, the probability of obtaining neither is given by P(~E)=1¡P(E)=1¡2 36=17 18: 2 In the above coin tossing and the dice rolling experiments, we have assigned an equal probability to each outcome. That is, in each example, we have chosen theuniform distribution function. These are the natural choices provided the coin is afair one and the dice are not loaded. However, the decision as to which distributionfunction to select to describe an experiment is nota part of the basic mathemat- ical theory of probability. The latter begins only when the sample space and thedistribution function have already been deflned. Determination of Probabilities It is important to consider ways in which probability distributions are determinedin practice. One way is by symmetry. For the case of the toss of a coin, we do not see any physical difierence between the two sides of a coin that should afiect thechance of one side or the other turning up. Similarly, with an ordinary die thereis no essential difierence between any two sides of the die, and so by symmetry weassign the same probability for any possible outcome. In general, considerationsof symmetry often suggest the uniform distribution function. Care must be usedhere. We should not always assume that, just because we do not know any reasonto suggest that one outcome is more likely than another, it is appropriate to assignequal probabilities. For example, consider the experiment of guessing the sex ofa newborn child. It has been observed that the proportion of newborn childrenwho are boys is about .513. Thus, it is more appropriate to assign a distributionfunction which assigns probability .513 to the outcome boyand probability .487 to the outcome girlthan to assign probability 1/2 to each outcome. This is an example where we use statistical observations to determine probabilities. Note that theseprobabilities may change with new studies and may vary from country to country.Genetic engineering might even allow an individual to in°uence this probability fora particular case. Odds Statistical estimates for probabilities are flne if the experiment under considerationcan be repeated a number of times under similar circumstances. However, assumethat, at the beginning of a football season, you want to assign a probability to theevent that Dartmouth will beat Harvard. You really do not have data that relates tothis year’s football team. However, you can determine your own personal probability 28 CHAPTER 1. DISCRETE PROBABILITY DISTRIBUTIONS by seeing what kind of a bet you would be willing to make. For example, suppose that you are willing to mak e a 1 dollar bet giving 2 to 1 odds that Dartmouth will win. Then you are willing to pay 2 dollars if Dartmouth loses in return for receiving1 dollar if Dartmouth wins. This means that you think the appropriate probabilityfor Dartmouth winning is 2/3. Let us look more carefully at the relation between odds and probabilities. Sup- pose that we make a bet at rto 1 odds that an event Eoccurs. This means that we think that it is rtimes as likely that Ewill occur as that Ewill not occur. In general,rtosodds will be taken to mean the same thing as r=sto 1, i.e., the ratio between the two numbers is the only quantity of importance when stating odds. N o wi fi ti s rtimes as likely that Ewill occur as that Ewill not occur, then the probability that Eoccurs must be r=(r+ 1), since we have P(E)=rP(~E) and P(E)+P(~E)=1: In general, the statement that the odds are rtosin favor of an event Eoccurring is equivalent to the statement that P(E)=r=s (r=s)+1 =r r+s: If we letP(E)=p, then the above equation can easily be solved for r=sin terms of p; we obtain r=s=p=(1¡p). We summarize the above discussion in the following deflnition. Deflnition 1.4 IfP(E)=p, the odds in favor of the event Eoccurring are r:s(r tos) wherer=s=p=(1¡p). Ifrandsare given, then pcan be found by using the equationp=r=(r+s). 2 Example 1.12 (Example 1.9 continued) In Example 1.9 we assigned probability 1/5 to the event that candidate C wins the race. Thus the odds in favor of Cwinning are 1 =5:4=5. These odds could equally well have been written as 1 : 4, 2 : 8, and so forth. A bet that C wins is fair if we receive 4 dollars if C wins andpay 1 dollar if C loses. 2 Inflnite Sample Spaces If a sample space has an inflnite number of points, then the way that a distribution function is deflned depends upon whether or not the sample space is countable. Asample space is countably inflnite if the elements can be counted, i.e., can be put in one-to-one correspondence with the positive integers, and uncountably inflnite 1.2. DISCRETE PROBABILITY DISTRIBUTIONS 29 otherwise. Inflnite sample spaces require new concepts in general (see Chapter 2), but countably inflnite spaces do not. If ›=f!1;!2;!3;:::g is a countably inflnite sample space, then a distribution function is deflned exactly as in Deflnition 1.2, except that the sum must now be a convergent inflnite sum. Theorem 1.1 is still true, as are its extensions Theorems 1.2 and 1.4. One thing wecannot do on a countably inflnite sample space that we could do on a flnite samplespace is to deflne a uniform distribution function as in Deflnition 1.3. You are asked in Exercise 20 to explain why this is not possible. Example 1.13 A coin is tossed until the flrst time that a head turns up. Let the outcome of the experiment, !, be the flrst time that a head turns up. Then the possible outcomes of our experiment are ›=f1;2;3;:::g: Note that even though the coin could come up tails every time we have not allowed for this possibility. We will explain why in a moment. The probability that headscomes up on the flrst toss is 1/2. The probability that tails comes up on the flrsttoss and heads on the second is 1/4. The probability that we have two tails followedby a head is 1/8, and so forth. This suggests assigning the distribution functionm(n)=1=2 nforn= 1 , 2 , 3 , .... T o s e e that this is a distribution function we must show thatX !m(!)=1 2+1 4+1 8+¢¢¢=1: That this is true follows from the formula for the sum of a geometric series, 1+r+r2+r3+¢¢¢=1 1¡r; or r+r2+r3+r4+¢¢¢=r 1¡r; (1.2) for¡1<r< 1. Puttingr=1=2, we see that we have a probability of 1 that the coin eventu- ally turns up heads. The possible outcome of tails every time has to be assignedprobability 0, so we omit it from our sample space of possible outcomes. LetEbe the event that the flrst time a head turns up is after an even number of tosses. Then E=f2;4;6;8;:::g; and P(E)=1 4+1 16+1 64+¢¢¢: Puttingr=1=4 in Equation 1.2 see that P(E)=1=4 1¡1=4=1 3: Thus the probability that a head turns up for the flrst time after an even number of tosses is 1/3 and after an odd number of tosses is 2/3. 2 30 CHAPTER 1. DISCRETE PROBABILITY DISTRIBUTIONS Historical Remarks An interesting question in the history of science is: Why was probability not devel- oped until the sixteenth century? We know that in the sixteenth century problemsin gambling and games of chance made people start to think about probability. Butgambling and games of chance are almost as old as civilization itself. In ancientEgypt (at the time of the First Dynasty, ca. 3500 B.C.) a game now called \Hounds and Jackals" was played. In this game the movement of the hounds and jackals wasbased on the outcome of the roll of four-sided dice made out of animal bones calledastragali. Six-sided dice made of a variety of materials date back to the sixteenthcentury B.C.Gambling was widespread in ancient Greece and Rome. Indeed, in the Roman Empire it was sometimes found necessary to invoke laws against gambling.Why, then, were probabilities not calculated until the sixteenth century? Several explanations have been advanced for this late development. One is that the relevant mathematics was not developed and was not easy to develop. Theancient mathematical notation made numerical calculation complicated, and ourfamiliar algebraic notation was not developed until the sixteenth century. However,as we shall see, many of the combinatorial ideas needed to calculate probabilitieswere discussed long before the sixteenth century. Since many of the chance eventsof those times had to do with lotteries relating to religious afiairs, it has beensuggested that there may have been religious barriers to the study of chance andgambling. Another suggestion is that a stronger incentive, such as the developmentof commerce, was necessary. However, none of these explanations seems completelysatisfactory, and people still wonder why it took so long for probability to be studiedseriously. An interesting discussion of this problem can be found in Hacking. 14 The flrst person to calculate probabilities systematically was Gerolamo Cardano (1501{1576) in his book Liber de Ludo Aleae. This was translated from the Latin by Gould and appears in the book Cardano: The Gambling Scholar by Ore.15Ore provides a fascinating discussion of the life of this colorful scholar with accountsof his interests in many difierent flelds, including medicine, astrology, and mathe-matics. You will also flnd there a detailed account of Cardano’s famous battle withTartaglia over the solution to the cubic equation. In his book on probability Cardano dealt only with the special case that we have called the uniform distribution function. This restriction to equiprobable outcomeswas to continue for a long time. In this case Cardano realized that the probabilitythat an event occurs is the ratio of the number of favorable outcomes to the totalnumber of outcomes. Many of Cardano’s examples dealt with rolling dice. Here he realized that the outcomes for two rolls should be taken to be the 36 ordered pairs ( i;j) rather than the 21 unordered pairs. This is a subtle point that was still causing problems muchlater for other writers on probability. For example, in the eighteenth century thefamous French mathematician d’Alembert, author of several works on probability,claimed that when a coin is tossed twice the number of heads that turn up would 14I. Hacking, The Emergence of Probability (Cambridge: Cambridge University Press, 1975). 15O. Ore, Cardano: The Gambling Scholar (Princeton: Princeton University Press, 1953). 1.2. DISCRETE PROBABILITY DISTRIBUTIONS 31 be 0, 1, or 2, and hence we should assign equal probabilities for these three possible outcomes.16Cardano chose the correct sample space for his dice problems and calculated the correct probabilities for a variety of events. Cardano’s mathematical work is interspersed with a lot of advice to the potential gambler in short paragraphs, entitled, for example: \Who Should Play and When,"\Why Gambling Was Condemned by Aristotle," \Do Those Who Teach Also PlayWell?" and so forth. In a paragraph entitled \The Fundamental Principle of Gam-bling," Cardano writes: The most fundamental principle of all in gambling is simply equal con- ditions, e.g., of opponents, of bystanders, of money, of situation, of thedice box, and of the die itself. To the extent to which you depart fromthat equality, if it is in your opponent’s favor, you are a fool, and if inyour own, you are unjust. 17 Cardano did make mistakes, and if he realized it later he did not go back and change his error. For example, for an event that is favorable in three out of fourcases, Cardano assigned the correct odds 3 : 1 that the event will occur. But then heassigned odds by squaring these numbers (i.e., 9 : 1) for the event to happen twice ina row. Later, by considering the case where the odds are 1 : 1, he realized that thiscannot be correct and was led to the correct result that when fout ofnoutcomes are favorable, the odds for a favorable outcome twice in a row are f 2:n2¡f2. Ore points out that this is equivalent to the realization that if the probability that anevent happens in one experiment is p, the probability that it happens twice is p 2. Cardano proceeded to establish that for three successes the formula should be p3 and for four successes p4, making it clear that he understood that the probability ispnfornsuccesses in nindependent repetitions of such an experiment. This will follow from the concept of independence that we introduce in Section 4.1. Cardano’s work was a remarkable flrst attempt at writing down the laws of probability, but it was not the spark that started a systematic study of the subject.This came from a famous series of letters between Pascal and Fermat. This corre-spondence was initiated by Pascal to consult Fermat about problems he had beengiven by Chevalier de M¶ er¶e, a well-known writer, a prominent flgure at the court of Louis XIV, and an ardent gambler. The flrst problem de M¶ er¶e posed was a dice problem. The story goes that he had been betting that at least one six would turn up in four rolls of a die and winningtoo often, so he then bet that a pair of sixes would turn up in 24 rolls of a pairof dice. The probability of a six with one die is 1/6 and, by the product law forindependent experiments, the probability of two sixes when a pair of dice is thrownis (1=6)(1=6 )=1=36. Ore 18claims that a gambling rule of the time suggested that, since four repetitions was favorable for the occurrence of an event with probability1/6, for an event six times as unlikely, 6 ¢4 = 24 repetitions would be su–cient for 16J. d’Alembert, \Croix ou Pile," in L’Encyclop¶ edie, ed. Diderot, vol. 4 (Paris, 1754). 17O. Ore, op. cit., p. 189. 18O. Ore, \Pascal and the Invention of Probability Theory," American Mathematics Monthly , vol. 67 (1960), pp. 409{419. 32 CHAPTER 1. DISCRETE PROBABILITY DISTRIBUTIONS a favorable bet. Pascal showed, by exact calculation, that 25 rolls are required for a favorable bet for a pair of sixes. The second problem was a much harder one: it was an old problem and con- cerned the determination of a fair division of the stakes in a tournament when theseries, for some reason, is interrupted before it is completed. This problem is nowreferred to as the problem of points. The problem had been a standard problem inmathematical texts; it appeared in Fra Luca Paccioli’s book summa de Arithmetica, Geometria, Proportioni et Proportionalitµ a, printed in Venice in 1494, 19in the form: A team plays ball such that a total of 60 points are required to win the game, and each inning counts 10 points. The stakes are 10 ducats. Bysome incident they cannot flnish the game and one side has 50 pointsand the other 20. One wants to know what share of the prize moneybelongs to each side. In this case I have found that opinions difier fromone to another but all seem to me insu–cient in their arguments, but Ishall state the truth and give the correct way. Reasonable solutions, such as dividing the stakes according to the ratio of games won by each player, had been proposed, but no correct solution had been found atthe time of the Pascal-Fermat correspondence. The letters deal mainly with theattempts of Pascal and Fermat to solve this problem. Blaise Pascal (1623{1662)was a child prodigy, having published his treatise on conic sections at age sixteen,and having invented a calculating machine at age eighteen. At the time of theletters, his demonstration of the weight of the atmosphere had already establishedhis position at the forefront of contemporary physicists. Pierre de Fermat (1601{1665) was a learned jurist in Toulouse, who studied mathematics in his spare time.He has been called by some the prince of amateurs and one of the greatest puremathematicians of all times. The letters, translated by Maxine Merrington, appear in Florence David’s fasci- nating historical account of probability, Games, Gods and Gambling . 20In a letter dated Wednesday, 29th July, 1654, Pascal writes to Fermat: Sir, Like you, I am equally impatient, and although I am again ill in bed, I cannot help telling you that yesterday evening I received from M. deCarcavi your letter on the problem of points, which I admire more thanI can possibly say. I have not the leisure to write at length, but, in aword, you have solved the two problems of points, one with dice and theother with sets of games with perfect justness; I am entirely satisfledwith it for I do not doubt that I was in the wrong, seeing the admirableagreement in which I flnd myself with you now. . . Your method is very sound and is the one which flrst came to my mind in this research; but because the labour of the combination is excessive,I have found a short cut and indeed another method which is much 19ibid., p. 414. 20F. N. David, Games, Gods and Gambling (London: G. Gri–n, 1962), p. 230 fi. 1.2. DISCRETE PROBABILITY DISTRIBUTIONS 33 0123 0 12 30 00 81 63 2 6 4 20 32 48 64 64 32 44 56 Number of games A has wonNumber of games B has won Figure 1.9: Pascal’s table. quicker and neater, which I would like to tell you here in a few words: for henceforth I would like to open my heart to you, if I may, as I am sooverjoyed with our agreement. I see that truth is the same in Toulouseas in Paris. Here, more or less, is what I do to show the fair value of each game, when two opponents play, for example, in three games and each personhas staked 32 pistoles. Let us say that the flrst man had won twice and the other once; now they play another game, in which the conditions are that, if the flrstwins, he takes all the stakes; that is 64 pistoles; if the other wins it,then they have each won two games, and therefore, if they wish to stopplaying, they must each take back their own stake, that is, 32 pistoleseach. Then consider, Sir, if the flrst man wins, he gets 64 pistoles; if he loses he gets 32. Thus if they do not wish to risk this last game but wish toseparate without playing it, the flrst man must say: ‘I am certain to get32 pistoles, even if I lost I still get them; but as for the other 32, perhapsI will get them, perhaps you will get them, the chances are equal. Letus then divide these 32 pistoles in half and give one half to me as wellas my 32 which are mine for sure.’ He will then have 48 pistoles and theother 16. . . Pascal’s argument produces the table illustrated in Figure 1.9 for the amount due player A at any quitting point. Each entry in the table is the average of the numbers just above and to the right of the number. This fact, together with the known values when the tournament iscompleted, determines all the values in this table. If player A wins the flrst game, 34 CHAPTER 1. DISCRETE PROBABILITY DISTRIBUTIONS then he needs two games to win and B needs three games to win; and so, if the tounament is called ofi, A should receive 44 pistoles. The letter in which Fermat presented his solution has been lost; but fortunately, Pascal describes Fermat’s method in a letter dated Monday, 24th August, 1654.From Pascal’s letter: 21 This is your procedure when there are two players: If two players, play-ing several games, flnd themselves in that position when the flrst manneeds twogames and second needs three , then to flnd the fair division of stakes, you say that one must know in how many games the play willbe absolutely decided. It is easy to calculate that this will be in four games, from which you can conclude that it is necessary to see in how many ways four games can bearranged between two players, and one must see how many combinationswould make the flrst man win and how many the second and to shareout the stakes in this proportion. I would have found it di–cult tounderstand this if I had not known it myself already; in fact you hadexplained it with this idea in mind. Fermat realized that the number of ways that the game might be flnished may not be equally likely. For example, if A needs two more games and B needs three towin, two possible ways that the tournament might go for A to win are WLW andLWLW. These two sequences do not have the same chance of occurring. To avoidthis di–culty, Fermat extended the play, adding flctitious plays, so that all the waysthat the games might go have the same length, namely four. He was shrewd enoughto realize that this extension would not change the winner and that he now couldsimply count the number of sequences favorable to each player since he had madethem all equally likely. If we list all possible ways that the extended game of fourplays might go, we obtain the following 16 possible outcomes of the play: WWWW WLWW LWWW LLWW WWWL WLWL LWWL LLWL WWLW WLLW LWLW LLLW WWLL WLLL LWLL LLLL . Player A wins in the cases where there are at least two wins (the 11 underlined cases), and B wins in the cases where there are at least three losses (the other5 cases). Since A wins in 11 of the 16 possible cases Fermat argued that theprobability that A wins is 11/16. If the stakes are 64 pistoles, A should receive44 pistoles in agreement with Pascal’s result. Pascal and Fermat developed moresystematic methods for counting the number of favorable outcomes for problemslike this, and this will be one of our central problems. Such counting methods fallunder the subject of combinatorics , which is the topic of Chapter 3. 21ibid., p. 239fi. 1.2. DISCRETE PROBABILITY DISTRIBUTIONS 35 We see that these two mathematicians arrived at two very difierent ways to solve the problem of points. Pascal’s method was to develop an algorithm and use it tocalculate the fair division. This method is easy to implement on a computer and easyto generalize. Fermat’s method, on the other hand, was to change the problem intoan equivalent problem for which he could use counting or combinatorial methods.We will see in Chapter 3 that, in fact, Fermat used what has become known asPascal’s triangle! In our study of probability today we shall flnd that both thealgorithmic approach and the combinatorial approach share equal billing, just asthey did 300 years ago when probability got its start. Exercises 1Let › =fa;b;cgbe a sample space. Let m(a)=1=2,m(b)=1=3, and m(c)=1=6. Find the probabilities for all eight subsets of ›. 2Give a possible sample space › for each of the following experiments: (a) An election decides between two candidates A and B. (b) A two-sided coin is tossed. (c) A student is asked for the month of the year and the day of the week on which her birthday falls. (d) A student is chosen at random from a class of ten students. (e) You receive a grade in this course. 3For which of the cases in Exercise 2 would it be reasonable to assign the uniform distribution function? 4Describe in words the events specifled by the following subsets of ›=fH H H ;H H T ;H T H ;H T T ;T H H ;T H T ;T T H ;T T T g (see Example 1.6). (a)E=fHHH,HHT,HTH,HTT g. (b)E=fHHH,TTTg. (c)E=fHHT,HTH,THHg. (d)E=fHHT,HTH,HTT,THH,THT,TTH,TTT g. 5What are the probabilities of the events described in Exercise 4? 6A die is loaded in such a way that the probability of each face turning up is proportional to the number of dots on that face. (For example, a six isthree times as probable as a two.) What is the probability of getting an evennumber in one throw? 7LetAandBbe events such that P(A\B)=1=4,P(~A)=1=3, andP(B)= 1=2. What is P(A[B)? 36 CHAPTER 1. DISCRETE PROBABILITY DISTRIBUTIONS 8A student must choose one of the subjects, art, geology, or psychology, as an elective. She is equally likely to choose art or psychology and twice as likelyto choose geology. What are the respective probabilities that she chooses art,geology, and psychology? 9A student must choose exactly two out of three electives: art, French, and mathematics. He chooses art with probability 5/8, French with probability5/8, and art and French together with probability 1/4. What is the probabilitythat he chooses mathematics? What is the probability that he chooses eitherart or French? 10For a bill to come before the president of the United States, it must be passed by both the House of Representatives and the Senate. Assume that, of thebills presented to these two bodies, 60 percent pass the House, 80 percentpass the Senate, and 90 percent pass at least one of the two. Calculate theprobability that the next bill presented to the two groups will come before thepresident. 11What odds should a person give in favor of the following events? (a) A card chosen at random from a 52-card deck is an ace. (b) Two heads will turn up when a coin is tossed twice. (c) Boxcars (two sixes) will turn up when two dice are rolled. 12You ofier 3 : 1 odds that your friend Smith will be elected mayor of your city. What probability are you assigning to the event that Smith wins? 13In a horse race, the odds that Romance will win are listed as 2 : 3 and that Downhill will win are 1 : 2. What odds should be given for the event thateither Romance or Downhill wins? 14LetXbe a random variable with distribution function m X(x) deflned by mX(¡1 )=1=5;mX( 0 )=1=5;mX( 1 )=2=5;mX( 2 )=1=5: (a) LetYbe the random variable deflned by the equation Y=X+ 3. Find the distribution function mY(y)o fY. (b) LetZbe the random variable deflned by the equation Z=X2. Find the distribution function mZ(z)o fZ. *15 John and Mary are taking a mathematics course. The course has only three grades: A, B, and C. The probability that John get saBi s. 3 . T h e probability that Mary get saBi s. 4 .T h e probability that neither gets an A but at least o n eg e t saBi s. 1 . What is the probability that at least one gets a B but neither gets a C? 16In a flerce battle, not less than 70 percent of the soldiers lost one eye, not less than 75 percent lost one ear, not less than 80 percent lost one hand, and not 1.2. DISCRETE PROBABILITY DISTRIBUTIONS 37 less than 85 percent lost one leg. What is the minimal possible percentage of those who simultaneously lost one ear, one eye, one hand, and one leg?22 *17 Assume that the probability of a \success" on a single experiment with n outcomes is 1 =n. Letmbe the number of experiments necessary to make it a favorable bet that at least one success will occur (see Exercise 1.1.5). (a) Show that the probability that, in mtrials, there are no successes is (1¡1=n)m. (b) (de Moivre) Show that if m=nlog 2 then lim n!1µ 1¡1 n¶m =1 2: Hint: lim n!1µ 1¡1 n¶n =e¡1: Hence for large nwe should choose mto be about nlog 2. (c) Would DeMoivre have been led to the correct answer for de M¶ er¶e’s two bets if he had used his approximation? 18(a) For events A1, ...,An, prove that P(A1[¢¢¢[An)•P(A1)+¢¢¢+P(An): (b) For events AandB, prove that P(A\B)‚P(A)+P(B)¡1: 19IfA,B, andCare any three events, show that P(A[B[C)=P(A)+P(B)+P(C) ¡P(A\B)¡P(B\C)¡P(C\A) +P(A\B\C): 20Explain why it is not possible to deflne a uniform distribution function (see Deflnition 1.3) on a countably inflnite sample space. Hint: Assumem(!)=a for all!, where 0•a•1. Doesm(!) have all the properties of a distribution function? 21In Example 1.13 flnd the probability that the coin turns up heads for the flrst time on the tenth, eleventh, or twelfth toss. 22A die is rolled until the flrst time that a six turns up. We shall see that the probability that this occurs on the nth roll is (5=6)n¡1¢(1=6). Using this fact, describe the appropriate inflnite sample space and distribution function forthe experiment of rolling a die until a six turns up for the flrst time. Verifythat for your distribution functionP !m(!)=1 . 22See Knot X, in Lewis Carroll, Mathematical Recreations, vol. 2 (Dover, 1958). 38 CHAPTER 1. DISCRETE PROBABILITY DISTRIBUTIONS 23Let › be the sample space ›=f0;1;2;:::g; and deflne a distribution function by m(j)=( 1¡r)jr; for some flxed r,0<r< 1, and for j=0;1;2;:::. Show that this is a distribution function for ›. 24Our calendar has a 400-year cycle. B. H. Brown noticed that the number of times the thirteenth of the month falls on each of the days of the week in the4800 months of a cycle is as follows: Sunday 687Monday 685Tuesday 685Wednesday 687Thursday 684Friday 688Saturday 684From this he deduced that the thirteenth was more likely to fall on Friday than on any other day. Explain what he meant by this. 25Tversky and Kahneman 23asked a group of subjects to carry out the following task. They are told that: Linda is 31, single, outspoken, and very bright. She majored in philosophy in college. As a student, she was deeply concerned withracial discrimination and other social issues, and participated inanti-nuclear demonstrations. The subjects are then asked to rank the likelihood of various alternatives, such as:(1) Linda is active in the feminist movement.(2) Linda is a bank teller.(3) Linda is a bank teller and active in the feminist movement. Tversky and Kahneman found that between 85 and 90 percent of the subjects rated alternative (1) most likely, but alternative (3) more likely than alterna-tive (2). Is it? They call this phenomenon the conjunction fallacy, and note that it appears to be unafiected by prior training in probability or statistics.Explain why this is a fallacy. Can you give a possible explanation for thesubjects’ choices? 23K. McKean, \Decisions, Decisions," pp. 22{31. 1.2. DISCRETE PROBABILITY DISTRIBUTIONS 39 26Two cards are drawn successively from a deck of 52 cards. Find the probability that the second card is higher in rank than the flrst card. Hint: Show that 1 = P(higher) +P(lower) +P(same) and use the fact that P(higher) =P(lower). 27Alife table is a table that lists for a given number of births the estimated number of people who will live to a given age. In Appendix C we give a lifetable based upon 100,000 births for ages from 0 to 85, both for women and formen. Show how from this table you can estimate the probability m(x) that a person born in 1981 would live to age x. Write a program to plot m(x) both for men and for women, and comment on the difierences that you see in thetwo cases. *28 Here is an attempt to get around the fact that we cannot choose a \random integer." (a) What, intuitively, is the probability that a \randomly chosen" positive integer is a multiple of 3? (b) LetP 3(N) be the probability that an integer, chosen at random between 1 andN, is a multiple of 3 (since the sample space is flnite, this is a legitimate probability). Show that the limit P3= lim N!1P3(N) exists and equals 1/3. This formalizes the intuition in (a), and gives us a way to assign \probabilities" to certain events that are inflnite subsetsof the positive integers. (c) IfAis any set of positive integers, let A(N) mean the number of elements ofAwhich are less than or equal to N. Then deflne the \probability" of Aas P(A) = lim N!1A(N)=N ; provided this limit exists. Show that this deflnition would assign prob- ability 0 to any flnite set and probability 1 to the set of all positiveintegers. Thus, the probability of the set of all integers is not the sum ofthe probabilities of the individual integers in this set. This means thatthe deflnition of probability given here is not a completely satisfactorydeflnition. (d) LetAbe the set of all positive integers with an odd number of dig- its. Show that P(A) does not exist. This shows that under the above deflnition of probability, not all sets have probabilities. 29(from Sholander 24) In a standard clover-leaf interchange, there are four ramps for making right-hand turns, and inside these four ramps, there are four moreramps for making left-hand turns. Your car approaches the interchange fromthe south. A mechanism has been installed so that at each point where thereexists a choice of directions, the car turns to the right with flxed probability r. 24M. Sholander, Problem #1034, Mathematics Magazine, vol. 52, no. 3 (May 1979), p. 183. 40 CHAPTER 1. DISCRETE PROBABILITY DISTRIBUTIONS (a) Ifr=1=2, what is your chance of emerging from the interchange going west? (b) Find the value of rthat maximizes your chance of a westward departure from the interchange. 30(from Benkoski25) Consider a \pure" cloverleaf interchange in which there are no ramps for right-hand turns, but only the two intersecting straighthighways with cloverleaves for left-hand turns. (Thus, to turn right in such an interchange, one must make three left-hand turns.) As in the precedingproblem, your car approaches the interchange from the south. What is thevalue ofrthat maximizes your chances of an eastward departure from the interchange? 31(from vos Savant 26) A reader of Marilyn vos Savant’s column wrote in with the following question: My dad heard this story on the radio. At Duke University, two students had received A’s in chemistry all semester. But on thenight before the flnal exam, they were partying in another stateand didn’t get back to Duke until it was over. Their excuse to theprofessor was that they had a °at tire, and they asked if they couldtake a make-up test. The professor agreed, wrote out a test and sentthe two to separate rooms to take it. The flrst question (on one sideof the paper) was worth 5 points, and they answered it easily. Thenthey °ipped the paper over and found the second question, worth95 points: ‘Which tire was it?’ What was the probability that bothstudents would say the same thing? My dad and I think it’s 1 in16. Is that right?" (a) Is the answer 1/16? (b) The following question was asked of a class of students. \I was driving to school today, and one of my tires went °at. Which tire do you thinkit was?" The responses were as follows: right front, 58%, left front, 11%,right rear, 18%, left rear, 13%. Suppose that this distribution holds inthe general population, and assume that the two test-takers are randomlychosen from the general population. What is the probability that theywill give the same answer to the second question? 25S. Benkoski, Comment on Problem #1034, Mathematics Magazine, vol. 52, no. 3 (May 1979), pp. 183-184. 26M. vos Savant, Parade Magazine , 3 March 1996, p. 14. Chapter 2 Continuous Probability Densities 2.1 Simulation of Continuous Probabilities In this section we shall show how we can use computer simulations for experiments that have a whole continuum of possible outcomes. Probabilities Example 2.1 We begin by constructing a spinner, which consists of a circle of unit circumference and a pointer as shown in Figure 2.1. We pick a point on the circle and label it 0, and then label every other point on the circle with the distance, sayx, from 0 to that point, measured counterclockwise. The experiment consists of spinning the pointer and recording the label of the point at the tip of the pointer.We let the random variable Xdenote the value of this outcome. The sample space is clearly the interval [0 ;1). We would like to construct a probability model in which each outcome is equally likely to occur. If we proceed as we did in Chapter 1 for experiments with a flnite number of possible outcomes, then we must assign the probability 0 to each outcome, sinceotherwise, the sum of the probabilities, over all of the possible outcomes, wouldnot equal 1. (In fact, summing an uncountable number of real numbers is a trickybusiness; in particular, in order for such a sum to have any meaning, at mostcountably many of the summands can be difierent than 0.) However, if all of theassigned probabilities are 0, then the sum is 0, not 1, as it should be. In the next section, we will show how to construct a probability model in this situation. At present, we will assume that such a model can be constructed. Wewill also assume that in this model, if Eis an arc of the circle, and Eis of length p, then the model will assign the probability ptoE. This means that if the pointer is spun, the probability that it ends up pointing to a point in Eequalsp, which is certainly a reasonable thing to expect. 41 42 CHAPTER 2. CONTINUOUS PROBABILITY DENSITIES 0x Figure 2.1: A spinner. To simulate this experiment on a computer is an easy matter. Many computer software packages have a function which returns a random real number in the in-terval [0;1]. Actually, the returned value is always a rational number, and the values are determined by an algorithm, so a sequence of such values is not trulyrandom. Nevertheless, the sequences produced by such algorithms behave muchlike theoretically random sequences, so we can use such sequences in the simulationof experiments. On occasion, we will need to refer to such a function. We will callthis function rnd. 2 Monte Carlo Procedure and Areas It is sometimes desirable to estimate quantities whose exact values are di–cult or impossible to calculate exactly. In some of these cases, a procedure involving chance,called a Monte Carlo p rocedure , can be used to provide such an estimate. Example 2.2 In this example we show how simulation can be used to estimate areas of plane flgures. Suppose that we program our computer to provide a pair(x;y) or numbers, each chosen independently at random from the interval [0 ;1]. Then we can interpret this pair ( x;y) as the coordinates of a point chosen at random from the unit square. Events are subsets of the unit square. Our experience withExample 2.1 suggests that the point is equally likely to fall in subsets of equal area.Since the total area of the square is 1, the probability of the point falling in a speciflcsubsetEof the unit square should be equal to its area. Thus, we can estimate the area of any subset of the unit square by estimating the probability that a pointchosen at random from this square falls in the subset. We can use this method to estimate the area of the region Eunder the curve y=x 2in the unit square (see Figure 2.2). We choose a large number of points ( x;y) at random and record what fraction of them fall in the region E=f(x;y):y•x2g. The program MonteCarlo will carry out this experiment for us. Running this program for 10,000 experiments gives an estimate of .325 (see Figure 2.3). From these experiments we would estimate the area to be about 1/3. Of course, 2.1. SIMULATION OF CONTINUOUS PROBABILITIES 43 1x1y y = x2 E Figure 2.2: Area under y=x2: for this simple region we can flnd the exact area by calculus. In fact, Area ofE=Z1 0x2dx=1 3: We have remarked in Chapter 1 that, when we simulate an experiment of this type ntimes to estimate a probability, we can expect the answer to be in error by at most 1=pnat least 95 percent of the time. For 10,000 experiments we can expect an accuracy of 0.01, and our simulation did achieve this accuracy. This same argument works for any region Eof the unit square. For example, supposeEis the circle with center (1 =2;1=2) and radius 1/2. Then the probability that our random point ( x;y) lies inside the circle is equal to the area of the circle, that is, P(E)=…‡1 2·2 =… 4: If we did not know the value of …, we could estimate the value by performing this experiment a large number of times! 2 The above example is not the only way of estimating the value of …by a chance experiment. Here is another way, discovered by Bufion.1 1G. L. Bufion, in \Essai d’Arithm¶ etique Morale," Oeuvres Complµ etes de Bufion avec Supple- ments, tome iv, ed. Dum¶ enil (Paris, 1836). 44 CHAPTER 2. CONTINUOUS PROBABILITY DENSITIES 111000 trials Estimate of area is .325y = x2 E Figure 2.3: Computing the area by simulation. Bufion’s Needle Example 2.3 Suppose that we take a card table and draw across the top surface a set of parallel lines a unit distance apart. We then drop a common needle ofunit length at random on this surface and observe whether or not the needle liesacross one of the lines. We can describe the possible outcomes of this experimentby coordinates as follows: Let dbe the distance from the center of the needle to the nearest line. Next, let Lbe the line determined by the needle, and deflne µas the acute angle that the line Lmakes with the set of parallel lines. (The reader should certainly be wary of this description of the sample space. We are attempting tocoordinatize a set of line segments. To see why one must be careful in the choiceof coordinates, see Example 2.6.) Using this description, we have 0 •d•1=2, and 0•µ•…=2. Moreover, we see that the needle lies across the nearest line if and only if the hypotenuse of the triangle (see Figure 2.4) is less than half the length ofthe needle, that is, d sinµ<1 2: Now we assume that when the needle drops, the pair ( µ;d) is chosen at random from the rectangle 0 •µ•…=2, 0•d•1=2. We observe whether the needle lies across the nearest line (i.e., whether d•(1=2) sinµ). The probability of this event Eis the fraction of the area of the rectangle which lies inside E(see Figure 2.5). 2.1. SIMULATION OF CONTINUOUS PROBABILITIES 45 d1/2 θ Figure 2.4: Bufion’s experiment. θ01/2 0d π/2E Figure 2.5: Set Eof pairs (µ;d) withd<1 2sinµ. Now the area of the rectangle is …=4, while the area of Eis Area =Z…=2 01 2sinµdµ=1 2: Hence, we get P(E)=1=2 …=4=2 …: The program BufionsNeedle simulates this experiment. In Figure 2.6, we show the position of every 100th needle in a run of the program in which 10,000 needleswere \dropped." Our flnal estimate for …is 3.139. While this was within 0.003 of the true value for …we had no right to expect such accuracy. The reason for this is that our simulation estimates P(E). While we can expect this estimate to be in error by at most 0.001, a small error in P(E) gets magnifled when we use this to compute…=2=P(E). Perlman and Wichura, in their article \Sharpening Bufion’s 46 CHAPTER 2. CONTINUOUS PROBABILITY DENSITIES 0.005.00 0.501.001.502.002.503.003.504.004.505.0010000 3.139 Figure 2.6: Simulation of Bufion’s needle experiment. Needle,"2show that we can expect to have an error of not more than 5 =pnabout 95 percent of the time. Here nis the number of needles dropped. Thus for 10,000 needles we should expect an error of no more than 0.05, and that was the case here.We see that a large number of experiments is necessary to get a decent estimate for…. 2 In each of our examples so far, events of the same size are equally likely. Here is an example where they are not. We will see many other such examples later. Example 2.4 Suppose that we choose two random real numbers in [0 ;1] and add them together. Let Xbe the sum. How is Xdistributed? To help understand the answer to this question, we can use the program Are- abargraph . This program produces a bar graph with the property that on each interval, the area, rather than the height, of the bar is equal to the fraction of out- comes that fell in the corresponding interval. We have carried out this experiment1000 times; the data is shown in Figure 2.7. It appears that the function deflnedby f(x)=‰x; if 0•x•1, 2¡x;if 1<x•2 flts the data very well. (It is shown in the flgure.) In the next section, we will see that this function is the \right" function. By this we mean that if aandbare any two real numbers between 0 and 2, with a•b, then we can use this function to calculate the probability that a•X•b. To understand how this calculation might be performed, we again consider Figure 2.7. Because of the way the barswere constructed, the sum of the areas of the bars corresponding to the interval 2M. D. Perlman and M. J. Wichura, \Sharpening Bufion’s Needle," The American Statistician, vol. 29, no. 4 (1975), pp. 157{163. 2.1. SIMULATION OF CONTINUOUS PROBABILITIES 47 0 0.5 1 1.5 200.20.40.60.81 Figure 2.7: Sum of two random numbers. [a;b] approximates the probability that a•X•b. But the sum of the areas of these bars also approximates the integral Zb af(x)dx : This suggests that for an experiment with a continuum of possible outcomes, if we flnd a function with the above property, then we will be able to use it to calculateprobabilities. In the next section, we will show how to determine the functionf(x). 2 Example 2.5 Suppose that we choose 100 random numbers in [0 ;1], and let X represent their sum. How is Xdistributed? We have carried out this experiment 10000 times; the results are shown in Figure 2.8. It is not so clear what functionflts the bars in this case. It turns out that the type of function which does the jobis called a normal density function. This type of function is sometimes referred to as a \bell-shaped" curve. It is among the most important functions in the subjectof probability, and will be formally deflned in Section 5.2 of Chapter 4.3. 2 Our last example explores the fundamental question of how probabilities are assigned. Bertrand’s Paradox Example 2.6 A chord of a circle is a line segment both of whose endpoints lie on the circle. Suppose that a chord is drawn at random in a unit circle. What is the probability that its length exceedsp 3? Our answer will depend on what we mean by random, which will depend, in turn, on what we choose for coordinates. The sample space › is the set of all possiblechords in the circle. To flnd coordinates for these chords, we flrst introduce a 48 CHAPTER 2. CONTINUOUS PROBABILITY DENSITIES 40 45 50 55 6000.020.040.060.080.10.120.14 Figure 2.8: Sum of 100 random numbers. xy A BM θ βα Figure 2.9: Random chord. rectangular coordinate system with origin at the center of the circle (see Figure 2.9). We note that a chord of a circle is perpendicular to the radial line containing themidpoint of the chord. We can describe each chord by giving: 1. The rectangular coordinates ( x;y) of the midpoint M,o r 2. The polar coordinates ( r;µ) of the midpoint M,o r 3. The polar coordinates (1 ;fi) and (1;fl) of the endpoints AandB. In each case we shall interpret at random to mean: choose these coordinates at random. We can easily estimate this probability by computer simulation. In programming this simulation, it is convenient to include certain simpliflcations, which we describein turn: 2.1. SIMULATION OF CONTINUOUS PROBABILITIES 49 1. To simulate this case, we choose values for xandyfrom [¡1;1] at random. Then we check whether x2+y2•1. If not, the point M=(x;y) lies outside the circle and cannot be the midpoint of any chord, and we ignore it. Oth-erwise,Mlies inside the circle and is the midpoint of a unique chord, whose lengthLis given by the formula: L=2p 1¡(x2+y2): 2. To simulate this case, we take account of the fact that any rotation of the circle does not change the length of the chord, so we might as well assume inadvance that the chord is horizontal. Then we choose rfrom [¡1;1] at random, and compute the length of the resulting chord with midpoint ( r;…=2) by the formula: L=2p 1¡r2: 3. To simulate this case, we assume that one endpoint, say B, lies at (1;0) (i.e., thatfl= 0). Then we choose a value for fifrom [0;2…] at random and compute the length of the resulting chord, using the Law of Cosines, by the formula: L=p 2¡2 cosfi: The program BertrandsParadox carries out this simulation. Running this program produces the results shown in Figure 2.10. In the flrst circle in this flgure,a smaller circle has been drawn. Those chords which intersect this smaller circlehave length at leastp 3. In the second circle in the flgure, the vertical line intersects all chords of length at leastp 3. In the third circle, again the vertical line intersects all chords of length at leastp 3. In each case we run the experiment a large number of times and record the fraction of these lengths that exceedp 3. We have printed the results of every 100th trial up to 10,000 trials. It is interesting to observe that these fractions are notthe same in the three cases; they depend on our choice of coordinates. This phenomenon was flrst observed byBertrand, and is now known as Bertrand’s paradox. 3It is actually not a paradox at all; it is merely a re°ection of the fact that difierent choices of coordinates will leadto difierent assignments of probabilities. Which assignment is \correct" depends onwhat application or interpretation of the model one has in mind. One can imagine a real experiment involving throwing long straws at a circle drawn on a card table. A \correct" assignment of coordinates should not dependon where the circle lies on the card table, or where the card table sits in the room.Jaynes 4has shown that the only assignment which meets this requirement is (2). In this sense, the assignment (2) is the natural, or \correct" one (see Exercise 11). We can easily see in each case what the true probabilities are if we note thatp 3 is the length of the side of an inscribed equilateral triangle. Hence, a chord has 3J. Bertrand, Calcul des Probabilit¶ es(Paris: Gauthier-Villars, 1889). 4E. T. Jaynes, \The Well-Posed Problem," in Papers on Probability, Statistics and Statistical Physics, R. D. Rosencrantz, ed. (Dordrecht: D. Reidel, 1983), pp. 133{148. 50 CHAPTER 2. CONTINUOUS PROBABILITY DENSITIES .01.0 .2 .4 .6 .81.0 .488 .227 .01.0 .2 .4 .6 .81.0 .01.0 .2 .4 .6 .81.0 .33210000 1000010000 Figure 2.10: Bertrand’s paradox. lengthL>p 3 if its midpoint has distance d<1=2 from the origin (see Figure 2.9). The following calculations determine the probability that L>p 3i ne a c ho ft h e three cases. 1.L>p 3 if(x;y) lies inside a circle of radius 1/2, which occurs with probability p=…(1=2)2 …(1)2=1 4: 2.L>p 3i fjrj<1=2, which occurs with probability 1=2¡(¡1=2) 1¡(¡1)=1 2: 3.L>p 3i f2…=3<fi< 4…=3, which occurs with probability 4…=3¡2…=3 2…¡0=1 3: We see that our simulations agree quite well with these theoretical values. 2 Historical Remarks G. L. Bufion (1707{1788) was a natural scientist in the eighteenth century who applied probability to a number of his investigations. His work is found in hismonumental 44-volume Histoire Naturelle and its supplements. 5For example, he 5G. L. Bufion, Histoire Naturelle, Generali et Particular avec le Descripti¶ on du Cabinet du Roy, 44 vols. (Paris: L‘Imprimerie Royale, 1749{1803). 2.1. SIMULATION OF CONTINUOUS PROBABILITIES 51 Length of Number of Number of Estimate Experimenter needle casts crossings for … Wolf, 1850 .8 5000 2532 3.1596 Smith, 1855 .6 3204 1218.5 3.1553De Morgan, c.1860 1.0 600 382.5 3.137Fox, 1864 .75 1030 489 3.1595Lazzerini, 1901 .83 3408 1808 3.1415929Reina, 1925 .5419 2520 869 3.1795 Table 2.1: Bufion needle experiments to estimate …. presented a number of mortality tables and used them to compute, for each age group, the expected remaining lifetime. From his table he observed: the expectedremaining lifetime of an infant of one year is 33 years, while that of a man of 21years is also approximately 33 years. Thus, a father who is not yet 21 can hope tolive longer than his one year old son, but if the father is 40, the odds are already 3to 2 that his son will outlive him. 6 Bufion wanted to show that not all probability calculations rely only on algebra, but that some rely on geometrical calculations. One such problem was his famous\needle problem" as discussed in this chapter. 7In his original formulation, Bufion describes a game in which two gamblers drop a loaf of French bread on a wide-board°oor and bet on whether or not the loaf falls across a crack in the °oor. Bufionasked: what length Lshould the bread loaf be, relative to the width Wof the °oorboards, so that the game is fair. He found the correct answer ( L=(…=4)W) using essentially the methods described in this chapter. He also considered the caseof a checkerboard °oor, but gave the wrong answer in this case. The correct answerwas given later by Laplace. The literature contains descriptions of a number of experiments that were actu- ally carried out to estimate …by this method of dropping needles. N. T. Gridgeman 8 discusses the experiments shown in Table 2.1. (The halves for the number of cross- ing comes from a compromise when it could not be decided if a crossing had actuallyoccurred.) He observes, as we have, that 10,000 casts could do no more than estab-lish the flrst decimal place of …with reasonable confldence. Gridgeman points out that, although none of the experiments used even 10,000 casts, they are surprisinglygood, and in some cases, too good. The fact that the number of casts is not alwaysa round number would suggest that the authors might have resorted to clever stop-ping to get a good answer. Gridgeman comments that Lazzerini’s estimate turnedout to agree with a well-known approximation to …, 355=1 1 3=3:1415929, discov- ered by the flfth-century Chinese mathematician, Tsu Ch’ungchih. Gridgeman saysthat he did not have Lazzerini’s original report, and while waiting for it (knowing 6G. L. Bufion, \Essai d’Arithm¶ etique Morale," p. 301. 7ibid., pp. 277{278. 8N. T. Gridgeman, \Geometric Probability and the Number …"Scripta Mathematika, vol. 25, no. 3, (1960), pp. 183{195. 52 CHAPTER 2. CONTINUOUS PROBABILITY DENSITIES only the needle crossed a line 1808 times in 3408 casts) deduced that the length of the needle must have been 5/6. He calculated this from Bufion’s formula, assuming…= 355=113: L=…P(E) 2=1 2µ355 113¶µ1808 3408¶ =5 6=:8333: Even with careful planning one would have to be extremely lucky to be able to stop so cleverly. The second author likes to trace his interest in probability theory to the Chicago World’s Fair of 1933 where he observed a mechanical device dropping needles anddisplaying the ever-changing estimates for the value of …. (The flrst author likes to trace his interest in probability theory to the second author.) Exercises *1In the spinner problem (see Example 2.1) divide the unit circumference into three arcs of length 1/2, 1/3, and 1/6. Write a program to simulate thespinner experiment 1000 times and print out what fraction of the outcomesfall in each of the three arcs. Now plot a bar graph whose bars have width 1/2,1/3, and 1/6, and areas equal to the corresponding fractions as determinedby your simulation. Show that the heights of the bars are all nearly the same. 2Do the same as in Exercise 1, but divide the unit circumference into flve arcs of length 1/3, 1/4, 1/5, 1/6, and 1/20. 3Alter the program MonteCarlo to estimate the area of the circle of radius 1/2 with center at (1 =2;1=2) inside the unit square by choosing 1000 points at random. Compare your results with the true value of …=4. Use your results to estimate the value of …. How accurate is your estimate? 4Alter the program MonteCarlo to estimate the area under the graph of y= sin…xinside the unit square by choosing 10,000 points at random. Now calculate the true value of this area and use your results to estimate the valueof…. How accurate is your estimate? 5Alter the program MonteCarlo to estimate the area under the graph of y=1=(x+ 1) in the unit square in the same way as in Exercise 4. Calculate the true value of this area and use your simulation results to estimate thevalue of log 2. How accurate is your estimate? 6To simulate the Bufion’s needle problem we choose independently the dis- tancedand the angle µat random, with 0 •d•1=2 and 0•µ•…=2, and check whether d•(1=2) sinµ. Doing this a large number of times, we estimate…as 2=a, whereais the fraction of the times that d•(1=2) sinµ. Write a program to estimate …by this method. Run your program several times for each of 100, 1000, and 10,000 experiments. Does the accuracy ofthe experimental approximation for …improve as the number of experiments increases? 2.1. SIMULATION OF CONTINUOUS PROBABILITIES 53 7For Bufion’s needle problem, Laplace9considered a grid with horizontal and vertical lines one unit apart. He showed that the probability that a needle of lengthL•1 crosses at least one line is p=4L¡L2 …: To simulate this experiment we choose at random an angle µbetween 0 and …=2 and independently two numbers d1andd2between 0 and L=2. (The two numbers represent the distance from the center of the needle to the nearesthorizontal and vertical line.) The needle crosses a line if either d 1•(L=2) sinµ ord2•(L=2) cosµ. We do this a large number of times and estimate …as „…=4L¡L2 a; whereais the proportion of times that the needle crosses at least one line. Write a program to estimate …by this method, run your program for 100, 1000, and 10,000 experiments, and compare your results with Bufion’s methoddescribed in Exercise 6. (Take L= 1.) 8A long needle of length Lmuch bigger than 1 is dropped on a grid with horizontal and vertical lines one unit apart. We will see (in Exercise 6.3.28)that the average number aof lines crossed is approximately a=4L …: To estimate …by simulation, pick an angle µat random between 0 and …=2 and computeLsinµ+Lcosµ. This may be used for the number of lines crossed. Repeat this many times and estimate …by „…=4L a; whereais the average number of lines crossed per experiment. Write a pro- gram to simulate this experiment and run your program for the number ofexperiments equal to 100, 1000, and 10,000. Compare your results with themethods of Laplace or Bufion for the same number of experiments. (UseL= 100.) The following exercises involve experiments in which not all outcomes are equally likely. We shall consider such experiments in detail in the next section,but we invite you to explore a few simple cases here. 9A large number of waiting time problems have an exponential distribution of outcomes. We shall see (in Section 5.2) that such outcomes are simulated bycomputing (¡1=‚) log(rnd), where ‚>0. For waiting times produced in this way, the average waiting time is 1 =‚. For example, the times spent waiting for 9P. S. Laplace, Th¶eorie Analytique des Probabilit¶ es(Paris: Courcier, 1812). 54 CHAPTER 2. CONTINUOUS PROBABILITY DENSITIES a car to pass on a hig hway, or the times between emissions of particles from a radioactive source, are simulated by a sequence of random numbers, each ofwhich is chosen by computing ( ¡1=‚) log(rnd), where 1 =‚is the average time between cars or emissions. Write a program to simulate the times betweencars when the average time between cars is 30 seconds. Have your programcompute an area bar graph for these times by breaking the time interval from0 to 120 into 24 subintervals. On the same pair of axes, plot the functionf(x)=( 1=30)e ¡(1=30)x. Does the function flt the bar graph well? 10In Exercise 9, the distribution came \out of a hat." In this problem, we will again consider an experiment whose outcomes are not equally likely. We willdetermine a function f(x) which can be used to determine the probability of certain events. Let Tbe the right triangle in the plane with vertices at the points (0;0);(1;0);and (0;1). The experiment consists of picking a point at random in the interior of T, and recording only the x-coordinate of the point. Thus, the sample space is the set [0 ;1], but the outcomes do not seem to be equally likely. We can simulate this experiment by asking a computer toreturn two random real numbers in [0 ;1], and recording the flrst of these two numbers if their sum is less than 1. Write this program and run it for 10,000trials. Then make a bar graph of the result, breaking the interval [0 ;1] into 10 intervals. Compare the bar graph with the function f(x)=2¡2x.N o w show that there is a constant csuch that the height of Tat thex-coordinate valuexisctimesf(x) for every xin [0;1]. Finally, show that Z 1 0f(x)dx=1: How might one use the function f(x) to determine the probability that the outcome is between :2 and:5? 11Here is another way to pick a chord at random on the circle of unit radius. Imagine that we have a card table whose sides are of length 100. We placecoordinate axes on the table in such a way that each side of the table is parallelto one of the axes, and so that the center of the table is the origin. We nowplace a circle of unit radius on the table so that the center of the circle is theorigin. Now pick out a point ( x 0;y0) at random in the square, and an angle µ at random in the interval ( ¡…=2;…=2). Letm= tanµ. Then the equation of the line passing through ( x0;y0) with slope mis y=y0+m(x¡x0); and the distance of this line from the center of the circle (i.e., the origin) is d=flflflfly 0¡mx0p m2+1flflflfl: We can use this distance formula to check whether the line intersects the circle (i.e., whether d<1). If so, we consider the resulting chord a random chord. 2.2. CONTINUOUS DENSITY FUNCTIONS 55 This describes an experiment of dropping a long straw at random on a table on which a circle is drawn. Write a program to simulate this experiment 10000 times and estimate the probability that the length of the chord is greater thanp 3. How does your estimate compare with the results of Example 2.6? 2.2 Continuous Density Functions In the previous section we have seen how to simulate experiments with a wholecontinuum of possible outcomes and have gained some experience in thinking aboutsuch experiments. Now we turn to the general problem of assigning probabilities tothe outcomes and events in such experiments. We shall restrict our attention hereto those experiments whose sample space can be taken as a suitably chosen subsetof the line, the plane, or some other Euclidean space. We begin with some simpleexamples. Spinners Example 2.7 The spinner experiment described in Example 2.1 has the interval [0;1) as the set of possible outcomes. We would like to construct a probability model in which each outcome is equally likely to occur. We saw that in such amodel, it is necessary to assign the probability 0 to each outcome. This does not atall mean that the probability of every event must be zero. On the contrary, if we let the random variable Xdenote the outcome, then the probability P(0•X•1) that the head of the spinner comes to rest somewhere in the circle, should be equal to 1. Also, the probability that it comes to rest in the upper half of the circle shouldbe the same as for the lower half, so that Pµ 0•X<1 2¶ =Pµ1 2•X< 1¶ =1 2: More generally, in our model, we would like the equation P(c•X<d )=d¡c to be true for every choice of candd. If we letE=[c;d], then we can write the above formula in the form P(E)=Z Ef(x)dx ; wheref(x) is the constant function with value 1. This should remind the reader of the corresponding formula in the discrete case for the probability of an event: P(E)=X !2Em(!): 56 CHAPTER 2. CONTINUOUS PROBABILITY DENSITIES 0 0.2 0.4 0.6 0.8 100.20.40.60.81 Figure 2.11: Spinner experiment. The difierence is that in the continuous case, the quantity being integrated, f(x), is not the probability of the outcome x. (However, if one uses inflnitesimals, one can consider f(x)dxas the probability of the outcome x.) In the continuous case, we will use the following convention. If the set of out- comes is a set of real numbers, then the individual outcomes will be referred toby small Roman letters such as x. If the set of outcomes is a subset of R 2, then the individual outcomes will be denoted by ( x;y). In either case, it may be more convenient to refer to an individual outcome by using !, as in Chapter 1. Figure 2.11 shows the results of 1000 spins of the spinner. The function f(x) is also shown in the flgure. The reader will note that the area under f(x) and above a given interval is approximately equal to the fraction of outcomes that fellin that interval. The function f(x) is called the density function of the random variableX. The fact that the area under f(x) and above an interval corresponds to a probability is the deflning property of density functions. A precise deflnitionof density functions will be given shortly. 2 Darts Example 2.8 A game of darts involves throwing a dart at a circular target of unit radius. Suppose we throw a dart once so that it hits the target, and we observe where it lands. To describe the possible outcomes of this experiment, it is natural to take as our sample space the set › of all the points in the target. It is convenient to describethese points by their rectangular coordinates, relative to a coordinate system withorigin at the center of the target, so that each pair ( x;y) of coordinates with x 2+y2• 1 describes a possible outcome of the experiment. Then › = f(x;y):x2+y2•1g is a subset of the Euclidean plane, and the event E=f(x;y):y>0g, for example, corresponds to the statement that the dart lands in the upper half of the target,and so forth. Unless there is reason to believe otherwise (and with experts at the 2.2. CONTINUOUS DENSITY FUNCTIONS 57 game there may well be!), it is natural to assume that the coordinates are chosen at random. (When doing this with a computer, each coordinate is chosen uniformly from the interval [ ¡1;1]. If the resulting point does not lie inside the unit circle, the point is not counted.) Then the arguments used in the preceding example showthat the probability of any elementary event, consisting of a single outcome, mustbe zero, and suggest that the probability of the event that the dart lands in anysubsetEof the target should be determined by what fraction of the target area lies inE. Thus, P(E)=area ofE area of target=area ofE …: This can be written in the form P(E)=Z Ef(x)dx ; wheref(x) is the constant function with value 1 =…. In particular, if E=f(x;y): x2+y2•a2gis the event that the dart lands within distance a<1 of the center of the target, then P(E)=…a2 …=a2: For example, the probability that the dart lies within a distance 1/2 of the center is 1/4. 2 Example 2.9 In the dart game considered above, suppose that, instead of observ- ing where the dart lands, we observe how far it lands from the center of the target. In this case, we take as our sample space the set › of all circles with centers at the center of the target. It is convenient to describe these circles by their radii, sothat each circle is identifled by its radius r,0•r•1. In this way, we may regard › as the subset [0 ;1] of the real line. What probabilities should we assign to the events Eof ›? If E=fr:0•r•ag; thenEoccurs if the dart lands within a distance aof the center, that is, within the circle of radius a, and we saw in the previous example that under our assumptions the probability of this event is given by P([0;a]) =a 2: More generally, if E=fr:a•r•bg; then by our basic assumptions, P(E)=P([a;b]) =P([0;b])¡P([0;a]) =b2¡a2 =(b¡a)(b+a) =2 (b¡a)(b+a) 2: 58 CHAPTER 2. CONTINUOUS PROBABILITY DENSITIES 00.20.40.60.8 100.511.52 0 0.2 0.4 0.6 0.8 1 2 1.5 1 0.5 0 Figure 2.12: Distribution of dart distances in 400 throws. Thus,P(E) =2(length of E)(midpoint of E). Here we see that the probability assigned to the interval Edepends not only on its length but also on its midpoint (i.e., not only on how long it is, but also on where it is). Roughly speaking, in thisexperiment, events of the form E=[a;b] are more likely if they are near the rim of the target and less likely if they are near the center. (A common experience forbeginners! The conclusion might well be difierent if the beginner is replaced by anexpert.) Again we can simulate this by computer. We divide the target area into ten concentric regions of equal thickness. The computer program Darts throwsndarts and records what fraction of the total falls in each of these concentric regions. The program Areabargraph then plots a bar graph with the area of theith bar equal to the fraction of the total falling in the ith region. Running the program for 1000 darts resulted in the bar graph of Figure 2.12. Note that here the heights of the bars are not all equal, but grow approximately linearly with r. In fact, the linear function y=2rappears to flt our bar graph quite well. This suggests that the probability that the dart falls within a distance aof the center should be given by the area under the graph of the function y=2rbetween 0 anda. This area is a 2, which agrees with the probability we have assigned above to this event. 2 Sample Space Coordinates These examples suggest that for continuous experiments of this sort we should assign probabilities for the outcomes to fall in a given interval by means of the area undera suitable function. More generally, we suppose that suitable coordinates can be introduced into the sample space ›, so that we can regard › as a subset of R n. We call such a sample space a continuous sample space. We letXbe a random variable which represents the outcome of the experiment. Such a random variable is called a continuous random variable. We then deflne a density function for Xas follows. 2.2. CONTINUOUS DENSITY FUNCTIONS 59 Density Functions of Continuous Random Variables Deflnition 2.1 LetXbe a continuous real-valued random variable. A density function forXis a real-valued function fwhich satisfles P(a•X•b)=Zb af(x)dx for alla; b2R. 2 We note that it is notthe case that all continuous real-valued random variables possess density functions. However, in this book, we will only consider continuousrandom variables for which density functions exist. In terms of the density f(x), ifEis a subset of R, then P(X2E)=Z Ef(x)dx : The notation here assumes that Eis a subset of Rfor whichR Ef(x)dxmakes sense. Example 2.10 (Example 2.7 continued) In the spinner experiment, we choose for our set of outcomes the interval 0 •x<1, and for our density function f(x)=‰1;if 0•x<1, 0;otherwise. IfEis the event that the head of the spinner falls in the upper half of the circle, thenE=fx:0•x•1=2g, and so P(E)=Z1=2 01dx=1 2: More generally, if Eis the event that the head falls in the interval [ a;b], then P(E)=Zb a1dx=b¡a: 2 Example 2.11 (Example 2.8 continued) In the flrst dart game experiment, we choose for our sample space a disc of unit radius in the plane and for our densityfunction the function f(x;y)=‰ 1=…; ifx 2+y2•1, 0; otherwise. The probability that the dart lands inside the subset Eis then given by P(E)=ZZ E1 …dxdy =1 …¢(area ofE): 2 60 CHAPTER 2. CONTINUOUS PROBABILITY DENSITIES In these two examples, the density function is constant and does not depend on the particular outcome. It is often the case that experiments in which thecoordinates are chosen at random can be described by constant density functions, and, as in Section 1.2, we call such density functions uniform orequiprobable. Not all experiments are of this type, however. Example 2.12 (Example 2.9 continued) In the second dart game experiment, we choose for our sample space the unit interval on the real line and for our densitythe function f(r)=‰ 2r;if 0<r< 1, 0;otherwise. Then the probability that the dart lands at distance r,a•r•b, from the center of the target is given by P([a;b]) =Z b a2rdr =b2¡a2: Here again, since the density is small when ris near 0 and large when ris near 1, we see that in this experiment the dart is more likely to land near the rim of the targetthan near the center. In terms of the bar graph of Example 2.9, the heights of thebars approximate the density function, while the areas of the bars approximate theprobabilities of the subintervals (see Figure 2.12). 2 We see in this example that, unlike the case of discrete sample spaces, the valuef(x) of the density function for the outcome xisnotthe probability of x occurring (we have seen that this probability is always 0) and in general f(x)i snot a probability at all. In this example, if we take ‚= 2 thenf(3=4 )=3=2, which being bigger than 1, cannot be a probability. Nevertheless, the density function fdoes contain all the probability information about the experiment, since the probabilities of all events can be derived from it.In particular, the probability that the outcome of the experiment falls in an interval[a;b] is given by P([a;b]) =Z b af(x)dx ; that is, by the area under the graph of the density function in the interval [ a;b]. Thus, there is a close connection here between probabilities and areas. We havebeen guided by this close connection in making up our bar graphs; each bar is chosenso that its area, and not its height, represents the relative frequency of occurrence, and hence estimates the probability of the outcome falling in the associated interval. In the language of the calculus, we can say that the probability of occurrence of an event of the form [ x;x+dx], wheredxis small, is approximately given by P([x;x+dx])…f(x)dx ; that is, by the area of the rectangle under the graph of f. Note that as dx!0, this probability!0, so that the probability P(fxg) of a single point is again 0, as in Example 2.7. 2.2. CONTINUOUS DENSITY FUNCTIONS 61 A glance at the graph of a density function tells us immediately which events of an experiment are more likely. Roughly speaking, we can say that where the densityis large the events are more likely, and where it is small the events are less likely.In Example 2.4 the density function is largest at 1. Thus, given the two intervals[0;a] and [1;1+a], whereais a small positive real number, we see that Xis more likely to take on a value in the second interval than in the flrst. Cumulative Distribution Functions of Continuous Random Variables We have seen that density functions are useful when considering continuous ran- dom variables. There is another kind of function, closely related to these densityfunctions, which is also of great importance. These functions are called cumulative distribution functions. Deflnition 2.2 LetXbe a continuous real-valued random variable. Then the cumulative distribution function of Xis deflned by the equation F X(x)=P(X•x): 2 IfXis a continuous real-valued random variable which possesses a density function, then it also has a cumulative distribution function, and the following theorem showsthat the two functions are related in a very nice way. Theorem 2.1 LetXbe a continuous real-valued random variable with density functionf(x). Then the function deflned by F(x)=Z x ¡1f(t)dt is the cumulative distribution function of X. Furthermore, we have d dxF(x)=f(x): Proof. By deflnition, F(x)=P(X•x): LetE=(¡1;x]. Then P(X•x)=P(X2E); which equalsZx ¡1f(t)dt : Applying the Fundamental Theorem of Calculus to the flrst equation in the statement of the theorem yields the second statement. 2 62 CHAPTER 2. CONTINUOUS PROBABILITY DENSITIES -1 -0.5 0 0.5 1 1.5 20.250.50.7511.251.51.752 f (x) F (x)X X Figure 2.13: Distribution and density for X=U2. In many experiments, the density function of the relevant random variable is easy to write down. However, it is quite often the case that the cumulative distributionfunction is easier to obtain than the density function. (Of course, once we havethe cumulative distribution function, the density function can easily be obtained bydifierentiation, as the above theorem shows.) We now give some examples whichexhibit this phenomenon. Example 2.13 A real number is chosen at random from [0 ;1] with uniform prob- ability, and then this number is squared. Let Xrepresent the result. What is the cumulative distribution function of X? What is the density of X? We begin by letting Urepresent the chosen real number. Then X=U 2.I f 0•x•1, then we have FX(x)=P(X•x) =P(U2•x) =P(U•px) =px: It is clear that Xalways takes on a value between 0 and 1, so the cumulative distribution function of Xis given by FX(x)=8 < :0; ifx•0;px;if 0•x•1; 1; ifx‚1: From this we easily calculate that the density function of Xis fX(x)=8 < :0; ifx•0; 1=(2px);if 0•x•1; 0; ifx>1: Note thatFX(x) is continuous, but fX(x) is not. (See Figure 2.13.) 2 2.2. CONTINUOUS DENSITY FUNCTIONS 63 0.2 0.4 0.6 0.8 10.20.40.60.81 E.8 Figure 2.14: Calculation of distribution function for Example 2.14. When referring to a continuous random variable X(say with a uniform density function), it is customary to say that \ Xis uniformly distributed on the interval [a;b]." It is also customary to refer to the cumulative distribution function of Xas the distribution function of X. Thus, the word \distribution" is being used in sev- eral difierent ways in the subject of probability. (Recall that it also has a meaningwhen discussing discrete random variables.) When referring to the cumulative dis-tribution function of a continuous random variable X, we will always use the word \cumulative" as a modifler, unless the use of another modifler, such as \normal" or\exponential," makes it clear. Since the phrase \uniformly densitied on the interval[a;b]" is not acceptable English, we will have to say \uniformly distributed" instead. Example 2.14 In Example 2.4, we considered a random variable, deflned to be the sum of two random real numbers chosen uniformly from [0 ;1]. Let the random variablesXandYdenote the two chosen real numbers. Deflne Z=X+Y.W e will now derive expressions for the cumulative distribution function and the densityfunction of Z. Here we take for our sample space › the unit square in R 2with uniform density. A point!2› then consists of a pair ( x;y) of numbers chosen at random. Then 0•Z•2. LetEzdenote the event that Z•z. In Figure 2.14, we show the set E:8. The event Ez, for anyzbetween 0 and 1, looks very similar to the shaded set in the flgure. For 1 <z•2, the setEzlooks like the unit square with a triangle removed from the upper right-hand corner. We can now calculate the probabilitydistribution F ZofZ;i ti sg i v e nb y FZ(z)=P(Z•z) = Area of Ez 64 CHAPTER 2. CONTINUOUS PROBABILITY DENSITIES -1 1 2 30.20.40.60.81 -1 1 2 30.20.40.60.81FZ (z) f (z)Z Figure 2.15: Distribution and density functions for Example 2.14. 1 E Z Figure 2.16: Calculation of Fzfor Example 2.15. =8 >>< >>:0; ifz<0; (1=2)z2; if 0•z•1; 1¡(1=2)(2¡z)2;if 1•z•2; 1; if 2<z: The density function is obtained by difierentiating this function: fZ(z)=8 >>< >>:0; ifz<0; z; if 0•z•1; 2¡z;if 1•z•2; 0; if 2<z: The reader is referred to Figure 2.15 for the graphs of these functions. 2 Example 2.15 In the dart game described in Example 2.8, what is the distribution of the distance of the dart from the center of the target? What is its density? Here, as before, our sample space › is the unit disk in R2, with coordinates (X;Y ). LetZ=p X2+Y2represent the distance from the center of the target. Let 2.2. CONTINUOUS DENSITY FUNCTIONS 65 -1 -0.5 0.5 1 1.5 20.20.40.60.81 -1 -0.5 0 0.5 1 1.5 20.250.50.7511.251.51.752 F (z)Z f (z)Z Figure 2.17: Distribution and density for Z=p X2+Y2. Ebe the eventfZ•zg. Then the distribution function FZofZ(see Figure 2.16) is given by FZ(z)=P(Z•z) =Area ofE Area of target: Thus, we easily compute that FZ(z)=8 < :0;ifz•0; z2;if 0•z•1; 1;ifz>1: The density fZ(z) is given again by the derivative of FZ(z): fZ(z)=8 < :0;ifz•0; 2z;if 0•z•1; 0;ifz>1: The reader is referred to Figure 2.17 for the graphs of these functions. We can verify this result by simulation, as follows: We choose values for Xand Yat random from [0 ;1] with uniform distribution, calculate Z=p X2+Y2, check whether 0•Z•1, and present the results in a bar graph (see Figure 2.18). 2 Example 2.16 Suppose Mr. and Mrs. Lockhorn agree to meet at the Hanover Inn between 5:00 and 6:00 P.M. on Tuesday. Suppose each arrives at a time between 5:00 and 6:00 chosen at random with uniform probability. What is the distributionfunction for the length of time that the flrst to arrive has to wait for the other?What is the density function? Here again we can take the unit square to represent the sample space, and ( X;Y ) as the arrival times (after 5:00 P.M.) for the Lockhorns. Let Z=jX¡Yj. Then we haveFX(x)=xandFY(y)=y. Moreover (see Figure 2.19), FZ(z)=P(Z•z) =P(jX¡Yj•z) = Area of E: 66 CHAPTER 2. CONTINUOUS PROBABILITY DENSITIES 0 0.2 0.4 0.6 0.8 100.511.52 Figure 2.18: Simulation results for Example 2.15. Thus, we have FZ(z)=8 < :0; ifz•0; 1¡(1¡z)2;if 0•z•1; 1; ifz>1: The density fZ(z) is again obtained by difierentiation: fZ(z)=8 < :0; ifz•0; 2(1¡z);if 0•z•1; 0; ifz>1: 2 Example 2.17 There are many occasions where we observe a sequence of occur- rences which occur at \random" times. For example, we might be observing emis-sions of a radioactive isotope, or cars passing a milepost on a hig hway, or light bulbs burning out. In such cases, we might deflne a random variable Xto denote the time between successive occurrences. Clearly, Xis a continuous random variable whose range consists of the non-negative real numbers. It is often the case that we canmodelXby using the exponential density . This density is given by the formula f(t)=‰ ‚e ¡‚t;ift‚0; 0; ift<0: The number ‚is a non-negative real number, and represents the reciprocal of the average value of X. (This will be shown in Chapter 6.) Thus, if the average time between occurrences is 30 minutes, then ‚=1=30. A graph of this density function with‚=1=30 is shown in Figure 2.20. One can see from the flgure that even though the average value is 30, occasionally much larger values are taken on by X. Suppose that we have bought a computer that contains a Warp 9 hard drive. The salesperson says that the average time between breakdowns of this type of harddrive is 30 months. It is often assumed that the length of time between breakdowns 2.2. CONTINUOUS DENSITY FUNCTIONS 67 E1 - z 1 - z1 - z1 - z E Figure 2.19: Calculation of FZ. 20 40 60 80 100 1200.0050.010.0150.020.0250.03 f (t) = (1/30) e - (1/30) t Figure 2.20: Exponential density with ‚=1=30. 68 CHAPTER 2. CONTINUOUS PROBABILITY DENSITIES 0 20 40 60 80 10000.0050.010.0150.020.0250.03 Figure 2.21: Residual lifespan of a hard drive. is distributed according to the exponential density. We will assume that this model applies here, with ‚=1=30. Now suppose that we have been operating our computer for 15 months. We assume that the original hard drive is still running. We ask how long we shouldexpect the hard drive to continue to run. One could reasonably expect that thehard drive will run, on the average, another 15 months. (One might also guessthat it will run more than 15 months, since the fact that it has already run for 15months implies that we don’t have a lemon.) The time which we have to wait isa new random variable, which we will call Y. Obviously, Y=X¡15. We can write a computer program to produce a sequence of simulated Y-values. To do this, we flrst produce a sequence of X’s, and discard those values which are less than or equal to 15 (these values correspond to the cases where the hard drive has quitrunning before 15 months). To simulate a value of X, we compute the value of the expression ‡ ¡1 ‚· log(rnd); whererndrepresents a random real number between 0 and 1. (That this expression has the exponential density will be shown in Chapter 4.3.) Figure 2.21 shows anarea bar graph of 10,000 simulated Y-values. The average value of Yin this simulation is 29.74, which is closer to the original average life span of 30 months than to the value of 15 months which was guessedabove. Also, the distribution of Yis seen to be close to the distribution of X. It is in fact the case that XandYhave the same distribution. This property is called the memoryless property , because the amount of time that we have to wait for an occurrence does not depend on how long we have already waited. The onlycontinuous density function with this property is the exponential density. 2 2.2. CONTINUOUS DENSITY FUNCTIONS 69 Assignment of Probabilities A fundamental question in practice is: How shall we choose the probability density function in describing any given experiment? The answer depends to a great extenton the amount and kind of information available to us about the experiment. Insome cases, we can see that the outcomes are equally likely. In some cases, we cansee that the experiment resembles another already described by a known density.In some cases, we can run the experiment a large number of times and make areasonable guess at the density on the basis of the observed distribution of outcomes,as we did in Chapter 1. In general, the problem of choosing the right density functionfor a given experiment is a central problem for the experimenter and is not alwayseasy to solve (see Example 2.6). We shall not examine this question in detail herebut instead shall assume that the right density is already known for each of theexperiments under study. The introduction of suitable coordinates to describe a continuous sample space, and a suitable density to describe its probabilities, is not always so obvious, as ourflnal example shows. Inflnite Tree Example 2.18 Consider an experiment in which a fair coin is tossed repeatedly, without stopping. We have seen in Example 1.6 that, for a coin tossed ntimes, the natural sample space is a binary tree with nstages. On this evidence we expect that for a coin tossed repeatedly, the natural sample space is a binary tree with aninflnite number of stages, as indicated in Figure 2.22. It is surprising to learn that, although the n-stage tree is obviously a flnite sample space, the unlimited tree can be described as a continuous sample space. To see howthis comes about, let us agree that a typical outcome of the unlimited coin tossingexperiment can be described by a sequence of the form !=fHHTHTTH :::g. If we write 1 for H and 0 for T, then !=f1101001 :::g. In this way, each outcome is described by a sequence of 0’s and 1’s. Now suppose we think of this sequence of 0’s and 1’s as the binary expansion of some real number x=:1101001¢¢¢lying between 0 and 1. (A binary expansion is like a decimal expansion but based on 2 instead of 10.) Then each outcome isdescribed by a value of x, and in this way xbecomes a coordinate for the sample space, taking on all real values between 0 and 1. (We note that it is possible fortwo difierent sequences to correspond to the same real number; for example, thesequencesfTHHHHH :::gandfHTTTTT :::gboth correspond to the real number 1=2. We will not concern ourselves with this apparent problem here.) What probabilities should be assigned to the events of this sample space? Con- sider, for example, the event Econsisting of all outcomes for which the flrst toss comes up heads and the second tails. Every such outcome has the form :10⁄⁄⁄⁄¢¢¢ , where⁄can be either 0 or 1. Now if xis our real-valued coordinate, then the value ofxfor every such outcome must lie between 1 =2=:10000¢¢¢and 3=4=:11000¢¢¢, and moreover, every value of xbetween 1/2 and 3/4 has a binary expansion of the 70 CHAPTER 2. CONTINUOUS PROBABILITY DENSITIES 01 1 1 1 1 1 11 1 00000 00 0 (start)1 1111 1 1 00 0000 0 Figure 2.22: Tree for inflnite number of tosses of a coin. form:10⁄⁄⁄⁄¢¢¢ . This means that !2Eif and only if 1 =2•x<3=4, and in this way we see that we can describe Eby the interval [1 =2;3=4). More generally, every event consisting of outcomes for which the results of the flrst ntosses are prescribed is described by a binary interval of the form [ k=2n;(k+1 )=2n). We have already seen in Section 1.2 that in the experiment involving ntosses, the probability of any one outcome must be exactly 1 =2n. It follows that in the unlimited toss experiment, the probability of any event consisting of outcomes forwhich the results of the flrst ntosses are prescribed must also be 1 =2 n. But 1=2nis exactly the length of the interval of x-values describing E! Thus we see that, just as with the spinner experiment, the probability of an event Eis determined by what fraction of the unit interval lies in E. Consider again the statement: The probability is 1/2 that a fair coin will turn up heads when tossed. We have suggested that one interpretation of this statement isthat if we toss the coin indeflnitely the proportion of heads will approach 1/2. Thatis, in our correspondence with binary sequences we expect to get a binary sequencewith the proportion of 1’s tending to 1/2. The event Eof binary sequences for which this is true is a proper subset of the set of all possible binary sequences. It doesnot contain, for example, the sequence 011011011 :::(i.e., (011) repeated again and again). The event Eis actually a very complicated subset of the binary sequences, but its probability can be determined as a limit of probabilities for events with aflnite number of outcomes whose probabilities are given by flnite tree measures.When the probability of Eis computed in this way, its value is found to be 1. This remarkable result is known as the Strong Law of Large Numbers (orLaw of Averages ) and is one justiflcation for our frequency concept of probability. We shall prove a weak form of this theorem in Chapter 8. 2 2.2. CONTINUOUS DENSITY FUNCTIONS 71 Exercises 1Suppose you choose at random a real number Xfrom the interval [2 ;10]. (a) Find the density function f(x) and the probability of an event Efor this experiment, where Eis a subinterval [ a;b]o f[ 2;10]. (b) From (a), flnd the probability that X> 5, that 5<X< 7, and that X2¡12X+3 5>0. 2Suppose you choose a real number Xfrom the interval [2 ;10] with a density function of the form f(x)=Cx; whereCis a constant. (a) FindC. (b) FindP(E), whereE=[a;b] is a subinterval of [2 ;10]. (c) FindP(X> 5),P(X< 7), andP(X2¡12X+3 5>0). 3Same as Exercise 2, but suppose f(x)=C x: 4Suppose you throw a dart at a circular target of radius 10 inches. Assuming that you hit the target and that the coordinates of the outcomes are chosenat random, flnd the probability that the dart falls (a) within 2 inches of the center. (b) within 2 inches of the rim. (c) within the flrst quadrant of the target. (d) within the flrst quadrant and within 2 inches of the rim. 5Suppose you are watching a radioactive source that emits particles at a rate described by the exponential density f(t)=‚e ¡‚t; where‚= 1, so that the probability P(0;T) that a particle will appear in the nextTseconds isP([0;T]) =RT 0‚e¡‚tdt. Find the probability that a particle (not necessarily the flrst) will appear (a) within the next second. (b) within the next 3 seconds. (c) between 3 and 4 seconds from now. (d) after 4 seconds from now. 72 CHAPTER 2. CONTINUOUS PROBABILITY DENSITIES 6Assume that a new light bulb will burn out after thours, where tis chosen from [0;1) with an exponential density f(t)=‚e¡‚t: In this context, ‚is often called the failure rate of the bulb. (a) Assume that ‚=0:01, and flnd the probability that the bulb will not burn out before Thours. This probability is often called the reliability of the bulb. (b) For what Tis the reliability of the bulb = 1 =2? 7Choose a number Bat random from the interval [0 ;1] with uniform density. Find the probability that (a) 1=3<B< 2=3. (b)jB¡1=2j•1=4. (c)B< 1=4o r1¡B< 1=4. (d) 3B2<B. 8Choose independently two numbers BandCat random from the interval [0 ;1] with uniform density. Note that the point ( B;C) is then chosen at random in the unit square. Find the probability that (a)B+C< 1=2. (b)BC< 1=2. (c)jB¡Cj<1=2. (d) maxfB;Cg<1=2. (e) minfB;Cg<1=2. (f)B< 1=2 and 1¡C< 1=2. (g) conditions (c) and (f) both hold. (h)B2+C2•1=2. (i) (B¡1=2)2+(C¡1=2)2<1=4. 9Suppose that we have a sequence of occurrences. We assume that the time Xbetween occurrences is exponentially distributed with ‚=1=10, so on the average, there is one occurrence every 10 minutes (see Example 2.17). Youcome upon this system at time 100, and wait until the next occurrence. Makea conjecture concerning how long, on the average, you will have to wait. Writea program to see if your conjecture is right. 10As in Exercise 9, assume that we have a sequence of occurrences, but now assume that the time Xbetween occurrences is uniformly distributed between 5 and 15. As before, you come upon this system at time 100, and wait untilthe next occurrence. Make a conjecture concerning how long, on the average,you will have to wait. Write a program to see if your conjecture is right. 2.2. CONTINUOUS DENSITY FUNCTIONS 73 11For examples such as those in Exercises 9 and 10, it might seem that at least you should not have to wait on average more than 10 minutes if the average time between occurrences is 10 minutes. Alas, even this is not true. To seewhy, consider the following assumption about the times between occurrences.Assume that the time between occurrences is 3 minutes with probability .9and 73 minutes with probability .1. Show by simulation that the average timebetween occurrences is 10 minutes, but that if you come upon this system attime 100, your average waiting time is more than 10 minutes. 12Take a stick of unit length and break it into three pieces, choosing the break points at random. (The break points are assumed to be chosen simultane-ously.) What is the probability that the three pieces can be used to form atriangle? Hint: The sum of the lengths of any two pieces must exceed the length of the third, so each piece must have length <1=2. Now use Exer- cise 8(g). 13Take a stick of unit length and break it into two pieces, choosing the break point at random. Now break the longer of the two pieces at a random point.What is the probability that the three pieces can be used to form a triangle? 14Choose independently two numbers BandCat random from the interval [¡1;1] with uniform distribution, and consider the quadratic equation x 2+Bx+C=0: Find the probability that the roots of this equation (a) are both real. (b) are both positive. Hints : (a) requires 0•B2¡4C, (b) requires 0•B2¡4C,B•0, 0•C. 15At the Tunbridge World’s Fair, a coin toss game works as follows. Quarters are tossed onto a checkerboard. The management keeps all the quarters, butfor each quarter landing entirely within one square of the checkerboard themanagement pays a dollar. Assume that the edge of each square is twice thediameter of a quarter, and that the outcomes are described by coordinateschosen at random. Is this a fair game? 16Three points are chosen at random on a circle of unit circumference. What is the probability that the triangle deflned by these points as vertices has threeacute angles? Hint: One of the angles is obtuse if and only if all three points lie in the same semicircle. Take the circumference as the interval [0 ;1]. Take one point at 0 and the others at BandC. 17Write a program to choose a random number Xin the interval [2 ;10] 1000 times and record what fraction of the outcomes satisfy X> 5, what fraction satisfy 5<X< 7, and what fraction satisfy x 2¡12x+3 5>0. How do these results compare with Exercise 1? 74 CHAPTER 2. CONTINUOUS PROBABILITY DENSITIES 18Write a program to choose a point ( X;Y )at random in a square of side 20 inches, doing this 10,000 times, and recording what fraction of the outcomesfall within 19 inches of the center; of these, what fraction fall between 8 and 10inches of the center; and, of these, what fraction fall within the flrst quadrantof the square. How do these results compare with those of Exercise 4? 19Write a program to simulate the problem describe in Exercise 7 (see Exer- cise 17). How do the simulation results compare with the results of Exercise 7? 20Write a program to simulate the problem described in Exercise 12. 21Write a program to simulate the problem described in Exercise 16. 22Write a program to carry out the following experiment. A coin is tossed 100 times and the number of heads that turn up is recorded. This experimentis then repeated 1000 times. Have your program plot a bar graph for theproportion of the 1000 experiments in which the number of heads is n, for eachnin the interval [35 ;65]. Does the bar graph look as though it can be flt with a normal curve? 23Write a program that picks a random number between 0 and 1 and computes the negative of its logarithm. Repeat this process a large number of times andplot a bar graph to give the number of times that the outcome falls in eachinterval of length 0.1 in [0 ;10]. On this bar graph plot a graph of the density f(x)=e ¡x. How well does this density flt your graph? Chapter 3 Combinatorics 3.1 Permutations Many problems in probability theory require that we count the number of ways that a particular event can occur. For this, we study the topics of permutations and combinations. We consider permutations in this section and combinations in the next section. Before discussing permutations, it is useful to introduce a general counting tech- nique that will enable us to solve a variety of counting problems, including theproblem of counting the number of possible permutations of nobjects. Counting Problems Consider an experiment that takes place in several stages and is such that the number of outcomes mat thenth stage is independent of the outcomes of the previous stages. The number mmay be difierent for difierent stages. We want to count the number of ways that the entire experiment can be carried out. Example 3.1 You are eating at ¶Emile’s restaurant and the waiter informs you that you have (a) two choices for appetizers: soup or juice; (b) three for the maincourse: a meat, flsh, or vegetable dish; and (c) two for dessert: ice cream or cake.How many possible choices do you have for your complete meal? We illustrate thepossible meals by a tree diagram shown in Figure 3.1. Your menu is decided in threestages|at each stage the number of possible choices does not depend on what ischosen in the previous stages: two choices at the flrst stage, three at the second,and two at the third. From the tree diagram we see that the total number of choicesis the product of the number of choices at each stage. In this examples we have2¢3¢2 = 12 possible menus. Our menu example is an example of the following general counting technique. 2 75 76 CHAPTER 3. COMBINATORICS ice cream cake ice cream cake ice cream cake ice cream cake ice cream cake ice cream cake(start)soupmeat fish vegetable juicemeat fish vegetable Figure 3.1: Tree for your menu. A Counting Technique A task is to be carried out in a sequence of rstages. There are n1ways to carry out the flrst stage; for each of these n1ways, there are n2ways to carry out the second stage; for each of these n2ways, there are n3ways to carry out the third stage, and so forth. Then the total number of ways in which the entire task can beaccomplished is given by the product N=n 1¢n2¢:::¢nr. Tree Diagrams It will often be useful to use a tree diagram when studying probabilities of events relating to experiments that take place in stages and for which we are given theprobabilities for the outcomes at each stage. For example, assume that the ownerof¶Emile’s restaurant has observed that 80 percent of his customers choose the soup for an appetizer and 20 percent choose juice. Of those who choose soup, 50 percentchoose meat, 30 percent choose flsh, and 20 percent choose the vegetable dish. Ofthose who choose juice for an appetizer, 30 percent choose meat, 40 percent chooseflsh, and 30 percent choose the vegetable dish. We can use this to estimate theprobabilities at the flrst two stages as indicated on the tree diagram of Figure 3.2. We choose for our sample space the set › of all possible paths !=! 1,!2, ...,!6through the tree. How should we assign our probability distribution? For example, what probability should we assign to the customer choosing soup and thenthe meat? If 8/10 of the customers choose soup and then 1/2 of these choose meat,a proportion 8 =10¢1=2=4=10 of the customers choose soup and then meat. This suggests choosing our probability distribution for each path through the tree to betheproduct of the probabilities at each of the stages along the path. This results in the probability measure for the sample points !indicated in Figure 3.2. (Note thatm(! 1)+¢¢¢+m(!6) = 1.) From this we see, for example, that the probability 3.1. PERMUTATIONS 77 (start)soupmeat fish vegetable juice.8 .2.2 .3.3 .4.5 .3meat fish vegetableω (ω) ω ω ω ω ω ω .4 .24 .16 .06 .08 .06m 1 2 3 4 5 6 Figure 3.2: Two-stage probability assignment. that a customer chooses meat is m(!1)+m(!4)=:46. We shall say more about these tree measures when we discuss the concept of conditional probability in Chapter 4. We return now to more counting problems. Example 3.2 We can show that there are at least two people in Columbus, Ohio, who have the same three initials. Assuming that each person has three initials,there are 26 possibilities for a person’s flrst initial, 26 for the second, and 26 for thethird. Therefore, there are 26 3=1 7;576 possible sets of initials. This number is smaller than the number of people living in Columbus, Ohio; hence, there must beat least two people with the same three initials. 2 We consider next the celebrated birthday problem|often used to show that naive intuition cannot always be trusted in probability. Birthday Problem Example 3.3 How many people do we need to have in a room to make it a favorable bet (probability of success greater than 1/2) that two people in the room will havethe same birthday? Since there are 365 possible birthdays, it is tempting to guess that we would need about 1/2 this number, or 183. You would surely win this bet. In fact, thenumber required for a favorable bet is only 23. To show this, we flnd the probabilityp rthat, in a room with rpeople, there is no duplication of birthdays; we will have a favorable bet if this probability is less than one half. 78 CHAPTER 3. COMBINATORICS Number of people Probability that all birthdays are difierent 20 .5885616 21 .556311722 .524304723 .492702824 .461655725 .4313003 Table 3.1: Birthday problem. Assume that there are 365 possible birthdays for each person (we ignore leap years). Order the people from 1 to r. For a sample point !, we choose a possible sequence of length rof birthdays each chosen as one of the 365 possible dates. There are 365 possibilities for the flrst element of the sequence, and for each ofthese choices there are 365 for the second, and so forth, making 365 rpossible sequences of birthdays. We must flnd the number of these sequences that have noduplication of birthdays. For such a sequence, we can choose any of the 365 daysfor the flrst element, then any of the remaining 364 for the second, 363 for the third,and so forth, until we make rchoices. For the rth choice, there will be 365 ¡r+1 possibilities. Hence, the total number of sequences with no duplications is 365¢364¢363¢:::¢(365¡r+1 ): Thus, assuming that each sequence is equally likely, p r=365¢364¢:::¢(365¡r+1 ) 365r: We denote the product (n)(n¡1)¢¢¢(n¡r+1 ) by (n)r(read \ndownr," or \nlowerr"). Thus, pr=(365)r (365)r: The program Birthday carries out this computation and prints the probabilities forr= 20 to 25. Running this program, we get the results shown in Table 3.1. As we asserted above, the probability for no duplication changes from greater than onehalf to less than one half as we move from 22 to 23 people. To see how unlikely it isthat we would lose our bet for larger numbers of people, we have run the programagain, printing out values from r=1 0t or= 100 in steps of 10. We see that in a room of 40 people the odds already heavily favor a duplication, and in a roomof 100 the odds are overwhelmingly in favor of a duplication. We have assumedthat birthdays are equally likely to fall on any particular day. Statistical evidencesuggests that this is not true. However, it is intuitively clear (but not easy to prove)that this makes it even more likely to have a duplication with a group of 23 people.(See Exercise 19 to flnd out what happens on planets with more or fewer than 365days per year.) 2 3.1. PERMUTATIONS 79 Number of people Probability that all birthdays are difierent 10 .8830518 20 .588561630 .293683840 .108768250 .029626460 .005877370 .000840480 .000085790 .0000062 100 .0000003 Table 3.2: Birthday problem. We now turn to the topic of permutations. Permutations Deflnition 3.1 LetAbe any flnite set. A permutation of Ais a one-to-one mapping ofAonto itself. 2 To specify a particular permutation we list the elements of Aand, under them, show where each element is sent by the one-to-one mapping. For example, if A= fa;b;cga possible permutation ¾would be ¾=µabc bca¶ : By the permutation ¾,ais sent tob,bis sent toc, andcis sent toa. The condition that the mapping be one-to-one means that no two elements of Aare sent, by the mapping, into the same element of A. We can put the elements of our set in some order and rename them 1, 2, ...,n. Then, a typical permutation of the set A=fa1;a2;a3;a4gcan be written in the form ¾=µ1234 2143¶ ; indicating that a1went toa2,a2toa1,a3toa4, anda4toa3. I fw ea l w a y sc h o o s et h et o pr o wt ob e1234 then, to prescribe the permutation, we need only give the bottom row, with the understanding that this tells us where 1goes, 2 goes, and so forth, under the mapping. When this is done, the permutationis often called a rearrangement of thenobjects 1, 2, 3, ...,n. For example, all possible permutations, or rearrangements, of the numbers A=f1;2;3gare: 123;132;213;231;312;321: It is an easy matter to count the number of possible permutations of nobjects. By our general counting principle, there are nways to assign the flrst element, for 80 CHAPTER 3. COMBINATORICS nn ! 01 11223642 45 1206 7207 50408 403209 362880 10 3628800 Table 3.3: Values of the factorial function. each of these we have n¡1 ways to assign the second object, n¡2 for the third, and so forth. This proves the following theorem. Theorem 3.1 The total number of permutations of a set Aofnelements is given byn¢(n¡1)¢(n¡2)¢:::¢1. 2 It is sometimes helpful to consider orderings of subsets of a given set. This prompts the following deflnition. Deflnition 3.2 LetAbe ann-element set, and let kbe an integer between 0 and n. Then ak-permutation of Ais an ordered listing of a subset of Aof sizek. 2 Using the same techniques as in the last theorem, the following result is easily proved. Theorem 3.2 The total number of k-permutations of a set Aofnelements is given byn¢(n¡1)¢(n¡2)¢:::¢(n¡k+ 1). 2 Factorials The number given in Theorem 3.1 is called nfactorial, and is denoted by n!. The expression 0! is deflned to be 1 to make certain formulas come out simpler. Theflrst few values of this function are shown in Table 3.3. The reader will note thatthis function grows very rapidly. The expression n! will enter into many of our calculations, and we shall need to have some estimate of its magnitude when nis large. It is clearly not practical to make exact calculations in this case. We shall instead use a result called Stirling’s formula. Before stating this formula we need a deflnition. 3.1. PERMUTATIONS 81 nn ! Approximation Ratio 1 1 .922 1.084 2 2 1.919 1.0423 6 5.836 1.0284 24 23.506 1.0215 120 118.019 1.0166 720 710.078 1.0137 5040 4980.396 1.0118 40320 39902.395 1.0109 362880 359536.873 1.009 10 3628800 3598696.619 1.008 Table 3.4: Stirling approximations to the factorial function. Deflnition 3.3 Leta nandbnbe two sequences of numbers. We say that anis asymptotically equal to bn, and write an»bn,i f lim n!1an bn=1: 2 Example 3.4 Ifan=n+pnandbn=nthen, since an=bn=1+1=pnand this ratio tends to 1 as ntends to inflnity, we have an»bn. 2 Theorem 3.3 (Stirling’s Formula) The sequence n! is asymptotically equal to nne¡np 2…n: 2 The proof of Stirling’s formula may be found in most analysis texts. Let us verify this approximation by using the computer. The program StirlingApprox- imations printsn!, the Stirling approximation, and, flnally, the ratio of these two numbers. Sample output of this program is shown in Table 3.4. Note that, whilethe ratio of the numbers is getting closer to 1, the difierence between the exactvalue and the approximation is increasing, and indeed, this difierence will tend toinflnity asntends to inflnity, even though the ratio tends to 1. (This was also true in our Example 3.4 where n+p n»n, but the difierence ispn.) Generating Random Permutations We now consider the question of generating a random permutation of the integers between 1 and n. Consider the following experiment. We start with a deck of n cards, labelled 1 through n. We choose a random card out of the deck, note its label, and put the card aside. We repeat this process until all ncards have been chosen. It is clear that each permutation of the integers from 1 to ncan occur as a sequence 82 CHAPTER 3. COMBINATORICS Number of flxed points Fraction of permutations n=1 0 n=2 0 n=3 0 0 .362 .370 .358 1 .368 .396 .358 2 .202 .164 .192 3 .052 .060 .070 4 .012 .008 .020 5 .004 .002 .002 Average number of flxed points .996 .948 1.042 Table 3.5: Fixed point distributions. of labels in this experiment, and that each sequence of labels is equally likely to occur. In our implementations of the computer algorithms, the above procedure iscalled RandomPermutation . Fixed Points There are many interesting problems that relate to properties of a permutation chosen at random from the set of all permutations of a given flnite set. For example,since a permutation is a one-to-one mapping of the set onto itself, it is interesting toask how many points are mapped onto themselves. We call such points flxed points of the mapping. Letp k(n) be the probability that a random permutation of the set f1;2;:::;ng has exactly kflxed points. We will attempt to learn something about these prob- abilities using simulation. The program FixedPoints uses the procedure Ran- domPermutation to generate random permutations and count flxed points. The program prints the proportion of times that there are kflxed points as well as the average number of flxed points. The results of this program for 500 simulations forthe casesn= 10, 20, and 30 are shown in Table 3.5. Notice the rather surprising fact that our estimates for the probabilities do not seem to depend very heavily onthe number of elements in the permutation. For example, the probability that thereare no flxed points, when n=1 0;20;or 30 is estimated to be between .35 and .37. We shall see later (see Example 3.12) that for n‚10 the exact probabilities p n(0) are, to six decimal place accuracy, equal to 1 =e…:367879. Thus, for all practi- cal purposes, after n= 10 the probability that a random permutation of the set f1;2;:::;ngdoes not depend upon n. These simulations also suggest that the av- erage number of flxed points is close to 1. It can be shown (see Example 6.8) thatthe average is exactly equal to 1 for all n. More picturesque versions of the flxed-point problem are: You have arranged the books on your book shelf in alphabetical order by author and they get returnedto your shelf at random; what is the probability that exactly kof the books end up in their correct position? (The library problem.) In a restaurant nhats are checked and they are hopelessly scrambled; what is the probability that no one gets his ownhat back? (The hat check problem.) In the Historical Remarks at the end of thissection, we give one method for solving the hat check problem exactly. Another 3.1. PERMUTATIONS 83 Date Snowfall in inches 1974 75 1975 881976 721977 1101978 851979 301980 551981 861982 511983 64 Table 3.6: Snowfall in Hanover. Y e a r 123 4567891 0 Ranking 6951 0 71382 4 Table 3.7: Ranking of total snowfall. method is given in Example 3.12. Records Here is another interesting probability problem that involves permutations. Esti- mates for the amount of measured snow in inches in Hanover, New Hampshire, inthe ten years from 1974 to 1983 are shown in Table 3.6. Suppose we have startedkeeping records in 1974. Then our flrst year’s snowfall could be considered a recordsnowfall starting from this year. A new record was established in 1975; the nextrecord was established in 1977, and there were no new records established afterthis year. Thus, in this ten-year period, there were three records established: 1974,1975, and 1977. The question that we ask is: How many records should we expectto be established in such a ten-year period? We can count the number of recordsin terms of a permutation as follows: We number the years from 1 to 10. Theactual amounts of snowfall are not important but their relative sizes are. We can,therefore, change the numbers measuring snowfalls to numbers 1 to 10 by replacingthe smallest number by 1, the next smallest by 2, and so forth. (We assume thatthere are no ties.) For our example, we obtain the data shown in Table 3.7. This gives us a permutation of the numbers from 1 to 10 and, from this per- mutation, we can read ofi the records; they are in years 1, 2, and 4. Thus we candeflne records for a permutation as follows: Deflnition 3.4 Let¾be a permutation of the set f1;2;:::;ng. Theniis arecord of¾if eitheri=1o r¾(j)<¾(i) for every j=1;:::;i¡1. 2 Now if we regard all rankings of snowfalls over an n-year period to be equally likely (and allow no ties), we can estimate the probability that there will be k records innyears as well as the average number of records by simulation. 84 CHAPTER 3. COMBINATORICS We have written a program Records that counts the number of records in ran- domly chosen permutations. We have run this program for the cases n= 10, 20, 30. Forn= 10 the average number of records is 2.968, for 20 it is 3.656, and for 30 it is 3.960. We see now that the averages increase, but very slowly. We shall seelater (see Example 6.11) that the average number is approximately log n. Since log 10 = 2:3, log 20 = 3, and log 30 = 3 :4, this is consistent with the results of our simulations. As remarked earlier, we shall be able to obtain formulas for exact results of certain problems of the above type. However, only minor changes in the problemmake this impossible. The power of simulation is that minor changes in a problemdo not make the simulation much more di–cult. (See Exercise 20 for an interestingvariation of the hat check problem.) List of Permutations Another method to solve problems that is not sensitive to small changes in theproblem is to have the computer simply list all possible permutations and count thefraction that have the desired property. The program AllPermutations produces a list of all of the permutations of n. When we try running this program, we run into a limitation on the use of the computer. The number of permutations of n increases so rapidly that even to list all permutations of 20 objects is impractical. Historical Remarks Our basic counting principle stated that if you can do one thing in rways and for each of these another thing in sways, then you can do the pair in rsways. This is such a self-evident result that you might expect that it occurred very early inmathematics. N. L. Biggs suggests that we might trace an example of this principleas follows: First, he relates a popular nursery rhyme dating back to at least 1730: As I was going to St. Ives, I met a man with seven wives,Each wife had seven sacks,Each sack had seven cats,Each cat had seven kits.Kits, cats, sacks and wives,How many were going to St. Ives? (You need our principle only if you are not clever enough to realize that you are supposed to answer one, since only the narrator is going to St. Ives; the others are going in the other direction!) He also gives a problem appearing on one of the oldest surviving mathematical manuscripts of about 1650 B.C., roughly translated as: 3.1. PERMUTATIONS 85 Houses 7 Cats 49Mice 343Wheat 2401Hekat 16807 19607 The following interpretation has been suggested: there are seven houses, each with seven cats; each cat kills seven mice; each mouse would have eaten seven headsof wheat, each of which would have produced seven hekat measures of grain. Withthis interpretation, the table answers the question of how many hekat measureswere saved by the cats’ actions. It is not clear why the writer of the table wantedto add the numbers together. 1 One of the earliest uses of factorials occurred in Euclid’s proof that there are inflnitely many prime numbers. Euclid argued that there must be a prime numberbetweennandn! + 1 as follows: n! andn! + 1 cannot have common factors. Either n! + 1 is prime or it has a proper factor. In the latter case, this factor cannot divide n! and hence must be between nandn! + 1. If this factor is not prime, then it has a factor that, by the same argument, must be bigger than n. In this way, we eventually reach a prime bigger than n, and this holds for all n. The \n!" rule for the number of permutations seems to have occurred flrst in India. Examples have been found as early as 300 B.C., and by the eleventh century the general formula seems to have been well known in India and then in the Arabcountries. The hat check problem is found in an early probability book written by de Mont- mort and flrst printed in 1708. 2It appears in the form of a game called Treize. In a simplifled version of this game considered by de Montmort one turns over cardsnumbered 1 to 13, calling out 1, 2, ...,1 3a st h e cards are examined. De Montmort asked for the probability that no card that is turned up agrees with the numbercalled out. This probability is the same as the probability that a random permutation of 13 elements has no flxed point. De Montmort solved this problem by the use of arecursion relation as follows: let w nbe the number of permutations of nelements with no flxed point (such permutations are called derangements ). Thenw1= 0 and w2=1 . Now assume that n‚3 and choose a derangement of the integers between 1 and n. Letkbe the integer in the flrst position in this derangement. By the deflnition of derangement, we have k6= 1. There are two possibilities of interest concerning the position of 1 in the derangement: either 1 is in the kth position or it is elsewhere. In the flrst case, the n¡2 remaining integers can be positioned in wn¡2ways without resulting in any flxed points. In the second case, we consider the set of integersf1;2;:::;k¡1;k+1;:::;ng. The numbers in this set must occupy the positions f2;3;:::;ngso that none of the numbers other than 1 in this set are flxed, and 1N. L. Biggs, \The Roots of Combinatorics," Historia Mathematica, vol. 6 (1979), pp. 109{136. 2P. R. de Montmort, Essay d’Analyse sur des Jeux de Hazard, 2d ed. (Paris: Quillau, 1713). 86 CHAPTER 3. COMBINATORICS also so that 1 is not in position k. The number of ways of achieving this kind of arrangement is just wn¡1. Since there are n¡1 possible values of k, we see that wn=(n¡1)wn¡1+(n¡1)wn¡2 forn‚3. One might conjecture from this last equation that the sequence fwng grows like the sequence fn!g. In fact, it is easy to prove by induction that wn=nwn¡1+(¡1)n: Thenpi=wi=i! satisfles pi¡pi¡1=(¡1)i i!: If we sum from i=2t on, and use the fact that p1= 0, we obtain pn=1 2!¡1 3!+¢¢¢+(¡1)n n!: This agrees with the flrst n+ 1 terms of the expansion for exforx=¡1 and hence for largenis approximately e¡1…:368. David remarks that this was possibly the flrst use of the exponential function in probability.3We shall see another way to derive de Montmort’s result in the next section, using a method known as theInclusion-Exclusion method. Recently, a related problem appeared in a column of Marilyn vos Savant. 4 Charles Price wrote to ask about his experience playing a certain form of solitaire,sometimes called \frustration solitaire." In this particular game, a deck of cardsis shu†ed, and then dealt out, one card at a time. As the cards are being dealt,the player counts from 1 to 13, and then starts again at 1. (Thus, each number iscounted four times.) If a number that is being counted coincides with the rank ofthe card that is being turned up, then the player loses the game. Price found thathe he rarely won and wondered how often he should win. Vos Savant remarked thatthe expected number of matches is 4 so it should be di–cult to win the game. Finding the chance of winning is a harder problem than the one that de Mont- mort solved because, when one goes through the entire deck, there are difierentpatterns for the matches that might occur. For example matches may occur for twocards of the same rank, say two aces, or for two difierent ranks, say a two and athree. A discussion of this problem can be found in Riordan. 5In this book, it is shown that asn!1 , the probability of no matches tends to 1 =e4. The original game of Treize is more di–cult to analyze than frustration solitaire. The game of Treize is played as follows. One person is chosen as dealer and theothers are players. Each player, other than the dealer, puts up a stake. The dealershu†es the cards and turns them up one at a time calling out, \Ace, two, three,..., 3F. N. David, Games, Gods and Gambling (London: Gri–n, 1962), p. 146. 4M. vos Savant, Ask Marilyn, Parade Magazine, Boston Globe , 21 August 1994. 5J. Riordan, An Introduction to Combinatorial Analysis, (New York: John Wiley & Sons, 1958). 3.1. PERMUTATIONS 87 king," just as in frustration solitaire. If the dealer goes through the 13 cards without a match he pays the players an amount equal to their stake, and the deal passes tosomeone else. If there is a match the dealer collects the players’ stakes; the playersput up new stakes, and the dealer continues through the deck, calling out, \Ace,two, three, ...." If the dealer runs out of cards he reshu†es and continues the countwhere he left ofi. He continues until there is a run of 13 without a match and thena new dealer is chosen. The question at this point is how much money can the dealer expect to win from each player. De Montmort found that if each player puts up a stake of 1, say, thenthe dealer will win approximately .801 from each player. Peter Doyle calculated the exact amount that the dealer can expect to win. The answer is: 26516072156010218582227607912734182784642120482136091446715371962089931 523113435417245543349128705414402992392516076941135000807759178185120138217687665356317385287455585936725463200947740372739557280745938434274787664965076063990538261189388143513547366316017004945507201764278828306601171079536331427343824779227098352817532990359885814136883676558331132447615331072062747416971930180664915269870408438391421790790695497603628528211590140316202120601549126920880824913325553882692055427830810368578188612087582488006809786404381185828348775425609555506628789271230482699760170011623359279330829753364219350507454026892568319388782130144270519791882/33036929133582592220117220713156071114975101149831063364072138969878007996472047088253033875258922365813230156280056211434272906256589744339716571945412290800708628984130608756130281899116735786362375606718498649135353553622197448890223267101158801016285931351979294387223277033396967797970699334758024236769498736616051840314775615603933802570709707119596964126824245501331987974705469351780938375059348885869867236484695053988868628582609905586271001318150621134407056983214740221851567706672080945865893784594327998687063341618129886304963272872548184588793530244980032242558644674104814772093410806135061350385697304897121306393704051559533731591. This is .803 to 3 decimal places. A description of the algorithm used to flnd this answer can be found on his Web page. 6A discussion of this problem and other problems can be found in Doyle et al.7 The birthday problem does not seem to have a very old history. Problems of this type were flrst discussed by von Mises.8It was made popular in the 1950s by Feller’s book.9 6P. Doyle, \Solution to Montmort’s Probleme du Treize," http://math.ucsd.edu/~doyle/. 7P. Doyle, C. Grinstead, and J. Snell, \Frustration Solitaire," UMAP Journal , vol. 16, no. 2 (1995), pp. 137-145. 8R. von Mises, \ ˜Uber Aufteilungs- und Besetzungs-Wahrscheinlichkeiten," Revue de la Facult¶ e des Sciences de l’Universit¶ e d’Istanbul, N. S. vol. 4 (1938-39), pp. 145-163. 9W. Feller, Introduction to Probability Theory and Its Applications, vol. 1, 3rd ed. (New York: 88 CHAPTER 3. COMBINATORICS Stirling presented his formula n!»p 2…n‡n e·n in his work Methodus Difierentialis published in 1730.10This approximation was used by de Moivre in establishing his celebrated central limit theorem that wewill study in Chapter 9. De Moivre himself had independently established thisapproximation, but without identifying the constant …. Having established the approximation 2B pn for the central term of the binomial distribution, where the constant Bwas deter- mined by an inflnite series, de Moivre writes: . . . my worthy and learned Friend, Mr. James Stirling, who had applied himself after me to that inquiry, found that the Quantity Bdid denote the Square-root of the Circumference of a Circle whose Radius is Unity,so that if that Circumference be called cthe Ratio of the middle Term to the Sum of all Terms will be expressed by 2 =p nc....11 Exercises 1Four people are to be arranged in a row to have their picture taken. In how many ways can this be done? 2An automobile manufacturer has four colors available for automobile exteri- ors and three for interiors. How many difierent color combinations can heproduce? 3In a digital computer, a bitis one of the integers f0,1g, and a word is any string of 32 bits. How many difierent words are possible? 4What is the probability that at least 2 of the presidents of the United States have died on the same day of the year? If you bet this has happened, wouldyou win your bet? 5There are three difierent routes connecting city A to city B. How many ways can a round trip be made from A to B and back? How many ways if it isdesired to take a difierent route on the way back? 6In arranging people around a circular table, we take into account their seats relative to each other, not the actual position of any one person. Show thatnpeople can be arranged around a circular table in ( n¡1)! ways. John Wiley & Sons, 1968). 10J. Stirling, Methodus Difierentialis, (London: Bowyer, 1730). 11A. de Moivre, The Doctrine of Chances, 3rd ed. (London: Millar, 1756). 3.1. PERMUTATIONS 89 7Five people get on an elevator that stops at flve °oors. Assuming that each has an equal probability of going to any one °oor, flnd the probability thatthey all get ofi at difierent °oors. 8A flnite set › has nelements. Show that if we count the empty set and › as subsets, there are 2 nsubsets of ›. 9A more reflned inequality for approximating n! is given by p 2…n‡n e·n e1=(12n+1)<n!<p 2…n‡n e·n e1=(12n): Write a computer program to illustrate this inequality for n= 1 to 9. 10A deck of ordinary cards is shu†ed and 13 cards are dealt. What is the probability that the last card dealt is an ace? 11There arenapplicants for the director of computing. The applicants are inter- viewed independently by each member of the three-person search committeeand ranked from 1 to n. A candidate will be hired if he or she is ranked flrst by at least two of the three interviewers. Find the probability that a candidatewill be accepted if the members of the committee really have no ability at allto judge the candidates and just rank the candidates randomly. In particular,compare this probability for the case of three candidates and the case of tencandidates. 12A symphony orchestra has in its repertoire 30 Haydn symphonies, 15 modern works, and 9 Beethoven symphonies. Its program always consists of a Haydnsymphony followed by a modern work, and then a Beethoven symphony. (a) How many difierent programs can it play? (b) How many difierent programs are there if the three pieces can be played in any order? (c) How many difierent three-piece programs are there if more than one piece from the same category can be played and they can be played inany order? 13A certain state has license plates showing three numbers and three letters. How many difierent license plates are possible (a) if the numbers must come before the letters? (b) if there is no restriction on where the letters and numbers appear? 14The door on the computer center has a lock which has flve buttons numbered from 1 to 5. The combination of numbers that opens the lock is a sequenceof flve numbers and is reset every week. (a) How many combinations are possible if every button must be used once? 90 CHAPTER 3. COMBINATORICS (b) Assume that the lock can also have combinations that require you to push two buttons simultaneously and then the other three one at a time.How many more combinations does this permit? 15A computing center has 3 processors that receive njobs, with the jobs assigned to the processors purely at random so that all of the 3 npossible assignments are equally likely. Find the probability that exactly one processor has no jobs. 16Prove that at least two people in Atlanta, Georgia, have the same initials, assuming no one has more than four initials. 17Find a formula for the probability that among a set of npeople, at least two have their birthdays in the same month of the year (assuming the months areequally likely for birthdays). 18Consider the problem of flnding the probability of more than one coincidence of birthdays in a group of npeople. These include, for example, three people with the same birthday, or two pairs of people with the same birthday, orlarger coincidences. Show how you could compute this probability, and writea computer program to carry out this computation. Use your program to flndthe smallest number of people for which it would be a favorable bet that therewould be more than one coincidence of birthdays. *19 Suppose that on planet Zorg a year has ndays, and that the lifeforms there are equally likely to have hatched on any day of the year. We would liketo estimate d, which is the minimum number of lifeforms needed so that the probability of at least two sharing a birthday exceeds 1/2. (a) In Example 3.3, it was shown that in a set of dlifeforms, the probability that no two life forms share a birthday is (n) d nd; where (n)d=(n)(n¡1)¢¢¢(n¡d+ 1). Thus, we would like to set this equal to 1/2 and solve for d. (b) Using Stirling’s Formula, show that (n)d nd»µ 1+d n¡d¶n¡d+1=2 e¡d: (c) Now take the logarithm of the right-hand expression, and use the fact that for small values of x,w eh a v e log(1 +x)»x¡x2 2: (We are implicitly using the fact that dis of smaller order of magnitude thann. We will also use this fact in part (d).) 3.1. PERMUTATIONS 91 (d) Set the expression found in part (c) equal to ¡log(2), and solve for das a function of n, thereby showing that d»p 2(log 2)n: Hint: If all three summands in the expression found in part (b) are used, one obtains a cubic equation in d. If the smallest of the three terms is thrown away, one obtains a quadratic equation in d. (e) Use a computer to calculate the exact values of dfor various values of n. Compare these values with the approximate values obtained by using the answer to part d). 20At a mathematical conference, ten participants are randomly seated around a circular table for meals. Using simulation, estimate the probability that notwo people sit next to each other at both lunch and dinner. Can you make anintelligent conjecture for the case of nparticipants when nis large? 21Modify the program AllPermutations to count the number of permutations ofnobjects that have exactly jflxed points for j= 0 , 1 , 2 , ..., n. Run your program for n= 2 to 6. Make a conjecture for the relation between the number that have 0 flxed points and the number that have exactly 1 flxedpoint. A proof of the correct conjecture can be found in Wilf. 12 22Mr. Wimply Dimple, one of London’s most prestigious watch makers, has come to Sherlock Holmes in a panic, having discovered that someone hasbeen producing and selling crude counterfeits of his best selling watch. The 16counterfeits so far discovered bear stamped numbers, all of which fall between1 and 56, and Dimple is anxious to know the extent of the forger’s work. Allpresent agree that it seems reasonable to assume that the counterfeits thusfar produced bear consecutive numbers from 1 to whatever the total numberis. \Chin up, Dimple," opines Dr. Watson. \I shouldn’t worry overly much if I were you; the Maximum Likelihood Principle, which estimates the totalnumber as precisely that which gives the highest probability for the seriesof numbers found, suggests that we guess 56 itself as the total. Thus, yourforgers are not a big operation, and we shall have them safely behind barsbefore your business sufiers signiflcantly." \Stufi, nonsense, and bother your fancy principles, Watson," counters Holmes. \Anyone can see that, of course, there must be quite a few more than 56watches|why the odds of our having discovered precisely the highest num-bered watch made are laughably negligible. A much better guess would betwice 56." (a) Show that Watson is correct that the Maximum Likelihood Principle gives 56. 12H. S. Wilf, \A Bijection in the Theory of Derangements," Mathematics Magazine, vol. 57, no. 1 (1984), pp. 37{40. 92 CHAPTER 3. COMBINATORICS (b) Write a computer program to compare Holmes’s and Watson’s guessing strategies as follows: flx a total Nand choose 16 integers randomly between 1 and N. Letmdenote the largest of these. Then Watson’s guess forNism, while Holmes’s is 2 m. See which of these is closer to N. Repeat this experiment (with Nstill flxed) a hundred or more times, and determine the proportion of times that each comes closer. Whoseseems to be the better strategy? 23Barbara Smith is interviewing candidates to be her secretary. As she inter- views the candidates, she can determine the relative rank of the candidatesbut not the true rank. Thus, if there are six candidates and their true rank is6, 1, 4, 2, 3, 5, (where 1 is best) then after she had interviewed the flrst threecandidates she would rank them 3, 1, 2. As she interviews each candidate,she must either accept or reject the candidate. If she does not accept thecandidate after the interview, the candidate is lost to her. She wants to de-cide on a strategy for deciding when to stop and accept a candidate that willmaximize the probability of getting the best candidate. Assume that therearencandidates and they arrive in a random rank order. (a) What is the probability that Barbara gets the best candidate if she inter- views all of the candidates? What is it if she chooses the flrst candidate? (b) Assume that Barbara decides to interview the flrst half of the candidates and then continue interviewing until getting a candidate better than anycandidate seen so far. Show that she has a better than 25 percent chanceof ending up with the best candidate. 24For the task described in Exercise 23, it can be shown 13that the best strategy is to pass over the flrst k¡1 candidates where kis the smallest integer for which1 k+1 k+1+¢¢¢+1 n¡1•1: Using this strategy the probability of getting the best candidate is approxi- mately 1=e=:368. Write a program to simulate Barbara Smith’s interviewing if she uses this optimal strategy, using n= 10, and see if you can verify that the probability of success is approximately 1 =e. 3.2 Combinations Having mastered permutations, we now consider combinations. Let Ube a set with nelements; we want to count the number of distinct subsets of the set Uthat have exactlyjelements. The empty set and the set Uare considered to be subsets of U. The empty set is usually denoted by `. 13E. B. Dynkin and A. A. Yushkevich, Markov P rocesses: Theorems and Problems, trans. J. S. Wood (New York: Plenum, 1969). 3.2. COMBINATIONS 93 Example 3.5 LetU=fa;b;cg. The subsets of Uare `;fag;fbg;fcg;fa;bg;fa;cg;fb;cg;fa;b;cg: 2 Binomial Coe–cients The number of distinct subsets with jelements that can be chosen from a set with nelements is denoted by¡n j¢ , and is pronounced \ nchoosej." The number¡n j¢ is called a binomial coe–cient. This terminology comes from an application to algebra which will be discussed later in this section. In the above example, there is one subset with no elements, three subsets with exactly 1 element, three subsets with exactly 2 elements, and one subset with exactly3 elements. Thus,¡ 3 0¢ =1 ,¡3 1¢ =3 ,¡3 2¢ = 3, and¡3 3¢ = 1. Note that there are 23= 8 subsets in all. (We have already seen that a set with nelements has 2n subsets; see Exercise 3.1.8.) It follows that µ3 0¶ +µ3 1¶ +µ3 2¶ +µ3 3¶ =23=8; µn 0¶ =µn n¶ =1: Assume that n>0. Then, since there is only one way to choose a set with no elements and only one way to choose a set with nelements, the remaining values of¡n j¢ are determined by the following recurrence relation : Theorem 3.4 For integers nandj, with 0<j<n , the binomial coe–cients satisfy:µn j¶ =µn¡1 j¶ +µn¡1 j¡1¶ : (3.1) Proof. We wish to choose a subset of jelements. Choose an element uofU. Assume flrst that we do not want uin the subset. Then we must choose the j elements from a set of n¡1 elements; this can be done in¡n¡1 j¢ ways. On the other hand, assume that we do want uin the subset. Then we must choose the other j¡1 elements from the remaining n¡1 elements of U; this can be done in¡n¡1 j¡1¢ ways. Since uis either in our subset or not, the number of ways that we can choose a subset of jelements is the sum of the number of subsets of jelements which have uas a member and the number which do not|this is what Equation 3.1 states. 2 The binomial coe–cient¡n j¢ is deflned to be 0, if j<0o ri fj>n . With this deflnition, the restrictions on jin Theorem 3.4 are unnecessary. 94 CHAPTER 3. COMBINATORICS n = 0 1 10 1 10 45 120 210 252 210 120 45 10 1 9 1 9 36 84 126 126 84 36 9 1 8 1 8 28 56 70 56 28 8 17 1 7 21 35 35 21 7 16 1 6 15 20 15 6 15 1 5 10 10 5 14 1 4 6 4 13 1 3 3 12 1 2 11 1 1j = 0 1 2 3 4 5 6 7 8 9 10 Figure 3.3: Pascal’s triangle. Pascal’s Triangle The relation 3.1, together with the knowledge that µn 0¶ =µn n¶ =1; determines completely the numbers¡n j¢ . We can use these relations to determine the famous triangle of Pascal, which exhibits all these numbers in matrix form (see Figure 3.3). Thenth row of this triangle has the entries¡n 0¢ ,¡n 1¢ , ...,¡n n¢ . We know that the flrst and last of these numbers are 1. The remaining numbers are determined bythe recurrence relation Equation 3.1; that is, the entry¡ n j¢ for 0<j<n in the nth row of Pascal’s triangle is the sum of the entry immediately above and the one immediately to its left in the ( n¡1)st row. For example,¡5 2¢ =6+4=1 0 . This algorithm for constructing Pascal’s triangle can be used to write a computer program to compute the binomial coe–cients. You are asked to do this in Exercise 4. While Pascal’s triangle provides a way to construct recursively the binomial coe–cients, it is also possible to give a formula for¡n j¢ . Theorem 3.5 The binomial coe–cients are given by the formula µn j¶ =(n)j j!: (3.2) Proof. Each subset of size jof a set of size ncan be ordered in j! ways. Each of these orderings is a j-permutation of the set of size n. The number of j-permutations is (n)j, so the number of subsets of size jis (n)j j!: This completes the proof. 2 3.2. COMBINATIONS 95 The above formula can be rewritten in the form µn j¶ =n! j!(n¡j)!: This immediately shows that µn j¶ =µn n¡j¶ : When using Equation 3.2 in the calculation of¡n j¢ , if one alternates the multi- plications and divisions, then all of the intermediate values in the calculation areintegers. Furthermore, none of these intermediate values exceed the flnal value.(See Exercise 40.) Another point that should be made concerning Equation 3.2 is that if it is used todeflne the binomial coe–cients, then it is no longer necessary to require nto be a positive integer. The variable jmust still be a non-negative integer under this deflnition. This idea is useful when extending the Binomial Theorem to generalexponents. (The Binomial Theorem for non-negative integer exponents is givenbelow as Theorem 3.7.) Poker Hands Example 3.6 Poker players sometimes wonder why a four of a kind beats a full house. A poker hand is a random subset of 5 elements from a deck of 52 cards. A hand has four of a kind if it has four cards with the same value|for example,four sixes or four kings. It is a full house if it has three of one value and two of asecond|for example, three twos and two queens. Let us see which hand is morelikely. How many hands have four of a kind? There are 13 ways that we can specifythe value for the four cards. For each of these, there are 48 possibilities for the flfthcard. Thus, the number of four-of-a-kind hands is 13 ¢48 = 624. Since the total number of possible hands is¡ 52 5¢ = 2598960, the probability of a hand with four of a kind is 624 =2598960 =:00024. Now consider the case of a full house; how many such hands are there? There are 13 choices for the value which occurs three times; for each of these there are¡4 3¢ = 4 choices for the particular three cards of this value that are in the hand. Having picked these three cards, there are 12 possibilities for the value which occurstwice; for each of these there are¡ 4 2¢ = 6 possibilities for the particular pair of this value. Thus, the number of full houses is 13 ¢4¢12¢6 = 3744, and the probability of obtaining a hand with a full house is 3744 =2598960 = :0014. Thus, while both types of hands are unlikely, you are six times more likely to obtain a full house thanfour of a kind. 2 96 CHAPTER 3. COMBINATORICS (start)S FF F FS S SSSS F F Fp qp p q p q q p q p q q q qq qppp pp pq q pp qm (ω) ω ω ω ω ω ω ω ω ω 23 32222 21 2 3 4 5 6 7 8 Figure 3.4: Tree diagram of three Bernoulli trials. Bernoulli Trials Our principal use of the binomial coe–cients will occur in the study of one of the important chance processes called Bernoulli trials. Deflnition 3.5 ABernoulli trials p rocess is a sequence of nchance experiments such that 1. Each experiment has two possible outcomes, which we may call success and failure. 2. The probability pof success on each experiment is the same for each ex- periment, and this probability is not afiected by any knowledge of previousoutcomes. The probability qof failure is given by q=1¡p. 2 Example 3.7 The following are Bernoulli trials processes: 1. A coin is tossed ten times. The two possible outcomes are heads and tails. The probability of heads on any one toss is 1/2. 2. An opinion poll is carried out by asking 1000 people, randomly chosen from the population, if they favor the Equal Rights Amendment|the two outcomesbeing yes and no. The probability pof a yes answer (i.e., a success) indicates the proportion of people in the entire population that favor this amendment. 3. A gambler makes a sequence of 1-dollar bets, betting each time on black at roulette at Las Vegas. Here a success is winning 1 dollar and a failure is losing 3.2. COMBINATIONS 97 1 dollar. Since in American roulette the gambler wins if the ball stops on one of 18 out of 38 positions and loses otherwise, the probability of winning isp=1 8=38 =:474. 2 To analyze a Bernoulli trials process, we choose as our sample space a binary tree and assign a probability measure to the paths in this tree. Suppose, for example,that we have three Bernoulli trials. The possible outcomes are indicated in thetree diagram shown in Figure 3.4. We deflne Xto be the random variable which represents the outcome of the process, i.e., an ordered triple of S’s and F’s. Theprobabilities assigned to the branches of the tree represent the probability for eachindividual trial. Let the outcome of the ith trial be denoted by the random variable X i, with distribution function mi. Since we have assumed that outcomes on any one trial do not afiect those on another, we assign the same probabilities at eachlevel of the tree. An outcome !for the entire experiment will be a path through the tree. For example, ! 3represents the outcomes SFS. Our frequency interpretation of probability would lead us to expect a fraction pof successes on the flrst experiment; of these, a fraction qof failures on the second; and, of these, a fraction pof successes on the third experiment. This suggests assigning probability pqpto the outcome !3. More generally, we assign a distribution function m(!) for paths!by deflning m(!) to be the product of the branch probabilities along the path !. Thus, the probability that the three events S on the flrst trial, F on the second trial, and S on the thirdtrial occur is the product of the probabilities for the individual events. We shallsee in the next chapter that this means that the events involved are independent in the sense that the knowledge of one event does not afiect our prediction for theoccurrences of the other events. Binomial Probabilities We shall be particularly interested in the probability that in nBernoulli trials there are exactly jsuccesses. We denote this probability by b(n;p;j ). Let us calculate the particular value b(3;p;2) from our tree measure. We see that there are three paths which have exactly two successes and one failure, namely !2,!3, and!5. Each of these paths has the same probability p2q.T h u sb(3;p;2 )=3p2q. Considering all possible numbers of successes we have b(3;p;0) =q3; b(3;p;1 )=3pq2; b(3;p;2 )=3p2q; b(3;p;3) =p3: We can, in the same manner, carry out a tree measure for nexperiments and determineb(n;p;j ) for the general case of nBernoulli trials. 98 CHAPTER 3. COMBINATORICS Theorem 3.6 GivennBernoulli trials with probability pof success on each exper- iment, the probability of exactly jsuccesses is b(n;p;j )=µn j¶ pjqn¡j whereq=1¡p. Proof. We construct a tree measure as described above. We want to flnd the sum of the probabilities for all paths which have exactly jsuccesses and n¡jfailures. Each such path is assigned a probability pjqn¡j. How many such paths are there? To specify a path, we have to pick, from the npossible trials, a subset of jto be successes, with the remaining n¡joutcomes being failures. We can do this in¡n j¢ ways. Thus the sum of the probabilities is b(n;p;j )=µn j¶ pjqn¡j: 2 Example 3.8 A fair coin is tossed six times. What is the probability that exactly three heads turn up? The answer is b(6;:5;3) =µ6 3¶µ1 2¶3µ1 2¶3 =2 0¢1 64=:3125: 2 Example 3.9 A die is rolled four times. What is the probability that we obtain exactly one 6? We treat this as Bernoulli trials with success = \rolling a 6" and failure = \rolling some number other than a 6." Then p=1=6, and the probability of exactly one success in four trials is b(4;1=6;1) =µ4 1¶µ1 6¶1µ5 6¶3 =:386: 2 To compute binomial probabilities using the computer, multiply the function choose(n;k)b ypkqn¡k. The program BinomialProbabilities prints out the bi- nomial probabilities b(n;p;k ) forkbetweenkmin andkmax , and the sum of these probabilities. We have run this program for n= 100,p=1=2,kmin = 45, and kmax = 55; the output is shown in Table 3.8. Note that the individual probabilities are quite small. The probability of exactly 50 heads in 100 tosses of a coin is about.08. Our intuition tells us that this is the most likely outcome, which is correct;but, all the same, it is not a very likely outcome. 3.2. COMBINATIONS 99 kb (n;p;k ) 45 .0485 46 .058047 .066648 .073549 .078050 .079651 .078052 .073553 .066654 .058055 .0485 Table 3.8: Binomial probabilities for n= 100;p=1=2. Binomial Distributions Deflnition 3.6 Letnbe a positive integer, and let pbe a real number between 0 and 1. Let Bbe the random variable which counts the number of successes in a Bernoulli trials process with parameters nandp. Then the distribution b(n;p;k ) ofBis called the binomial distribution . 2 We can get a better idea about the binomial distribution by graphing this dis- tribution for difierent values of nandp(see Figure 3.5). The plots in this flgure were generated using the program BinomialPlot . We have run this program for p=:5 andp=:3. Note that even for p=:3 the graphs are quite symmetric. We shall have an explanation for this in Chapter 9. Wealso note that the highest probability occurs around the value np, but that these highest probabilities get smaller as nincreases. We shall see in Chapter 6 that np is the mean orexpected value of the binomial distribution b(n;p;k ). The following example gives a nice way to see the binomial distribution, when p=1=2. Example 3.10 AGalton board is a board in which a large number of BB-shots are dropped from a chute at the top of the board and de°ected ofi a number of pins ontheir way down to the bottom of the board. The flnal position of each slot is theresult of a number of random de°ections either to the left or the right. We havewritten a program GaltonBoard to simulate this experiment. We have run the program for the case of 20 rows of pins and 10,000 shots being dropped. We show the result of this simulation in Figure 3.6. Note that if we write 0 every time the shot is de°ected to the left, and 1 every time it is de°ected to the right, then the path of the shot can be described by asequence of 0’s and 1’s of length n, just as for the n-fold coin toss. The distribution shown in Figure 3.6 is an example of an empirical distribution, in the sense that it comes about by means of a sequence of experiments. As expected, 100 CHAPTER 3. COMBINATORICS 0 20 40 60 80 100 12000.0250.050.0750.10.1250.150 20 40 60 80 1000.020.040.060.080.10.12 p = .5 n = 40 n = 80 n = 160 n = 30 n = 120 n = 270p = .30 Figure 3.5: Binomial distributions. 3.2. COMBINATIONS 101 Figure 3.6: Simulation of the Galton board. this empirical distribution resembles the corresponding binomial distribution with parameters n= 20 andp=1=2. 2 Hypothesis Testing Example 3.11 Suppose that ordinary aspirin has been found efiective against headaches 60 percent of the time, and that a drug company claims that its newaspirin with a special headache additive is more efiective. We can test this claimas follows: we call their claim the alternate hypothesis, and its negation, that the additive has no appreciable efiect, the null hypothesis. Thus the null hypothesis is thatp=:6, and the alternate hypothesis is that p>: 6, wherepis the probability that the new aspirin is efiective. We give the aspirin to npeople to take when they have a headache. We want to flnd a number m, called the critical value for our experiment, such that we reject the null hypothesis if at least mpeople are cured, and otherwise we accept it. How should we determine this critical value? First note that we can make two kinds of errors. The flrst, often called a type 1 error in statistics, is to reject the null hypothesis when in fact it is true. The second, called a type 2 error, is to accept the null hypothesis when it is false. To determine the probability of both these types of errors we introduce a function fi(p), deflned to be the probability that we reject the null hypothesis, where this probability iscalculated under the assumption that the null hypothesis is true. In the presentcase, we have fi(p)=X m•k•nb(n;p;k ): 102 CHAPTER 3. COMBINATORICS Note thatfi(:6) is the probability of a type 1 error, since this is the probability of a high number of successes for an inefiective additive. So for a given nwe want to choosemso as to make fi(:6) quite small, to reduce the likelihood of a type 1 error. But as mincreases above the most probable value np=:6n,fi(:6), being the upper tail of a binomial distribution, approaches 0. Thus increasingmmakes a type 1 error less likely. Now suppose that the additive really is efiective, so that pis appreciably greater than .6; say p=:8. (This alternative value of pis chosen arbitrarily; the following calculations depend on this choice.) Then choosing mwell below np=:8nwill increasefi(:8), since now fi(:8) is all but the lower tail of a binomial distribution. Indeed, if we put fl(:8 )=1¡fi(:8), thenfl(:8) gives us the probability of a type 2 error, and so decreasingmmakes a type 2 error less likely. The manufacturer would like to guard against a type 2 error, since if such an error is made, then the test does not show that the new drug is better, when infact it is. If the alternative value of pis chosen closer to the value of pgiven in the null hypothesis (in this case p=:6), then for a given test population, the value offlwill increase. So, if the manufacturer’s statistician chooses an alternative value forpwhich is close to the value in the null hypothesis, then it will be an expensive proposition (i.e., the test population will have to be large) to reject thenull hypothesis with a small value of fl. What we hope to do then, for a given test population n, is to choose a value ofm, if possible, which makes both these probabilities small. If we make a type 1 error we end up buying a lot of essentially ordinary aspirin at an in°ated price; atype 2 error means we miss a bargain on a superior medication. Let us say thatwe want our critical number mto make each of these undesirable cases less than 5 percent probable. We write a program PowerCurve to plot, for n= 100 and selected values of m, the function fi(p), forpranging from .4 to 1. The result is shown in Figure 3.7. We include in our graph a box (in dotted lines) from .6 to .8, with bottom and top atheights .05 and .95. Then a value for msatisfles our requirements if and only if the graph offienters the box from the bottom, and leaves from the top (why?|which is the type 1 and which is the type 2 criterion?). As mincreases, the graph of fi moves to the right. A few experiments have shown us that m= 69 is the smallest value formthat thwarts a type 1 error, while m= 73 is the largest which thwarts a type 2. So we may choose our critical value between 69 and 73. If we’re more intenton avoiding a type 1 error we favor 73, and similarly we favor 69 if we regard atype 2 error as worse. Of course, the drug company may not be happy with havinga sm u c ha sa5p ercent chance of an error. They might insist on havin ga1p ercent chance of an error. For this we would have to increase the number nof trials (see Exercise 28). 2 Binomial Expansion We next remind the reader of an application of the binomial coe–cients to algebra. This is the binomial expansion, from which we get the term binomial coe–cient. 3.2. COMBINATIONS 103 .4 1 .5 .6 .7 .8 .9 1 .01.0 .1 .2 .3 .4 .5 .6 .7 .8 .91.0 .4 1 .5 .6 .7 .8 .9 1 .01.0 .1 .2 .3 .4 .5 .6 .7 .8 .91.0 Figure 3.7: The power curve. Theorem 3.7 (Binomial Theorem) The quantity ( a+b)ncan be expressed in the form (a+b)n=nX j=0µn j¶ ajbn¡j: Proof. To see that this expansion is correct, write (a+b)n=(a+b)(a+b)¢¢¢(a+b): When we multiply this out we will have a sum of terms each of which results from a choice of an aorbfor each of nfactors. When we choose ja’s and (n¡j)b’s, we obtain a term of the form ajbn¡j. To determine such a term, we have to specify jof thenterms in the product from which we choose the a. This can be done in¡n j¢ ways. Thus, collecting these terms in the sum contributes a term¡n j¢ ajbn¡j.2 For example, we have (a+b)0=1 (a+b)1=a+b (a+b)2=a2+2ab+b2 (a+b)3=a3+3a2b+3ab2+b3: We see here that the coe–cients of successive powers do indeed yield Pascal’s tri- angle. Corollary 3.1 The sum of the elements in the nth row of Pascal’s triangle is 2n. If the elements in the nth row of Pascal’s triangle are added with alternating signs, the sum is 0. 104 CHAPTER 3. COMBINATORICS Proof. The flrst statement in the corollary follows from the fact that 2n=( 1+1 )n=µn 0¶ +µn 1¶ +µn 2¶ +¢¢¢+µn n¶ ; and the second from the fact that 0=( 1¡1)n=µn 0¶ ¡µn 1¶ +µn 2¶ ¡¢¢¢ +(¡1)nµn n¶ : 2 The flrst statement of the corollary tells us that the number of subsets of a set ofnelements is 2n. We shall use the second statement in our next application of the binomial theorem. We have seen that, when AandBare any two events (cf. Section 1.2), P(A[B)=P(A)+P(B)¡P(A\B): We now extend this theorem to a more general version, which will enable us to flnd the probability that at least one of a number of events occurs. Inclusion-Exclusion Principle Theorem 3.8 LetPbe a probability measure on a sample space ›, and let fA1;A2; :::; Angbe a flnite set of events. Then P(A1[A2[¢¢¢[An)=nX i=1P(Ai)¡X 1•i<j•nP(Ai\Aj) +X 1•i<j<k•nP(Ai\Aj\Ak)¡¢¢¢:(3.3) That is, to flnd the probability that at least one of neventsAioccurs, flrst add the probability of each event, then subtract the probabilities of all possible two-wayintersections, add the probability of all three-way intersections, and so forth. Proof. If the outcome !occurs in at least one of the events A i, its probability is added exactly once by the left side of Equation 3.3. We must show that it is addedexactly once by the right side of Equation 3.3. Assume that !is in exactly kof the sets. Then its probability is added ktimes in the flrst term, subtracted¡ k 2¢ times in the second, added¡k 3¢ times in the third term, and so forth. Thus, the total number of times that it is added is µk 1¶ ¡µk 2¶ +µk 3¶ ¡¢¢¢ (¡1)k¡1µk k¶ : But 0=( 1¡1)k=kX j=0µk j¶ (¡1)j=µk 0¶ ¡kX j=1µk j¶ (¡1)j¡1: 3.2. COMBINATIONS 105 Hence, 1=µk 0¶ =kX j=1µk j¶ (¡1)j¡1: If the outcome !is not in any of the events Ai, then it is not counted on either side of the equation. 2 Hat Check Problem Example 3.12 We return to the hat check problem discussed in Section 3.1, that is, the problem of flnding the probability that a random permutation contains atleast one flxed point. Recall that a permutation is a one-to-one map of a setA=fa 1;a2;:::;angonto itself. Let Aibe the event that the ith elementairemains flxed under this map. If we require that aiis flxed, then the map of the remaining n¡1 elements provides an arbitrary permutation of ( n¡1) objects. Since there are (n¡1)! such permutations, P(Ai)=(n¡1)!=n!=1=n. Since there are nchoices forai, the flrst term of Equation 3.3 is 1. In the same way, to have a particular pair (ai;aj) flxed, we can choose any permutation of the remaining n¡2 elements; there are (n¡2)! such choices and thus P(Ai\Aj)=(n¡2)! n!=1 n(n¡1): The number of terms of this form in the right side of Equation 3.3 is µn 2¶ =n(n¡1) 2!: Hence, the second term of Equation 3.3 is ¡n(n¡1) 2!¢1 n(n¡1)=¡1 2!: Similarly, for any speciflc three events Ai,Aj,Ak, P(Ai\Aj\Ak)=(n¡3)! n!=1 n(n¡1)(n¡2); and the number of such terms is µn 3¶ =n(n¡1)(n¡2) 3!; making the third term of Equation 3.3 equal to 1/3!. Continuing in this way, we obtain P(at least one flxed point) = 1 ¡1 2!+1 3!¡¢¢¢ (¡1)n¡11 n! and P(no flxed point) =1 2!¡1 3!+¢¢¢(¡1)n1 n!: 106 CHAPTER 3. COMBINATORICS Probability that no one n gets his own hat back 3 .3333334 .3755 .3666676 .3680567 .3678578 .3678829 .367879 10 .367879 Table 3.9: Hat check problem. From calculus we learn that ex=1+x+1 2!x2+1 3!x3+¢¢¢+1 n!xn+¢¢¢: Thus, ifx=¡1, we have e¡1=1 2!¡1 3!+¢¢¢+(¡1)n n!+¢¢¢ =:3678794: Therefore, the probability that there is no flxed point, i.e., that none of the npeople gets his own hat back, is equal to the sum of the flrst nterms in the expression for e¡1. This series converges very fast. Calculating the partial sums for n= 3 to 10 gives the data in Table 3.9. Aftern= 9 the probabilities are essentially the same to six signiflcant flgures. Interestingly, the probability of no flxed point alternately increases and decreasesasnincreases. Finally, we note that our exact results are in good agreement with our simulations reported in the previous section. 2 Choosing a Sample Space We now have some of the tools needed to accurately describe sample spaces and to assign probability functions to those sample spaces. Nevertheless, in some cases,the description and assignment process is somewhat arbitrary. Of course, it is tobe hoped that the description of the sample space and the subsequent assignmentof a probability function will yield a model which accurately predicts what wouldhappen if the experiment were actually carried out. As the following examples show,there are situations in which \reasonable" descriptions of the sample space do notproduce a model which flts the data. In Feller’s book, 14a pair of models is given which describe arrangements of certain kinds of elementary particles, such as photons and protons. It turns out thatexperiments have shown that certain types of elementary particles exhibit behavior 14W. Feller, Introduction to Probability Theory and Its Applications vol. 1, 3rd ed. (New York: John Wiley and Sons, 1968), p. 41 3.2. COMBINATIONS 107 which is accurately described by one model, called \Bose-Einstein statistics," while other types of elementary particles can be modelled using \Fermi-Dirac statistics." Feller says: We have here an instructive example of the impossibility of selecting or justifying probability models by a priori arguments. In fact, no pure reasoning could tell that photons and protons would not obey the sameprobability laws. We now give some examples of this description and assignment process. Example 3.13 In the quantum mechanical model of the helium atom, various parameters can be used to classify the energy states of the atom. In the tripletspin state ( S= 1) with orbital angular momentum 1 ( L= 1), there are three possibilities, 0, 1, or 2, for the total angular momentum ( J). (It is not assumed that the reader knows what any of this means; in fact, the example is more illustrativeif the reader does notknow anything about quantum mechanics.) We would like to assign probabilities to the three possibilities for J. The reader is undoubtedly resisting the idea of assigning the probability of 1 =3 to each of these outcomes. She should now ask herself why she is resisting this assignment. The answer is probablybecause she does not have any \intuition" (i.e., experience) about the way in whichhelium atoms behave. In fact, in this example, the probabilities 1 =9;3=9;and 5=9 are assigned by the theory. The theory gives these assignments because these frequencies were observed in experiments and further parameters were developed in the theory to allow these frequencies to be predicted. 2 Example 3.14 Suppose two pennies are °ipped once each. There are several \rea- sonable" ways to describe the sample space. One way is to count the number ofheads in the outcome; in this case, the sample space can be written f0;1;2g. An- other description of the sample space is the set of all ordered pairs of H’s andT’s, i.e., f(H;H );(H;T);(T;H);(T;T)g: Both of these descriptions are accurate ones, but it is easy to see that (at most) one of these, if assigned a constant probability function, can claim to accurately modelreality. In this case, as opposed to the preceding example, the reader will probablysay that the second description, with each outcome being assigned a probability of1=4, is the \right" description. This conviction is due to experience; there is no proof that this is the way reality works. 2 The reader is also referred to Exercise 26 for another example of this process. Historical Remarks The binomial coe–cients have a long and colorful history leading up to Pascal’s Treatise on the Arithmetical Triangle,15where Pascal developed many important 15B. Pascal, Trait¶ e du Triangle Arithm¶ etique (Paris: Desprez, 1665). 108 CHAPTER 3. COMBINATORICS 1 111 1 111 1 1 1 234 5 678 913 61 0 1 5 2 12 83 6141 02 0 3 5 5 68 4151 53 5 7 01 2 6162 15 61 2 6172 88 4183 6191 Table 3.10: Pascal’s triangle. natural numbers 123456789 triangular numbers 1361 0 1 5 2 1 2 8 3 6 4 5 tetrahedral numbers 1 4 10 20 35 56 84 120 165 Table 3.11: Figurate numbers. properties of these numbers. This history is set forth in the book Pascal’s Arith- metical Triangle by A. W. F. Edwards. 16Pascal wrote his triangle in the form shown in Table 3.10. Edwards traces three difierent ways that the binomial coe–cients arose. He refers to these as the flgurate numbers, thecombinatorial numbers, and the binomial numbers. They are all names for the same thing (which we have called binomial coe–cients) but that they are all the same was not appreciated until the sixteenthcentury. The flgurate numbers date back to the Pythagorean interest in number pat- terns around 540 BC.The Pythagoreans considered, for example, triangular patterns shown in Figure 3.8. The sequence of numbers 1;3;6;10;::: obtained as the number of points in each triangle are called triangular numbers. From the triangles it is clear that the nth triangular number is simply the sum of the flrstnintegers. The tetrahedral numbers are the sums of the triangular numbers and were obtained by the Greek mathematicians Theon and Nicomachus at thebeginning of the second century BC.The tetrahedral number 10, for example, has the geometric representation shown in Figure 3.9. The flrst three types of flguratenumbers can be represented in tabular form as shown in Table 3.11. These numbers provide the flrst four rows of Pascal’s triangle, but the table was not to be completed in the West until the sixteenth century. In the East, Hindu mathematicians began to encounter the binomial coe–cients in combinatorial problems. Bhaskara in his Lilavati of 1150 gave a rule to flnd the 16A. W. F. Edwards, Pascal’s Arithmetical Triangle (London: Gri–n, 1987). 3.2. COMBINATIONS 109 1 36 10 Figure 3.8: Pythagorean triangular patterns. Figure 3.9: Geometric representation of the tetrahedral number 10. 110 CHAPTER 3. COMBINATORICS 11 12 2213 23 3314 24 34 4415 25 35 45 5516 26 36 46 56 66 Table 3.12: Outcomes for the roll of two dice. number of medicinal preparations using 1, 2, 3, 4, 5, or 6 possible ingredients. 17His rule is equivalent to our formula µn r¶ =(n)r r!: The binomial numbers as coe–cients of ( a+b)nappeared in the works of math- ematicians in China around 1100. There are references about this time to \thetabulation system for unlocking binomial coe–cients." The triangle to provide thecoe–cients up to the eighth power is given by Chu Shih-chieh in a book writtenaround 1303 (see Figure 3.10). 18The original manuscript of Chu’s book has been lost, but copies have survived. Edwards notes that there is an error in this copy ofChu’s triangle. Can you flnd it? ( Hint: Two numbers which should be equal are not.) Other copies do not show this error. The flrst appearance of Pascal’s triangle in the West seems to have come from calculations of Tartaglia in calculating the number of possible ways that ndice might turn up. 19For one die the answer is clearly 6. For two dice the possibilities may be displayed as shown in Table 3.12. Displaying them this way suggests the sixth triangular number 1+2+3+4+ 5 + 6 = 21 for the throw of 2 dice. Tartaglia \on the flrst day of Lent, 1523, inVerona, having thought about the problem all night," 20realized that the extension of the flgurate table gave the answers for ndice. The problem had suggested itself to Tartaglia from watching people casting their own horoscopes by means of a Book of Fortune, selecting verses by a process which included noting the numbers on the faces of three dice. The 56 ways that three dice can fall were set out on each page.The way the numbers were written in the book did not suggest the connection withflgurate numbers, but a method of enumeration similar to the one we used for 2dice does. Tartaglia’s table was not published until 1556. A table for the binomial coe–cients was published in 1554 by the German mathe- matician Stifel. 21Pascal’s triangle appears also in Cardano’s Opus novum of 1570.22 17ibid., p. 27. 18J. Needham, Science and Civilization in China, vol. 3 (New York: Cambridge University Press, 1959), p. 135. 19N. Tartaglia, General Trattato di Numeri et Misure (Vinegia, 1556). 20Quoted in Edwards, op. cit., p. 37. 21M. Stifel, Arithmetica Integra (Norimburgae, 1544). 22G. Cardano, Opus Novum de Proportionibus Numerorum (Basilea, 1570). 3.2. COMBINATIONS 111 Figure 3.10: Chu Shih-chieh’s triangle. [From J. Needham, Science and Civilization in China, vol. 3 (New York: Cambridge University Press, 1959), p. 135. Reprinted with permission.] 112 CHAPTER 3. COMBINATORICS Cardano was interested in the problem of flnding the number of ways to choose r objects out of n. Thus by the time of Pascal’s work, his triangle had appeared as a result of looking at the flgurate numbers, the combinatorial numbers, and thebinomial numbers, and the fact that all three were the same was presumably prettywell understood. Pascal’s interest in the binomial numbers came from his letters with Fermat concerning a problem known as the problem of points. This problem, and thecorrespondence between Pascal and Fermat, were discussed in Chapter 1. Thereader will recall that this problem can be described as follows: Two players A andB are playing a sequence of games and the flrst player to win ngames wins the match. It is desired to flnd the probability that A wins the match at a time whenA has won agames and B has won bgames. (See Exercises 4.1.40-4.1.42.) Pascal solved the problem by backward induction, much the way we would do today in writing a computer program for its solution. He referred to the combina-torial method of Fermat which proceeds as follows: If A needs cgames and B needs dgames to win, we require that the players continue to play until they have played c+d¡1 games. The winner in this extended series will be the same as the winner in the original series. The probability that A wins in the extended series and hencein the original series is c+d¡1X r=c1 2c+d¡1µc+d¡1 r¶ : Even at the time of the letters Pascal seemed to understand this formula. Suppose that the flrst player to win ngames wins the match, and suppose that each player has put up a stake of x. Pascal studied the value of winning a particular game. By this he meant the increase in the expected winnings of the winner of theparticular game under consideration. He showed that the value of the flrst game is 1¢3¢5¢:::¢(2n¡1) 2¢4¢6¢:::¢(2n)x: His proof of this seems to use Fermat’s formula and the fact that the above ratio of products of odd to products of even numbers is equal to the probability of exactlynheads in 2ntosses of a coin. (See Exercise 39.) Pascal presented Fermat with the table shown in Table 3.13. He states: You will see as always, that the value of the flrst game is equal to that of the second which is easily shown by combinations. You will see, inthe same way, that the numbers in the flrst line are always increasing;so also are those in the second; and those in the third. But those in thefourth line are decreasing, and those in the flfth, etc. This seems odd. 23 The student can pursue this question further using the computer and Pascal’s backward iteration method for computing the expected payofi at any point in theseries. 23F. N. David, op. cit., p. 235. 3.2. COMBINATIONS 113 if each one staken 256 in From my opponent’s 256 654321 positions I get, for the games games games games games games 1st game 63 70 80 96 128 256 2nd game 63 70 80 96 128 3rd game 56 60 64 64 4th game 42 40 325th game 24 166th game 8 Table 3.13: Pascal’s solution for the problem of points. In his treatise, Pascal gave a formal proof of Fermat’s combinatorial formula as well as proofs of many other basic properties of binomial numbers. Many of hisproofs involved induction and represent some of the flrst proofs by this method.His book brought together all the difierent aspects of the numbers in the Pascaltriangle as known in 1654, and, as Edwards states, \That the Arithmetical Triangleshould bear Pascal’s name cannot be disputed." 24 The flrst serious study of the binomial distribution was undertaken by James Bernoulli in his Ars Conjectandi published in 1713.25We shall return to this work in the historical remarks in Chapter 8. Exercises 1Compute the following: (a)¡6 3¢ (b)b(5;:2;4) (c)¡7 2¢ (d)¡26 26¢ (e)b(4;:2;3) (f)¡6 2¢ (g)¡10 9¢ (h)b(8;:3;5) 2In how many ways can we choose flve people from a group of ten to form a committee? 3How many seven-element subsets are there in a set of nine elements? 4Using the relation Equation 3.1 write a program to compute Pascal’s triangle, putting the results in a matrix. Have your program print the triangle forn= 10. 24A. W. F. Edwards, op. cit., p. ix. 25J. Bernoulli, Ars Conjectandi (Basil: Thurnisiorum, 1713). 114 CHAPTER 3. COMBINATORICS 5Use the program BinomialProbabilities to flnd the probability that, in 100 tosses of a fair coin, the number of heads that turns up lies between 35 and65, between 40 and 60, and between 45 and 55. 6Charles claims that he can distinguish between beer and ale 75 percent of the time. Ruth bets that he cannot and, in fact, just guesses. To settle this, a betis made: Charles is to be given ten small glasses, each having been fllled withbeer or ale, chosen by tossing a fair coin. He wins the bet if he gets seven ormore correct. Find the probability that Charles wins if he has the ability thathe claims. Find the probability that Ruth wins if Charles is guessing. 7Show that b(n;p;j )=p qµn¡j+1 j¶ b(n;p;j¡1); forj‚1. Use this fact to determine the value or values of jwhich give b(n;p;j ) its greatest value. Hint: Consider the successive ratios as jincreases. 8A die is rolled 30 times. What is the probability tha t a 6 turns up exactly 5 times? What is the most probable number of times tha t a 6 will turn up? 9Find integers nandrsuch that the following equation is true: µ13 5¶ +2µ13 6¶ +µ13 7¶ =µn r¶ : 10In a ten-question true-false exam, flnd the probability that a student gets a grade of 70 percent or better by guessing. Answer the same question if thetest has 30 questions, and if the test has 50 questions. 11A restaurant ofiers apple and blueberry pies and stocks an equal number of each kind of pie. Each day ten customers request pie. They choose, withequal probabilities, one of the two kinds of pie. How many pieces of each kindof pie should the owner provide so that the probability is about .95 that eachcustomer gets the pie of his or her own choice? 12A poker hand is a set of 5 cards randomly chosen from a deck of 52 cards. Find the probability of a (a) royal °ush (ten, jack, queen, king, ace in a single suit). (b) straight °ush (flve in a sequence in a single suit, but not a royal °ush). (c) four of a kind (four cards of the same face value). (d) full house (one pair and one triple, each of the same face value). (e) °ush (flve cards in a single suit but not a straight or royal °ush). (f) straight (flve cards in a sequence, not all the same suit). (Note that in straights, an ace counts high or low.) 13If a set has 2 nelements, show that it has more subsets with nelements than with any other number of elements. 3.2. COMBINATIONS 115 14Letb(2n;:5;n) be the probability that in 2 ntosses of a fair coin exactly nheads turn up. Using Stirling’s formula (Theorem 3.3), show that b(2n;:5;n)» 1=p…n. Use the program BinomialProbabilities to compare this with the exact value for n=1 0t o2 5 . 15A baseball player, Smith, has a batting average of :300 and in a typical game comes to bat three times. Assume that Smith’s hits in a game can be consid-ered to be a Bernoulli trials process with probability .3 for success. Find the probability that Smith gets 0, 1, 2, and 3 hits. 16The Siwash University football team plays eight games in a season, winning three, losing three, and ending two in a tie. Show that the number of waysthat this can happen isµ8 3¶µ5 3¶ =8! 3! 3! 2!: 17Using the technique of Exercise 16, show that the number of ways that one can putndifierent objects into three boxes with ain the flrst, bin the second, andcin the third is n!=(a!b!c!). 18Baumgartner, Prosser, and Crowell are grading a calculus exam. There is a true-false question with ten parts. Baumgartner notices that one student hasonly two out of the ten correct and remarks, \The student was not even brightenough to have °ipped a coin to determine his answers." \Not so clear," saysProsser. \With 340 students I bet that if they all °ipped coins to determinetheir answers there would be at least one exam with two or fewer answerscorrect." Crowell says, \I’m with Prosser. In fact, I bet that we should expectat least one exam in which no answer is correct if everyone is just guessing."Who is right in all of this? 19A gin hand consists of 10 cards from a deck of 52 cards. Find the probability that a gin hand has (a) all 10 cards of the same suit. (b) exactly 4 cards in one suit and 3 in two other suits. (c) a 4, 3, 2, 1, distribution of suits. 20A six-card hand is dealt from an ordinary deck of cards. Find the probability that: (a) All six cards are hearts. (b) There are three aces, two kings, and one queen. (c) There are three cards of one suit and three of another suit. 21A lady wishes to color her flngernails on one hand using at most two of the colors red, yellow, and blue. How many ways can she do this? 116 CHAPTER 3. COMBINATORICS 22How many ways can six indistinguishable letters be put in three mail boxes? Hint: One representation of this is given by a sequence jLLjLjLLLjwhere the j’s represent the partitions for the boxes and the L’s the letters. Any possible way can be so described. Note that we need two bars at the ends and theremaining two bars and the six L’s can be put in any order. 23Using the method for the hint in Exercise 22, show that rindistinguishable objects can be put in nboxes in µn+r¡1 n¡1¶ =µn+r¡1 r¶ difierent ways. 24A travel bureau estimates that when 20 tourists go to a resort with ten hotels they distribute themselves as if the bureau were putting 20 indistinguishableobjects into ten distinguishable boxes. Assuming this model is correct, flndthe probability that no hotel is left vacant when the flrst group of 20 touristsarrives. 25An elevator takes on six passengers and stops at ten °oors. We can assign two difierent equiprobable measures for the ways that the passengers are dis-charged: (a) we consider the passengers to be distinguishable or (b) we con-sider them to be indistinguishable (see Exercise 23 for this case). For eachcase, calculate the probability that all the passengers get ofi at difierent °oors. 26You are playing heads or tails with Prosser but you suspect that his coin is unfair. Von Neumann suggested that you proceed as follows: Toss Prosser’scoin twice. If the outcome is HT call the result win. if it is TH call the result lose. If it is TT or HH ignore the outcome and toss Prosser’s coin twice again. Keep going until you get either an HT or a TH and call the result win or losein a single play. Repeat this procedure for each play. Assume that Prosser’scoin turns up heads with probability p. (a) Find the probability of HT, TH, HH, TT with two tosses of Prosser’s coin. (b) Using part (a), show that the probability of a win on any one play is 1/2, no matter what pis. 27John claims that he has extrasensory powers and can tell which of two symbols is on a card turned face down (see Example 3.11). To test his ability he isasked to do this for a sequence of trials. Let the null hypothesis be that he isjust guessing, so that the probability is 1/2 of his getting it right each time,and let the alternative hypothesis be that he can name the symbol correctlymore than half the time. Devise a test with the property that the probabilityof a type 1 error is less than .05 and the probability of a type 2 error is lessthan .05 if John can name the symbol correctly 75 percent of the time. 3.2. COMBINATIONS 117 28In Example 3.11 assume the alternative hypothesis is that p=:8 and that it is desired to have the probability of each type of error less than .01. Use theprogram PowerCurve to determine values of nandmthat will achieve this. Choosenas small as possible. 29A drug is assumed to be efiective with an unknown probability p. To estimate pthe drug is given to npatients. It is found to be efiective for mpatients. The method of maximum likelihood for estimating pstates that we should choose the value for pthat gives the highest probability of getting what we got on the experiment. Assuming that the experiment can be considered as aBernoulli trials process with probability pfor success, show that the maximum likelihood estimate for pis the proportion m=n of successes. 30Recall that in the World Series the flrst team to win four games wins the series. The series can go at most seven games. Assume that the Red Soxand the Mets are playing the series. Assume that the Mets win each gamewith probability p. Fermat observed that even though the series might not go seven games, the probability that the Mets win the series is the same as theprobability that they win four or more game in a series that was forced to goseven games no matter who wins the individual games. (a) Using the program PowerCurve of Example 3.11 flnd the probability that the Mets win the series for the cases p=:5,p=:6,p=:7. (b) Assume that the Mets have probability .6 of winning each game. Use the program PowerCurve to flnd a value of nso that, if the series goes to the flrst team to win more than half the games, the Mets will have a95 percent chance of winning the series. Choose nas small as possible. 31Each of the four engines on an airplane functions correctly on a given °ight with probability .99, and the engines function independently of each other.Assume that the plane can make a safe landing if at least two of its enginesare functioning correctly. What is the probability that the engines will allowfor a safe landing? 32A small boy is lost coming down Mount Washington. The leader of the search team estimates that there is a probability pthat he came down on the east side and a probability 1 ¡pthat he came down on the west side. He has n people in his search team who will search independently and, if the boy ison the side being searched, each member will flnd the boy with probabilityu. Determine how he should divide the npeople into two groups to search the two sides of the mountain so that he will have the highest probability offlnding the boy. How does this depend on u? *33 2nballs are chosen at random from a total of 2 nred balls and 2 nblue balls. Find a combinatorial expression for the probability that the chosen balls areequally divided in color. Use Stirling’s formula to estimate this probability. 118 CHAPTER 3. COMBINATORICS Using BinomialProbabilities , compare the exact value with Stirling’s ap- proximation for n= 20. 34Assume that every time you buy a box of Wheaties, you receive one of the pictures of the nplayers on the New York Yankees. Over a period of time, you buym‚nboxes of Wheaties. (a) Use Theorem 3.8 to show that the probability that you get all npictures is 1¡µn 1¶µn¡1 n¶m +µn 2¶µn¡2 n¶m ¡¢¢¢ +(¡1)n¡1µn n¡1¶µ1 n¶m : Hint: LetEkbe the event that you do not get the kth player’s picture. (b) Write a computer program to compute this probability. Use this program to flnd, for given n, the smallest value of mwhich will give probability ‚:5 of getting all npictures. Consider n= 50, 100, and 150 and show thatm=nlogn+nlog 2 is a good estimate for the number of boxes needed. (For a derivation of this estimate, see Feller.26) *35 Prove the following binomial identity µ2n n¶ =nX j=0µn j¶2 : Hint: Consider an urn with nred balls and nblue balls inside. Show that each side of the equation equals the number of ways to choose nballs from the urn. 36Letjandnbe positive integers, with j•n. An experiment consists of choosing, at random, a j-tuple of positive integers whose sum is at most n. (a) Find the size of the sample space. Hint: Consider nindistinguishable balls placed in a row. Place jmarkers between consecutive pairs of balls, with no two markers between the same pair of balls. (We also allow oneof thenmarkers to be placed at the end of the row of balls.) Show that there is a 1-1 correspondence between the set of possible positions forthe markers and the set of j-tuples whose size we are trying to count. (b) Find the probability that the j-tuple selected contains at least one 1. 37Letn(modm) denote the remainder when the integer nis divided by the integerm. Write a computer program to compute the numbers¡ n j¢ (modm) where¡n j¢ is a binomial coe–cient and mis an integer. You can do this by using the recursion relations for generating binomial coe–cients, doing all the 26W. Feller, Introduction to Probability Theory and its Applications, vol. I, 3rd ed. (New York: John Wiley & Sons, 1968), p. 106. 3.2. COMBINATIONS 119 arithmetic using the basic function mod( n;m). Try to write your program to make as large a table as possible. Run your program for the cases m= 2 to 7. Do you see any patterns? In particular, for the case m= 2 andna power of 2, verify that all the entries in the ( n¡1)st row are 1. (The corresponding binomial numbers are odd.) Use your pictures to explain why this is true. 38Lucas27proved the following general result relating to Exercise 37. If pis any prime number, then¡n j¢ (modp) can be found as follows: Expand n andjin basepasn=s0+s1p+s2p2+¢¢¢+skpkandj=r0+r1p+ r2p2+¢¢¢+rkpk, respectively. (Here kis chosen large enough to represent all numbers from 0 to nin basepusingkdigits.) Let s=(s0;s1;s2;:::;sk) and r=(r0;r1;r2;:::;rk). Then µn j¶ (modp)=kY i=0µsi ri¶ (modp): For example, if p=7 ,n= 12, andj= 9, then 1 2=5¢70+1¢71; 9=2¢70+1¢71; so that s=( 5;1); r=( 2;1); and this result states that µ12 9¶ (modp)=µ5 2¶µ1 1¶ (mod 7): Since¡12 9¢ = 220 = 3 (mod 7), and¡5 2¢ = 10 = 3 (mod 7), we see that the result is correct for this example. Show that this result implies that, for p= 2, the (pk¡1)st row of your triangle in Exercise 37 has no zeros. 39Prove that the probability of exactly nheads in 2ntosses of a fair coin is given by the product of the odd numbers up to 2 n¡1 divided by the product of the even numbers up to 2 n. 40Letnbe a positive integer, and assume that jis a positive integer not exceed- ingn=2. Show that in Theorem 3.5, if one alternates the multiplications and divisions, then all of the intermediate values in the calculation are integers.Show also that none of these intermediate values exceed the flnal value. 27E. Lucas, \Th¶ eorie des Functions Num¶ eriques Simplement Periodiques," American J. Math., vol. 1 (1878), pp. 184-240, 289-321. 120 CHAPTER 3. COMBINATORICS 3.3 Card Shu†ing Much of this section is based upon an article by Brad Mann,28which is an exposition of an article by David Bayer and Persi Diaconis.29 Ri†e Shu†es Given a deck of ncards, how many times must we shu†e it to make it \random"? Of course, the answer depends upon the method of shu†ing which is used and whatwe mean by \random." We shall begin the study of this question by considering astandard model for the ri†e shu†e. We begin with a deck of ncards, which we will assume are labelled in increasing order with the integers from 1 to n. A ri†e shu†e consists of a cut of the deck into two stacks and an interleaving of the two stacks. For example, if n= 6, the initial ordering is (1 ;2;3;4;5;6), and a cut might occur between cards 2 and 3. This gives rise to two stacks, namely (1 ;2) and (3;4;5;6). These are interleaved to form a new ordering of the deck. For example, these two stacks might form the ordering(1;3;4;2;5;6). In order to discuss such shu†es, we need to assign a probability measure to the set of all possible shu†es. There are several reasonable ways inwhich this can be done. We will give several difierent assignment strategies, andshow that they are equivalent. (This does not mean that this assignment is theonly reasonable one.) First, we assign the binomial probability b(n;1=2;k)t ot h e event that the cut occurs after the kth card. Next, we assume that all possible interleavings, given a cut, are equally likely. Thus, to complete the assignmentof probabilities, we need to determine the number of possible interleavings of twostacks of cards, with kandn¡kcards, respectively. We begin by writing the second stack in a line, with spaces in between each pair of consecutive cards, and with spaces at the beginning and end (so there aren¡k+ 1 spaces). We choose, with replacement, kof these spaces, and place the cards from the flrst stack in the chosen spaces. This can be done in µn k¶ ways. Thus, the probability of a given interleaving should be 1 ¡n k¢: Next, we note that if the new ordering is not the identity ordering, it is the result of a unique cut-interleaving pair. If the new ordering is the identity, it is theresult of any one of n+ 1 cut-interleaving pairs. We deflne a rising sequence in an ordering to be a maximal subsequence of consecutive integers in increasing order. For example, in the ordering (2;3;5;1;4;7;6); 28B. Mann, \How Many Times Should You Shu†e a Deck of Cards?", UMAP Journal , vol. 15, no. 4 (1994), pp. 303{331. 29D. Bayer and P. Diaconis, \Trailing the Dovetail Shu†e to its Lair," Annals of Applied Prob- ability , vol. 2, no. 2 (1992), pp. 294{313. 3.3. CARD SHUFFLING 121 there are 4 rising sequences; they are (1), (2 ;3;4), (5;6), and (7). It is easy to see that an ordering is the result of a ri†e shu†e applied to the identity ordering ifand only if it has no more than two rising sequences. (If the ordering has two risingsequences, then these rising sequences correspond to the two stacks induced by thecut, and if the ordering has one rising sequence, then it is the identity ordering.)Thus, the sample space of orderings obtained by applying a ri†e shu†e to theidentity ordering is naturally described as the set of all orderings with at most tworising sequences. It is now easy to assign a probability measure to this sample space. Each ordering with two rising sequences is assigned the value b(n;1=2;k) ¡n k¢ =1 2n; and the identity ordering is assigned the value n+1 2n: There is another way to view a ri†e shu†e. We can imagine starting with a deck cut into two stacks as before, with the same probabilities assignment as beforei.e., the binomial distribution. Once we have the two stacks, we take cards, one byone, ofi of the bottom of the two stacks, and place them onto one stack. If therearek 1andk2cards, respectively, in the two stacks at some point in this process, then we make the assumption that the probabilities that the next card to be takencomes from a given stack is proportional to the current stack size. This implies thatthe probability that we take the next card from the flrst stack equals k 1 k1+k2; and the corresponding probability for the second stack is k2 k1+k2: We shall now show that this process assigns the uniform probability to each of the possible interleavings of the two stacks. Suppose, for example, that an interleaving came about as the result of choosing cards from the two stacks in some order. The probability that this result occurredis the product of the probabilities at each point in the process, since the choiceof card at each point is assumed to be independent of the previous choices. Eachfactor of this product is of the form k i k1+k2; wherei= 1 or 2, and the denominator of each factor equals the number of cards left to be chosen. Thus, the denominator of the probability is just n!. At the moment when a card is chosen from a stack that has icards in it, the numerator of the 122 CHAPTER 3. COMBINATORICS corresponding factor in the probability is i, and the number of cards in this stack decreases by 1. Thus, the numerator is seen to be k!(n¡k)!, since all cards in both stacks are eventually chosen. Therefore, this process assigns the probability 1¡n k¢ to each possible interleaving. We now turn to the question of what happens when we ri†e shu†e stimes. It should be clear that if we start with the identity ordering, we obtain an orderingwith at most 2 srising sequences, since a ri†e shu†e creates at most two rising sequences from every rising sequence in the starting ordering. In fact, it is not hardto see that each such ordering is the result of sri†e shu†es. The question becomes, then, in how many ways can an ordering with rrising sequences come about by applyingsri†e shu†es to the identity ordering? In order to answer this question, we turn to the idea of an a-shu†e. a-Shu†es There are several ways to visualize an a-shu†e. One way is to imagine a creature withahands who is given a deck of cards to ri†e shu†e. The creature naturally cuts the deck into astacks, and then ri†es them together. (Imagine that!) Thus, the ordinary ri†e shu†e is a 2-shu†e. As in the case of the ordinary 2-shu†e, weallow some of the stacks to have 0 cards. Another way to visualize an a-shu†e is to think about its inverse, called an a-unshu†e. This idea is described in the proof of the next theorem. We will now show that an a-shu†e followed by a b-shu†e is equivalent to an ab- shu†e. This means, in particular, that sri†e shu†es in succession are equivalent to one 2 s-shu†e. This equivalence is made precise by the following theorem. Theorem 3.9 Letaandbbe two positive integers. Let Sa;bbe the set of all ordered pairs in which the flrst entry is an a-shu†e and the second entry is a b-shu†e. Let Sabbe the set of all ab-shu†es. Then there is a 1-1 correspondence between Sa;b andSabwith the following property. Suppose that ( T1;T2) corresponds to T3.I f T1is applied to the identity ordering, and T2is applied to the resulting ordering, then the flnal ordering is the same as the ordering that is obtained by applying T3 to the identity ordering. Proof. The easiest way to describe the required correspondence is through the idea of an unshu†e. An a-unshu†e begins with a deck of ncards. One by one, cards are taken from the top of the deck and placed, with equal probability, on the bottomof any one of astacks, where the stacks are labelled from 0 to a¡1. After all of the cards have been distributed, we combine the stacks to form one stack by placingstackion top of stack i+1, for 0•i•a¡1. It is easy to see that if one starts with a deck, there is exactly one way to cut the deck to obtain the astacks generated by thea-unshu†e, and with these astacks, there is exactly one way to interleave them 3.3. CARD SHUFFLING 123 to obtain the deck in the order that it was in before the unshu†e was performed. Thus, this a-unshu†e corresponds to a unique a-shu†e, and this a-shu†e is the inverse of the original a-unshu†e. If we apply an ab-unshu†eU3to a deck, we obtain a set of abstacks, which are then combined, in order, to form one stack. We label these stacks with orderedpairs of integers, where the flrst coordinate is between 0 and a¡1, and the second coordinate is between 0 and b¡1. Then we label each card with the label of its stack. The number of possible labels is ab, as required. Using this labelling, we can describe how to flnd a b-unshu†e and an a-unshu†e, such that if these two unshu†es are applied in this order to the deck, we obtain the same set of abstacks as were obtained by the ab-unshu†e. To obtain the b-unshu†eU 2, we sort the deck into bstacks, with the ith stack containing all of the cards with second coordinate i, for 0•i•b¡1. Then these stacks are combined to form one stack. The a-unshu†eU1proceeds in the same manner, except that the flrst coordinates of the labels are used. The resulting a stacks are then combined to form one stack. The above description shows that the cards ending up on top are all those labelled (0;0). These are followed by those labelled (0 ;1);(0;2); :::; (0;b¡ 1);(1;0);(1;1);:::; (a¡1;b¡1). Furthermore, the relative order of any pair of cards with the same labels is never altered. But this is exactly the same as anab-unshu†e, if, at the beginning of such an unshu†e, we label each of the cards with one of the labels (0 ;0);(0;1);:::; (0;b¡1);(1;0);(1;1);:::; (a¡1;b¡1). This completes the proof. 2 In Figure 3.11, we show the labels for a 2-unshu†e of a deck with 10 cards. There are 4 cards with the label 0 and 6 cards with the label 1, so if the 2-unshu†eis performed, the flrst stack will have 4 cards and the second stack will have 6 cards.When this unshu†e is performed, the deck ends up in the identity ordering. In Figure 3.12, we show the labels for a 4-unshu†e of the same deck (because there are four labels being used). This flgure can also be regarded as an example ofa pair of 2-unshu†es, as described in the proof above. The flrst 2-unshu†e will usethe second coordinate of the labels to determine the stacks. In this case, the twostacks contain the cards whose values are f5;1;6;2;7gandf8;9;3;4;10g: After this 2-unshu†e has been performed, the deck is in the order shown in Fig- ure 3.11, as the reader should check. If we wish to perform a 4-unshu†e on thedeck, using the labels shown, we sort the cards lexicographically, obtaining the fourstacks f1;2g;f3;4g;f5;6;7g;andf8;9;10g: When these stacks are combined, we once again obtain the identity ordering of the deck. The point of the above theorem is that both sorting procedures always leadto the same initial ordering. 124 CHAPTER 3. COMBINATORICS Figure 3.11: Before a 2-unshu†e. Figure 3.12: Before a 4-unshu†e. 3.3. CARD SHUFFLING 125 Theorem 3.10 IfDis any ordering that is the result of applying an a-shu†e and then ab-shu†e to the identity ordering, then the probability assigned to Dby this pair of operations is the same as the probability assigned to Dby the process of applying an ab-shu†e to the identity ordering. Proof. Call the sample space of a-shu†esSa. If we label the stacks by the integers from 0 toa¡1, then each cut-interleaving pair, i.e., shu†e, corresponds to exactly onen-digit base ainteger, where the ith digit in the integer is the stack of which theith card is a member. Thus, the number of cut-interleaving pairs is equal to the number of n-digit base aintegers, which is an. Of course, not all of these pairs leads to difierent orderings. The number of pairs leading to a given orderingwill be discussed later. For our purposes it is enough to point out that it is thecut-interleaving pairs that determine the probability assignment. The previous theorem shows that there is a 1-1 correspondence between S a;band Sab. Furthermore, corresponding elements give the same ordering when applied to the identity ordering. Given any ordering D, letm1be the number of elements ofSa;bwhich, when applied to the identity ordering, result in D. Letm2be the number of elements of Sabwhich, when applied to the identity ordering, result in D. The previous theorem implies that m1=m2. Thus, both sets assign the probability m1 (ab)n toD. This completes the proof. 2 Connection with the Birthday Problem There is another point that can be made concerning the labels given to the cards by the successive unshu†es. Suppose that we 2-unshu†e an n-card deck until the labels on the cards are all difierent. It is easy to see that this process produceseach permutation with the same probability, i.e., this is a random process. To seethis, note that if the labels become distinct on the sth 2-unshu†e, then one can think of this sequence of 2-unshu†es as one 2 s-unshu†e, in which all of the stacks determined by the unshu†e have at most one card in them (remember, the stackscorrespond to the labels). If each stack has at most one card in it, then given anytwo cards in the deck, it is equally likely that the flrst card has a lower or a higherlabel than the second card. Thus, each possible ordering is equally likely to resultfrom this 2 s-unshu†e. LetTbe the random variable that counts the number of 2-unshu†es until all labels are distinct. One can think of Tas giving a measure of how long it takes in the unshu†ing process until randomness is reached. Since shu†ing and unshu†ingare inverse processes, Talso measures the number of shu†es necessary to achieve randomness. Suppose that we have an n-card deck, and we ask for P(T•s). This equals 1¡P(T>s ). ButT>s if and only if it is the case that not all of the labels after s2-unshu†es are distinct. This is just the birthday problem; we are asking for the probability that at least two people have the same birthday, given 126 CHAPTER 3. COMBINATORICS that we have npeople and there are 2spossible birthdays. Using our formula from Example 3.3, we flnd that P(T>s )=1¡µ2s n¶n! 2sn: (3.4) In Chapter 6, we will deflne the average value of a random variable. Using this idea, and the above equation, one can calculate the average value of the randomvariableT(see Exercise 6.1.41). For example, if n= 52, then the average value of Tis about 11.7. This means that, on the average, about 12 ri†e shu†es are needed for the process to be considered random. Cut-Interleaving Pairs and Orderings As was noted in the proof of Theorem 3.10, not all of the cut-interleaving pairs leadto difierent orderings. However, there is an easy formula which gives the number ofsuch pairs that lead to a given ordering. Theorem 3.11 If an ordering of length nhasrrising sequences, then the number of cut-interleaving pairs under an a-shu†e of the identity ordering which lead to the ordering isµn+a¡r n¶ : Proof. To see why this is true, we need to count the number of ways in which the cut in ana-shu†e can be performed which will lead to a given ordering with rrising sequences. We can disregard the interleavings, since once a cut has been made, atmost one interleaving will lead to a given ordering. Since the given ordering hasrrising sequences, r¡1 of the division points in the cut are determined. The remaininga¡1¡(r¡1) =a¡rdivision points can be placed anywhere. The number of places to put these remaining division points is n+ 1 (which is the number of spaces between the consecutive pairs of cards, including the positions atthe beginning and the end of the deck). These places are chosen with repetitionallowed, so the number of ways to make these choices is µn+a¡r a¡r¶ =µn+a¡r n¶ : In particular, this means that if Dis an ordering that is the result of applying ana-shu†e to the identity ordering, and if Dhasrrising sequences, then the probability assigned to Dby this process is ¡ n+a¡r n¢ an: This completes the proof. 2 3.3. CARD SHUFFLING 127 The above theorem shows that the essential information about the probability assigned to an ordering under an a-shu†e is just the number of rising sequences in the ordering. Thus, if we determine the number of orderings which contain exactlyrrising sequences, for each rbetween 1 and n, then we will have determined the distribution function of the random variable which consists of applying a randoma-shu†e to the identity ordering. The number of orderings of f1;2;:::;ngwithrrising sequences is denoted by A(n;r), and is called an Eulerian number. There are many ways to calculate the values of these numbers; the following theorem gives one recursive method whichfollows immediately from what we already know about a-shu†es. Theorem 3.12 Letaandnbe positive integers. Then a n=aX r=1µn+a¡r n¶ A(n;r): (3.5) Thus, A(n;a)=an¡a¡1X r=1µn+a¡r n¶ A(n;r): In addition, A(n;1 )=1: Proof. The second equation can be used to calculate the values of the Eulerian numbers, and follows immediately from the Equation 3.5. The last equation isa consequence of the fact that the only ordering of f1;2;:::;ngwith one rising sequence is the identity ordering. Thus, it remains to prove Equation 3.5. We willcount the set of a-shu†es of a deck with ncards in two ways. First, we know that there area nsuch shu†es (this was noted in the proof of Theorem 3.10). But there areA(n;r) orderings off1;2;:::;ngwithrrising sequences, and Theorem 3.11 states that for each such ordering, there are exactly µn+a¡r n¶ cut-interleaving pairs that lead to the ordering. Therefore, the right-hand side of Equation 3.5 counts the set of a-shu†es of an n-card deck. This completes the proof. 2 Random Orderings and Random Processes We now turn to the second question that was asked at the beginning of this section: What do we mean by a \random" ordering? It is somewhat misleading to thinkabout a given ordering as being random or not random. If we want to choose arandom ordering from the set of all orderings of f1;2;:::;ng, we mean that we want every ordering to be chosen with the same probability, i.e., any ordering is as\random" as any other. 128 CHAPTER 3. COMBINATORICS The word \random" should really be used to describe a process. We will say that a process that produces an object from a (flnite) set of objects is a random processif each object in the set is produced with the same probability by the process. Inthe present situation, the objects are the orderings, and the process which producesthese objects is the shu†ing process. It is easy to see that no a-shu†e is really a random process, since if T 1andT2are two orderings with a difierent number of rising sequences, then they are produced by an a-shu†e, applied to the identity ordering, with difierent probabilities. Variation Distance Instead of requiring that a sequence of shu†es yield a process which is random, wewill deflne a measure that describes how far away a given pro cess is from a random process. Let Xbe any process which produces an ordering of f1;2;:::;ng. Deflne f X(…) be the probability that Xproduces the ordering …. (Thus,Xcan be thought of as a random variable with distribution function f.) Let › nbe the set of all orderings off1;2;:::;ng. Finally, let u(…)=1=j›njfor all…2›n. The function uis the distribution function of a process which produces orderings and which is random. For each ordering …2›n, the quantity jfX(…)¡u(…)j is the difierence between the actual and desired probabilities that Xproduces….I f we sum this over all orderings …and call this sum S, we see that S= 0 if and only ifXis random, and otherwise Sis positive. It is easy to show that the maximum value ofSis 2, so we will multiply the sum by 1 =2 so that the value falls in the interval [0;1]. Thus, we obtain the following sum as the formula for the variation distance between the two processes: kfX¡uk=1 2X …2›njfX(…)¡u(…)j: Now we apply this idea to the case of shu†ing. We let Xbe the process of s successive ri†e shu†es applied to the identity ordering. We know that it is alsopossible to think of Xas one 2 s-shu†e. We also know that fXis constant on the set of all orderings with rrising sequences, where ris any positive integer. Finally, we know the value of fXon an ordering with rrising sequences, and we know how many such orderings there are. Thus, in this speciflc case, we have kfX¡uk=1 2nX r=1A(n;r)flflflflµ2 s+n¡r n¶ =2ns¡1 n!flflflfl: Since this sum has only nsummands, it is easy to compute this for moderate sized values ofn.F o rn= 52, we obtain the list of values given in Table 3.14. To help in understanding these data, they are shown in graphical form in Fig- ure 3.13. The program VariationList produces the data shown in both Table 3.14 and Figure 3.13. One sees that until 5 shu†es have occurred, the output of Xis 3.3. CARD SHUFFLING 129 Number of Ri†e Shu†es Variation Distance 11 21314 0.99999953345 0.92373292946 0.61354959667 0.33406099958 0.16715864199 0.0854201934 10 0.042945548911 0.021502376012 0.010754893513 0.005377910114 0.0026890130 Table 3.14: Distance to the random process. 5 10 15 200.20.40.60.81 Figure 3.13: Distance to the random process. 130 CHAPTER 3. COMBINATORICS very far from random. After 5 shu†es, the distance from the random process is essentially halved each time a shu†e occurs. Given the distribution functions fX(…) andu(…) as above, there is another way to view the variation distance kfX¡uk. Given any event T(which is a subset ofSn), we can calculate its probability under the process Xand under the uniform process. For example, we can imagine that Trepresents the set of all permutations in which the flrst player in a 7-player poker game is dealt a straight°ush (flve consecutive cards in the same suit). It is interesting to consider howmuch the probability of this event after a certain number of shu†es difiers from theprobability of this event if all permutations are equally likely. This difierence canbe thought of as describing how close the process Xis to the random process with respect to the event T. Now consider the event Tsuch that the absolute value of the difierence between these two probabilities is as large as possible. It can be shown that this absolutevalue is the variation distance between the process Xand the uniform process. (The reader is asked to prove this fact in Exercise 4.) We have just seen that, for a deck of 52 cards, the variation distance between the 7-ri†e shu†e process and the random process is about :334. It is of interest to flnd an event Tsuch that the difierence between the probabilities that the two processes produce Tis close to:334. An event with this property can be described in terms of the game called New-Age Solitaire. New-Age Solitaire This game was invented by Peter Doyle. It is played with a standard 52-card deck.We deal the cards face up, one at a time, onto a discard pile. If an ace is encountered,say the ace of Hearts, we use it to start a Heart pile. Each suit pile must be builtup in order, from ace to king, using only subsequently dealt cards. Once we havedealt all of the cards, we pick up the discard pile and continue. We deflne the Yinsuits to be Hearts and Clubs, and the Yang suits to be Diamonds and Spades. Thegame ends when either both Yin suit piles have been completed, or both Yang suitpiles have been completed. It is clear that if the ordering of the deck is producedby the random process, then the probability that the Yin suit piles are completedflrst is exactly 1/2. Now suppose that we buy a new deck of cards, break the seal on the package, and ri†e shu†e the deck 7 times. If one tries this, one flnds that the Yin suits winabout 75% of the time. This is 25% more than we would get if the deck were intruly random order. This deviation is reasonably close to the theoretical maximumof 33:4% obtained above. Why do the Yin suits win so often? In a brand new deck of cards, the suits are in the following order, from top to bottom: ace through king of Hearts, ace throughking of Clubs, king through ace of Diamonds, and king through ace of Spades. Notethat if the cards were not shu†ed at all, then the Yin suit piles would be completedon the flrst pass, before any Yang suit cards are even seen. If we were to continueplaying the game until the Yang suit piles are completed, it would take 13 passes 3.3. CARD SHUFFLING 131 through the deck to do this. Thus, one can see that in a new deck, the Yin suits are in the most advantageous order and the Yang suits are in the least advantageousorder. Under 7 ri†e shu†es, the relative advantage of the Yin suits over the Yangsuits is preserved to a certain extent. Exercises 1Given any ordering ¾off1;2;:::;ng, we can deflne ¾¡1, the inverse ordering of¾, to be the ordering in which the ith element is the position occupied by iin¾. For example, if ¾=( 1;3;5;2;4;7;6), then¾¡1=( 1;4;2;5;3;7;6). (If one thinks of these orderings as permutations, then ¾¡1is the inverse of ¾.) Afalloccurs between two positions in an ordering if the left position is occu- pied by a larger number than the right position. It will be convenient to saythat every ordering has a fall after the last position. In the above example,¾ ¡1has four falls. They occur after the second, fourth, sixth, and seventh positions. Prove that the number of rising sequences in an ordering ¾equals the number of falls in ¾¡1. 2Show that if we start with the identity ordering of f1;2;:::;ng, then the prob- ability that an a-shu†e leads to an ordering with exactly rrising sequences equals¡n+a¡r n¢ anA(n;r); for 1•r•a. 3LetDb ead e c ko f ncards. We have seen that there are ana-shu†es of D. A coding of the set of a-unshu†es was given in the proof of Theorem 3.9. We will now give a coding of the a-shu†es which corresponds to the coding of thea-unshu†es. Let Sbe the set of all n-tuples of integers, each between 0 anda¡1. LetM=(m1;m2;:::;mn) be any element of S. Letnibe the number ofi’s inM, for 0•i•a¡1. Suppose that we start with the deck in increasing order (i.e., the cards are numbered from 1 to n). We label the flrstn0cards with a 0, the next n1cards with a 1, etc. Then the a-shu†e corresponding to Mis the shu†e which results in the ordering in which the cards labelled iare placed in the positions in Mcontaining the label i. The cards with the same label are placed in these positions in increasing order oftheir numbers. For example, if n= 6 anda= 3, letM=( 1;0;2;2;0;2). Thenn 0=2;n1=1;andn2= 3. So we label cards 1 and 2 with a 0, card 3 with a 1, and cards 4, 5, and 6 with a 2. Then cards 1 and 2 are placedin positions 2 and 5, card 3 is placed in position 1, and cards 4, 5, and 6 areplaced in positions 3, 4, and 6, resulting in the ordering (3 ;1;4;5;2;6). (a) Using this coding, show that the probability that in an a-shu†e, the flrst card (i.e., card number 1) moves to the ith position, is given by the following expression: (a¡1) i¡1an¡i+(a¡2)i¡1(a¡1)n¡i+¢¢¢+1i¡12n¡i an: 132 CHAPTER 3. COMBINATORICS (b) Give an accurate estimate for the probability that in three ri†e shu†es of a 52-card deck, the flrst card ends up in one of the flrst 26 positions.Using a computer, accurately estimate the probability of the same eventafter seven ri†e shu†es. 4LetXdenote a particular process that produces elements of S n, and letU denote the uniform process. Let the distribution functions of these processesbe denoted by f Xandu, respectively. Show that the variation distance kfX¡ukis equal to max T‰SnX …2T‡ fX(…)¡u(…)· : Hint: Write the permutations in Snin decreasing order of the difierence fX(…)¡u(…). 5Consider the process described in the text in which an n-card deck is re- peatedly labelled and 2-unshu†ed, in the manner described in the proof ofTheorem 3.9. (See Figures 3.10 and 3.13.) The process continues until thelabels are all difierent. Show that the process never terminates until at leastdlog 2(n)eunshu†es have been done. Chapter 4 Conditional Probability 4.1 Discrete Conditional Probability Conditional Probability In this section we ask and answer the following question. Suppose we assign a distribution function to a sample space and then learn that an event Ehas occurred. How should we change the probabilities of the remaining events? We shall call thenew probability for an event Ftheconditional probability of FgivenEand denote it byP(FjE). Example 4.1 An experiment consists of rolling a die once. Let Xbe the outcome. LetFbe the eventfX=6g, and letEbe the eventfX> 4g. We assign the distribution function m(!)=1=6 for!=1;2;:::; 6. Thus,P(F)=1=6. Now suppose that the die is rolled and we are told that the event Ehas occurred. This leaves only two possible outcomes: 5 and 6. In the absence of any other information,we would still regard these outcomes to be equally likely, so the probability of F becomes 1/2, making P(FjE)=1=2. 2 Example 4.2 In the Life Table (see Appendix C), one flnds that in a population of 100,000 females, 89.835% can expect to live to age 60, while 57.062% can expectto live to age 80. Given that a woman is 60, what is the probability that she livesto age 80? This is an example of a conditional probability. In this case, the original sample space can be thought of as a set of 100,000 females. The events EandFare the subsets of the sample space consisting of all women who live at least 60 years, andat least 80 years, respectively. We consider Eto be the new sample space, and note thatFis a subset of E. Thus, the size of Eis 89,835, and the size of Fis 57,062. So, the probability in question equals 57 ;062=89;835 =:6352. Thus, a woman who is 60 has a 63.52% chance of living to age 80. 2 133 134 CHAPTER 4. CONDITIONAL PROBABILITY Example 4.3 Consider our voting example from Section 1.2: three candidates A, B, and C are running for o–ce. We decided that A and B have an equal chance ofwinning and C is only 1/2 as likely to win as A. Let Abe the event \A wins," B that \B wins," and Cthat \C wins." Hence, we assigned probabilities P(A)=2=5, P(B)=2=5, andP(C)=1=5. Suppose that before the election is held, Adrops out of the race. As in Exam- ple 4.1, it would be natural to assign new probabilities to the events BandCwhich are proportional to the original probabilities. Thus, we would have P(BjA)=2=3, andP(CjA)=1=3. It is important to note that any time we assign probabilities to real-life events, the resulting distribution is only useful if we take into accountall relevant information. In this example, we may have knowledge that most voterswho favorAwill vote for CifAis no longer in the race. This will clearly make the probability that Cwins greater than the value of 1/3 that was assigned above. 2 In these examples we assigned a distribution function and then were given new information that determined a new sample space, consisting of the outcomes thatare still possible, and caused us to assign a new distribution function to this space. We want to make formal the procedure carried out in these examples. Let ›=f! 1;!2;:::;!rgbe the original sample space with distribution function m(!j) assigned. Suppose we learn that the event Ehas occurred. We want to assign a new distribution function m(!jjE) to › to re°ect this fact. Clearly, if a sample point !j is not inE, we wantm(!jjE) = 0. Moreover, in the absence of information to the contrary, it is reasonable to assume that the probabilities for !kinEshould have the same relative magnitudes that they had before we learned that Ehad occurred. For this we require that m(!kjE)=cm(!k) for all!kinE, withcsome positive constant. But we must also have X Em(!kjE)=cX Em(!k)=1: Thus, c=1P Em(!k)=1 P(E): (Note that this requires us to assume that P(E)>0.) Thus, we will deflne m(!kjE)=m(!k) P(E) for!kinE. We will call this new distribution the conditional distribution givenE. For a general event F, this gives P(FjE)=X F\Em(!kjE)=X F\Em(!k) P(E)=P(F\E) P(E): We callP(FjE) the conditional probability of Foccurring given that Eoccurs, and compute it using the formula P(FjE)=P(F\E) P(E): 4.1. DISCRETE CONDITIONAL PROBABILITY 135 (start)p (ω) ω ω ω ω ω 1/2 1/2l ll2/5 3/5 1/21/2bw wb1/5 3/10 1/4 1/4Urn Color of ball 1 2 3 4 Figure 4.1: Tree diagram. Example 4.4 (Example 4.1 continued) Let us return to the example of rolling a die. Recall that Fis the event X= 6, andEis the event X> 4. Note that E\F is the event F. So, the above formula gives P(FjE)=P(F\E) P(E) =1=6 1=3 =1 2; in agreement with the calculations performed earlier. 2 Example 4.5 We have two urns, I and II. Urn I contains 2 black balls and 3 white balls. Urn II contains 1 black ball and 1 white ball. An urn is drawn at randomand a ball is chosen at random from it. We can represent the sample space of thisexperiment as the paths through a tree as shown in Figure 4.1. The probabilitiesassigned to the paths are also shown. LetBbe the event \a black ball is drawn," and Ithe event \urn I is chosen." Then the branch weight 2/5, which is shown on one branch in the flgure, can nowbe interpreted as the conditional probability P(BjI). Suppose we wish to calculate P(IjB). Using the formula, we obtain P(IjB)=P(I\B) P(B) =P(I\B) P(B\I)+P(B\II) =1=5 1=5+1=4=4 9: 2 136 CHAPTER 4. CONDITIONAL PROBABILITY (start)p (ω) ω ω ω ω ω 9/20 11/20b w4/9 5/9 5/116/11III III1/5 3/101/4 1/4Urn Color of ball 1 3 2 4 Figure 4.2: Reverse tree diagram. Bayes Probabilities Our original tree measure gave us the probabilities for drawing a ball of a given color, given the urn chosen. We have just calculated the inverse probability that a particular urn was chosen, given the color of the ball. Such an inverse probability iscalled a Bayes probability and may be obtained by a formula that we shall develop later. Bayes probabilities can also be obtained by simply constructing the treemeasure for the two-stage experiment carried out in reverse order. We show thistree in Figure 4.2. The paths through the reverse tree are in one-to-one correspondence with those in the forward tree, since they correspond to individual outcomes of the experiment,and so they are assigned the same probabilities. From the forward tree, we flnd thatthe probability of a black ball is 1 2¢2 5+1 2¢1 2=9 20: The probabilities for the branches at the second level are found by simple divi- sion. For example, if xis the probability to be assigned to the top branch at the second level, we must have 9 20¢x=1 5 orx=4=9. Thus,P(IjB)=4=9, in agreement with our previous calculations. The reverse tree then displays all of the inverse, or Bayes, probabilities. Example 4.6 We consider now a problem called the Monty Hall problem. This has long been a favorite problem but was revived by a letter from Craig Whitakerto Marilyn vos Savant for consideration in her column in Parade Magazine . 1Craig wrote: 1Marilyn vos Savant, Ask Marilyn, Parade Magazine , 9 September; 2 December; 17 February 1990, reprinted in Marilyn vos Savant, Ask Marilyn , St. Martins, New York, 1992. 4.1. DISCRETE CONDITIONAL PROBABILITY 137 Suppose you’re on Monty Hall’s Let’s Make a Deal! You are given the choice of three doors, behind one door is a car, the others, goats. Youpick a door, say 1, Monty opens another door, say 3, which has a goat.Monty says to you \Do you want to pick door 2?" Is it to your advantageto switch your choice of doors? Marilyn gave a solution concluding that you should switch, and if you do, your probability of winning is 2/3. Several irate readers, some of whom identifled them-selves as having a PhD in mathematics, said that this is absurd since after Montyhas ruled out one door there are only two possible doors and they should still eachhave the same probability 1/2 so there is no advantage to switching. Marilyn stuckto her solution and encouraged her readers to simulate the game and draw their ownconclusions from this. We also encourage the reader to do this (see Exercise 11). Other readers complained that Marilyn had not described the problem com- pletely. In particular, the way in which certain decisions were made during a playof the game were not specifled. This aspect of the problem will be discussed in Sec-tion 4.3. We will assume that the car was put behind a door by rolling a three-sideddie which made all three choices equally likely. Monty knows where the car is, andalways opens a door with a goat behind it. Finally, we assume that if Monty hasa choice of doors (i.e., the contestant has picked the door with the car behind it),he chooses each door with probability 1/2. Marilyn clearly expected her readers toassume that the game was played in this manner. As is the case with most apparent paradoxes, this one can be resolved through careful analysis. We begin by describing a simpler, related question. We say thata contestant is using the \stay" strategy if he picks a door, and, if ofiered a chanceto switch to another door, declines to do so (i.e., he stays with his original choice).Similarly, we say that the contestant is using the \switch" strategy if he picks a door,and, if ofiered a chance to switch to another door, takes the ofier. Now supposethat a contestant decides in advance to play the \stay" strategy. His only actionin this case is to pick a door (and decline an invitation to switch, if one is ofiered).What is the probability that he wins a car? The same question can be asked aboutthe \switch" strategy. Using the \stay" strategy, a contestant will win the car with probability 1/3, since 1/3 of the time the door he picks will have the car behind it. On the otherhand, if a contestant plays the \switch" strategy, then he will win whenever thedoor he originally picked does not have the car behind it, which happens 2/3 of thetime. This very simple analysis, though correct, does not quite solve the problem that Craig posed. Craig asked for the conditional probability that you win if youswitch, given that you have chosen door 1 and that Monty has chosen door 3. Tosolve this problem, we set up the problem before getting this information and thencompute the conditional probability given this information. This is a process thattakes place in several stages; the car is put behind a door, the contestant picks adoor, and flnally Monty opens a door. Thus it is natural to analyze this using atree measure. Here we make an additional assumption that if Monty has a choice 138 CHAPTER 4. CONDITIONAL PROBABILITY 1/3 1/31/31/31/31/3 1/3 1/31/31/311/2 1/21/21/21/2 1 1/2111 1Door opened by MontyDoor chosen by contestantPath probabilitiesPlacement of car 1 2 31 2 3 11 22 3 32 3 323 3 1 1 21 121/18 1/18 1/181/181/18 1/91/9 1/9 1/18 1/9 1/9 1/91/31/3 Figure 4.3: The Monty Hall problem. of doors (i.e., the contestant has picked the door with the car behind it) then he picks each door with probability 1/2. The assumptions we have made determine thebranch probabilities and these in turn determine the tree measure. The resultingtree and tree measure are shown in Figure 4.3. It is tempting to reduce the tree’ssize by making certain assumptions such as: \Without loss of generality, we willassume that the contestant always picks door 1." We have chosen not to make anysuch assumptions, in the interest of clarity. Now the given information, namely that the contestant chose door 1 and Monty chose door 3, means only two paths through the tree are possible (see Figure 4.4).For one of these paths, the car is behind door 1 and for the other it is behind door2. The path with the car behind door 2 is twice as likely as the one with the carbehind door 1. Thus the conditional probability is 2/3 that the car is behind door 2and 1/3 that it is behind door 1, so if you switch you have a 2/3 chance of winningthe car, as Marilyn claimed. At this point, the reader may think that the two problems above are the same, since they have the same answers. Recall that we assumed in the original problem 4.1. DISCRETE CONDITIONAL PROBABILITY 139 1/3 1/31/31/2 1Door opened by MontyDoor chosen by contestantUnconditional probabilityPlacement of car 1 21 133 1/18 1/9 1/3Conditional probability 1/3 2/3 Figure 4.4: Conditional probabilities for the Monty Hall problem. if the contestant chooses the door with the car, so that Monty has a choice of two doors, he chooses each of them with probability 1/2. Now suppose instead thatin the case that he has a choice, he chooses the door with the larger number withprobability 3/4. In the \switch" vs. \stay" problem, the probability of winningwith the \switch" strategy is still 2/3. However, in the original problem, if thecontestant switches, he wins with probability 4/7. The reader can check this bynoting that the same two paths as before are the only two possible paths in thetree. The path leading to a win, if the contestant switches, has probability 1/3,while the path which leads to a loss, if the contestant switches, has probability 1/4.2 Independent Events It often happens that the knowledge that a certain event Ehas occurred has no efiect on the probability that some other event Fhas occurred, that is, that P(FjE)= P(F). One would expect that in this case, the equation P(EjF)=P(E) would also be true. In fact (see Exercise 1), each equation implies the other. If theseequations are true, we might say the Fisindependent ofE. For example, you would not expect the knowledge of the outcome of the flrst toss of a coin to changethe probability that you would assign to the possible outcomes of the second toss,that is, you would not expect that the second toss depends on the flrst. This ideais formalized in the following deflnition of independent events. Deflnition 4.1 Two events EandFareindependent if bothEandFhave positive probability and if P(EjF)=P(E); and P(FjE)=P(F): 2 140 CHAPTER 4. CONDITIONAL PROBABILITY As noted above, if both P(E) andP(F) are positive, then each of the above equations imply the other, so that to see whether two events are independent, onlyone of these equations must be checked (see Exercise 1). The following theorem provides another way to check for independence. Theorem 4.1 IfP(E)>0 andP(F)>0, thenEandFare independent if and only if P(E\F)=P(E)P(F): Proof. Assume flrst that EandFare independent. Then P(EjF)=P(E), and so P(E\F)=P(EjF)P(F) =P(E)P(F): Assume next that P(E\F)=P(E)P(F). Then P(EjF)=P(E\F) P(F)=P(E): Also, P(FjE)=P(F\E) P(E)=P(F): Therefore,EandFare independent. 2 Example 4.7 Suppose that we have a coin which comes up heads with probability p, and tails with probability q. Now suppose that this coin is tossed twice. Using a frequency interpretation of probability, it is reasonable to assign to the outcome(H;H ) the probability p 2, to the outcome ( H;T) the probability pq, and so on. Let Ebe the event that heads turns up on the flrst toss and Fthe event that tails turns up on the second toss. We will now check that with the above probabilityassignments, these two events are independent, as expected. We have P(E)= p 2+pq=p,P(F)=pq+q2=q. FinallyP(E\F)=pq,s oP(E\F)= P(E)P(F). 2 Example 4.8 It is often, but not always, intuitively clear when two events are independent. In Example 4.7, let Abe the event \the flrst toss is a head" and B the event \the two outcomes are the same." Then P(BjA)=P(B\A) P(A)=PfHHg PfHH,HTg=1=4 1=2=1 2=P(B): Therefore,AandBare independent, but the result was not so obvious. 2 4.1. DISCRETE CONDITIONAL PROBABILITY 141 Example 4.9 Finally, let us give an example of two events that are not indepen- dent. In Example 4.7, let Ibe the event \heads on the flrst toss" and Jthe event \two heads turn up." Then P(I)=1=2 andP(J)=1=4. The event I\Jis the event \heads on both tosses" and has probability 1 =4. Thus,IandJare not independent sinceP(I)P(J)=1=86=P(I\J). 2 We can extend the concept of independence to any flnite set of events A1,A2, ...,An. Deflnition 4.2 A set of eventsfA1;A2;:::;Angis said to be mutually indepen- dent if for any subsetfAi;Aj;:::; Amgof these events we have P(Ai\Aj\¢¢¢\Am)=P(Ai)P(Aj)¢¢¢P(Am); or equivalently, if for any sequence „A1,„A2, ..., „Anwith „Aj=Ajor~Aj, P(„A1\„A2\¢¢¢\ „An)=P(„A1)P(„A2)¢¢¢P(„An): (For a proof of the equivalence in the case n= 3, see Exercise 33.) 2 Using this terminology, it is a fact that any sequence (S ;S;F;F;S;:::; S) of possible outcomes of a Bernoulli trials process forms a sequence of mutually independentevents. It is natural to ask: If all pairs of a set of events are independent, is the whole set mutually independent? The answer is not necessarily, and an example is given in Exercise 7. It is important to note that the statement P(A 1\A2\¢¢¢\An)=P(A1)P(A2)¢¢¢P(An) does not imply that the events A1,A2, ...,Anare mutually independent (see Exercise 8). Joint Distribution Functions and Independence of Random Variables It is frequently the case that when an experiment is performed, several difierent quantities concerning the outcomes are investigated. Example 4.10 Suppose we toss a coin three times. The basic random variable „Xcorresponding to this experiment has eight possible outcomes, which are the ordered triples consisting of H’s and T’s. We can also deflne the random variableX i, fori=1;2;3, to be the outcome of the ith toss. If the coin is fair, then we should assign the probability 1/8 to each of the eight possible outcomes. Thus, thedistribution functions of X 1,X2, andX3are identical; in each case they are deflned bym(H)=m(T)=1=2. 2 142 CHAPTER 4. CONDITIONAL PROBABILITY If we have several random variables X1;X2;:::;Xnwhich correspond to a given experiment, then we can consider the joint random variable „X=(X1;X2;:::;Xn) deflned by taking an outcome !of the experiment, and writing, as an n-tuple, the corresponding noutcomes for the random variables X1;X2;:::;Xn. Thus, if the random variable Xihas, as its set of possible outcomes the set Ri, then the set of possible outcomes of the joint random variable „Xis the Cartesian product of the Ri’s, i.e., the set of all n-tuples of possible outcomes of the Xi’s. Example 4.11 (Example 4.10 continued) In the coin-tossing example above, let Xidenote the outcome of the ith toss. Then the joint random variable „X= (X1;X2;X3) has eight possible outcomes. Suppose that we now deflne Yi, fori=1;2;3, as the number of heads which occur in the flrst itosses. Then Yihasf0;1;:::;igas possible outcomes, so at flrst glance, the set of possible outcomes of the joint random variable „Y=(Y1;Y2;Y3) should be the set f(a1;a2;a3):0•a1•1;0•a2•2;0•a3•3g: However, the outcome (1 ;0;1) cannot occur, since we must have a1•a2•a3. The solution to this problem is to deflne the probability of the outcome (1 ;0;1) to be 0. We now illustrate the assignment of probabilities to the various outcomes for the joint random variables „Xand „Y. In the flrst case, each of the eight outcomes should be assigned the probability 1/8, since we are assuming that we have a faircoin. In the second case, since Y ihasi+ 1 possible outcomes, the set of possible outcomes has size 24. Only eight of these 24 outcomes can actually occur, namelythe ones satisfying a 1•a2•a3. Each of these outcomes corresponds to exactly one of the outcomes of the random variable „X, so it is natural to assign probability 1/8 to each of these. We assign probability 0 to the other 16 outcomes. In eachcase, the probability function is called a joint distribution function. 2 We collect the above ideas in a deflnition. Deflnition 4.3 LetX 1;X2;:::;Xnbe random variables associated with an exper- iment. Suppose that the sample space (i.e., the set of possible outcomes) of Xiis the setRi. Then the joint random variable „X=(X1;X2;:::;Xn) is deflned to be the random variable whose outcomes consist of ordered n-tuples of outcomes, with theith coordinate lying in the set Ri. The sample space › of „Xis the Cartesian product of the Ri’s: ›=R1£R1£¢¢¢£Rn: The joint distribution function of „Xis the function which gives the probability of each of the outcomes of „X. 2 Example 4.12 (Example 4.10 continued) We now consider the assignment of prob- abilities in the above example. In the case of the random variable „X, the probabil- ity of any outcome ( a1;a2;a3) is just the product of the probabilities P(Xi=ai), 4.1. DISCRETE CONDITIONAL PROBABILITY 143 Not smoke Smoke Total Not cancer 40 10 50 Cancer 73 10 Totals 47 13 60 Table 4.1: Smoking and cancer. S 01 040/60 10/60 C 1 7/60 3/60 Table 4.2: Joint distribution. fori=1;2;3. However, in the case of „Y, the probability assigned to the outcome (1;1;0) is not the product of the probabilities P(Y1= 1),P(Y2= 1), andP(Y3= 0). The difierence between these two situations is that the value of Xidoes not afiect the value of Xj,i fi6=j, while the values of YiandYjafiect one another. For example, if Y1= 1, thenY2cannot equal 0. This prompts the next deflnition. 2 Deflnition 4.4 The random variables X1,X2, ...,Xnaremutually independent if P(X1=r1;X2=r2;:::;Xn=rn) =P(X1=r1)P(X2=r2)¢¢¢P(Xn=rn) for any choice of r1;r2;:::;rn. Thus, ifX1;X2;:::;Xnare mutually independent, then the joint distribution function of the random variable „X=(X1;X2;:::;Xn) is just the product of the individual distribution functions. When two random variables are mutually independent, we shall say more brie°y that they are indepen- dent. 2 Example 4.13 In a group of 60 people, the numbers who do or do not smoke and do or do not have cancer are reported as shown in Table 4.1. Let › be the samplespace consisting of these 60 people. A person is chosen at random from the group.LetC(!) = 1 if this person has cancer and 0 if not, and S(!) = 1 if this person smokes and 0 if not. Then the joint distribution of fC;Sgis given in Table 4.2. For exampleP(C=0;S=0 )=4 0=60,P(C=0;S=1 )=1 0=60, and so forth. The distributions of the individual random variables are called marginal distributions. The marginal distributions of CandSare: p C=µ01 50=60 10=60¶ ; 144 CHAPTER 4. CONDITIONAL PROBABILITY pS=µ01 47=60 13=60¶ : The random variables SandCare not independent, since P(C=1;S=1 ) =3 60=:05; P(C=1 )P(S=1 ) =10 60¢13 60=:036: Note that we would also see this from the fact that P(C=1jS=1 ) =3 13=:23; P(C=1 ) =1 6=:167: 2 Independent Trials Processes The study of random variables proceeds by considering special classes of random variables. One such class that we shall study is the class of independent trials. Deflnition 4.5 A sequence of random variables X1,X2,...,Xnthat are mutually independent and that have the same distribution is called a sequence of independenttrials or an independent trials p rocess. Independent trials processes arise naturally in the following way. We have a single experiment with sample space R=fr 1;r2;:::;rsgand a distribution function mX=µr1r2¢¢¢rs p1p2¢¢¢ps¶ : We repeat this experiment ntimes. To describe this total experiment, we choose as sample space the space ›=R£R£¢¢¢£R; consisting of all possible sequences !=(!1;!2;:::;!n) where the value of each !j is chosen from R. We assign a distribution function to be the product distribution m(!)=m(!1)¢:::¢m(!n); withm(!j)=pkwhen!j=rk. Then we let Xjdenote the jth coordinate of the outcome (r1;r2;:::;rn). The random variables X1, ...,Xnform an independent trials process. 2 Example 4.14 An experiment consists of rolling a die three times. Let Xirepre- sent the outcome of the ith roll, for i=1;2;3. The common distribution function is mi=µ123456 1=61=61=61=61=61=6¶ : 4.1. DISCRETE CONDITIONAL PROBABILITY 145 The sample space is R3=R£R£RwithR=f1;2;3;4;5;6g.I f!=( 1;3;6), thenX1(!)=1 ,X2(!)=3 ,a n d X3(!) = 6 indicating that the flrst roll was a 1, the second was a 3, and the third was a 6. The probability assigned to any samplepoint is m(!)=1 6¢1 6¢1 6=1 216: 2 Example 4.15 Consider next a Bernoulli trials process with probability pfor suc- cess on each experiment. Let Xj(!) = 1 if the jth outcome is success and Xj(!)=0 if it is a failure. Then X1,X2, ...,Xnis an independent trials process. Each Xj has the same distribution function mj=µ01 qp¶ ; whereq=1¡p. IfSn=X1+X2+¢¢¢+Xn, then P(Sn=j)=µn j¶ pjqn¡j; andSnhas, as distribution, the binomial distribution b(n;p;j ). 2 Bayes’ Formula In our examples, we have considered conditional probabilities of the following form: Given the outcome of the second stage of a two-stage experiment, flnd the proba-bility for an outcome at the flrst stage. We have remarked that these probabilitiesare called Bayes probabilities. We return now to the calculation of more general Bayes probabilities. Suppose we have a set of events H 1;H2, ...,Hmthat are pairwise disjoint and such that ›=H1[H2[¢¢¢[Hm: We call these events hypotheses. We also have an event Ethat gives us some information about which hypothesis is correct. We call this event evidence. Before we receive the evidence, then, we have a set of prior probabilities P(H1), P(H2) ,...,P(Hm) for the hypotheses. If we know the correct hypothesis, we know the probability for the evidence. That is, we know P(EjHi) for alli. W ew a n tt o flnd the probabilities for the hypotheses given the evidence. That is, we want to flndthe conditional probabilities P(H ijE). These probabilities are called the posterior probabilities. To flnd these probabilities, we write them in the form P(HijE)=P(Hi\E) P(E): (4.1) 146 CHAPTER 4. CONDITIONAL PROBABILITY Number having The results Disease this disease ++ +{ {+ {{ d1 3215 2110 301 704 100 d2 2125 396 132 1187 410 d3 4660 510 3568 73 509 Total 10000 Table 4.3: Diseases data. We can calculate the numerator from our given information by P(Hi\E)=P(Hi)P(EjHi): (4.2) Since one and only one of the events H1,H2, ...,Hmcan occur, we can write the probability of Eas P(E)=P(H1\E)+P(H2\E)+¢¢¢+P(Hm\E): Using Equation 4.2, the above expression can be seen to equal P(H1)P(EjH1)+P(H2)P(EjH2)+¢¢¢+P(Hm)P(EjHm): (4.3) Using (4.1), (4.2), and (4.3) yields Bayes’ formula : P(HijE)=P(Hi)P(EjHi)Pm k=1P(Hk)P(EjHk): Although this is a very famous formula, we will rarely use it. If the number of hypotheses is small, a simple tree measure calculation is easily carried out, as wehave done in our examples. If the number of hypotheses is large, then we shoulduse a computer. Bayes probabilities are particularly appropriate for medical diagnosis. A doctor is anxious to know which of several diseases a patient might have. She collectsevidence in the form of the outcomes of certain tests. From statistical studies thedoctor can flnd the prior probabilities of the various diseases before the tests, andthe probabilities for speciflc test outcomes, given a particular disease. What thedoctor wants to know is the posterior probability for the particular disease, giventhe outcomes of the tests. Example 4.16 A doctor is trying to decide if a patient has one of three diseases d 1,d2,o rd3. Two tests are to be carried out, each of which results in a positive (+) or a negative ( ¡) outcome. There are four possible test patterns ++, + ¡, ¡+, and¡¡. National records have indicated that, for 10,000 people having one of these three diseases, the distribution of diseases and test results are as in Table 4.3. From this data, we can estimate the prior probabilities for each of the diseases and, given a particular disease, the probability of a particular test outcome. Forexample, the prior probability of disease d 1may be estimated to be 3215 =10;000 = :3215. The probability of the test result + ¡, given disease d1, may be estimated to be 301=3125 =:094. 4.1. DISCRETE CONDITIONAL PROBABILITY 147 d1d2d3 + + .700 .132 .168 + { .076 .033 .891 { + .357 .605 .038{ { .098 .405 .497 Table 4.4: Posterior probabilities. We can now use Bayes’ formula to compute various posterior probabilities. The computer program Bayes computes these posterior probabilities. The results for this example are shown in Table 4.4. We note from the outcomes that, when the test result is ++, the disease d 1has a signiflcantly higher probability than the other two. When the outcome is + ¡, this is true for disease d3. When the outcome is ¡+, this is true for disease d2. Note that these statements might have been guessed by looking at the data. If theoutcome is¡¡, the most probable cause is d 3, but the probability that a patient hasd2is only slightly smaller. If one looks at the data in this case, one can see that it might be hard to guess which of the two diseases d2andd3is more likely. 2 Our flnal example shows that one has to be careful when the prior probabilities are small. Example 4.17 A doctor gives a patient a test for a particular cancer. Before the results of the test, the only evidence the doctor has to go on is that 1 womanin 1000 has this cancer. Experience has shown that, in 99 percent of the cases inwhich cancer is present, the test is positive; and in 95 percent of the cases in whichit is not present, it is negative. If the test turns out to be positive, what probabilityshould the doctor assign to the event that cancer is present? An alternative formof this question is to ask for the relative frequencies of false positives and cancers. We are given that prior(cancer) = :001 and prior(not cancer) = :999. We know also that P(+jcancer) = :99,P(¡jcancer) = :01,P(+jnot cancer) = :05, andP(¡jnot cancer) = :95. Using this data gives the result shown in Figure 4.5. We see now that the probability of cancer given a positive test has only increased from .001 to .019. While this is nearly a twenty-fold increase, the probability thatthe patient has the cancer is still small. Stated in another way, among the positiveresults, 98.1 percent are false positives, and 1.9 percent are cancers. When a groupof second-year medical students was asked this question, over half of the studentsincorrectly guessed the probability to be greater than .5. 2 Historical Remarks Conditional probability was used long before it was formally deflned. Pascal and Fermat considered the problem of points : given that team A has won mgames and team B has won ngames, what is the probability that A will win the series? (See Exercises 40{42.) This is clearly a conditional probability problem. In his book, Huygens gave a number of problems, one of which was: 148 CHAPTER 4. CONDITIONAL PROBABILITY .001 can not.01 .95.05+ -.001 0 .05 .949+ -.051 .949+ -.981 10can not.001 .05 0 .949can not.019Original Tree Reverse Tree .99 .999 Figure 4.5: Forward and reverse tree diagrams. Three gamblers, A, B and C, take 12 balls of which 4 are white and 8 black. They play with the rules that the drawer is blindfolded, A is todraw flrst, then B and then C, the winner to be the one who flrst drawsa white ball. What is the ratio of their chances? 2 From his answer it is clear that Huygens meant that each ball is replaced after drawing. However, John Hudde, the mayor of Amsterdam, assumed that he meantto sample without replacement and corresponded with Huygens about the difierencein their answers. Hacking remarks that \Neither party can understand what theother is doing." 3 By the time of de Moivre’s book, The Doctrine of Chances, these distinctions were well understood. De Moivre deflned independence and dependence as follows: Two Events are independent, when they have no connexion one with the other, and that the happening of one neither forwards nor obstructsthe happening of the other. Two Events are dependent, when they are so connected together as that the Probability of either’s happening is altered by the happening of theother. 4 De Moivre used sampling with and without replacement to illustrate that the probability that two independent events both happen is the product of their prob-abilities, and for dependent events that: 2Quoted in F. N. David, Games, Gods and Gambling (London: Gri–n, 1962), p. 119. 3I. Hacking, The Emergence of Probability (Cambridge: Cambridge University Press, 1975), p. 99. 4A. de Moivre, The Doctrine of Chances, 3rd ed. (New York: Chelsea, 1967), p. 6. 4.1. DISCRETE CONDITIONAL PROBABILITY 149 The Probability of the happening of two Events dependent, is the prod- uct of the Probability of the happening of one of them, by the Probabilitywhich the other will have of happening, when the flrst is considered ashaving happened; and the same Rule will extend to the happening of asmany Events as may be assigned. 5 The formula that we call Bayes’ formula, and the idea of computing the proba- bility of a hypothesis given evidence, originated in a famous essay of Thomas Bayes.Bayes was an ordained minister in Tunbridge Wells near London. His mathemat-ical interests led him to be elected to the Royal Society in 1742, but none of hisresults were published within his lifetime. The work upon which his fame rests,\An Essay Toward Solving a Problem in the Doctrine of Chances," was publishedin 1763, three years after his death. 6Bayes reviewed some of the basic concepts of probability and then considered a new kind of inverse probability problem requiringthe use of conditional probability. Bernoulli, in his study of processes that we now call Bernoulli trials, had proven his famous law of large numbers which we will study in Chapter 8. This theoremassured the experimenter that if he knew the probability pfor success, he could predict that the proportion of successes would approach this value as he increasedthe number of experiments. Bernoulli himself realized that in most interesting casesyou do not know the value of pand saw his theorem as an important step in showing that you could determine pby experimentation. To study this problem further, Bayes started by assuming that the probability p for success is itself determined by a random experiment. He assumed in fact that thisexperiment was such that this value for pis equally likely to be any value between 0 and 1. Without knowing this value we carry out nexperiments and observe m successes. Bayes proposed the problem of flnding the conditional probability thatthe unknown probability plies between aandb. He obtained the answer: P(a•p<bjmsuccesses in ntrials) =R b axm(1¡x)n¡mdx R1 0xm(1¡x)n¡mdx: We shall see in the next section how this result is obtained. Bayes clearly wanted to show that the conditional distribution function, given the outcomes of more andmore experiments, becomes concentrated around the true value of p. Thus, Bayes was trying to solve an inverse problem. The computation of the integrals was too di–cult for exact solution except for small values of jandn, and so Bayes tried approximate methods. His methods were not very satisfactory and it has beensuggested that this discouraged him from publishing his results. However, his paper was the flrst in a series of important studies carried out by Laplace, Gauss, and other great mathematicians to solve inverse problems. Theystudied this problem in terms of errors in measurements in astronomy. If an as-tronomer were to know the true value of a distance and the nature of the random 5ibid, p. 7. 6T. Bayes, \An Essay Toward Solving a Problem in the Doctrine of Chances," Phil. Trans. Royal Soc. London, vol. 53 (1763), pp. 370{418. 150 CHAPTER 4. CONDITIONAL PROBABILITY errors caused by his measuring device he could predict the probabilistic nature of his measurements. In fact, however, he is presented with the inverse problem ofknowing the nature of the random errors, and the values of the measurements, andwanting to make inferences about the unknown true value. As Maistrov remarks, the formula that we have called Bayes’ formula does not appear in his essay. Laplace gave it this name when he studied these inverse prob-lems. 7The computation of inverse probabilities is fundamental to statistics and has led to an important branch of statistics called Bayesian analysis, assuring Bayeseternal fame for his brief essay. Exercises 1Assume that EandFare two events with positive probabilities. Show that ifP(EjF)=P(E), thenP(FjE)=P(F). 2A coin is tossed three times. What is the probability that exactly two heads occur, given that (a) the flrst outcome was a head? (b) the flrst outcome was a tail? (c) the flrst two outcomes were heads? (d) the flrst two outcomes were tails? (e) the flrst outcome was a head and the third outcome was a head? 3A die is rolled twice. What is the probability that the sum of the faces is greater than 7, given that (a) the flrst outcome was a 4? (b) the flrst outcome was greater than 3? (c) the flrst outcome was a 1? (d) the flrst outcome was less than 5? 4A card is drawn at random from a deck of cards. What is the probability that (a) it is a heart, given that it is red? (b) it is higher than a 10, given that it is a heart? (Interpret J, Q, K, A as 11, 12, 13, 14.) (c) it is a jack, given that it is red? 5A coin is tossed three times. Consider the following events A: Heads on the flrst toss. B: Tails on the second. C: Heads on the third toss. D: All three outcomes the same (HHH or TTT). E: Exactly one head turns up. 7L. E. Maistrov, Probability Theory: A Historical Sketch, trans. and ed. Samual Kotz (New York: Academic Press, 1974), p. 100. 4.1. DISCRETE CONDITIONAL PROBABILITY 151 (a) Which of the following pairs of these events are independent? (1)A,B (2)A,D (3)A,E (4)D,E (b) Which of the following triples of these events are independent? (1)A,B,C (2)A,B,D (3)C,D,E 6From a deck of flve cards numbered 2, 4, 6, 8, and 10, respectively, a card is drawn at random and replaced. This is done three times. What is theprobability that the card numbered 2 was drawn exactly two times, giventhat the sum of the numbers on the three draws is 12? 7A coin is tossed twice. Consider the following events. A: Heads on the flrst toss. B: Heads on the second toss. C: The two tosses come out the same. (a) Show that A,B,Care pairwise independent but not independent. (b) Show that Cis independent of AandBbut not ofA\B. 8Let › =fa;b;c;d;e;fg. Assume that m(a)=m(b)=1=8 andm(c)= m(d)=m(e)=m(f)=3=16. LetA,B, andCbe the events A=fd;e;ag, B=fc;e;ag,C=fc;d;ag. Show that P(A\B\C)=P(A)P(B)P(C) but no two of these events are independent. 9What is the probability that a family of two children has (a) two boys given that it has at least one boy? (b) two boys given that the flrst child is a boy? 10In Example 4.2, we used the Life Table (see Appendix C) to compute a con- ditional probability. The number 93,753 in the table, corresponding to 40-year-old males, means that of all the males born in the United States in 1950,93.753% were alive in 1990. Is it reasonable to use this as an estimate for theprobability of a male, born this year, surviving to age 40? 11Simulate the Monty Hall problem. Carefully state any assumptions that you have made when writing the program. Which version of the problem do youthink that you are simulating? 12In Example 4.17, how large must the prior probability of cancer be to give a posterior probability of .5 for cancer given a positive test? 13Two cards are drawn from a bridge deck. What is the probability that the second card drawn is red? 152 CHAPTER 4. CONDITIONAL PROBABILITY 14IfP(~B)=1=4 andP(AjB)=1=2, what isP(A\B)? 15(a) What is the probability that your bridge partner has exactly two aces, given that she has at least one ace? (b) What is the probability that your bridge partner has exactly two aces, given that she has the ace of spades? 16Prove that for any three events A,B,C, each having positive probability, P(A\B\C)=P(A)P(BjA)P(CjA\B): 17Prove that if AandBare independent so are (a)Aand ~B. (b) ~Aand ~B. 18A doctor assumes that a patient has one of three diseases d1,d2,o rd3. Before any test, he assumes an equal probability for each disease. He carries out atest that will be positive with probability .8 if the patient has d 1,. 6i fh eh a s diseased2, and .4 if he has disease d3. Given that the outcome of the test was positive, what probabilities should the doctor now assign to the three possiblediseases? 19In a poker hand, John has a very strong hand and bets 5 dollars. The prob- ability that Mary has a better hand is .04. If Mary had a better hand shewould raise with probability .9, but with a poorer hand she would only raisewith probability .1. If Mary raises, what is the probability that she has abetter hand than John does? 20The Polya urn model for contagion is as follows: We start with an urn which contains one white ball and one black ball. At each second we choose a ballat random from the urn and replace this ball and add one more of the colorchosen. Write a program to simulate this model, and see if you can makeany predictions about the proportion of white balls in the urn after a largenumber of draws. Is there a tendency to have a large fraction of balls of thesame color in the long run? 21It is desired to flnd the probability that in a bridge deal each player receives an ace. A student argues as follows. It does not matter where the flrst ace goes.The second ace must go to one of the other three players and this occurs withprobability 3/4. Then the next must go to one of two, an event of probability1/2, and flnally the last ace must go to the player who does not have an ace.This occurs with probability 1/4. The probability that all these events occuris the product (3 =4)(1=2)(1=4 )=3=32. Is this argument correct? 22One coin in a collection of 65 has two heads. The rest are fair. If a coin, chosen at random from the lot and then tossed, turns up heads 6 times in arow, what is the probability that it is the two-headed coin? 4.1. DISCRETE CONDITIONAL PROBABILITY 153 23You are given two urns and flfty balls. Half of the balls are white and half are black. You are asked to distribute the balls in the urns with no restrictionplaced on the number of either type in an urn. How should you distributethe balls in the urns to maximize the probability of obtaining a white ball ifan urn is chosen at random and a ball drawn out at random? Justify youranswer. 24A fair coin is thrown ntimes. Show that the conditional probability of a head on any specifled trial, given a total of kheads over the ntrials, isk=n(k>0). 25(Johnsonbough 8) A coin with probability pfor heads is tossed ntimes. LetE be the event \a head is obtained on the flrst toss’ and Fkthe event ‘exactly k heads are obtained." For which pairs ( n;k) areEandFkindependent? 26Suppose that AandBare events such that P(AjB)=P(BjA) andP(A[B)= 1 andP(A\B)>0. Prove that P(A)>1=2. 27(Chung9) In London, half of the days have some rain. The weather forecaster is correct 2/3 of the time, i.e., the probability that it rains, given that she haspredicted rain, and the probability that it does not rain, given that she haspredicted that it won’t rain, are both equal to 2/3. When rain is forecast,Mr. Pickwick takes his umbrella. When rain is not forecast, he takes it withprobability 1/3. Find (a) the probability that Pickwick has no umbrella, given that it rains. (b) the probability that it doesn’t rain, given that he brings his umbrella. 28Probability theory was used in a famous court case: People v. Collins. 10In this case a purse was snatched from an elderly person in a Los Angeles suburb.A couple seen running from the scene were described as a black man with abeard and a mustache and a blond girl with hair in a ponytail. Witnesses saidthey drove ofi in a partly yellow car. Malcolm and Janet Collins were arrested.He was black and though clean shaven when arrested had evidence of recentlyhaving had a beard and a mustache. She was blond and usually wore her hairin a ponytail. They drove a partly yellow Lincoln. The prosecution called aprofessor of mathematics as a witness who suggested that a conservative set ofprobabilities for the characteristics noted by the witnesses would be as shownin Table 4.5. The prosecution then argued that the probability that all of these character- istics are met by a randomly chosen couple is the product of the probabilitiesor 1/12,000,000, which is very small. He claimed this was proof beyond a rea-sonable doubt that the defendants were guilty. The jury agreed and handeddown a verdict of guilty of second-degree robbery. 8R. Johnsonbough, \Problem #103," Two Year College Math Journal, vol. 8 (1977), p. 292. 9K. L. Chung, Elementary Probability Theory With Stochastic P rocesses, 3rd ed. (New York: Springer-Verlag, 1979), p. 152. 10M. W. Gray, \Statistics and the Law," Mathematics Magazine, vol. 56 (1983), pp. 67{81. 154 CHAPTER 4. CONDITIONAL PROBABILITY man with mustache 1/4 girl with blond hair 1/3girl with ponytail 1/10black man with beard 1/10interracial couple in a car 1/1000partly yellow car 1/10 Table 4.5: Collins case probabilities. If you were the lawyer for the Collins couple how would you have countered the above argument? (The appeal of this case is discussed in Exercise 5.1.34.) 29A student is applying to Harvard and Dartmouth. He estimates that he has a probability of .5 of being accepted at Dartmouth and .3 of being acceptedat Harvard. He further estimates the probability that he will be accepted byboth is .2. What is the probability that he is accepted by Dartmouth if he isaccepted by Harvard? Is the event \accepted at Harvard" independent of theevent \accepted at Dartmouth"? 30Luxco, a wholesale lightbulb manufacturer, has two factories. Factory A sells bulbs in lots that consists of 1000 regular and 2000 softglow bulbs each. Ran- dom sampling has shown that on the average there tend to be about 2 badregular bulbs and 11 bad softglow bulbs per lot. At factory B the lot size isreversed|there are 2000 regular and 1000 softglow per lot|and there tendto be 5 bad regular and 6 bad softglow bulbs per lot. The manager of factory A asserts, \We’re obviously the better producer; our bad bulb rates are .2 percent and .55 percent compared to B’s .25 percent and.6 percent. We’re better at both regular and softglow bulbs by half of a tenthof a percent each." \Au contraire," counters the manager of B, \each of our 3000 bulb lots con- tains only 11 bad bulbs, while A’s 3000 bulb lots contain 13. So our .37percent bad bulb rate beats their .43 percent." Who is right? 31Using the Life Table for 1981 given in Appendix C, flnd the probability that a male of age 60 in 1981 lives to age 80. Find the same probability for a female. 32(a) There has been a blizzard and Helen is trying to drive from Woodstock to Tunbridge, which are connected like the top graph in Figure 4.6. Herepandqare the probabilities that the two roads are passable. What is the probability that Helen can get from Woodstock to Tunbridge? (b) Now suppose that Woodstock and Tunbridge are connected like the mid- dle graph in Figure 4.6. What now is the probability that she can getfromWtoT? Note that if we think of the roads as being components of a system, then in (a) and (b) we have computed the reliability of a system whose components are (a) in series and (b) in parallel. 4.1. DISCRETE CONDITIONAL PROBABILITY 155 Woodstock Tunbridgep q C DT W .8 .9.9 .8 .95WTp q(a) (b) (c) Figure 4.6: From Woodstock to Tunbridge. (c) Now suppose WandTare connected like the bottom graph in Figure 4.6. Find the probability of Helen’s getting from WtoT.Hint: If the road fromCtoDis impassable, it might as well not be there at all; if it is passable, then flgure out how to use part (b) twice. 33LetA1,A2, andA3be events, and let Birepresent either Aior its complement ~Ai. Then there are eight possible choices for the triple ( B1;B2;B3). Prove that the events A1,A2,A3are independent if and only if P(B1\B2\B3)=P(B1)P(B2)P(B3); for all eight of the possible choices for the triple ( B1;B2;B3). 34Four women, A, B, C, and D, check their hats, and the hats are returned in a random manner. Let › be the set of all possible permutations of A, B, C, D.LetX j= 1 if thejth woman gets her own hat back and 0 otherwise. What is the distribution of Xj? Are theXi’s mutually independent? 35A box has numbers from 1 to 10. A number is drawn at random. Let X1be the number drawn. This number is replaced, and the ten numbers mixed. Asecond number X 2is drawn. Find the distributions of X1andX2. AreX1 andX2independent? Answer the same questions if the flrst number is not replaced before the second is drawn. 156 CHAPTER 4. CONDITIONAL PROBABILITY Y - 1012 X -1 0 1/36 1/6 1/12 01/18 0 1/18 0 1 0 1/36 1/6 1/12 21/12 0 1/12 1/6 Table 4.6: Joint distribution. 36A die is thrown twice. Let X1andX2denote the outcomes. Deflne X= min(X1;X2). Find the distribution of X. *37 Given that P(X=a)=r,P(max(X;Y )=a)=s, andP(min(X;Y )=a)= t, show that you can determine u=P(Y=a) in terms of r,s, andt. 38A fair coin is tossed three times. Let Xbe the number of heads that turn up on the flrst two tosses and Ythe number of heads that turn up on the third toss. Give the distribution of (a) the random variables XandY. (b) the random variable Z=X+Y. (c) the random variable W=X¡Y. 39Assume that the random variables XandYhave the joint distribution given in Table 4.6. (a) What is P(X‚1 andY•0)? (b) What is the conditional probability that Y•0 given that X=2 ? (c) AreXandYindependent? (d) What is the distribution of Z=XY? 40In the problem of points , discussed in the historical remarks in Section 3.2, two players, A and B, play a series of points in a game with player A winning eachpoint with probability pand player B winning each point with probability q=1¡p. The flrst player to win Npoints wins the game. Assume that N= 3. LetXbe a random variable that has the value 1 if player A wins the series and 0 otherwise. Let Ybe a random variable with value the number of points played in a game. Find the distribution of XandYwhenp=1=2. AreXandYindependent in this case? Answer the same questions for the casep=2=3. 41The letters between Pascal and Fermat, which are often credited with having started probability theory, dealt mostly with the problem of points described in Exercise 40. Pascal and Fermat considered the problem of flnding a fairdivision of stakes if the game must be called ofi when the flrst player has wonrgames and the second player has won sgames, with r<N ands<N . Let P(r;s) be the probability that player A wins the game if he has already won rpoints and player B has won spoints. Then 4.1. DISCRETE CONDITIONAL PROBABILITY 157 (a)P(r;N)=0i fr<N , (b)P(N;s)=1i fs<N , (c)P(r;s)=pP(r+1;s)+qP(r;s+1 )i fr<N ands<N ; and (1), (2), and (3) determine P(r;s) forr•Nands•N. Pascal used these facts to flnd P(r;s) by working backward: He flrst obtained P(N¡1;j) forj=N¡1,N¡2 , ..., 0 ; then, from these values, he obtained P(N¡2;j) forj=N¡1,N¡2, . . . , 0 and, continuing backward, obtained all the valuesP(r;s). Write a program to compute P(r;s) for given N,a,b, andp. Warning : Follow Pascal and you will be able to run N= 100; use recursion and you will not be able to run N= 20. 42Fermat solved the problem of points (see Exercise 40) as follows: He realized that the problem was di–cult because the possible ways the play might go arenot equally likely. For example, when the flrst player needs two more gamesand the second needs three to win, two possible ways the series might go forthe flrst player are WLW and LWLW. These sequences are not equally likely.To avoid this di–culty, Fermat extended the play, adding flctitious plays sothat the series went the maximum number of games needed (four in this case).He obtained equally likely outcomes and used, in efiect, the Pascal triangle tocalculateP(r;s). Show that this leads to a formula forP(r;s) even for the casep6=1=2. 43The Yankees are playing the Dodgers in a world series. The Yankees win each game with probability .6. What is the probability that the Yankees win theseries? (The series is won by the flrst team to win four games.) 44C. L. Anderson 11has used Fermat’s argument for the problem of points to prove the following result due to J. G. Kingston. You are playing the game of points (see Exercise 40) but, at each point, when you serve you win with probability p, and when your opponent serves you win with probability „ p. You will serve flrst, but you can choose one of the following two conventionsfor serving: for the flrst convention you alternate service (tennis), and for thesecond the person serving continues to serve until he loses a point and thenthe other player serves (racquetball). The flrst player to win Npoints wins the game. The problem is to show that the probability of winning the gameis the same under either convention. (a) Show that, under either convention, you will serve at most Npoints and your opponent at most N¡1 points. (b) Extend the number of points to 2 N¡1 so that you serve Npoints and your opponent serves N¡1. For example, you serve any additional points necessary to make Nserves and then your opponent serves any additional points necessary to make him serve N¡1 points. The winner 11C. L. Anderson, \Note on the Advantage of First Serve," Journal of Combinatorial Theory, Series A, vol. 23 (1977), p. 363. 158 CHAPTER 4. CONDITIONAL PROBABILITY is now the person, in the extended game, who wins the most points. Show that playing these additional points has not changed the winner. (c) Show that (a) and (b) prove that you have the same probability of win- ning the game under either convention. 45In the previous problem, assume that p=1¡„p. (a) Show that under either service convention, the flrst player will win more often than the second player if and only if p>: 5. (b) In volleyball, a team can only win a point while it is serving. Thus, any individual \play" either ends with a point being awarded to the servingteam or with the service changing to the other team. The flrst team towinNpoints wins the game. (We ignore here the additional restriction that the winning team must be ahead by at least two points at the end ofthe game.) Assume that each team has the same probability of winningthe play when it is serving, i.e., that p=1¡„p. Show that in this case, the team that serves flrst will win more than half the time, as long asp>0. (Ifp= 0, then the game never ends.) Hint: Deflnep 0to be the probability that a team wins the next point, given that it is serving. Ifwe writeq=1¡p, then one can show that p 0=p 1¡q2: If one now considers this game in a slightly difierent way, one can see that the second service convention in the preceding problem can be used,withpreplaced by p 0. 46A poker hand consists of 5 cards dealt from a deck of 52 cards. Let Xand Ybe, respectively, the number of aces and kings in a poker hand. Find the joint distribution of XandY. 47LetX1andX2be independent random variables and let Y1=`1(X1) and Y2=`2(X2). (a) Show that P(Y1=r;Y2=s)=X `1(a)=r `2(b)=sP(X1=a;X 2=b): (b) Using (a), show that P(Y1=r;Y2=s)=P(Y1=r)P(Y2=s) so that Y1andY2are independent. 48Let › be the sample space of an experiment. Let Ebe an event with P(E)>0 and deflne mE(!)b ymE(!)=m(!jE). Prove that mE(!) is a distribution function on E, that is, that mE(!)‚0 and thatP !2›mE(!)=1 . T h e functionmEis called the conditional distribution given E. 4.1. DISCRETE CONDITIONAL PROBABILITY 159 49You are given two urns each containing two biased coins. The coins in urn I come up heads with probability p1, and the coins in urn II come up heads with probability p26=p1. You are given a choice of (a) choosing an urn at random and tossing the two coins in this urn or (b) choosing one coin fromeach urn and tossing these two coins. You win a prize if both coins turn upheads. Show that you are better ofi selecting choice (a). 50Prove that, if A 1,A2, ...,Anare independent events deflned on a sample space › and if 0 <P(Aj)<1 for allj, then › must have at least 2npoints. 51Prove that if P(AjC)‚P(BjC) andP(Aj~C)‚P(Bj~C); thenP(A)‚P(B). 52A coin is in one of nboxes. The probability that it is in the ith box ispi. If you search in the ith box and it is there, you flnd it with probability ai. Show that the probability pthat the coin is in the jth box, given that you have looked in the ith box and not found it, is p=‰pj=(1¡aipi); ifj6=i; (1¡ai)pi=(1¡aipi);ifj=i: 53George Wolford has suggested the following variation on the Linda problem (see Exercise 1.2.25). The registrar is carrying John and Mary’s registrationcards and drops them in a puddle. When he pickes them up he cannot read thenames but on the flrst card he picked up he can make out Mathematics 23 andGovernment 35, and on the second card he can make out only Mathematics23. He asks you if you can help him decide which card belongs to Mary. Youknow that Mary likes government but does not like mathematics. You knownothing about John and assume that he is just a typical Dartmouth student.From this you estimate: P(Mary takes Government 35) = :5; P(Mary takes Mathematics 23) = :1; P(John takes Government 35) = :3; P(John takes Mathematics 23) = :2: Assume that their choices for courses are independent events. Show that the card with Mathematics 23 and Government 35 showing is more likelyto be Mary’s than John’s. The conjunction fallacy referred to in the Lindaproblem would be to assume that the event \Mary takes Mathematics 23 andGovernment 35" is more likely than the event \Mary takes Mathematics 23."Why are we not making this fallacy here? 160 CHAPTER 4. CONDITIONAL PROBABILITY 54(Suggested by Eisenberg and Ghosh12) A deck of playing cards can be de- scribed as a Cartesian product Deck = Suit£Rank; where Suit =f|;};~;˜gand Rank =f2;3;:::; 10;J;Q;K;Ag. This just means that every card may be thought of as an ordered pair like ( };2). By asuit event we mean any event Acontained in Deck which is described in terms of Suit alone. For instance, if Ais \the suit is red," then A=f};~g£ Rank; so thatAconsists of all cards of the form ( };r)o r(~;r) whereris any rank. Similarly, a rank event is any event described in terms of rank alone. (a) Show that if Ais any suit event and Bany rank event, then AandBare independent. (We can express this brie°y by saying that suit and rank are independent.) (b) Throw away the ace of spades. Show that now no nontrivial (i.e., neither empty nor the whole space) suit event Ais independent of any nontrivial rank event B.Hint: Here independence comes down to c=5 1=(a=51)¢(b=51); wherea,b,care the respective sizes of A,BandA\B. It follows that 51 must divide ab, hence that 3 must divide one of aandb, and 17 the other. But the possible sizes for suit and rank events preclude this. (c) Show that the deck in (b) nevertheless does have pairs A,Bof nontrivial independent events. Hint: Find 2 events AandBof sizes 3 and 17, respectively, which intersect in a single point. (d) Add a joker to a full deck. Show that now there is no pair A,Bof nontrivial independent events. Hint: See the hint in (b); 53 is prime. The following problems are suggested by Stanley Gudder in his article \Do Good Hands Attract?"13He says that event Aattracts eventBifP(BjA)> P(B) and repelsBifP(BjA)<P(B). 55LetRibe the event that the ith player in a poker game has a royal °ush. Show that a royal °ush (A,K,Q,J,10 of one suit) attracts another royal °ush,that isP(R 2jR1)>P(R2). Show that a royal °ush repels full houses. 56Prove that AattractsBif and only if BattractsA. Hence we can say that AandBaremutually attractive ifAattractsB. 12B. Eisenberg and B. K. Ghosh, \Independent Events in a Discrete Uniform Probability Space," The American Statistician, vol. 41, no. 1 (1987), pp. 52{56. 13S. Gudder, \Do Good Hands Attract?" Mathematics Magazine, vol. 54, no. 1 (1981), pp. 13{ 16. 4.1. DISCRETE CONDITIONAL PROBABILITY 161 57Prove that Aneither attracts nor repels Bif and only if AandBare inde- pendent. 58Prove that AandBare mutually attractive if and only if P(BjA)>P(Bj~A). 59Prove that if AattractsB, thenArepels ~B. 60Prove that if Aattracts both BandC, andArepelsB\C, thenAattracts B[C. Is there any example in which Aattracts both BandCand repels B[C? 61Prove that if B1,B2,...,Bnare mutually disjoint and collectively exhaustive, and ifAattracts some Bi, thenAmust repel some Bj. 62(a) Suppose that you are looking in your desk for a letter from some time ago. Your desk has eight drawers, and you assess the probability that itis in any particular drawer is 10% (so there is a 20% chance that it is notin the desk at all). Suppose now that you start searching systematicallythrough your desk, one drawer at a time. In addition, suppose thatyou have not found the letter in the flrst idrawers, where 0 •i•7. Letp idenote the probability that the letter will be found in the next drawer, and let qidenote the probability that the letter will be found in some subsequent drawer (both piandqiare conditional probabilities, since they are based upon the assumption that the letter is not in theflrstidrawers). Show that the p i’s increase and the qi’s decrease. (This problem is from Falk et al.14) (b) The following data appeared in an article in the Wall Street Journal.15 For the ages 20, 30, 40, 50, and 60, the probability of a woman in the U.S. developing cancer in the next ten years is 0.5%, 1.2%, 3.2%, 6.4%,and 10.8%, respectively. At the same set of ages, the probability of awoman in the U.S. eventually developing cancer is 39.6%, 39.5%, 39.1%,37.5%, and 34.2%, respectively. Do you think that the problem in part(a) gives an explanation for these data? 63Here are two variations of the Monty Hall problem that are discussed by Granberg. 16 (a) Suppose that everything is the same except that Monty forgot to flnd out in advance which door has the car behind it. In the spirit of \theshow must go on," he makes a guess at which of the two doors to openand gets lucky, opening a door behind which stands a goat. Now shouldthe contestant switch? 14R. Falk, A. Lipson, and C. Konold, \The ups and downs of the hope function in a fruitless search," in Subjective Probability, G. Wright and P. Ayton, (eds.) (Chichester: Wiley, 1994), pgs. 353-377. 15C. Crossen, \Fright by the numbers: Alarming disease data are frequently °awed," Wall Street Journal, 11 April 1996, p. B1. 16D. Granberg, \To switch or not to switch," in The power of logical thinking, M. vos Savant, (New York: St. Martin’s 1996). 162 CHAPTER 4. CONDITIONAL PROBABILITY (b) You have observed the show for a long time and found that the car is put behind door A 45% of the time, behind door B 40% of the time andbehind door C 15% of the time. Assume that everything else about theshow is the same. Again you pick door A. Monty opens a door with agoat and ofiers to let you switch. Should you? Suppose you knew inadvance that Monty was going to give you a chance to switch. Shouldyou have initially chosen door A? 4.2 Continuous Conditional Probability In situations where the sample space is continuous we will follow the same procedureas in the previous section. Thus, for example, if Xis a continuous random variable with density function f(x), and ifEis an event with positive probability, we deflne a conditional density function by the formula f(xjE)=‰f(x)=P(E);ifx2E; 0; ifx62E: Then for any event F,w eh a v e P(FjE)=Z Ff(xjE)dx : The expression P(FjE) is called the conditional probability of FgivenE. A si nt h e previous section, it is easy to obtain an alternative expression for this probability: P(FjE)=Z Ff(xjE)dx=Z E\Ff(x) P(E)dx=P(E\F) P(E): We can think of the conditional density function as being 0 except on E, and normalized to have integral 1 over E. Note that if the original density is a uniform density corresponding to an experiment in which all events of equal size are equally likely, then the same will be true for the conditional density. Example 4.18 In the spinner experiment (cf. Example 2.1), suppose we know that the spinner has stopped with head in the upper half of the circle, 0 •x•1=2. What is the probability that 1 =6•x•1=3? HereE=[ 0;1=2],F=[ 1=6;1=3], andF\E=F. Hence P(FjE)=P(F\E) P(E) =1=6 1=2 =1 3; which is reasonable, since Fis 1/3 the size of E. The conditional density function here is given by 4.2. CONTINUOUS CONDITIONAL PROBABILITY 163 f(xjE)=‰2;if 0•x<1=2; 0;if 1=2•x<1: Thus the conditional density function is nonzero only on [0 ;1=2], and is uniform there. 2 Example 4.19 In the dart game (cf. Example 2.8), suppose we know that the dart lands in the upper half of the target. What is the probability that its distance fromthe center is less than 1/2? HereE=f(x;y):y‚0g, andF=f(x;y):x 2+y2<(1=2)2g. Hence, P(FjE)=P(F\E) P(E)=(1=…)[(1=2)(…=4)] (1=…)(…=2) =1=4: Here again, the size of F\Eis 1/4 the size of E. The conditional density function is f((x;y)jE)=‰f(x;y)=P(E)=2=…; if (x;y)2E; 0; if (x;y)62E: 2 Example 4.20 We return to the exponential density (cf. Example 2.17). We sup- pose that we are observing a lump of plutonium-239. Our experiment consists ofwaiting for an emission, then starting a clock, and recording the length of time X that passes until the next emission. Experience has shown that Xhas an expo- nential density with some parameter ‚, which depends upon the size of the lump. Suppose that when we perform this experiment, we notice that the clock reads r seconds, and is still running. What is the probability that there is no emission in afurthersseconds? LetG(t) be the probability that the next particle is emitted after time t. Then G(t)=Z 1 t‚e¡‚xdx =¡e¡‚xflfl1 t=e¡‚t: LetEbe the event \the next particle is emitted after time r" andFthe event \the next particle is emitted after time r+s." Then P(FjE)=P(F\E) P(E) =G(r+s) G(r) =e¡‚(r+s) e¡‚r =e¡‚s: 164 CHAPTER 4. CONDITIONAL PROBABILITY This tells us the rather surprising fact that the probability that we have to wait sseconds more for an emission, given that there has been no emission in rseconds, isindependent of the time r. This property (called the memoryless property) was introduced in Example 2.17. When trying to model various phenomena, thisproperty is helpful in deciding whether the exponential density is appropriate. The fact that the exponential density is memoryless means that it is reasonable to assume if one comes upon a lump of a radioactive isotope at some random time,then the amount of time until the next emission has an exponential density withthe same parameter as the time between emissions. A well-known example, knownas the \bus paradox," replaces the emissions by buses. The apparent paradox arisesfrom the following two facts: 1) If you know that, on the average, the buses comeby every 30 minutes, then if you come to the bus stop at a random time, you shouldonly have to wait, on the average, for 15 minutes for a bus, and 2) Since the busesarrival times are being modelled by the exponential density, then no matter whenyou arrive, you will have to wait, on the average, for 30 minutes for a bus. The reader can now see that in Exercises 2.2.9, 2.2.10, and 2.2.11, we were asking for simulations of conditional probabilities, under various assumptions onthe distribution of the interarrival times. If one makes a reasonable assumptionabout this distribution, such as the one in Exercise 2.2.10, then the average waitingtime is more nearly one-half the average interarrival time. 2 Independent Events IfEandFare two events with positive probability in a continuous sample space, then, as in the case of discrete sample spaces, we deflne EandFto be independent ifP(EjF)=P(E) andP(FjE)=P(F). As before, each of the above equations imply the other, so that to see whether two events are independent, only one of theseequations must be checked. It is also the case that, if EandFare independent, thenP(E\F)=P(E)P(F). Example 4.21 (Example 4.18 continued) In the dart game (see Example 4.18), let Ebe the event that the dart lands in the upper half of the target ( y‚0) andFthe event that the dart lands in the right half of the target ( x‚0). ThenP(E\F)i s the probability that the dart lies in the flrst quadrant of the target, and P(E\F)=1 …Z E\F1dxdy = Area (E\F) = Area (E) Area (F) =µ1 …Z E1dxdy¶µ1 …Z F1dxdy¶ =P(E)P(F) so thatEandFare independent. What makes this work is that the events Eand Fare described by restricting difierent coordinates. This idea is made more precise below. 2 4.2. CONTINUOUS CONDITIONAL PROBABILITY 165 Joint Density and Cumulative Distribution Functions In a manner analogous with discrete random variables, we can deflne joint density functions and cumulative distribution functions for multi-dimensional continuousrandom variables. Deflnition 4.6 LetX 1;X 2;:::; Xnbe continuous random variables associated with an experiment, and let „X=(X1;X2;:::; Xn). Then the joint cumulative distribution function of „Xis deflned by F(x1;x2;:::;xn)=P(X1•x1;X2•x2;:::;Xn•xn): The joint density function of „Xsatisfles the following equation: F(x1;x2;:::;xn)=Zx1 ¡1Zx2 ¡1¢¢¢Zxn ¡1f(t1;t2;:::tn)dtndtn¡1:::dt 1: 2 It is straightforward to show that, in the above notation, f(x1;x2;:::;xn)=@nF(x1;x2;:::;xn) @x1@x2¢¢¢@xn: (4.4) Independent Random Variables As with discrete random variables, we can deflne mutual independence of continuous random variables. Deflnition 4.7 LetX1,X2,...,Xnbe continuous random variables with cumula- tive distribution functions F1(x);F2(x);:::; Fn(x). Then these random variables aremutually independent if F(x1;x2;:::;xn)=F1(x1)F2(x2)¢¢¢Fn(xn) for any choice of x1;x2;:::;xn. Thus, if X1;X 2;:::; Xnare mutually inde- pendent, then the joint cumulative distribution function of the random variable „X=(X1;X2;:::;Xn) is just the product of the individual cumulative distribution functions. When two random variables are mutually independent, we shall say morebrie°y that they are independent. 2 Using Equation 4.4, the following theorem can easily be shown to hold for mu- tually independent continuous random variables. Theorem 4.2 LetX 1,X2, ...,Xnbe continuous random variables with density functionsf1(x);f2(x);:::; fn(x). Then these random variables are mutually in- dependent if and only if f(x1;x2;:::;xn)=f1(x1)f2(x2)¢¢¢fn(xn) for any choice of x1;x2;:::;xn. 2 166 CHAPTER 4. CONDITIONAL PROBABILITY 11 r r 0ω ωE2 1 12 1 Figure 4.7: X1andX2are independent. Let’s look at some examples. Example 4.22 In this example, we deflne three random variables, X1;X2, and X3. We will show that X1andX2are independent, and that X1andX3are not independent. Choose a point !=(!1;!2) at random from the unit square. Set X1=!2 1,X2=!2 2, andX3=!1+!2. Find the joint distributions F12(r1;r2) and F23(r2;r3). We have already seen (see Example 2.13) that F1(r1)=P(¡1<X 1•r1) =pr1; if 0•r1•1; and similarly, F2(r2)=pr2; if 0•r2•1. Now we have (see Figure 4.7) F12(r1;r2)=P(X1•r1andX2•r2) =P(!1•pr1and!2•pr2) = Area (E1) =pr1pr2 =F1(r1)F2(r2): In this case F12(r1;r2)=F1(r1)F2(r2) so thatX1andX2are independent. On the other hand, if r1=1=4 andr3= 1, then (see Figure 4.8) F13(1=4;1) =P(X1•1=4;X3•1) 4.2. CONTINUOUS CONDITIONAL PROBABILITY 167 11 0ω ωω + ω = 1 1/212 12 Ε2 Figure 4.8: X1andX3are not independent. =P(!1•1=2;!1+!2•1) = Area (E2) =1 2¡1 8=3 8: Now recalling that F3(r3)=8 >>< >>:0; ifr 3<0; (1=2)r2 3; if 0•r3•1; 1¡(1=2)(2¡r3)2;if 1•r3•2; 1; if 2<r3; (see Example 2.14), we have F1(1=4)F3( 1 )=( 1=2)(1=2 )=1=4. Hence,X1andX3 are not independent random variables. A similar calculation shows that X2andX3 are not independent either. 2 Although we shall not prove it here, the following theorem is a useful one. The statement also holds for mutually independent discrete random variables. A proofmay be found in R¶ enyi. 17 Theorem 4.3 LetX1;X2;:::;Xnbe mutually independent continuous random variables and let `1(x);`2(x);:::;`n(x) be continuous functions. Then `1(X1); `2(X2);:::;`n(Xn) are mutually independent. 2 Independent Trials Using the notion of independence, we can now formulate for continuous sample spaces the notion of independent trials (see Deflnition 4.5). 17A. R¶ enyi, Probability Theory (Budapest: Akad¶ emiai Kiad¶ o, 1970), p. 183. 168 CHAPTER 4. CONDITIONAL PROBABILITY 0.2 0.4 0.6 0.8 10.511.522.53 α = β =.5 α = β =1 α = β = 2 0 Figure 4.9: Beta density for fi=fl=:5;1;2: Deflnition 4.8 A sequence X1,X2, ...,Xnof random variables Xithat are mutually independent and have the same density is called an independent trials process. 2 As in the case of discrete random variables, these independent trials processes arise naturally in situations where an experiment described by a single randomvariable is repeated ntimes. Beta Density We consider next an example which involves a sample space with both discrete and continuous coordinates. For this example we shall need a new density functioncalled the beta density. This density has two parameters fi,fland is deflned by B(fi;fl;x )=‰(1=B(fi;fl))x fi¡1(1¡x)fl¡1;if 0•x•1; 0; otherwise: Herefiandflare any positive numbers, and the beta function B(fi;fl) is given by the area under the graph of xfi¡1(1¡x)fl¡1between 0 and 1: B(fi;fl)=Z1 0xfi¡1(1¡x)fl¡1dx : Note that when fi=fl= 1 the beta density if the uniform density. When fiand flare greater than 1 the density is bell-shaped, but when they are less than 1 it is U-shaped as suggested by the examples in Figure 4.9. We shall need the values of the beta function only for integer values of fiandfl, and in this case B(fi;fl)=(fi¡1)! (fl¡1)! (fi+fl¡1)!: Example 4.23 In medical problems it is often assumed that a drug is efiective with a probability xeach time it is used and the various trials are independent, so that 4.2. CONTINUOUS CONDITIONAL PROBABILITY 169 one is, in efiect, tossing a biased coin with probability xfor heads. Before further experimentation, you do not know the value xbut past experience might give some information about its possible values. It is natural to represent this informationby sketching a density function to determine a distribution for x. Thus, we are considering xto be a continuous random variable, which takes on values between 0 and 1. If you have no knowledge at all, you would sketch the uniform density.If past experience suggests that xis very likely to be near 2/3 you would sketch a density with maximum at 2/3 and a spread re°ecting your uncertainly in theestimate of 2/3. You would then want to flnd a density function that reasonablyflts your sketch. The beta densities provide a class of densities that can be flt tomost sketches you might make. For example, for fi>1 andfl>1 it is bell-shaped with the parameters fiandfldetermining its peak and its spread. Assume that the experimenter has chosen a beta density to describe the state of his knowledge about xbefore the experiment. Then he gives the drug to nsubjects and records the number iof successes. The number iis a discrete random variable, so we may conveniently describe the set of possible outcomes of this experiment byreferring to the ordered pair ( x;i). We letm(ijx) denote the probability that we observe isuccesses given the value ofx. By our assumptions, m(ijx) is the binomial distribution with probability x for success: m(ijx)=b(n;x;i )=µn i¶ x i(1¡x)j; wherej=n¡i. Ifxis chosen at random from [0 ;1] with a beta density B(fi;fl;x ), then the density function for the outcome of the pair ( x;i)i s f(x;i)=m(ijx)B(fi;fl;x ) =µn i¶ xi(1¡x)j1 B(fi;fl)xfi¡1(1¡x)fl¡1 =µn i¶1 B(fi;fl)xfi+i¡1(1¡x)fl+j¡1: Now letm(i) be the probability that we observe isuccesses notknowing the value ofx. Then m(i)=Z1 0m(ijx)B(fi;fl;x )dx =µn i¶1 B(fi;fl)Z1 0xfi+i¡1(1¡x)fl+j¡1dx =µn i¶B(fi+i;fl+j) B(fi;fl): Hence, the probability density f(xji) forx, given that isuccesses were observed, is f(xji)=f(x;i) m(i) 170 CHAPTER 4. CONDITIONAL PROBABILITY =xfi+i¡1(1¡x)fl+j¡1 B(fi+i;fl+j); (4.5) that is,f(xji) is another beta density. This says that if we observe isuccesses and jfailures innsubjects, then the new density for the probability that the drug is efiective is again a beta density but with parameters fi+i,fl+j. Now we assume that before the experiment we choose a beta density with pa- rametersfiandfl, and that in the experiment we obtain isuccesses in ntrials. We have just seen that in this case, the new density for xis a beta density with parameters fi+iandfl+j. Now we wish to calculate the probability that the drug is efiective on the next subject. For any particular real number tbetween 0 and 1, the probability that x has the value tis given by the expression in Equation 4.5. Given that xhas the valuet, the probability that the drug is efiective on the next subject is just t. Thus, to obtain the probability that the drug is efiective on the next subject, we integratethe product of the expression in Equation 4.5 and tover all possible values of t.W e obtain: 1 B(fi+i;fl+j)Z1 0t¢tfi+i¡1(1¡t)fl+j¡1dt =B(fi+i+1;fl+j) B(fi+i;fl+j) =(fi+i)! (fl+j¡1)! (fi+fl+i+j)!¢(fi+fl+i+j¡1)! (fi+i¡1)! (fl+j¡1)! =fi+i fi+fl+n: Ifnis large, then our estimate for the probability of success after the experiment is approximately the proportion of successes observed in the experiment, which iscertainly a reasonable conclusion. 2 The next example is another in which the true probabilities are unknown and must be estimated based upon experimental data. Example 4.24 (Two-armed bandit problem) You are in a casino and confronted by two slot machines. Each machine pays ofi either 1 dollar or nothing. The probabilitythat the flrst machine pays ofi a dollar is xand that the second machine pays ofi a dollar isy. We assume that xandyare random numbers chosen independently from the interval [0 ;1] and unknown to you. You are permitted to make a series of ten plays, each time choosing one machine or the other. How should you choose tomaximize the number of times that you win? One strategy that sounds reasonable is to calculate, at every stage, the prob- ability that each machine will pay ofi and choose the machine with the higherprobability. Let win( i), fori= 1 or 2, be the number of times that you have won on theith machine. Similarly, let lose( i) be the number of times you have lost on theith machine. Then, from Example 4.23, the probability p(i) that you win if you 4.2. CONTINUOUS CONDITIONAL PROBABILITY 171 0.2 0.4 0.6 0.8 10.511.522.5 Machine Result 1 W 1 L 2 L 1 L 1 W 1 L 1 L 1 L 2 W 2 L 0 0 Figure 4.10: Play the best machine. choose the ith machine is p(i)=win(i)+1 win(i) + lose(i)+2: Thus, ifp(1)>p(2) you would play machine 1 and otherwise you would play machine 2. We have written a program TwoArm to simulate this experiment. In the program, the user specifles the initial values for xandy(but these are unknown to the experimenter). The program calculates at each stage the two conditionaldensities for xandy, given the outcomes of the previous trials, and then computes p(i), fori= 1, 2. It then chooses the machine with the highest value for the probability of winning for the next play. The program prints the machine chosenon each play and the outcome of this play. It also plots the new densities for x (solid line) and y(dotted line), showing only the current densities. We have run the program for ten plays for the case x=:6 andy=:7. The result is shown in Figure 4.10. The run of the program shows the weakness of this strategy. Our initial proba- bility for winning on the better of the two machines is .7. We start with the poorermachine and our outcomes are such that we always have a probability greater than.6 of winning and so we just keep playing this machine even though the other ma-chine is better. If we had lost on the flrst play we would have switched machines.Our flnal density for yis the same as our initial density, namely, the uniform den- sity. Our flnal density for xis difierent and re°ects a much more accurate knowledge aboutx. The computer did pretty well with this strategy, winning seven out of the ten trials, but ten trials are not enough to judge whether this is a good strategy inthe long run. Another popular strategy is the play-the-winner strategy. As the name suggests, for this strategy we choose the same machine when we win and switch machineswhen we lose. The program TwoArm will simulate this strategy as well. In Figure 4.11, we show the results of running this program with the play-the-winnerstrategy and the same true probabilities of .6 and .7 for the two machines. Afterten plays our densities for the unknown probabilities of winning suggest to us thatthe second machine is indeed the better of the two. We again won seven out of theten trials. 172 CHAPTER 4. CONDITIONAL PROBABILITY 0.2 0.4 0.6 0.8 10.511.52 Machine Result 1 W 1 W 1 L 2 L 1 W 1 W 1 L 2 L 1 L 2 W Figure 4.11: Play the winner. Neither of the strategies that we simulated is the best one in terms of maximizing our average winnings. This best strategy is very complicated but is reasonably ap-proximated by the play-the-winner strategy. Variations on this example have playedan important role in the problem of clinical tests of drugs where experimenters facea similar situation. 2 Exercises 1Pick a point xat random (with uniform density) in the interval [0 ;1]. Find the probability that x>1=2, given that (a)x>1=4. (b)x<3=4. (c)jx¡1=2j<1=4. (d)x2¡x+2=9<0. 2A radioactive material emits fi-particles at a rate described by the density function f(t)=:1e¡:1t: Find the probability that a particle is emitted in the flrst 10 seconds, given that (a) no particle is emitted in the flrst second. (b) no particle is emitted in the flrst 5 seconds. (c) a particle is emitted in the flrst 3 seconds. (d) a particle is emitted in the flrst 20 seconds. 3The Acme Super light bulb is known to have a useful life described by the density function f(t)=:01e¡:01t; where time tis measured in hours. 4.2. CONTINUOUS CONDITIONAL PROBABILITY 173 (a) Find the failure rate of this bulb (see Exercise 2.2.6). (b) Find the reliability of this bulb after 20 hours. (c) Given that it lasts 20 hours, flnd the probability that the bulb lasts another 20 hours. (d) Find the probability that the bulb burns out in the forty-flrst hour, given that it lasts 40 hours. 4Suppose you toss a dart at a circular target of radius 10 inches. Given that the dart lands in the upper half of the target, flnd the probability that (a) it lands in the right half of the target. (b) its distance from the center is less than 5 inches. (c) its distance from the center is greater than 5 inches. (d) it lands within 5 inches of the point (0 ;5). 5Suppose you choose two numbers xandy, independently at random from the interval [0 ;1]. Given that their sum lies in the interval [0 ;1], flnd the probability that (a)jx¡yj<1. (b)xy< 1=2. (c) maxfx;yg<1=2. (d)x2+y2<1=4. (e)x>y . 6Find the conditional density functions for the following experiments. (a) A number xis chosen at random in the interval [0 ;1], given that x>1=4. (b) A number tis chosen at random in the interval [0 ;1) with exponential densitye¡t, given that 1 <t< 10. (c) A dart is thrown at a circular target of radius 10 inches, given that it falls in the upper half of the target. (d) Two numbers xandyare chosen at random in the interval [0 ;1], given thatx>y . 7Letxandybe chosen at random from the interval [0 ;1]. Show that the events x>1=3 andy>2=3 are independent events. 8Letxandybe chosen at random from the interval [0 ;1]. Which pairs of the following events are independent? (a)x>1=3. (b)y>2=3. (c)x>y . 174 CHAPTER 4. CONDITIONAL PROBABILITY (d)x+y<1. 9Suppose that XandYare continuous random variables with density functions fX(x) andfY(y), respectively. Let f(x;y) denote the joint density function of (X;Y ). Show thatZ1 ¡1f(x;y)dy=fX(x); andZ1 ¡1f(x;y)dx=fY(y): *10 In Exercise 2.2.12 you proved the following: If you take a stick of unit length and break it into three pieces, choosing the breaks at random (i.e., choosingtwo real numbers independently and uniformly from [0, 1]), then the prob-ability that the three pieces form a triangle is 1/4. Consider now a similarexperiment: First break the stick at random, then break the longer pieceat random. Show that the two experiments are actually quite difierent, asfollows: (a) Write a program which simulates both cases for a run of 1000 trials, prints out the proportion of successes for each run, and repeats this process tentimes. (Call a trial a success if the three pieces do form a triangle.) Haveyour program pick ( x;y) at random in the unit square, and in each case usexandyto flnd the two breaks. For each experiment, have it plot (x;y)i f(x;y) gives a success. (b) Show that in the second experiment the theoretical probability of success is actually 2 log 2 ¡1. 11A coin has an unknown bias pthat is assumed to be uniformly distributed between 0 and 1. The coin is tossed ntimes and heads turns up jtimes and tails turns up ktimes. We have seen that the probability that heads turns up next time is j+1 n+2: Show that this is the same as the probability that the next ball is black for the Polya urn model of Exercise 4.1.20. Use this result to explain why, in thePolya urn model, the proportion of black balls does not tend to 0 or 1 as onemight expect but rather to a uniform distribution on the interval [0 ;1]. 12Previous experience with a drug suggests that the probability pthat the drug is efiective is a random quantity having a beta density with parameters fi=2 andfl= 3. The drug is used on ten subjects and found to be successful in four out of the ten patients. What density should we now assign to theprobability p? What is the probability that the drug will be successful the next time it is used? 4.3. PARADOXES 175 13Write a program to allow you to compare the strategies play-the-winner and play-the-best-machine for the two-armed bandit problem of Example 4.24.Have your program determine the initial payofi probabilities for each machineby choosing a pair of random numbers between 0 and 1. Have your programcarry out 20 plays and keep track of the number of wins for each of the twostrategies. Finally, have your program make 1000 repetitions of the 20 playsand compute the average winning per 20 plays. Which strategy seems tobe the best? Repeat these simulations with 20 replaced by 100. Does youranswer to the above question change? 14Consider the two-armed bandit problem of Example 4.24. Bruce Barnes pro- posed the following strategy, which is a variation on the play-the-best-machinestrategy. The machine with the greatest probability of winning is played un- lessthe following two conditions hold: (a) the difierence in the probabilities for winning is less than .08, and (b) the ratio of the number of times playedon the more often played machine to the number of times played on the lessoften played machine is greater than 1.4. If the above two conditions hold,then the machine with the smaller probability of winning is played. Write aprogram to simulate this strategy. Have your program choose the initial payofiprobabilities at random from the unit interval [0 ;1], make 20 plays, and keep track of the number of wins. Repeat this experiment 1000 times and obtainthe average number of wins per 20 plays. Implement a second strategy|forexample, play-the-best-machine or one of your own choice, and see how thissecond strategy compares with Bruce’s on average wins. 4.3 Paradoxes Much of this section is based on an article by Snell and Vanderbei.18 One must be very careful in dealing with problems involving conditional prob- ability. The reader will recall that in the Monty Hall problem (Example 4.6), ifthe contestant chooses the door with the car behind it, then Monty has a choice ofdoors to open. We made an assumption that in this case, he will choose each doorwith probability 1/2. We then noted that if this assumption is changed, the answerto the original question changes. In this section, we will study other examples ofthe same phenomenon. Example 4.25 Consider a family with two children. Given that one of the children is a boy, what is the probability that both children are boys? One way to approach this problem is to say that the other child is equally likely to be a boy or a girl, so the probability that both children are boys is 1/2. The \text-book" solution would be to draw the tree diagram and then form the conditionaltree by deleting paths to leave only those paths that are consistent with the given 18J. L. Snell and R. Vanderbei, \Three Bewitching Paradoxes," in Topics in Contemporary Probability and Its Applications , CRC Press, Boca Raton, 1995. 176 CHAPTER 4. CONDITIONAL PROBABILITY First childSecond child Conditional probabilityFirst childSecond childUnconditional probability b g b g b g b1/4 1/41/4 1/4 1/3 1/3 1/3 1/41/41/41/2 1/21/2 1/21/2b g1/2 1/21/2 gb Unconditional probability1/21/2 1/2 Figure 4.12: Tree for Example 4.25. information. The result is shown in Figure 4.12. We see that the probability of two boys given a boy in the family is not 1/2 but rather 1/3. 2 This problem and others like it are discussed in Bar-Hillel and Falk.19These authors stress that the answer to conditional probabilities of this kind can changedepending upon how the information given was actually obtained. For example,they show that 1/2 is the correct answer for the following scenario. Example 4.26 Mr. Smith is the father of two. We meet him walking along the street with a young boy whom he proudly introduces as his son. What is theprobability that Mr. Smith’s other child is also a boy? As usual we have to make some additional assumptions. For example, we will assume that if Mr. Smith has a boy and a girl, he is equally likely to choose eitherone to accompany him on his walk. In Figure 4.13 we show the tree analysis of thisproblem and we see that 1/2 is, indeed, the correct answer. 2 Example 4.27 It is not so easy to think of reasonable scenarios that would lead to the classical 1/3 answer. An attempt was made by Stephen Geller in proposing thisproblem to Marilyn vos Savant. 20Geller’s problem is as follows: A shopkeeper says she has two new baby beagles to show you, but she doesn’t know whether they’reboth male, both female, or one of each sex. You tell her that you want only a male,and she telephones the fellow who’s giving them a bath. \Is at least one a male?" 19M. Bar-Hillel and R. Falk, \Some teasers concerning conditional probabilities," Cognition , vol. 11 (1982), pgs. 109-122. 20M. vos Savant, \Ask Marilyn," Parade Magazine , 9 September; 2 December; 17 February 1990, reprinted in Marilyn vos Savant, Ask Marilyn , St. Martins, New York, 1992. 4.3. PARADOXES 177 Mr.Smith's childrenWalking with Mr.SmithUnconditional probability Mr.Smith's childrenWalking with Mr. SmithUnconditional probabilitybbbb g g gb1/4 1/8 1/81/81/8 1/41/41/4 1/41/41 1/21/21/21/2 1bg gb gg bbbb b1/4 1/81/8 1/41/41/41 1/21/2 bg gbConditional probability 1/2 1/4 1/4 Figure 4.13: Tree for Example 4.26. 178 CHAPTER 4. CONDITIONAL PROBABILITY she asks. \Yes," she informs you with a smile. What is the probability that the other one is male? The reader is asked to decide whether the model which gives an answer of 1/3 is a reasonable one to use in this case. 2 In the preceding examples, the apparent paradoxes could easily be resolved by clearly stating the model that is being used and the assumptions that are beingmade. We now turn to some examples in which the paradoxes are not so easilyresolved. Example 4.28 Two envelopes each contain a certain amount of money. One en- velope is given to Ali and the other to Baba and they are told that one envelopecontains twice as much money as the other. However, neither knows who has thelarger prize. Before anyone has opened their envelope, Ali is asked if she would liketo trade her envelope with Baba. She reasons as follows: Assume that the amountin my envelope is x. If I switch, I will end up with x=2 with probability 1/2, and 2xwith probability 1/2. If I were given the opportunity to play this game many times, and if I were to switch each time, I would, on average, get 1 2x 2+1 22x=5 4x: This is greater than my average winnings if I didn’t switch. Of course, Baba is presented with the same opportunity and reasons in the same way to conclude that he too would like to switch. So they switch and each thinksthat his/her net worth just went up by 25%. Since neither has yet opened any envelope, this process can be repeated and so again they switch. Now they are back with their original envelopes and yet theythink that their fortune has increased 25% twice. By this reasoning, they couldconvince themselves that by repeatedly switching the envelopes, they could becomearbitrarily wealthy. Clearly, something is wrong with the above reasoning, butwhere is the mistake? One of the tricks of making paradoxes is to make them slightly more di–cult than is necessary to further befuddle us. As John Finn has suggested, in this paradox wecould just have well started with a simpler problem. Suppose Ali and Baba knowthat I am going to give then either an envelope with $5 or one with $10 and I amgoing to toss a coin to decide which to give to Ali, and then give the other to Baba.Then Ali can argue that Baba has 2 xwith probability 1 =2 andx=2 with probability 1=2. This leads Ali to the same conclusion as before. But now it is clear that this is nonsense, since if Ali has the envelope containing $5, Baba cannot possibly havehalf of this, namely $2.50, since that was not even one of the choices. Similarly, ifAli has $10, Baba cannot have twice as much, namely $20. In fact, in this simplerproblem the possibly outcomes are given by the tree diagram in Figure 4.14. Fromthe diagram, it is clear that neither is made better ofi by switching. 2 In the above example, Ali’s reasoning is incorrect because he infers that if the amount in his envelope is x, then the probability that his envelope contains the 4.3. PARADOXES 179 $5 $10$10 $5 1/21/21/21 1 1/2In Ali's envelopeIn Baba's envelope Figure 4.14: John Finn’s version of Example 4.28. smaller amount is 1/2, and the probability that her envelope contains the larger amount is also 1/2. In fact, these conditional probabilities depend upon the distri-bution of the amounts that are placed in the envelopes. For deflniteness, let Xdenote the positive integer-valued random variable which represents the smaller of the two amounts in the envelopes. Suppose, in addition,that we are given the distribution of X, i.e., for each positive integer x, we are given the value of p x=P(X=x): (In Finn’s example, p5= 1, andpn= 0 for all other values of n.) Then it is easy to calculate the conditional probability that an envelope contains the smaller amount,given that it contains xdollars. The two possible sample points are ( x;x= 2) and (x;2x). Ifxis odd, then the flrst sample point has probability 0, since x=2i sn o t an integer, so the desired conditional probability is 1 that xis the smaller amount. Ifxis even, then the two sample points have probabilities p x=2andpx, respectively, so the conditional probability that xis the smaller amount is px px=2+px; which is not necessarily equal to 1/2. Steven Brams and D. Marc Kilgour21study the problem, for difierent distri- butions, of whether or not one should switch envelopes, if one’s objective is tomaximize the long-term average winnings. Let xbe the amount in your envelope. They show that for any distribution of X, there is at least one value of xsuch that you should switch. They give an example of a distribution for which there isexactly one value of xsuch that you should switch (see Exercise 5). Perhaps the most interesting case is a distribution in which you should always switch. We nowgive this example. Example 4.29 Suppose that we have two envelopes in front of us, and that one envelope contains twice the amount of money as the other (both amounts are pos-itive integers). We are given one of the envelopes, and asked if we would like toswitch. 21S. J. Brams and D. M. Kilgour, \The Box Problem: To Switch or Not to Switch," Mathematics Magazine , vol. 68, no. 1 (1995), p. 29. 180 CHAPTER 4. CONDITIONAL PROBABILITY As above, we let Xdenote the smaller of the two amounts in the envelopes, and let px=P(X=x): We are now in a position where we can calculate the long-term average winnings, if we switch. (This long-term average is an example of a probabilistic concept knownas expectation, and will be discussed in Chapter 6.) Given that one of the twosample points has occurred, the probability that it is the point ( x;x= 2) is p x=2 px=2+px; and the probability that it is the point ( x;2x)i s px px=2+px: Thus, if we switch, our long-term average winnings are px=2 px=2+pxx 2+px px=2+px2x: If this is greater than x, then it pays in the long run for us to switch. Some routine algebra shows that the above expression is greater than xif and only if px=2 px=2+px<2 3: (4.6) It is interesting to consider whether there is a distribution on the positive integers such that the inequality 4.6 is true for all even values of x. Brams and Kilgour22 give the following example. We deflnepxas follows: px=( 1 3‡ 2 3·k¡1 ;ifx=2k; 0; otherwise. It is easy to calculate (see Exercise 4) that for all relevant values of x,w eh a v e px=2 px=2+px=3 5; which means that the inequality 4.6 is always true. 2 So far, we have been able to resolve paradoxes by clearly stating the assumptions being made and by precisely stating the models being used. We end this section bydescribing a paradox which we cannot resolve. Example 4.30 Suppose that we have two envelopes in front of us, and we are told that the envelopes contain XandYdollars, respectively, where XandYare difierent positive integers. We randomly choose one of the envelopes, and we open 22ibid. 4.3. PARADOXES 181 it, revealing X, say. Is it possible to determine, with probability greater than 1/2, whetherXis the smaller of the two dollar amounts? Even if we have no knowledge of the joint distribution of XandY, the surprising answer is yes! Here’s how to do it. Toss a fair coin until the flrst time that headsturns up. Let Zdenote the number of tosses required plus 1/2. If Z>X , then we say thatXis the smaller of the two amounts, and if Z<X , then we say that Xis the larger of the two amounts. First, ifZlies between XandY, then we are sure to be correct. Since Xand Yare unequal, Zlies between them with positive probability. Second, if Zis not betweenXandY, thenZis either greater than both XandY, or is less than both XandY. In either case, Xis the smaller of the two amounts with probability 1/2, by symmetry considerations (remember, we chose the envelope at random). Thus,the probability that we are correct is greater than 1/2. 2 Exercises 1One of the flrst conditional probability paradoxes was provided by Bertrand.23 It is called the Box Paradox . A cabinet has three drawers. In the flrst drawer there are two gold balls, in the second drawer there are two silver balls, andin the third drawer there is one silver and one gold ball. A drawer is picked atrandom and a ball chosen at random from the two balls in the drawer. Giventhat a gold ball was drawn, what is the probability that the drawer with thetwo gold balls was chosen? In the next two problems, the reader is asked to assume that the deck has only four cards: the ace of hearts, the ace of spades, the king of hearts, andthe king of spades. (The reader is also invited to solve these problems usinga standard 52-card deck.) 2The following problem is called the two aces problem . This problem, dat- ing back to 1936, has been attributed to the English mathematician J. H.C. Whitehead (see Gridgeman 24). This problem was also submitted to Mar- ilyn vos Savant by the master of mathematical puzzles Martin Gardner, whoremarks that it is one of his favorites. A bridge hand has been dealt. Are the following two conditional probabilities equal? If a given hand has an ace, what is the probability that the given handhas a second ace? Given that the hand has the ace of hearts, what is theprobability that the hand has a second ace? 3In the preceding exercise, it is natural to ask \How do we get the information that the given hand has an ace?" Gridgeman considers two difierent ways thatwe might get this information. 23J. Bertrand, Calcul des Probabilit¶ es, Gauthier-Uillars, 1888. 24N. T. Gridgeman, Letter, American Statistician , 21 (1967), pgs. 38-39. 182 CHAPTER 4. CONDITIONAL PROBABILITY (a) Assume that the person holding the hand is asked to \Name an ace in your hand" and answers \The ace of hearts." What is the probabilitythat he has a second ace? (b) Suppose the person holding the hand is asked the more direct question \Do you have the ace of hearts?" and the answer is yes. What is theprobability that he has a second ace? 4Using the notation introduced in Example 4.29, show that in the example of Brams and Kilgour, if xis a positive power of 2, then p x=2 px=2+px=3 5: 5Using the notation introduced in Example 4.29, let px=( 2 3‡ 1 3·k ;ifx=2k; 0; otherwise. Show that there is exactly one value of xsuch that if your envelope contains x, then you should switch. *6(For bridge players only. From Sutherland.25) Suppose that we are the de- clarer in a hand of bridge, and we have the king, 9, 8, 7, and 2 of a certainsuit, while the dummy has the ace, 10, 5, and 4 of the same suit. Supposethat we want to play this suit in such a way as to maximize the probabilityof having no losers in the suit. We begin by leading the 2 to the ace, and wenote that the queen drops on our left. We then lead the 10 from the dummy,and our right-hand opponent plays the six (after playing the three on the flrstround). Should we flnesse or play for the drop? 25E. Sutherland, \Restricted Choice | Fact or Fiction?", Canadian Master Point , November 1, 1993. Chapter 5 Important Distributions and Densities 5.1 Important Distributions In this chapter, we describe the discrete probability distributions and the continuous probability densities that occur most often in the analysis of experiments. We willalso show how one simulates these distributions and densities on a computer. Discrete Uniform Distribution In Chapter 1, we saw that in many cases, we assume that all outcomes of an exper-iment are equally likely. If Xis a random variable which represents the outcome of an experiment of this type, then we say that Xis uniformly distributed. If the sample space Sis of sizen, where 0<n<1, then the distribution function m(!) is deflned to be 1 =nfor all!2S. As is the case with all of the discrete probabil- ity distributions discussed in this chapter, this experiment can be simulated on acomputer using the program GeneralSimulation . However, in this case, a faster algorithm can be used instead. (This algorithm was described in Chapter 1; werepeat the description here for completeness.) The expression 1+bn(rnd)c takes on as a value each integer between 1 and nwith probability 1 =n(the notation bxcdenotes the greatest integer not exceeding x). Thus, if the possible outcomes of the experiment are labelled ! 1!2;:::;!n, then we use the above expression to represent the subscript of the output of the experiment. If the sample space is a countably inflnite set, such as the set of positive integers, then it is not possible to have an experiment which is uniform on this set (seeExercise 3). If the sample space is an uncountable set, with positive, flnite length,such as the interval [0 ;1], then we use continuous density functions (see Section 5.2). 183 184 CHAPTER 5. DISTRIBUTIONS AND DENSITIES Binomial Distribution The binomial distribution with parameters n,p, andkwas deflned in Chapter 3. It is the distribution of the random variable which counts the number of heads whichoccur when a coin is tossed ntimes, assuming that on any one toss, the probability that a head occurs is p. The distribution function is given by the formula b(n;p;k )=µn k¶ p kqn¡k; whereq=1¡p. One straightforward way to simulate a binomial random variable Xis to compute the sum ofnindependent 0¡1 random variables, each of which take on the value 1 with probability p. This method requires ncalls to a random number generator to obtain one value of the random variable. When nis relatively large (say at least 30), the Central Limit Theorem (see Chapter 9) implies that the binomial distribution iswell-approximated by the corresponding normal density function (which is deflnedin Section 5.2) with parameters „=npand¾=p npq. Thus, in this case we can compute a value Yof a normal random variable with these parameters, and if ¡1=2•Y<n +1=2, we can use the value bY+1=2c to represent the random variable X.I fY<¡1=2o rY>n +1=2, we reject Yand compute another value. We will see in the next section how we can quickly simulatenormal random variables. Geometric Distribution Consider a Bernoulli trials process continued for an inflnite number of trials; forexample, a coin tossed an inflnite sequence of times. We showed in Section 2.2how to assign a probability measure to the inflnite tree. Thus, we can determinethe distribution for any random variable Xrelating to the experiment provided P(X=a) can be computed in terms of a flnite number of trials. For example, let Tbe the number of trials up to and including the flrst success. Then P(T=1 ) =p; P(T=2 ) =qp; P(T=3 ) =q 2p; and in general, P(T=n)=qn¡1p: To show that this is a distribution, we must show that p+qp+q2p+¢¢¢=1: 5.1. IMPORTANT DISTRIBUTIONS 185 0 5 10 15 2000.20.40.60.81 p = .5 0 5 10 15 2000.050.10.150.20.25 p = .2 Figure 5.1: Geometric distributions. The left-hand expression is just a geometric series with flrst term pand common ratioq, so its sum isp 1¡q which equals 1. In Figure 5.1 we have plotted this distribution using the program Geometric- Plot for the cases p=:5 andp=:2. We see that as pdecreases we are more likely to get large values for T, as would be expected. In both cases, the most probable value forTis 1. This will always be true since P(T=j+1 ) P(T=j)=q<1: In general, if 0 <p< 1, andq=1¡p, then we say that the random variable T has a geometric distribution if P(T=j)=qj¡1p; forj=1;2;3; ::: . To simulate the geometric distribution with parameter p, we can simply compute a sequence of random numbers in [0 ;1), stopping when an entry does not exceed p. However, for small values of p, this is time-consuming (taking, on the average, 1 =p steps). We now describe a method whose running time does not depend upon thesize ofp. LetXbe a geometrically distributed random variable with parameter p, where 0<p< 1. Now, deflne Yto be the smallest integer satisfying the inequality 1¡q Y‚rnd: (5.1) Then we have P(Y=j)=P‡ 1¡qj‚rnd> 1¡qj¡1· =qj¡1¡qj =qj¡1(1¡q) =qj¡1p: 186 CHAPTER 5. DISTRIBUTIONS AND DENSITIES Thus,Yis geometrically distributed with parameter p. To generate Y, all we have to do is solve Equation 5.1 for Y. We obtain Y=$ log(1¡rnd) logq% : Since log(1¡rnd) and log(rnd) are identically distributed, Ycan also be generated using the equation Y=$ logrnd logq% : Example 5.1 The geometric distribution plays an important role in the theory of queues, or waiting lines. For example, suppose a line of customers waits for serviceat a counter. It is often assumed that, in each small time unit, either 0 or 1 newcustomers arrive at the counter. The probability that a customer arrives is pand that no customer arrives is q=1¡p. Then the time Tuntil the next arrival has a geometric distribution. It is natural to ask for the probability that no customerarrives in the next ktime units, that is, for P(T>k ). This is given by P(T>k )= 1X j=k+1qj¡1p=qk(p+qp+q2p+¢¢¢) =qk: This probability can also be found by noting that we are asking for no successes (i.e., arrivals) in a sequence of kconsecutive time units, where the probability of a success in any one time unit is p. Thus, the probability is just qk, since arrivals in any two time units are independent events. It is often assumed that the length of time required to service a customer also has a geometric distribution but with a difierent value for p. This implies a rather special property of the service time. To see this, let us compute the conditionalprobability P(T>r +sjT>r )=P(T>r +s) P(T>r )=qr+s qr=qs: Thus, the probability that the customer’s service takes smore time units is inde- pendent of the length of time rthat the customer has already been served. Because of this interpretation, this property is called the \memoryless" property, and is alsoobeyed by the exponential distribution. (Fortunately, not too many service stationshave this property.) 2 Negative Binomial Distribution Suppose we are given a coin which has probability pof coming up heads when it is tossed. We flx a positive integer k, and toss the coin until the kth head appears. We letXrepresent the number of tosses. When k=1 ,Xis geometrically distributed. 5.1. IMPORTANT DISTRIBUTIONS 187 For a general k, we say that Xhas a negative binomial distribution. We now calculate the probability distribution of X.I fX=x, then it must be true that there were exactly k¡1 heads thrown in the flrst x¡1 tosses, and a head must have been thrown on the xth toss. There are µx¡1 k¡1¶ sequences of length xwith these properties, and each of them is assigned the same probability, namely pk¡1qx¡k: Therefore, if we deflne u(x;k;p )=P(X=x); then u(x;k;p )=µx¡1 k¡1¶ pkqx¡k: One can simulate this on a computer by simulating the tossing of a coin. The following algorithm is, in general, much faster. We note that Xcan be understood as the sum of koutcomes of a geometrically distributed experiment with parameter p. Thus, we can use the following sum as a means of generating X: kX j=1$ logrndj logq% : Example 5.2 A fair coin is tossed until the second time a head turns up. The distribution for the number of tosses is u(x;2;p). Thus the probability that xtosses are needed to obtain two heads is found by letting k= 2 in the above formula. We obtain u(x;2;1=2) =µx¡1 1¶1 2x; forx=2;3;:::. In Figure 5.2 we give a graph of the distribution for k= 2 andp=:25. Note that the distribution is quite asymmetric, with a long tail re°ecting the fact thatlarge values of xare possible. 2 Poisson Distribution The Poisson distribution arises in many situations. It is safe to say that it is one of the three most important discrete probability distributions (the other two being theuniform and the binomial distributions). The Poisson distribution can be viewedas arising from the binomial distribution or from the exponential density. We shallnow explain its connection with the former; its connection with the latter will beexplained in the next section. Suppose that we have a situation in which a certain kind of occurrence happens at random over a period of time. For example, the occurrences that we are interested 188 CHAPTER 5. DISTRIBUTIONS AND DENSITIES 5 10 15 20 25 3000.020.040.060.080.1 Figure 5.2: Negative binomial distribution with k= 2 andp=:25. in might be incoming telephone calls to a police station in a large city. We want to model this situation so that we can consider the probabilities of events suchas more than 10 phone calls occurring in a 5-minute time interval. Presumably,in our example, there would be more incoming calls between 6:00 and 7:00 P.M. than between 4:00 and 5:00 A.M., and this fact would certainly afiect the above probability. Thus, to have a hope of computing such probabilities, we must assumethat the average rate, i.e., the average number of occurrences per minute, is aconstant. This rate we will denote by ‚. (Thus, in a given 5-minute time interval, we would expect about 5 ‚occurrences.) This means that if we were to apply our model to the two time periods given above, we would simply use difierent ratesfor the two time periods, thereby obtaining two difierent probabilities for the givenevent. Our next assumption is that the number of occurrences in two non-overlapping time intervals are independent. In our example, this means that the events thatthere arejcalls between 5:00 and 5:15 P.M. andkcalls between 6:00 and 6:15 P.M. on the same day are independent. We can use the binomial distribution to model this situation. We imagine that a given time interval is broken up into nsubintervals of equal length. If the subin- tervals are su–ciently short, we can assume that two or more occurrences happenin one subinterval with a probability which is negligible in comparison with theprobability of at most one occurrence. Thus, in each subinterval, we are assumingthat there is either 0 or 1 occurrence. This means that the sequence of subintervalscan be thought of as a sequence of Bernoulli trials, with a success corresponding toan occurrence in the subinterval. To decide upon the proper value of p, the probability of an occurrence in a given subinterval, we reason as follows. On the average, there are ‚toccurrences in a 5.1. IMPORTANT DISTRIBUTIONS 189 time interval of length t. If this time interval is divided into nsubintervals, then we would expect, using the Bernoulli trials interpretation, that there should be np occurrences. Thus, we want ‚t=np ; so p=‚t n: We now wish to consider the random variable X, which counts the number of occurrences in a given time interval. We want to calculate the distribution of X. For ease of calculation, we will assume that the time interval is of length 1; for timeintervals of arbitrary length t, see Exercise 11. We know that P(X=0 )=b(n;p;0 )=( 1¡p) n=‡ 1¡‚ n·n : For largen, this is approximately e¡‚. It is easy to calculate that for any flxed k, we have b(n;p;k ) b(n;p;k¡1)=‚¡(k¡1)p kq which, for large n(and therefore small p) is approximately ‚=k. Thus, we have P(X=1 )…‚e¡‚; and in general, P(X=k)…‚k k!e¡‚: (5.2) The above distribution is the Poisson distribution. We note that it must be checked that the distribution given in Equation 5.2 really isa distribution, i.e., that its values are non-negative and sum to 1. (See Exercise 12.) The Poisson distribution is used as an approximation to the binomial distribu- tion when the parameters nandpare large and small, respectively (see Examples 5.3 and 5.4). However, the Poisson distribution also arises in situations where it maynot be easy to interpret or measure the parameters nandp(see Example 5.5). Example 5.3 A typesetter makes, on the average, one mistake per 1000 words. Assume that he is setting a book with 100 words to a page. Let S 100be the number of mistakes that he makes on a single page. Then the exact probability distributionforS 100would be obtained by considering S100as a result of 100 Bernoulli trials withp=1=1000. The expected value of S100is‚= 100(1=1000) =:1. The exact probability that S100=jisb(100;1=1000;j), and the Poisson approximation is e¡:1(:1)j j!: In Table 5.1 we give, for various values of nandp, the exact values computed by the binomial distribution and the Poisson approximation. 2 190 CHAPTER 5. DISTRIBUTIONS AND DENSITIES Poisson Binomial Poisson Binomial Poisson Binomial n= 100 n= 100 n= 1000 j‚=:1p=:001‚=1p=:01‚=1 0p=:01 0 .9048 .9048 .3679 .3660 .0000 .0000 1 .0905 .0905 .3679 .3697 .0005 .0004 2 .0045 .0045 .1839 .1849 .0023 .0022 3 .0002 .0002 .0613 .0610 .0076 .0074 4 .0000 .0000 .0153 .0149 .0189 .0186 5 .0031 .0029 .0378 .0374 6 .0005 .0005 .0631 .0627 7 .0001 .0001 .0901 .0900 8 .0000 .0000 .1126 .1128 9 .1251 .1256 10 .1251 .1257 11 .1137 .1143 12 .0948 .0952 13 .0729 .0731 14 .0521 .0520 15 .0347 .0345 16 .0217 .0215 17 .0128 .0126 18 .0071 .0069 19 .0037 .0036 20 .0019 .0018 21 .0009 .0009 22 .0004 .0004 23 .0002 .0002 24 .0001 .0001 25 .0000 .0000 Table 5.1: Poisson approximation to the binomial distribution. 5.1. IMPORTANT DISTRIBUTIONS 191 Example 5.4 In his book,1Feller discusses the statistics of °ying bomb hits in the south of London during the Second World War. Assume that you live in a district of size 10 blocks by 10 blocks so that the total district is divided into 100 small squares. How likely is it that the square in whichyou live will receive no hits if the total area is hit by 400 bombs? We assume that a particular bomb will hit your square with probability 1/100. Since there are 400 bombs, we can regard the number of hits that your squarereceives as the number of successes in a Bernoulli trials process with n= 400 and p=1=100. Thus we can use the Poisson distribution with ‚= 400¢1=100 = 4 to approximate the probability that your square will receive jhits. This probability isp(j)=e ¡44j=j!. The expected number of squares that receive exactly jhits is then 100¢p(j). It is easy to write a program LondonBombs to simulate this situation and compare the expected number of squares with jhits with the observed number. In Exercise 26 you are asked to compare the actual observed data withthat predicted by the Poisson distribution. In Figure 5.3, we have shown the simulated hits, together with a spike graph showing both the observed and predicted frequencies. The observed frequencies areshown as squares, and the predicted frequencies are shown as dots. 2 If the reader would rather not consider °ying bombs, he is invited to instead consider an analogous situation involving cookies and raisins. We assume that we have madeenough cookie dough for 500 cookies. We put 600 raisins in the dough, and mix itthoroughly. One way to look at this situation is that we have 500 cookies, and afterplacing the cookies in a grid on the table, we throw 600 raisins at the cookies. (SeeExercise 22.) Example 5.5 Suppose that in a certain flxed amount Aof blood, the average human has 40 white blood cells. Let Xbe the random variable which gives the number of white blood cells in a random sample of size Afrom a random individual. We can think of Xas binomially distributed with each white blood cell in the body representing a trial. If a given white blood cell turns up in the sample, then thetrial corresponding to that blood cell was a success. Then pshould be taken as the ratio of Ato the total amount of blood in the individual, and nwill be the number of white blood cells in the individual. Of course, in practice, neither ofthese parameters is very easy to measure accurately, but presumably the number40 is easy to measure. But for the average human, we then have 40 = np,s ow e can think of Xas being Poisson distributed, with parameter ‚= 40. In this case, it is easier to model the situation using the Poisson distribution than the binomialdistribution. 2 To simulate a Poisson random variable on a computer, a good way is to take advantage of the relationship between the Poisson distribution and the exponentialdensity. This relationship and the resulting simulation algorithm will be describedin the next section. 1ibid., p. 161. 192 CHAPTER 5. DISTRIBUTIONS AND DENSITIES 0 2 4 6 8 1000.050.10.150.2 Figure 5.3: Flying bomb hits. 5.1. IMPORTANT DISTRIBUTIONS 193 Hypergeometric Distribution Suppose that we have a set of Nballs, of which kare red and N¡kare blue. We choosenof these balls, without replacement, and deflne Xto be the number of red balls in our sample. The distribution of Xis called the hypergeometric distribution. We note that this distribution depends upon three parameters, namely N,k, and n. There does not seem to be a standard notation for this distribution; we will use the notation h(N;k;n;x ) to denote P(X=x). This probability can be found by noting that there areµN n¶ difierent samples of size n, and the number of such samples with exactly xred balls is obtained by multiplying the number of ways of choosing xred balls from the set ofkred balls and the number of ways of choosing n¡xblue balls from the set of N¡kblue balls. Hence, we have h(N;k;n;x )=¡k x¢¡N¡k n¡x¢ ¡N n¢: This distribution can be generalized to the case where there are more than two types of objects. (See Exercise 40.) If we letNandktend to1, in such a way that the ratio k=N remains flxed, then the hypergeometric distribution tends to the binomial distribution with parametersnandp=k=N. This is reasonable because if Nandkare much larger than n, then whether we choose our sample with or without replacement should not afiect theprobabilities very much, and the experiment consisting of choosing with replacementyields a binomially distributed random variable (see Exercise 44). An example of how this distribution might be used is given in Exercises 36 and 37. We now give another example involving the hypergeometric distribution. Itillustrates a statistical test called Fisher’s Exact Test. Example 5.6 It is often of interest to consider two traits, such as eye color and hair color, and to ask whether there is an association between the two traits. Twotraits are associated if knowing the value of one of the traits for a given personallows us to predict the value of the other trait for that person. The stronger theassociation, the more accurate the predictions become. If there is no associationbetween the traits, then we say that the traits are independent. In this example, wewill use the traits of gender and political party, and we will assume that there areonly two possible genders, female and male, and only two possible political parties,Democratic and Republican. Suppose that we have collected data concerning these traits. To test whether there is an association between the traits, we flrst assume that there is no associationbetween the two traits. This gives rise to an \expected" data set, in which knowledgeof the value of one trait is of no help in predicting the value of the other trait. Ourcollected data set usually difiers from this expected data set. If it difiers by quite abit, then we would tend to reject the assumption of independence of the traits. To 194 CHAPTER 5. DISTRIBUTIONS AND DENSITIES Democrat Republican Female 24 4 28 Male 8 14 22 32 18 50 Table 5.2: Observed data. Democrat Republican Female s11 s12 t11 Male s21 s22 t12 t21 t22 n Table 5.3: General data table. nail down what is meant by \quite a bit," we decide which possible data sets difier from the expected data set by at least as much as ours does, and then we computethe probability that any of these data sets would occur under the assumption ofindependence of traits. If this probability is small, then it is unlikely that thedifierence between our collected data set and the expected data set is due entirelyto chance. Suppose that we have collected the data shown in Table 5.2. The row and column sums are called marginal totals, or marginals. In what follows, we will denote therow sums by t 11andt12, and the column sums by t21andt22. Theijth entry in the table will be denoted by sij. Finally, the size of the data set will be denoted byn. Thus, a general data table will look as shown in Table 5.3. We now explain the model which will be used to construct the \expected" data set. In the model,we assume that the two traits are independent. We then put t 21yellow balls and t22green balls, corresponding to the Democratic and Republican marginals, into an urn. We draw t11balls, without replacement, from the urn, and call these balls females. The t12balls remaining in the urn are called males. In the speciflc case under consideration, the probability of getting the actual data under this model isgiven by the expression ¡ 32 24¢¡18 4¢ ¡5028¢; i.e., a value of the hypergeometric distribution. We are now ready to construct the expected data set. If we choose 28 balls out of 50, we should expect to see, on the average, the same percentage of yellowballs in our sample as in the urn. Thus, we should expect to see, on the average,28(32=5 0 )=1 7:92…18 yellow balls in our sample. (See Exercise 36.) The other expected values are computed in exactly the same way. Thus, the expected dataset is shown in Table 5.4. We note that the value of s 11determines the other three values in the table, since the marginals are all flxed. Thus, in consideringthe possible data sets that could appear in this model, it is enough to consider thevarious possible values of s 11. In the speciflc case at hand, what is the probability 5.1. IMPORTANT DISTRIBUTIONS 195 Democrat Republican Female 18 10 28 Male 14 8 22 32 18 50 Table 5.4: Expected data. of drawing exactly ayellow balls, i.e., what is the probability that s11=a?I ti s ¡32 a¢¡18 28¡a¢ ¡50 28¢: (5.3) We are now ready to decide whether our actual data difiers from the expected data set by an amount which is greater than could be reasonably attributed tochance alone. We note that the expected number of female Democrats is 18, butthe actual number in our data is 24. The other data sets which difier from theexpected data set by more than ours correspond to those where the number offemale Democrats equals 25, 26, 27, or 28. Thus, to obtain the required probability,we sum the expression in (5.3) from a=2 4t oa= 28. We obtain a value of :000395. Thus, we should reject the hypothesis that the two traits are independent. 2 Finally, we turn to the question of how to simulate a hypergeometric random variableX. Let us assume that the parameters for XareN,k, andn. We imagine that we have a set of Nballs, labelled from 1 to N. We decree that the flrst kof these balls are red, and the rest are blue. Suppose that we have chosen mballs, and thatjof them are red. Then there are k¡jred balls left, and N¡mballs left. Thus, our next choice will be red with probability k¡j N¡m: So at this stage, we choose a random number in [0 ;1], and report that a red ball has been chosen if and only if the random number does not exceed the above expression.Then we update the values of mandj, and continue until nballs have been chosen. Benford Distribution Our next example of a distribution comes from the study of leading digits in data sets. It turns out that many data sets that occur \in real life" have the property thatthe flrst digits of the data are not uniformly distributed over the set f1;2;:::; 9g. Rather, it appears that the digit 1 is most likely to occur, and that the distributionis monotonically decreasing on the set of possible digits. The Benford distributionappears, in many cases, to flt such data. Many explanations have been given for theoccurrence of this distribution. Possibly the most convincing explanation is thatthis distribution is the only one that is invariant under a change of scale. If onethinks of certain data sets as somehow \naturally occurring," then the distributionshould be unafiected by which units are chosen in which to represent the data, i.e.,the distribution should be invariant under change of scale. 196 CHAPTER 5. DISTRIBUTIONS AND DENSITIES 2 4 6 800.050.10.150.20.250.3 Figure 5.4: Leading digits in President Clinton’s tax returns. Theodore Hill2gives a general description of the Benford distribution, when one considers the flrst ddigits of integers in a data set. We will restrict our attention to the flrst digit. In this case, the Benford distribution has distribution function f(k) = log10(k+1 )¡log10(k); for 1•k•9. Mark Nigrini3has advocated the use of the Benford distribution as a means of testing suspicious flnancial records such as bookkeeping entries, checks, and taxreturns. His idea is that if someone were to \make up" numbers in these cases,the person would probably produce numbers that are fairly uniformly distributed,while if one were to use the actual numbers, the leading digits would roughly followthe Benford distribution. As an example, Negrini analyzed President Clinton’s taxreturns for a 13-year period. In Figure 5.4, the Benford distribution values areshown as squares, and the President’s tax return data are shown as circles. Onesees that in this example, the Benford distribution flts the data very well. This distribution was discovered by the astronomer Simon Newcomb who stated the following in his paper on the subject: \That the ten digits do not occur withequal frequency must be evident to anyone making use of logarithm tables, andnoticing how much faster the flrst pages wear out than the last ones. The flrstsigniflcant flgure is oftener 1 than any other digit, and the frequency diminishes upto 9." 4 2T. P. Hill, \The Signiflcant Digit Phenomenon," American Mathematical Monthly, vol. 102, no. 4 (April 1995), pgs. 322-327. 3M. Nigrini, \Detecting Biases and Irregularities in Tabulated Data," working paper 4S. Newcomb, \Note on the frequency of use of the difierent digits in natural numbers," Amer- ican Journal of Mathematics, vol. 4 (1881), pgs. 39-40. 5.1. IMPORTANT DISTRIBUTIONS 197 Exercises 1For which of the following random variables would it be appropriate to assign a uniform distribution? (a) LetXrepresent the roll of one die. (b) LetXrepresent the number of heads obtained in three tosses of a coin. (c) A roulette wheel has 38 possible outcomes: 0, 00, and 1 through 36. Let Xrepresent the outcome when a roulette wheel is spun. (d) LetXrepresent the birthday of a randomly chosen person. (e) LetXrepresent the number of tosses of a coin necessary to achieve a head for the flrst time. 2Letnbe a positive integer. Let Sbe the set of integers between 1 and n. Consider the following process: We remove a number from Sand write it down. We repeat this until Sis empty. The result is a permutation of the integers from 1 to n. LetXdenote this permutation. Is Xuniformly distributed? 3LetXbe a random variable which can take on countably many values. Show thatXcannot be uniformly distributed. 4Suppose we are attending a college which has 3000 students. We wish to choose a subset of size 100 from the student body. Let Xrepresent the subset, chosen using the following possible strategies. For which strategies would itbe appropriate to assign the uniform distribution to X? If it is appropriate, what probability should we assign to each outcome? (a) Take the flrst 100 students who enter the cafeteria to eat lunch. (b) Ask the Registrar to sort the students by their Social Security number, and then take the flrst 100 in the resulting list. (c) Ask the Registrar for a set of cards, with each card containing the name of exactly one student, and with each student appearing on exactly onecard. Throw the cards out of a third-story window, then walk outsideand pick up the flrst 100 cards that you flnd. 5Under the same conditions as in the preceding exercise, can you describe a procedure which, if used, would produce each possible outcome with thesame probability? Can you describe such a procedure that does not rely on acomputer or a calculator? 6LetX 1;X2; :::; Xnbenmutually independent random variables, each of which is uniformly distributed on the integers from 1 to k. LetYdenote the minimum of the Xi’s. Find the distribution of Y. 7A die is rolled until the flrst time Tthat a six turns up. (a) What is the probability distribution for T? 198 CHAPTER 5. DISTRIBUTIONS AND DENSITIES (b) FindP(T>3). (c) FindP(T>6jT>3). 8If a coin is tossed a sequence of times, what is the probability that the flrst head will occur after the flfth toss, given that it has not occurred in the flrsttwo tosses? 9A worker for the Department of Fish and Game is assigned the job of esti- mating the number of trout in a certain lake of modest size. She proceeds asfollows: She catches 100 trout, tags each of them, and puts them back in thelake. One month later, she catches 100 more trout, and notes that 10 of themhave tags. (a) Without doing any fancy calculations, give a rough estimate of the num- ber of trout in the lake. (b) LetNbe the number of trout in the lake. Find an expression, in terms ofN, for the probability that the worker would catch 10 tagged trout out of the 100 trout that she caught the second time. (c) Find the value of Nwhich maximizes the expression in part (b). This value is called the maximum likelihood estimate for the unknown quantity N.Hint: Consider the ratio of the expressions for successive values of N. 10A census in the United States is an attempt to count everyone in the country. It is inevitable that many people are not counted. The U. S. Census Bureauproposed a way to estimate the number of people who were not counted bythe latest census. Their proposal was as follows: In a given locality, let N denote the actual number of people who live there. Assume that the censuscountedn 1people living in this area. Now, another census was taken in the locality, and n2people were counted. In addition, n12people were counted both times. (a) GivenN,n1, andn2, letXdenote the number of people counted both times. Find the probability that X=k, wherekis a flxed positive integer between 0 and n2. (b) Now assume that X=n12. Find the value of Nwhich maximizes the expression in part (a). Hint: Consider the ratio of the expressions for successive values of N. 11Suppose that Xis a random variable which represents the number of calls coming in to a police station in a one-minute interval. In the text, we showedthatXcould be modelled using a Poisson distribution with parameter ‚, where this parameter represents the average number of incoming calls perminute. Now suppose that Yis a random variable which represents the num- ber of incoming calls in an interval of length t. Show that the distribution of Yis given by P(Y=k)=e ¡‚t(‚t)k k!; 5.1. IMPORTANT DISTRIBUTIONS 199 i.e.,Yis Poisson with parameter ‚t.Hint: Suppose a Martian were to observe the police station. Let us also assume that the basic time interval used onMars is exactly tEarth minutes. Finally, we will assume that the Martian understands the derivation of the Poisson distribution in the text. Whatwould she write down for the distribution of Y? 12Show that the values of the Poisson distribution given in Equation 5.2 sum to 1. 13The Poisson distribution with parameter ‚=:3 has been assigned for the outcome of an experiment. Let Xbe the outcome function. Find P(X= 0), P(X= 1), andP(X> 1). 14On the average, only 1 person in 1000 has a particular rare blood type. (a) Find the probability that, in a city of 10,000 people, no one has this blood type. (b) How many people would have to be tested to give a probability greater than 1/2 of flnding at least one person with this blood type? 15Write a program for the user to input n,p,jand have the program print out the exact value of b(n;p;k ) and the Poisson approximation to this value. 16Assume that, during each second, a Dartmouth switchboard receives one call with probability .01 and no calls with probability .99. Use the Poisson ap-proximation to estimate the probability that the operator will miss at mostone call if she takes a 5-minute cofiee break. 17The probability of a royal °ush in a poker hand is p=1=649;740. How large mustnbe to render the probability of having no royal °ush in nhands smaller than 1=e? 18A baker blends 600 raisins and 400 chocolate chips into a dough mix and, from this, makes 500 cookies. (a) Find the probability that a randomly picked cookie will have no raisins. (b) Find the probability that a randomly picked cookie will have exactly two chocolate chips. (c) Find the probability that a randomly chosen cookie will have at least two bits (raisins or chips) in it. 19The probability that, in a bridge deal, one of the four hands has all hearts is approximately 6 :3£10 ¡12. In a city with about 50,000 bridge players the resident probability expert is called on the average once a year (usually late atnight) and told that the caller has just been dealt a hand of all hearts. Shouldshe suspect that some of these callers are the victims of practical jokes? 200 CHAPTER 5. DISTRIBUTIONS AND DENSITIES 20An advertiser drops 10,000 lea°ets on a city which has 2000 blocks. Assume that each lea°et has an equal chance of landing on each block. What is theprobability that a particular block will receive no lea°ets? 21In a class of 80 students, the professor calls on 1 student chosen at random for a recitation in each class period. There are 32 class periods in a term. (a) Write a formula for the exact probability that a given student is called uponjtimes during the term. (b) Write a formula for the Poisson approximation for this probability. Using your formula estimate the probability that a given student is called uponmore than twice. 22Assume that we are making raisin cookies. We put a box of 600 raisins into our dough mix, mix up the dough, then make from the dough 500 cookies.We then ask for the probability that a randomly chosen cookie will have0, 1, 2, . . . raisins. Consider the cookies as trials in an experiment, andletXbe the random variable which gives the number of raisins in a given cookie. Then we can regard the number of raisins in a cookie as the resultofn= 600 independent trials with probability p=1=500 for success on each trial. Since nis large and pis small, we can use the Poisson approximation with‚= 600(1=500) = 1:2. Determine the probability that a given cookie will have at least flve raisins. 23For a certain experiment, the Poisson distribution with parameter ‚=mhas been assigned. Show that a most probable outcome for the experiment isthe integer value ksuch thatm¡1•k•m. Under what conditions will there be two most probable values? Hint: Consider the ratio of successive probabilities. 24When John Kemeny was chair of the Mathematics Department at Dartmouth College, he received an average of ten letters each day. On a certain weekdayhe received no mail and wondered if it was a holiday. To decide this hecomputed the probability that, in ten years, he would have at least 1 daywithout any mail. He assumed that the number of letters he received on agiven day has a Poisson distribution. What probability did he flnd? Hint: Apply the Poisson distribution twice. First, to flnd the probability that, in3000 days, he will have at least 1 day without mail, assuming each year hasabout 300 days on which mail is delivered. 25Reese Prosser never puts money in a 10-cent parking meter in Hanover. He assumes that there is a probability of .05 that he will be caught. The flrstofiense costs nothing, the second costs 2 dollars, and subsequent ofienses cost5 dollars each. Under his assumptions, how does the expected cost of parking100 times without paying the meter compare with the cost of paying the metereach time? 5.1. IMPORTANT DISTRIBUTIONS 201 Number of deaths Number of corps with xdeaths in a given year 0 144 1 91 2 32 3 11 4 2 Table 5.5: Mule kicks. 26Feller5discusses the statistics of °ying bomb hits in an area in the south of London during the Second World War. The area in question was divided into24£24 = 576 small areas. The total number of hits was 537. There were 229 squares with 0 hits, 211 with 1 hit, 93 with 2 hits, 35 with 3 hits, 7 with4 hits, and 1 with 5 or more. Assuming the hits were purely random, use thePoisson approximation to flnd the probability that a particular square wouldhave exactly khits. Compute the expected number of squares that would have 0, 1, 2, 3, 4, and 5 or more hits and compare this with the observedresults. 27Assume that the probability that there is a signiflcant accident in a nuclear power plant during one year’s time is .001. If a country has 100 nuclear plants,estimate the probability that there is at least one such accident during a givenyear. 28An airline flnds that 4 percent of the passengers that make reservations on a particular °ight will not show up. Consequently, their policy is to sell 100reserved seats on a plane that has only 98 seats. Find the probability thatevery person who shows up for the °ight will flnd a seat available. 29The king’s coinmaster boxes his coins 500 to a box and puts 1 counterfeit coin in each box. The king is suspicious, but, instead of testing all the coins in1 box, he tests 1 coin chosen at random out of each of 500 boxes. What is theprobability that he flnds at least one fake? What is it if the king tests 2 coinsfrom each of 250 boxes? 30(From Kemeny 6) Show that, if you make 100 bets on the number 17 at roulette at Monte Carlo (see Example 6.13), you will have a probability greaterthan 1/2 of coming out ahead. What is your expected winning? 31In one of the flrst studies of the Poisson distribution, von Bortkiewicz 7con- sidered the frequency of deaths from kicks in the Prussian army corps. Fromthe study of 14 corps over a 20-year period, he obtained the data shown inTable 5.5. Fit a Poisson distribution to this data and see if you think thatthe Poisson distribution is appropriate. 5ibid., p. 161. 6Private communication. 7L. von Bortkiewicz, Das Gesetz der Kleinen Zahlen (Leipzig: Teubner, 1898), p. 24. 202 CHAPTER 5. DISTRIBUTIONS AND DENSITIES 32It is often assumed that the auto tra–c that arrives at the intersection during a unit time period has a Poisson distribution with expected value m. Assume that the number of cars Xthat arrive at an intersection from the north in unit time has a Poisson distribution with parameter ‚=mand the number Ythat arrive from the west in unit time has a Poisson distribution with parameter‚=„m.I fXandYare independent, show that the total number X+Y that arrive at the intersection in unit time has a Poisson distribution withparameter‚=m+„m. 33Cars coming along Magnolia Street come to a fork in the road and have to choose either Willow Street or Main Street to continue. Assume that thenumber of cars that arrive at the fork in unit time has a Poisson distributionwith parameter ‚= 4. A car arriving at the fork chooses Main Street with probability 3/4 and Willow Street with probability 1/4. Let Xbe the random variable which counts the number of cars that, in a given unit of time, passby Joe’s Barber Shop on Main Street. What is the distribution of X? 34In the appeal of the People v. Collins case (see Exercise 4.1.28), the counsel for the defense argued as follows: Suppose, for example, there are 5,000,000couples in the Los Angeles area and the probability that a randomly chosencouple flts the witnesses’ description is 1/12,000,000. Then the probabilitythat there are two such couples given that there is at least one is not at allsmall. Find this probability. (The California Supreme Court overturned theinitial guilty verdict.) 35A manufactured lot of brass turnbuckles has Sitems of which Dare defective. A sample of sitems is drawn without replacement. Let Xbe a random variable that gives the number of defective items in the sample. Let p(d)=P(X=d). (a) Show that p(d)=¡ D d¢¡S¡D s¡d¢ ¡S s¢: Thus, X is hypergeometric. (b) Prove the following identity, known as Euler’s formula : min(D;s)X d=0µD d¶µS¡D s¡d¶ =µS s¶ : 36A bin of 1000 turnbuckles has an unknown number Dof defectives. A sample of 100 turnbuckles has 2 defectives. The maximum likelihood estimate forD is the number of defectives which gives the highest probability for obtainingthe number of defectives observed in the sample. Guess this number Dand then write a computer program to verify your guess. 37There are an unknown number of moose on Isle Royale (a National Park in Lake Superior). To estimate the number of moose, 50 moose are captured and 5.1. IMPORTANT DISTRIBUTIONS 203 tagged. Six months later 200 moose are captured and it is found that 8 of these were tagged. Estimate the number of moose on Isle Royale from thesedata, and then verify your guess by computer program (see Exercise 36). 38A manufactured lot of buggy whips has 20 items, of which 5 are defective. A random sample of 5 items is chosen to be inspected. Find the probability thatthe sample contains exactly one defective item (a) if the sampling is done with replacement. (b) if the sampling is done without replacement. 39Suppose that Nandktend to1in such a way that k=N remains flxed. Show that h(N;k;n;x )!b(n;k=N;x ): 40A bridge deck has 52 cards with 13 cards in each of four suits: spades, hearts, diamonds, and clubs. A hand of 13 cards is dealt from a shu†ed deck. Findthe probability that the hand has (a) a distribution of suits 4, 4, 3, 2 (for example, four spades, four hearts, three diamonds, two clubs). (b) a distribution of suits 5, 3, 3, 2. 41Write a computer algorithm that simulates a hypergeometric random variable with parameters N,k, andn. 42You are presented with four difierent dice. The flrst one has two sides marked 0 and four sides marked 4. The second one ha sa3o ne v e r y side. The third one h a sa2o n four sides an da6o nt w o sides, and the fourth one ha sa1o n three sides an da5o n three sides. You allow your friend to pick any of the four dice he wishes. Then you pick one of the remaining three and you each rollyour die. The person with the largest number showing wins a dollar. Showthat you can choose your die so that you have probability 2/3 of winning nomatter which die your friend picks. (See Tenney and Foster. 8) 43The students in a certain class were classifled by hair color and eye color. The conventions used were: Brown and black hair were considered dark, and redand blonde hair were considered light; black and brown eyes were considereddark, and blue and green eyes were considered light. They collected the datashown in Table 5.6. Are these traits independent? (See Example 5.6.) 44Suppose that in the hypergeometric distribution, we let Nandktend to1in such a way that the ratio k=N approaches a real number pbetween 0 and 1. Show that the hypergeometric distribution tends to the binomial distributionwith parameters nandp. 8R. L. Tenney and C. C. Foster, Non-transitive Dominance , Math. Mag. 49 (1976) no. 3, pgs. 115-120. 204 CHAPTER 5. DISTRIBUTIONS AND DENSITIES Dark Eyes Light Eyes Dark Hair 28 15 43 Light Hair 9 23 32 37 38 75 Table 5.6: Observed data. 0 10 20 30 400500100015002000250030003500 Figure 5.5: Distribution of choices in the Powerball lottery. 45(a) Compute the leading digits of the flrst 100 powers of 2, and see how well these data flt the Benford distribution. (b) Multiply each number in the data set of part (a) by 3, and compare the distribution of the leading digits with the Benford distribution. 46In the Powerball lottery, contestants pick 5 difierent integers between 1 and 45, and in addition, pick a bonus integer from the same range (the bonus integercan equal one of the flrst flve integers chosen). Some contestants choose thenumbers themselves, and others let the computer choose the numbers. Thedata shown in Table 5.7 are the contestant-chosen numbers in a certain stateon May 3, 1996. A spike graph of the data is shown in Figure 5.5. The goal of this problem is to check the hypothesis that the chosen numbers are uniformly distributed. To do this, compute the value vof the random variable ´ 2given in Example 5.10. In the present case, this random variable has 44 degrees of freedom. One can flnd, in a ´2table, the value v0=5 9:43 , which represents a number with the property that a ´2-distributed random variable takes on values that exceed v0only 5% of the time. Does your computed value ofvexceedv0? If so, you should reject the hypothesis that the contestants’ choices are uniformly distributed. 5.2. IMPORTANT DENSITIES 205 Integer Times Integer Times Integer Times Chosen Chosen Chosen 1 2646 2 2934 3 33524 3000 5 3357 6 28927 3657 8 3025 9 336210 2985 11 3138 12 304313 2690 14 2423 15 255616 2456 17 2479 18 227619 2304 20 1971 21 254322 2678 23 2729 24 241425 2616 26 2426 27 238128 2059 29 2039 30 229831 2081 32 1508 33 188734 1463 35 1594 36 135437 1049 38 1165 39 124840 1493 41 1322 42 142343 1207 44 1259 45 1224 Table 5.7: Numbers chosen by contestants in the Powerball lottery. 5.2 Important Densities In this section, we will introduce some important probability density functions and give some examples of their use. We will also consider the question of how onesimulates a given density using a computer. Continuous Uniform Density The simplest density function corresponds to the random variable Uwhose value represents the outcome of the experiment consisting of choosing a real number atrandom from the interval [ a;b]. f(!)=‰1=(b¡a);ifa•!•b; 0; otherwise. It is easy to simulate this density on a computer. We simply calculate the expression (b¡a)rnd+a: Exponential and Gamma Densities The exponential density function is deflned by f(x)=‰‚e¡‚x;if 0•x<1; 0; otherwise: Here‚is any positive constant, depending on the experiment. The reader has seen this density in Example 2.17. In Figure 5.6 we show graphs of several exponen-tial densities for difierent choices of ‚. The exponential density is often used to 206 CHAPTER 5. DISTRIBUTIONS AND DENSITIES 0 2 4 6 8 10λ=1λ=2 λ=1/2 Figure 5.6: Exponential densities. describe experiments involving a question of the form: How long until something happens? For example, the exponential density is often used to study the timebetween emissions of particles from a radioactive source. The cumulative distribution function of the exponential density is easy to com- pute. LetTbe an exponentially distributed random variable with parameter ‚.I f x‚0, then we have F(x)=P(T•x) =Z x 0‚e¡‚tdt =1¡e¡‚x: Both the exponential density and the geometric distribution share a property known as the \memoryless" property. This property was introduced in Example 5.1;it says that P(T>r +sjT>r )=P(T>s ): This can be demonstrated to hold for the exponential density by computing both sides of this equation. The right-hand side is just 1¡F(s)=e ¡‚s; while the left-hand side is P(T>r +s) P(T>r )=1¡F(r+s) 1¡F(s) 5.2. IMPORTANT DENSITIES 207 =e¡‚(r+s) e¡‚r =e¡‚s: There is a very important relationship between the exponential density and the Poisson distribution. We begin by deflning X1;X2; ::: to be a sequence of independent exponentially distributed random variables with parameter ‚.W e might think of Xias denoting the amount of time between the ith and (i+ 1)st emissions of a particle by a radioactive source. (As we shall see in Chapter 6, wecan think of the parameter ‚as representing the reciprocal of the average length of time between emissions. This parameter is a quantity that might be measured inan actual experiment of this type.) We now consider a time interval of length t, and we let Ydenote the random variable which counts the number of emissions that occur in the time interval. Wewould like to calculate the distribution function of Y(clearly,Yis a discrete random variable). If we let S ndenote the sum X1+X2+¢¢¢+Xn, then it is easy to see that P(Y=n)=P(Sn•tandSn+1>t): Since the event Sn+1•tis a subset of the event Sn•t, the above probability is seen to be equal to P(Sn•t)¡P(Sn+1•t): (5.4) We will show in Chapter 7 that the density of Snis given by the following formula: gn(x)=( ‚(‚x)n¡1 (n¡1)!e¡‚x;ifx>0, 0; otherwise. This density is an example of a gamma density with parameters ‚andn. The general gamma density allows nto be any positive real number. We shall not discuss this general density. It is easy to show by induction on nthat the cumulative distribution function ofSnis given by: Gn(x)=8 < :1¡e¡‚xµ 1+‚x 1!+¢¢¢+(‚x)n¡1 (n¡1)!¶ ;ifx>0; 0; otherwise. Using this expression, the quantity in (5.4) is easy to compute; we obtain e¡‚t(‚t)n n!; which the reader will recognize as the probability that a Poisson-distributed random variable, with parameter ‚t, takes on the value n. The above relationship will allow us to simulate a Poisson distribution, once we have found a way to simulate an exponential density. The following randomvariable does the job: Y=¡1 ‚log(rnd): (5.5) 208 CHAPTER 5. DISTRIBUTIONS AND DENSITIES Using Corollary 5.2 (below), one can derive the above expression (see Exercise 3). We content ourselves for now with a short calculation that should convince thereader that the random variable Yhas the required property. We have P(Y•y)=P‡ ¡1 ‚log(rnd)•y· =P(log(rnd)‚¡‚y) =P(rnd‚e¡‚y) =1¡e¡‚y: This last expression is seen to be the cumulative distribution function of an expo- nentially distributed random variable with parameter ‚. To simulate a Poisson random variable Wwith parameter ‚, we simply generate a sequence of values of an exponentially distributed random variable with the sameparameter, and keep track of the subtotals S kof these values. We stop generating the sequence when the subtotal flrst exceeds ‚. Assume that we flnd that Sn•‚<Sn+1: Then the value nis returned as a simulated value for W. Example 5.7 (Queues) Suppose that customers arrive at random times at a service station with one server, and suppose that each customer is served immediately ifno one is ahead of him, but must wait his turn in line otherwise. How long shouldeach customer expect to wait? (We deflne the waiting time of a customer to be thelength of time between the time that he arrives and the time that he begins to beserved.) Let us assume that the interarrival times between successive customers are given by random variables X 1,X2,...,Xnthat are mutually independent and identically distributed with an exponential cumulative distribution function given by FX(t)=1¡e¡‚t: Let us assume, too, that the service times for successive customers are given by random variables Y1,Y2,...,Ynthat again are mutually independent and identically distributed with another exponential cumulative distribution function given by FY(t)=1¡e¡„t: The parameters ‚and„represent, respectively, the reciprocals of the average time between arrivals of customers and the average service time of the customers.Thus, for example, the larger the value of ‚, the smaller the average time between arrivals of customers. We can guess that the length of time a customer will spendin the queue depends on the relative sizes of the average interarrival time and theaverage service time. It is easy to verify this conjecture by simulation. The program Queue simulates this queueing process. Let N(t) be the number of customers in the queue at time t. 5.2. IMPORTANT DENSITIES 209 2000 4000 6000 8000 10000102030405060 2000 4000 6000 8000 1000020040060080010001200λ = 1 λ = 1 µ = .9 µ = 1.1 Figure 5.7: Queue sizes. 0 10 20 30 40 5000.010.020.030.040.050.060.07 Figure 5.8: Waiting times. Then we plot N(t) as a function of tfor difierent choices of the parameters ‚and „(see Figure 5.7). We note that when ‚<„ , then 1=‚ > 1=„, so the average interarrival time is greater than the average service time, i.e., customers are served more quickly, onaverage, than new ones arrive. Thus, in this case, it is reasonable to expect thatN(t) remains small. However, if ‚>„ then customers arrive more quickly than they are served, and, as expected, N(t) appears to grow without limit. We can now ask: How long will a customer have to wait in the queue for service? To examine this question, we let W ibe the length of time that the ith customer has to remain in the system (waiting in line and being served). Then we can presentthese data in a bar graph, using the program Queue , to give some idea of how the W iare distributed (see Figure 5.8). (Here ‚= 1 and„=1:1.) We see that these waiting times appear to be distributed exponentially. This is always the case when ‚<„ . The proof of this fact is too complicated to give here, but we can verify it by simulation for difierent choices of ‚and„,a sa b o v e . 2 210 CHAPTER 5. DISTRIBUTIONS AND DENSITIES Functions of a Random Variable Before continuing our list of important densities, we pause to consider random variables which are functions of other random variables. We will prove a generaltheorem that will allow us to derive expressions such as Equation 5.5. Theorem 5.1 LetXbe a continuous random variable, and suppose that `(x)i sa strictly increasing function on the range of X. DeflneY=`(X). Suppose that X andYhave cumulative distribution functions F XandFYrespectively. Then these functions are related by FY(y)=FX(`¡1(y)): If`(x) is strictly decreasing on the range of X, then FY(y)=1¡FX(`¡1(y)): Proof. Since`is a strictly increasing function on the range of X, the events (X•`¡1(y)) and (`(X)•y) are equal. Thus, we have FY(y)=P(Y•y) =P(`(X)•y) =P(X•`¡1(y)) =FX(`¡1(y)): If`(x) is strictly decreasing on the range of X, then we have FY(y)=P(Y•y) =P(`(X)•y) =P(X‚`¡1(y)) =1¡P(X<`¡1(y)) =1¡FX(`¡1(y)): This completes the proof. 2 Corollary 5.1 LetXbe a continuous random variable, and suppose that `(x)i sa strictly increasing function on the range of X. DeflneY=`(X). Suppose that the density functions of XandYarefXandfY, respectively. Then these functions are related by fY(y)=fX(`¡1(y))d dy`¡1(y): If`(x) is strictly decreasing on the range of X, then fY(y)=¡fX(`¡1(y))d dy`¡1(y): 5.2. IMPORTANT DENSITIES 211 Proof. This result follows from Theorem 5.1 by using the Chain Rule. 2 If the function `is neither strictly increasing nor strictly decreasing, then the situation is somewhat more complicated but can be treated by the same methods.For example, suppose that Y=X 2, Then`(x)=x2, and FY(y)=P(Y•y) =P(¡py•X•+py) =P(X•+py)¡P(X•¡py) =FX(py)¡FX(¡py): Moreover, fY(y)=d dyFY(y) =d dy(FX(py)¡FX(¡py)) =‡ fX(py)+fX(¡py)·1 2py: We see that in order to express FYin terms of FXwhenY=`(X), we have to expressP(Y•y) in terms of P(X•x), and this process will depend in general upon the structure of `. Simulation Theorem 5.1 tells us, among other things, how to simulate on the computer a random variableYwith a prescribed cumulative distribution function F. We assume that F(y) is strictly increasing for those values of ywhere 0<F(y)<1. For this purpose, let Ube a random variable which is uniformly distributed on [0 ;1]. Then Uhas cumulative distribution function FU(u)=u.N o w , i fFis the prescribed cumulative distribution function for Y, then to write Yin terms of Uwe flrst solve the equation F(y)=u foryin terms of u. We obtain y=F¡1(u). Note that since Fis an increasing function this equation always has a unique solution (see Figure 5.9). Then we setZ=F ¡1(U) and obtain, by Theorem 5.1, FZ(y)=FU(F(y)) =F(y); sinceFU(u)=u. Therefore, ZandYhave the same cumulative distribution func- tion. Summarizing, we have the following. 212 CHAPTER 5. DISTRIBUTIONS AND DENSITIES y = φ(x)x = FY(y)Y(y) Graph of Fx y1 0 Figure 5.9: Converting a uniform distribution FUinto a prescribed distribution FY. Corollary 5.2 IfF(y) is a given cumulative distribution function that is strictly increasing when 0 <F(y)<1 and ifUis a random variable with uniform distribu- tion on [0;1], then Y=F¡1(U) has the cumulative distribution F(y). 2 Thus, to simulate a random variable with a given cumulative distribution Fwe need only set Y=F¡1(rnd). Normal Density We now come to the most important density function, the normal density function. We have seen in Chapter 3 that the binomial distribution functions are bell-shaped,even for moderate size values of n. We recall that a binomially-distributed random variable with parameters nandpcan be considered to be the sum of nmutually independent 0-1 random variables. A very important theorem in probability theory,called the Central Limit Theorem, states that under very general conditions, if wesum a large number of mutually independent random variables, then the distributionof the sum can be closely approximated by a certain speciflc continuous density,called the normal density. This theorem will be discussed in Chapter 9. The normal density function with parameters „and¾is deflned as follows: f X(x)=1p 2…¾e¡(x¡„)2=2¾2: The parameter „represents the \center" of the density (and in Chapter 6, we will show that it is the average, or expected, value of the density). The parameter ¾ is a measure of the \spread" of the density, and thus it is assumed to be positive.(In Chapter 6, we will show that ¾is the standard deviation of the density.) We note that it is not at all obvious that the above function is a density, i.e., that its 5.2. IMPORTANT DENSITIES 213 -4 -2 2 40.10.20.30.4 σ = 1 σ = 2 Figure 5.10: Normal density for two sets of parameter values. integral over the real line equals 1. The cumulative distribution function is given by the formula FX(x)=Zx ¡11p 2…¾e¡(u¡„)2=2¾2du : In Figure 5.10 we have included for comparison a plot of the normal density for the cases„= 0 and¾= 1, and„= 0 and¾=2 . One cannot write FXin terms of simple functions. This leads to several prob- lems. First of all, values of FXmust be computed using numerical integration. Extensive tables exist containing values of this function (see Appendix A). Sec-ondly, we cannot write F ¡1 Xin closed form, so we cannot use Corollary 5.2 to help us simulate a normal random variable. For this reason, special methods have beendeveloped for simulating a normal distribution. One such method relies on the factthat ifUandVare independent random variables with uniform densities on [0 ;1], then the random variables X=p ¡2 logUcos 2…V and Y=p ¡2 logUsin 2…V are independent, and have normal density functions with parameters „= 0 and ¾= 1. (This is not obvious, nor shall we prove it here. See Box and Muller.9) LetZbe a normal random variable with parameters „= 0 and¾=1 . A normal random variable with these parameters is said to be a standard normal random variable. It is an important and useful fact that if we write X=¾Z+„; thenXis a normal random variable with parameters „and¾. To show this, we will use Theorem 5.1. We have `(z)=¾z+„,`¡1(x)=(x¡„)=¾, and FX(x)=FZµx¡„ ¾¶ ; 9G. E. P. Box and M. E. Muller, A Note on the Generation of Random Normal Deviates , Ann. of Math. Stat. 29 (1958), pgs. 610-611. 214 CHAPTER 5. DISTRIBUTIONS AND DENSITIES fX(x)=fZµx¡„ ¾¶ ¢1 ¾ =1p 2…¾e¡(x¡„)2=2¾2: The reader will note that this last expression is the density function with parameters „and¾, as claimed. We have seen above that it is possible to simulate a standard normal random variableZ. If we wish to simulate a normal random variable Xwith parameters „ and¾, then we need only transform the simulated values for Zusing the equation X=¾Z+„. Suppose that we wish to calculate the value of a cumulative distribution function for the normal random variable X, with parameters „and¾. We can reduce this calculation to one concerning the standard normal random variable Zas follows: FX(x)=P(X•x) =Pµ Z•x¡„ ¾¶ =FZµx¡„ ¾¶ : This last expression can be found in a table of values of the cumulative distribution function for a standard normal random variable. Thus, we see that it is unnecessaryto make tables of normal distribution functions with arbitrary „and¾. The process of changing a normal random variable to a standard normal ran- dom variable is known as standardization. If Xhas a normal distribution with parameters „and¾and if Z=X¡„ ¾; thenZis said to be the standardized version of X. The following example shows how we use the standardized version of a normal random variable Xto compute speciflc probabilities relating to X. Example 5.8 Suppose that Xis a normally distributed random variable with pa- rameters„= 10 and¾= 3. Find the probability that Xis between 4 and 16. To solve this problem, we note that Z=(X¡10)=3 is the standardized version ofX. So, we have P(4•X•16) =P(X•16)¡P(X•4) =FX(16)¡FX(4) =FZµ16¡10 3¶ ¡FZµ4¡10 3¶ =FZ(2)¡FZ(¡2): 5.2. IMPORTANT DENSITIES 215 0 1 2 3 4 500.10.20.30.40.50.6 Figure 5.11: Distribution of dart distances in 1000 drops. This last expression can be evaluated by using tabulated values of the standard normal distribution function (see 12.3); when we use this table, we flnd that FZ(2) = :9772 andFZ(¡2) =:0228. Thus, the answer is .9544. In Chapter 6, we will see that the parameter „is the mean, or average value, of the random variable X. The parameter ¾is a measure of the spread of the random variable, and is called the standard deviation. Thus, the question asked in thisexample is of a typical type, namely, what is the probability that a random variablehas a value within two standard deviations of its average value. 2 Maxwell and Rayleigh Densities Example 5.9 Suppose that we drop a dart on a large table top, which we consider as thexy-plane, and suppose that the xandycoordinates of the dart point are independent and have a normal distribution with parameters „= 0 and¾=1 . How is the distance of the point from the origin distributed? This problem arises in physics when it is assumed that a moving particle in Rnhas components of the velocity that are mutually independent and normally distributed and it is desired to flnd the density of the speed of the particle. Thedensity in the case n= 3 is called the Maxwell density. The density in the case n= 2 (i.e. the dart board experiment described above) is called the Rayleigh density. We can simulate this case by picking independently apair of coordinates ( x;y), each from a normal distribution with „= 0 and¾=1o n (¡1;1), calculating the distance r=p x2+y2of the point ( x;y) from the origin, repeating this process a large number of times, and then presenting the results in abar graph. The results are shown in Figure 5.11. 216 CHAPTER 5. DISTRIBUTIONS AND DENSITIES Female Male A 37 56 93 B 63 60 123 C 47 43 90 Below C 5 8 13 152 167 319 Table 5.8: Calculus class data. Female Male A 44.3 48.7 93 B 58.6 64.4 123 C 42.9 47.1 90 Below C 6.2 6.8 13 152 167 319 Table 5.9: Expected data. We have also plotted the theoretical density f(r)=re¡r2=2: This will be derived in Chapter 7; see Example 7.7. 2 Chi-Squared Density We return to the problem of independence of traits discussed in Example 5.6. It is frequently the case that we have two traits, each of which have several difierentvalues. As was seen in the example, quite a lot of calculation was needed evenin the case of two values for each trait. We now give another method for testingindependence of traits, which involves much less calculation. Example 5.10 Suppose that we have the data shown in Table 5.8 concerning grades and gender of students in a Calculus class. We can use the same sort ofmodel in this situation as was used in Example 5.6. We imagine that we have anurn with 319 balls of two colors, say blue and red, corresponding to females andmales, respectively. We now draw 93 balls, without replacement, from the urn.These balls correspond to the grade of A. We continue by drawing 123 balls, whichcorrespond to the grade of B. When we flnish, we have four sets of balls, with eachball belonging to exactly one set. (We could have stipulated that the balls wereof four colors, corresponding to the four possible grades. In this case, we woulddraw a subset of size 152, which would correspond to the females. The balls re-maining in the urn would correspond to the males. The choice does not afiect theflnal determination of whether we should reject the hypothesis of independence oftraits.) The expected data set can be determined in exactly the same way as in Exam- ple 5.6. If we do this, we obtain the expected values shown in Table 5.9. Even if 5.2. IMPORTANT DENSITIES 217 the traits are independent, we would still expect to see some difierences between the numbers in corresponding boxes in the two tables. However, if the difierencesare large, then we might suspect that the two traits are not independent. In Ex-ample 5.6, we used the probability distribution of the various possible data sets tocompute the probability of flnding a data set that difiers from the expected dataset by at least as much as the actual data set does. We could do the same in thiscase, but the amount of computation is enormous. Instead, we will describe a single number which does a good job of measuring how far a given data set is from the expected one. To quantify how far apart the twosets of numbers are, we could sum the squares of the difierences of the correspondingnumbers. (We could also sum the absolute values of the difierences, but we wouldnot want to sum the difierences.) Suppose that we have data in which we expectto see 10 objects of a certain type, but instead we see 18, while in another case weexpect to see 50 objects of a certain type, but instead we see 58. Even though thetwo difierences are about the same, the flrst difierence is more surprising than thesecond, since the expected number of outcomes in the second case is quite a bitlarger than the expected number in the flrst case. One way to correct for this isto divide the individual squares of the difierences by the expected number for thatbox. Thus, if we label the values in the eight boxes in the flrst table by O i(for observed values) and the values in the eight boxes in the second table by Ei(for expected values), then the following expression might be a reasonable one to use tomeasure how far the observed data is from what is expected: 8X i=1(Oi¡Ei)2 Ei: This expression is a random variable, which is usually denoted by the symbol ´2, pronounced \ki-squared." It is called this because, under the assumption of inde-pendence of the two traits, the density of this random variable can be computed andis approximately equal to a density called the chi-squared density. We choose notto give the explicit expression for this density, since it involves the gamma function,which we have not discussed. The chi-squared density is, in fact, a special case ofthe general gamma density. In applying the chi-squared density, tables of values of this density are used, as in the case of the normal density. The chi-squared density has one parameter n, which is called the number of degrees of freedom. The number nis usually easy to determine from the problem at hand. For example, if we are checking two traits forindependence, and the two traits have aandbvalues, respectively, then the number of degrees of freedom of the random variable ´ 2is (a¡1)(b¡1). So, in the example at hand, the number of degrees of freedom is 3. We recall that in this example, we are trying to test for independence of the two traits of gender and grades. If we assume these traits are independent, thenthe ball-and-urn model given above gives us a way to simulate the experiment.Using a computer, we have performed 1000 experiments, and for each one, we havecalculated a value of the random variable ´ 2. The results are shown in Figure 5.12, together with the chi-squared density function with three degrees of freedom. 218 CHAPTER 5. DISTRIBUTIONS AND DENSITIES 0 2 4 6 8 10 1200.050.10.150.2 Figure 5.12: Chi-squared density with three degrees of freedom. As we stated above, if the value of the random variable ´2is large, then we would tend not to believe that the two traits are independent. But how large islarge? The actual value of this random variable for the data above is 4.13. InFigure 5.12, we have shown the chi-squared density with 3 degrees of freedom. Itcan be seen that the value 4.13 is larger than most of the values taken on by thisrandom variable. Typically, a statistician will compute the value vof the random variable ´ 2, just as we have done. Then, by looking in a table of values of the chi-squareddensity, a value v 0is determined which is only exceeded 5% of the time. If v‚v0, the statistician rejects the hypothesis that the two traits are independent. In thepresent case, v 0=7:815, so we would not reject the hypothesis that the two traits are independent. 2 Cauchy Density The following example is from Feller.10 Example 5.11 Suppose that a mirror is mounted on a vertical axis, and is free to revolve about that axis. The axis of the mirror is 1 foot from a straight wallof inflnite length. A pulse of light is shown onto the mirror, and the re°ected rayhits the wall. Let `be the angle between the re°ected ray and the line that is perpendicular to the wall and that runs through the axis of the mirror. We assumethat`is uniformly distributed between ¡…=2 and…=2. LetXrepresent the distance between the point on the wall that is hit by the re°ected ray and the point on thewall that is closest to the axis of the mirror. We now determine the density of X. LetBbe a flxed positive quantity. Then X‚Bif and only if tan( `)‚B, which happens if and only if `‚arctan(B). This happens with probability …=2¡arctan(B) …: 10W. Feller, An Introduction to Probability Theory and Its Applications, , vol. 2, (New York: Wiley, 1966) 5.2. IMPORTANT DENSITIES 219 Thus, for positive B, the cumulative distribution function of Xis F(B)=1¡…=2¡arctan(B) …: Therefore, the density function for positive Bis f(B)=1 …(1 +B2): Since the physical situation is symmetric with respect to `= 0, it is easy to see that the above expression for the density is correct for negative values of Bas well. The Law of Large Numbers, which we will discuss in Chapter 8, states that in many cases, if we take the average of independent values of a random variable,then the average approaches a speciflc number as the number of values increases.It turns out that if one does this with a Cauchy-distributed random variable, theaverage does not approach any speciflc number. 2 Exercises 1Choose a number Ufrom the unit interval [0 ;1] with uniform distribution. Find the cumulative distribution and density for the random variables (a)Y=U+2 . (b)Y=U3. 2Choose a number Ufrom the interval [0 ;1] with uniform distribution. Find the cumulative distribution and density for the random variables (a)Y=1=(U+ 1). (b)Y= log(U+ 1). 3Use Corollary 5.2 to derive the expression for the random variable given in Equation 5.5. Hint: The random variables 1 ¡rndandrndare identically distributed. 4Suppose we know a random variable Yas a function of the uniform random variableU:Y=`(U), and suppose we have calculated the cumulative dis- tribution function FY(y) and thence the density fY(y). How can we check whether our answer is correct? An easy simulation provides the answer: Makea bar graph of Y=`(rnd) and compare the result with the graph of f Y(y). These graphs should look similar. Check your answers to Exercises 1 and 2by this method. 5Choose a number Ufrom the interval [0 ;1] with uniform distribution. Find the cumulative distribution and density for the random variables (a)Y=jU¡1=2j. (b)Y=(U¡1=2) 2. 220 CHAPTER 5. DISTRIBUTIONS AND DENSITIES 6Check your results for Exercise 5 by simulation as described in Exercise 4. 7Explain how you can generate a random variable whose cumulative distribu- tion function is F(x)=8 < :0;ifx<0; x2;if 0•x•1; 1;ifx>1: 8Write a program to generate a sample of 1000 random outcomes each of which is chosen from the distribution given in Exercise 7. Plot a bar graph of yourresults and compare this empirical density with the density for the cumulativedistribution given in Exercise 7. 9LetU,Vbe random numbers chosen independently from the interval [0 ;1] with uniform distribution. Find the cumulative distribution and density ofeach of the variables (a)Y=U+V. (b)Y=jU¡Vj. 10LetU,Vbe random numbers chosen independently from the interval [0 ;1]. Find the cumulative distribution and density for the random variables (a)Y= max(U;V). (b)Y= min(U;V). 11Write a program to simulate the random variables of Exercises 9 and 10 and plot a bar graph of the results. Compare the resulting empirical density withthe density found in Exercises 9 and 10. 12A numberUis chosen at random in the interval [0 ;1]. Find the probability that (a)R=U 2<1=4. (b)S=U(1¡U)<1=4. (c)T=U=(1¡U)<1=4. 13Find the cumulative distribution function Fand the density function ffor each of the random variables R,S, andTin Exercise 12. 14A pointPin the unit square has coordinates XandYchosen at random in the interval [0 ;1]. LetDbe the distance from Pto the nearest edge of the square, and Ethe distance to the nearest corner. What is the probability that (a)D< 1=4? (b)E< 1=4? 15In Exercise 14 flnd the cumulative distribution Fand density ffor the random variableD. 5.2. IMPORTANT DENSITIES 221 16LetXbe a random variable with density function fX(x)=‰ cx(1¡x);if 0<x< 1; 0; otherwise. (a) What is the value of c? (b) What is the cumulative distribution function FXforX? (c) What is the probability that X< 1=4? 17LetXbe a random variable with cumulative distribution function F(x)=8 < :0; ifx<0; sin2(…x=2);if 0•x•1; 1; if 1<x: (a) What is the density function fXforX? (b) What is the probability that X< 1=4? 18LetXbe a random variable with cumulative distribution function FX, and letY=X+b,Z=aX, andW=aX+b, whereaandbare any constants. Find the cumulative distribution functions FY,FZ, andFW.Hint: The cases a>0,a= 0, anda<0 require difierent arguments. 19LetXbe a random variable with density function fX, and letY=X+b, Z=aX, andW=aX+b, wherea6= 0. Find the density functions fY,fZ, andfW. (See Exercise 18.) 20LetXbe a random variable uniformly distributed over [ c;d], and letY= aX+b. For what choice of aandbisYuniformly distributed over [0 ;1]? 21LetXbe a random variable with cumulative distribution function Fstrictly increasing on the range of X. LetY=F(X). Show that Yis uniformly distributed in the interval [0 ;1]. (The formula X=F¡1(Y) then tells us how to construct Xfrom a uniform random variable Y.) 22LetXbe a random variable with cumulative distribution function F. The median ofXis the value mfor whichF(m)=1=2. ThenX<m with probability 1/2 and X>m with probability 1/2. Find mifXis (a) uniformly distributed over the interval [ a;b]. (b) normally distributed with parameters „and¾. (c) exponentially distributed with parameter ‚. 23LetXbe a random variable with density function fX. The mean ofXis the value„=R xfx(x)dx. Then„gives an average value for X(see Sec- tion 6.3). Find „ifXis distributed uniformly, normally, or exponentially, as in Exercise 22. 222 CHAPTER 5. DISTRIBUTIONS AND DENSITIES Test Score Letter grade „+¾<x A „<x<„ +¾ B „¡¾<x<„ C „¡2¾<x<„¡¾ D x<„¡2¾ F Table 5.10: Grading on the curve. 24LetXbe a random variable with density function fX. The mode ofXis the valueMfor whichf(M) is maximum. Then values of XnearMare most likely to occur. Find MifXis distributed normally or exponentially, as in Exercise 22. What happens if Xis distributed uniformly? 25LetXbe a random variable normally distributed with parameters „= 70, ¾= 10. Estimate (a)P(X> 50). (b)P(X< 60). (c)P(X> 90). (d)P(60<X< 80). 26Bridies’ Bearing Works manufactures bearing shafts whose diameters are nor- mally distributed with parameters „=1 ,¾=:002. The buyer’s speciflcations require these diameters to be 1 :000§:003 cm. What fraction of the manu- facturer’s shafts are likely to be rejected? If the manufacturer improves herquality control, she can reduce the value of ¾. What value of ¾will ensure that no more than 1 percent of her shafts are likely to be rejected? 27A flnal examination at Podunk University is constructed so that the test scores are approximately normally distributed, with parameters „and¾. The instructor assigns letter grades to the test scores as shown in Table 5.10 (thisis the process of \grading on the curve"). What fraction of the class gets A, B, C, D, F? 28(Ross 11) An expert witness in a paternity suit testifles that the length (in days) of a pregnancy, from conception to delivery, is approximately normallydistributed, with parameters „= 270,¾= 10. The defendant in the suit is able to prove that he was out of the country during the period from 290 to 240days before the birth of the child. What is the probability that the defendantwas in the country when the child was conceived? 29Suppose that the time (in hours) required to repair a car is an exponentially distributed random variable with parameter ‚=1=2. What is the probabil- ity that the repair time exceeds 4 hours? If it exceeds 4 hours what is theprobability that it exceeds 8 hours? 11S. Ross, A First Course in Probability Theory, 2d ed. (New York: Macmillan, 1984). 5.2. IMPORTANT DENSITIES 223 30Suppose that the number of years a car will run is exponentially distributed with parameter „=1=4. If Prosser buys a used car today, what is the probability that it will still run after 4 years? 31LetUbe a uniformly distributed random variable on [0 ;1]. What is the probability that the equation x2+4Ux+1=0 has two distinct real roots x1andx2? 32Write a program to simulate the random variables whose densities are given by the following, making a suitable bar graph of each and comparing the exactdensity with the bar graph. (a)f X(x)=e¡xon [0;1) (but just do it on [0 ;10]): (b)fX(x)=2xon [0;1]: (c)fX(x)=3x2on [0;1]: (d)fX(x)=4jx¡1=2jon [0;1]: 33Suppose we are observing a process such that the time between occurrences is exponentially distributed with ‚=1=30 (i.e., the average time between occurrences is 30 minutes). Suppose that the process starts at a certain timeand we start observing the process 3 hours later. Write a program to simulatethis process. Let Tdenote the length of time that we have to wait, after we start our observation, for an occurrence. Have your program keep track of T. What is an estimate for the average value of T? 34Jones puts in two new lightbulbs: a 60 watt bulb and a 100 watt bulb. It is claimed that the lifetime of the 60 watt bulb has an exponential densitywith average lifetime 200 hours ( ‚=1=200). The 100 watt bulb also has an exponential density but with average lifetime of only 100 hours ( ‚=1=100). Jones wonders what is the probability that the 100 watt bulb will outlast the60 watt bulb. IfXandYare two independent random variables with exponential densities f(x)=‚e ¡‚xandg(x)=„e¡„x, respectively, then the probability that Xis less thanYis given by P(X<Y )=Z1 0f(x)(1¡G(x))dx; whereG(x) is the cumulative distribution function for g(x). Explain why this is the case. Use this to show that P(X<Y )=‚ ‚+„ and to answer Jones’s question. 224 CHAPTER 5. DISTRIBUTIONS AND DENSITIES 35Consider the simple queueing process of Example 5.7. Suppose that you watch the size of the queue. If there are jpeople in the queue the next time the queue size changes it will either decrease to j¡1 or increase to j+ 1. Use the result of Exercise 34 to show that the probability that the queue sizedecreases to j¡1i s„=(„+‚) and the probability that it increases to j+1 is‚=(„+‚). When the queue size is 0 it can only increase to 1. Write a program to simulate the queue size. Use this simulation to help formulate aconjecture containing conditions on „and‚that will ensure that the queue will have times when it is empty. 36LetXbe a random variable having an exponential density with parameter ‚. Find the density for the random variable Y=rX, whereris a positive real number. 37LetXbe a random variable having a normal density and consider the random variableY=e X. ThenYhas a log normal density. Find this density of Y. 38LetX1andX2be independent random variables and for i=1;2, letYi= `i(Xi), where`iis strictly increasing on the range of Xi. Show that Y1and Y2are independent. Note that the same result is true without the assumption that the`i’s are strictly increasing, but the proof is more di–cult. Chapter 6 Expected Value and Variance 6.1 Expected Value of Discrete Random Variables When a large collection of numbers is assembled, as in a census, we are usually interested not in the individual numbers, but rather in certain descriptive quantitiessuch as the average or the median. In general, the same is true for the probabilitydistribution of a numerically-valued random variable. In this and in the next section,we shall discuss two such descriptive quantities: the expected value and the variance. Both of these quantities apply only to numerically-valued random variables, and sowe assume, in these sections, that all random variables have numerical values. Togive some intuitive justiflcation for our deflnition, we consider the following game. Average Value A die is rolled. If an odd number turns up, we win an amount equal to this number;if an even number turns up, we lose an amount equal to this number. For example,if a two turns up we lose 2, and if a three comes up we win 3. We want to decide ifthis is a reasonable game to play. We flrst try simulation. The program Diecarries out this simulation. The program prints the frequency and the relative frequency with which each outcome occurs. It also calculates the average winnings. We have run the programtwice. The results are shown in Table 6.1. In the flrst run we have played the game 100 times. In this run our average gain is¡:57. It looks as if the game is unfavorable, and we wonder how unfavorable it really is. To get a better idea, we have played the game 10,000 times. In this caseour average gain is ¡:4949. We note that the relative frequency of each of the six possible outcomes is quite close to the probability 1/6 for this outcome. This corresponds to our frequencyinterpretation of probability. It also suggests that for very large numbers of plays,our average gain should be „=1‡1 6· ¡2‡1 6· +3‡1 6· ¡4‡1 6· +5‡1 6· ¡6‡1 6· 225 226 CHAPTER 6. EXPECTED VALUE AND VARIANCE n = 100 n = 10000 Winning Frequency Relative Frequency Relative Frequency Frequency 1 17 .17 1681 .1681 -2 17 .17 1678 .1678 3 16 .16 1626 .1626 -4 18 .18 1696 .1696 5 16 .16 1686 .1686 -6 16 .16 1633 .1633 Table 6.1: Frequencies for dice game. =9 6¡12 6=¡3 6=¡:5: This agrees quite well with our average gain for 10,000 plays. We note that the value we have chosen for the average gain is obtained by taking the possible outcomes, multiplying by the probability, and adding the results. Thissuggests the following deflnition for the expected outcome of an experiment. Expected Value Deflnition 6.1 LetXbe a numerically-valued discrete random variable with sam- ple space › and distribution function m(x). The expected value E(X) is deflned by E(X)=X x2›xm(x); provided this sum converges absolutely. We often refer to the expected value as themean, and denote E(X)b y„for short. If the above sum does not converge absolutely, then we say that Xdoes not have an expected value. 2 Example 6.1 Let an experiment consist of tossing a fair coin three times. Let Xdenote the number of heads which appear. Then the possible values of Xare 0;1;2 and 3. The corresponding probabilities are 1 =8;3=8;3=8;and 1=8. Thus, the expected value of Xequals 0µ1 8¶ +1µ3 8¶ +2µ3 8¶ +3µ1 8¶ =3 2: Later in this section we shall see a quicker way to compute this expected value, based on the fact that Xcan be written as a sum of simpler random variables. 2 Example 6.2 Suppose that we toss a fair coin until a head flrst comes up, and let Xrepresent the number of tosses which were made. Then the possible values of X are 1;2;:::, and the distribution function of Xis deflned by m(i)=1 2i: 6.1. EXPECTED VALUE 227 (This is just the geometric distribution with parameter 1 =2.) Thus, we have E(X)=1X i=1i1 2i =1X i=11 2i+1X i=21 2i+¢¢¢ =1 +1 2+1 22+¢¢¢ =2: 2 Example 6.3 (Example 6.2 continued) Suppose that we °ip a coin until a head flrst appears, and if the number of tosses equals n, then we are paid 2ndollars. What is the expected value of the payment? We letYrepresent the payment. Then, P(Y=2n)=1 2n; forn‚1. Thus, E(Y)=1X n=12n1 2n; which is a divergent sum. Thus, Yhas no expectation. This example is called theSt. Petersburg Paradox . The fact that the above sum is inflnite suggests that a player should be willing to pay any flxed amount per game for the privilege ofplaying this game. The reader is asked to consider how much he or she would bewilling to pay for this privilege. It is unlikely that the reader’s answer is more than10 dollars; therein lies the paradox. In the early history of probability, various mathematicians gave ways to resolve this paradox. One idea (due to G. Cramer) consists of assuming that the amountof money in the world is flnite. He thus assumes that there is some flxed value ofnsuch that if the number of tosses equals or exceeds n, the payment is 2 ndollars. The reader is asked to show in Exercise 20 that the expected value of the paymentis now flnite. Daniel Bernoulli and Cramer also considered another way to assign value to the payment. Their idea was that the value of a payment is some function of thepayment; such a function is now called a utility function. Examples of reasonableutility functions might include the square-root function or the logarithm function.In both cases, the value of 2 ndollars is less than twice the value of ndollars. It can easily be shown that in both cases, the expected utility of the payment is flnite(see Exercise 20). 2 228 CHAPTER 6. EXPECTED VALUE AND VARIANCE Example 6.4 LetTbe the time for the flrst success in a Bernoulli trials process. Then we take as sample space › the integers 1 ;2; ::: and assign the geometric distribution m(j)=P(T=j)=qj¡1p: Thus, E(T)=1¢p+2qp+3q2p+¢¢¢ =p( 1+2q+3q2+¢¢¢): Now ifjxj<1, then 1+x+x2+x3+¢¢¢=1 1¡x: Difierentiating this formula, we get 1+2x+3x2+¢¢¢=1 (1¡x)2; so E(T)=p (1¡q)2=p p2=1 p: In particular, we see that if we toss a fair coin a sequence of times, the expected time until the flrst heads is 1/(1/2) = 2. If we roll a die a sequence of times, theexpected number of rolls until the flrst six is 1/(1/6) = 6. 2 Interpretation of Expected Value In statistics, one is frequently concerned with the average value of a set of data. The following example shows that the ideas of average value and expected value arevery closely related. Example 6.5 The heights, in inches, of the women on the Swarthmore basketball team are 5’ 9", 5’ 9", 5’ 6", 5’ 8", 5’ 11", 5’ 5", 5’ 7", 5’ 6", 5’ 6", 5’ 7", 5’ 10", and6’ 0". A statistician would compute the average height (in inches) as follows: 6 9+6 9+6 6+6 8+7 1+6 5+6 7+6 6+6 6+6 7+7 0+7 2 12=6 7:9: One can also interpret this number as the expected value of a random variable. To see this, let an experiment consist of choosing one of the women at random, and letXdenote her height. Then the expected value of Xequals 67.9. 2 Of course, just as with the frequency interpretation of probability, to interpret expected value as an average outcome requires further justiflcation. We know thatfor any flnite experiment the average of the outcomes is not predictable. However,we shall eventually prove that the average will usually be close to E(X) if we repeat the experiment a large number of times. We flrst need to develop some properties ofthe expected value. Using these properties, and those of the concept of the variance 6.1. EXPECTED VALUE 229 XY HHH 1 HHT 2HTH 3 HTT 2 THH 2 THT 3TTH 2 TTT 1 Table 6.2: Tossing a coin three times. to be introduced in the next section, we shall be able to prove the L a wo fL a r g e Numbers. This theorem will justify mathematically both our frequency concept of probability and the interpretation of expected value as the average value to beexpected in a large number of experiments. Expectation of a Function of a Random Variable Suppose that Xis a discrete random variable with sample space ›, and `(x)i s a real-valued function with domain ›. Then `(X) is a real-valued random vari- able. One way to determine the expected value of `(X) is to flrst determine the distribution function of this random variable, and then use the deflnition of expec-tation. However, there is a better way to compute the expected value of `(X), as demonstrated in the next example. Example 6.6 Suppose a coin is tossed 9 times, with the result HHHTTTTHT : The flrst set of three heads is called a run. There are three more runs in this sequence, namely the next four tails, the next head, and the next tail. We do notconsider the flrst two tosses to constitute a run, since the third toss has the samevalue as the flrst two. Now suppose an experiment consists of tossing a fair coin three times. Find the expected number of runs. It will be helpful to think of two random variables, X andY, associated with this experiment. We let Xdenote the sequence of heads and tails that results when the experiment is performed, and Ydenote the number of runs in the outcome X. The possible outcomes of Xand the corresponding values ofYare shown in Table 6.2. To calculate E(Y) using the deflnition of expectation, we flrst must flnd the distribution function m(y)o fYi.e., we group together those values of Xwith a common value of Yand add their probabilities. In this case, we calculate that the distribution function of Yis:m( 1 )=1=4;m( 2 )=1=2;andm( 3 )=1=4. One easily flnds thatE(Y)=2 . 230 CHAPTER 6. EXPECTED VALUE AND VARIANCE Now suppose we didn’t group the values of Xwith a common Y-value, but instead, for each X-valuex, we multiply the probability of xand the corresponding value ofY, and add the results. We obtain 1µ1 8¶ +2µ1 8¶ +3µ1 8¶ +2µ1 8¶ +2µ1 8¶ +3µ1 8¶ +2µ1 8¶ +1µ1 8¶ ; which equals 2. This illustrates the following general principle. If XandYare two random variables, and Ycan be written as a function of X, then one can compute the expected value of Yusing the distribution function of X. 2 Theorem 6.1 IfXis a discrete random variable with sample space › and distri- bution function m(x), and if`:›!R is a function, then E(`(X)) =X x2›`(x)m(x); provided the series converges absolutely. 2 The proof of this theorem is straightforward, involving nothing more than group- ing values of Xwith a common Y-value, as in Example 6.6. The Sum of Two Random Variables Many important results in probability theory concern sums of random variables. We flrst consider what it means to add two random variables. Example 6.7 We °ip a coin and let Xhave the value 1 if the coin comes up heads and 0 if the coin comes up tails. Then, we roll a die and let Ydenote the face that comes up. What does X+Ymean, and what is its distribution? This question is easily answered in this case, by considering, as we did in Chapter 4, the jointrandom variable Z=(X;Y ), whose outcomes are ordered pairs of the form ( x;y), where 0•x•1 and 1•y•6. The description of the experiment makes it reasonable to assume that XandYare independent, so the distribution function ofZis uniform, with 1 =12 assigned to each outcome. Now it is an easy matter to flnd the set of outcomes of X+Y, and its distribution function. 2 In Example 6.1, the random variable Xdenoted the number of heads which occur when a fair coin is tossed three times. It is natural to think of Xas the sum of the random variables X 1;X2;X3, whereXiis deflned to be 1 if the ith toss comes up heads, and 0 if the ith toss comes up tails. The expected values of the Xi’s are extremely easy to compute. It turns out that the expected value of Xcan be obtained by simply adding the expected values of the Xi’s. This fact is stated in the following theorem. 6.1. EXPECTED VALUE 231 Theorem 6.2 LetXandYbe random variables with flnite expected values. Then E(X+Y)=E(X)+E(Y); and ifcis any constant, then E(cX)=cE(X): Proof. Let the sample spaces of XandYbe denoted by › Xand ›Y, and suppose that ›X=fx1;x2;:::g and ›Y=fy1;y2;:::g: Then we can consider the random variable X+Yto be the result of applying the function`(x;y)=x+yto the joint random variable ( X;Y ). Then, by Theorem 6.1, we have E(X+Y)=X jX k(xj+yk)P(X=xj;Y=yk) =X jX kxjP(X=xj;Y=yk)+X jX kykP(X=xj;Y=yk) =X jxjP(X=xj)+X kykP(Y=yk): The last equality follows from the fact that X kP(X=xj;Y=yk)=P(X=xj) andX jP(X=xj;Y=yk)=P(Y=yk): Thus, E(X+Y)=E(X)+E(Y): Ifcis any constant, E(cX)=X jcxjP(X=xj) =cX jxjP(X=xj) =cE(X): 2 232 CHAPTER 6. EXPECTED VALUE AND VARIANCE XY abc 3 acb 1 bac 1 bca 0 cab 0 cba 1 Table 6.3: Number of flxed points. It is easy to prove by mathematical induction that the expected value of the sum of any flnite number of random variables is the sum of the expected values of theindividual random variables. It is important to note that mutual independence of the summands was not needed as a hypothesis in the Theorem 6.2 and its generalization. The fact thatexpectations add, whether or not the summands are mutually independent, is some-times referred to as the First Fundamental Mystery of Probability. Example 6.8 LetYbe the number of flxed points in a random permutation of the setfa;b;cg. To flnd the expected value of Y, it is helpful to consider the basic random variable associated with this experiment, namely the random variable X which represents the random permutation. There are six possible outcomes of X, and we assign to each of them the probability 1 =6 see Table 6.3. Then we can calculateE(Y) using Theorem 6.1, as 3‡1 6· +1‡1 6· +1‡1 6· +0‡1 6· +0‡1 6· +1‡1 6· =1: We now give a very quick way to calculate the average number of flxed points in a random permutation of the set f1;2;3;:::;ng. LetZdenote the random permutation. For each i,1•i•n, letXiequal 1 ifZflxesi, and 0 otherwise. So if we letFdenote the number of flxed points in Z, then F=X1+X2+¢¢¢+Xn: Therefore, Theorem 6.2 implies that E(F)=E(X1)+E(X2)+¢¢¢+E(Xn): But it is easy to see that for each i, E(Xi)=1 n; so E(F)=1: This method of calculation of the expected value is frequently very useful. It applies whenever the random variable in question can be written as a sum of simpler randomvariables. We emphasize again that it is not necessary that the summands bemutually independent. 2 6.1. EXPECTED VALUE 233 Bernoulli Trials Theorem 6.3 LetSnbe the number of successes in nBernoulli trials with prob- abilitypfor success on each trial. Then the expected number of successes is np. That is, E(Sn)=np : Proof. LetXjbe a random variable which has the value 1 if the jth outcome is a success and 0 if it is a failure. Then, for each Xj, E(Xj)=0¢(1¡p)+1¢p=p: Since Sn=X1+X2+¢¢¢+Xn; and the expected value of the sum is the sum of the expected values, we have E(Sn)=E(X1)+E(X2)+¢¢¢+E(Xn) =np : 2 Poisson Distribution Recall that the Poisson distribution with parameter ‚was obtained as a limit of binomial distributions with parameters nandp, where it was assumed that np=‚, andn!1 . Since for each n, the corresponding binomial distribution has expected value‚, it is reasonable to guess that the expected value of a Poisson distribution with parameter ‚also has expectation equal to ‚. This is in fact the case, and the reader is invited to show this (see Exercise 21). Independence IfXandYare two random variables, it is not true in general that E(X¢Y)= E(X)E(Y). However, this is true if XandYareindependent. Theorem 6.4 IfXandYare independent random variables, then E(X¢Y)=E(X)E(Y): Proof. Suppose that ›X=fx1;x2;:::g and ›Y=fy1;y2;:::g 234 CHAPTER 6. EXPECTED VALUE AND VARIANCE are the sample spaces of XandY, respectively. Using Theorem 6.1, we have E(X¢Y)=X jX kxjykP(X=xj;Y=yk): But ifXandYare independent, P(X=xj;Y=yk)=P(X=xj)P(Y=yk): Thus, E(X¢Y)=X jX kxjykP(X=xj)P(Y=yk) =0 @X jxjP(X=xj)1 AˆX kykP(Y=yk)! =E(X)E(Y): 2 Example 6.9 A coin is tossed twice. Xi= 1 if theith toss is heads and 0 otherwise. We know that X1andX2are independent. They each have expected value 1/2. ThusE(X1¢X2)=E(X1)E(X2)=( 1=2)(1=2 )=1=4. 2 We next give a simple example to show that the expected values need not mul- tiply if the random variables are not independent. Example 6.10 Consider a single toss of a coin. We deflne the random variable X to be 1 if heads turns up and 0 if tails turns up, and we set Y=1¡X. Then E(X)=E(Y)=1=2. ButX¢Y= 0 for either outcome. Hence, E(X¢Y)=06= E(X)E(Y). 2 We return to our records example of Section 3.1 for another application of the result that the expected value of the sum of random variables is the sum of theexpected values of the individual random variables. Records Example 6.11 We start keeping snowfall records this year and want to flnd the expected number of records that will occur in the next nyears. The flrst year is necessarily a record. The second year will be a record if the snowfall in the secondyear is greater than that in the flrst year. By symmetry, this probability is 1/2.More generally, let X jb e1i ft h e jth year is a record and 0 otherwise. To flnd E(Xj), we need only flnd the probability that the jth year is a record. But the record snowfall for the flrst jyears is equally likely to fall in any one of these years, 6.1. EXPECTED VALUE 235 soE(Xj)=1=j. Therefore, if Snis the total number of records observed in the flrstnyears, E(Sn)=1+1 2+1 3+¢¢¢+1 n: This is the famous divergent harmonic series. It is easy to show that E(Sn)»logn asn!1 . Therefore, in ten years the expected number of records is approximately log 10 = 2:3; the exact value is the sum of the flrst ten terms of the harmonic series which is 2.9. We see that, even for such a small value as n= 10, lognis not a bad approximation. 2 Craps Example 6.12 In the game of craps, the player makes a bet and rolls a pair of dice. If the sum of the numbers is 7 or 11 the player wins, if it is 2, 3, or 12 theplayer loses. If any other number results, say r, thenrbecomes the player’s point and he continues to roll until either ror 7 occurs. If rcomes up flrst he wins, and if 7 comes up flrst he loses. The program Craps simulates playing this game a number of times. We have run the program for 1000 plays in which the player bets 1 dollar each time. The player’s average winnings were ¡:006. The game of craps would seem to be only slightly unfavorable. Let us calculate the expected winnings on a singleplay and see if this is the case. We construct a two-stage tree measure as shown inFigure 6.1. The flrst stage represents the possible sums for his flrst roll. The second stage represents the possible outcomes for the game if it has not ended on the flrst roll. Inthis stage we are representing the possible outcomes of a sequence of rolls requiredto determine the flnal outcome. The branch probabilities for the flrst stage arecomputed in the usual way assuming all 36 possibilites for outcomes for the pair ofdice are equally likely. For the second stage we assume that the game will eventuallyend, and we compute the conditional probabilities for obtaining either the point ora 7. For example, assume that the player’s point is 6. Then the game will end whenone of the eleven pairs, (1 ;5), (2;4), (3;3), (4;2), (5;1), (1;6), (2;5), (3;4), (4;3), (5;2), (6;1), occurs. We assume that each of these possible pairs has the same probability. Then the player wins in the flrst flve cases and loses in the last six.Thus the probability of winning is 5/11 and the probability of losing is 6/11. Fromthe path probabilities, we can flnd the probability that the player wins 1 dollar; itis 244/495. The probability of losing is then 251/495. Thus if Xis his winning for a dollar bet, E(X)=1‡244 495· +(¡1)‡251 495· =¡7 495…¡:0141: 236 CHAPTER 6. EXPECTED VALUE AND VARIANCE W L W L W L W L W L W L (2,3,12) L1098654(7,11) W 1/3 2/3 2/5 3/5 5/11 6/11 5/11 6/11 2/5 3/5 1/3 2/32/9 1/12 1/9 5/36 5/36 1/91/12 1/9 1/36 2/362/45 3/4525/39630/39625/396 30/3962/45 3/451/36 2/36 Figure 6.1: Tree measure for craps. 6.1. EXPECTED VALUE 237 The game is unfavorable, but only slightly. The player’s expected gain in nplays is ¡n(:0141). Ifnis not large, this is a small expected loss for the player. The casino makes a large number of plays and so can afiord a small average gain per play andstill expect a large proflt. 2 Roulette Example 6.13 In Las Vegas, a roulette wheel has 38 slots numbered 0, 00, 1, 2, ..., 3 6 . T h e 0 a n d 0 0 slots are green, and half of the remaining 36 slots are red and half are black. A croupier spins the wheel and throws an ivory ball. If you bet1 dollar on red, you win 1 dollar if the ball stops in a red slot, and otherwise youlose a dollar. We wish to calculate the expected value of your winnings, if you bet1 dollar on red. LetXbe the random variable which denotes your winnings i n a 1 dollar bet on red in Las Vegas roulette. Then the distribution of Xis given by m X=µ¡11 20=38 18=38¶ ; and one can easily calculate (see Exercise 5) that E(X)…¡:0526: We now consider the roulette game in Monte Carlo, and follow the treatment of Sagan.1In the roulette game in Monte Carlo there is only one 0. If you bet 1 franc on red an d a 0 turns up, then, depending upon the casino, one or more of the following options may be ofiered:(a) You get 1/2 of your bet back, and the casino gets the other half of your bet.(b) Your bet is put \in prison," which we will denote by P 1. If red comes up on the next turn, you get your bet back (but you don’t win any money). If black or 0comes up, you lose your bet.(c) Your bet is put in prison P 1, as before. If red comes up on the next turn, you get your bet back, and if black comes up on the next turn, then you lose your bet.If a 0 comes up on the next turn, then your bet is put into double prison, which wewill denote by P 2. If your bet is in double prison, and if red comes up on the next turn, then your bet is moved back to prison P1and the game proceeds as before. If your bet is in double prison, and if black or 0 come up on the next turn, thenyou lose your bet. We refer the reader to Figure 6.2, where a tree for this option isshown. In this flgure, Sis the starting position, Wmeans that you win your bet, Lmeans that you lose your bet, and Emeans that you break even. It is interesting to compare the expected winnings o f a 1 franc bet on red, under each of these three options. We leave the flrst two calculations as an exercise (seeExercise 37). Suppose that you choose to play alternative (c). The calculation forthis case illustrates the way that the early French probabilists worked problems likethis. 1H. Sagan, Markov Chains in Monte Carlo, Math. Mag., vol. 54, no. 1 (1981), pp. 3-10. 238 CHAPTER 6. EXPECTED VALUE AND VARIANCE SW LE L LL L LLE P1 P1 P1 P2P2P2 Figure 6.2: Tree for 2-prison Monte Carlo roulette. Suppose you bet on red, you choose alternative (c), an d a 0 comes up. Your possible future outcomes are shown in the tree diagram in Figure 6.3. Assume thatyour money is in the flrst prison and let xbe the probability that you lose your franc. From the tree diagram we see that x=18 37+1 37P(you lose your franc jyour franc is in P2): Also, P(you lose your franc jyour franc is in P2)=19 37+18 37x: So, we have x=18 37+1 37‡19 37+18 37x· : Solving for x, we obtain x= 685=1351. Thus, starting at S, the probability that you lose your bet equals 18 37+1 37x=25003 49987: To flnd the probability that you win when you bet on red, note that you can only win if red comes up on the flrst turn, and this happens with probability 18/37.Thus your expected winnings are 1¢18 37¡1¢25003 49987=¡687 49987…¡:0137: It is interesting to note that the more romantic option (c) is less favorable than option (a) (see Exercise 37). 6.1. EXPECTED VALUE 239 PWL PPL18/37 18/37 1/3719/37 18/37 11 2 Figure 6.3: Your money is put in prison. If you bet 1 dollar on the number 17, then the distribution function for your winningsXis PX=µ¡13 5 36=37 1=37¶ ; and the expected winnings are ¡1¢36 37+3 5¢1 37=¡1 37…¡:027: Thus, at Monte Carlo difierent bets have difierent expected values. In Las Vegas almost all bets have the same expected value of ¡2=38 =¡:0526 (see Exercises 4 and 5). 2 Conditional Expectation Deflnition 6.2 IfFis any event and Xis a random variable with sample space ›=fx1;x2;:::g, then the conditional expectation given Fis deflned by E(XjF)=X jxjP(X=xjjF): Conditional expectation is used most often in the form provided by the following theorem. 2 Theorem 6.5 LetXbe a random variable with sample space ›. If F1,F2,...,Fr are events such that Fi\Fj=;fori6=jand › =[jFj, then E(X)=X jE(XjFj)P(Fj): 240 CHAPTER 6. EXPECTED VALUE AND VARIANCE Proof. We have X jE(XjFj)P(Fj)=X jX kxkP(X=xkjFj)P(Fj) =X jX kxkP(X=xkandFjoccurs) =X kX jxkP(X=xkandFjoccurs) =X kxkP(X=xk) =E(X): 2 Example 6.14 (Example 6.12 continued) Let Tbe the number of rolls in a single play of craps. We can think of a single play as a two-stage process. The flrst stageconsists of a single roll of a pair of dice. The play is over if this roll is a 2, 3, 7,11, or 12. Otherwise, the player’s point is established, and the second stage begins.This second stage consists of a sequence of rolls which ends when either the player’sp o i n to ra7i s rolled. We record the outcomes of this two-stage experiment using the random variables XandS, whereXdenotes the flrst roll, and Sdenotes the number of rolls in the second stage of the experiment (of course, Sis sometimes equal to 0). Note that T=S+ 1. Then by Theorem 6.5 E(T)= 12X j=2E(TjX=j)P(X=j): Ifj= 7, 11 or 2, 3, 12, then E(TjX=j)=1 . I fj=4;5;6;8;9;or 10, we can use Example 6.4 to calculate the expected value of S. In each of these cases, we continue rolling until we get either a jor a 7. Thus, Sis geometrically distributed with parameter p, which depends upon j.I fj= 4, for example, the value of pis 3=3 6+6=3 6=1=4. Thus, in this case, the expected number of additional rolls is 1=p=4 ,s oE(TjX= 4) = 1 + 4 = 5. Carrying out the corresponding calculations for the other possible values of jand using Theorem 6.5 gives E(T)=1‡12 36· +‡ 1+36 3+6·‡3 36· +‡ 1+36 4+6·‡4 36· +‡ 1+36 5+6·‡5 36· +‡ 1+36 5+6·‡5 36· +‡ 1+36 4+6·‡4 36· +‡ 1+36 3+6·‡3 36· =557 165 …3:375::: : 2 6.1. EXPECTED VALUE 241 Martingales We can extend the notion of fairness to a player playing a sequence of games by using the concept of conditional expectation. Example 6.15 LetS1,S2,...,Snbe Peter’s accumulated fortune in playing heads or tails (see Example 1.4). Then E(SnjSn¡1=a;:::;S 1=r)=1 2(a+1 )+1 2(a¡1) =a: We note that Peter’s expected fortune after the next play is equal to his present fortune. When this occurs, we say the game is fair. A fair game is also called a martingale. If the coin is biased and comes up heads with probability pand tails with probability q=1¡p, then E(SnjSn¡1=a;:::;S 1=r)=p(a+1 )+q(a¡1) =a+p¡q: Thus, ifp<q , this game is unfavorable, and if p>q , it is favorable. 2 If you are in a casino, you will see players adopting elaborate systems of play to try to make unfavorable games favorable. Two such systems, the martingaledoubling system and the more conservative Labouchere system, were described inExercises 1.1.9 and 1.1.10. Unfortunately, such systems cannot change even a fairgame into a favorable game. Even so, it is a favorite pastime of many people to develop systems of play for gambling games and for other games such as the stock market. We close this sectionwith a simple illustration of such a system. Stock Prices Example 6.16 Let us assume that a stock increases or decreases in value each day by 1 dollar, each with probability 1/2. Then we can identify this simplifledmodel with our familiar game of heads or tails. We assume that a buyer, Mr. Ace,adopts the following strategy. He buys the stock on the flrst day at its price V. He then waits until the price of the stock increases by one to V+ 1 and sells. He then continues to watch the stock until its price falls back to V. He buys again and waits until it goes up to V+ 1 and sells. Thus he holds the stock in intervals during which it increases by 1 dollar. In each such interval, he makes a proflt of 1 dollar.However, we assume that he can do this only for a flnite number of trading days.Thus he can lose if, in the last interval that he holds the stock, it does not getb a c ku pt o V+ 1; and this is the only we he can lose. In Figure 6.4 we illustrate a typical history if Mr. Ace must stop in twenty days. Mr. Ace holds the stock underhis system during the days indicated by broken lines. We note that for the historyshown in Figure 6.4, his system nets him a gain of 4 dollars. We have written a program StockSystem to simulate the fortune of Mr. Ace if he uses his sytem over an n-day period. If one runs this program a large number 242 CHAPTER 6. EXPECTED VALUE AND VARIANCE 5 10 15 20 -1-0.50.511.52 Figure 6.4: Mr. Ace’s system. of times, for n= 20, say, one flnds that his expected winnings are very close to 0, but the probability that he is ahead after 20 days is signiflcantly greater than 1/2.For small values of n, the exact distribution of winnings can be calculated. The distribution for the case n= 20 is shown in Figure 6.5. Using this distribution, it is easy to calculate that the expected value of his winnings is exactly 0. Thisis another instance of the fact that a fair game (a martingale) remains fair underquite general systems of play. Although the expected value of his winnings is 0, the probability that Mr. Ace is ahead after 20 days is about .610. Thus, he would be able to tell his friends that hissystem gives him a better chance of being ahead than that of someone who simplybuys the stock and holds it, if our simple random model is correct. There have beena number of studies to determine how random the stock market is. 2 Historical Remarks With the Law of Large Numbers to bolster the frequency interpretation of proba- bility, we flnd it natural to justify the deflnition of expected value in terms of theaverage outcome over a large number of repetitions of the experiment. The conceptof expected value was used before it was formally deflned; and when it was used,it was considered not as an average value but rather as the appropriate value for agamble. For example, recall Pascal’s way of flnding the value of a three-game seriesthat had to be called ofi before it is flnished. Pascal flrst observed that if each player has only one game to win, then the stake of 64 pistoles should be divided evenly. Then he considered the case whereone player has won two games and the other one. Then consider, Sir, if the flrst man wins, he gets 64 pistoles, if he loses he gets 32. Thus if they do not wish to risk this last game, but wishto separate without playing it, the flrst man must say: \I am certain 6.1. EXPECTED VALUE 243 -20 -15 -10 -5 0 5 1000.050.10.150.2 Figure 6.5: Winnings distribution for n= 20. to get 32 pistoles, even if I lose I still get them; but as for the other 32 pistoles, perhaps I will get them, perhaps you will get them, thechances are equal. Let us then divide these 32 pistoles in half and giveone half to me as well as my 32 which are mine for sure." He will thenhave 48 pistoles and the other 16. 2 Note that Pascal reduced the problem to a symmetric bet in which each player gets the same amount and takes it as obvious that in this case the stakes should bedivided equally. The flrst systematic study of expected value appears in Huygens’ book. Like Pascal, Huygens flnd the value of a gamble by assuming that the answer is obviousfor certain symmetric situations and uses this to deduce the expected for the generalsituation. He does this in steps. His flrst proposition is Prop. I. If I expect aorb, either of which, with equal probability, may fall to me, then my Expectation is worth ( a+b)=2, that is, the half Sum ofaandb. 3 Huygens proved this as follows: Assume that two player A and B play a game in which each player puts up a stake of ( a+b)=2 with an equal chance of winning the total stake. Then the value of the game to each player is ( a+b)=2. For example, if the game had to be called ofi clearly each player should just get back his originalstake. Now, by symmetry, this value is not changed if we add the condition thatthe winner of the game has to pay the loser an amount bas a consolation prize. Then for player A the value is still ( a+b)=2. But what are his possible outcomes for the modifled game? If he wins he gets the total stake a+band must pay B an 2Quoted in F. N. David, Games, Gods and Gambling (London: Gri–n, 1962), p. 231. 3C. Huygens, Calculating in Games of Chance, translation attributed to John Arbuthnot (Lon- don, 1692), p. 34. 244 CHAPTER 6. EXPECTED VALUE AND VARIANCE amountbso ends up with a. If he loses he gets an amount bfrom player B. Thus player A wins aorbwith equal chances and the value to him is ( a+b)=2. Huygens illustrated this proof in terms of an example. If you are ofiered a game in which you have an equal chance of winning 2 or 8, the expected value is 5, sincethis game is equivalent to the game in which each player stakes 5 and agrees to paythe lose r3|a game in which the value is obviously 5. Huygens’ second proposition is Prop. II. If I expect a,b,o rc, either of which, with equal facility, may happen, then the Value of my Expectation is ( a+b+c)=3, or the third of the Sum of a,b, andc. 4 His argument here is similar. Three players, A, B, and C, each stake (a+b+c)=3 in a game they have an equal chance of winning. The value of this game to player A is clearly the amount he has staked. Further, this value is not changed if A entersinto an agreement with B that if one of them wins he pays the other a consolationprize ofband with C that if one of them wins he pays the other a consolation prize ofc. By symmetry these agreements do not change the value of the game. In this modifled game, if A wins he wins the total stake a+b+cminus the consolation prizesb+cgiving him a flnal winning of a. If B wins, A wins band if C wins, A winsc. Thus A flnds himself in a game with value ( a+b+c)=3 and with outcomes a,b, andcoccurring with equal chance. This proves Proposition II. More generally, this reasoning shows that if there are noutcomes a 1;a2; :::; an; all occurring with the same probability, the expected value is a1+a2+¢¢¢+an n: In his third proposition Huygens considered the case where you win aorbbut with unequal probabilities. He assumed there are pchances of winning a, andq chances of winning b, all having the same probability. He then showed that the expected value is E=p p+q¢a+q p+q¢b: This follows by considering an equivalent gamble with p+qoutcomes all occurring with the same probability and with a payofi of ainpof the outcomes and binqof the outcomes. This allowed Huygens to compute the expected value for experimentswith unequal probabilities, at least when these probablities are rational numbers. Thus, instead of deflning the expected value as a weighted average, Huygens assumed that the expected value of certain symmetric gambles are known and de-duced the other values from these. Although this requires a good deal of clever 4ibid., p. 35. 6.1. EXPECTED VALUE 245 manipulation, Huygens ended up with values that agree with those given by our modern deflnition of expected value. One advantage of this method is that it givesa justiflcation for the expected value in cases where it is not reasonable to assumethat you can repeat the experiment a large number of times, as for example, inbetting that at least two presidents died on the same day of the year. (In fact,three did; all were signers of the Declaration of Independence, and all three died onJuly 4.) In his book, Huygens calculated the expected value of games using techniques similar to those which we used in computing the expected value for roulette atMonte Carlo. For example, his proposition XIV is: Prop. XIV. If I were playing with another by turns, with two Dice, on this Condition, that if I thro w 7 I gain, and if he throws 6 he gains allowing him the flrst Throw: To flnd the proportion of my Hazard tohis. 5 A modern description of this game is as follows. Huygens and his opponent take turns rolling a die. The game is over if Huygens roll sa7o rh i so p p onent rolls a 6. His opponent rolls flrst. What is the probability that Huygens wins the game? To solve this problem Huygens let xbe his chance of winning when his opponent threw flrst and yhis chance of winning when he threw flrst. Then on the flrst roll his opponent wins on 5 out of the 36 possibilities. Thus, x=31 36¢y: But when Huygens rolls he wins on 6 out of the 36 possible outcomes, and in the other 30, he is led back to where his chances are x.T h u s y=6 36+30 36¢x: From these two equations Huygens found that x=3 1=61. Another early use of expected value appeared in Pascal’s argument to show that a rational person should believe in the existence of God.6Pascal said that we have to make a wager whether to believe or not to believe. Let pdenote the probability that God does not exist. His discussion suggests that we are playing a game withtwo strategies, believe and not believe, with payofis as shown in Table 6.4. Here¡urepresents the cost to you of passing up some worldly pleasures as a consequence of believing that God exists. If you do not believe, and God is avengeful God, you will lose x. If God exists and you do believe you will gain v. Now to determine which strategy is best you should compare the two expectedvalues p(¡u)+( 1¡p)v andp0+( 1¡p)(¡x); 5ibid., p. 47. 6Quoted in I. Hacking, The Emergence of Probability (Cambridge: Cambridge Univ. Press, 1975). 246 CHAPTER 6. EXPECTED VALUE AND VARIANCE God does not exist God exists p 1¡p believe ¡u v not believe 0 ¡x Table 6.4: Payofis. Age Survivors 0 100 66 4 16 4026 2536 1646 1056 666 376 1 Table 6.5: Graunt’s mortality data. and choose the larger of the two. In general, the choice will depend upon the value of p. But Pascal assumed that the value of vis inflnite and so the strategy of believing is best no matter what probability you assign for the existence of God. This exampleis considered by some to be the beginning of decision theory. Decision analyses ofthis kind appear today in many flelds, and, in particular, are an important part ofmedical diagnostics and corporate business decisions. Another early use of expected value was to decide the price of annuities. The study of statistics has its origins in the use of the bills of mortality kept in theparishes in London from 1603. These records kept a weekly tally of christeningsand burials. From these John Graunt made estimates for the population of Londonand also provided the flrst mortality data, 7shown in Table 6.5. As Hacking observes, Graunt apparently constructed this table by assuming that after the age of 6 there is a constant probability of about 5/8 of survivingfor another decade. 8For example, of the 64 people who survive to age 6, 5/8 of 64 or 40 survive to 16, 5/8 of these 40 or 25 survive to 26, and so forth. Of course,he rounded ofi his flgures to the nearest whole person. Clearly, a constant mortality rate cannot be correct throughout the whole range, and later tables provided by Halley were more realistic in this respect. 9 7ibid., p. 108. 8ibid., p. 109. 9E. Halley, \An Estimate of The Degrees of Mortality of Mankind," Phil. Trans. Royal. Soc., 6.1. EXPECTED VALUE 247 Aterminal annuity provides a flxed amount of money during a period of nyears. To determine the price of a terminal annuity one needs only to know the appropriateinterest rate. A life annuity provides a flxed amount during each year of the buyer’s life. The appropriate price for a life annuity is the expected value of the terminalannuity evaluated for the random lifetime of the buyer. Thus, the work of Huygensin introducing expected value and the work of Graunt and Halley in determiningmortality tables led to a more rational method for pricing annuities. This was oneof the flrst serious uses of probability theory outside the gambling houses. Although expected value plays a role now in every branch of science, it retains its importance in the casino. In 1962, Edward Thorp’s book Beat the Dealer 10 provided the reader with a strategy for playing the popular casino game of blackjack that would assure the player a positive expected winning. This book forevermorechanged the belief of the casinos that they could not be beat. Exercises 1A card is drawn at random from a deck consisting of cards numbered 2 through 10. A player wins 1 dollar if the number on the card is odd andloses 1 dollar if the number if even. What is the expected value of his win-nings? 2A card is drawn at random from a deck of playing cards. If it is red, the player wins 1 dollar; if it is black, the player loses 2 dollars. Find the expected valueof the game. 3In a class there are 20 students: 3 are 5’ 6", 5 are 5’8", 4 are 5’10", 4 are 6’, and 4 are 6’ 2". A student is chosen at random. What is the student’sexpected height? 4In Las Vegas the roulette wheel ha s a 0 and a 00 and then the numbers 1 to 36 marked on equal slots; the wheel is spun and a ball stops randomly in oneslot. When a player bets 1 dollar on a number, he receives 36 dollars if theball stops on this number, for a net gain of 35 dollars; otherwise, he loses hisdollar bet. Find the expected value for his winnings. 5In a second version of roulette in Las Vegas, a player bets on red or black. Half of the numbers from 1 to 36 are red, and half are black. If a player betsa dollar on black, and if the ball stops on a black number, he gets his dollarback and another dollar. If the ball stops on a red number or on 0 or 00 heloses his dollar. Find the expected winnings for this bet. 6A die is rolled twice. Let Xdenote the sum of the two numbers that turn up, andYthe difierence of the numbers (speciflcally, the number on the flrst roll minus the number on the second). Show that E(XY)=E(X)E(Y). AreX andYindependent? vol. 17 (1693), pp. 596{610; 654{656. 10E. Thorp, Beat the Dealer (New York: Random House, 1962). 248 CHAPTER 6. EXPECTED VALUE AND VARIANCE *7Show that, if XandYare random variables taking on only two values each, and ifE(XY)=E(X)E(Y), thenXandYare independent. 8A royal family has children until it has a boy or until it has three children, whichever comes flrst. Assume that each child is a boy with probability 1/2.Find the expected number of boys in this royal family and the expected num-ber of girls. 9If the flrst roll in a game of craps is neither a natural nor craps, the player can make an additional bet, equal to his original one, that he will make hispoint before a seven turns up. If his point is four or ten he is paid ofi at 2 : 1odds; if it is a flve or nine he is paid ofi at odds 3 : 2; and if it is a six or eighthe is paid ofi at odds 6 : 5. Find the player’s expected winnings if he makesthis additional bet when he has the opportunity. 10In Example 6.16 assume that Mr. Ace decides to buy the stock and hold it until it goes up 1 dollar and then sell and not buy again. Modify the programStockSystem to flnd the distribution of his proflt under this system after a twenty-day period. Find the expected proflt and the probability that hecomes out ahead. 11On September 26, 1980, the New York Times reported that a mysterious stranger strode into a Las Vegas casino, placed a single bet of 777,000 dollarson the \don’t pass" line at the crap table, and walked away with more than 1.5 million dollars. In the \don’t pass" bet, the bettor is essentially bettingwith the house. An exception occurs if the roller rolls a 12 on the flrst roll.In this case, the roller loses and the \don’t pass" better just gets back themoney bet instead of winning. Show that the \don’t pass" bettor has a morefavorable bet than the roller. 12Recall that in the martingale doubling system (see Exercise 1.1.10), the player doubles his bet each time he loses and quits the flrst time he is ahead. Supposethat you are playing roulette in a fair casino where there are no 0’s, and you bet on red each time. You then win with probability 1/2 each time. Assumethat you start with a 1-dollar bet and employ the martingale system. Sinceyou entered the casino with 100 dollars, you also quit in the unlikely eventthat black turns up six times in a row so that you are down 63 dollars andcannot make the required 64-dollar bet. Find your expected winnings underthis system of play. 13You have 80 dollars and play the following game. An urn contains two white balls and two black balls. You draw the balls out one at a time withoutreplacement until all the balls are gone. On each draw, you bet half of yourpresent fortune that you will draw a white ball. What is your flnal fortune? 14In the hat check problem (see Example 3.12), it was assumed that Npeople check their hats and the hats are handed back at random. Let X j= 1 if the 6.1. EXPECTED VALUE 249 jth person gets his or her hat and 0 otherwise. Find E(Xj) andE(Xj¢Xk) forjnot equal to k. AreXjandXkindependent? 15A box contains two gold balls and three silver balls. You are allowed to choose successively balls from the box at random. You win 1 dollar each time youdraw a gold ball and lose 1 dollar each time you draw a silver ball. After adraw, the ball is not replaced. Show that, if you draw until you are ahead by1 dollar or until there are no more gold balls, this is a favorable game. 16Gerolamo Cardano in his book, The Gambling Scholar, written in the early 1500s, considers the following carnival game. There are six dice. Each of thedice has flve blank sides. The sixth side has a number between 1 and 6|adifierent number on each die. The six dice are rolled and the player wins aprize depending on the total of the numbers which turn up. (a) Find, as Cardano did, the expected total without flnding its distribution. (b) Large prizes were given for large totals with a modest fee to play the game. Explain why this could be done. 17LetXbe the flrst time that a failure occurs in an inflnite sequence of Bernoulli trials with probability pfor success. Let p k=P(X=k) fork= 1 , 2 , .... Show that pk=pk¡1qwhereq=1¡p. Show thatP kpk= 1. Show that E(X)=1=q. What is the expected number of tosses of a coin required to obtain the flrst tail? 18Exactly one of six similar keys opens a certain door. If you try the keys, one after another, what is the expected number of keys that you will have to trybefore success? 19A multiple choice exam is given. A problem has four possible answers, and exactly one answer is correct. The student is allowed to choose a subset ofthe four possible answers as his answer. If his chosen subset contains thecorrect answer, the student receives three points, but he loses one point foreach wrong answer in his chosen subset. Show that if he just guesses a subsetuniformly and randomly his expected score is zero. 20You are ofiered the following game to play: a fair coin is tossed until heads turns up for the flrst time (see Example 6.3). If this occurs on the flrst tossyou receive 2 dollars, if it occurs on the second toss you receive 2 2= 4 dollars and, in general, if heads turns up for the flrst time on the nth toss you receive 2ndollars. (a) Show that the expected value of your winnings does not exist (i.e., is given by a divergent sum) for this game. Does this mean that this gameis favorable no matter how much you pay to play it? (b) Assume that you only receive 2 10dollars if any number greater than or equal to ten tosses are required to obtain the flrst head. Show that yourexpected value for this modifled game is flnite and flnd its value. 250 CHAPTER 6. EXPECTED VALUE AND VARIANCE (c) Assume that you pay 10 dollars for each play of the original game. Write a program to simulate 100 plays of the game and see how you do. (d) Now assume that the utility of ndollars ispn. Write an expression for the expected utility of the payment, and show that this expression has aflnite value. Estimate this value. Repeat this exercise for the case thatthe utility function is log( n). 21LetXbe a random variable which is Poisson distributed with parameter ‚. Show thatE(X)=‚.Hint: Recall that e x=1+x+x2 2!+x3 3!+¢¢¢: 22Recall that in Exercise 1.1.14, we considered a town with two hospitals. In the large hospital about 45 babies are born each day, and in the smallerhospital about 15 babies are born each day. We were interested in guessingwhich hospital would have on the average the largest number of days withthe property that more than 60 percent of the children born on that day areboys. For each hospital flnd the expected number of days in a year that havethe property that more than 60 percent of the children born on that day wereboys. 23An insurance company has 1,000 policies on men of age 50. The company estimates that the probability that a man of age 50 dies within a year is .01.Estimate the number of claims that the company can expect from beneflciariesof these men within a year. 24Using the life table for 1981 in Appendix C, write a program to compute the expected lifetime for males and females of each possible age from 1 to 85.Compare the results for males and females. Comment on whether life insur-ance should be priced difierently for males and females. *25 A deck of ESP cards consists of 20 cards each of two types: say ten stars, ten circles (normally there are flve types). The deck is shu†ed and the cardsturned up one at a time. You, the alleged percipient, are to name the symbolon each card before it is turned up. Suppose that you are really just guessing at the cards. If you do not get to see each card after you have made your guess, then it is easy to calculate theexpected number of correct guesses, namely ten. If, on the other hand, you are guessing with information, that is, if you see each card after your guess, then, of course, you might expect to get a higherscore. This is indeed the case, but calculating the correct expectation is nolonger easy. But it is easy to do a computer simulation of this guessing with information, so we can get a good idea of the expectation by simulation. (This is similar tothe way that skilled blackjack players make blackjack into a favorable gameby observing the cards that have already been played. See Exercise 29.) 6.1. EXPECTED VALUE 251 (a) First, do a simulation of guessing without information, repeating the experiment at least 1000 times. Estimate the expected number of correctanswers and compare your result with the theoretical expectation. (b) What is the best strategy for guessing with information? (c) Do a simulation of guessing with information, using the strategy in (b). Repeat the experiment at least 1000 times, and estimate the expectationin this case. (d) LetSbe the number of stars and Cthe number of circles in the deck. Let h(S;C) be the expected winnings using the optimal guessing strategy in (b). Show that h(S;C) satisfles the recursion relation h(S;C)=S S+Ch(S¡1;C)+C S+Ch(S;C¡1) +max(S;C) S+C; andh(0;0) =h(¡1;0) =h(0;¡1) = 0. Using this relation, write a program to compute h(S;C) and flndh(10;10). Compare the computed value ofh(10;10) with the result of your simulation in (c). For more about this exercise and Exercise 26 see Diaconis and Graham.11 *26 Consider the ESP problem as described in Exercise 25. You are again guessing with information, and you are using the optimal guessing strategy of guessingstar if the remaining deck has more stars, circle if more circles, and tossing a coin if the number of stars and circles are equal. Assume that S‚C, where Sis the number of stars and Cthe number of circles. We can plot the results of a typical game on a graph, where the horizontal axis represents the number of steps and the vertical axis represents the difierence between the number of stars and the number of circles that have been turnedup. A typical game is shown in Figure 6.6. In this particular game, the orderin which the cards were turned up is ( C;S;S;S;S;C;C;S;S;C ). Thus, in this particular game, there were six stars and four circles in the deck. This means,in particular, that every game played with this deck would have a graph whichends at the point (10 ;2). We deflne the line Lto be the horizontal line which goes through the ending point on the graph (so its vertical coordinate is justthe difierence between the number of stars and circles in the deck). (a) Show that, when the random walk is below the line L, the player guesses right when the graph goes up (star is turned up) and, when the walk isabove the line, the player guesses right when the walk goes down (circleturned up). Show from this property that the subject is sure to have atleastScorrect guesses. (b) When the walk is at a point ( x;x)onthe lineLthe number of stars and circles remaining is the same, and so the subject tosses a coin. Show that 11P. Diaconis and R. Graham, \The Analysis of Sequential Experiments with Feedback to Sub- jects," Annals of Statistics, vol. 9 (1981), pp. 3{23. 252 CHAPTER 6. EXPECTED VALUE AND VARIANCE 2 1 12345 678 9 10(10,2)L Figure 6.6: Random walk for ESP. the probability that the walk reaches ( x;x)i s ¡S x¢¡C x¢ ¡S+C 2x¢: Hint: The outcomes of 2 xcards is a hypergeometric distribution (see Section 5.1). (c) Using the results of (a) and (b) show that the expected number of correct guesses under intelligent guessing is S+CX x=11 2¡S x¢¡C x¢ ¡S+C 2x¢: 27It has been said12that a Dr. B. Muriel Bristol declined a cup of tea stating that she preferred a cup into which milk had been poured flrst. The famousstatistician R. A. Fisher carried out a test to see if she could tell whether milkwas put in before or after the tea. Assume that for the test Dr. Bristol wasgiven eight cups of tea|four in which the milk was put in before the tea andfour in which the milk was put in after the tea. (a) What is the expected number of correct guesses the lady would make if she had no information after each test and was just guessing? (b) Using the result of Exercise 26 flnd the expected number of correct guesses if she was told the result of each guess and used an optimalguessing strategy. 28In a popular computer game the computer picks an integer from 1 to nat random. The player is given kchances to guess the number. After each guess the computer responds \correct," \too small," or \too big." 12J. F. Box, R. A. Fisher, The Life of a Scientist (New York: John Wiley and Sons, 1978). 6.1. EXPECTED VALUE 253 (a) Show that if n•2k¡1, then there is a strategy that guarantees you will correctly guess the number in ktries. (b) Show that if n‚2k¡1, there is a strategy that assures you of identifying one of 2k¡1 numbers and hence gives a probability of (2k¡1)=nof winning. Why is this an optimal strategy? Illustrate your result interms of the case n= 9 andk=3 . 29In the casino game of blackjack the dealer is dealt two cards, one face up and one face down, and each player is dealt two cards, both face down. If thedealer is showing an ace the player can look at his down cards and then makea bet called an insurance bet. (Expert players will recognize why it is called insurance.) If you make this bet you will win the bet if the dealer’s secondcard is a ten card : namely, a ten, jack, queen, or king. If you win, you are paid twice your insurance bet; otherwise you lose this bet. Show that, if theonly cards you can see are the dealer’s ace and your two cards and if yourcards are not ten cards, then the insurance bet is an unfavorable bet. Show,however, that if you are playing two hands simultaneously, and you have noten cards, then it is a favorable bet. (Thorp 13has shown that the game of blackjack is favorable to the player if he or she can keep good enough trackof the cards that have been played.) 30Assume that, every time you buy a box of Wheaties, you receive a picture of one of thenplayers for the New York Yankees (see Exercise 3.2.34). Let X k be the number of additional boxes you have to buy, after you have obtained k¡1 difierent pictures, in order to obtain the next new picture. Thus X1=1 , X2is the number of boxes bought after this to obtain a picture difierent from the flrst pictured obtained, and so forth. (a) Show that Xkhas a geometric distribution with p=(n¡k+1 )=n. (b) Simulate the experiment for a team with 26 players (25 would be more accurate but we want an even number). Carry out a number of simula-tions and estimate the expected time required to get the flrst 13 playersand the expected time to get the second 13. How do these expectationscompare? (c) Show that, if there are 2 nplayers, the expected time to get the flrst half of the players is 2nµ1 2n+1 2n¡1+¢¢¢+1 n+1¶ ; and the expected time to get the second half is 2nµ1 n+1 n¡1+¢¢¢+1¶ : 13E. Thorp, Beat the Dealer (New York: Random House, 1962). 254 CHAPTER 6. EXPECTED VALUE AND VARIANCE (d) In Section 3.1 we showed that 1+1 2+1 3+¢¢¢+1 n»logn+:5772 +1 2n: Use this to estimate the expression in (c). Compare these estimates with the exact values and also with your estimates obtained by simulation forthe casen= 26. *31 (Feller 14) A large number, N, of people are subjected to a blood test. This can be administered in two ways: (1) Each person can be tested separately,in this case Ntest are required, (2) the blood samples of kpersons can be pooled and analyzed together. If this test is negative, this one test su–ces for thekpeople. If the test is positive, each of the kpersons must be tested separately, and in all, k+ 1 tests are required for the kpeople. Assume that the probability pthat a test is positive is the same for all people and that these events are independent. (a) Find the probability that the test for a pooled sample of kpeople will be positive. (b) What is the expected value of the number Xof tests necessary under plan (2)? (Assume that Nis divisible by k.) (c) For small p, show that the value of kwhich will minimize the expected number of tests under the second plan is approximately 1 =p p. 32Write a program to add random numbers chosen from [0 ;1] until the flrst time the sum is greater than one. Have your program repeat this experimenta number of times to estimate the expected number of selections necessaryin order that the sum of the chosen numbers flrst exceeds 1. On the basis ofyour experiments, what is your estimate for this number? *33 The following related discrete problem also gives a good clue for the answer to Exercise 32. Randomly select with replacement t 1,t2, ...,trfrom the set (1=n;2=n;:::;n=n ). LetXbe the smallest value of rsatisfying t1+t2+¢¢¢+tr>1: ThenE(X)=( 1+1=n)n. To prove this, we can just as well choose t1,t2, ...,trrandomly with replacement from the set (1 ;2;:::;n ) and letXbe the smallest value of rfor which t1+t2+¢¢¢+tr>n: (a) Use Exercise 3.2.36 to show that P(X‚j+1 )=µn j¶‡1 n·j : 14W. Feller, Introduction to Probability Theory and Its Applications, 3rd ed., vol. 1 (New York: John Wiley and Sons, 1968), p. 240. 6.1. EXPECTED VALUE 255 (b) Show that E(X)=nX j=0P(X‚j+1 ): (c) From these two facts, flnd an expression for E(X). This proof is due to Harris Schultz.15 *34 (Banach’s Matchbox16) A man carries in each of his two front pockets a box of matches originally containing Nmatches. Whenever he needs a match, he chooses a pocket at random and removes one from that box. One day hereaches into a pocket and flnds the box empty. (a) Letp rdenote the probability that the other pocket contains rmatches. Deflne a sequence of counter random variables as follows: Let Xi=1i f theith draw is from the left pocket, and 0 if it is from the right pocket. Interpretprin terms of Sn=X1+X2+¢¢¢+Xn. Find a binomial expression for pr. (b) Write a computer program to compute the pr, as well as the probability that the other pocket contains at least rmatches, for N= 100 and r from 0 to 50. (c) Show that ( N¡r)pr=( 1=2)(2N+1 )pr+1¡(1=2)(r+1 )pr+1. (d) EvaluateP rpr. (e) Use (c) and (d) to determine the expectation Eof the distribution fprg. (f) Use Stirling’s formula to obtain an approximation for E. How many matches must each box contain to ensure a value of about 13 for theexpectation E?( T a k e…=2 2=7.) 35A coin is tossed until the flrst time a head turns up. If this occurs on the nth toss andnis odd you win 2 n=n, but ifnis even then you lose 2n=n. Then if your expected winnings exist they are given by the convergent series 1¡1 2+1 3¡1 4+¢¢¢ called the alternating harmonic series. It is tempting to say that this should be the expected value of the experiment. Show that if we were to do this, theexpected value of an experiment would depend upon the order in which theoutcomes are listed. 36Suppose we have an urn containing cyellow balls and dgreen balls. We draw kballs, without replacement, from the urn. Find the expected number of yellow balls drawn. Hint: Write the number of yellow balls drawn as the sum ofcrandom variables. 15H. Schultz, \An Expected Value Problem," Two-Year Mathematics Journal, vol. 10, no. 4 (1979), pp. 277{78. 16W. Feller, Introduction to Probability Theory, vol. 1, p. 166. 256 CHAPTER 6. EXPECTED VALUE AND VARIANCE 37The reader is referred to Example 6.13 for an explanation of the various op- tions available in Monte Carlo roulette. (a) Compute the expected winnings o f a 1 franc bet on red under option (a). (b) Repeat part (a) for option (b). (c) Compare the expected winnings for all three options. *38 (from Pittel17) Telephone books, nin number, are kept in a stack. The probability that the book numbered i(where 1•i•n) is consulted for a given phone call is pi>0, where the pi’s sum to 1. After a book is used, it is placed at the top of the stack. Assume that the calls are independentand evenly spaced, and that the system has been employed indeflnitely farinto the past. Let d ibe the average depth of book iin the stack. Show that di•djwheneverpi‚pj. Thus, on the average, the more popular books have a tendency to be closer to the top of the stack. Hint: Letpijdenote the probability that book iis above book j. Show that pij=pij(1¡pj)+pjipi. *39 (from Propp18) In the previous problem, let Pbe the probability that at the present time, each book is in its proper place, i.e., book iisith from the top. Find a formula for Pin terms of the pi’s. In addition, flnd the least upper bound onP, if thepi’s are allowed to vary. Hint: First flnd the probability that book 1 is in the right place. Then flnd the probability that book 2 is inthe right place, given that book 1 is in the right place. Continue. *40 (from H. Shultz and B. Leonard 19) A sequence of random numbers in [0 ;1) is generated until the sequence is no longer monotone increasing. The num-bers are chosen according to the uniform distribution. What is the expectedlength of the sequence? (In calculating the length, the term that destroysmonotonicity is included.) Hint: Leta 1;a2; ::: be the sequence and let X denote the length of the sequence. Then P(X>k )=P(a1<a2<¢¢¢<ak); and the probability on the right-hand side is easy to calculate. Furthermore, one can show that E(X)=1+P(X> 1) +P(X> 2) +¢¢¢: 41LetTbe the random variable that counts the number of 2-unshu†es per- formed on an n-card deck until all of the labels on the cards are distinct. This random variable was discussed in Section 3.3. Using Equation 3.4 in thatsection, together with the formula E(T)=1X s=0P(T>s ) 17B. Pittel, Problem #1195, Mathematics Magazine, vol. 58, no. 3 (May 1985), pg. 183. 18J. Propp, Problem #1159, Mathematics Magazine vol. 57, no. 1 (Feb. 1984), pg. 50. 19H. Shultz and B. Leonard, \Unexpected Occurrences of the Number e,"Mathematics Magazine vol. 62, no. 4 (October, 1989), pp. 269-271. 6.2. VARIANCE OF DISCRETE RANDOM VARIABLES 257 that was proved in Exercise 33, show that E(T)=1X s=0µ 1¡µ2s n¶n! 2sn¶ : Show that for n= 52, this expression is approximately equal to 11.7. (As was stated in Chapter 3, this means that on the average, almost 12 ri†e shu†es ofa 52-card deck are required in order for the process to be considered random.) 6.2 Variance of Discrete Random Variables The usefulness of the expected value as a prediction for the outcome of an ex-periment is increased when the outcome is not likely to deviate too much from theexpected value. In this section we shall introduce a measure of this deviation, calledthe variance. Variance Deflnition 6.3 LetXbe a numerically valued random variable with expected value „=E(X). Then the variance ofX, denoted by V(X), is V(X)=E((X¡„)2): 2 Note that, by Theorem 6.1, V(X) is given by V(X)=X x(x¡„)2m(x); (6.1) wheremis the distribution function of X. Standard Deviation The standard deviation ofX, denoted by D(X), isD(X)=p V(X). We often write¾forD(X) and¾2forV(X). Example 6.17 Consider one roll of a die. Let Xbe the number that turns up. To flndV(X), we must flrst flnd the expected value of X. This is „=E(X)=1‡1 6· +2‡1 6· +3‡1 6· +4‡1 6· +5‡1 6· +6‡1 6· =7 2: To flnd the variance of X, we form the new random variable ( X¡„)2and compute its expectation. We can easily do this using the following table. 258 CHAPTER 6. EXPECTED VALUE AND VARIANCE xm (x)(x¡7=2)2 1 1/6 25/4 2 1/6 9/43 1/6 1/44 1/6 1/45 1/6 9/46 1/6 25/4 Table 6.6: Variance calculation. From this table we flnd E((X¡„)2)i s V(X)=1 6µ25 4+9 4+1 4+1 4+9 4+25 4¶ =35 12; and the standard deviation D(X)=p 35=12…1:707. 2 Calculation of Variance We next prove a theorem that gives us a useful alternative form for computing the variance. Theorem 6.6 IfXis any random variable with E(X)=„, then V(X)=E(X2)¡„2: Proof. We have V(X)=E((X¡„)2)=E(X2¡2„X+„2) =E(X2)¡2„E(X)+„2=E(X2)¡„2: 2 Using Theorem 6.6, we can compute the variance of the outcome of a roll of a die by flrst computing E(X2)=1‡1 6· +4‡1 6· +9‡1 6· +1 6‡1 6· +2 5‡1 6· +3 6‡1 6· =91 6; and, V(X)=E(X2)¡„2=91 6¡‡7 2·2 =35 12; in agreement with the value obtained directly from the deflnition of V(X). 6.2. VARIANCE OF DISCRETE RANDOM VARIABLES 259 Properties of Variance The variance has properties very difierent from those of the expectation. If cis any constant,E(cX)=cE(X) andE(X+c)=E(X)+c. These two statements imply that the expectation is a linear function. However, the variance is not linear, asseen in the next theorem. Theorem 6.7 IfXis any random variable and cis any constant, then V(cX)=c 2V(X) and V(X+c)=V(X): Proof. Let„=E(X). ThenE(cX)=c„, and V(cX)=E((cX¡c„)2)=E(c2(X¡„)2) =c2E((X¡„)2)=c2V(X): To prove the second assertion, we note that, to compute V(X+c), we would replace!by!+cand„by„+cin Equation 6.1. But the c’s would cancel, leaving V(X). 2 We turn now to some general properties of the variance. Recall that if XandY are any two random variables, E(X+Y)=E(X)+E(Y). This is not always true for the case of the variance. For example, let Xbe a random variable with V(X)6=0 , and deflne Y=¡X. ThenV(X)=V(Y), so thatV(X)+V(Y)=2V(X). But X+Yis always 0 and hence has variance 0. Thus V(X+Y)6=V(X)+V(Y). In the important case of mutually independent random variables, however, the variance of the sum is the sum of the variances. Theorem 6.8 LetXandYbe two independent random variables. Then V(X+Y)=V(X)+V(Y): Proof. LetE(X)=aandE(Y)=b. Then V(X+Y)=E((X+Y)2)¡(a+b)2 =E(X2)+2E(XY)+E(Y2)¡a2¡2ab¡b2: SinceXandYare independent, E(XY)=E(X)E(Y)=ab. Thus, V(X+Y)=E(X2)¡a2+E(Y2)¡b2=V(X)+V(Y): 2 260 CHAPTER 6. EXPECTED VALUE AND VARIANCE It is easy to extend this proof, by mathematical induction, to show that the variance of the sum of any number of mutually independent random variables is thesum of the individual variances. Thus we have the following theorem. Theorem 6.9 LetX 1,X2,...,Xnbe an independent trials process with E(Xj)= „andV(Xj)=¾2. Let Sn=X1+X2+¢¢¢+Xn be the sum, and An=Sn n be the average. Then E(Sn)=n„ ; V(Sn)=n¾2; E(An)=„; V(An)=¾2 n: Proof. Since all the random variables Xjhave the same expected value, we have E(Sn)=E(X1)+¢¢¢+E(Xn)=n„ ; and V(Sn)=V(X1)+¢¢¢+V(Xn)=n¾2: We have seen that, if we multiply a random variable Xwith mean„and variance ¾2by a constant c, the new random variable has expected value c„and variance c2¾2. Thus, E(An)=EµSn n¶ =n„ n=„; and V(An)=VµSn n¶ =V(Sn) n2=n¾2 n2=¾2 n: Finally, the standard deviation of Anis given by ¾(An)=¾pn: 2 The last statement in the above proof implies that in an independent trials pro- cess, if the individual summands have flnite variance, then the standard deviationof the average goes to 0 as n!1 . Since the standard deviation tells us something about the spread of the distribution around the mean, we see that for large values ofn, the value of A nis usually very close to the mean of An, which equals „, as shown above. This statement is made precise in Chapter 8, where it is called the Law ofLarge Numbers. For example, let Xrepresent the roll of a fair die. In Figure 6.7, we show the distribution of a random variable A ncorresponding to X, forn=1 0 andn= 100. 6.2. VARIANCE OF DISCRETE RANDOM VARIABLES 261 1 2 3 4 5 600.10.20.30.40.50.6 2 2.5 3 3.5 4 4.5 500.511.52 n = 10 n = 100 Figure 6.7: Empirical distribution of An. Example 6.18 Considernrolls of a die. We have seen that, if Xjis the outcome if thejth roll, then E(Xj)=7=2 andV(Xj)=3 5=12. Thus, if Snis the sum of the outcomes, and An=Sn=nis the average of the outcomes, we have E(An)=7=2 and V(An) = (35=12)=n. Therefore, as nincreases, the expected value of the average remains constant, but the variance tends to 0. If the variance is a measure of theexpected deviation from the mean this would indicate that, for large n, we can expect the average to be very near the expected value. This is in fact the case, andwe shall justify it in Chapter 8. 2 Bernoulli Trials Consider next the general Bernoulli trials process. As usual, we let Xj= 1 if the jth outcome is a success and 0 if it is a failure. If pis the probability of a success, andq=1¡p, then E(Xj)=0q+1p=p; E(X2 j)=02q+12p=p; and V(Xj)=E(X2 j)¡(E(Xj))2=p¡p2=pq : Thus, for Bernoulli trials, if Sn=X1+X2+¢¢¢+Xnis the number of successes, thenE(Sn)=np,V(Sn)=npq, andD(Sn)=pnpq: IfAn=Sn=nis the average number of successes, then E(An)=p,V(An)=pq=n , andD(An)=p pq=n .W e see that the expected proportion of successes remains pand the variance tends to 0. This suggests that the frequency interpretation of probability is a correct one. Weshall make this more precise in Chapter 8. Example 6.19 LetTdenote the number of trials until the flrst success in a Bernoulli trials process. Then Tis geometrically distributed. What is the vari- ance ofT? In Example 4.15, we saw that m T=µ12 3¢¢¢ pq pq2p¢¢¢¶ : 262 CHAPTER 6. EXPECTED VALUE AND VARIANCE In Example 6.4, we showed that E(T)=1=p : Thus, V(T)=E(T2)¡1=p2; so we need only flnd E(T2)=1p+4qp+9q2p+¢¢¢ =p( 1+4q+9q2+¢¢¢): To evaluate this sum, we start again with 1+x+x2+¢¢¢=1 1¡x: Difierentiating, we obtain 1+2x+3x2+¢¢¢=1 (1¡x)2: Multiplying by x, x+2x2+3x3+¢¢¢=x (1¡x)2: Difierentiating again gives 1+4x+9x2+¢¢¢=1+x (1¡x)3: Thus, E(T2)=p1+q (1¡q)3=1+q p2 and V(T)=E(T2)¡(E(T))2 =1+q p2¡1 p2=q p2: For example, the variance for the number of tosses of a coin until the flrst head turns up is (1 =2)=(1=2)2= 2. The variance for the number of rolls of a die until the flrst six turns up is (5 =6)=(1=6)2= 30. Note that, as pdecreases, the variance increases rapidly. This corresponds to the increased spread of the geometricdistribution as pdecreases (noted in Figure 5.1). 2 Poisson Distribution Just as in the case of expected values, it is easy to guess the variance of the Poisson distribution with parameter ‚. We recall that the variance of a binomial distribution with parameters nandpequalsnpq. We also recall that the Poisson distribution could be obtained as a limit of binomial distributions, if ngoes to1andpgoes to 0 in such a way that their product is kept flxed at the value ‚. In this case, npq=‚qapproaches ‚, sinceqgoes to 1. So, given a Poisson distribution with parameter‚, we should guess that its variance is ‚. The reader is asked to show this in Exercise 30. 6.2. VARIANCE OF DISCRETE RANDOM VARIABLES 263 Exercises 1A number is chosen at random from the set S=f¡1;0;1g. LetXbe the number chosen. Find the expected value, variance, and standard deviation ofX. 2A random variable Xhas the distribution p X=µ0124 1=31=31=61=6¶ : Find the expected value, variance, and standard deviation of X. 3You place a 1-dollar bet on the number 17 at Las Vegas, and your friend places a 1-dollar bet on black (see Exercises 1.1.6 and 1.1.7). Let Xbe your winnings and Ybe her winnings. Compare E(X),E(Y), andV(X),V(Y). What do these computations tell you about the nature of your winnings ifyou and your friend make a sequence of bets, with you betting each time ona number and your friend betting on a color? 4Xis a random variable with E(X) = 100 and V(X) = 15. Find (a)E(X 2). (b)E(3X+ 10). (c)E(¡X). (d)V(¡X). (e)D(¡X). 5In a certain manufacturing process, the (Fahrenheit) temperature never varies by more than 2–from 62–. The temperature is, in fact, a random variable F with distribution PF=µ60 61 62 63 64 1=10 2=10 4=10 2=10 1=10¶ : (a) FindE(F) andV(F). (b) Deflne T=F¡62. FindE(T) andV(T), and compare these answers with those in part (a). (c) It is decided to report the temperature readings on a Celsius scale, that is,C=( 5=9)(F¡32). What is the expected value and variance for the readings now? 6Write a computer program to calculate the mean and variance of a distribution which you specify as data. Use the program to compare the variances for thefollowing densities, both having expected value 0: p X=µ¡2¡1 012 3=11 2=11 1=11 2=11 3=11¶ ; pY=µ¡2¡1 012 1=11 2=11 5=11 2=11 1=11¶ : 264 CHAPTER 6. EXPECTED VALUE AND VARIANCE 7A coin is tossed three times. Let Xbe the number of heads that turn up. FindV(X) andD(X). 8A random sample of 2400 people are asked if they favor a government pro- posal to develop new nuclear power plants. If 40 percent of the people in thecountry are in favor of this proposal, flnd the expected value and the stan-dard deviation for the number S 2400of people in the sample who favored the proposal. 9A die is loaded so that the probability of a face coming up is proportional to the number on that face. The die is rolled with outcome X. FindV(X) and D(X). 10Prove the following facts about the standard deviation. (a)D(X+c)=D(X). (b)D(cX)=jcjD(X). 11A number is chosen at random from the integers 1, 2, 3, ...,n. LetXbe the number chosen. Show that E(X)=(n+1 )=2 andV(X)=(n¡1)(n+1 )=12. Hint: The following identity may be useful: 12+22+¢¢¢+n2=(n)(n+ 1)(2n+1 ) 6: 12LetXbe a random variable with „=E(X) and¾2=V(X). DeflneX⁄= (X¡„)=¾. The random variable X⁄is called the standardized random variable associated with X. Show that this standardized random variable has expected value 0 and variance 1. 13Peter and Paul play Heads or Tails (see Example 1.4). Let Wnbe Peter’s winnings after nmatches. Show that E(Wn)=0a n dV(Wn)=n. 14Find the expected value and the variance for the number of boys and the number of girls in a royal family that has children until there is a boy or untilthere are three children, whichever comes flrst. 15Suppose that npeople have their hats returned at random. Let X i= 1 if the ith person gets his or her own hat back and 0 otherwise. Let Sn=Pn i=1Xi. ThenSnis the total number of people who get their own hats back. Show that (a)E(X2 i)=1=n. (b)E(Xi¢Xj)=1=n(n¡1) fori6=j. (c)E(S2 n) = 2 (using (a) and (b)). (d)V(Sn)=1 . 6.2. VARIANCE OF DISCRETE RANDOM VARIABLES 265 16LetSnbe the number of successes in nindependent trials. Use the program BinomialProbabilities (Section 3.2) to compute, for given n,p, andj, the probability P(¡jpnpq<Sn¡np<jpnpq): (a) Letp=:5, and compute this probability for j=1 ,2 ,3a n d n= 10, 30, 50. Do the same for p=:2. (b) Show that the standardized random variable S⁄ n=(Sn¡np)=pnpqhas expected value 0 and variance 1. What do your results from (a) tell youabout this standardized quantity S ⁄ n? 17LetXbe the outcome of a chance experiment with E(X)=„andV(X)= ¾2. When„and¾2are unknown, the statistician often estimates them by repeating the experiment ntimes with outcomes x1,x2, ...,xn, estimating „by the sample mean „x=1 nnX i=1xi; and¾2by the sample variance s2=1 nnX i=1(xi¡„x)2: Thensis the sample standard deviation. These formulas should remind the reader of the deflnitions of the theoretical mean and variance. (Many statisti-cians deflne the sample variance with the coe–cient 1 =nreplaced by 1 =(n¡1). If this alternative deflnition is used, the expected value of s 2is equal to ¾2. See Exercise 18, part (d).) Write a computer program that will roll a die ntimes and compute the sample mean and sample variance. Repeat this experiment several times for n=1 0 andn= 1000. How well do the sample mean and sample variance estimate the true mean 7/2 and variance 35/12? 18Show that, for the sample mean „ xand sample variance s2as deflned in Exer- cise 17, (a)E(„x)=„. (b)E¡ („x¡„)2¢ =¾2=n. (c)E(s2)=n¡1 n¾2.Hint: For (c) write nX i=1(xi¡„x)2=nX i=1¡ (xi¡„)¡(„x¡„)¢2 =nX i=1(xi¡„)2¡2(„x¡„)nX i=1(xi¡„)+n(„x¡„)2 =nX i=1(xi¡„)2¡n(„x¡„)2; 266 CHAPTER 6. EXPECTED VALUE AND VARIANCE and take expectations of both sides, using part (b) when necessary. (d) Show that if, in the deflnition of s2in Exercise 17, we replace the coe–- cient 1=nby the coe–cient 1 =(n¡1), thenE(s2)=¾2. (This shows why many statisticians use the coe–cient 1 =(n¡1). The number s2is used to estimate the unknown quantity ¾2. If an estimator has an average value which equals the quantity being estimated, then the estimator issaid to be unbiased . Thus, the statement E(s 2)=¾2says thats2is an unbiased estimator of ¾2.) 19LetXbe a random variable taking on values a1,a2,...,arwith probabilities p1,p2, ...,prand withE(X)=„. Deflne the spread ofXas follows: „¾=rX i=1jai¡„jpi: This, like the standard deviation, is a way to quantify the amount that a random variable is spread out around its mean. Recall that the variance of asum of mutually independent random variables is the sum of the individualvariances. The square of the spread corresponds to the variance in a mannersimilar to the correspondence between the spread and the standard deviation.Show by an example that it is not necessarily true that the square of thespread of the sum of two independent random variables is the sum of thesquares of the individual spreads. 20We have two instruments that measure the distance between two points. The measurements given by the two instruments are random variables X 1and X2that are independent with E(X1)=E(X2)=„, where„is the true distance. From experience with these instruments, we know the values of thevariances¾ 2 1and¾2 2. These variances are not necessarily the same. From two measurements, we estimate „by the weighted average „ „=wX1+( 1¡w)X2. Herewis chosen in [0 ;1] to minimize the variance of „ „. (a) What is E(„„)? (b) How should wbe chosen in [0 ;1] to minimize the variance of „ „? 21LetXbe a random variable with E(X)=„andV(X)=¾2. Show that the functionf(x) deflned by f(x)=X !(X(!)¡x)2p(!) has its minimum value when x=„. 22LetXandYbe two random variables deflned on the flnite sample space ›. Assume that X,Y,X+Y, andX¡Yall have the same distribution. Prove thatP(X=Y=0 )=1 . 6.2. VARIANCE OF DISCRETE RANDOM VARIABLES 267 23IfXandYare any two random variables, then the covariance ofXandYis deflned by Cov( X;Y )=E((X¡E(X))(Y¡E(Y))). Note that Cov( X;X )= V(X). Show that, if XandYare independent, then Cov( X;Y )=0 ;a n d show, by an example, that we can have Cov( X;Y )=0a n d XandYnot independent. *24 A professor wishes to make up a true-false exam with nquestions. She assumes that she can design the problems in such a way that a student will answerthejth problem correctly with probability p j, and that the answers to the various problems may be considered independent experiments. Let Snbe the number of problems that a student will get correct. The professor wishes tochoosep jso thatE(Sn)=:7nand so that the variance of Snis as large as possible. Show that, to achieve this, she should choose pj=:7 for allj; that is, she should make all the problems have the same di–culty. 25(Lamperti20) An urn contains exactly 5000 balls, of which an unknown number Xare white and the rest red, where Xis a random variable with a probability distribution on the integers 0, 1, 2, ...,5000. (a) Suppose we know that E(X)=„. Show that this is enough to allow us to calculate the probability that a ball drawn at random from the urnwill be white. What is this probability? (b) Suppose the variance of Xis¾ 2. What is the probability of drawing two white balls in part (b)? 26We draw a ball from the urn, examine its color, replace it, and then draw another. Under what conditions, if any, are the results of the two drawingsindependent; that is, does P(white;white) =P(white) 2? 27For a sequence of Bernoulli trials, let X1be the number of trials until the flrst success. For j‚2, letXjbe the number of trials after the ( j¡1)st success until thejth success. It can be shown that X1,X2,...i sa n independent trials process. (a) What is the common distribution, expected value, and variance for Xj? (b) LetTn=X1+X2+¢¢¢+Xn. ThenTnis the time until the nth success. FindE(Tn) andV(Tn). (c) Use the results of (b) to flnd the expected value and variance for the number of tosses of a coin until the nth occurrence of a head. 28Referring to Exercise 6.1.30, flnd the variance for the number of boxes of Wheaties bought before getting half of the players’ pictures and the variancefor the number of additional boxes needed to get the second half of the players’pictures. 20Private communication. 268 CHAPTER 6. EXPECTED VALUE AND VARIANCE 29In Example 5.3, assume that the book in question has 1000 pages. Let Xbe the number of pages with no mistakes. Show that E(X) = 905 and V(X)= 86. Using these results, show that the probability is •:05 that there will be more than 924 pages without errors or fewer than 866 pages without errors. 30LetXbe Poisson distributed with parameter ‚. Show that V(X)=‚. 6.3 Continuous Random Variables In this section we consider the properties of the expected value and the variance of a continuous random variable. These quantities are deflned just as for discreterandom variables and share the same properties. Expected Value Deflnition 6.4 LetXbe a real-valued random variable with density function f(x). The expected value „=E(X) is deflned by „=E(X)=Z+1 ¡1xf(x)dx ; provided the integralZ+1 ¡1jxjf(x)dx is flnite. 2 The reader should compare this deflnition with the corresponding one for discrete random variables in Section 6.1. Intuitively, we can interpret E(X), as we did in the previous sections, as the value that we should expect to obtain if we perform alarge number of independent experiments and average the resulting values of X. We can summarize the properties of E(X) as follows (cf. Theorem 6.2). Theorem 6.10 IfXandYare real-valued random variables and cis any constant, then E(X+Y)=E(X)+E(Y); E(cX)=cE(X): The proof is very similar to the proof of Theorem 6.2, and we omit it. 2 More generally, if X 1,X2,...,Xnarenreal-valued random variables, and c1,c2, ...,cnarenconstants, then E(c1X1+c2X2+¢¢¢+cnXn)=c1E(X1)+c2E(X2)+¢¢¢+cnE(Xn): 6.3. CONTINUOUS RANDOM VARIABLES 269 Example 6.20 LetXbe uniformly distributed on the interval [0 ;1]. Then E(X)=Z1 0xdx =1=2: It follows that if we choose a large number Nof random numbers from [0 ;1] and take the average, then we can expect that this average should be close to the expectedvalue of 1/2. 2 Example 6.21 LetZ=(x;y) denote a point chosen uniformly and randomly from the unit disk, as in the dart game in Example 2.8 and let X=(x 2+y2)1=2be the distance from Zto the center of the disk. The density function of Xcan easily be shown to equal f(x)=2x, so by the deflnition of expected value, E(X)=Z1 0xf(x)dx =Z1 0x(2x)dx =2 3: 2 Example 6.22 In the example of the couple meeting at the Inn (Example 2.16), each person arrives at a time which is uniformly distributed between 5:00 and 6:00PM. The random variable Zunder consideration is the length of time the flrst person has to wait until the second one arrives. It was shown that f Z(z) = 2(1¡z); for 0•z•1. Hence, E(Z)=Z1 0zfZ(z)dz =Z1 02z(1¡z)dz =h z2¡2 3z3i1 0 =1 3: 2 Expectation of a Function of a Random Variable Suppose that Xis a real-valued random variable and `(x) is a continuous function fromRtoR. The following theorem is the continuous analogue of Theorem 6.1. 270 CHAPTER 6. EXPECTED VALUE AND VARIANCE Theorem 6.11 IfXis a real-valued random variable and if `:R!Ris a continuous real-valued function with domain [ a;b], then E(`(X)) =Z+1 ¡1`(x)fX(x)dx ; provided the integral exists. 2 For a proof of this theorem, see Ross.21 Expectation of the Product of Two Random Variables In general, it is not true that E(XY)=E(X)E(Y), since the integral of a product is not the product of integrals. But if XandYare independent, then the expectations multiply. Theorem 6.12 LetXandYbe independent real-valued continuous random vari- ables with flnite expected values. Then we have E(XY)=E(X)E(Y): Proof. We will prove this only in the case that the ranges of XandYare contained in the intervals [ a;b] and [c;d], respectively. Let the density functions of XandY be denoted by fX(x) andfY(y), respectively. Since XandYare independent, the joint density function of XandYis the product of the individual density functions. Hence E(XY)=Zb aZd cxyfX(x)fY(y)dydx =Zb axfX(x)dxZd cyfY(y)dy =E(X)E(Y): The proof in the general case involves using sequences of bounded random vari- ables that approach XandY, and is somewhat technical, so we will omit it. 2 In the same way, one can show that if X1,X2, ...,Xnarenmutually indepen- dent real-valued random variables, then E(X1X2¢¢¢Xn)=E(X1)E(X2)¢¢¢E(Xn): Example 6.23 LetZ=(X;Y ) be a point chosen at random in the unit square. LetA=X2andB=Y2. Then Theorem 4.3 implies that AandBare independent. Using Theorem 6.11, the expectations of AandBare easy to calculate: E(A)=E(B)=Z1 0x2dx =1 3: 21S. Ross, A First Course in Probability, (New York: Macmillan, 1984), pgs. 241-245. 6.3. CONTINUOUS RANDOM VARIABLES 271 Using Theorem 6.12, the expectation of ABis just the product of E(A) andE(B), or 1/9. The usefulness of this theorem is demonstrated by noting that it is quite abit more di–cult to calculate E(AB) from the deflnition of expectation. One flnds that the density function of ABis f AB(t)=¡log(t) 4p t; so E(AB)=Z1 0tfAB(t)dt =1 9: 2 Example 6.24 Again letZ=(X;Y ) be a point chosen at random in the unit square, and let W=X+Y. ThenYandWare not independent, and we have E(Y)=1 2; E(W)=1; E(YW)=E(XY+Y2)=E(X)E(Y)+1 3=7 126=E(Y)E(W): 2 We turn now to the variance. Variance Deflnition 6.5 LetXbe a real-valued random variable with density function f(x). The variance¾2=V(X) is deflned by ¾2=V(X)=E((X¡„)2): 2 The next result follows easily from Theorem 6.1. There is another way to calculate the variance of a continuous random variable, which is usually slightly easier. It isgiven in Theorem 6.15. Theorem 6.13 IfXis a real-valued random variable with E(X)=„, then ¾ 2=Z1 ¡1(x¡„)2f(x)dx : 2 272 CHAPTER 6. EXPECTED VALUE AND VARIANCE The properties listed in the next three theorems are all proved in exactly the same way that the corresponding theorems for discrete random variables wereproved in Section 6.2. Theorem 6.14 IfXis a real-valued random variable deflned on › and cis any constant, then (cf. Theorem 6.7) V(cX)=c 2V(X); V(X+c)=V(X): 2 Theorem 6.15 IfXis a real-valued random variable with E(X)=„, then (cf. Theorem 6.6) V(X)=E(X2)¡„2: 2 Theorem 6.16 IfXandYare independent real-valued random variables on ›, then (cf. Theorem 6.8) V(X+Y)=V(X)+V(Y): 2 Example 6.25 (continuation of Example 6.20) If Xis uniformly distributed on [0;1], then, using Theorem 6.15, we have V(X)=Z1 0‡ x¡1 2·2 dx=1 12: 2 Example 6.26 LetXbe an exponentially distributed random variable with pa- rameter‚. Then the density function of Xis fX(x)=‚e¡‚x: From the deflnition of expectation and integration by parts, we have E(X)=Z1 0xfX(x)dx =‚Z1 0xe¡‚xdx =¡xe¡‚xflflflfl1 0+Z1 0e¡‚xdx =0 +e¡‚x ¡‚flflflfl1 0=1 ‚: 6.3. CONTINUOUS RANDOM VARIABLES 273 Similarly, using Theorems 6.11 and 6.15, we have V(X)=Z1 0x2fX(x)dx¡1 ‚2 =‚Z1 0x2e¡‚xdx¡1 ‚2 =¡x2e¡‚xflflflfl1 0+2Z1 0xe¡‚xdx¡1 ‚2 =¡x2e¡‚xflflflfl1 0¡2xe¡‚x ‚flflflfl1 0¡2 ‚2e¡‚xflflflfl1 0¡1 ‚2=2 ‚2¡1 ‚2=1 ‚2: In this case, both E(X) andV(X) are flnite if ‚>0. 2 Example 6.27 LetZbe a standard normal random variable with density function fZ(x)=1p 2…e¡x2=2: Since this density function is symmetric with respect to the y-axis, then it is easy to show that Z1 ¡1xfZ(x)dx has value 0. The reader should recall however, that the expectation is deflned to be the above integral only if the integral Z1 ¡1jxjfZ(x)dx is flnite. This integral equals 2Z1 0xfZ(x)dx ; which one can easily show is flnite. Thus, the expected value of Zis 0. To calculate the variance of Z, we begin by applying Theorem 6.15: V(Z)=Z+1 ¡1x2fZ(x)dx¡„2: If we write x2asx¢x, and integrate by parts, we obtain 1p 2…(¡xe¡x2=2)flflflfl+1 ¡1+1p 2…Z+1 ¡1e¡x2=2dx : The flrst summand above can be shown to equal 0, since as x!§1 ,e¡x2=2gets small more quickly than xgets large. The second summand is just the standard normal density integrated over its domain, so the value of this summand is 1.Therefore, the variance of the standard normal density equals 1. 274 CHAPTER 6. EXPECTED VALUE AND VARIANCE Now letXbe a (not necessarily standard) normal random variable with param- eters„and¾. Then the density function of Xis fX(x)=1p 2…¾e¡(x¡„)2=2¾2: We can write X=¾Z+„, whereZis a standard normal random variable. Since E(Z)=0a n dV(Z) = 1 by the calculation above, Theorems 6.10 and 6.14 imply that E(X)=E(¾Z+„)=„; V(X)=V(¾Z+„)=¾2: 2 Example 6.28 LetXbe a continuous random variable with the Cauchy density function fX(x)=a …1 a2+x2: Then the expectation of Xdoes not exist, because the integral a …Z+1 ¡1jxjdx a2+x2 diverges. Thus the variance of Xalso fails to exist. Densities whose variance is not deflned, like the Cauchy density, behave quite difierently in a number of importantrespects from those whose variance is flnite. We shall see one instance of thisdifierence in Section 8.2. 2 Independent Trials Corollary 6.1 IfX1,X2, ...,Xnis an independent trials process of real-valued random variables, with E(Xi)=„andV(Xi)=¾2, and if Sn=X1+X2+¢¢¢+Xn; An=Sn n; then E(Sn)=n„ ; E(An)=„; V(Sn)=n¾2; V(An)=¾2 n: It follows that if we set S⁄ n=Sn¡n„p n¾2; 6.3. CONTINUOUS RANDOM VARIABLES 275 then E(S⁄ n)=0; V(S⁄ n)=1: We say that S⁄ nis astandardized version of Sn(see Exercise 12 in Section 6.2). 2 Queues Example 6.29 Let us consider again the queueing problem, that is, the problem of the customers waiting in a queue for service (see Example 5.7). We suppose againthat customers join the queue in such a way that the time between arrivals is anexponentially distributed random variable Xwith density function f X(t)=‚e¡‚t: Then the expected value of the time between arrivals is simply 1 =‚(see Exam- ple 6.26), as was stated in Example 5.7. The reciprocal ‚of this expected value is often referred to as the arrival rate. The service time of an individual who is flrst in line is deflned to be the amount of time that the person stays at the headof the line before leaving. We suppose that the customers are served in such a waythat the service time is another exponentially distributed random variable Ywith density function f X(t)=„e¡„t: Then the expected value of the service time is E(X)=Z1 0tfX(t)dt=1 „: The reciprocal „if this expected value is often referred to as the service rate. We expect on grounds of our everyday experience with queues that if the service rate is greater than the arrival rate, then the average queue size will tend to stabilize,but if the service rate is less than the arrival rate, then the queue will tend to increasein length without limit (see Figure 5.7). The simulations in Example 5.7 tend tobear out our everyday experience. We can make this conclusion more precise if weintroduce the tra–c intensity as the product ‰= (arrival rate)(average service time) =‚ „=1=„ 1=‚: The tra–c intensity is also the ratio of the average service time to the average time between arrivals. If the tra–c intensity is less than 1 the queue will performreasonably, but if it is greater than 1 the queue will grow indeflnitely large. In thecritical case of ‰= 1, it can be shown that the queue will become large but there will always be times at which the queue is empty. 22 22L. Kleinrock, Queueing Systems, vol. 2 (New York: John Wiley and Sons, 1975). 276 CHAPTER 6. EXPECTED VALUE AND VARIANCE In the case that the tra–c intensity is less than 1 we can consider the length of the queue as a random variable Zwhose expected value is flnite, E(Z)=N: The time spent in the queue by a single customer can be considered as a random variableWwhose expected value is flnite, E(W)=T: Then we can argue that, when a customer joins the queue, he expects to flnd N people ahead of him, and when he leaves the queue, he expects to flnd ‚Tpeople behind him. Since, in equilibrium, these should be the same, we would expect toflnd that N=‚T : This last relationship is called Little’s law for queues. 23We will not prove it here. A proof may be found in Ross.24Note that in this case we are counting the waiting time of all customers, even those that do not have to wait at all. In our simulationin Section 4.2, we did not consider these customers. If we knew the expected queue length then we could use Little’s law to obtain the expected waiting time, since T=N ‚: The queue length is a random variable with a discrete distribution. We can estimate this distribution by simulation, keeping track of the queue lengths at the times atwhich a customer arrives. We show the result of this simulation (using the programQueue ) in Figure 6.8. We note that the distribution appears to be a geometric distribution. In the study of queueing theory it is shown that the distribution for the queue length inequilibrium is indeed a geometric distribution with s j=( 1¡‰)‰jforj=0;1;2;::: ; if‰<1. The expected value of a random variable with this distribution is N=‰ (1¡‰) (see Example 6.4). Thus by Little’s result the expected waiting time is T=‰ ‚(1¡‰)=1 „¡‚; where„is the service rate, ‚the arrival rate, and ‰the tra–c intensity. 23ibid., p. 17. 24S. M. Ross, Applied Probability Models with Optimization Applications, (San Francisco: Holden-Day, 1970) 6.3. CONTINUOUS RANDOM VARIABLES 277 0 10 20 30 40 5000.020.040.060.08 Figure 6.8: Distribution of queue lengths. In our simulation, the arrival rate is 1 and the service rate is 1.1. Thus, the tra–c intensity is 1 =1:1=1 0=11, the expected queue size is 10=11 (1¡10=11)=1 0; and the expected waiting time is 1 1:1¡1=1 0: In our simulation the average queue size was 8.19 and the average waiting time was 7.37. In Figure 6.9, we show the histogram for the waiting times. This histogramsuggests that the density for the waiting times is exponential with parameter „¡‚, and this is the case. 2 Exercises 1LetXbe a random variable with range [ ¡1;1] and letfX(x) be the density function of X. Find„(X) and¾2(X) if, forjxj<1, (a)fX(x)=1=2. (b)fX(x)=jxj. (c)fX(x)=1¡jxj. (d)fX(x)=( 3=2)x2. 2LetXbe a random variable with range [ ¡1;1] andfXits density function. Find„(X) and¾2(X) if, forjxj>1,fX(x) = 0, and forjxj<1, (a)fX(x)=( 3=4)(1¡x2). 278 CHAPTER 6. EXPECTED VALUE AND VARIANCE 0 10 20 30 40 5000.020.040.060.08 Figure 6.9: Distribution of queue waiting times. (b)fX(x)=(…=4) cos(…x=2). (c)fX(x)=(x+1 )=2. (d)fX(x)=( 3=8)(x+1 )2. 3The lifetime, measure in hours, of the ACME super light bulb is a random variableTwith density function fT(t)=‚2te¡‚t, where‚=:05. What is the expected lifetime of this light bulb? What is its variance? 4LetXbe a random variable with range [ ¡1;1] and density function fX(x)= ax+bifjxj<1. (a) Show that ifR+1 ¡1fX(x)dx= 1, thenb=1=2. (b) Show that if fX(x)‚0, then¡1=2•a•1=2. (c) Show that „=( 2=3)a, and hence that ¡1=3•„•1=3. (d) Show that ¾2(X)=( 2=3)b¡(4=9)a2=1=3¡(4=9)a2. 5LetXbe a random variable with range [ ¡1;1] and density function fX(x)= ax2+bx+cifjxj<1 and 0 otherwise. (a) Show that 2 a=3+2c= 1 (see Exercise 4). (b) Show that 2 b=3=„(X). (c) Show that 2 a=5+2c=3=¾2(X). (d) Finda,b, andcif„(X)=0 ,¾2(X)=1=15, and sketch the graph of fX. (e) Finda,b, andcif„(X)=0 ,¾2(X)=1=2, and sketch the graph of fX. 6LetTbe a random variable with range [0 ;1] andfTits density function. Find„(T) and¾2(T) if, fort<0,fT(t) = 0, and for t>0, 6.3. CONTINUOUS RANDOM VARIABLES 279 (a)fT(t)=3e¡3t. (b)fT(t)=9te¡3t=2. (c)fT(t)=3=(1 +t)4. 7LetXbe a random variable with density function fX. Show, using elementary calculus, that the function `(a)=E((X¡a)2) takes its minimum value when a=„(X), and in that case `(a)=¾2(X). 8LetXbe a random variable with mean „and variance ¾2. LetY=aX2+ bX+c. Find the expected value of Y. 9LetX,Y, andZbe independent random variables, each with mean „and variance¾2. (a) Find the expected value and variance of S=X+Y+Z. (b) Find the expected value and variance of A=( 1=3)(X+Y+Z). (c) Find the expected value of S2andA2. 10LetXandYbe independent random variables with uniform density functions on [0;1]. Find (a)E(jX¡Yj). (b)E(max(X;Y )). (c)E(min(X;Y )). (d)E(X2+Y2). (e)E((X+Y)2). 11The Pilsdorfi Beer Company runs a °eet of trucks along the 100 mile road from Hangtown to Dry Gulch. The trucks are old, and are apt to breakdown at any point along the road with equal probability. Where should thecompany locate a garage so as to minimize the expected distance from atypical breakdown to the garage? In other words, if Xis a random variable giving the location of the breakdown, measured, say, from Hangtown, and b gives the location of the garage, what choice of bminimizesE(jX¡bj)? Now supposeXis not distributed uniformly over [0 ;100], but instead has density functionf X(x)=2x=10;000. Then what choice of bminimizesE(jX¡bj)? 12FindE(XY), whereXandYare independent random variables which are uniform on [0 ;1]. Then verify your answer by simulation. 13LetXbe a random variable that takes on nonnegative values and has distri- bution function F(x). Show that E(X)=Z1 0(1¡F(x))dx : 280 CHAPTER 6. EXPECTED VALUE AND VARIANCE Hint: Integrate by parts. Illustrate this result by calculating E(X) by this method if Xhas an expo- nential distribution F(x)=1¡e¡‚xforx‚0, andF(x) = 0 otherwise. 14LetXbe a continuous random variable with density function fX(x). Show that ifZ+1 ¡1x2fX(x)dx<1; thenZ+1 ¡1jxjfX(x)dx<1: Hint: Except on the interval [ ¡1;1], the flrst integrand is greater than the second integrand. 15LetXbe a random variable distributed uniformly over [0 ;20]. Deflne a new random variable YbyY=bXc(the greatest integer in X). Find the expected value ofY. Do the same for Z=bX+:5c. Compute E¡ jX¡Yj¢ and E¡ jX¡Zj¢ . (Note that Yis the value of Xrounded ofi to the nearest smallest integer, while Zis the value of Xrounded ofi to the nearest integer. Which method of rounding ofi is better? Why?) 16Assume that the lifetime of a diesel engine part is a random variable Xwith densityfX. When the part wears out, it is replaced by another with the same density. Let N(t) be the number of parts that are used in time t. We want to study the random variable N(t)=t. Since parts are replaced on the average everyE(X) time units, we expect about t=E(X) parts to be used in time t. That is, we expect that lim t!1E‡N(t) t· =1 E(X): This result is correct but quite di–cult to prove. Write a program that will allow you to specify the density fX, and the time t, and simulate this experi- ment to flnd N(t)=t. Have your program repeat the experiment 500 times and plot a bar graph for the random outcomes of N(t)=t. From this data, estimate E(N(t)=t) and compare this with 1 =E(X). In particular, do this for t= 100 with the following two densities: (a)fX=e¡t. (b)fX=te¡t. 17LetXandYbe random variables. The covariance Cov(X;Y) is deflned by (see Exercise 6.2.23) cov(X;Y) = E((X¡„(X))(Y¡„(Y))): (a) Show that cov(X ;Y) = E(XY)¡E(X)E(Y). 6.3. CONTINUOUS RANDOM VARIABLES 281 (b) Using (a), show that cov( X;Y )=0 ,i fXandYare independent. (Cau- tion: the converse is notalways true.) (c) Show that V(X+Y)=V(X)+V(Y) + 2cov(X;Y ). 18LetXandYbe random variables with positive variance. The correlation of XandYis deflned as ‰(X;Y )=cov(X;Y )p V(X)V(Y): (a) Using Exercise 17(c), show that 0•VµX ¾(X)+Y ¾(Y)¶ = 2(1 +‰(X;Y )): (b) Now show that 0•VµX ¾(X)¡Y ¾(Y)¶ = 2(1¡‰(X;Y )): (c) Using (a) and (b), show that ¡1•‰(X;Y )•1: 19LetXandYbe independent random variables with uniform densities in [0 ;1]. LetZ=X+YandW=X¡Y. Find (a)‰(X;Y ) (see Exercise 18). (b)‰(X;Z). (c)‰(Y;W ). (d)‰(Z;W ). *20 When studying certain physiological data, such as heights of fathers and sons, it is often natural to assume that these data (e.g., the heights of the fathersand the heights of the sons) are described by random variables with normaldensities. These random variables, however, are not independent but ratherare correlated. For example, a two-dimensional standard normal density forcorrelated random variables has the form f X;Y(x;y)=1 2…p 1¡‰2¢e¡(x2¡2‰xy+y2)=2(1¡‰2): (a) Show that XandYeach have standard normal densities. (b) Show that the correlation of XandY(see Exercise 18) is ‰. *21 For correlated random variables XandYit is natural to ask for the expected value forXgivenY. For example, Galton calculated the expected value of the height of a son given the height of the father. He used this to show 282 CHAPTER 6. EXPECTED VALUE AND VARIANCE that tall men can be expected to have sons who are less tall on the average. Similarly, students who do very well on one exam can be expected to do lesswell on the next exam, and so forth. This is called regression on the mean. To deflne this conditional expected value, we flrst deflne a conditional densityofXgivenY=yby f XjY(xjy)=fX;Y(x;y) fY(y); wherefX;Y(x;y) is the joint density of XandY, andfYis the density for Y. Then the conditional expected value of XgivenYis E(XjY=y)=Zb axfXjY(xjy)dx : For the normal density in Exercise 20, show that the conditional density of fXjY(xjy) is normal with mean ‰yand variance 1¡‰2. From this we see that ifXandYare positively correlated (0 <‰< 1), and ify>E (Y), then the expected value for XgivenY=ywill be less than y(i.e., we have regression on the mean). 22A pointYis chosen at random from [0 ;1]. A second point Xis then chosen from the interval [0 ;Y]. Find the density for X.Hint: Calculate fXjYas in Exercise 21 and then use fX(x)=Z1 xfXjY(xjy)fY(y)dy : Can you also derive your result geometrically? *23 LetXandVbe two standard normal random variables. Let ‰be a real number between -1 and 1. (a) LetY=‰X+p 1¡‰2V. Show that E(Y) = 0 andVar(Y)=1 . W e shall see later (see Example 7.5 and Example 10.17), that the sum of twoindependent normal random variables is again normal. Thus, assumingthis fact, we have shown that Yis standard normal. (b) Using Exercises 17 and 18, show that the correlation of XandYis‰. (c) In Exercise 20, the joint density function f X;Y(x;y) for the random vari- able (X;Y ) is given. Now suppose that we want to know the set of points (x;y)i nt h exy-plane such that fX;Y(x;y)=Cfor some constant C. This set of points is called a set of constant density. Roughly speak- ing, a set of constant density is a set of points where the outcomes ( X;Y ) are equally likely to fall. Show that for a given C, the set of points of constant density is a curve whose equation is x2¡2‰xy+y2=D; whereDis a constant which depends upon C. (This curve is an ellipse.) 6.3. CONTINUOUS RANDOM VARIABLES 283 (d) One can plot the ellipse in part (c) by using the parametric equations x=rcosµp 2(1¡‰)+rsinµp 2(1 +‰); y=rcosµp 2(1¡‰)¡rsinµp 2(1 +‰): Write a program to plot 1000 pairs ( X;Y ) for‰=¡1=2;0;1=2. For each plot, have your program plot the above parametric curves for r=1;2;3. *24 Following Galton, let us assume that the fathers and sons have heights that are dependent normal random variables. Assume that the average height is68 inches, standard deviation is 2.7 inches, and the correlation coe–cient is .5(see Exercises 20 and 21). That is, assume that the heights of the fathersand sons have the form 2 :7X+ 68 and 2 :7Y+ 68, respectively, where X andYare correlated standardized normal random variables, with correlation coe–cient .5. (a) What is the expected height for the son of a father whose height is 72 inches? (b) Plot a scatter diagram of the heights of 1000 father and son pairs. Hint: You can choose standardized pairs as in Exercise 23 and then plot (2 :7X+ 68;2:7Y+ 68). *25 When we have pairs of data ( x i;yi) that are outcomes of the pairs of dependent random variables X,Ywe can estimate the coorelation coe–cient ‰by „r=P i(xi¡„x)(yi¡„y) (n¡1)sXsY; where „xand „yare the sample means for XandY, respectively, and sXandsY are the sample standard deviations for XandY(see Exercise 6.2.17). Write a program to compute the sample means, variances, and correlation for suchdependent data. Use your program to compute these quantities for Galton’sdata on heights of parents and children given in Appendix B. Plot the equal density ellipses as deflned in Exercise 23 for r= 4, 6, and 8, and on the same graph print the values that appear in the table at the appropriatepoints. For example, print 12 at the point (70 :5;68:2), indicating that there were 12 cases where the parent’s height was 70.5 and the child’s was 68.12.See if Galton’s data is consistent with the equal density ellipses. 26(from Hamming 25) Suppose you are standing on the bank of a straight river. (a) Choose, at random, a direction which will keep you on dry land, and walk 1 km in that direction. Let Pdenote your position. What is the expected distance from Pto the river? 25R. W. Hamming, The Art of Probability for Scientists and Engineers (Redwood City: Addison-Wesley, 1991), p. 192. 284 CHAPTER 6. EXPECTED VALUE AND VARIANCE (b) Now suppose you proceed as in part (a), but when you get to P, you pick a random direction (from among alldirections) and walk 1 km. What is the probability that you will reach the river before the second walk iscompleted? 27(from Hamming 26) A game is played as follows: A random number Xis chosen uniformly from [0 ;1]. Then a sequence Y1;Y2;:::of random numbers is chosen independently and uniformly from [0 ;1]. The game ends the flrst time that Yi>X . You are then paid ( i¡1) dollars. What is a fair entrance fee for this game? 28A long needle of length Lmuch bigger than 1 is dropped on a grid with horizontal and vertical lines one unit apart. Show that the average number a of lines crossed is approximately a=4L …: 26ibid., pg. 205. Chapter 7 Sums of Independent Random Variables 7.1 Sums of Discrete Random Variables In this chapter we turn to the important question of determining the distribution of a sum of independent random variables in terms of the distributions of the individualconstituents. In this section we consider only sums of discrete random variables,reserving the case of continuous random variables for the next section. We consider here only random variables whose values are integers. Their distri- bution functions are then deflned on these integers. We shall flnd it convenient toassume here that these distribution functions are deflned for allintegers, by deflning them to be 0 where they are not otherwise deflned. Convolutions SupposeXandYare two independent discrete random variables with distribution functionsm1(x) andm2(x). LetZ=X+Y. We would like to determine the dis- tribution function m3(x)o fZ. To do this, it is enough to determine the probability thatZtakes on the value z, wherezis an arbitrary integer. Suppose that X=k, wherekis some integer. Then Z=zif and only if Y=z¡k. So the event Z=z is the union of the pairwise disjoint events (X=k) and (Y=z¡k); wherekruns over the integers. Since these events are pairwise disjoint, we have P(Z=z)=1X k=¡1P(X=k)¢P(Y=z¡k): Thus, we have found the distribution function of the random variable Z. This leads to the following deflnition. 285 286 CHAPTER 7. SUMS OF RANDOM VARIABLES Deflnition 7.1 LetXandYbe two independent integer-valued random variables, with distribution functions m1(x) andm2(x) respectively. Then the convolution of m1(x) andm2(x) is the distribution function m3=m1⁄m2given by m3(j)=X km1(k)¢m2(j¡k); forj=:::;¡2;¡1;0;1;2; :::. The function m3(x) is the distribution function of the random variable Z=X+Y. 2 It is easy to see that the convolution operation is commutative, and it is straight- forward to show that it is also associative. Now letSn=X1+X2+¢¢¢+Xnbe the sum of nindependent random variables of an independent trials process with common distribution function mdeflned on the integers. Then the distribution function of S1ism. We can write Sn=Sn¡1+Xn: Thus, since we know the distribution function of Xnism, we can flnd the distribu- tion function of Snby induction. Example 7.1 A die is rolled twice. Let X1andX2be the outcomes, and let S2=X1+X2be the sum of these outcomes. Then X1andX2have the common distribution function: m=µ123456 1=61=61=61=61=61=6¶ : The distribution function of S2is then the convolution of this distribution with itself. Thus, P(S2=2 ) =m(1)m(1) =1 6¢1 6=1 36; P(S2=3 ) =m(1)m(2) +m(2)m(1) =1 6¢1 6+1 6¢1 6=2 36; P(S2=4 ) =m(1)m(3) +m(2)m(2) +m(3)m(1) =1 6¢1 6+1 6¢1 6+1 6¢1 6=3 36: Continuing in this way we would flnd P(S2=5 )=4=36,P(S2=6 )=5=36, P(S2=7 )=6=36,P(S2=8 )=5=36,P(S2=9 )=4=36,P(S2=1 0 )=3=36, P(S2=1 1 )=2=36, andP(S2=1 2 )=1=36. The distribution for S3would then be the convolution of the distribution for S2 with the distribution for X3.T h u s P(S3=3 ) =P(S2=2 )P(X3=1 ) 7.1. SUMS OF DISCRETE RANDOM VARIABLES 287 =1 36¢1 6=1 216; P(S3=4 ) =P(S2=3 )P(X3=1 )+P(S2=2 )P(X3=2 ) =2 36¢1 6+1 36¢1 6=3 216; and so forth. This is clearly a tedious job, and a program should be written to carry out this calculation. To do this we flrst write a program to form the convolution of twodensitiespandqand return the density r. We can then write a program to flnd the density for the sum S nofnindependent random variables with a common density p, at least in the case that the random variables have a flnite number of possible values. Running this program for the example of rolling a die ntimes forn=1 0;20;30 results in the distributions shown in Figure 7.1. We see that, as in the case ofBernoulli trials, the distributions become bell-shaped. We shall discuss in Chapter 9a very general theorem called the Central Limit Theorem that will explain this phenomenon. 2 Example 7.2 A well-known method for evaluating a bridge hand is: an ace is assigned a value of 4, a king 3, a queen 2, and a jack 1. All other cards are assigneda value of 0. The point count of the hand is then the sum of the values of the cards in the hand. (It is actually more complicated than this, taking into accountvoids in suits, and so forth, but we consider here this simplifled form of the pointcount.) If a card is dealt at random to a player, then the point count for this cardhas distribution p X=µ0 1234 36=52 4=52 4=52 4=52 4=52¶ : Let us regard the total hand of 13 cards as 13 independent trials with this common distribution. (Again this is not quite correct because we assume here thatwe are always choosing a card from a full deck.) Then the distribution for the pointcountCfor the hand can be found from the program NFoldConvolution by using the distribution for a single card and choosing n= 13. A player with a point count of 13 or more is said to have an opening bid. The probability of having an opening bid is then P(C‚13): Since we have the distribution of C, it is easy to compute this probability. Doing this we flnd that P(C‚13) =:2845; so that about one in four hands should be an opening bid according to this simplifled model. A more realistic discussion of this problem can be found in Epstein, The Theory of Gambling and Statistical Logic. 12 1R. A. Epstein, The Theory of Gambling and Statistical Logic, rev. ed. (New York: Academic Press, 1977). 288 CHAPTER 7. SUMS OF RANDOM VARIABLES 20 40 60 80 100 120 14000.010.020.030.040.050.060.070.08 20 40 60 80 100 120 14000.010.020.030.040.050.060.070.08 20 40 60 80 100 120 14000.010.020.030.040.050.060.070.08n = 10 n = 20 n = 30 Figure 7.1: Density of Snfor rolling a die ntimes. 7.1. SUMS OF DISCRETE RANDOM VARIABLES 289 For certain special distributions it is possible to flnd an expression for the dis- tribution that results from convoluting the distribution with itself ntimes. The convolution of two binomial distributions, one with parameters mandp and the other with parameters nandp, is a binomial distribution with parameters (m+n) andp. This fact follows easily from a consideration of the experiment which consists of flrst tossing a coin mtimes, and then tossing it nmore times. The convolution of kgeometric distributions with common parameter pis a negative binomial distribution with parameters pandk. This can be seen by con- sidering the experiment which consists of tossing a coin until the kth head appears. Exercises 1A die is rolled three times. Find the probability that the sum of the outcomes is (a) greater than 9. (b) an odd number. 2The price of a stock on a given trading day changes according to the distri- bution pX=µ¡1 012 1=41=21=81=8¶ : Find the distribution for the change in stock price after two (independent) trading days. 3LetX1andX2be independent random variables with common distribution pX=µ012 1=83=81=2¶ : Find the distribution of the sum X1+X2. 4In one play of a certain game you win an amount Xwith distribution pX=µ123 1=41=41=2¶ : Using the program NFoldConvolution flnd the distribution for your total winnings after ten (independent) plays. Plot this distribution. 5Consider the following two experiments: the flrst has outcome Xtaking on the values 0, 1, and 2 with equal probabilities; the second results in an (in-dependent) outcome Ytaking on the value 3 with probability 1/4 and 4 with probability 3/4. Find the distribution of (a)Y+X. (b)Y¡X. 290 CHAPTER 7. SUMS OF RANDOM VARIABLES 6People arrive at a queue according to the following scheme: During each minute of time either 0 or 1 person arrives. The probability that 1 personarrives ispand that no person arrives is q=1¡p. LetC rbe the number of customers arriving in the flrst rminutes. Consider a Bernoulli trials process with a success if a person arrives in a unit time and failure if no person arrivesin a unit time. Let T rbe the number of failures before the rth success. (a) What is the distribution for Tr? (b) What is the distribution for Cr? (c) Find the mean and variance for the number of customers arriving in the flrstrminutes. 7(a) A die is rolled three times with outcomes X1,X2, andX3. LetY3be the maximum of the values obtained. Show that P(Y3•j)=P(X1•j)3: Use this to flnd the distribution of Y3.D o e sY3have a bell-shaped dis- tribution? (b) Now let Ynbe the maximum value when ndice are rolled. Find the distribution of Yn. Is this distribution bell-shaped for large values of n? 8A baseball player is to play in the World Series. Based upon his season play, you estimate that if he comes to bat four times in a game the number of hitshe will get has a distribution p X=µ01234 :4:2:2:1:1¶ : Assume that the player comes to bat four times in each game of the series. (a) LetXdenote the number of hits that he gets in a series. Using the program NFoldConvolution , flnd the distribution of Xfor each of the possible series lengths: four-game, flve-game, six-game, seven-game. (b) Using one of the distribution found in part (a), flnd the probability that his batting average exceeds .400 in a four-game series. (The battingaverage is the number of hits divided by the number of times at bat.) (c) Given the distribution p X, what is his long-term batting average? 9Prove that you cannot load two dice in such a way that the probabilities for any sum from 2 to 12 are the same. (Be sure to consider the case where oneor more sides turn up with probability zero.) 10(L¶evy 2) Assume that nis an integer, not prime. Show that you can flnd two distributions aandbon the nonnegative integers such that the convolution of 2See M. Krasner and B. Ranulae, \Sur une Propriet¶ e des Polynomes de la Division du Circle"; and the following note by J. Hadamard, in C. R. Acad. Sci., vol. 204 (1937), pp. 397{399. 7.2. SUMS OF CONTINUOUS RANDOM VARIABLES 291 aandbis the equiprobable distribution on the set 0, 1, 2, ...,n¡1. Ifnis prime this is not possible, but the proof is not so easy. (Assume that neitheranorbis concentrated at 0.) 11Assume that you are playing craps with dice that are loaded in the following way: faces two, three, four, and flve all come up with the same probability(1=6) +r. Faces one and six come up with probability (1 =6)¡2r, with 0< r<: 02. Write a computer program to flnd the probability of winning at craps with these dice, and using your program flnd which values of rmake craps a favorable game for the player with these dice. 7.2 Sums of Continuous Random Variables In this section we consider the continuous version of the problem posed in theprevious section: How are sums of independent random variables distributed? Convolutions Deflnition 7.2 LetXandYbe two continuous random variables with density functionsf(x) andg(y), respectively. Assume that both f(x) andg(y) are deflned for all real numbers. Then the convolution f⁄goffandgis the function given by (f⁄g)(z)=Z+1 ¡1f(z¡y)g(y)dy =Z+1 ¡1g(z¡x)f(x)dx : 2 This deflnition is analogous to the deflnition, given in Section 7.1, of the con- volution of two distribution functions. Thus it should not be surprising that if X andYare independent, then the density of their sum is the convolution of their densities. This fact is stated as a theorem below, and its proof is left as an exercise(see Exercise 1). Theorem 7.1 LetXandYbe two independent random variables with density functionsf X(x) andfY(y) deflned for all x. Then the sum Z=X+Yis a random variable with density function fZ(z), wherefZis the convolution of fXandfY.2 To get a better understanding of this important result, we will look at some examples. 292 CHAPTER 7. SUMS OF RANDOM VARIABLES Sum of Two Independent Uniform Random Variables Example 7.3 Suppose we choose independently two numbers at random from the interval [0;1] with uniform probability density. What is the density of their sum? LetXandYbe random variables describing our choices and Z=X+Ytheir sum. Then we have fX(x)=fY(x)=‰ 1i f 0•x•1, 0 otherwise; and the density function for the sum is given by fZ(z)=Z+1 ¡1fX(z¡y)fY(y)dy : SincefY(y)=1i f0•y•1 and 0 otherwise, this becomes fZ(z)=Z1 0fX(z¡y)dy : Now the integrand is 0 unless 0 •z¡y•1 (i.e., unless z¡1•y•z) and then it is 1. So if 0•z•1, we have fZ(z)=Zz 0dy=z; while if 1<z•2, we have fZ(z)=Z1 z¡1dy=2¡z; and ifz<0o rz>2w eh a v efZ(z) = 0 (see Figure 7.2). Hence, fZ(z)=8 < :z; if 0•z•1; 2¡z;if 1<z•2; 0; otherwise. Note that this result agrees with that of Example 2.4. 2 Sum of Two Independent Exponential Random Variables Example 7.4 Suppose we choose two numbers at random from the interval [0 ;1) with an exponential density with parameter ‚. What is the density of their sum? LetX,Y, andZ=X+Ydenote the relevant random variables, and fX,fY, andfZtheir densities. Then fX(x)=fY(x)=‰ ‚e¡‚x;ifx‚0; 0; otherwise; 7.2. SUMS OF CONTINUOUS RANDOM VARIABLES 293 0.5 1 1.5 20.20.40.60.81 Figure 7.2: Convolution of two uniform densities. 1 2 3 4 5 60.050.10.150.20.250.30.35 Figure 7.3: Convolution of two exponential densities with ‚=1 . and so, ifz>0, fZ(z)=Z+1 ¡1fX(z¡y)fY(y)dy =Zz 0‚e¡‚(z¡y)‚e¡‚ydy =Zz 0‚2e¡‚zdy =‚2ze¡‚z; while ifz<0,fZ(z) = 0 (see Figure 7.3). Hence, fZ(z)=‰ ‚2ze¡‚z;ifz‚0; 0; otherwise. 2 294 CHAPTER 7. SUMS OF RANDOM VARIABLES Sum of Two Independent Normal Random Variables Example 7.5 It is an interesting and important fact that the convolution of two normal densities with means „1and„2and variances ¾1and¾2is again a normal density, with mean „1+„2and variance ¾2 1+¾2 2. We will show this in the special case that both random variables are standard normal. The general case can be donein the same way, but the calculation is messier. Another way to show the generalresult is given in Example 10.17. SupposeXandYare two independent random variables, each with the standard normal density (see Example 5.8). We have f X(x)=fY(y)=1p 2…e¡x2=2; and so fZ(z)=fX⁄fY(z) =1 2…Z+1 ¡1e¡(z¡y)2=2e¡y2=2dy =1 2…e¡z2=4Z+1 ¡1e¡(y¡z=2)2dy =1 2…e¡z2=4p…•1p…Z1 ¡1e¡(y¡z=2)2dy‚ : The expression in the brackets equals 1, since it is the integral of the normal density function with „= 0 and¾=p 2. So, we have fZ(z)=1p 4…e¡z2=4: 2 Sum of Two Independent Cauchy Random Variables Example 7.6 Choose two numbers at random from the interval ( ¡1;+1) with the Cauchy density with parameter a= 1 (see Example 5.10). Then fX(x)=fY(x)=1 …(1 +x2); andZ=X+Yhas density fZ(z)=1 …2Z+1 ¡11 1+(z¡y)21 1+y2dy : 7.2. SUMS OF CONTINUOUS RANDOM VARIABLES 295 This integral requires some efiort, and we give here only the result (see Section 10.3, or Dwass3): fZ(z)=2 …(4 +z2): Now, suppose that we ask for the density function of the average A=( 1=2)(X+Y) ofXandY. ThenA=( 1=2)Z. Exercise 5.2.19 shows that if UandVare two continuous random variables with density functions fU(x) andfV(x), respectively, and ifV=aU, then fV(x)=µ1 a¶ fUµx a¶ : Thus, we have fA(z)=2fZ(2z)=1 …(1 +z2): Hence, the density function for the average of two random variables, each having a Cauchy density, is again a random variable with a Cauchy density; this remarkableproperty is a peculiarity of the Cauchy density. One consequence of this is if theerror in a certain measurement process had a Cauchy density and you averageda number of measurements, the average could not be expected to be any moreaccurate than any one of your individual measurements! 2 Rayleigh Density Example 7.7 SupposeXandYare two independent standard normal random variables. Now suppose we locate a point Pin thexy-plane with coordinates ( X;Y ) and ask: What is the density of the square of the distance of Pfrom the origin? (We have already simulated this problem in Example 5.9.) Here, with the precedingnotation, we have f X(x)=fY(x)=1p 2…e¡x2=2: Moreover, if X2denotes the square of X, then (see Theorem 5.1 and the discussion following) fX2(r)=‰1 2pr(fX(pr)+fX(¡pr)) ifr>0; 0 otherwise. =‰1p 2…r(e¡r=2)i f r>0; 0 otherwise. 3M. Dwass, \On the Convolution of Cauchy Distributions," American Mathematical Monthly, vol. 92, no. 1, (1985), pp. 55{57; see also R. Nelson, letters to the Editor, ibid., p. 679. 296 CHAPTER 7. SUMS OF RANDOM VARIABLES This is a gamma density with ‚=1=2,fl=1=2 (see Example 7.4). Now let R2=X2+Y2. Then fR2(r)=Z+1 ¡1fX2(r¡s)fY2(s)ds =1 4…Z+1 ¡1e¡(r¡s)=2r¡s 2¡1=2 e¡ss 2¡1=2 ds ; =‰1 2e¡r2=2;ifr‚0; 0; otherwise. Hence,R2has a gamma density with ‚=1=2,fl= 1. We can interpret this result as giving the density for the square of the distance of Pfrom the center of a target if its coordinates are normally distributed. The density of the random variable Ris obtained from that of R2in the usual way (see Theorem 5.1), and we flnd fR(r)=‰1 2e¡r2=2¢2r=re¡r2=2;ifr‚0; 0; otherwise. Physicists will recognize this as a Rayleigh density. Our result here agrees with our simulation in Example 5.9. 2 Chi-Squared Density More generally, the same method shows that the sum of the squares of nindependent normally distributed random variables with mean 0 and standard deviation 1 hasa gamma density with ‚=1=2 andfl=n=2. Such a density is called a chi-squared density withndegrees of freedom. This density was introduced in Chapter 4.3. In Example 5.10, we used this density to test the hypothesis that two traits wereindependent. Another important use of the chi-squared density is in comparing experimental data with a theoretical discrete distribution, to see whether the data supports thetheoretical model. More speciflcally, suppose that we have an experiment with aflnite set of outcomes. If the set of outcomes is countable, we group them into flnitelymany sets of outcomes. We propose a theoretical distribution which we think willmodel the experiment well. We obtain some data by repeating the experiment anumber of times. Now we wish to check how well the theoretical distribution fltsthe data. LetXbe the random variable which represents a theoretical outcome in the model of the experiment, and let m(x) be the distribution function of X.I n a manner similar to what was done in Example 5.10, we calculate the value of theexpression V=X x(ox¡n¢m(x))2 n¢m(x); where the sum runs over all possible outcomes x,nis the number of data points, andoxdenotes the number of outcomes of type xobserved in the data. Then 7.2. SUMS OF CONTINUOUS RANDOM VARIABLES 297 Outcome Observed Frequency 1 15 2 8 3 7 4 5 5 7 6 18 Table 7.1: Observed data. for moderate or large values of n, the quantity Vis approximately chi-squared distributed, with ”¡1 degrees of freedom, where ”represents the number of possible outcomes. The proof of this is beyond the scope of this book, but we will illustratethe reasonableness of this statement in the next example. If the value of Vis very large, when compared with the appropriate chi-squared density function, then wewould tend to reject the hypothesis that the model is an appropriate one for theexperiment at hand. We now give an example of this procedure. Example 7.8 Suppose we are given a single die. We wish to test the hypothesis that the die is fair. Thus, our theoretical distribution is the uniform distribution onthe integers between 1 and 6. So, if we roll the die ntimes, the expected number of data points of each type is n=6. Thus, if o idenotes the actual number of data points of type i, for 1•i•6, then the expression V=6X i=1(oi¡n=6)2 n=6 is approximately chi-squared distributed with 5 degrees of freedom. Now suppose that we actually roll the die 60 times and obtain the data in Table 7.1. If we calculate Vfor this data, we obtain the value 13.6. The graph of the chi-squared density with 5 degrees of freedom is shown in Figure 7.4. One seesthat values as large as 13.6 are rarely taken on by Vif the die is fair, so we would reject the hypothesis that the die is fair. (When using this test, a statistician willreject the hypothesis if the data gives a value of Vwhich is larger than 95% of the values one would expect to obtain if the hypothesis is true.) In Figure 7.5, we show the results of rolling a die 60 times, then calculating V, and then repeating this experiment 1000 times. The program that performs thesecalculations is called DieTest . We have superimposed the chi-squared density with 5 degrees of freedom; one can see that the data values flt the curve fairly well, whichsupports the statement that the chi-squared density is the correct one to use. 2 So far we have looked at several important special cases for which the convolution integral can be evaluated explicitly. In general, the convolution of two continuousdensities cannot be evaluated explicitly, and we must resort to numerical methods.Fortunately, these prove to be remarkably efiective, at least for bounded densities. 298 CHAPTER 7. SUMS OF RANDOM VARIABLES 5 10 15 200.0250.050.0750.10.1250.15 Figure 7.4: Chi-squared density with 5 degrees of freedom. 0 5 10 15 20 25 3000.0250.050.0750.10.1250.151000 experiments 60 rolls per experiment Figure 7.5: Rolling a fair die. 7.2. SUMS OF CONTINUOUS RANDOM VARIABLES 299 1 2 3 4 5 6 7 800.20.40.60.81n = 2 n = 4 n = 6 n = 8 n = 10 Figure 7.6: Convolution of nuniform densities. Independent Trials We now consider brie°y the distribution of the sum of nindependent random vari- ables, all having the same density function. If X1,X2, ...,Xnare these random variables and Sn=X1+X2+¢¢¢+Xnis their sum, then we will have fSn(x)=(fX1⁄fX2⁄¢¢¢⁄fXn)(x); where the right-hand side is an n-fold convolution. It is possible to calculate this density for general values of nin certain simple cases. Example 7.9 Suppose the Xiare uniformly distributed on the interval [0 ;1]. Then fXi(x)=‰1;if 0•x•1; 0;otherwise, andfSn(x) is given by the formula4 fSn(x)=‰1 (n¡1)!P 0•j•x(¡1)j¡n j¢ (x¡j)n¡1;if 0<x<n; 0; otherwise. The density fSn(x) forn= 2, 4, 6, 8, 10 is shown in Figure 7.6. If theXiare distributed normally, with mean 0 and variance 1, then (cf. Exam- ple 7.5) fXi(x)=1p 2…e¡x2=2; 4J. B. Uspensky, Introduction to Mathematical Probability (New York: McGraw-Hill, 1937), p. 277. 300 CHAPTER 7. SUMS OF RANDOM VARIABLES -15 -10 -5 5 10 150.0250.050.0750.10.1250.150.175 n = 5 n = 10 n = 15 n = 20 n = 25 Figure 7.7: Convolution of nstandard normal densities. and fSn(x)=1p 2…ne¡x2=2n: Here the density fSnforn= 5, 10, 15, 20, 25 is shown in Figure 7.7. If theXiare all exponentially distributed, with mean 1 =‚, then fXi(x)=‚e¡‚x; and fSn(x)=‚e¡‚x(‚x)n¡1 (n¡1)!: In this case the density fSnforn= 2, 4, 6, 8, 10 is shown in Figure 7.8. 2 Exercises 1LetXandYbe independent real-valued random variables with density func- tionsfX(x) andfY(y), respectively. Show that the density function of the sumX+Yis the convolution of the functions fX(x) andfY(y).Hint: Let „X be the joint random variable ( X;Y ). Then the joint density function of „Xis fX(x)fY(y), sinceXandYare independent. Now compute the probability thatX+Y•z, by integrating the joint density function over the appropriate region in the plane. This gives the cumulative distribution function of Z.N o w difierentiate this function with respect to zto obtain the density function of z. 2LetXandYbe independent random variables deflned on the space ›, with density functions fXandfY, respectively. Suppose that Z=X+Y. Find the density fZofZif 7.2. SUMS OF CONTINUOUS RANDOM VARIABLES 301 5 10 15 200.050.10.150.20.250.30.35n = 2 n = 4 n = 6 n = 8 n = 10 Figure 7.8: Convolution of nexponential densities with ‚=1 . (a) fX(x)=fY(x)=‰1=2;if¡1•x•+1; 0; otherwise. (b) fX(x)=fY(x)=‰1=2;if 3•x•5; 0; otherwise. (c) fX(x)=‰1=2;if¡1•x•1; 0; otherwise. fY(x)=‰1=2;if 3•x•5; 0; otherwise. (d) What can you say about the set E=fz:fZ(z)>0gin each case? 3Suppose again that Z=X+Y. FindfZif (a) fX(x)=fY(x)=‰ x=2;if 0<x< 2; 0; otherwise: (b) fX(x)=fY(x)=‰ (1=2)(x¡3);if 3<x< 5; 0; otherwise: (c) fX(x)=‰1=2;if 0<x< 2; 0; otherwise; 302 CHAPTER 7. SUMS OF RANDOM VARIABLES fY(x)=‰ x=2;if 0<x< 2; 0; otherwise: (d) What can you say about the set E=fz:fZ(z)>0gin each case? 4LetX,Y, andZbe independent random variables with fX(x)=fY(x)=fZ(x)=‰ 1;if 0<x< 1; 0;otherwise. Suppose that W=X+Y+Z. FindfWdirectly, and compare your answer with that given by the formula in Example 7.9. Hint: See Example 7.3. 5Suppose that XandYare independent and Z=X+Y. FindfZif (a) fX(x)=‰ ‚e¡‚x;ifx>0; 0; otherwise. fY(x)=‰„e¡„x;ifx>0; 0; otherwise. (b) fX(x)=‰ ‚e¡‚x;ifx>0; 0; otherwise. fY(x)=‰1;if 0<x< 1; 0;otherwise. 6Suppose again that Z=X+Y. FindfZif fX(x)=1p 2…¾1e¡(x¡„1)2=2¾2 1 fY(x)=1p 2…¾2e¡(x¡„2)2=2¾2 2: *7Suppose that R2=X2+Y2. FindfR2andfRif fX(x)=1p 2…¾1e¡(x¡„1)2=2¾2 1 fY(x)=1p 2…¾2e¡(x¡„2)2=2¾2 2: 8Suppose that R2=X2+Y2. FindfR2andfRif fX(x)=fY(x)=‰ 1=2;if¡1•x•1; 0; otherwise. 9Assume that the service time for a customer at a bank is exponentially dis- tributed with mean service time 2 minutes. Let Xbe the total service time for 10 customers. Estimate the probability that X> 22 minutes. 7.2. SUMS OF CONTINUOUS RANDOM VARIABLES 303 10LetX1,X2, ...,Xnbenindependent random variables each of which has an exponential density with mean „. LetMbe the minimum value of the Xj. Show that the density for Mis exponential with mean „=n.Hint: Use cumulative distribution functions. 11A company buys 100 lightbulbs, each of which has an exponential lifetime of 1000 hours. What is the expected time for the flrst of these bulbs to burnout? (See Exercise 10.) 12An insurance company assumes that the time between claims from each of its homeowners’ policies is exponentially distributed with mean „. It would like to estimate „by averaging the times for a number of policies, but this is not very practical since the time between claims is about 30 years. At Galambos’ 5 suggestion the company puts its customers in groups of 50 and observes thetime of the flrst claim within each group. Show that this provides a practicalway to estimate the value of „. 13Particles are subject to collisions that cause them to split into two parts with each part a fraction of the parent. Suppose that this fraction is uniformlydistributed between 0 and 1. Following a single particle through several split-tings we obtain a fraction of the original particle Z n=X1¢X2¢:::¢Xnwhere eachXjis uniformly distributed between 0 and 1. Show that the density for the random variable Znis fn(z)=1 (n¡1)!(¡logz)n¡1: Hint: Show that Yk=¡logXkis exponentially distributed. Use this to flnd the density function for Sn=Y1+Y2+¢¢¢+Yn, and from this the cumulative distribution and density of Zn=e¡Sn. 14Assume that X1andX2are independent random variables, each having an exponential density with parameter ‚. Show that Z=X1¡X2has density fZ(z)=( 1=2)‚e¡‚jzj: 15Suppose we want to test a coin for fairness. We °ip the coin ntimes and record the number of times X0that the coin turns up tails and the number of timesX1=n¡X0that the coin turns up heads. Now we set Z=1X i=0(Xi¡n=2)2 n=2: Then for a fair coin Zhas approximately a chi-squared distribution with 2¡1 = 1 degree of freedom. Verify this by computer simulation flrst for a fair coin (p=1=2) and then for a biased coin ( p=1=3). 5J. Galambos, Introductory Probability Theory (New York: Marcel Dekker, 1984), p. 159. 304 CHAPTER 7. SUMS OF RANDOM VARIABLES 16Verify your answers in Exercise 2(a) by computer simulation: Choose Xand Yfrom [¡1;1] with uniform density and calculate Z=X+Y. Repeat this experiment 500 times, recording the outcomes in a bar graph on [ ¡2;2] with 40 bars. Does the density fZcalculated in Exercise 2(a) describe the shape of your bar graph? Try this for Exercises 2(b) and Exercise 2(c), too. 17Verify your answers to Exercise 3 by computer simulation. 18Verify your answer to Exercise 4 by computer simulation. 19The support of a function f(x) is deflned to be the set fx:f(x)>0g: Suppose that XandYare two continuous random variables with density functionsfX(x) andfY(y), respectively, and suppose that the supports of these density functions are the intervals [ a;b] and [c;d], respectively. Find the support of the density function of the random variable X+Y. 20LetX1,X2,...,Xnbe a sequence of independent random variables, all having a common density function fXwith support [ a;b] (see Exercise 19). Let Sn=X1+X2+¢¢¢+Xn, with density function fSn. Show that the support offSnis the interval [ na;nb ].Hint: WritefSn=fSn¡1⁄fX. Now use Exercise 19 to establish the desired result by induction. 21LetX1,X2,...,Xnbe a sequence of independent random variables, all having a common density function fX. LetA=Sn=nbe their average. Find fAif (a)fX(x)=( 1=p 2…)e¡x2=2(normal density). (b)fX(x)=e¡x(exponential density). Hint: WritefA(x) in terms of fSn(x). Chapter 8 Law of Large Numbers 8.1 Law of Large Numbers for Discrete Random Variables We are now in a position to prove our flrst fundamental theorem of probability. We have seen that an intuitive way to view the probability of a certain outcomeis as the frequency with which that outcome occurs in the long run, when the ex-periment is repeated a large number of times. We have also deflned probabilitymathematically as a value of a distribution function for the random variable rep-resenting the experiment. The Law of Large Numbers, which is a theorem provedabout the mathematical model of probability, shows that this model is consistentwith the frequency interpretation of probability. This theorem is sometimes calledthelaw of averages. To flnd out what would happen if this law were not true, see the article by Robert M. Coates. 1 Chebyshev Inequality To discuss the Law of Large Numbers, we flrst need an important inequality calledtheChebyshev Inequality. Theorem 8.1 (Chebyshev Inequality) LetXbe a discrete random variable with expected value „=E(X), and let†>0 be any positive real number. Then P(jX¡„j‚†)•V(X) †2: Proof. Letm(x) denote the distribution function of X. Then the probability that Xdifiers from „by at least †is given by P(jX¡„j‚†)=X jx¡„j‚†m(x): 1R. M. Coates, \The Law," The World of Mathematics, ed. James R. Newman (New York: Simon and Schuster, 1956. 305 306 CHAPTER 8. LAW OF LARGE NUMBERS We know that V(X)=X x(x¡„)2m(x); and this is clearly at least as large as X jx¡„j‚†(x¡„)2m(x); since all the summands are positive and we have restricted the range of summation in the second sum. But this last sum is at least X jx¡„j‚††2m(x)=†2X jx¡„j‚†m(x) =†2P(jX¡„j‚†): So, P(jX¡„j‚†)•V(X) †2: 2 Note thatXin the above theorem can be any discrete random variable, and †any positive number. Example 8.1 LetXby any random variable with E(X)=„andV(X)=¾2. Then, if†=k¾, Chebyshev’s Inequality states that P(jX¡„j‚k¾)•¾2 k2¾2=1 k2: Thus, for any random variable, the probability of a deviation from the mean of more thankstandard deviations is •1=k2. If, for example, k=5 ,1=k2=:04. 2 Chebyshev’s Inequality is the best possible inequality in the sense that, for any †>0, it is possible to give an example of a random variable for which Chebyshev’s Inequality is in fact an equality. To see this, given †>0, chooseXwith distribution pX=µ¡†+† 1=21=2¶ : ThenE(X)=0 ,V(X)=†2, and P(jX¡„j‚†)=V(X) †2=1: We are now prepared to state and prove the Law of Large Numbers. 8.1. DISCRETE RANDOM VARIABLES 307 Law of Large Numbers Theorem 8.2 (Law of Large Numbers) LetX1,X2,...,Xnbe an independent trials process, with flnite expected value „=E(Xj) and flnite variance ¾2=V(Xj). LetSn=X1+X2+¢¢¢+Xn. Then for any †>0, PµflflflflS n n¡„flflflfl‚†¶ !0 asn!1 . Equivalently, PµflflflflS n n¡„flflflfl<†¶ !1 asn!1 . Proof. SinceX 1,X2,...,Xnare independent and have the same distributions, we can apply Theorem 6.9. We obtain V(Sn)=n¾2; and V(Sn n)=¾2 n: Also we know that E(Sn n)=„: By Chebyshev’s Inequality, for any †>0, PµflflflflS n n¡„flflflfl‚†¶ •¾ 2 n†2: Thus, for flxed †, PµflflflflS n n¡„flflflfl‚†¶ !0 asn!1 , or equivalently, PµflflflflS n n¡„flflflfl<†¶ !1 asn!1 . 2 Law of Averages Note thatSn=nis an average of the individual outcomes, and one often calls the Law of Large Numbers the \law of averages." It is a striking fact that we can start witha random experiment about which little can be predicted and, by taking averages,obtain an experiment in which the outcome can be predicted with a high degreeof certainty. The Law of Large Numbers, as we have stated it, is often called the\Weak Law of Large Numbers" to distinguish it from the \Strong Law of LargeNumbers" described in Exercise 15. 308 CHAPTER 8. LAW OF LARGE NUMBERS Consider the important special case of Bernoulli trials with probability pfor success. Let Xj= 1 if thejth outcome is a success and 0 if it is a failure. Then Sn=X1+X2+¢¢¢+Xnis the number of successes in ntrials and„=E(X1)=p. The Law of Large Numbers states that for any †>0 PµflflflflS n n¡pflflflfl<†¶ !1 asn!1 . The above statement says that, in a large number of repetitions of a Bernoulli experiment, we can expect the proportion of times the event will occur tobe nearp. This shows that our mathematical model of probability agrees with our frequency interpretation of probability. Coin Tossing Let us consider the special case of tossing a coin ntimes with Snthe number of heads that turn up. Then the random variable Sn=nrepresents the fraction of times heads turns up and will have values between 0 and 1. The Law of Large Numberspredicts that the outcomes for this random variable will, for large n, be near 1/2. In Figure 8.1, we have plotted the distribution for this example for increasing values ofn. We have marked the outcomes between .45 and .55 by dots at the top of the spikes. We see that as nincreases the distribution gets more and more con- centrated around .5 and a larger and larger percentage of the total area is containedwithin the interval ( :45;:55), as predicted by the Law of Large Numbers. Die Rolling Example 8.2 Considernrolls of a die. Let Xjbe the outcome of the jth roll. ThenSn=X1+X2+¢¢¢+Xnis the sum of the flrst nrolls. This is an independent trials process with E(Xj)=7=2. Thus, by the Law of Large Numbers, for any †>0 PµflflflflS n n¡7 2flflflfl‚†¶ !0 asn!1 . An equivalent way to state this is that, for any †>0, PµflflflflS n n¡7 2flflflfl<†¶ !1 asn!1 . 2 Numerical Comparisons It should be emphasized that, although Chebyshev’s Inequality proves the Law of Large Numbers, it is actually a very crude inequality for the probabilities involved.However, its strength lies in the fact that it is true for any random variable at all,and it allows us to prove a very powerful theorem. In the following example, we compare the estimates given by Chebyshev’s In- equality with the actual values. 8.1. DISCRETE RANDOM VARIABLES 309 0 0.2 0.4 0.6 0.8 100.020.040.060.080.1 0 0.2 0.4 0.6 0.8 100.020.040.060.080 0.2 0.4 0.6 0.8 100.020.040.060.080.10.120.14 0 0.2 0.4 0.6 0.8 100.020.040.060.080.10.1200.2 0.4 0.6 0.8 100.050.10.150.20.25 0 0.2 0.4 0.6 0.8 100.0250.050.0750.10.1250.150.175n=10 n=20 n=40 n=30 n=60 n=100 Figure 8.1: Bernoulli trials distributions. 310 CHAPTER 8. LAW OF LARGE NUMBERS Example 8.3 LetX1,X2,...,Xnbe a Bernoulli trials process with probability .3 for success and .7 for failure. Let Xj= 1 if the jth outcome is a success and 0 otherwise. Then, E(Xj)=:3 andV(Xj)=(:3)(:7) =:21. If An=Sn n=X1+X2+¢¢¢+Xn n is the average of theXi, thenE(An)=:3 andV(An)=V(Sn)=n2=:21=n. Chebyshev’s Inequality states that if, for example, †=:1, P(jAn¡:3j‚:1)•:21 n(:1)2=21 n: Thus, ifn= 100, P(jA100¡:3j‚:1)•:21; or ifn= 1000, P(jA1000¡:3j‚:1)•:021: These can be rewritten as P(:2<A 100<:4)‚:79; P(:2<A 1000<:4)‚:979: These values should be compared with the actual values, which are (to six decimal places) P(:2<A 100<:4)…:962549 P(:2<A 1000<:4)…1: The program Law can be used to carry out the above calculations in a systematic way. 2 Historical Remarks The Law of Large Numbers was flrst proved by the Swiss mathematician James Bernoulli in the fourth part of his work Ars Conjectandi published posthumously in 1713.2As often happens with a flrst proof, Bernoulli’s proof was much more di–cult than the proof we have presented using Chebyshev’s inequality. Cheby-shev developed his inequality to prove a general form of the Law of Large Numbers(see Exercise 12). The inequality itself appeared much earlier in a work by Bien-aym¶e, and in discussing its history Maistrov remarks that it was referred to as the Bienaym¶ e-Chebyshev Inequality for a long time. 3 InArs Conjectandi Bernoulli provides his reader with a long discussion of the meaning of his theorem with lots of examples. In modern notation he has an event 2J. Bernoulli, The Art of Conjecturing IV, trans. Bing Sung, Technical Report No. 2, Dept. of Statistics, Harvard Univ., 1966 3L. E. Maistrov, Probability Theory: A Historical Approach, trans. and ed. Samual Kotz, (New York: Academic Press, 1974), p. 202 8.1. DISCRETE RANDOM VARIABLES 311 that occurs with probability pbut he does not know p. He wants to estimate p by the fraction „ pof the times the event occurs when the experiment is repeated a number of times. He discusses in detail the problem of estimating, by this method,the proportion of white balls in an urn that contains an unknown number of whiteand black balls. He would do this by drawing a sequence of balls from the urn,replacing the ball drawn after each draw, and estimating the unknown proportionof white balls in the urn by the proportion of the balls drawn that are white. Heshows that, by choosing nlarge enough he can obtain any desired accuracy and reliability for the estimate. He also provides a lively discussion of the applicabilityof his theorem to estimating the probability of dying of a particular disease, ofdifierent kinds of weather occurring, and so forth. In speaking of the number of trials necessary for making a judgement, Bernoulli observes that the \man on the street" believes the \law of averages." Further, it cannot escape anyone that for judging in this way about any event at all, it is not enough to use one or two trials, but rather a greatnumber of trials is required. And sometimes the stupidest man|bysome instinct of nature per se and by no previous instruction (this is truly amazing)| knows for sure that the more observations of this sortthat are taken, the less the danger will be of straying from the mark. 4 But he goes on to say that he must contemplate another possibility. Something futher must be contemplated here which perhaps no one has thought about till now. It certainly remains to be inquired whetherafter the number of observations has been increased, the probability isincreased of attaining the true ratio between the number of cases inwhich some event can happen and in which it cannot happen, so thatthis probability flnally exceeds any given degree of certainty; or whetherthe problem has, so to speak, its own asymptote|that is, whether somedegree of certainty is given which one can never exceed. 5 Bernoulli recognized the importance of this theorem, writing: Therefore, this is the problem which I now set forth and make known after I have already pondered over it for twenty years. Both its noveltyand its very great usefullness, coupled with its just as great di–culty,can exceed in weight and value all the remaining chapters of this thesis. 6 Bernoulli concludes his long proof with the remark: Whence, flnally, this one thing seems to follow: that if observations of all events were to be continued throughout all eternity, (and hence theultimate probability would tend toward perfect certainty), everything in 4Bernoulli, op. cit., p. 38. 5ibid., p. 39. 6ibid., p. 42. 312 CHAPTER 8. LAW OF LARGE NUMBERS the world would be perceived to happen in flxed ratios and according to a constant law of alternation, so that even in the most accidental andfortuitous occurrences we would be bound to recognize, as it were, acertain necessity and, so to speak, a certain fate. I do now know whether Plato wished to aim at this in his doctrine of the universal return of things, according to which he predicted that allthings will return to their original state after countless ages have past. 7 Exercises 1A fair coin is tossed 100 times. The expected number of heads is 50, and the standard deviation for the number of heads is (100 ¢1=2¢1=2)1=2= 5. What does Chebyshev’s Inequality tell you about the probability that the numberof heads that turn up deviates from the expected number 50 by three or morestandard deviations (i.e., by at least 15)? 2Write a program that uses the function binomial( n;p;x ) to compute the exact probability that you estimated in Exercise 1. Compare the two results. 3Write a program to toss a coin 10,000 times. Let S nbe the number of heads in the flrst ntosses. Have your program print out, after every 1000 tosses, Sn¡n=2. On the basis of this simulation, is it correct to say that you can expect heads about half of the time when you toss a coin a large number oftimes? 4A 1-dollar bet on craps has an expected winning of ¡:0141. What does the Law of Large Numbers say about your winnings if you make a large numberof 1-dollar bets at the craps table? Does it assure you that your losses will besmall? Does it assure you that if nis very large you will lose? 5LetXbe a random variable with E(X)=0a n dV(X) = 1. What integer valuekwill assure us that P(jXj‚k)•:01? 6LetS nbe the number of successes in nBernoulli trials with probability pfor success on each trial. Show, using Chebyshev’s Inequality, that for any †>0 PµflflflflS n n¡pflflflfl‚†¶ •p(1¡p) n†2: 7Find the maximum possible value for p(1¡p)i f0<p< 1. Using this result and Exercise 6, show that the estimate PµflflflflS n n¡pflflflfl‚†¶ •1 4n†2 is valid for any p. 7ibid., pp. 65{66. 8.1. DISCRETE RANDOM VARIABLES 313 8A fair coin is tossed a large number of times. Does the Law of Large Numbers assure us that, if nis large enough, with probability >:99 the number of heads that turn up will not deviate from n=2 by more than 100? 9In Exercise 6.2.15, you showed that, for the hat check problem, the number Snof people who get their own hats back has E(Sn)=V(Sn) = 1. Using Chebyshev’s Inequality, show that P(Sn‚11)•:01 for anyn‚11. 10LetXby any random variable which takes on values 0, 1, 2, ...,nand has E(X)=V(X) = 1. Show that, for any integer k, P(X‚k+1 )•1 k2: 11We have two coins: one is a fair coin and the other is a coin that produces heads with probability 3/4. One of the two coins is picked at random, and thiscoin is tossed ntimes. LetS nbe the number of heads that turns up in these ntosses. Does the Law of Large Numbers allow us to predict the proportion of heads that will turn up in the long run? After we have observed a largenumber of tosses, can we tell which coin was chosen? How many tosses su–ceto make us 95 percent sure? 12(Chebyshev 8) Assume that X1,X2,...,Xnare independent random variables with possibly difierent distributions and let Snbe their sum. Let mk=E(Xk), ¾2 k=V(Xk), andMn=m1+m2+¢¢¢+mn. Assume that ¾2 k<R for allk. Prove that, for any †>0, PµflflflflS n n¡Mn nflflflfl<†¶ !1 asn!1 . 13A fair coin is tossed repeatedly. Before each toss, you are allowed to decide whether to bet on the outcome. Can you describe a betting system withinflnitely many bets which will enable you, in the long run, to win morethan half of your bets? (Note that we are disallowing a betting system thatsays to bet until you are ahead, then quit.) Write a computer program thatimplements this betting system. As stated above, your program must decidewhether to bet on a particular outcome before that outcome is determined.For example, you might select only outcomes that come after there have beenthree tails in a row. See if you can get more than 50% heads by your \system." *14 Prove the following analogue of Chebyshev’s Inequality: P(jX¡E(X)j‚†)•1 †E(jX¡E(X)j): 8P. L. Chebyshev, \On Mean Values," J. Math. Pure. Appl., vol. 12 (1867), pp. 177{184. 314 CHAPTER 8. LAW OF LARGE NUMBERS *15 We have proved a theorem often called the \Weak Law of Large Numbers." Most people’s intuition and our computer simulations suggest that, if we tossa coin a sequence of times, the proportion of heads will really approach 1/2;that is, ifS nis the number of heads in ntimes, then we will have An=Sn n!1 2 asn!1 . Of course, we cannot be sure of this since we are not able to toss the coin an inflnite number of times, and, if we could, the coin could come upheads every time. However, the \Strong Law of Large Numbers," proved inmore advanced courses, states that PµS n n!1 2¶ =1: Describe a sample space › that would make it possible for us to talk about the event E=‰ !:Sn n!1 2¾ : Could we assign the equiprobable measure to this space? (See Example 2.18.) *16 In this problem, you will construct a sequence of random variables which satisfles the Weak Law of Large Numbers, but not the Strong Law of LargeNumbers (see Exercise 15). For each positive integer n, let the random variable X nbe deflned by P(Xn=§n2n)=f(n); P(Xn= 0 )=1¡2f(n); wheref(n) is a function that will be chosen later (and which satisfles 0 • f(n)•1=2 for all positive integers n). LetSn=X1+X2+¢¢¢+Xn. (a) Show that „(Sn) = 0 for all n. (b) Show that if Xn>0, thenSn‚2n. (c) Use part (b) to show that Sn=n!0a sn!1 if and only if there exists ann0such thatXk= 0 for all k‚n0. Show that this happens with probability 0 if we require that f(n)<1=2 for alln. This shows that the sequencefXngdoes not satisfy the Strong Law of Large Numbers. (d) We now turn our attention to the Weak Law of Large Numbers. Given a positive†, we wish to estimate PˆflflflflS n nflflflfl‚†! : Suppose that X k= 0 form<k•n. Show that jSnj•22m: 8.1. DISCRETE RANDOM VARIABLES 315 (e) Show that if we deflne g(n)=( 1=2) log2(†n), then 22m<†n: This shows that if Xk= 0 forg(n)<k•n, then jSnj<†n; or flflflflS n nflflflfl<†: We wish to show that the probability of this event tends to 1 as n!1 , or equivalently, that the probability of the complementary event tendst o0a sn!1 . The complementary event is the event that X k6=0 for somekwithg(n)<k•n. Show that the probability of this event equals 1¡nY k=dg(n)e¡ 1¡2f(n)¢ ; and show that this expression is less than 1¡1Y k=dg(n)e¡ 1¡2f(n)¢ : (f) Show that by making f(n)!0 rapidly enough, the expression in part (e) can be made to approach 1 as n!1 . This shows that the sequence fXngsatisfles the Weak Law of Large Numbers. *17 Let us toss a biased coin that comes up heads with probability pand assume the validity of the Strong Law of Large Numbers as described in Exercise 15.Then, with probability 1, S n n!p asn!1 .I ff(x) is a continuous function on the unit interval, then we also have fµSn n¶ !f(p): Finally, we could hope that Eµ fµSn n¶¶ !E(f(p)) =f(p): Show that, if all this is correct, as in fact it is, we would have proven that any continuous function on the unit interval is a limit of polynomial func-tions. This is a sketch of a probabilistic proof of an important theorem inmathematics called the Weierstrass approximation theorem. 316 CHAPTER 8. LAW OF LARGE NUMBERS 8.2 Law of Large Numbers for Continuous Ran- dom Variables In the previous section we discussed in some detail the Law of Large Numbers for discrete probability distributions. This law has a natural analogue for continuousprobability distributions, which we consider somewhat more brie°y here. Chebyshev Inequality Just as in the discrete case, we begin our discussion with the Chebyshev Inequality. Theorem 8.3 (Chebyshev Inequality) LetXbe a continuous random variable with density function f(x). Suppose Xhas a flnite expected value „=E(X) and flnite variance ¾2=V(X). Then for any positive number †>0w eh a v e P(jX¡„j‚†)•¾2 †2: 2 The proof is completely analogous to the proof in the discrete case, and we omit it. Note that this theorem says nothing if ¾2=V(X) is inflnite. Example 8.4 LetXbe any continuous random variable with E(X)=„and V(X)=¾2. Then, if†=k¾=kstandard deviations for some integer k, then P(jX¡„j‚k¾)•¾2 k2¾2=1 k2; just as in the discrete case. 2 Law of Large Numbers With the Chebyshev Inequality we can now state and prove the Law of Large Numbers for the continuous case. Theorem 8.4 (Law of Large Numbers) LetX1,X2,...,Xnbe an independent trials process with a continuous density function f, flnite expected value „, and flnite variance¾2. LetSn=X1+X2+¢¢¢+Xnbe the sum of the Xi. Then for any real number†>0w eh a v e lim n!1PµflflflflS n n¡„flflflfl‚†¶ =0; or equivalently, lim n!1PµflflflflS n n¡„flflflfl<†¶ =1: 2 8.2. CONTINUOUS RANDOM VARIABLES 317 Note that this theorem is not necessarily true if ¾2is inflnite (see Example 8.8). As in the discrete case, the Law of Large Numbers says that the average value ofnindependent trials tends to the expected value as n!1 , in the precise sense that, given †>0, the probability that the average value and the expected value difier by more than †tends to 0 as n!1 . Once again, we suppress the proof, as it is identical to the proof in the discrete case. Uniform Case Example 8.5 Suppose we choose at random nnumbers from the interval [0 ;1] with uniform distribution. Then if Xidescribes the ith choice, we have „=E(Xi)=Z1 0xdx =1 2; ¾2=V(Xi)=Z1 0x2dx¡„2 =1 3¡1 4=1 12: Hence, EµSn n¶ =1 2; VµSn n¶ =1 12n; and for any †>0, PµflflflflS n n¡1 2flflflfl‚†¶ •1 12n†2: This says that if we choose nnumbers at random from [0 ;1], then the chances are better than 1 ¡1=(12n†2) that the difierence jSn=n¡1=2jis less than †. Note that†plays the role of the amount of error we are willing to tolerate: If we choose †=0:1, say, then the chances that jSn=n¡1=2jis less than 0.1 are better than 1¡100=(12n). Forn= 100, this is about .92, but if n= 1000, this is better than .99 and ifn=1 0;000, this is better than .999. We can illustrate what the Law of Large Numbers says for this example graph- ically. The density for An=Sn=nis determined by fAn(x)=nfSn(nx): We have seen in Section 7.2, that we can compute the density fSn(x) for the sum ofnuniform random variables. In Figure 8.2 we have used this to plot the density for Anfor various values of n. We have shaded in the area for which An would lie between .45 and .55. We see that as we increase n, we obtain more and more of the total area inside the shaded region. The Law of Large Numbers tells usthat we can obtain as much of the total area as we please inside the shaded regionby choosing nlarge enough (see also Figure 8.1). 2 318 CHAPTER 8. LAW OF LARGE NUMBERS n=2 n=5 n=10 n=20 n=30 n=50 Figure 8.2: Illustration of Law of Large Numbers | uniform case. Normal Case Example 8.6 Suppose we choose nreal numbers at random, using a normal dis- tribution with mean 0 and variance 1. Then „=E(Xi)=0; ¾2=V(Xi)=1: Hence, EµSn n¶ =0; VµSn n¶ =1 n; and, for any †>0, PµflflflflS n n¡0flflflfl‚†¶ •1 n†2: In this case it is possible to compare the Chebyshev estimate for P(jSn=n¡„j‚†) in the Law of Large Numbers with exact values, since we know the density functionforS n=nexactly (see Example 7.9). The comparison is shown in Table 8.1, for †=:1. The data in this table was produced by the program LawContinuous .W e see here that the Chebyshev estimates are in general notvery accurate. 2 8.2. CONTINUOUS RANDOM VARIABLES 319 nP(jSn=nj‚:1) Chebyshev 100 .31731 1.00000 200 .15730 .50000 300 .08326 .33333 400 .04550 .25000 500 .02535 .20000 600 .01431 .16667 700 .00815 .14286 800 .00468 .12500 900 .00270 .11111 1000 .00157 .10000 Table 8.1: Chebyshev estimates. Monte Carlo Method Here is a somewhat more interesting example. Example 8.7 Letg(x) be a continuous function deflned for x2[0;1] with values in [0;1]. In Section 2.1, we showed how to estimate the area of the region under the graph of g(x) by the Monte Carlo method, that is, by choosing a large number of random values for xandywith uniform distribution and seeing what fraction of the pointsP(x;y) fell inside the region under the graph (see Example 2.2). Here is a better way to estimate the same area (see Figure 8.3). Let us choose a large number of independent values Xnat random from [0 ;1] with uniform density, setYn=g(Xn), and flnd the average value of the Yn. Then this average is our estimate for the area. To see this, note that if the density function for Xnis uniform, „=E(Yn)=Z1 0g(x)f(x)dx =Z1 0g(x)dx = average value of g(x); while the variance is ¾2=E((Yn¡„)2)=Z1 0(g(x)¡„)2dx< 1; since for all xin [0;1],g(x)i si n[ 0;1], hence„is in [0;1], and sojg(x)¡„j•1. Now letAn=( 1=n)(Y1+Y2+¢¢¢+Yn). Then by Chebyshev’s Inequality, we have P(jAn¡„j‚†)•¾2 n†2<1 n†2: This says that to get within †of the true value for „=R1 0g(x)dxwith probability at leastp, we should choose nso that 1=n†2•1¡p(i.e., so that n‚1=†2(1¡p)). Note that this method tells us how large to take nto get a desired accuracy. 2 320 CHAPTER 8. LAW OF LARGE NUMBERS Y XY = g (x) 01 1 Figure 8.3: Area problem. The Law of Large Numbers requires that the variance ¾2of the original under- lying density be flnite: ¾2<1. In cases where this fails to hold, the Law of Large Numbers may fail, too. An example follows. Cauchy Case Example 8.8 Suppose we choose nnumbers from (¡1;+1) with a Cauchy den- sity with parameter a= 1. We know that for the Cauchy density the expected value and variance are undeflned (see Example 6.28). In this case, the density functionfor A n=Sn n is given by (see Example 7.6) fAn(x)=1 …(1 +x2); that is, the density function for Anis the same for all n.In this case, as nincreases, the density function does not change at all, and the Law of Large Numbers doesnot hold. 2 Exercises 1LetXbe a continuous random variable with mean „= 10 and variance ¾2= 100=3. Using Chebyshev’s Inequality, flnd an upper bound for the following probabilities. 8.2. CONTINUOUS RANDOM VARIABLES 321 (a)P(jX¡10j‚2). (b)P(jX¡10j‚5). (c)P(jX¡10j‚9). (d)P(jX¡10j‚20). 2LetXbe a continuous random variable with values unformly distributed over the interval [0 ;20]. (a) Find the mean and variance of X. (b) Calculate P(jX¡10j‚2),P(jX¡10j‚5),P(jX¡10j‚9), and P(jX¡10j‚20) exactly. How do your answers compare with those of Exercise 1? How good is Chebyshev’s Inequality in this case? 3LetXbe the random variable of Exercise 2. (a) Calculate the function f(x)=P(jX¡10j‚x). (b) Now graph the function f(x), and on the same axes, graph the Chebyshev functiong(x) = 100=(3x2). Show that f(x)•g(x) for allx> 0, but thatg(x) is not a very good approximation for f(x). 4LetXbe a continuous random variable with values exponentially distributed over [0;1) with parameter ‚=0:1. (a) Find the mean and variance of X. (b) Using Chebyshev’s Inequality, flnd an upper bound for the following probabilities: P(jX¡10j‚2),P(jX¡10j‚5),P(jX¡10j‚9), and P(jX¡10j‚20). (c) Calculate these probabilities exactly, and compare with the bounds in (b). 5LetXbe a continuous random variable with values normally distributed over (¡1;+1) with mean „= 0 and variance ¾2=1 . (a) Using Chebyshev’s Inequality, flnd upper bounds for the following prob- abilities:P(jXj‚1),P(jXj‚2), andP(jXj‚3). (b) The area under the normal curve between ¡1 and 1 is .6827, between ¡2 and 2 is .9545, and between ¡3 and 3 it is .9973 (see the table in Appendix A). Compare your bounds in (a) with these exact values. Howgood is Chebyshev’s Inequality in this case? 6IfXis normally distributed, with mean „and variance ¾ 2, flnd an upper bound for the following probabilities, using Chebyshev’s Inequality. (a)P(jX¡„j‚¾). (b)P(jX¡„j‚2¾). (c)P(jX¡„j‚3¾). 322 CHAPTER 8. LAW OF LARGE NUMBERS (d)P(jX¡„j‚4¾). Now flnd the exact value using the program NormalArea or the normal table in Appendix A, and compare. 7IfXis a random variable with mean „6= 0 and variance ¾2, deflne the relative deviationDofXfrom its mean by D=flflflflX¡„ „flflflfl: (a) Show that P(D‚a)•¾ 2=(„2a2). (b) IfXis the random variable of Exercise 1, flnd an upper bound for P(D‚ :2),P(D‚:5),P(D‚:9), andP(D‚2). 8LetXbe a continuous random variable and deflne the standardized version X⁄ofXby: X⁄=X¡„ ¾: (a) Show that P(jX⁄j‚a)•1=a2. (b) IfXis the random variable of Exercise 1, flnd bounds for P(jX⁄j‚2), P(jX⁄j‚5), andP(jX⁄j‚9). 9(a) Suppose a number Xis chosen at random from [0 ;20] with uniform probability. Find a lower bound for the probability that Xlies between 8 and 12, using Chebyshev’s Inequality. (b) Now suppose 20 real numbers are chosen independently from [0 ;20] with uniform probability. Find a lower bound for the probability that theiraverage lies between 8 and 12. (c) Now suppose 100 real numbers are chosen independently from [0 ;20]. Find a lower bound for the probability that their average lies between8 and 12. 10A student’s score on a particular calculus flnal is a random variable with values of [0;100], mean 70, and variance 25. (a) Find a lower bound for the probability that the student’s score will fall between 65 and 75. (b) If 100 students take the flnal, flnd a lower bound for the probability that the class average will fall between 65 and 75. 11The Pilsdorfi beer company runs a °eet of trucks along the 100 mile road from Hangtown to Dry Gulch, and maintains a garage halfway in between. Eachof the trucks is apt to break down at a point Xmiles from Hangtown, where Xis a random variable uniformly distributed over [0 ;100]. (a) Find a lower bound for the probability P(jX¡50j•10). 8.2. CONTINUOUS RANDOM VARIABLES 323 (b) Suppose that in one bad week, 20 trucks break down. Find a lower bound for the probability P(jA20¡50j•10), where A20is the average of the distances from Hangtown at the time of breakdown. 12A share of common stock in the Pilsdorfi beer company has a price Ynon thenth business day of the year. Finn observes that the price change Xn= Yn+1¡Ynappears to be a random variable with mean „= 0 and variance ¾2=1=4. IfY1= 30, flnd a lower bound for the following probabilities, under the assumption that the Xn’s are mutually independent. (a)P(25•Y2•35). (b)P(25•Y11•35). (c)P(25•Y101•35). 13Suppose one hundred numbers X1,X2,...,X100are chosen independently at random from [0 ;20]. LetS=X1+X2+¢¢¢+X100be the sum, A=S=100 the average, and S⁄=(S¡1000)=(10=p 3) the standardized sum. Find lower bounds for the probabilities (a)P(jS¡1000j•100). (b)P(jA¡10j•1). (c)P(jS⁄j•p 3). 14LetXbe a continuous random variable normally distributed on ( ¡1;+1) with mean 0 and variance 1. Using the normal table provided in Appendix A,or the program NormalArea , flnd values for the function f(x)=P(jXj‚x) asxincreases from 0 to 4.0 in steps of .25. Note that for x‚0 the table gives NA(0;x)=P(0•X•x) and thus P(jXj‚x)=2 (:5¡NA(0;x). Plot by hand the graph of f(x) using these values, and the graph of the Chebyshev functiong(x)=1=x 2, and compare (see Exercise 3). 15Repeat Exercise 14, but this time with mean 10 and variance 3. Note that the table in Appendix A presents values for a standard normal variable. Findthe standardized version X ⁄forX, flnd values for f⁄(x)=P(jX⁄j‚x)a si n Exercise 14, and then rescale these values for f(x)=P(jX¡10j‚x). Graph and compare this function with the Chebyshev function g(x)=3=x2. 16LetZ=X=Y whereXandYhave normal densities with mean 0 and standard deviation 1. Then it can be shown that Zhas a Cauchy density. (a) Write a program to illustrate this result by plotting a bar graph of 1000 samples obtained by forming the ratio of two standard normal outcomes.Compare your bar graph with the graph of the Cauchy density. Depend-ing upon which computer language you use, you may or may not need totell the computer how to simulate a normal random variable. A methodfor doing this was described in Section 5.2. 324 CHAPTER 8. LAW OF LARGE NUMBERS (b) We have seen that the Law of Large Numbers does not apply to the Cauchy density (see Example 8.8). Simulate a large number of experi-ments with Cauchy density and compute the average of your results. Dothese averages seem to be approaching a limit? If so can you explainwhy this might be? 17Show that, if X‚0, thenP(X‚a)•E(X)=a. 18(Lamperti 9) LetXbe a non-negative random variable. What is the best upper bound you can give for P(X‚a) if you know (a)E(X) = 20. (b)E(X)=2 0a n d V(X) = 25. (c)E(X) = 20,V(X) = 25, and Xis symmetric about its mean. 9Private communication. Chapter 9 Central Limit Theorem 9.1 Central Limit Theorem for Bernoulli Trials The second fundamental theorem of probability is the Central Limit Theorem. This theorem says that if Snis the sum of nmutually independent random variables, then the distribution function of Snis well-approximated by a certain type of continuous function known as a normal density function, which is given by the formula f„;¾(x)=1p 2…¾e¡(x¡„)2=(2¾2); as we have seen in Chapter 4.3. In this section, we will deal only with the case that „= 0 and¾= 1. We will call this particular normal density function the standard normal density, and we will denote it by `(x): `(x)=1p 2…e¡x2=2: A graph of this function is given in Figure 9.1. It can be shown that the area under any normal density equals 1. The Central Limit Theorem tells us, quite generally, what happens when we have the sum of a large number of independent random variables each of which con-tributes a small amount to the total. In this section we shall discuss this theoremas it applies to the Bernoulli trials and in Section 9.2 we shall consider more generalprocesses. We will discuss the theorem in the case that the individual random vari-ables are identically distributed, but the theorem is true, under certain conditions,even if the individual random variables have difierent distributions. Bernoulli Trials Consider a Bernoulli trials process with probability pfor success on each trial. LetXi= 1 or 0 according as the ith outcome is a success or failure, and let Sn=X1+X2+¢¢¢+Xn. ThenSnis the number of successes in ntrials. We know thatSnhas as its distribution the binomial probabilities b(n;p;j ). In Section 3.2, 325 326 CHAPTER 9. CENTRAL LIMIT THEOREM -4 -2 0 2 400.10.20.30.4 Figure 9.1: Standard normal density. we plotted these distributions for p=:3 andp=:5 for various values of n(see Figure 3.5). We note that the maximum values of the distributions appeared near the ex- pected value np, which causes their spike graphs to drift ofi to the right as nin- creased. Moreover, these maximum values approach 0 as nincreased, which causes the spike graphs to °atten out. Standardized Sums We can prevent the drifting of these spike graphs by subtracting the expected num-ber of successes npfromS n, obtaining the new random variable Sn¡np. Now the maximum values of the distributions will always be near 0. To prevent the spreading of these spike graphs, we can normalize Sn¡npto have variance 1 by dividing by its standard deviationpnpq(see Exercise 6.2.12 and Ex- ercise 6.2.16). Deflnition 9.1 The standardized sum ofSnis given by S⁄ n=Sn¡nppnpq: S⁄ nalways has expected value 0 and variance 1. 2 Suppose we plot a spike graph with the spikes placed at the possible values of S⁄ n: x0,x1, ...,xn, where xj=j¡nppnpq: (9.1) We make the height of the spike at xjequal to the distribution value b(n;p;j ). An example of this standardized spike graph, with n= 270 and p=:3, is shown in Figure 9.2. This graph is beautifully bell-shaped. We would like to flt a normaldensity to this spike graph. The obvious choice to try is the standard normal density,since it is centered at 0, just as the standardized spike graph is. In this flgure, we 9.1. BERNOULLI TRIALS 327 -4 -2 0 2 400.10.20.30.4 Figure 9.2: Normalized binomial distribution and standard normal density. have drawn this standard normal density. The reader will note that a horrible thing has occurred: Even though the shapes of the two graphs are the same, the heightsare quite difierent. If we want the two graphs to flt each other, we must modify one of them; we choose to modify the spike graph. Since the shapes of the two graphs look fairlyclose, we will attempt to modify the spike graph without changing its shape. Thereason for the difiering heights is that the sum of the heights of the spikes equals1, while the area under the standard normal density equals 1. If we were to draw acontinuous curve through the top of the spikes, and flnd the area under this curve,we see that we would obtain, approximately, the sum of the heights of the spikesmultiplied by the distance between consecutive spikes, which we will call †. Since the sum of the heights of the spikes equals one, the area under this curve would beapproximately †. Thus, to change the spike graph so that the area under this curve has value 1, we need only multiply the heights of the spikes by 1 =†. It is easy to see from Equation 9.1 that †=1 pnpq: In Figure 9.3 we show the standardized sum S⁄ nforn= 270 and p=:3, after correcting the heights, together with the standard normal density. (This flgure wasproduced with the program CLTBernoulliPlot .) The reader will note that the standard normal flts the height-corrected spike graph extremely well. In fact, oneversion of the Central Limit Theorem (see Theorem 9.1) says that as nincreases, the standard normal density will do an increasingly better job of approximatingthe height-corrected spike graphs corresponding to a Bernoulli trials process withnsummands. Let us flx a value xon thex-axis and let nbe a flxed positive integer. Then, using Equation 9.1, the point x jthat is closest to xhas a subscript jgiven by the 328 CHAPTER 9. CENTRAL LIMIT THEOREM -4 -2 0 2 400.10.20.30.4 Figure 9.3: Corrected spike graph with standard normal density. formula j=hnp+xpnpqi; wherehaimeans the integer nearest to a. Thus the height of the spike above xj will bepnpqb (n;p;j )=pnpqb (n;p;hnp+xjpnpqi): For largen, we have seen that the height of the spike is very close to the height of the normal density at x. This suggests the following theorem. Theorem 9.1 (Central Limit Theorem for Binomial Distributions) For the binomial distribution b(n;p;j )w eh a v e lim n!1pnpqb (n;p;hnp+xpnpqi)=`(x); where`(x) is the standard normal density. The proof of this theorem can be carried out using Stirling’s approximation from Section 3.1. We indicate this method of proof by considering the case x=0 . I n this case, the theorem states that lim n!1pnpqb (n;p;hnpi)=1p 2…=:3989::: : In order to simplify the calculation, we assume that npis an integer, so that hnpi= np. Then pnpqb (n;p;np )=pnpqpnpqnqn! (np)! (nq)!: Recall that Stirling’s formula (see Theorem 3.3) states that n!»p 2…nnne¡nasn!1: 9.1. BERNOULLI TRIALS 329 Using this, we have pnpqb (n;p;np )»pnpqpnpqnqp 2…nnne¡n p2…npp2…nq(np)np(nq)nqe¡npe¡nq; which simplifles to 1 =p 2…. 2 Approximating Binomial Distributions We can use Theorem 9.1 to flnd approximations for the values of binomial distri- bution functions. If we wish to flnd an approximation for b(n;p;j ), we set j=np+xpnpq and solve for x, obtaining x=j¡nppnpq: Theorem 9.1 then says thatpnpqb (n;p;j ) is approximately equal to `(x), so b(n;p;j )…`(x)pnpq =1pnpq`µj¡nppnpq¶ : Example 9.1 Let us estimate the probability of exactly 55 heads in 100 tosses of a coin. For this case np= 100¢1=2=5 0a n dpnpq=p 100¢1=2¢1=2=5 . T h u s x55= (55¡50)=5=1a n d P(S100= 55)»`(1) 5=1 5µ1p 2…e¡1=2¶ =:0484: To four decimal places, the actual value is .0485, and so the approximation is very good. 2 The program CLTBernoulliLocal illustrates this approximation for any choice ofn,p, andj. We have run this program for two examples. The flrst is the probability of exactly 50 heads in 100 tosses of a coin; the estimate is .0798, while theactual value, to four decimal places, is .0796. The second example is the probabilityof exactly eight sixes in 36 rolls of a die; here the estimate is .1093, while the actualvalue, to four decimal places, is .1196. 330 CHAPTER 9. CENTRAL LIMIT THEOREM The individual binomial probabilities tend to 0 as ntends to inflnity. In most applications we are not interested in the probability that a speciflc outcome occurs,but rather in the probability that the outcome lies in a given interval, say the interval[a;b]. In order to flnd this probability, we add the heights of the spike graphs for values ofjbetweenaandb. This is the same as asking for the probability that the standardized sum S ⁄ nlies between a⁄andb⁄, wherea⁄andb⁄are the standardized values ofaandb. But asntends to inflnity the sum of these areas could be expected to approach the area under the standard normal density between a⁄andb⁄. The Central Limit Theorem states that this does indeed happen. Theorem 9.2 (Central Limit Theorem for Bernoulli Trials) LetSnbe the number of successes in nBernoulli trials with probability pfor success, and let a andbbe two flxed real numbers. Deflne a⁄=a¡nppnpq and b⁄=b¡nppnpq: Then lim n!1P(a•Sn•b)=Zb⁄ a⁄`(x)dx : 2 This theorem can be proved by adding together the approximations to b(n;p;k ) given in Theorem 9.1.It is also a special case of the more general Central LimitTheorem (see Section 10.3). We know from calculus that the integral on the right side of this equation is equal to the area under the graph of the standard normal density `(x) between aandb. We denote this area by NA( a ⁄;b⁄). Unfortunately, there is no simple way to integrate the function e¡x2=2, and so we must either use a table of values or else a numerical integration program. (See Figure 9.4 for values of NA(0 ;z). A more extensive table is given in Appendix A.) It is clear from the symmetry of the standard normal density that areas such as that between¡2 and 3 can be found from this table by adding the area from 0 to 2 (same as that from ¡2 to 0) to the area from 0 to 3. Approximation of Binomial Probabilities Suppose that Snis binomially distributed with parameters nandp. We have seen that the above theorem shows how to estimate a probability of the form P(i•Sn•j); (9.2) whereiandjare integers between 0 and n. As we have seen, the binomial distri- bution can be represented as a spike graph, with spikes at the integers between 0andn, and with the height of the kth spike given by b(n;p;k ). For moderate-sized 9.1. BERNOULLI TRIALS 331 NA (0,z) = area of shaded region 0z z NA(z) z NA(z) z NA(z) z NA(z) .0 .0000 1.0 .3413 2.0 .4772 3.0 .4987 .1 .0398 1.1 .3643 2.1 .4821 3.1 .4990 .2 .0793 1.2 .3849 2.2 .4861 3.2 .4993 .3 .1179 1.3 .4032 2.3 .4893 3.3 .4995 .4 .1554 1.4 .4192 2.4 .4918 3.4 .4997.5 .1915 1.5 .4332 2.5 .4938 3.5 .4998.6 .2257 1.6 .4452 2.6 .4953 3.6 .4998.7 .2580 1.7 .4554 2.7 .4965 3.7 .4999.8 .2881 1.8 .4641 2.8 .4974 3.8 .4999.9 .3159 1.9 .4713 2.9 .4981 3.9 .5000 Figure 9.4: Table of values of NA(0 ;z), the normal area from 0 to z. 332 CHAPTER 9. CENTRAL LIMIT THEOREM values ofn, if we standardize this spike graph, and change the heights of its spikes, in the manner described above, the sum of the heights of the spikes is approximatedby the area under the standard normal density between i ⁄andj⁄. It turns out that a slightly more accurate approximation is afiorded by the area under the standardnormal density between the standardized values corresponding to ( i¡1=2) and (j+1=2); these values are i ⁄=i¡1=2¡nppnpq and j⁄=j+1=2¡nppnpq: Thus, P(i•Sn•j)…NAˆ i¡1 2¡nppnpq;j+1 2¡nppnpq! : We now illustrate this idea with some examples. Example 9.2 A coin is tossed 100 times. Estimate the probability that the number of heads lies between 40 and 60 (the word \between" in mathematics means inclusiveof the endpoints). The expected number of heads is 100 ¢1=2 = 50, and the standard deviation for the number of heads isp 100¢1=2¢1=2 = 5. Thus, since n= 100 is reasonably large, we have P(40•Sn•60)…Pµ39:5¡50 5•S⁄ n•60:5¡50 5¶ =P(¡2:1•S⁄ n•2:1) …NA(¡2:1;2:1) = 2NA(0;2:1) …:9642: The actual value is .96480, to flve decimal places. Note that in this case we are asking for the probability that the outcome will not deviate by more than two standard deviations from the expected value. Hadwe asked for the probability that the number of successes is between 35 and 65, thiswould have represented three standard deviations from the mean, and, using our1/2 correction, our estimate would be the area under the standard normal curvebetween¡3:1 and 3.1, or 2NA(0 ;3:1) =:9980. The actual answer in this case, to flve places, is .99821. 2 It is important to work a few problems by hand to understand the conversion from a given inequality to an inequality relating to the standardized variable. Afterthis, one can then use a computer program that carries out this conversion, includingthe 1/2 correction. The program CLTBernoulliGlobal is such a program for estimating probabilities of the form P(a•S n•b). 9.1. BERNOULLI TRIALS 333 Example 9.3 Dartmouth College would like to have 1050 freshmen. This college cannot accommodate more than 1060. Assume that each applicant accepts withprobability .6 and that the acceptances can be modeled by Bernoulli trials. If thecollege accepts 1700, what is the probability that it will have too many acceptances? If it accepts 1700 students, the expected number of students who matricu- late is:6¢1700 = 1020. The standard deviation for the number that accept isp 1700¢:6¢:4…20. Thus we want to estimate the probability P(S1700>1060) = P(S1700‚1061) =Pµ S⁄ 1700‚1060:5¡1020 20¶ =P(S⁄ 1700‚2:025): From Table 9.4, if we interpolate, we would estimate this probability to be :5¡:4784 =:0216. Thus, the college is fairly safe using this admission policy. 2 Applications to Statistics There are many important questions in the fleld of statistics that can be answered using the Central Limit Theorem for independent trials processes. The followingexample is one that is encountered quite frequently in the news. Another exampleof an application of the Central Limit Theorem to statistics is given in Section 9.2. Example 9.4 One frequently reads that a poll has been taken to estimate the proportion of people in a certain population who favor one candidate over anotherin a race with two candidates. (This model also applies to races with more thantwo candidates AandB, and to ballot propositions.) Clearly, it is not possible for pollsters to ask everyone for their preference. What is done instead is to pick asubset of the population, called a sample, and ask everyone in the sample for theirpreference. Let pbe the actual proportion of people in the population who are in favor of candidate Aand letq=1¡p. If we choose a sample of size nfrom the pop- ulation, the preferences of the people in the sample can be represented by randomvariablesX 1;X2;:::;Xn, whereXi= 1 if person iis in favor of candidate A, and Xi= 0 if person iis in favor of candidate B. LetSn=X1+X2+¢¢¢+Xn. If each subset of size nis chosen with the same probability, then Snis hypergeometrically distributed. If nis small relative to the size of the population (which is typically true in practice), then Snis approximately binomially distributed, with parameters nandp. The pollster wants to estimate the value p. An estimate for pis provided by the value „p=Sn=n, which is the proportion of people in the sample who favor candidate B. The Central Limit Theorem says that the random variable „ pis approximately normally distributed. (In fact, our version of the Central Limit Theorem says thatthe distribution function of the random variable S ⁄ n=Sn¡nppnpq 334 CHAPTER 9. CENTRAL LIMIT THEOREM is approximated by the standard normal density.) But we have „p=Sn¡nppnpqrpq n+p; i.e., „pis just a linear function of S⁄ n. Since the distribution of S⁄ nis approximated by the standard normal density, the distribution of the random variable „ pmust also be bell-shaped. We also know how to write the mean and standard deviation of „ p in terms of pandn. The mean of „ pis justp, and the standard deviation is rpq n: Thus, it is easy to write down the standardized version of „ p;i ti s „p⁄=„p¡pp pq=n: Since the distribution of the standardized version of „ pis approximated by the standard normal density, we know, for example, that 95% of its values will lie withintwo standard deviations of its mean, and the same is true of „ p.S ow eh a v e Pµ p¡2r pq n<„p<p +2rpq n¶ …:954: Now the pollster does not know porq, but he can use „ pand „q=1¡„pin their place without too much danger. With this idea in mind, the above statement isequivalent to the statement Pˆ „p¡2r „p„q n<p< „p+2r „p„q n! …:954: The resulting intervalµ „p¡2p„p„qpn;„p+2p„p„qpn¶ is called the 95 percent confldence interval for the unknown value of p. The name is suggested by the fact that if we use this method to estimate pin a large number of samples we should expect that in about 95 percent of the samples the true valueofpis contained in the confldence interval obtained from the sample. In Exercise 11 you are asked to write a program to illustrate that this does indeed happen. The pollster has control over the value of n. Thus, if he wants to create a 95% confldence interval with length 6%, then he should choose a value of nso that 2p „p„qpn•:03: Using the fact that „ p„q•1=4, no matter what the value of „ pis, it is easy to show that if he chooses a value of nso that 1pn•:03; 9.1. BERNOULLI TRIALS 335 0.48 0.5 0.52 0.54 0.56 0.58 0.60510152025 Figure 9.5: Polling simulation. he will be safe. This is equivalent to choosing n‚1111: So if the pollster chooses nto be 1200, say, and calculates „ pusing his sample of size 1200, then 19 times out of 20 (i.e., 95% of the time), his confldence interval, whichis of length 6%, will contain the true value of p. This type of confldence interval is typically reported in the news as follows: this survey has a 3% margin of error.In fact, most of the surveys that one sees reported in the paper will have samplesizes around 1000. A somewhat surprising fact is that the size of the population hasapparently no efiect on the sample size needed to obtain a 95% confldence intervalforpwith a given margin of error. To see this, note that the value of nthat was needed depended only on the number .03, which is the margin of error. In otherwords, whether the population is of size 100,000 or 100,000,000, the pollster needsonly to choose a sample of size 1200 or so to get the same accuracy of estimate ofp. (We did use the fact that the sample size was small relative to the population size in the statement that S nis approximately binomially distributed.) In Figure 9.5, we show the results of simulating the polling process. The popula- tion is of size 100,000, and for the population, p=:54. The sample size was chosen to be 1200. The spike graph shows the distribution of „ pfor 10,000 randomly chosen samples. For this simulation, the program kept track of the number of samples forwhich „pwas within 3% of .54. This number was 9648, which is close to 95% of the number of samples used. Another way to see what the idea of confldence intervals means is shown in Figure 9.6. In this flgure, we show 100 confldence intervals, obtained by computing „pfor 100 difierent samples of size 1200 from the same population as before. The reader can see that most of these confldence intervals (96, to be exact) contain thetrue value of p. 336 CHAPTER 9. CENTRAL LIMIT THEOREM 0.48 0.5 0.52 0.54 0.56 0.58 0.6 Figure 9.6: Confldence interval simulation. The Gallup Poll has used these polling techniques in every Presidential election since 1936 (and in innumerable other elections as well). Table 9.11shows the results of their efiorts. The reader will note that most of the approximations to pare within 3% of the actual value of p. The sample sizes for these polls were typically around 1500. (In the table, both the predicted and actual percentages for the winningcandidate refer to the percentage of the vote among the \major" political parties.In most elections, there were two major parties, but in several elections, there werethree.) This technique also plays an important role in the evaluation of the efiectiveness of drugs in the medical profession. For example, it is sometimes desired to knowwhat proportion of patients will be helped by a new drug. This proportion canbe estimated by giving the drug to a subset of the patients, and determining theproportion of this sample who are helped by the drug. 2 Historical Remarks The Central Limit Theorem for Bernoulli trials was flrst proved by Abraham de Moivre and appeared in his book, The Doctrine of Chances, flrst published in 1718.2 De Moivre spent his years from age 18 to 21 in prison in France because of his Protestant background. When he was released he left France for England, wherehe worked as a tutor to the sons of noblemen. Newton had presented a copy ofhisPrincipia Mathematica to the Earl of Devonshire. The story goes that, while 1The Gallup Poll Monthly, November 1992, No. 326, p. 33. Supplemented with the help of Lydia K. Saab, The Gallup Organization. 2A. de Moivre, The Doctrine of Chances, 3d ed. (London: Millar, 1756). 9.1. BERNOULLI TRIALS 337 Year Winning Gallup Final Election Deviation Candidate Survey Result 1936 Roosevelt 55.7% 62.5% 6.8%1940 Roosevelt 52.0% 55.0% 3.0%1944 Roosevelt 51.5% 53.3% 1.8%1948 Truman 44.5% 49.9% 5.4%1952 Eisenhower 51.0% 55.4% 4.4%1956 Eisenhower 59.5% 57.8% 1.7%1960 Kennedy 51.0% 50.1% 0.9%1964 Johnson 64.0% 61.3% 2.7%1968 Nixon 43.0% 43.5% 0.5%1972 Nixon 62.0% 61.8% 0.2%1976 Carter 48.0% 50.0% 2.0%1980 Reagan 47.0% 50.8% 3.8%1984 Reagan 59.0% 59.1% 0.1%1988 Bush 56.0% 53.9% 2.1%1992 Clinton 49.0% 43.2% 5.8%1996 Clinton 52.0% 50.1% 1.9% Table 9.1: Gallup Poll accuracy record. de Moivre was tutoring at the Earl’s house, he came upon Newton’s work and found that it was beyond him. It is said that he then bought a copy of his own and toreit into separate pages, learning it page by page as he walked around London to histutoring jobs. De Moivre frequented the cofieehouses in London, where he startedhis probability work by calculating odds for gamblers. He also met Newton at such acofieehouse and they became fast friends. De Moivre dedicated his book to Newton. The Doctrine of Chances provides the techniques for solving a wide variety of gambling problems. In the midst of these gambling problems de Moivre rathermodestly introduces his proof of the Central Limit Theorem, writing A Method of approximating the Sum of the Terms of the Binomial (a+b) nexpanded into a Series, from whence are deduced some prac- tical Rules to estimate the Degree of Assent which is to be given toExperiments. 3 De Moivre’s proof used the approximation to factorials that we now call Stirling’sformula. De Moivre states that he had obtained this formula before Stirling butwithout determining the exact value of the constantp 2…. While he says it is not really necessary to know this exact value, he concedes that knowing it \has spreada singular Elegancy on the Solution." The complete proof and an interesting discussion of the life of de Moivre can be found in the book Games, Gods and Gambling by F. N. David. 4 3ibid., p. 243. 4F. N. David, Games, Gods and Gambling (London: Gri–n, 1962). 338 CHAPTER 9. CENTRAL LIMIT THEOREM Exercises 1LetS100be the number of heads that turn up in 100 tosses of a fair coin. Use the Central Limit Theorem to estimate (a)P(S100•45). (b)P(45<S100<55). (c)P(S100>63). (d)P(S100<57). 2LetS200be the number of heads that turn up in 200 tosses of a fair coin. Estimate (a)P(S200= 100). (b)P(S200= 90). (c)P(S200= 80). 3A true-false examination has 48 questions. June has probability 3/4 of an- swering a question correctly. April just guesses on each question. A passingscore is 30 or more correct answers. Compare the probability that June passesthe exam with the probability that April passes it. 4LetSbe the number of heads in 1,000,000 tosses of a fair coin. Use (a) Cheby- shev’s inequality, and (b) the Central Limit Theorem, to estimate the prob-ability that Slies between 499,500 and 500,500. Use the same two methods to estimate the probability that Slies between 499,000 and 501,000, and the probability that Slies between 498,500 and 501,500. 5A rookie is brought to a baseball club on the assumption that he will have a .300 batting average. (Batting average is the ratio of the number of hits to thenumber of times at bat.) In the flrst year, he comes to bat 300 times and hisbatting average is .267. Assume that his at bats can be considered Bernoullitrials with probability .3 for success. Could such a low average be consideredjust bad luck or should he be sent back to the minor leagues? Comment onthe assumption of Bernoulli trials in this situation. 6Once upon a time, there were two railway trains competing for the passenger tra–c of 1000 people leaving from Chicago at the same hour and going to LosAngeles. Assume that passengers are equally likely to choose each train. Howmany seats must a train have to assure a probability of .99 or better of havinga seat for each passenger? 7Assume that, as in Example 9.3, Dartmouth admits 1750 students. What is the probability of too many acceptances? 8A club serves dinner to members only. They are seated at 12-seat tables. The manager observes over a long period of time that 95 percent of the time thereare between six and nine full tables of members, and the remainder of the 9.1. BERNOULLI TRIALS 339 time the numbers are equally likely to fall above or below this range. Assume that each member decides to come with a given probability p, and that the decisions are independent. How many members are there? What is p? 9LetSnbe the number of successes in nBernoulli trials with probability .8 for success on each trial. Let An=Sn=nbe the average number of successes. In each case give the value for the limit, and give a reason for your answer. (a) limn!1P(An=:8). (b) limn!1P(:7n<Sn<:9n). (c) limn!1P(Sn<:8n+:8pn). (d) limn!1P(:79<An<:81). 10Find the probability that among 10,000 random digits the digit 3 appears not more than 931 times. 11Write a computer program to simulate 10,000 Bernoulli trials with probabil- ity .3 for success on each trial. Have the program compute the 95 percentconfldence interval for the probability of success based on the proportion ofsuccesses. Repeat the experiment 100 times and see how many times the truevalue of .3 is included within the confldence limits. 12A balanced coin is °ipped 400 times. Determine the number xsuch that the probability that the number of heads is between 200 ¡xand 200 +xis approximately .80. 13A noodle machine in Spumoni’s spaghetti factory makes about 5 percent de- fective noodles even when properly adjusted. The noodles are then packedin crates containing 1900 noodles each. A crate is examined and found tocontain 115 defective noodles. What is the approximate probability of flndingat least this many defective noodles if the machine is properly adjusted? 14A restaurant feeds 400 customers per day. On the average 20 percent of the customers order apple pie. (a) Give a range (called a 95 percent confldence interval) for the number of pieces of apple pie ordered on a given day such that you can be 95 percentsure that the actual number will fall in this range. (b) How many customers must the restaurant have, on the average, to be at least 95 percent sure that the number of customers ordering pie on thatday falls in the 19 to 21 percent range? 15Recall that if Xis a random variable, the cumulative distribution function ofXis the function F(x) deflned by F(x)=P(X•x): (a) LetS nbe the number of successes in nBernoulli trials with probability p for success. Write a program to plot the cumulative distribution for Sn. 340 CHAPTER 9. CENTRAL LIMIT THEOREM (b) Modify your program in (a) to plot the cumulative distribution F⁄ n(x)o f the standardized random variable S⁄ n=Sn¡nppnpq: (c) Deflne the normal distribution N(x) to be the area under the normal curve up to the value x. Modify your program in (b) to plot the normal distribution as well, and compare it with the cumulative distributionofS ⁄ n. Do this for n=1 0;50, and 100. 16In Example 3.11, we were interested in testing the hypothesis that a new form of aspirin is efiective 80 percent of the time rather than the 60 percent of thetime as reported for standard aspirin. The new aspirin is given to npeople. If it is efiective in mor more cases, we accept the claim that the new drug is efiective 80 percent of the time and if not we reject the claim. Using theCentral Limit Theorem, show that you can choose the number of trials nand the critical value mso that the probability that we reject the hypothesis when it is true is less than .01 and the probability that we accept it when it is falseis also less than .01. Find the smallest value of nthat will su–ce for this. 17In an opinion poll it is assumed that an unknown proportion pof the people are in favor of a proposed new law and a proportion 1 ¡pare against it. A sample of npeople is taken to obtain their opinion. The proportion „ pin favor in the sample is taken as an estimate of p. Using the Central Limit Theorem, determine how large a sample will ensure that the estimate will,with probability .95, be correct to within .01. 18A description of a poll in a certain newspaper says that one can be 95% confldent that error due to sampling will be no more than plus or minus 3percentage points. A poll in the New York Times taken in Iowa says that\according to statistical theory, in 19 out of 20 cases the results based on suchsamples will difier by no more than 3 percentage points in either directionfrom what would have been obtained by interviewing all adult Iowans." Theseare both attempts to explain the concept of confldence intervals. Do bothstatements say the same thing? If not, which do you think is the more accuratedescription? 9.2 Central Limit Theorem for Discrete Indepen- dent Trials We have illustrated the Central Limit Theorem in the case of Bernoulli trials, but this theorem applies to a much more general class of chance processes. In particular,it applies to any independent trials process such that the individual trials have flnitevariance. For such a process, both the normal approximation for individual termsand the Central Limit Theorem are valid. 9.2. DISCRETE INDEPENDENT TRIALS 341 LetSn=X1+X2+¢¢¢+Xnbe the sum of nindependent discrete random variables of an independent trials process with common distribution function m(x) deflned on the integers, with mean „and variance ¾2. We have seen in Section 7.2 that the distributions for such independent sums have shapes resembling the nor-mal curve, but the largest values drift to the right and the curves °atten out (seeFigure 7.6). We can prevent this just as we did for Bernoulli trials. Standardized Sums Consider the standardized random variable S⁄ n=Sn¡n„p n¾2: This standardizes Snto have expected value 0 and variance 1. If Sn=j, then S⁄ nhas the value xjwith xj=j¡n„p n¾2: We can construct a spike graph just as we did for Bernoulli trials. Each spike is centered at some xj. The distance between successive spikes is b=1p n¾2; and the height of the spike is h=p n¾2P(Sn=j): The case of Bernoulli trials is the special case for which Xj= 1 if the jth outcome is a success and 0 otherwise; then „=pand¾2=ppq. We now illustrate this process for two difierent discrete distributions. The flrst is the distribution m, given by m=µ12345 :2:2:2:2:2¶ : In Figure 9.7 we show the standardized sums for this distribution for the cases n= 2 andn= 10. Even for n= 2 the approximation is surprisingly good. For our second discrete distribution, we choose m=µ12345 :4:3:1:1:1¶ : This distribution is quite asymmetric and the approximation is not very good forn= 3, but by n= 10 we again have an excellent approximation (see Figure 9.8). Figures 9.7 and 9.8 were produced by the program CLTIndTrialsPlot . 342 CHAPTER 9. CENTRAL LIMIT THEOREM -4 -2 0 2 400.10.20.30.4 -4 -2 0 2 400.10.20.30.4 n = 2 n = 10 Figure 9.7: Distribution of standardized sums. -4 -2 0 2 400.10.20.30.4 -4 -2 0 2 400.10.20.30.4 n = 3 n = 10 Figure 9.8: Distribution of standardized sums. Approximation Theorem As in the case of Bernoulli trials, these graphs suggest the following approximation theorem for the individual probabilities. Theorem 9.3 LetX1,X2, ...,Xnbe an independent trials process and let Sn= X1+X2+¢¢¢+Xn. Assume that the greatest common divisor of the difierences of all the values that the Xjcan take on is 1. Let E(Xj)=„andV(Xj)=¾2. Then fornlarge, P(Sn=j)»`(xj)p n¾2; wherexj=(j¡n„)=p n¾2, and`(x) is the standard normal density. 2 The program CLTIndTrialsLocal implements this approximation. When we run this program for 6 rolls of a die, and ask for the probability that the sum of therolls equals 21, we obtain an actual value of .09285, and a normal approximationvalue of .09537. If we run this program for 24 rolls of a die, and ask for theprobability that the sum of the rolls is 72, we obtain an actual value of .01724and a normal approximation value of .01705. These results show that the normalapproximations are quite good. 9.2. DISCRETE INDEPENDENT TRIALS 343 Central Limit Theorem for a Discrete Independent Trials Pro- cess The Central Limit Theorem for a discrete independent trials process is as follows. Theorem 9.4 (Central Limit Theorem) LetSn=X1+X2+¢¢¢+Xnbe the sum ofndiscrete independent random variables with common distribution having expected value „and variance ¾2. Then, for a<b , lim n!1Pµ a<Sn¡n„p n¾2<b¶ =1p 2…Zb ae¡x2=2dx : 2 We will give the proofs of Theorems 9.3 and Theorem 9.4 in Section 10.3. Here we consider several examples. Examples Example 9.5 A die is rolled 420 times. What is the probability that the sum of the rolls lies between 1400 and 1550? The sum is a random variable S420=X1+X2+¢¢¢+X420; where each Xjhas distribution mX=µ123456 1=61=61=61=61=61=6¶ We have seen that „=E(X)=7=2 and¾2=V(X)=3 5=12. Thus,E(S420)= 420¢7=2 = 1470,¾2(S420) = 420¢35=12 = 1225, and ¾(S420) = 35. Therefore, P(1400•S420•1550)…Pµ1399:5¡1470 35•S⁄ 240•1550:5¡1470 35¶ =P(¡2:01•S⁄ 420•2:30) …NA(¡2:01;2:30) =:9670: We note that the program CLTIndTrialsGlobal could be used to calculate these probabilities. 2 Example 9.6 A student’s grade point average is the average of his grades in 30 courses. The grades are based on 100 possible points and are recorded as integers.Assume that, in each course, the instructor makes an error in grading of kwith probabilityjp=kj, wherek=§1,§2,§3,§4,§5. The probability of no error is then 1¡(137=30)p. (The parameter prepresents the inaccuracy of the instructor’s grading.) Thus, in each course, there are two grades for the student, namely the 344 CHAPTER 9. CENTRAL LIMIT THEOREM \correct" grade and the recorded grade. So there are two average grades for the student, namely the average of the correct grades and the average of the recordedgrades. We wish to estimate the probability that these two average grades difier by less than .05 for a given student. We now assume that p=1=20. We also assume that the total error is the sum S 30of 30 independent random variables each with distribution mX:‰¡5¡4¡3¡2¡10 1234 5 1 1001 801 601 401 20463 6001 201 401 601 801 100¾ : One can easily calculate that E(X)=0a n d¾2(X)=1:5. Then we have P¡ ¡:05•S30 30•:05¢ =P(¡1:5•S30•1:5) =P‡ ¡1:5p 30¢1:5•S⁄ 30•1:5p 30¢1:5· =P(¡:224•S⁄ 30•:224) …NA(¡:224;:224) =:1772: This means that there is only a 17.7% chance that a given student’s grade point average is accurate to within .05. (Thus, for example, if two candidates for valedic-torian have recorded averages of 97.1 and 97.2, there is an appreciable probabilitythat their correct averages are in the reverse order.) For a further discussion of thisexample, see the article by R. M. Kozelka. 52 A More General Central Limit Theorem In Theorem 9.4, the discrete random variables that were being summed were as-sumed to be independent and identically distributed. It turns out that the assump-tion of identical distributions can be substantially weakened. Much work has beendone in this area, with an important contribution being made by J. W. Lindeberg.Lindeberg found a condition on the sequence fX ngwhich guarantees that the dis- tribution of the sum Snis asymptotically normally distributed. Feller showed that Lindeberg’s condition is necessary as well, in the sense that if the condition doesnot hold, then the sum S nis not asymptotically normally distributed. For a pre- cise statement of Lindeberg’s Theorem, we refer the reader to Feller.6A su–cient condition that is stronger (but easier to state) than Lindeberg’s condition, and isweaker than the condition in Theorem 9.4, is given in the following theorem. 5R. M. Kozelka, \Grade-Point Averages and the Central Limit Theorem," American Math. Monthly, vol. 86 (Nov 1979), pp. 773-777. 6W. Feller, Introduction to Probability Theory and its Applications, vol. 1, 3rd ed. (New York: John Wiley & Sons, 1968), p. 254. 9.2. DISCRETE INDEPENDENT TRIALS 345 Theorem 9.5 (Central Limit Theorem) LetX1;X2; :::; Xn; ::: be a se- quence of independent discrete random variables, and let Sn=X1+X2+¢¢¢+Xn. For eachn, denote the mean and variance of Xnby„nand¾2 n, respectively. De- flne the mean and variance of Snto bemnands2 n, respectively, and assume that sn!1 . If there exists a constant A, such thatjXnj•Afor alln, then fora<b , lim n!1Pµ a<Sn¡mn sn<b¶ =1p 2…Zb ae¡x2=2dx : 2 The condition that jXnj•Afor allnis sometimes described by saying that the sequencefXngis uniformly bounded. The condition that sn!1 is necessary (see Exercise 15). We illustrate this theorem by generating a sequence of nrandom distributions on the interval [ a;b]. We then convolute these distributions to flnd the distribution of the sum ofnexperiments governed by these distributions. Finally, we standardized the distribution for the sum to have mean 0 and standard deviation 1 and compareit with the normal density. The program CLTGeneral carries out this procedure. In Figure 9.9 we show the result of running this program for [ a;b]=[¡2;4], and n=1;4;and 10. We see that our flrst random distribution is quite asymmetric. By the time we choose the sum of ten such experiments we have a very good flt tothe normal curve. The above theorem essentially says that anything that can be thought of as being made up as the sum of many small independent pieces is approximately normallydistributed. This brings us to one of the most important questions that was askedabout genetics in the 1800’s. The Normal Distribution and Genetics When one looks at the distribution of heights of adults of one sex in a given pop-ulation, one cannot help but notice that this distribution looks like the normaldistribution. An example of this is shown in Figure 9.10. This flgure shows thedistribution of heights of 9593 women between the ages of 21 and 74. These datacome from the Health and Nutrition Examination Survey I (HANES I). For thissurvey, a sample of the U.S. civilian population was chosen. The survey was carriedout between 1971 and 1974. A natural question to ask is \How does this come about?". Francis Galton, an English scientist in the 19th century, studied this question, and other relatedquestions, and constructed probability models that were of great importance inexplaining the genetic efiects on such attributes as height. In fact, one of the mostimportant ideas in statistics, the idea of regression to the mean, was invented byGalton in his attempts to understand these genetic efiects. Galton was faced with an apparent contradiction. On the one hand, he knew that the normal distribution arises in situations in which many small independentefiects are being summed. On the other hand, he also knew that many quantitativeattributes, such as height, are strongly in°uenced by genetic factors: tall parents 346 CHAPTER 9. CENTRAL LIMIT THEOREM -4 -2 0 2 400.10.20.30.40.50.6 -4 -2 0 2 400.10.20.30.4 -4 -2 0 2 400.10.20.30.4 Figure 9.9: Sums of randomly chosen random variables. 9.2. DISCRETE INDEPENDENT TRIALS 347 50 55 60 65 70 75 8000.0250.050.0750.10.1250.15 Figure 9.10: Distribution of heights of adult women. tend to have tall ofispring. Thus in this case, there seem to be two large efiects, namely the parents. Galton was certainly aware of the fact that non-genetic factorsplayed a role in determining the height of an individual. Nevertheless, unless thesenon-genetic factors overwhelm the genetic ones, thereby refuting the hypothesisthat heredity is important in determining height, it did not seem possible for sets ofparents of given heights to have ofispring whose heights were normally distributed. One can express the above problem symbolically as follows. Suppose that we choose two speciflc positive real numbers xandy, and then flnd all pairs of parents one of whom is xunits tall and the other of whom is yunits tall. We then look at all of the ofispring of these pairs of parents. One can postulate the existence ofa function f(x;y) which denotes the genetic efiect of the parents’ heights on the heights of the ofispring. One can then let Wdenote the efiects of the non-genetic factors on the heights of the ofispring. Then, for a given set of heights fx;yg, the random variable which represents the heights of the ofispring is given by H=f(x;y)+W; wherefis a deterministic function, i.e., it gives one output for a pair of inputs fx;yg. If we assume that the efiect of fis large in comparison with the efiect of W, then the variance of Wis small. But since f is deterministic, the variance of H equals the variance of W, so the variance of His small. However, Galton observed from his data that the variance of the heights of the ofispring of a given pair ofparent heights is not small. This would seem to imply that inheritance plays asmall role in the determination of the height of an individual. Later in this section,we will describe the way in which Galton got around this problem. We will now consider the modern explanation of why certain traits, such as heights, are approximately normally distributed. In order to do so, we need tointroduce some terminology from the fleld of genetics. The cells in a living organismthat are not directly involved in the transmission of genetic material to ofispringare called somatic cells, and the remaining cells are called germ cells. Organisms ofa given species have their genetic information encoded in sets of physical entities, 348 CHAPTER 9. CENTRAL LIMIT THEOREM called chromosomes. The chromosomes are paired in each somatic cell. For example, human beings have 23 pairs of chromosomes in each somatic cell. The sex cellscontain one chromosome from each pair. In sexual reproduction, two sex cells, onefrom each parent, contribute their chromosomes to create the set of chromosomesfor the ofispring. Chromosomes contain many subunits, called genes. Genes consist of molecules of DNA, and one gene has, encoded in its DNA, information that leads to the reg-ulation of proteins. In the present context, we will consider those genes containinginformation that has an efiect on some physical trait, such as height, of the organ-ism. The pairing of the chromosomes gives rise to a pairing of the genes on thechromosomes. In a given species, each gene can be any one of several forms. These various forms are called alleles. One should think of the difierent alleles as potentiallyproducing difierent efiects on the physical trait in question. Of the two alleles thatare found in a given gene pair in an organism, one of the alleles came from oneparent and the other allele came from the other parent. The possible types of pairsof alleles (without regard to order) are called genotypes. If we assume that the height of a human being is largely controlled by a speciflc gene, then we are faced with the same di–culty that Galton was. We are assumingthat each parent has a pair of alleles which largely controls their heights. Sinceeach parent contributes one allele of this gene pair to each of its ofispring, there arefour possible allele pairs for the ofispring at this gene location. The assumption isthat these pairs of alleles largely control the height of the ofispring, and we are alsoassuming that genetic factors outweigh non-genetic factors. It follows that amongthe ofispring we should see several modes in the height distribution of the ofispring,one mode corresponding to each possible pair of alleles. This distribution does notcorrespond to the observed distribution of heights. An alternative hypothesis, which does explain the observation of normally dis- tributed heights in ofispring of a given sex, is the multiple-gene hypothesis. Underthis hypothesis, we assume that there are many genes that afiect the height of anindividual. These genes may difier in the amount of their efiects. Thus, we canrepresent each gene pair by a random variable X i, where the value of the random variable is the allele pair’s efiect on the height of the individual. Thus, for example,if each parent has two difierent alleles in the gene pair under consideration, thenthe ofispring has one of four possible pairs of alleles at this gene location. Now theheight of the ofispring is a random variable, which can be expressed as H=X 1+X2+¢¢¢+Xn+W; if there are ngenes that afiect height. (Here, as before, the random variable Wde- notes non-genetic efiects.) Although nis flxed, if it is fairly large, then Theorem 9.5 implies that the sum X1+X2+¢¢¢+Xnis approximately normally distributed. Now, if we assume that the Xi’s have a signiflcantly larger cumulative efiect than Wdoes, thenHis approximately normally distributed. Another observed feature of the distribution of heights of adults of one sex in a population is that the variance does not seem to increase or decrease from one 9.2. DISCRETE INDEPENDENT TRIALS 349 generation to the next. This was known at the time of Galton, and his attempts to explain this led him to the idea of regression to the mean. This idea will bediscussed further in the historical remarks at the end of the section. (The reasonthat we only consider one sex is that human heights are clearly sex-linked, and ingeneral, if we have two populations that are each normally distributed, then theirunion need not be normally distributed.) Using the multiple-gene hypothesis, it is easy to explain why the variance should be constant from generation to generation. We begin by assuming that for a speciflcgene location, there are kalleles, which we will denote by A 1;A2; :::; Ak.W e assume that the ofispring are produced by random mating. By this we mean thatgiven any ofispring, it is equally likely that it came from any pair of parents in thepreceding generation. There is another way to look at random mating that makesthe calculations easier. We consider the set Sof all of the alleles (at the given gene location) in all of the germ cells of all of the individuals in the parent generation.In terms of the set S, by random mating we mean that each pair of alleles in Sis equally likely to reside in any particular ofispring. (The reader might object to thisway of thinking about random mating, as it allows two alleles from the same parentto end up in an ofispring; but if the number of individuals in the parent populationis large, then whether or not we allow this event does not afiect the probabilitiesvery much.) For 1•i•k, we letp idenote the proportion of alleles in the parent population that are of type Ai. It is clear that this is the same as the proportion of alleles in the germ cells of the parent population, assuming that each parent produces roughlythe same number of germs cells. Consider the distribution of alleles in the ofispring.Since each germ cell is equally likely to be chosen for any particular ofispring, thedistribution of alleles in the ofispring is the same as in the parents. We next consider the distribution of genotypes in the two generations. We will prove the following fact: the distribution of genotypes in the ofispring generationdepends only upon the distribution of alleles in the parent generation (in particular,it does not depend upon the distribution of genotypes in the parent generation).Consider the possible genotypes; there are k(k+1 )=2 of them. Under our assump- tions, the genotype A iAiwill occur with frequency p2 i, and the genotype AiAj, withi6=j, will occur with frequency 2 pipj. Thus, the frequencies of the genotypes depend only upon the allele frequencies in the parent generation, as claimed. This means that if we start with a certain generation, and a certain distribution of alleles, then in all generations after the one we started with, both the alleledistribution and the genotype distribution will be flxed. This last statement isknown as the Hardy-Weinberg Law. We can describe the consequences of this law for the distribution of heights among adults of one sex in a population. We recall that the height of an ofispringwas given by a random variable H, where H=X 1+X2+¢¢¢+Xn+W; with theXi’s corresponding to the genes that afiect height, and the random variable Wdenoting non-genetic efiects. The Hardy-Weinberg Law states that for each Xi, 350 CHAPTER 9. CENTRAL LIMIT THEOREM the distribution in the ofispring generation is the same as the distribution in the parent generation. Thus, if we assume that the distribution of Wis roughly the same from generation to generation (or if we assume that its efiects are small), thenthe distribution of His the same from generation to generation. (In fact, dietary efiects are part of W, and it is clear that in many human populations, diets have changed quite a bit from one generation to the next in recent times. This change isthought to be one of the reasons that humans, on the average, are getting taller. Itis also the case that the efiects of Ware thought to be small relative to the genetic efiects of the parents.) Discussion Generally speaking, the Central Limit Theorem contains more information thanthe Law of Large Numbers, because it gives us detailed information about theshape of the distribution of S ⁄ n; for large nthe shape is approximately the same as the shape of the standard normal density. More speciflcally, the Central LimitTheorem says that if we standardize and height-correct the distribution of S n, then the normal density function is a very good approximation to this distribution whennis large. Thus, we have a computable approximation for the distribution for S n, which provides us with a powerful technique for generating answers for all sortsof questions about sums of independent random variables, even if the individualrandom variables have difierent distributions. Historical Remarks In the mid-1800’s, the Belgian mathematician Quetelet7had shown empirically that the normal distribution occurred in real data, and had also given a method for flttingthe normal curve to a given data set. Laplace 8had shown much earlier that the sum of many independent identically distributed random variables is approximatelynormal. Galton knew that certain physical traits in a population appeared to beapproximately normally distributed, but he did not consider Laplace’s result to bea good explanation of how this distribution comes about. We give a quote fromGalton that appears in the fascinating book by S. Stigler 9on the history of statistics: First, let me point out a fact which Quetelet and all writers who have followed in his paths have unaccountably overlooked, and which has anintimate bearing on our work to-night. It is that, although characteris-tics of plants and animals conform to the law, the reason of their doingso is as yet totally unexplained. The essence of the law is that difierencesshould be wholly due to the collective actions of a host of independentpetty in°uences in various combinations...Now the processes of hered- ity...are not petty in°uences, but very important ones...The conclusionis...that the processes of heredity must work harmoniously with the lawof deviation, and be themselves in some sense conformable to it. 7S. Stigler, The History of Statistics, (Cambridge: Harvard University Press, 1986), p. 203. 8ibid., p. 136 9ibid., p. 281. 9.2. DISCRETE INDEPENDENT TRIALS 351 Figure 9.11: Two-stage version of the quincunx. Galton invented a device known as a quincunx (now commonly called a Galton board), which we used in Example 3.10 to show how to physically obtain a binomialdistribution. Of course, the Central Limit Theorem says that for large values ofthe parameter n, the binomial distribution is approximately normal. Galton used the quincunx to explain how inheritance afiects the distribution of a trait amongofispring. We consider, as Galton did, what happens if we interrupt, at some intermediate height, the progress of the shot that is falling in the quincunx. The reader is referredto Figure 9.11. This flgure is a drawing of Karl Pearson, 10based upon Galton’s notes. In this flgure, the shot is being temporarily segregated into compartments atthe line AB. (The line A 0B0forms a platform on which the shot can rest.) If the line AB is not too close to the top of the quincunx, then the shot will be approximatelynormally distributed at this line. Now suppose that one compartment is opened, asshown in the flgure. The shot from that compartment will fall, forming a normaldistribution at the bottom of the quincunx. If now all of the compartments areopened, all of the shot will fall, producing the same distribution as would occur ifthe shot were not temporarily stopped at the line AB. But the action of stoppingthe shot at the line AB, and then releasing the compartments one at a time, is 10Karl Pearson, The Life, Letters and Labours of Francis Galton, vol. IIIB, (Cambridge at the University Press 1930.) p. 466. Reprinted with permission. 352 CHAPTER 9. CENTRAL LIMIT THEOREM just the same as convoluting two normal distributions. The normal distributions at the bottom, corresponding to each compartment at the line AB, are being mixed,with their weights being the number of shot in each compartment. On the otherhand, it is already known that if the shot are unimpeded, the flnal distribution isapproximately normal. Thus, this device shows that the convolution of two normaldistributions is again normal. Galton also considered the quincunx from another perspective. He segregated into seven groups, by weight, a set of 490 sweet pea seeds. He gave 10 seeds fromeach of the seven group to each of seven friends, who grew the plants from theseeds. Galton found that each group produced seeds whose weights were normallydistributed. (The sweet pea reproduces by self-pollination, so he did not need toconsider the possibility of interaction between difierent groups.) In addition, hefound that the variances of the weights of the ofispring were the same for eachgroup. This segregation into groups corresponds to the compartments at the lineAB in the quincunx. Thus, the sweet peas were acting as though they were beinggoverned by a convolution of normal distributions. He now was faced with a problem. We have shown in Chapter 7, and Galton knew, that the convolution of two normal distributions produces a normal distribu-tion with a larger variance than either of the original distributions. But his data onthe sweet pea seeds showed that the variance of the ofispring population was thesame as the variance of the parent population. His answer to this problem was topostulate a mechanism that he called reversion , and is now called regression to the mean . As Stigler puts it: 11 The seven groups of progeny were normally distributed, but not about their parents’ weight. Rather they were in every case distributed abouta value that was closer to the average population weight than was that ofthe parent. Furthermore, this reversion followed \the simplest possiblelaw," that is, it was linear. The average deviation of the progeny fromthe population average was in the same direction as that of the parent,but only a third as great. The mean progeny reverted to type, andthe increased variation was just su–cient to maintain the populationvariability. Galton illustrated reversion with the illustration shown in Figure 9.12. 12The parent population is shown at the top of the flgure, and the slanted lines are meantto correspond to the reversion efiect. The ofispring population is shown at thebottom of the flgure. Exercises 1A die is rolled 24 times. Use the Central Limit Theorem to estimate the probability that 11ibid., p. 282. 12Karl Pearson, The Life, Letters and Labours of Francis Galton, vol. IIIA, (Cambridge at the University Press 1930.) p. 9. Reprinted with permission. 9.2. DISCRETE INDEPENDENT TRIALS 353 Figure 9.12: Galton’s explanation of reversion. 354 CHAPTER 9. CENTRAL LIMIT THEOREM (a) the sum is greater than 84. (b) the sum is equal to 84. 2A random walker starts at 0 on the x-axis and at each time unit moves 1 step to the right or 1 step to the left with probability 1/2. Estimate theprobability that, after 100 steps, the walker is more than 10 steps from thestarting position. 3A piece of rope is made up of 100 strands. Assume that the breaking strength of the rope is the sum of the breaking strengths of the individual strands.Assume further that this sum may be considered to be the sum of an inde-pendent trials process with 100 experiments each having expected value of 10pounds and standard deviation 1. Find the approximate probability that therope will support a weight (a) of 1000 pounds. (b) of 970 pounds. 4Write a program to flnd the average of 1000 random digits 0, 1, 2, 3, 4, 5, 6, 7, 8, or 9. Have the program test to see if the average lies within three standarddeviations of the expected value of 4.5. Modify the program so that it repeatsthis simulation 1000 times and keeps track of the number of times the test ispassed. Does your outcome agree with the Central Limit Theorem? 5A die is thrown until the flrst time the total sum of the face values of the die is 700 or greater. Estimate the probability that, for this to happen, (a) more than 210 tosses are required. (b) less than 190 tosses are required. (c) between 180 and 210 tosses, inclusive, are required. 6A bank accepts rolls of pennies and gives 50 cents credit to a customer without counting the contents. Assume that a roll contains 49 pennies 30 percent ofthe time, 50 pennies 60 percent of the time, and 51 pennies 10 percent of thetime. (a) Find the expected value and the variance for the amount that the bank loses on a typical roll. (b) Estimate the probability that the bank will lose more than 25 cents in 100 rolls. (c) Estimate the probability that the bank will lose exactly 25 cents in 100 rolls. (d) Estimate the probability that the bank will lose any money in 100 rolls. (e) How many rolls does the bank need to collect to have a 99 percent chance of a net loss? 9.2. DISCRETE INDEPENDENT TRIALS 355 7A surveying instrument makes an error of ¡2,¡1, 0, 1, or 2 feet with equal probabilities when measuring the height of a 200-foot tower. (a) Find the expected value and the variance for the height obtained using this instrument once. (b) Estimate the probability that in 18 independent measurements of this tower, the average of the measurements is between 199 and 201, inclusive. 8For Example 9.6 estimate P(S30= 0). That is, estimate the probability that the errors cancel out and the student’s grade point average is correct. 9Prove the Law of Large Numbers using the Central Limit Theorem. 10Peter and Paul match pennies 10,000 times. Describe brie°y what each of the following theorems tells you about Peter’s fortune. (a) The Law of Large Numbers. (b) The Central Limit Theorem. 11A tourist in Las Vegas was attracted by a certain gambling game in which the customer stakes 1 dollar on each play; a win then pays the customer2 dollars plus the return of her stake, although a loss costs her only her stake.Las Vegas insiders, and alert students of probability theory, know that theprobability of winning at this game is 1/4. When driven from the tables byhunger, the tourist had played this game 240 times. Assuming that no nearmiracles happened, about how much poorer was the tourist upon leaving thecasino? What is the probability that she lost no money? 12We have seen that, in playing roulette at Monte Carlo (Example 6.13), betting 1 dollar on red or 1 dollar on 17 amounts to choosing between the distributions m X=µ¡1¡1=21 18=37 1=37 18=37¶ or mX=µ¡13 5 36=37 1=37¶ You plan to choose one of these methods and use it to make 100 1-dollar bets using the method chosen. Which gives you the higher probability of winningat least 20 dollars? Which gives you the higher probability of winning anymoney? 13In Example 9.6 flnd the largest value of pthat gives probability .954 that the flrst decimal place is correct. 14It has been suggested that Example 9.6 is unrealistic, in the sense that the probabilities of errors are too low. Make up your own (reasonable) estimatefor the distribution m(x), and determine the probability that a student’s grade point average is accurate to within .05. Also determine the probability thatit is accurate to within .5. 356 CHAPTER 9. CENTRAL LIMIT THEOREM 15Find a sequence of uniformly bounded discrete independent random variables fXngsuch that the variance of their sum does not tend to 1asn!1 , and such that their sum is not asymptotically normally distributed. 9.3 Central Limit Theorem for Continuous Inde- pendent Trials We have seen in Section 9.2 that the distribution function for the sum of a large numbernof independent discrete random variables with mean „and variance ¾2 tends to look like a normal density with mean n„and variance n¾2. What is remarkable about this result is that it holds for anydistribution with flnite mean and variance. We shall see in this section that the same result also holds true forcontinuous random variables having a common density function. Let us begin by looking at some examples to see whether such a result is even plausible. Standardized Sums Example 9.7 Suppose we choose nrandom numbers from the interval [0 ;1] with uniform density. Let X1,X2, ...,Xndenote these choices, and Sn=X1+X2+ ¢¢¢+Xntheir sum. We saw in Example 7.9 that the density function for Sntends to have a normal shape, but is centered at n=2 and is °attened out. In order to compare the shapes of these density functions for difierent values of n, we proceed as in the previous section: we standardize Snby deflning S⁄ n=Sn¡n„pn¾: Then we see that for all nwe have E(S⁄ n)=0; V(S⁄ n)=1: The density function for S⁄ nis just a standardized version of the density function forSn(see Figure 9.13). 2 Example 9.8 Let us do the same thing, but now choose numbers from the interval [0;+1) with an exponential density with parameter ‚. Then (see Example 6.26) „=E(Xi)=1 ‚; ¾2=V(Xj)=1 ‚2: 9.3. CONTINUOUS INDEPENDENT TRIALS 357 -3 -2 -1 1 2 30.10.20.30.4n = 2 n = 3 n = 4 n = 10 Figure 9.13: Density function for S⁄ n(uniform case, n=2;3;4;10). Here we know the density function for Snexplicitly (see Section 7.2). We can use Corollary 5.1 to calculate the density function for S⁄ n. We obtain fSn(x)=‚e¡‚x(‚x)n¡1 (n¡1)!; fS⁄ n(x)=pn ‚fSnµpnx+n ‚¶ : The graph of the density function for S⁄ nis shown in Figure 9.14. 2 These examples make it seem plausible that the density function for the nor- malized random variable S⁄ nfor largenwill look very much like the normal density with mean 0 and variance 1 in the continuous case as well as in the discrete case.The Central Limit Theorem makes this statement precise. Central Limit Theorem Theorem 9.6 (Central Limit Theorem) LetSn=X1+X2+¢¢¢+Xnbe the sum ofnindependent continuous random variables with common density function p having expected value „and variance ¾2. LetS⁄ n=(Sn¡n„)=pn¾. Then we have, for alla<b , lim n!1P(a<S⁄ n<b)=1p 2…Zb ae¡x2=2dx : 2 We shall give a proof of this theorem in Section 10.3. We will now look at some examples. 358 CHAPTER 9. CENTRAL LIMIT THEOREM -4 -2 20.10.20.30.40.5 n = 2 n = 3n = 10n = 30 Figure 9.14: Density function for S⁄ n(exponential case, n=2;3;10;30,‚= 1). Example 9.9 Suppose a surveyor wants to measure a known distance, say of 1 mile, using a transit and some method of triangulation. He knows that because of possiblemotion of the transit, atmospheric distortions, and human error, any one measure-ment is apt to be slightly in error. He plans to make several measurements and takean average. He assumes that his measurements are independent random variableswith a common distribution of mean „= 1 and standard deviation ¾=:0002 (so, if the errors are approximately normally distributed, then his measurements arewithin 1 foot of the correct distance about 65% of the time). What can he sayabout the average? He can say that if nis large, the average S n=nhas a density function that is approximately normal, with mean „= 1 mile, and standard deviation ¾=:0002=pn miles. How many measurements should he make to be reasonably sure that his average lies within .0001 of the true value? The Chebyshev inequality says PµflflflflS n n¡„flflflfl‚:0001¶ •(:0002) 2 n(10¡8)=4 n; so that we must have n‚80 before the probability that his error is less than .0001 exceeds .95. We have already noticed that the estimate in the Chebyshev inequality is not always a good one, and here is a case in point. If we assume that nis large enough so that the density for Snis approximately normal, then we have PµflflflflS n n¡„flflflfl<:0001¶ =P¡ ¡:5p n<S⁄ n<+:5pn¢ …1p 2…Z+:5pn ¡:5pne¡x2=2dx ; 9.3. CONTINUOUS INDEPENDENT TRIALS 359 and this last expression is greater than .95 if :5pn‚2:This says that it su–ces to taken= 16 measurements for the same results. This second calculation is stronger, but depends on the assumption that n= 16 is large enough to establish the normal density as a good approximation to S⁄ n, and hence to Sn. The Central Limit Theorem here says nothing about how large nhas to be. In most cases involving sums of independent random variables, a good rule of thumb is that forn‚30, the approximation is a good one. In the present case, if we assume that the errors are approximately normally distributed, then the approximation is probablyfairly good even for n= 16. 2 Estimating the Mean Example 9.10 (Continuation of Example 9.9) Now suppose our surveyor is mea- suring an unknown distance with the same instruments under the same conditions.He takes 36 measurements and averages them. How sure can he be that his mea-surement lies within .0002 of the true value? Again using the normal approximation, we get PµflflflflS n n¡„flflflfl<:0002¶ =P¡ jS ⁄ nj<:5pn¢ …2p 2…Z3 ¡3e¡x2=2dx …:997: This means that the surveyor can be 99.7 percent sure that his average is within .0002 of the true value. To improve his confldence, he can take more measurements,or require less accuracy, or improve the quality of his measurements (i.e., reducethe variance ¾ 2). In each case, the Central Limit Theorem gives quantitative infor- mation about the confldence of a measurement process, assuming always that thenormal approximation is valid. Now suppose the surveyor does not know the mean or standard deviation of his measurements, but assumes that they are independent. How should he proceed? Again, he makes several measurements of a known distance and averages them. As before, the average error is approximately normally distributed, but now withunknown mean and variance. 2 Sample Mean If he knows the variance ¾2of the error distribution is .0002, then he can estimate the mean„by taking the average, orsample mean of, say, 36 measurements: „„=x1+x2+¢¢¢+xn n; wheren= 36. Then, as before, E(„„)=„. Moreover, the preceding argument shows that P(j„„¡„j<:0002)…:997: 360 CHAPTER 9. CENTRAL LIMIT THEOREM The interval („ „¡:0002;„„+:0002) is called the 99.7% confldence interval for„(see Example 9.4). Sample Variance If he does not know the variance ¾2of the error distribution, then he can estimate ¾2by the sample variance : „¾2=(x1¡„„)2+(x2¡„„)2+¢¢¢+(xn¡„„)2 n; wheren= 36. The Law of Large Numbers, applied to the random variables ( Xi¡ „„)2, says that for large n, the sample variance „ ¾2lies close to the variance ¾2,s o that the surveyor can use „ ¾2in place of¾2in the argument above. Experience has shown that, in most practical problems of this type, the sample variance is a good estimate for the variance, and can be used in place of the varianceto determine confldence levels for the sample mean. This means that we can relyon the Law of Large Numbers for estimating the variance, and the Central LimitTheorem for estimating the mean. We can check this in some special cases. Suppose we know that the error distri- bution is normal, with unknown mean and variance. Then we can take a sample of nmeasurements, flnd the sample mean „ „and sample variance „ ¾ 2, and form T⁄ n=Sn¡n„„pn„¾; wheren= 36. We expect T⁄ nto be a good approximation for S⁄ nfor largen. t-Density The statistician W. S. Gosset13has shown that in this case T⁄ nhas a density function that is not normal but rather a t-density withndegrees of freedom. (The number nof degrees of freedom is simply a parameter which tells us which t-density to use.) In this case we can use the t-density in place of the normal density to determine confldence levels for „.A snincreases, the t-density approaches the normal density. Indeed, even for n= 8 thet-density and normal density are practically the same (see Figure 9.15). Exercises Notes on computer problems : (a) Simulation: Recall (see Corollary 5.2) that X=F¡1(rnd) 13W. S. Gosset discovered the distribution we now call the t-distribution while working for the Guinness Brewery in Dublin. He wrote under the pseudonym \Student." The results discussed hereflrst appeared in Student, \The Probable Error of a Mean," Biometrika, vol. 6 (1908), pp. 1-24. 9.3. CONTINUOUS INDEPENDENT TRIALS 361 -6 -4 -2 2 4 60.10.20.30.4 Figure 9.15: Graph of t¡density for n=1;3;8 and the normal density with „= 0;¾=1 . will simulate a random variable with density f(x) and distribution F(X)=Zx ¡1f(t)dt : In the case that f(x) is a normal density function with mean „and standard deviation¾, where neither FnorF¡1can be expressed in closed form, use instead X=¾p ¡2 log(rnd) cos 2…(rnd)+„: (b) Bar graphs: you should aim for about 20 to 30 bars (of equal width) in your graph. You can achieve this by a good choice of the range [ xmin;xmin] and the number of bars (for instance, [ „¡3¾;„+3¾] with 30 bars will work in many cases). Experiment! 1LetXbe a continuous random variable with mean „(X) and variance ¾2(X), and letX⁄=(X¡„)=¾be its standardized version. Verify directly that „(X⁄)=0a n d¾2(X⁄)=1 . 2LetfXkg,1•k•n, be a sequence of independent random variables, all with mean 0 and variance 1, and let Sn,S⁄ n, andAnbe their sum, standardized sum, and average, respectively. Verify directly that S⁄ n=Sn=pn=pnAn. 3LetfXkg,1•k•n, be a sequence of random variables, all with mean „and variance¾2, andYk=X⁄ kbe their standardized versions. Let SnandTnbe the sum of the XkandYk, andS⁄ nandT⁄ ntheir standardized version. Show thatS⁄ n=T⁄ n=Tn=pn. 362 CHAPTER 9. CENTRAL LIMIT THEOREM 4Suppose we choose independently 25 numbers at random (uniform density) from the interval [0 ;20]. Write the normal densities that approximate the densities of their sum S25, their standardized sum S⁄ 25, and their average A25. 5Write a program to choose independently 25 numbers at random from [0 ;20], compute their sum S25, and repeat this experiment 1000 times. Make a bar graph for the density of S25and compare it with the normal approximation of Exercise 4. How good is the flt? Now do the same for the standardizedsumS ⁄ 25and the average A25. 6In general, the Central Limit Theorem gives a better estimate than Cheby- shev’s inequality for the average of a sum. To see this, let A25be the average calculated in Exercise 5, and let Nbe the normal approximation forA25. Modify your program in Exercise 5 to provide a table of the function F(x)=P(jA25¡10j‚x) = fraction of the total of 1000 trials for which jA25¡10j‚x. Do the same for the function f(x)=P(jN¡10j‚x). (You can use the normal table, Table 9.4, or the procedure NormalArea for this.) Now plot on the same axes the graphs of F(x),f(x), and the Chebyshev functiong(x)=4=(3x2). How do f(x) andg(x) compare as estimates for F(x)? 7The Central Limit Theorem says the sums of independent random variables tend to look normal, no matter what crazy distribution the individual variableshave. Let us test this by a computer simulation. Choose independently 25numbers from the interval [0 ;1] with the probability density f(x) given below, and compute their sum S 25. Repeat this experiment 1000 times, and make up a bar graph of the results. Now plot on the same graph the density `(x)= normal (x;„(S25);¾(S25)). How well does the normal density flt your bar graph in each case? (a)f(x)=1 . (b)f(x)=2x. (c)f(x)=3x2. (d)f(x)=4jx¡1=2j. (e)f(x)=2¡4jx¡1=2j. 8Repeat the experiment described in Exercise 7 but now choose the 25 numbers from [0;1), usingf(x)=e¡x. 9How large must nbe beforeSn=X1+X2+¢¢¢+Xnis approximately normal? This number is often surprisingly small. Let us explore this question with acomputer simulation. Choose nnumbers from [0 ;1] with probability density f(x), wheren= 3, 6, 12, 20, and f(x) is each of the densities in Exercise 7. Compute their sum S n, repeat this experiment 1000 times, and make up a bar graph of 20 bars of the results. How large must nbe before you get a good flt? 9.3. CONTINUOUS INDEPENDENT TRIALS 363 10A surveyor is measuring the height of a clifi known to be about 1000 feet. He assumes his instrument is properly calibrated and that his measurementerrors are independent, with mean „= 0 and variance ¾ 2= 10. He plans to takenmeasurements and form the average. Estimate, using (a) Chebyshev’s inequality and (b) the normal approximation, how large nshould be if he wants to be 95 percent sure that his average falls within 1 foot of the truevalue. Now estimate, using (a) and (b), what value should ¾ 2have if he wants to make only 10 measurements with the same confldence? 11The price of one share of stock in the Pilsdorfi Beer Company (see Exer- cise 8.2.12) is given by Ynon thenth day of the year. Finn observes that the difierences Xn=Yn+1¡Ynappear to be independent random variables with a common distribution having mean „= 0 and variance ¾2=1=4. If Y1= 100, estimate the probability that Y365is (a)‚100. (b)‚110. (c)‚120. 12Test your conclusions in Exercise 11 by computer simulation. First choose 364 numbers Xiwith density f(x) = normal( x;0;1=4). Now form the sum Y365= 100 +X1+X2+¢¢¢+X364, and repeat this experiment 200 times. Make up a bar graph on [50 ;150] of the results, superimposing the graph of the approximating normal density. What does this graph say about your answersin Exercise 11? 13Physicists say that particles in a long tube are constantly moving back and forth along the tube, each with a velocity V k(in cm/sec) at any given moment that is normally distributed, with mean „= 0 and variance ¾2= 1. Suppose there are 1020particles in the tube. (a) Find the mean and variance of the average velocity of the particles. (b) What is the probability that the average velocity is ‚10¡9cm/sec? 14An astronomer makes nmeasurements of the distance between Jupiter and a particular one of its moons. Experience with the instruments used leadsher to believe that for the proper units the measurements will be normallydistributed with mean d, the true distance, and variance 16. She performs a series ofnmeasurements. Let A n=X1+X2+¢¢¢+Xn n be the average of these measurements. (a) Show that Pµ An¡8pn•d•An+8pn¶ …:95: 364 CHAPTER 9. CENTRAL LIMIT THEOREM (b) When nine measurements were taken, the average of the distances turned out to be 23.2 units. Putting the observed values in (a) gives the 95 per- cent confldence interval for the unknown distance d. Compute this in- terval. (c) Why not say in (b) more simply that the probability is .95 that the value ofdlies in the computed confldence interval? (d) What changes would you make in the above procedure if you wanted to compute a 99 percent confldence interval? 15Plot a bar graph similar to that in Figure 9.10 for the heights of the mid- parents in Galton’s data as given in Appendix B and compare this bar graphto the appropriate normal curve. Chapter 10 Generating Functions 10.1 Generating Functions for Discrete Distribu- tions So far we have considered in detail only the two most important attributes of a random variable, namely, the mean and the variance. We have seen how theseattributes enter into the fundamental limit theorems of probability, as well as intoall sorts of practical calculations. We have seen that the mean and variance ofa random variable contain important information about the random variable, or,more precisely, about the distribution function of that variable. Now we shall seethat the mean and variance do notcontain allthe available information about the density function of a random variable. To begin with, it is easy to give examples ofdifierent distribution functions which have the same mean and the same variance.For instance, suppose XandYare random variables, with distributions p X=µ12 34 56 01=41=2001 =4¶ ; pY=µ12 34 56 1=4001 =21=40¶ : Then with these choices, we have E(X)=E(Y)=7=2 andV(X)=V(Y)=9=4, and yet certainly pXandpYare quite difierent density functions. This raises a question: If Xis a random variable with range fx1;x2;:::gof at most countable size, and distribution function p=pX, and if we know its mean „=E(X) and its variance ¾2=V(X), then what else do we need to know to determinepcompletely? Moments A nice answer to this question, at least in the case that Xhas flnite range, can be given in terms of the moments ofX, which are numbers deflned as follows: 365 366 CHAPTER 10. GENERATING FUNCTIONS „k=kth moment of X =E(Xk) =1X j=1(xj)kp(xj); provided the sum converges. Here p(xj)=P(X=xj). In terms of these moments, the mean „and variance ¾2ofXare given simply by „=„1; ¾2=„2¡„2 1; so that a knowledge of the flrst two moments of Xgives us its mean and variance. But a knowledge of allthe moments of Xdetermines its distribution function p completely. Moment Generating Functions To see how this comes about, we introduce a new variable t, and deflne a function g(t) as follows: g(t)=E(etX) =1X k=0„ktk k! =Eˆ1X k=0Xktk k!! =1X j=1etxjp(xj): We callg(t) the moment generating function forX, and think of it as a convenient bookkeeping device for describing the moments of X. Indeed, if we difierentiate g(t)ntimes and then set t=0 ,w eg e t „n: dn dtng(t)flflflfl t=0=g(n)(0) =1X k=nk!„ktk¡n (k¡n)!k!flflflflfl t=0 =„n: It is easy to calculate the moment generating function for simple examples. 10.1. DISCRETE DISTRIBUTIONS 367 Examples Example 10.1 SupposeXhas rangef1;2;3;:::;ngandpX(j)=1=nfor 1•j•n (uniform distribution). Then g(t)=nX j=11 netj =1 n(et+e2t+¢¢¢+ent) =et(ent¡1) n(et¡1): If we use the expression on the right-hand side of the second line above, then it is easy to see that „1=g0(0) =1 n( 1+2+3+¢¢¢+n)=n+1 2; „2=g00(0) =1 n( 1+4+9+¢¢¢+n2)=(n+ 1)(2n+1 ) 6; and that„=„1=(n+1 )=2 and¾2=„2¡„2 1=(n2¡1)=12. 2 Example 10.2 Suppose now that Xhas rangef0;1;2;3;:::;ngandpX(j)=¡n j¢ pjqn¡jfor 0•j•n(binomial distribution). Then g(t)=nX j=0etjµn j¶ pjqn¡j =nX j=0µn j¶ (pet)jqn¡j =(pet+q)n: Note that „1=g0(0) =n(pet+q)n¡1petflfl t=0=np ; „2=g00(0) =n(n¡1)p2+np ; so that„=„1=np, and¾2=„2¡„2 1=np(1¡p), as expected. 2 Example 10.3 SupposeXhas rangef1;2;3;:::gandpX(j)=qj¡1pfor allj (geometric distribution). Then g(t)=1X j=1etjqj¡1p =pet 1¡qet: 368 CHAPTER 10. GENERATING FUNCTIONS Here „1=g0(0) =pet (1¡qet)2flflflfl t=0=1 p; „2=g00(0) =pet+pqe2t (1¡qet)3flflflfl t=0=1+q p2; „=„1=1=p, and¾2=„2¡„2 1=q=p2, as computed in Example 6.26. 2 Example 10.4 LetXhave rangef0;1;2;3;:::gand letpX(j)=e¡‚‚j=j! for allj (Poisson distribution with mean ‚). Then g(t)=1X j=0etje¡‚‚j j! =e¡‚1X j=0(‚et)j j! =e¡‚e‚et=e‚(et¡1): Then „1=g0(0) =e‚(et¡1)‚etflflfl t=0=‚; „2=g00(0) =e‚(et¡1)(‚2e2t+‚et)flflfl t=0=‚2+‚; „=„1=‚, and¾2=„2¡„2 1=‚. The variance of the Poisson distribution is easier to obtain in this way than directly from the deflnition (as was done in Exercise 6.2.30). 2 Moment Problem Using the moment generating function, we can now show, at least in the case of a discrete random variable with flnite range, that its distribution function is com-pletely determined by its moments. Theorem 10.1 LetXbe a discrete random variable with flnite range fx 1;x2;:::;xng; and moments „k=E(Xk). Then the moment series g(t)=1X k=0„ktk k! converges for all tto an inflnitely difierentiable function g(t). Proof. We know that „k=nX j=1(xj)kp(xj): 10.1. DISCRETE DISTRIBUTIONS 369 If we setM= maxjxjj, then we have j„kj•nX j=1jxjjkp(xj) •Mk¢nX j=1p(xj)=Mk: Hence, for all Nwe have NX k=0flflflfl„ ktk k!flflflfl•NX k=0(Mjtj)k k!•eMjtj; which shows that the moment series converges for all t. Since it is a power series, we know that its sum is inflnitely difierentiable. This shows that the „kdetermineg(t). Conversely, since „k=g(k)(0), we see thatg(t) determines the „k. 2 Theorem 10.2 LetXbe a discrete random variable with flnite range fx1;x2;:::; xng, distribution function p, and moment generating function g. Thengis uniquely determined by p, and conversely. Proof. We know that pdetermines g, since g(t)=nX j=1etxjp(xj): In this formula, we set aj=p(xj) and, after choosing nconvenient distinct values tioft,w es e tbi=g(ti). Then we have bi=nX j=1etixjaj; or, in matrix notation B=MA: Here B=(bi) and A=(aj) are column n-vectors, and M=(etixj)i sa nn£n matrix. We can solve this matrix equation for A: A=M¡1B; provided only that the matrix Misinvertible (i.e., provided that the determinant ofMis difierent from 0). We can always arrange for this by choosing the values ti=i¡1, since then the determinant of Mis the Vandermonde determinant det0 BBBB@111 ¢¢¢ 1 e tx1etx2etx3¢¢¢etxn e2tx1e2tx2e2tx3¢¢¢e2txn ¢¢¢ e(n¡1)tx1e(n¡1)tx2e(n¡1)tx3¢¢¢e(n¡1)txn1 CCCCA 370 CHAPTER 10. GENERATING FUNCTIONS of theexi, with valueQ i<j(exi¡exj). This determinant is always difierent from 0 if thexjare distinct. 2 If we delete the hypothesis that Xhave flnite range in the above theorem, then the conclusion is no longer necessarily true. Ordinary Generating Functions In the special but important case where the xjare all nonnegative integers, xj=j, we can prove this theorem in a simpler way. In this case, we have g(t)=nX j=0etjp(j); and we see that g(t)i sa polynomial inet. If we write z=et, and deflne the function hby h(z)=nX j=0zjp(j); thenh(z) is a polynomial in zcontaining the same information as g(t), and in fact h(z)=g(logz); g(t)=h(et): The function h(z) is often called the ordinary generating function forX. Note that h(1) =g( 0 )=1 ,h0(1) =g0(0) =„1, andh00(1) =g00(0)¡g0(0) =„2¡„1. It follows from all this that if we know g(t), then we know h(z), and if we know h(z), then we can flnd the p(j) by Taylor’s formula: p(j) = coe–cient of zjinh(z) =h(j)(0) j!: For example, suppose we know that the moments of a certain discrete random variableXare given by „0=1; „k=1 2+2k 4; fork‚1: Then the moment generating function gofXis g(t)=1X k=0„ktk k! =1 +1 21X k=1tk k!+1 41X k=1(2t)k k! =1 4+1 2et+1 4e2t: 10.1. DISCRETE DISTRIBUTIONS 371 This is a polynomial in z=et, and h(z)=1 4+1 2z+1 4z2: Hence,Xmust have range f0;1;2g, andpmust have values f1=4;1=2;1=4g. Properties Both the moment generating function gand the ordinary generating function hhave many properties useful in the study of random variables, of which we can consideronly a few here. In particular, if Xis any discrete random variable and Y=X+a, then g Y(t)=E(etY) =E(et(X+a)) =etaE(etX) =etagX(t); while ifY=bX, then gY(t)=E(etY) =E(etbX) =gX(bt): In particular, if X⁄=X¡„ ¾; then (see Exercise 11) gx⁄(t)=e¡„t=¾gXµt ¾¶ : IfXandYareindependent random variables and Z=X+Yis their sum, withpX,pY, andpZthe associated distribution functions, then we have seen in Chapter 7 that pZis the convolution ofpXandpY, and we know that convolution involves a rather complicated calculation. But for the generating functions we haveinstead the simple relations g Z(t)=gX(t)gY(t); hZ(z)=hX(z)hY(z); that is,gZis simply the product ofgXandgY, and similarly for hZ. To see this, flrst note that if XandYare independent, then etXandetYare independent (see Exercise 5.2.38), and hence E(etXetY)=E(etX)E(etY): 372 CHAPTER 10. GENERATING FUNCTIONS It follows that gZ(t)=E(etZ)=E(et(X+Y)) =E(etX)E(etY) =gX(t)gY(t); and, replacing tby logz, we also get hZ(z)=hX(z)hY(z): Example 10.5 IfXandYare independent discrete random variables with range f0;1;2;:::;ngand binomial distribution pX(j)=pY(j)=µn j¶ pjqn¡j; and ifZ=X+Y, then we know (cf. Section 7.1) that the range of Xis f0;1;2;:::; 2ng andXhas binomial distribution pZ(j)=(pX⁄pY)(j)=µ2n j¶ pjq2n¡j: Here we can easily verify this result by using generating functions. We know that gX(t)=gY(t)=nX j=0etjµn j¶ pjqn¡j =(pet+q)n; and hX(z)=hY(z)=(pz+q)n: Hence, we have gZ(t)=gX(t)gY(t)=(pet+q)2n; or, what is the same, hZ(z)=hX(z)hY(z)=(pz+q)2n =2nX j=0µ2n j¶ (pz)jq2n¡j; from which we can see that the coe–cient of zjis justpZ(j)=¡2n j¢ pjq2n¡j. 2 10.1. DISCRETE DISTRIBUTIONS 373 Example 10.6 IfXandYare independent discrete random variables with the non-negative integers f0;1;2;3;:::gas range, and with geometric distribution func- tion pX(j)=pY(j)=qjp; then gX(t)=gY(t)=p 1¡qet; and ifZ=X+Y, then gZ(t)=gX(t)gY(t) =p2 1¡2qet+q2e2t: If we replace etbyz,w eg e t hZ(z)=p2 (1¡qz)2 =p21X k=0(k+1 )qkzk; and we can read ofi the values of pZ(j) as the coe–cient of zjin this expansion forh(z), even though h(z) is not a polynomial in this case. The distribution pZis a negative binomial distribution (see Section 5.1). 2 Here is a more interesting example of the power and scope of the method of generating functions. Heads or Tails Example 10.7 In the coin-tossing game discussed in Example 1.4, we now consider the question \When is Peter flrst in the lead?" LetXkdescribe the outcome of the kth trial in the game Xk=‰+1;ifkth toss is heads ; ¡1;ifkth toss is tails. Then theXkare independent random variables describing a Bernoulli process. Let S0= 0, and, for n‚1, let Sn=X1+X2+¢¢¢+Xn: ThenSndescribes Peter’s fortune after ntrials, and Peter is flrst in the lead after ntrials ifSk•0 for 1•k<n andSn=1 . Now this can happen when n= 1, in which case S1=X1= 1, or when n>1, in which case S1=X1=¡1. In the latter case, Sk= 0 fork=n¡1, and perhaps for otherkbetween 1 and n. Letmbe the least such value of k; thenSm= 0 and 374 CHAPTER 10. GENERATING FUNCTIONS Sk<0 for 1•k<m . In this case Peter loses on the flrst trial, regains his initial position in the next m¡1 trials, and gains the lead in the next n¡mtrials. Letpbe the probability that the coin comes up heads, and let q=1¡p. Let rnbe the probability that Peter is flrst in the lead after ntrials. Then from the discussion above, we see that rn=0; ifneven; r1=p (= probability of heads in a single toss) ; rn=q(r1rn¡2+r3rn¡4+¢¢¢+rn¡2r1); ifn>1;nodd: Now letTdescribe the time (that is, the number of trials) required for Peter to take the lead. Then Tis a random variable, and since P(T=n)=rn,ris the distribution function for T. We introduce the generating function hT(z) forT: hT(z)=1X n=0rnzn: Then, by using the relations above, we can verify the relation hT(z)=pz+qz(hT(z))2: If we solve this quadratic equation for hT(z), we get hT(z)=1§p 1¡4pqz2 2qz=2pz 1¤p 1¡4pqz2: Of these two solutions, we want the one that has a convergent power series in z (i.e., that is flnite for z= 0). Hence we choose hT(z)=1¡p 1¡4pqz2 2qz=2pz 1+p 1¡4pqz2: Now we can ask: What is the probability that Peter is ever in the lead? This probability is given by (see Exercise 10) 1X n=0rn=hT(1) =1¡p 1¡4pq 2q =1¡jp¡qj 2q =‰ p=q; ifp<q; 1; ifp‚q; so that Peter is sure to be in the lead eventually if p‚q. How long will it take? That is, what is the expected value of T? This value is given by E(T)=h0 T(1) =‰1=(p¡q);ifp>q; 1; ifp=q: 10.1. DISCRETE DISTRIBUTIONS 375 This says that if p>q , then Peter can expect to be in the lead by about 1 =(p¡q) trials, but if p=q, he can expect to wait a long time. A related problem, known as the Gambler’s Ruin problem, is studied in Exer- cise 23 and in Section 12.2. 2 Exercises 1Find the generating functions, both ordinary h(z) and moment g(t), for the following discrete probability distributions. (a) The distribution describing a fair coin. (b) The distribution describing a fair die. (c) The distribution describing a die that always comes up 3. (d) The uniform distribution on the set fn;n+1;n+2;:::;n +kg. (e) The binomial distribution on fn;n+1;n+2;:::;n +kg. (f) The geometric distribution on f0;1;2;:::;gwithp(j)=2=3j+1. 2For each of the distributions (a) through (d) of Exercise 1 calculate the flrst and second moments, „1and„2, directly from their deflnition, and verify that h( 1 )=1 ,h0(1) =„1, andh00(1) =„2¡„1. 3Letpbe a probability distribution on f0;1;2gwith moments „1=1 ,„2=3=2. (a) Find its ordinary generating function h(z). (b) Using (a), flnd its moment generating function. (c) Using (b), flnd its flrst six moments. (d) Using (a), flnd p0,p1, andp2. 4In Exercise 3, the probability distribution is completely determined by its flrst two moments. Show that this is always true for any probability distributiononf0;1;2g.Hint: Given„ 1and„2, flndh(z) as in Exercise 3 and use h(z) to determine p. 5Letpandp0be the two distributions p=µ12 345 1=3002 =30¶ ; p0=µ123 45 02=3001 =3¶ : (a) Show that pandp0have the same flrst and second moments, but not the same third and fourth moments. (b) Find the ordinary and moment generating functions for pandp0. 376 CHAPTER 10. GENERATING FUNCTIONS 6Letpbe the probability distribution p=µ01 2 01=32=3¶ ; and letpn=p⁄p⁄¢¢¢⁄pbe then-fold convolution of pwith itself. (a) Findp2by direct calculation (see Deflnition 7.1). (b) Find the ordinary generating functions h(z) andh2(z) forpandp2, and verify that h2(z)=(h(z))2. (c) Findhn(z) fromh(z). (d) Find the flrst two moments, and hence the mean and variance, of pn fromhn(z). Verify that the mean of pnisntimes the mean of p. (e) Find those integers jfor whichpn(j)>0 fromhn(z). 7LetXbe a discrete random variable with values in f0;1;2;:::;ngand moment generating function g(t). Find, in terms of g(t), the generating functions for (a)¡X. (b)X+1 . (c) 3X. (d)aX+b. 8LetX1,X2, ...,Xnbe an independent trials process, with values in f0;1g and mean„=1=3. Find the ordinary and moment generating functions for the distribution of (a)S1=X1.Hint: First flnd X1explicitly. (b)S2=X1+X2. (c)Sn=X1+X2+¢¢¢+Xn. (d)An=Sn=n. (e)S⁄ n=(Sn¡n„)=p n¾2. 9LetXandYbe random variables with values in f1;2;3;4;5;6gwith distri- bution functions pXandpYgiven by pX(j)=aj; pY(j)=bj: (a) Find the ordinary generating functions hX(z) andhY(z) for these distri- butions. (b) Find the ordinary generating function hZ(z) for the distribution Z= X+Y. 10.2. BRANCHING PROCESSES 377 (c) Show that hZ(z) cannot ever have the form hZ(z)=z2+z3+¢¢¢+z12 11: Hint:hXandhYmust have at least one nonzero root, but hZ(z) in the form given has no nonzero real roots. It follows from this observation that there is no way to load two dice so that the probability that a given sum will turn up when they are tossed is the samefor all sums (i.e., that all outcomes are equally likely). 10Show that if h(z)=1¡p 1¡4pqz2 2qz; then h(1) =‰p=q; ifp•q; 1; ifp‚q; and h0(1) =‰1=(p¡q);ifp>q; 1; ifp=q: 11Show that if Xis a random variable with mean „and variance ¾2, and if X⁄=(X¡„)=¾is the standardized version of X, then gX⁄(t)=e¡„t=¾gXµt ¾¶ : 10.2 Branching Processes Historical Background In this section we apply the theory of generating functions to the study of an important chance process called a branching p rocess. Until recently it was thought that the theory of branching processes originated with the following problem posed by Francis Galton in the Educational Times in 1873.1 Problem 4001: A large nation, of whom we will only concern ourselveswith the adult males, Nin number, and who each bear separate sur- names, colonise a district. Their law of population is such that, in eachgeneration, a 0per cent of the adult males have no male children who reach adult life; a1have one such male child; a2have two; and so on up toa5who have flve. Find (1) what proportion of the surnames will have become extinct afterrgenerations; and (2) how many instances there will be of the same surname being held by mpersons. 1D. G. Kendall, \Branching Processes Since 1873," Journal of London Mathematics Society, vol. 41 (1966), p. 386. 378 CHAPTER 10. GENERATING FUNCTIONS The flrst attempt at a solution was given by Reverend H. W. Watson. Because of a mistake in algebra, he incorrectly concluded that a family name would alwaysdie out with probability 1. However, the methods that he employed to solve theproblems were, and still are, the basis for obtaining the correct solution. Heyde and Seneta discovered an earlier communication by Bienaym¶ e (1845) that anticipated Galton and Watson by 28 years. Bienaym¶ e showed, in fact, that he was aware of the correct solution to Galton’s problem. Heyde and Seneta in their bookI. J. Bienaym¶ e: Statistical Theory Anticipated, 2give the following translation from Bienaym¶ e’s paper: If . . . the mean of the number of male children who replace the number of males of the preceding generation were less than unity, it would beeasily realized that families are dying out due to the disappearance ofthe members of which they are composed. However, the analysis showsfurther that when this mean is equal to unity families tend to disappear,although less rapidly .... The analysis also shows clearly that if the mean ratio is greater than unity, the probability of the extinction of families with the passing oftime no longer reduces to certainty. It only approaches a flnite limit,which is fairly simple to calculate and which has the singular charac-teristic of being given by one of the roots of the equation (in whichthe number of generations is made inflnite) which is not relevant to thequestion when the mean ratio is less than unity. 3 Although Bienaym¶ e does not give his reasoning for these results, he did indicate that he intended to publish a special paper on the problem. The paper was neverwritten, or at least has never been found. In his communication Bienaym¶ e indicated that he was motivated by the same problem that occurred to Galton. The openingparagraph of his paper as translated by Heyde and Seneta says, A great deal of consideration has been given to the possible multipli- cation of the numbers of mankind; and recently various very curiousobservations have been published on the fate which allegedly hangs overthe aristocrary and middle classes; the families of famous men, etc. Thisfate, it is alleged, will inevitably bring about the disappearance of theso-called families ferm¶ ees. 4 A much more extensive discussion of the history of branching processes may be found in two papers by David G. Kendall.5 2C. C. Heyde and E. Seneta, I. J. Bienaym¶ e: Statistical Theory Anticipated (New York: Springer Verlag, 1977). 3ibid., pp. 117{118. 4ibid., p. 118. 5D. G. Kendall, \Branching Processes Since 1873," pp. 385{406; and \The Genealogy of Ge- nealogy: Branching Processes Before (and After) 1873," Bulletin London Mathematics Society, vol. 7 (1975), pp. 225{253. 10.2. BRANCHING PROCESSES 379 2 101/41/4 1/4 1/4 1/4 1/4 1/21/16 1/8 5/16 1/24 3 2 1 0 0121/64 1/325/64 1/81/16 1/16 1/16 1/16 1/2 Figure 10.1: Tree diagram for Example 10.8. Branching processes have served not only as crude models for population growth but also as models for certain physical processes such as chemical and nuclear chainreactions. Problem of Extinction We turn now to the flrst problem posed by Galton (i.e., the problem of flnding theprobability of extinction for a branching process). We start in the 0th generationwith 1 male parent. In the flrst generation we shall have 0, 1, 2, 3, . . . maleofispring with probabilities p 0,p1,p2,p3, .... I f i n t h e flrst generation there are k ofispring, then in the second generation there will be X1+X2+¢¢¢+Xkofispring, whereX1,X2, ...,Xkare independent random variables, each with the common distribution p0,p1,p2, .... This description enables us to construct a tree, and a tree measure, for any number of generations. Examples Example 10.8 Assume that p0=1=2,p1=1=4, andp2=1=4. Then the tree measure for the flrst two generations is shown in Figure 10.1. Note that we use the theory of sums of independent random variables to assign branch probabilities. For example, if there are two ofispring in the flrst generation,the probability that there will be two in the second generation is P(X 1+X2=2 ) =p0p2+p1p1+p2p0 =1 2¢1 4+1 4¢1 4+1 4¢1 2=5 16: We now study the probability that our process dies out (i.e., that at some generation there are no ofispring). 380 CHAPTER 10. GENERATING FUNCTIONS Letdmbe the probability that the process dies out by the mth generation. Of course,d0= 0. In our example, d1=1=2 andd2=1=2+1=8+1=16 = 11=16 (see Figure 10.1). Note that we must add the probabilities for all paths that lead to 0by themth generation. It is clear from the deflnition that 0=d 0•d1•d2•¢¢¢• 1: Hence,dmconverges to a limit d,0•d•1, anddis the probability that the process will ultimately die out. It is this value that we wish to determine. Webegin by expressing the value d min terms of all possible outcomes on the flrst generation. If there are jofispring in the flrst generation, then to die out by the mth generation, each of these lines must die out in m¡1 generations. Since they proceed independently, this probability is ( dm¡1)j. Therefore dm=p0+p1dm¡1+p2(dm¡1)2+p3(dm¡1)3+¢¢¢: (10.1) Leth(z) be the ordinary generating function for the pi: h(z)=p0+p1z+p2z2+¢¢¢: Using this generating function, we can rewrite Equation 10.1 in the form dm=h(dm¡1): (10.2) Sincedm!d, by Equation 10.2 we see that the value dthat we are looking for satisfles the equation d=h(d): (10.3) One solution of this equation is always d= 1, since 1=p0+p1+p2+¢¢¢: This is where Watson made his mistake. He assumed that 1 was the only solution to Equation 10.3. To examine this question more carefully, we flrst note that solutionsto Equation 10.3 represent intersections of the graphs of y=z and y=h(z)=p 0+p1z+p2z2+¢¢¢: Thus we need to study the graph of y=h(z). We note that h(0) =p0. Also, h0(z)=p1+2p2z+3p3z2+¢¢¢; (10.4) and h00(z)=2p2+3¢2p3z+4¢3p4z2+¢¢¢: From this we see that for z‚0,h0(z)‚0 andh00(z)‚0. Thus for nonnegative z,h(z) is an increasing function and is concave upward. Therefore the graph of 10.2. BRANCHING PROCESSES 381 1 1 11 11 0 00 00y z d > 1 d < 1 d = 1 0y = zyy zzy = h (z) 1 1 (a) (c) (b) Figure 10.2: Graphs of y=zandy=h(z). y=h(z) can intersect the line y=zin at most two points. Since we know it must intersect the line y=zat (1;1), we know that there are just three possibilities, as shown in Figure 10.2. In case (a) the equation d=h(d) has rootsfd;1gwith 0•d<1. In the second case (b) it has only the one root d= 1. In case (c) it has two roots f1;dgwhere 1<d. Since we are looking for a solution 0 •d•1, we see in cases (b) and (c) that our only solution is 1. In these cases we can conclude that the process will dieout with probability 1. However in case (a) we are in doubt. We must study thiscase more carefully. From Equation 10.4 we see that h 0(1) =p1+2p2+3p3+¢¢¢=m; wheremis the expected number of ofispring produced by a single parent. In case (a) we haveh0(1)>1, in (b)h0(1) = 1, and in (c) h0(1)<1. Thus our three cases correspond to m> 1,m= 1, andm< 1. We assume now that m> 1. Recall that d0=0 ,d1=h(d0)=p0,d2=h(d1) , ..., a n d dn=h(dn¡1). We can construct these values geometrically, as shown in Figure 10.3. We can see geometrically, as indicated for d0,d1,d2, andd3in Figure 10.3, that the points ( di;h(di)) will always lie above the line y=z. Hence, they must converge to the flrst intersection of the curves y=zandy=h(z) (i.e., to the root d<1). This leads us to the following theorem. 2 Theorem 10.3 Consider a branching process with generating function h(z) for the number of ofispring of a given parent. Let dbe the smallest root of the equation z=h(z). If the mean number mof ofispring produced by a single parent is •1, thend= 1 and the process dies out with probability 1. If m> 1 thend<1 and the process dies out with probability d. 2 We shall often want to know the probability that a branching process dies out by a particular generation, as well as the limit of these probabilities. Let dnbe 382 CHAPTER 10. GENERATING FUNCTIONS y = z y = h(z)y z1 p0 0 d= 0 1 d d d d 123 Figure 10.3: Geometric determination of d. the probability of dying out by the nth generation. Then we know that d1=p0. We know further that dn=h(dn¡1) whereh(z) is the generating function for the number of ofispring produced by a single parent. This makes it easy to computethese probabilities. The program Branch calculates the values of d n. We have run this program for 12 generations for the case that a parent can produce at most two ofispring andthe probabilities for the number produced are p 0=:2,p1=:5, andp2=:3. The results are given in Table 10.1. We see that the probability of dying out by 12 generations is about .6. We shall see in the next example that the probability of eventually dying out is 2/3, so thateven 12 generations is not enough to give an accurate estimate for this probability. We now assume that at most two ofispring can be produced. Then h(z)=p 0+p1z+p2z2: In this simple case the condition z=h(z) yields the equation d=p0+p1d+p2d2; which is satisfled by d= 1 andd=p0=p2. Thus, in addition to the root d=1w e have the second root d=p0=p2. The mean number mof ofispring produced by a single parent is m=p1+2p2=1¡p0¡p2+2p2=1¡p0+p2: Thus, ifp0>p2,m< 1 and the second root is >1. Ifp0=p2, we have a double rootd=1 . I fp0<p2,m> 1 and the second root dis less than 1 and represents the probability that the process will die out. 10.2. BRANCHING PROCESSES 383 Generation Probability of dying out 1. 2 2 .3123 .3852034 .4371165 .4758796 .5058787 .5297138 .5490359 .564949 10 .57822511 .58941612 .598931 Table 10.1: Probability of dying out. p 0=:2092 p1=:2584 p2=:2360 p3=:1593 p4=:0828 p5=:0357 p6=:0133 p7=:0042 p8=:0011 p9=:0002 p10=:0000 Table 10.2: Distribution of number of female children. Example 10.9 Keyfltz6compiled and analyzed data on the continuation of the female family line among Japanese women. His estimates at the basic probabilitydistribution for the number of female children born to Japanese women of ages45{49 in 1960 are given in Table 10.2. The expected number of girls in a family is then 1.837 so the probability dof extinction is less than 1. If we run the program Branch , we can estimate that dis in fact only about .324. 2 Distribution of Ofispring So far we have considered only the flrst of the two problems raised by Galton, namely the probability of extinction. We now consider the second problem, thatis, the distribution of the number Z nof ofispring in the nth generation. The exact form of the distribution is not known except in very special cases. We shall see, 6N. Keyfltz, Introduction to the Mathematics of Population, rev. ed. (Reading, PA: Addison Wesley, 1977). 384 CHAPTER 10. GENERATING FUNCTIONS however, that we can describe the limiting behavior of Znasn!1 . We flrst show that the generating function hn(z) of the distribution of Zncan be obtained from h(z) for any branching process. We recall that the value of the generating function at the value zfor any random variableXcan be written as h(z)=E(zX)=p0+p1z+p2z2+¢¢¢: That is,h(z) is the expected value of an experiment which has outcome zjwith probability pj. LetSn=X1+X2+¢¢¢+Xnwhere each Xjhas the same integer-valued distribution ( pj) with generating function k(z)=p0+p1z+p2z2+¢¢¢:Letkn(z) be the generating function of Sn. Then using one of the properties of ordinary generating functions discussed in Section 10.1, we have kn(z)=(k(z))n; since theXj’s are independent and all have the same distribution. Consider now the branching process Zn. Lethn(z) be the generating function ofZn. Then hn+1(z)=E(zZn+1) =X kE(zZn+1jZn=k)P(Zn=k): IfZn=k, thenZn+1=X1+X2+¢¢¢+XkwhereX1,X2, ...,Xkare independent random variables with common generating function h(z). Thus E(zZn+1jZn=k)=E(zX1+X2+¢¢¢+Xk)=(h(z))k; and hn+1(z)=X k(h(z))kP(Zn=k): But hn(z)=X kP(Zn=k)zk: Thus, hn+1(z)=hn(h(z)): (10.5) Hence the generating function for Z2ish2(z)=h(h(z)), forZ3is h3(z)=h(h(h(z))); and so forth. From this we see also that hn+1(z)=h(hn(z)): (10.6) If we difierentiate Equation 10.6 and use the chain rule we have h0 n+1(z)=h0(hn(z))h0 n(z): 10.2. BRANCHING PROCESSES 385 Puttingz= 1 and using the fact that hn(1) = 1 and h0 n(1) =mn= the mean number of ofispring in the n’th generation, we have mn+1=m¢mn: Thus,m2=m¢m=m2,m3=m¢m2=m3, and in general mn=mn: Thus, for a branching process with m> 1, the mean number of ofispring grows exponentially at a rate m. Examples Example 10.10 For the branching process of Example 10.8 we have h(z)=1=2+( 1=4)z+( 1=4)z2; h2(z)=h(h(z) )=1=2+( 1=4)[1=2+( 1=4)z+( 1=4)z2] = +(1=4)[1=2+( 1=4)z+( 1=4)z2]2 =1 1=1 6+( 1=8)z+( 9=64)z2+( 1=32)z3+( 1=64)z4: The probabilities for the number of ofispring in the second generation agree with those obtained directly from the tree measure (see Figure 1). 2 It is clear that even in the simple case of at most two ofispring, we cannot easily carry out the calculation of hn(z) by this method. However, there is one special case in which this can be done. Example 10.11 Assume that the probabilities p1,p2, . . . form a geometric series: pk=bck¡1,k=1 , 2 , ..., with 0<b•1¡cand p0=1¡p1¡p2¡¢¢¢ =1¡b¡bc¡bc2¡¢¢¢ =1¡b 1¡c: Then the generating function h(z) for this distribution is h(z)=p0+p1z+p2z2+¢¢¢ =1¡b 1¡c+bz+bcz2+bc2z3+¢¢¢ =1¡b 1¡c+bz 1¡cz: From this we flnd h0(z)=bcz (1¡cz)2+b 1¡cz=b (1¡cz)2 386 CHAPTER 10. GENERATING FUNCTIONS and m=h0(1) =b (1¡c)2: We know that if m•1 the process will surely die out and d= 1. To flnd the probability dwhenm> 1w em u s tfl n dar o o t d<1 of the equation z=h(z); or z=1¡b 1¡c+bz 1¡cz: This leads us to a quadratic equation. We know that z= 1 is one solution. The other is found to be d=1¡b¡c c(1¡c): It is easy to verify that d<1 just when m> 1. It is possible in this case to flnd the distribution of Zn. This is done by flrst flnding the generating function hn(z).7The result for m6= 1 is: hn(z)=1¡mn•1¡d mn¡d‚ +mnh 1¡d mn¡di2 z 1¡h mn¡1 mn¡di z: The coe–cients of the powers of zgive the distribution for Zn: P(Zn=0 )=1¡mn1¡d mn¡d=d(mn¡1) mn¡d and P(Zn=j)=mn‡1¡d mn¡d·2 ¢‡mn¡1 mn¡d·j¡1 ; forj‚1. 2 Example 10.12 Let us re-examine the Keyfltz data to see if a distribution of the type considered in Example 10.11 could reasonably be used as a model for thispopulation. We would have to estimate from the data the parameters bandcfor the formula p k=bck¡1. Recall that m=b (1¡c)2(10.7) and the probability dthat the process dies out is d=1¡b¡c c(1¡c): (10.8) Solving Equation 10.7 and 10.8 for bandcgives c=m¡1 m¡d 7T. E. Harris, The Theory of Branching P rocesses (Berlin: Springer, 1963), p. 9. 10.2. BRANCHING PROCESSES 387 Geometric pjData Model 0 .2092 .18161 .2584 .36662 .2360 .20283 .1593 .11224 .0828 .06215 .0357 .03446 .0133 .01907 .0042 .01058 .0011 .00589 .0002 .0032 10 .0000 .0018 Table 10.3: Comparison of observed and expected frequencies. and b=m‡1¡d m¡d·2 : We shall use the value 1.837 for mand .324 for dthat we found in the Keyfltz example. Using these values, we obtain b=:3666 andc=:5533. Note that (1¡c)2<b< 1¡c, as required. In Table 10.3 we give for comparison the probabilities p0throughp8as calculated by the geometric distribution versus the empirical values. The geometric model tends to favor the larger numbers of ofispring but is similar enough to show that this modifled geometric distribution might be appropriate touse for studies of this kind. Recall that if S n=X1+X2+¢¢¢+Xnis the sum of independent random variables with the same distribution then the Law of Large Numbers states thatS n=nconverges to a constant, namely E(X1). It is natural to ask if there is a similar limiting theorem for branching processes. Consider a branching process with Znrepresenting the number of ofispring after ngenerations. Then we have seen that the expected value of Znismn. T h u sw ec a n scale the random variable Znto have expected value 1 by considering the random variable Wn=Zn mn: In the theory of branching processes it is proved that this random variable Wn will tend to a limit as ntends to inflnity. However, unlike the case of the Law of Large Numbers where this limit is a constant, for a branching process the limitingvalue of the random variables W nis itself a random variable. Although we cannot prove this theorem here we can illustrate it by simulation. This requires a little care. When a branching process survives, the number ofofispring is apt to get very large. If in a given generation there are 1000 ofispring,the ofispring of the next generation are the result of 1000 chance events, and it willtake a while to simulate these 1000 experiments. However, since the flnal result is 388 CHAPTER 10. GENERATING FUNCTIONS 5 10 15 20 250.511.522.53 Figure 10.4: Simulation of Zn=mnfor the Keyfltz example. the sum of 1000 independent experiments we can use the Central Limit Theorem to replace these 1000 experiments by a single experiment with normal density havingthe appropriate mean and variance. The program BranchingSimulation carries out this process. We have run this program for the Keyfltz example, carrying out 10 simulations and graphing the results in Figure 10.4. The expected number of female ofispring per female is 1.837, so that we are graphing the outcome for the random variables W n=Zn=(1:837)n. For three of the simulations the process died out, which is consistent with the value d=:3 that we found for this example. For the other seven simulations the value of Wntends to a limiting value which is difierent for each simulation. 2 Example 10.13 We now examine the random variable Znmore closely for the casem< 1 (see Example 10.11). Fix a value t>0; let [tmn] be the integer part of tmn. Then P(Zn=[tmn]) =mn(1¡d mn¡d)2(mn¡1 mn¡d)[tmn]¡1 =1 mn(1¡d 1¡d=mn)2(1¡1=mn 1¡d=mn)tmn+a; wherejaj•2. Thus, as n!1 , mnP(Zn=[tmn])!(1¡d)2e¡t e¡td=( 1¡d)2e¡t(1¡d): Fort=0 , P(Zn=0 )!d: 10.2. BRANCHING PROCESSES 389 We can compare this result with the Central Limit Theorem for sums Snof integer- valued independent random variables (see Theorem 9.3), which states that if tis an integer and u=(t¡n„)=p ¾2n, then asn!1 , p ¾2nP(Sn=up ¾2n+„n)!1p 2…e¡u2=2: We see that the form of these statements are quite similar. It is possible to prove a limit theorem for a general class of branching processes that states that undersuitable hypotheses, as n!1 , m nP(Zn=[tmn])!k(t); fort>0, and P(Zn=0 )!d: However, unlike the Central Limit Theorem for sums of independent random vari- ables, the function k(t) will depend upon the basic distribution that determines the process. Its form is known for only a very few examples similar to the one we haveconsidered here. 2 Chain Letter Problem Example 10.14 An interesting example of a branching process was suggested by Free Huizinga.8In 1978, a chain letter called the \Circle of Gold," believed to have started in California, found its way across the country to the theater district of NewYork. The chain required a participant to buy a letter containing a list of 12 namesfor 100 dollars. The buyer gives 50 dollars to the person from whom the letter waspurchased and then sends 50 dollars to the person whose name is at the top of thelist. The buyer then crosses ofi the name at the top of the list and adds her ownname at the bottom in each letter before it is sold again. Let us flrst assume that the buyer may sell the letter only to a single person. If you buy the letter you will want to compute your expected winnings. (We areignoring here the fact that the passing on of chain letters through the mail is afederal ofiense with certain obvious resulting penalties.) Assume that each personinvolved has a probability pof selling the letter. Then you will receive 50 dollars with probability pand another 50 dollars if the letter is sold to 12 people, since then your name would have risen to the top of the list. This occurs with probability p 12, and so your expected winnings are ¡100 + 50p+5 0p12. Thus the chain in this situation is a highly unfavorable game. It would be more reasonable to allow each person involved to make a copy of the list and try to sell the letter to at least 2 other people. Then you would havea chance of recovering your 100 dollars on these sales, and if any of the letters issold 12 times you will receive a bonus of 50 dollars for each of these cases. We canconsider this as a branching process with 12 generations. The members of the flrst 8Private communication. 390 CHAPTER 10. GENERATING FUNCTIONS generation are the letters you sell. The second generation consists of the letters sold by members of the flrst generation, and so forth. Let us assume that the probabilities that each individual sells letters to 0, 1, or 2 others are p0,p1, andp2, respectively. Let Z1,Z2, ...,Z12be the number of letters in the flrst 12 generations of this branching process. Then your expectedwinnings are 50(E(Z 1)+E(Z12) )=5 0m+5 0m12; wherem=p1+2p2is the expected number of letters you sold. Thus to be favorable we just have 50m+5 0m12>100; or m+m12>2: But this will be true if and only if m> 1. We have seen that this will occur in the quadratic case if and only if p2>p0. Let us assume for example that p0=:2, p1=:5, andp2=:3. Thenm=1:1 and the chain would be a favorable game. Your expected proflt would be 50(1:1+1:112)¡100…112: The probability that you receive at least one payment from the 12th generation is 1¡d12. We flnd from our program Branch thatd12=:599. Thus, 1¡d12=:401 is the probability that you receive some bonus. The maximum that you could receivefrom the chain would be 50(2 + 2 12) = 204;900 if everyone were to successfully sell two letters. Of course you can not always expect to be so lucky. (What is theprobability of this happening?) To simulate this game, we need only simulate a branching process for 12 gen- erations. Using a slightly modifled version of our program BranchingSimulation we carried out twenty such simulations, giving the results shown in Table 10.4. Note that we were quite lucky on a few runs, but we came out ahead only a little less than half the time. The process died out by the twelfth generation in 12out of the 20 experiments, in good agreement with the probability d 12=:599 that we calculated using the program Branch . Let us modify the assumptions about our chain letter to let the buyer sell the letter to as many people as she can instead of to a maximum of two. We shallassume, in fact, that a person has a large number Nof acquaintances and a small probability pof persuading any one of them to buy the letter. Then the distribution for the number of letters that she sells will be a binomial distribution with meanm=Np. SinceNis large and pis small, we can assume that the probability p j that an individual sells the letter to jpeople is given by the Poisson distribution pj=e¡mmj j!: 10.2. BRANCHING PROCESSES 391 Z1Z2Z3Z4Z5Z6Z7Z8Z9Z10Z11Z12Proflt 10000000 0000 - 5 0 11232321 23362 5 0 00000000 0000 -100 24423443 2211 5 0 12354333 58662 5 0 00000000 0000 -100 23222123 33463 0 0 12111121 0000 - 5 0 00000000 0000 -100 10000000 0000 - 5 0 23233359 1 21 21 31 5 7 5 011100000 0000 - 5 0 12233000 0000 - 5 0 11112234 46452 0 0 11000000 0000 - 5 0 10000000 0000 - 5 0 10000000 0000 - 5 0 11233423 3332 5 0 124669 1 0 1 3 1 61 71 51 8 8 5 010000000 0000 - 5 0 Table 10.4: Simulation of chain letter (flnite distribution case). 392 CHAPTER 10. GENERATING FUNCTIONS Z1Z2Z3Z4Z5Z6Z7Z8Z9Z10Z11Z12Proflt 126778 1 19 76652 0 0 10000000 0000 - 5 0 10000000 0000 - 5 0 11100000 0000 - 5 0 00000000 0000 -100 11111124 97973 0 0 23342000 0000 0 10000000 0000 - 5 0 21000000 0000 0 3347 1 1 1 7 1 4 1 1 1 11 01 62 5 1300 00000000 0000 -100 12211310 0000 - 5 0 00000000 0000 -100 23100000 0000 0 31000000 0000 5 0 10000000 0000 - 5 0 3447 1 0 1 19 1 1 1 21 41 31 0 5 5 013349579 88631 0 0 104669 1 0 1 3 0000 - 5 0 10000000 0000 - 5 0 Table 10.5: Simulation of chain letter (Poisson case). The generating function for the Poisson distribution is h(z)=1X j=0e¡mmjzj j! =e¡m1X j=0mjzj j! =e¡memz=em(z¡1): The expected number of letters that an individual passes on is m, and again to be favorable we must have m> 1. Let us assume again that m=1:1. Then we can flnd again the probability 1 ¡d12of a bonus from Branch . The result is .232. Although the expected winnings are the same, the variance is larger in this case,and the buyer has a better chance for a reasonably large proflt. We again carriedout 20 simulations using the Poisson distribution with mean 1.1. The results areshown in Table 10.5. We note that, as before, we came out ahead less than half the time, but we also had one large proflt. In only 6 of the 20 cases did we receive any proflt. This isagain in reasonable agreement with our calculation of a probability .232 for thishappening. 2 10.2. BRANCHING PROCESSES 393 Exercises 1LetZ1,Z2, ...,ZNdescribe a branching process in which each parent has jofispring with probability pj. Find the probability dthat the process even- tually dies out if (a)p0=1=2,p1=1=4, andp2=1=4. (b)p0=1=3,p1=1=3, andp2=1=3. (c)p0=1=3,p1= 0, andp2=2=3. (d)pj=1=2j+1, forj=0 , 1 , 2 , .... (e)pj=( 1=3)(2=3)j, forj=0 , 1 , 2 , .... (f)pj=e¡22j=j!, forj= 0, 1, 2, . . . (estimate dnumerically). 2LetZ1,Z2, ...,ZNdescribe a branching process in which each parent has jofispring with probability pj. Find the probability dthat the process dies out if (a)p0=1=2,p1=p2= 0, andp3=1=2. (b)p0=p1=p2=p3=1=4. (c)p0=t,p1=1¡2t,p2= 0, andp3=t, wheret•1=2. 3In the chain letter problem (see Example 10.14) flnd your expected proflt if (a)p0=1=2,p1= 0, andp2=1=2. (b)p0=1=6,p1=1=2, andp2=1=3. Show that if p0>1=2, you cannot expect to make a proflt. 4LetSN=X1+X2+¢¢¢+XN, where the Xi’s are independent random variables with common distribution having generating function f(z). Assume thatNis an integer valued random variable independent of all of the Xjand having generating function g(z). Show that the generating function for SNis h(z)=g(f(z)).Hint: Use the fact that h(z)=E(zSN)=X kE(zSNjN=k)P(N=k): 5We have seen that if the generating function for the ofispring of a single parent isf(z), then the generating function for the number of ofispring after two generations is given by h(z)=f(f(z)). Explain how this follows from the result of Exercise 4. 6Consider a queueing process (see Example 5.7) such that in each minute either 0 or 1 customers arrive with probabilities porq=1¡p, respectively. (The numberpis called the arrival rate .) When a customer starts service she flnishes in the next minute with probability r. The number ris called the service rate .) Thus when a customer begins being served she will flnish being served injminutes with probability (1 ¡r)j¡1r, forj=1 , 2 , 3 , .... 394 CHAPTER 10. GENERATING FUNCTIONS (a) Find the generating function f(z) for the number of customers who arrive in one minute and the generating function g(z) for the length of time that a person spends in service once she begins service. (b) Consider a customer branching p rocess by considering the ofispring of a customer to be the customers who arrive while she is being served. UsingExercise 4, show that the generating function for our customer branchingprocess ish(z)=g(f(z)). (c) If we start the branching process with the arrival of the flrst customer, then the length of time until the branching process dies out will be thebusy period for the server. Find a condition in terms of the arrival rate and service rate that will assure that the server will ultimately have atime when he is not busy. 7LetNbe the expected total number of ofispring in a branching process. Let mbe the mean number of ofispring of a single parent. Show that N=1+‡X p k¢k· N=1+mN and hence that Nis flnite if and only if m< 1 and in that case N=1=(1¡m). 8Consider a branching process such that the number of ofispring of a parent is jwith probability 1 =2j+1forj=0 , 1 , 2 , .... (a) Using the results of Example 10.11 show that the probability that there arejofispring in the nth generation is p(n) j=‰1 n(n+1)(n n+1)j;ifj‚1; n n+1; ifj=0: (b) Show that the probability that the process dies out exactly at the nth generation is 1 =n(n+ 1). (c) Show that the expected lifetime is inflnite even though d=1 . 10.3 Generating Functions for Continuous Densi- ties In the previous section, we introduced the concepts of moments and moment gen- erating functions for discrete random variables. These concepts have natural ana-logues for continuous random variables, provided some care is taken in argumentsinvolving convergence. Moments IfXis a continuous random variable deflned on the probability space ›, with density function fX, then we deflne the nth moment of Xby the formula „n=E(Xn)=Z+1 ¡1xnfX(x)dx ; 10.3. CONTINUOUS DENSITIES 395 provided the integral „n=E(Xn)=Z+1 ¡1jxjnfX(x)dx ; is flnite. Then, just as in the discrete case, we see that „0=1 ,„1=„, and „2¡„2 1=¾2. Moment Generating Functions Now we deflne the moment generating function g(t) forXby the formula g(t)=1X k=0„ktk k!=1X k=0E(Xk)tk k! =E(etX)=Z+1 ¡1etxfX(x)dx ; provided this series converges. Then, as before, we have „n=g(n)(0): Examples Example 10.15 LetXbe a continuous random variable with range [0 ;1] and density function fX(x) = 1 for 0•x•1 (uniform density). Then „n=Z1 0xndx=1 n+1; and g(t)=1X k=0tk (k+ 1)! =et¡1 t: Here the series converges for all t. Alternatively, we have g(t)=Z+1 ¡1etxfX(x)dx =Z1 0etxdx=et¡1 t: Then (by L’H^ opital’s rule) „0=g(0) = lim t!0et¡1 t=1; „1=g0(0) = lim t!0tet¡et+1 t2=1 2; „2=g00(0) = lim t!0t3et¡2t2et+2tet¡2t t4=1 3: 396 CHAPTER 10. GENERATING FUNCTIONS In particular, we verify that „=g0( 0 )=1=2 and ¾2=g00(0)¡(g0(0))2=1 3¡1 4=1 12 as before (see Example 6.25). 2 Example 10.16 LetXhave range [ 0 ;1) and density function fX(x)=‚e¡‚x (exponential density with parameter ‚). In this case „n=Z1 0xn‚e¡‚xdx=‚(¡1)ndn d‚nZ1 0e¡‚xdx =‚(¡1)ndn d‚n[1 ‚]=n! ‚n; and g(t)=1X k=0„ktk k! =1X k=0[t ‚]k=‚ ‚¡t: Here the series converges only for jtj<‚. Alternatively, we have g(t)=Z1 0etx‚e¡‚xdx =‚e(t¡‚)x t¡‚flflflfl1 0=‚ ‚¡t: Now we can verify directly that „n=g(n)(0) =‚n! (‚¡t)n+1flflflfl t=0=n! ‚n: 2 Example 10.17 LetXhave range (¡1;+1) and density function fX(x)=1p 2…e¡x2=2 (normal density). In this case we have „n=1p 2…Z+1 ¡1xne¡x2=2dx =‰(2m)! 2mm!;ifn=2m, 0; ifn=2m+1 . 10.3. CONTINUOUS DENSITIES 397 (These moments are calculated by integrating once by parts to show that „n= (n¡1)„n¡2, and observing that „0= 1 and„1= 0.) Hence, g(t)=1X n=0„ntn n! =1X m=0t2m 2mm!=et2=2: This series converges for all values of t. Again we can verify that g(n)(0) =„n. LetXbe a normal random variable with parameters „and¾. It is easy to show that the moment generating function of Xis given by et„+(¾2=2)t2: Now suppose that XandYare two independent normal random variables with parameters „1,¾1, and„2,¾2, respectively. Then, the product of the moment generating functions of XandYis et(„1+„2)+((¾2 1+¾2 2)=2)t2: This is the moment generating function for a normal random variable with mean „1+„2and variance ¾2 1+¾2 2. Thus, the sum of two independent normal random variables is again normal. (This was proved for the special case that both summandsare standard normal in Example 7.5.) 2 In general, the series deflning g(t) will not converge for all t. But in the important special case where Xis bounded (i.e., where the range of Xis contained in a flnite interval), we can show that the series does converge for all t. Theorem 10.4 SupposeXis a continuous random variable with range contained in the interval [¡M;M ]. Then the series g(t)=1X k=0„ktk k! converges for all tto an inflnitely difierentiable function g(t), andg(n)(0) =„n. Proof. We have „k=Z+M ¡MxkfX(x)dx ; so j„kj•Z+M ¡MjxjkfX(x)dx •MkZ+M ¡MfX(x)dx=Mk: 398 CHAPTER 10. GENERATING FUNCTIONS Hence, for all Nwe have NX k=0flflflfl„ ktk k!flflflfl•NX k=0(Mjtj)k k!•eMjtj; which shows that the power series converges for all t. We know that the sum of a convergent power series is always difierentiable. 2 Moment Problem Theorem 10.5 IfXis a bounded random variable, then the moment generating functiongX(t)o fxdetermines the density function fX(x) uniquely. Sketch of the Proof. We know that gX(t)=1X k=0„ktk k! =Z+1 ¡1etxf(x)dx : If we replace tbyi¿, where¿is real and i=p¡1, then the series converges for all¿, and we can deflne the function kX(¿)=gX(i¿)=Z+1 ¡1ei¿xfX(x)dx : The function kX(¿) is called the characteristic function ofX, and is deflned by the above equation even when the series for gXdoes not converge. This equation says thatkXis the Fourier transform offX. It is known that the Fourier transform has an inverse, given by the formula fX(x)=1 2…Z+1 ¡1e¡i¿xkX(¿)d¿ ; suitably interpreted.9Here we see that the characteristic function kX, and hence the moment generating function gX, determines the density function fXuniquely under our hypotheses. 2 Sketch of the Proof of the Central Limit Theorem With the above result in mind, we can now sketch a proof of the Central Limit Theorem for bounded continuous random variables (see Theorem 9.6). To this end,letXbe a continuous random variable with density function f X, mean„= 0 and variance¾2= 1, and moment generating function g(t) deflned by its series for all t. 9H. Dym and H. P. McKean, Fourier Series and Integrals (New York: Academic Press, 1972). 10.3. CONTINUOUS DENSITIES 399 LetX1,X2, ...,Xnbe an independent trials process with each Xihaving density fX, and letSn=X1+X2+¢¢¢+Xn, andS⁄ n=(Sn¡n„)=p n¾2=Sn=pn. Then eachXihas moment generating function g(t), and since the Xiare independent, the sumSn, just as in the discrete case (see Section 10.1), has moment generating function gn(t)=(g(t))n; and the standardized sum S⁄ nhas moment generating function g⁄ n(t)=µ gµtpn¶¶n : We now show that, as n!1 ,g⁄ n(t)!et2=2, whereet2=2is the moment gener- ating function of the normal density n(x)=( 1=p 2…)e¡x2=2(see Example 10.17). To show this, we set u(t) = logg(t), and u⁄ n(t) = logg⁄ n(t) =nloggµtpn¶ =nuµtpn¶ ; and show that u⁄ n(t)!t2=2a sn!1 . First we note that u(0) = log gn( 0 )=0; u0(0) =g0(0) g(0)=„1 1=0; u00(0) =g00(0)g(0)¡(g0(0))2 (g(0))2 =„2¡„2 1 1=¾2=1: Now by using L’H^ opital’s rule twice, we get lim n!1u⁄ n(t) = lim s!1u(t=ps) s¡1 = lim s!1u0(t=ps)t 2s¡1=2 = lim s!1u00µtps¶t2 2=¾2t2 2=t2 2: Hence,g⁄ n(t)!et2=2asn!1 . Now to complete the proof of the Central Limit Theorem, we must show that if g⁄ n(t)!et2=2, then under our hypotheses the distribution functions F⁄ n(x)o ft h eS⁄ nmust converge to the distribution function F⁄ N(x) of the normal variable N; that is, that F⁄ n(a)=P(S⁄ n•a)!1p 2…Za ¡1e¡x2=2dx ; and furthermore, that the density functions f⁄ n(x)o ft h eS⁄ nmust converge to the density function for N; that is, that f⁄ n(x)!1p 2…e¡x2=2; 400 CHAPTER 10. GENERATING FUNCTIONS asn!1 . Since the densities, and hence the distributions, of the S⁄ nare uniquely deter- mined by their moment generating functions under our hypotheses, these conclu-sions are certainly plausible, but their proofs involve a detailed examination ofcharacteristic functions and Fourier transforms, and we shall not attempt themhere. In the same way, we can prove the Central Limit Theorem for bounded discrete random variables with integer values (see Theorem 9.4). Let Xbe a discrete random variable with density function p(j), mean„= 0, variance ¾ 2= 1, and moment generating function g(t), and letX1,X2,...,Xnform an independent trials process with common density p. LetSn=X1+X2+¢¢¢+XnandS⁄ n=Sn=pn, with densitiespnandp⁄ n, and moment generating functions gn(t) andg⁄ n(t)=‡ g(tpn)·n : Then we have g⁄ n(t)!et2=2; just as in the continuous case, and this implies in the same way that the distribution functionsF⁄ n(x) converge to the normal distribution; that is, that F⁄ n(a)=P(S⁄ n•a)!1p 2…Za ¡1e¡x2=2dx ; asn!1 . The corresponding statement about the distribution functions p⁄ n, however, re- quires a little extra care (see Theorem 9.3). The trouble arises because the dis-tributionp(x) is not deflned for all x, but only for integer x. It follows that the distribution p ⁄ n(x) is deflned only for xof the form j=pn, and these values change asnchanges. We can flx this, however, by introducing the function „ p(x), deflned by the for- mula „p(x)=‰p(j);ifj¡1=2•x<j +1=2, 0; otherwise: Then „p(x) is deflned for all x,„p(j)=p(j), and the graph of „ p(x) is the step function for the distribution p(j) (see Figure 3 of Section 9.1). In the same way we introduce the step function „ pn(x) and „p⁄ n(x) associated with the distributions pnandp⁄ n, and their moment generating functions „ gn(t) and „g⁄ n(t). If we can show that „ g⁄ n(t)!et2=2, then we can conclude that „p⁄ n(x)!1p 2…et2=2; asn!1 , for allx, a conclusion strongly suggested by Figure 9.3. Now „g(t) is given by „g(t)=Z+1 ¡1etx„p(x)dx =+NX j=¡NZj+1=2 j¡1=2etxp(j)dx 10.3. CONTINUOUS DENSITIES 401 =+NX j=¡Np(j)etjet=2¡e¡t=2 2t=2 =g(t)sinh(t=2) t=2; where we have put sinh(t=2) =et=2¡e¡t=2 2: In the same way, we flnd that „gn(t)=gn(t)sinh(t=2) t=2; „g⁄ n(t)=g⁄ n(t)sinh(t=2pn) t=2pn: Now, asn!1 , we know that g⁄ n(t)!et2=2, and, by L’H^ opital’s rule, lim n!1sinh(t=2pn) t=2pn=1: It follows that „g⁄ n(t)!et2=2; and hence that „p⁄ n(x)!1p 2…e¡x2=2; asn!1 . The astute reader will note that in this sketch of the proof of Theo- rem 9.3, we never made use of the hypothesis that the greatest common divisor ofthe difierences of all the values that the X ican take on is 1. This is a technical point that we choose to ignore. A complete proof may be found in Gnedenko andKolmogorov. 10 Cauchy Density The characteristic function of a continuous density is a useful tool even in cases whenthe moment series does not converge, or even in cases when the moments themselvesare not flnite. As an example, consider the Cauchy density with parameter a=1 (see Example 5.10) f(x)=1 …(1 +x2): IfXandYare independent random variables with Cauchy density f(x), then the averageZ=(X+Y)=2 also has Cauchy density f(x), that is, fZ(x)=f(x): 10B. V. Gnedenko and A. N. Kolomogorov, Limit Distributions for Sums of Independent Random Variables (Reading: Addison-Wesley, 1968), p. 233. 402 CHAPTER 10. GENERATING FUNCTIONS This is hard to check directly, but easy to check by using characteristic functions. Note flrst that „2=E(X2)=Z+1 ¡1x2 …(1 +x2)dx=1 so that„2is inflnite. Nevertheless, we can deflne the characteristic function kX(¿) ofxby the formula kX(¿)=Z+1 ¡1ei¿x 1 …(1 +x2)dx : This integral is easy to do by contour methods, and gives us kX(¿)=kY(¿)=e¡j¿j: Hence, kX+Y(¿)=(e¡j¿j)2=e¡2j¿j; and since kZ(¿)=kX+Y(¿=2); we have kZ(¿)=e¡2j¿=2j=e¡j¿j: This shows that kZ=kX=kY, and leads to the conclusions that fZ=fX=fY. It follows from this that if X1,X2,...,Xnis an independent trials process with common Cauchy density, and if An=X1+X2+¢¢¢+Xn n is the average of the Xi, thenAnhas the same density as do the Xi. This means that the Law of Large Numbers fails for this process; the distribution of the averageA nis exactly the same as for the individual terms. Our proof of the Law of Large Numbers fails in this case because the variance of Xiis not flnite. Exercises 1LetXbe a continuous random variable with values in [ 0 ;2] and density fX. Find the moment generating function g(t) forXif (a)fX(x)=1=2. (b)fX(x)=( 1=2)x. (c)fX(x)=1¡(1=2)x. (d)fX(x)=j1¡xj. (e)fX(x)=( 3=8)x2. Hint: Use the integral deflnition, as in Examples 10.15 and 10.16. 2For each of the densities in Exercise 1 calculate the flrst and second moments, „1and„2, directly from their deflnition and verify that g( 0 )=1 ,g0(0) =„1, andg00(0) =„2. 10.3. CONTINUOUS DENSITIES 403 3LetXbe a continuous random variable with values in [ 0 ;1) and density fX. Find the moment generating functions for Xif (a)fX(x)=2e¡2x. (b)fX(x)=e¡2x+( 1=2)e¡x. (c)fX(x)=4xe¡2x. (d)fX(x)=‚(‚x)n¡1e¡‚x=(n¡1)!. 4For each of the densities in Exercise 3, calculate the flrst and second moments, „1and„2, directly from their deflnition and verify that g( 0 )=1 ,g0(0) =„1, andg00(0) =„2. 5Find the characteristic function kX(¿) for each of the random variables Xof Exercise 1. 6LetXbe a continuous random variable whose characteristic function kX(¿) is kX(¿)=e¡j¿j;¡1<¿< +1: Show directly that the density fXofXis fX(x)=1 …(1 +x2): 7LetXbe a continuous random variable with values in [ 0 ;1], uniform density functionfX(x)·1 and moment generating function g(t)=(et¡1)=t. Find in terms of g(t) the moment generating function for (a)¡X. (b) 1 +X. (c) 3X. (d)aX+b. 8LetX1,X2, ...,Xnbe an independent trials process with uniform density. Find the moment generating function for (a)X1. (b)S2=X1+X2. (c)Sn=X1+X2+¢¢¢+Xn. (d)An=Sn=n. (e)S⁄ n=(Sn¡n„)=p n¾2. 9LetX1,X2, ...,Xnbe an independent trials process with normal density of mean 1 and variance 2. Find the moment generating function for (a)X1. (b)S2=X1+X2. 404 CHAPTER 10. GENERATING FUNCTIONS (c)Sn=X1+X2+¢¢¢+Xn. (d)An=Sn=n. (e)S⁄ n=(Sn¡n„)=p n¾2. 10LetX1,X2, ...,Xnbe an independent trials process with density f(x)=1 2e¡jxj;¡1<x< +1: (a) Find the mean and variance of f(x). (b) Find the moment generating function for X1,Sn,An, andS⁄ n. (c) What can you say about the moment generating function of S⁄ nasn! 1? (d) What can you say about the moment generating function of Anasn! 1? Chapter 11 Markov Chains 11.1 Introduction Most of our study of probability has dealt with independent trials processes. These processes are the basis of classical probability theory and much of statistics. Wehave discussed two of the principal theorems for these processes: the Law of LargeNumbers and the Central Limit Theorem. We have seen that when a sequence of chance experiments forms an indepen- dent trials process, the possible outcomes for each experiment are the same andoccur with the same probability. Further, knowledge of the outcomes of the pre-vious experiments does not in°uence our predictions for the outcomes of the nextexperiment. The distribution for the outcomes of a single experiment is su–cientto construct a tree and a tree measure for a sequence of nexperiments, and we can answer any probability question about these experiments by using this treemeasure. Modern probability theory studies chance processes for which the knowledge of previous outcomes in°uences predictions for future experiments. In principle,when we observe a sequence of chance experiments, all of the past outcomes couldin°uence our predictions for the next experiment. For example, this should be thecase in predicting a student’s grades on a sequence of exams in a course. But toallow this much generality would make it very di–cult to prove general results. In 1907, A. A. Markov began the study of an important new type of chance process. In this process, the outcome of a given experiment can afiect the outcomeof the next experiment. This type of process is called a Markov chain. Specifying a Markov Chain We describe a Markov chain as follows: We have a set of states,S=fs1;s2;:::;srg. The process starts in one of these states and moves successively from one state toanother. Each move is called a step. If the chain is currently in state s i, then it moves to state sjat the next step with a probability denoted by pij, and this probability does not depend upon which states the chain was in before the current 405 406 CHAPTER 11. MARKOV CHAINS state. The probabilities pijare called transition probabilities. The process can remain in the state it is in, and this occurs with probability pii. An initial probability distribution, deflned on S, specifles the starting state. Usually this is done by specifying a particular state as the starting state. R. A. Howard1provides us with a picturesque description of a Markov chain as a frog jumping on a set of lily pads. The frog starts on one of the pads and thenjumps from lily pad to lily pad with the appropriate transition probabilities. Example 11.1 According to Kemeny, Snell, and Thompson, 2the Land of Oz is blessed by many things, but not by good weather. They never have two nice daysin a row. If they have a nice day, they are just as likely to have snow as rain thenext day. If they have snow or rain, they have an even chance of having the samethe next day. If there is change from snow or rain, only half of the time is this achange to a nice day. With this information we form a Markov chain as follows.We take as states the kinds of weather R, N, and S. From the above informationwe determine the transition probabilities. These are most conveniently representedin a square array as P=0 @RNS R1=21=41=4 N1=201=2 S1=41=41=21 A: 2 Transition Matrix The entries in the flrst row of the matrix Pin Example 11.1 represent the proba- bilities for the various kinds of weather following a rainy day. Similarly, the entriesin the second and third rows represent the probabilities for the various kinds ofweather following nice and snowy days, respectively. Such a square array is calledthematrix of transition probabilities ,o rt h e transition matrix . We consider the question of determining the probability that, given the chain is in stateitoday, it will be in state jtwo days from now. We denote this probability byp (2) ij. In Example 11.1, we see that if it is rainy today then the event that it is snowy two days from now is the disjoint union of the following three events: 1)it is rainy tomorrow and snowy two days from now, 2) it is nice tomorrow andsnowy two days from now, and 3) it is snowy tomorrow and snowy two days fromnow. The probability of the flrst of these events is the product of the conditionalprobability that it is rainy tomorrow, given that it is rainy today, and the conditionalprobability that it is snowy two days from now, given that it is rainy tomorrow.Using the transition matrix P, we can write this product as p 11p13. The other two 1R. A. Howard, Dynamic Probabilistic Systems, vol. 1 (New York: John Wiley and Sons, 1971). 2J. G. Kemeny, J. L. Snell, G. L. Thompson, Introduction to Finite Mathematics, 3rd ed. (Englewood Clifis, NJ: Prentice-Hall, 1974). 11.1. INTRODUCTION 407 events also have probabilities that can be written as products of entries of P. Thus, we have p(2) 13=p11p13+p12p23+p13p33: This equation should remind the reader of a dot product of two vectors; we are dotting the flrst row of Pwith the third column of P. This is just what is done in obtaining the 1 ;3-entry of the product of Pwith itself. In general, if a Markov chain hasrstates, then p(2) ij=rX k=1pikpkj: The following general theorem is easy to prove by using the above observation and induction. Theorem 11.1 LetPbe the transition matrix of a Markov chain. The ijth en- tryp(n) ijof the matrix Pngives the probability that the Markov chain, starting in statesi, will be in state sjafternsteps. Proof. The proof of this theorem is left as an exercise (Exercise 17). 2 Example 11.2 (Example 11.1 continued) Consider again the weather in the Land of Oz. We know that the powers of the transition matrix give us interesting in-formation about the process as it evolves. We shall be particularly interested inthe state of the chain after a large number of steps. The program Matrix Powers computes the powers of P. We have run the program Matrix Powers for the Land of Oz example to com- pute the successive powers of Pfrom 1 to 6. The results are shown in Table 11.1. We note that after six days our weather predictions are, to three-decimal-place ac-curacy, independent of today’s weather. The probabilities for the three types ofweather, R, N, and S, are .4, .2, and .4 no matter where the chain started. Thisis an example of a type of Markov chain called a regular Markov chain. For this type of chain, it is true that long-range predictions are independent of the startingstate. Not all chains are regular, but this is an important class of chains that weshall study in detail later. 2 We now consider the long-term behavior of a Markov chain when it starts in a state chosen by a probability distribution on the set of states, which we will call aprobability vector . A probability vector with rcomponents is a row vector whose entries are non-negative and sum to 1. If uis a probability vector which represents the initial state of a Markov chain, then we think of the ith component of uas representing the probability that the chain starts in state s i. With this interpretation of random starting states, it is easy to prove the fol- lowing theorem. 408 CHAPTER 11. MARKOV CHAINS P1=0 @Rain Nice Snow Rain:500:250:250 Nice:500:000:500 Snow:250:250:5001 A P2=0 @Rain Nice Snow Rain:438:188:375 Nice:375:250:375 Snow:375:188:4381 A P3=0 @Rain Nice Snow Rain:406:203:391 Nice:406:188:406 Snow:391:203:4061 A P4=0 @Rain Nice Snow Rain:402:199:398 Nice:398:203:398 Snow:398:199:4021 A P5=0 @Rain Nice Snow Rain:400:200:399 Nice:400:199:400 Snow:399:200:4001 A P6=0 @Rain Nice Snow Rain:400:200:400 Nice:400:200:400 Snow:400:200:4001 A Table 11.1: Powers of the Land of Oz transition matrix. 11.1. INTRODUCTION 409 Theorem 11.2 LetPbe the transition matrix of a Markov chain, and let ube the probability vector which represents the starting distribution. Then the probabilitythat the chain is in state s iafternsteps is the ith entry in the vector u(n)=uPn: Proof. The proof of this theorem is left as an exercise (Exercise 18). 2 We note that if we want to examine the behavior of the chain under the assump- tion that it starts in a certain state si, we simply choose uto be the probability vector with ith entry equal to 1 and all other entries equal to 0. Example 11.3 In the Land of Oz example (Example 11.1) let the initial probability vector uequal (1=3;1=3;1=3). Then we can calculate the distribution of the states after three days using Theorem 11.2 and our previous calculation of P3. We obtain u(3)=uP3=( 1=3;1=3;1=3)0 @:406:203:391 :406:188:406 :391:203:4061 A =(:401;:188;:401 ): 2 Examples The following examples of Markov chains will be used throughout the chapter for exercises. Example 11.4 The President of the United States tells person A his or her in- tention to run or not to run in the next election. Then A relays the news to B,who in turn relays the message to C, and so forth, always to some new person. Weassume that there is a probability athat a person will change the answer from yes to no when transmitting it to the next person and a probability bthat he or she will change it from no to yes. We choose as states the message, either yes or no.The transition matrix is then P=µyes no yes 1¡aa nob 1¡b¶ : The initial state represents the President’s choice. 2 Example 11.5 Each time a certain horse runs in a three-horse race, he has proba- bility 1/2 of winning, 1/4 of coming in second, and 1/4 of coming in third, indepen-dent of the outcome of any previous race. We have an independent trials process, 410 CHAPTER 11. MARKOV CHAINS but it can also be considered from the point of view of Markov chain theory. The transition matrix is P=0 @WP S W:5:25:25 P:5:25:25 S:5:25:251 A: 2 Example 11.6 In the Dark Ages, Harvard, Dartmouth, and Yale admitted only male students. Assume that, at that time, 80 percent of the sons of Harvard menwent to Harvard and the rest went to Yale, 40 percent of the sons of Yale men wentto Yale, and the rest split evenly between Harvard and Dartmouth; and of the sonsof Dartmouth men, 70 percent went to Dartmouth, 20 percent to Harvard, and10 percent to Yale. We form a Markov chain with transition matrix P=0 @HYD H:8:20 Y:3:4:3 D:2:1:71 A: 2 Example 11.7 Modify Example 11.6 by assuming that the son of a Harvard man always went to Harvard. The transition matrix is now P=0 @HYD H100 Y:3:4:3 D:2:1:71 A: 2 Example 11.8 (Ehrenfest Model) The following is a special case of a model, called the Ehrenfest model, 3that has been used to explain difiusion of gases. The general model will be discussed in detail in Section 11.5. We have two urns that, betweenthem, contain four balls. At each step, one of the four balls is chosen at randomand moved from the urn that it is in into the other urn. We choose, as states, thenumber of balls in the flrst urn. The transition matrix is then P=0 BBBB@01234 0 01000 11=403=40 0 20 1=201=20 30 0 3 =401=4 4 000101 CCCCA: 2 3P. and T. Ehrenfest, \ ˜Uber zwei bekannte Einw˜ ande gegen das Boltzmannsche H-Theorem," Physikalishce Zeitschrift, vol. 8 (1907), pp. 311-314. 11.1. INTRODUCTION 411 Example 11.9 (Gene Model) The simplest type of inheritance of traits in animals occurs when a trait is governed by a pair of genes, each of which may be of two types,say G and g. An individual may have a GG combination or Gg (which is geneticallythe same as gG) or gg. Very often the GG and Gg types are indistinguishable inappearance, and then we say that the G gene dominates the g gene. An individualis called dominant if he or she has GG genes, recessive if he or she has gg, and hybrid with a Gg mixture. In the mating of two animals, the ofispring inherits one gene of the pair from each parent, and the basic assumption of genetics is that these genes are selected atrandom, independently of each other. This assumption determines the probabilityof occurrence of each type of ofispring. The ofispring of two purely dominant parentsmust be dominant, of two recessive parents must be recessive, and of one dominantand one recessive parent must be hybrid. In the mating of a dominant and a hybrid animal, each ofispring must get a G gene from the former and has an equal chance of getting G or g from the latter.Hence there is an equal probability for getting a dominant or a hybrid ofispring.Again, in the mating of a recessive and a hybrid, there is an even chance for gettingeither a recessive or a hybrid. In the mating of two hybrids, the ofispring has anequal chance of getting G or g from each parent. Hence the probabilities are 1/4for GG, 1/2 for Gg, and 1/4 for gg. Consider a process of continued matings. We start with an individual of known genetic character and mate it with a hybrid. We assume that there is at least oneofispring. An ofispring is chosen at random and is mated with a hybrid and thisprocess repeated through a number of generations. The genetic type of the chosenofispring in successive generations can be represented by a Markov chain. The statesare dominant, hybrid, and recessive, and indicated by GG, Gg, and gg respectively. The transition probabilities are P=0 @GG Gg gg GG:5:50 Gg:25:5:25 gg 0 :5:51 A: 2 Example 11.10 Modify Example 11.9 as follows: Instead of mating the oldest ofispring with a hybrid, we mate it with a dominant individual. The transitionmatrix is P=0 @GG Gg gg GG 1 0 0 Gg:5:50 gg 0 1 01 A: 2 412 CHAPTER 11. MARKOV CHAINS Example 11.11 We start with two animals of opposite sex, mate them, select two of their ofispring of opposite sex, and mate those, and so forth. To simplify theexample, we will assume that the trait under consideration is independent of sex. Here a state is determined by a pair of animals. Hence, the states of our process will be:s 1= (GG;GG),s2= (GG;Gg),s3= (GG;gg),s4= (Gg;Gg),s5= (Gg;gg), ands6= (gg;gg). We illustrate the calculation of transition probabilities in terms of the state s2. When the process is in this state, one parent has GG genes, the other Gg. Hence,the probability of a dominant ofispring is 1/2. Then the probability of transitiontos 1(selection of two dominants) is 1/4, transition to s2is 1/2, and to s4is 1/4. The other states are treated the same way. The transition matrix of this chain is: P1=0 BBBBBBB@GG,GG GG,Gg GG,gg Gg,Gg Gg,gg gg,gg GG,GG 1 :000:000:000:000:000:000 GG,Gg :250:500:000:250:000:000 GG,gg :000:000:000 1:000:000:000 Gg,Gg :062:250:125:250:250:062 Gg,gg :000:000:000:250:500:250 gg,gg :000:000:000:000:000 1:0001 CCCCCCCA: 2 Example 11.12 (Stepping Stone Model) Our flnal example is another example that has been used in the study of genetics. It is called the stepping stone model. 4 In this model we have an n-by-narray of squares, and each square is initially any one ofkdifierent colors. For each step, a square is chosen at random. This square then chooses one of its eight neighbors at random and assumes the color of thatneighbor. To avoid boundary problems, we assume that if a square Sis on the left-hand boundary, say, but not at a corner, it is adjacent to the square Ton the right-hand boundary in the same row as S, andSis also adjacent to the squares just above and below T. A similar assumption is made about squares on the upper and lower boundaries. (These adjacencies are much easier to understand if one imaginesmaking the array into a cylinder by gluing the top and bottom edge together, andthen making the cylinder into a doughnut by gluing the two circular boundariestogether.) With these adjacencies, each square in the array is adjacent to exactlyeight other squares. A state in this Markov chain is a description of the color of each square. For this Markov chain the number of states is k n2, which for even a small array of squares is enormous. This is an example of a Markov chain that is easy to simulate butdi–cult to analyze in terms of its transition matrix. The program SteppingStone simulates this chain. We have started with a random initial conflguration of twocolors with n= 20 and show the result after the process has run for some time in Figure 11.2. 4S. Sawyer, \Results for The Stepping Stone Model for Migration in Population Genetics," Annals of Probability, vol. 4 (1979), pp. 699{728. 11.1. INTRODUCTION 413 Figure 11.1: Initial state of the stepping stone model. Figure 11.2: State of the stepping stone model after 10,000 steps. This is an example of an absorbing Markov chain. This type of chain will be studied in Section 11.2. One of the theorems proved in that section, applied tothe present example, implies that with probability 1, the stones will eventually allbe the same color. By watching the program run, you can see that territories areestablished and a battle develops to see which color survives. At any time theprobability that a particular color will win out is equal to the proportion of thearray of this color. You are asked to prove this in Exercise 11.2.32. 2 Exercises 1It is raining in the Land of Oz. Determine a tree and a tree measure for the next three days’ weather. Find w(1);w(2);andw(3)and compare with the results obtained from P;P2;andP3. 2In Example 11.4, let a= 0 andb=1=2. Find P;P2;andP3:What would Pnbe? What happens to Pnasntends to inflnity? Interpret this result. 3In Example 11.5, flnd P,P2;andP3:What is Pn? 414 CHAPTER 11. MARKOV CHAINS 4For Example 11.6, flnd the probability that the grandson of a man from Har- vard went to Harvard. 5In Example 11.7, flnd the probability that the grandson of a man from Harvard went to Harvard. 6In Example 11.9, assume that we start with a hybrid bred to a hybrid. Find w(1);w(2);andw(3):What would w(n)be? 7Find the matrices P2;P3;P4;andPnfor the Markov chain determined by the transition matrix P=µ10 01¶ . Do the same for the transition matrix P=µ01 10¶ . Interpret what happens in each of these processes. 8A certain calculating machine uses only the digits 0 and 1. It is supposed to transmit one of these digits through several stages. However, at every stage,there is a probability pthat the digit that enters this stage will be changed when it leaves and a probability q=1¡pthat it won’t. Form a Markov chain to represent the process of transmission by taking as states the digits 0 and 1.What is the matrix of transition probabilities? 9For the Markov chain in Exercise 8, draw a tree and assign a tree measure assuming that the process begins in state 0 and moves through two stagesof transmission. What is the probability that the machine, after two stages,produces the digit 0 (i.e., the correct digit)? What is the probability that themachine never changed the digit from 0? Now let p=:1. Using the program Matrix Powers , compute the 100th power of the transition matrix. Interpret the entries of this matrix. Repeat this with p=:2. Why do the 100th powers appear to be the same? 10Modify the program Matrix Powers so that it prints out the average A nof the powers Pn, forn=1t oN. Try your program on the Land of Oz example and compare AnandPn: 11Assume that a man’s profession can be classifled as professional, skilled la- borer, or unskilled laborer. Assume that, of the sons of professional men,80 percent are professional, 10 percent are skilled laborers, and 10 percent areunskilled laborers. In the case of sons of skilled laborers, 60 percent are skilledlaborers, 20 percent are professional, and 20 percent are unskilled. Finally, inthe case of unskilled laborers, 50 percent of the sons are unskilled laborers,and 25 percent each are in the other two categories. Assume that every manhas at least one son, and form a Markov chain by following the profession ofa randomly chosen son of a given family through several generations. Set upthe matrix of transition probabilities. Find the probability that a randomlychosen grandson of an unskilled laborer is a professional man. 12In Exercise 11, we assumed that every man has a son. Assume instead that the probability that a man has at least one son is .8. Form a Markov chain 11.2. ABSORBING MARKOV CHAINS 415 with four states. If a man has a son, the probability that this son is in a particular profession is the same as in Exercise 11. If there is no son, theprocess moves to state four which represents families whose male line has diedout. Find the matrix of transition probabilities and flnd the probability thata randomly chosen grandson of an unskilled laborer is a professional man. 13Write a program to compute u (n)given uandP. Use this program to compute u(10)for the Land of Oz example, with u=( 0;1;0), and with u=( 1=3;1=3;1=3). 14Using the program Matrix Powers , flnd P1through P6for Examples 11.9 and 11.10. See if you can predict the long-range probability of flnding theprocess in each of the states for these examples. 15Write a program to simulate the outcomes of a Markov chain after nsteps, given the initial starting state and the transition matrix Pas data (see Ex- ample 11.12). Keep this program for use in later problems. 16Modify the program of Exercise 15 so that it keeps track of the proportion of times in each state in nsteps. Run the modifled program for difierent starting states for Example 11.1 and Example 11.8. Does the initial state afiect theproportion of time spent in each of the states if nis large? 17Prove Theorem 11.1. 18Prove Theorem 11.2. 19Consider the following process. We have two coins, one of which is fair, and the other of which has heads on both sides. We give these two coins to our friend,who chooses one of them at random (each with probability 1/2). During therest of the process, she uses only the coin that she chose. She now proceedsto toss the coin many times, reporting the results. We consider this processto consist solely of what she reports to us. (a) Given that she reports a head on the nth toss, what is the probability that a head is thrown on the ( n+ 1)st toss? (b) Consider this process as having two states, heads and tails. By computing the other three transition probabilities analogous to the one in part (a),write down a \transition matrix" for this process. (c) Now assume that the process is in state \heads" on both the ( n¡1)st and thenth toss. Find the probability that a head comes up on the (n+ 1)st toss. (d) Is this process a Markov chain? 11.2 Absorbing Markov Chains The subject of Markov chains is best studied by considering special types of Markov chains. The flrst type that we shall study is called an absorbing Markov chain. 416 CHAPTER 11. MARKOV CHAINS 1 2 3 0 41 11/2 1/2 1/2 1/2 1/2 1/2 Figure 11.3: Drunkard’s walk. Deflnition 11.1 A statesiof a Markov chain is called absorbing if it is impossible to leave it (i.e., pii= 1). A Markov chain is absorbing if it has at least one absorbing state, and if from every state it is possible to go to an absorbing state (not necessarilyin one step). 2 Deflnition 11.2 In an absorbing Markov chain, a state which is not absorbing is called transient. 2 Drunkard’s Walk Example 11.13 A man walks along a four-block stretch of Park Avenue (see Fig- ure 11.3). If he is at corner 1, 2, or 3, then he walks to the left or right with equalprobability. He continues until he reaches corner 4, which is a bar, or corner 0,which is his home. If he reaches either home or the bar, he stays there. We form a Markov chain with states 0, 1, 2, 3, and 4. States 0 and 4 are absorbing states. The transition matrix is then P=0 BBBB@01234 0 10000 11=201=20 0 20 1=201=20 30 0 1 =201=2 4 000011 CCCCA: The states 1, 2, and 3 are transient states, and from any of these it is possible to reach the absorbing states 0 and 4. Hence the chain is an absorbing chain. Whena process reaches an absorbing state, we shall say that it is absorbed . 2 The most obvious question that can be asked about such a chain is: What is the probability that the process will eventually reach an absorbing state? Otherinteresting questions include: (a) What is the probability that the process will endup in a given absorbing state? (b) On the average, how long will it take for theprocess to be absorbed? (c) On the average, how many times will the process be ineach transient state? The answers to all these questions depend, in general, on thestate from which the process starts as well as the transition probabilities. 11.2. ABSORBING MARKOV CHAINS 417 Canonical Form Consider an arbitrary absorbing Markov chain. Renumber the states so that the transient states come flrst. If there are rabsorbing states and ttransient states, the transition matrix will have the following canonical form P=0 B@TR. ABS. TR. Q R ABS. 0 I1 CA HereIis anr-by-rindentity matrix, 0is anr-by-tzero matrix, Ris a nonzero t-by-rmatrix, and Qis ant-by-tmatrix. The flrst tstates are transient and the lastrstates are absorbing. In Section 11.1, we saw that the entry p(n) ijof the matrix Pnis the probability of being in the state sjafternsteps, when the chain is started in state si. A standard matrix algebra argument shows that Pnis of the form Pn=0 B@TR. ABS. TR. Qn⁄ ABS. 0 I1 CA where the asterisk ⁄stands for the t-by-rmatrix in the upper right-hand corner ofPn:(This submatrix can be written in terms of QandR, but the expression is complicated and is not needed at this time.) The form of Pnshows that the entries of Qngive the probabilities for being in each of the transient states after n steps for each possible transient starting state. For our flrst theorem we prove thatthe probability of being in the transient states after nsteps approaches zero. Thus every entry of Q nmust approach zero as napproaches inflnity (i.e, Qn!0). In the following, if uandvare two vectors we say that u•vif all components ofuare less than or equal to the corresponding components of v. Similarly, if AandBare matrices then A•Bif each entry of Ais less than or equal to the corresponding entry of B. Probability of Absorption Theorem 11.3 In an absorbing Markov chain, the probability that the process will be absorbed is 1 (i.e., Qn!0asn!1 ). Proof. From each nonabsorbing state sjit is possible to reach an absorbing state. Letmjbe the minimum number of steps required to reach an absorbing state, starting from sj. Letpjbe the probability that, starting from sj, the process will not reach an absorbing state in mjsteps. Then pj<1. Letmbe the largest of the mjand letpbe the largest of pj. The probability of not being absorbed in msteps 418 CHAPTER 11. MARKOV CHAINS is less than or equal to p,i n2nsteps less than or equal to p2, etc. Since p<1 these probabilities tend to 0. Since the probability of not being absorbed in nsteps is monotone decreasing, these probabilities also tend to 0, hence lim n!1Qn=0:2 The Fundamental Matrix Theorem 11.4 For an absorbing Markov chain the matrix I¡Qhas an inverse NandN=I+Q+Q2+¢¢¢. Theij-entrynijof the matrix Nis the expected number of times the chain is in state sj, given that it starts in state si. The initial state is counted if i=j. Proof. Let (I¡Q)x= 0; that is x=Qx:Then, iterating this we see that x=Qnx:Since Qn!0, we have Qnx!0,s ox=0.T h u s ( I¡Q)¡1=N exists. Note next that (I¡Q)(I+Q+Q2+¢¢¢+Qn)=I¡Qn+1: Thus multiplying both sides by Ngives I+Q+Q2+¢¢¢+Qn=N(I¡Qn+1): Lettingntend to inflnity we have N=I+Q+Q2+¢¢¢: Letsiandsjbe two transient states, and assume throughout the remainder of the proof that iandjare flxed. Let X(k)be a random variable which equals 1 if the chain is in state sjafterksteps, and equals 0 otherwise. For each k, this random variable depends upon both iandj; we choose not to explicitly show this dependence in the interest of clarity. We have P(X(k)=1 )=q(k) ij; and P(X(k)=0 )=1¡q(k) ij; whereq(k) ijis theijth entry of Qk. These equations hold for k= 0 since Q0=I. Therefore, since X(k)is a 0-1 random variable, E(X(k))=q(k) ij. The expected number of times the chain is in state sjin the flrst nsteps, given that it starts in state si, is clearly E‡ X(0)+X(1)+¢¢¢+X(n)· =q(0) ij+q(1) ij+¢¢¢+q(n) ij: Lettingntend to inflnity we have E‡ X(0)+X(1)+¢¢¢· =q(0) ij+q(1) ij+¢¢¢=nij: 2 11.2. ABSORBING MARKOV CHAINS 419 Deflnition 11.3 For an absorbing Markov chain P, the matrix N=(I¡Q)¡1is called the fundamental matrix forP. The entry nijofNgives the expected number of times that the process is in the transient state sjif it is started in the transient statesi. 2 Example 11.14 (Example 11.13 continued) In the Drunkard’s Walk example, the transition matrix in canonical form is P=0 BBB@12304 10 1=201=20 21=201=200 30 1=20 01=2 0 000 10 4 000 011 CCCA: From this we see that the matrix Qis Q=0 @01=20 1=201=2 01=201 A; and I¡Q=0 @1¡1=20 ¡1=21¡1=2 0¡1=211 A: Computing ( I¡Q)¡1,w efl n d N=(I¡Q)¡1=0 @123 13=211=2 2 12131=213=21 A: From the middle row of N, we see that if we start in state 2, then the expected number of times in states 1, 2, and 3 before being absorbed are 1, 2, and 1. 2 Time to Absorption We now consider the question: Given that the chain starts in state si, what is the expected number of steps before the chain is absorbed? The answer is given in thenext theorem. Theorem 11.5 Lett ibe the expected number of steps before the chain is absorbed, given that the chain starts in state si, and let tbe the column vector whose ith entry isti. Then t=Nc; where cis a column vector all of whose entries are 1. 420 CHAPTER 11. MARKOV CHAINS Proof. If we add all the entries in the ith row of N, we will have the expected number of times in any of the transient states for a given starting state si, that is, the expected time required before being absorbed. Thus, tiis the sum of the entries in the ith row of N. If we write this statement in matrix form, we obtain the theorem. 2 Absorption Probabilities Theorem 11.6 Letbijbe the probability that an absorbing chain will be absorbed in the absorbing state sjif it starts in the transient state si. Let Bbe the matrix with entries bij. Then Bis ant-by-rmatrix, and B=NR; where Nis the fundamental matrix and Ris as in the canonical form. Proof. We have Bij=X nX kq(n) ikrkj =X kX nq(n) ikrkj =X knikrkj =(NR)ij: This completes the proof. 2 Another proof of this is given in Exercise 34. Example 11.15 (Example 11.14 continued) In the Drunkard’s Walk example, we found that N=0 @123 13=211=2 2 12131=213=21 A: Hence, t=Nc =0 @3=211=2 121 1=213=21 A0 @1 111 A =0 @3 431 A: 11.2. ABSORBING MARKOV CHAINS 421 Thus, starting in states 1, 2, and 3, the expected times to absorption are 3, 4, and 3, respectively. From the canonical form, R=0 @04 11=20 20 030 1=21 A: Hence, B=NR =0 @3=211=2 121 1=213=21 A¢0 @1=20 0001=21 A =0 @04 13=41=4 21=21=2 31=43=41 A: Here the flrst row tells us that, starting from state 1, there is probability 3/4 of absorption in state 0 and 1/4 of absorption in state 4. 2 Computation The fact that we have been able to obtain these three descriptive quantities in matrix form makes it very easy to write a computer program that determines thesequantities for a given absorbing chain matrix. The program AbsorbingChain calculates the basic descriptive quantities of an absorbing Markov chain. We have run the program AbsorbingChain for the example of the drunkard’s walk (Example 11.13) with 5 blocks. The results are as follows: Q=0 BB@1234 1:00:50:00:00 2:50:00:50:00 3:00:50:00:50 4:00:00:50:001 CCA; R=0 BB@05 1:50:00 2:00:00 3:00:00 4:00:501 CCA; 422 CHAPTER 11. MARKOV CHAINS N=0 BB@1234 11:60 1:20:80:40 21:20 2:40 1:60:80 3:80 1:60 2:40 1:20 4:40:80 1:20 1:601 CCA; t=0 BB@14:00 26:00 36:00 44:001 CCA; B=0 BB@05 1:80:20 2:60:40 3:40:60 4:20:801 CCA: Note that the probability of reaching the bar before reaching home, starting atx,i sx=5 (i.e., proportional to the distance of home from the starting point). (See Exercise 24.) Exercises 1In Example 11.4, for what values of aandbdo we obtain an absorbing Markov chain? 2Show that Example 11.7 is an absorbing Markov chain. 3Which of the genetics examples (Examples 11.9, 11.10, and 11.11) are ab- sorbing? 4Find the fundamental matrix Nfor Example 11.10. 5For Example 11.11, verify that the following matrix is the inverse of I¡Q and hence is the fundamental matrix N. N=0 BB@8=31=64=32=3 4=34=38=34=3 4=31=38=34=3 2=31=64=38=31 CCA: FindNcandNR. Interpret the results. 6In the Land of Oz example (Example 11.1), change the transition matrix by making R an absorbing state. This gives P=0 @RNS R 100 N1=201=2 S1=41=41=21 A: 11.2. ABSORBING MARKOV CHAINS 423 Find the fundamental matrix N, and also NcandNR. Interpret the results. 7In Example 11.8, make states 0 and 4 into absorbing states. Find the fun- damental matrix N, and also NcandNR, for the resulting absorbing chain. Interpret the results. 8In Example 11.13 (Drunkard’s Walk) of this section, assume that the proba- bility of a step to the right is 2/3, and a step to the left is 1/3. Find N;Nc, andNR. Compare these with the results of Example 11.15. 9A process moves on the integers 1, 2, 3, 4, and 5. It starts at 1 and, on each successive step, moves to an integer greater than its present position, movingwith equal probability to each of the remaining larger integers. State flve isan absorbing state. Find the expected number of steps to reach state flve. 10Using the result of Exercise 9, make a conjecture for the form of the funda- mental matrix if the process moves as in that exercise, except that it nowmoves on the integers from 1 to n. Test your conjecture for several difierent values ofn. Can you conjecture an estimate for the expected number of steps to reach state n, for largen? (See Exercise 11 for a method of determining this expected number of steps.) *11 Letb kdenote the expected number of steps to reach nfromn¡k, in the process described in Exercise 9. (a) Deflneb0= 0. Show that for k>0, we have bk=1+1 k¡ bk¡1+bk¡2+¢¢¢+b0¢ : (b) Let f(x)=b0+b1x+b2x2+¢¢¢: Using the recursion in part (a), show that f(x) satisfles the difierential equation (1¡x)2y0¡(1¡x)y+1=0: (c) Show that the general solution of the difierential equation in part (b) is y=¡log(1¡x) 1¡x+c 1¡x; wherecis a constant. (d) Use part (c) to show that bk=1+1 2+1 3+¢¢¢+1 k: 12Three tanks flght a three-way duel. Tank A has probability 1/2 of destroying the tank at which it flres, tank B has probability 1/3 of destroying the tank atwhich it flres, and tank C has probability 1/6 of destroying the tank at which 424 CHAPTER 11. MARKOV CHAINS it flres. The tanks flre together and each tank flres at the strongest opponent not yet destroyed. Form a Markov chain by taking as states the subsets of theset of tanks. Find N;Nc, and NR, and interpret your results. Hint: Take as states ABC, AC, BC, A, B, C, and none, indicating the tanks that couldsurvive starting in state ABC. You can omit AB because this state cannot bereached from ABC. 13Smith is in jail and has 3 dollars; he can get out on bail if he has 8 dollars. A guard agrees to make a series of bets with him. If Smith bets Adollars, he winsAdollars with probability .4 and loses Adollars with probability .6. Find the probability that he wins 8 dollars before losing all of his money if (a) he bets 1 dollar each time (timid strategy). (b) he bets, each time, as much as possible but not more than necessary to bring his fortune up to 8 dollars (bold strategy). (c) Which strategy gives Smith the better chance of getting out of jail? 14With the situation in Exercise 13, consider the strategy such that for i<4, Smith bets min( i;4¡i), and fori‚4, he bets according to the bold strategy, whereiis his current fortune. Find the probability that he gets out of jail using this strategy. How does this probability compare with that obtained forthe bold strategy? 15Consider the game of tennis when deuce is reached. If a player wins the next point, he has advantage. On the following point, he either wins the game or the game returns to deuce. Assume that for any point, player A has probability .6 of winning the point and player B has probability .4 of winning the point. (a) Set this up as a Markov chain with state 1: A wins; 2: B wins; 3: advantage A; 4: deuce; 5: advantage B. (b) Find the absorption probabilities. (c) At deuce, flnd the expected duration of the game and the probability that B will win. Exercises 16 and 17 concern the inheritance of color-blindness, which is a sex- linked characteristic. There is a pair of genes, g and G, of which the formertends to produce color-blindness, the latter normal vision. The G gene isdominant. But a man has only one gene, and if this is g, he is color-blind. Aman inherits one of his mother’s two genes, while a woman inherits one genefrom each parent. Thus a man may be of type G or g, while a woman may betype GG or Gg or gg. We will study a process of inbreeding similar to thatof Example 11.11 by constructing a Markov chain. 16List the states of the chain. Hint: There are six. Compute the transition probabilities. Find the fundamental matrix N,Nc, and NR. 11.2. ABSORBING MARKOV CHAINS 425 17Show that in both Example 11.11 and the example just given, the probability of absorption in a state having genes of a particular type is equal to theproportion of genes of that type in the starting state. Show that this canbe explained by the fact that a game in which your fortune is the number ofgenes of a particular type in the state of the Markov chain is a fair game. 5 18Assume that a student going to a certain four-year medical school in northern New England has, each year, a probability qof °unking out, a probability r of having to repeat the year, and a probability pof moving on to the next year (in the fourth year, moving on means graduating). (a) Form a transition matrix for this process taking as states F, 1, 2, 3, 4, and G where F stands for °unking out and G for graduating, and theother states represent the year of study. (b) For the case q=:1,r=:2, andp=:7 flnd the time a beginning student can expect to be in the second year. How long should this student expectto be in medical school? (c) Find the probability that this beginning student will graduate. 19(E. Brown 6) Mary and John are playing the following game: They have a three-card deck marked with the numbers 1, 2, and 3 and a spinner with thenumbers 1, 2, and 3 on it. The game begins by dealing the cards out so thatthe dealer gets one card and the other person gets two. A move in the gameconsists of a spin of the spinner. The person having the card with the numberthat comes up on the spinner hands that card to the other person. The gameends when someone has all the cards. (a) Set up the transition matrix for this absorbing Markov chain, where the states correspond to the number of cards that Mary has. (b) Find the fundamental matrix. (c) On the average, how many moves will the game last? (d) If Mary deals, what is the probability that John will win the game? 20Assume that an experiment has mequally probable outcomes. Show that the expected number of independent trials before the flrst occurrence of kconsec- utive occurrences of one of these outcomes is ( m k¡1)=(m¡1).Hint: Form an absorbing Markov chain with states 1, 2, ...,kwith stateirepresenting the length of the current run. The expected time until a run of kis 1 more than the expected time until absorption for the chain started in state 1. It hasbeen found that, in the decimal expansion of pi, starting with the 24,658,601stdigit, there is a run of nine 7’s. What would your result say about the ex-pected number of digits necessary to flnd such a run if the digits are producedrandomly? 5H. Gonshor, \An Application of Random Walk to a Problem in Population Genetics," Amer- ican Math Monthly, vol. 94 (1987), pp. 668{671 6Private communication. 426 CHAPTER 11. MARKOV CHAINS 21(Roberts7) A city is divided into 3 areas 1, 2, and 3. It is estimated that amountsu1,u2, andu3of pollution are emitted each day from these three areas. A fraction qijof the pollution from region iends up the next day at regionj. A fraction qi=1¡P jqij>0 goes into the atmosphere and escapes. Letw(n) ibe the amount of pollution in area iafterndays. (a) Show that w(n)=u+uQ+¢¢¢+uQn¡1. (b) Show that w(n)!w, and show how to compute wfromu. (c) The government wants to limit pollution levels to a prescribed level by prescribing w:Show how to determine the levels of pollution uwhich would result in a prescribed limiting value w. 22In the Leontief economic model,8there arenindustries 1, 2, ...,n. The ith industry requires an amount 0 •qij•1 of goods (in dollar value) from companyjto produce 1 dollar’s worth of goods. The outside demand on the industries, in dollar value, is given by the vector d=(d1;d2;:::;dn). Let Q be the matrix with entries qij. (a) Show that if the industries produce total amounts given by the vector x=(x1;x2;:::;xn) then the amounts of goods of each type that the industries will need just to meet their internal demands is given by thevector xQ. (b) Show that in order to meet the outside demand dand the internal de- mands the industries must produce total amounts given by a vectorx=(x 1;x2;:::;xn) which satisfles the equation x=xQ+d. (c) Show that if Qis the Q-matrix for an absorbing Markov chain, then it is possible to meet any outside demand d. (d) Assume that the row sums of Qare less than or equal to 1. Give an economic interpretation of this condition. Form a Markov chain by takingthe states to be the industries and the transition probabilites to be the q ij. Add one absorbing state 0. Deflne qi0=1¡X jqij: Show that this chain will be absorbing if every company is either making a proflt or ultimately depends upon a proflt-making company. (e) Deflne xcto be the gross national product. Find an expression for the gross national product in terms of the demand vector dand the vector tgiving the expected time to absorption. 23A gambler plays a game in which on each play he wins one dollar with prob- abilitypand loses one dollar with probability q=1¡p. The Gambler’s Ruin 7F. Roberts, Discrete Mathematical Models (Englewood Clifis, NJ: Prentice Hall, 1976). 8W. W. Leontief, Input-Output Economics (Oxford: Oxford University Press, 1966). 11.2. ABSORBING MARKOV CHAINS 427 problem is the problem of flnding the probability wxof winning an amount T before losing everything, starting with state x. Show that this problem may be considered to be an absorbing Markov chain with states 0, 1, 2, ...,Twith 0 andTabsorbing states. Suppose that a gambler has probability p=:48 of winning on each play. Suppose, in addition, that the gambler starts with50 dollars and that T= 100 dollars. Simulate this game 100 times and see how often the gambler is ruined. This estimates w 50. 24Show thatwxof Exercise 23 satisfles the following conditions: (a)wx=pwx+1+qwx¡1forx=1 , 2 , ..., T¡1. (b)w0=0 . (c)wT=1 . Show that these conditions determine wx. Show that, if p=q=1=2, then wx=x T satisfles (a), (b), and (c) and hence is the solution. If p6=q, show that wx=(q=p)x¡1 (q=p)T¡1 satisfles these conditions and hence gives the probability of the gambler win- ning. 25Write a program to compute the probability wxof Exercise 24 for given values ofx,p, andT. Study the probability that the gambler will ruin the bank in a game that is only slightly unfavorable, say p=:49, if the bank has signiflcantly more money than the gambler. *26 We considered the two examples of the Drunkard’s Walk corresponding to the casesn= 4 andn= 5 blocks (see Example 11.13). Verify that in these two examples the expected time to absorption, starting at x, is equal to x(n¡x). See if you can prove that this is true in general. Hint: Show that if f(x)i s the expected time to absorption then f(0) =f(n)=0a n d f(x)=( 1=2)f(x¡1 )+( 1=2)f(x+1 )+1 for 0<x<n . Show that if f1(x) andf2(x) are two solutions, then their difierenceg(x) is a solution of the equation g(x)=( 1=2)g(x¡1 )+( 1=2)g(x+1 ): Also,g(0) =g(n) = 0. Show that it is not possible for g(x) to have a strict maximum or a strict minimum at the point i, where 1•i•n¡1. Use this to show that g(i) = 0 for all i. This shows that there is at most one solution. Then verify that the function f(x)=x(n¡x) is a solution. 428 CHAPTER 11. MARKOV CHAINS 27Consider an absorbing Markov chain with state space S. Letfbe a function deflned onSwith the property that f(i)=X j2Spijf(j); or in vector form f=Pf: Thenfis called a harmonic function forP. If you imagine a game in which your fortune is f(i) when you are in state i, then the harmonic condition means that the game is fairin the sense that your expected fortune after one step is the same as it was before the step. (a) Show that for fharmonic f=Pnf for alln. (b) Show, using (a), that for fharmonic f=P1f; where P1= lim n!1Pn=µ0B 0I¶ : (c) Using (b), prove that when you start in a transient state iyour expected flnal fortuneX kbikf(k) is equal to your starting fortune f(i). In other words, a fair game on a flnite state space remains fair to the end. (Fair games in general arecalled martingales. Fair games on inflnite state spaces need not remain fair with an unlimited number of plays allowed. For example, considerthe game of Heads or Tails (see Example 1.4). Let Peter start with1 penny and play until he has 2. Then Peter will be sure to end up1 penny ahead.) 28A coin is tossed repeatedly. We are interested in flnding the expected number of tosses until a particular pattern, sa y B = HTH, occurs for the flrst time. If, for example, the outcomes of the tosses are HHTTHTH we say that thepattern B has occurred for the flrst time after 7 tosses. Let T Bbe the time to obtain pattern B for the flrst time. Li9gives the following method for determining E(TB). We are in a casino and, before each toss of the coin, a gambler enters, pays 1 dollar to play, and bets that the patter n B = HTH will occur on the next 9S-Y. R. Li, \A Martingale Approach to the Study of Occurrence of Sequence Patterns in Repeated Experiments," Annals of Probability, vol. 8 (1980), pp. 1171{1176. 11.2. ABSORBING MARKOV CHAINS 429 three tosses. If H occurs, he wins 2 dollars and bets this amount that the next outcome will be T. If he wins, he wins 4 dollars and bets this amount thatH will come up next time. If he wins, he wins 8 dollars and the pattern hasoccurred. If at any time he loses, he leaves with no winnings. Let A and B be two patterns. Let AB be the amount the gamblers win who arrive while the pattern A occurs and bet that B will occur. For example, ifA = HT and B = HTH then A B=2+4=6 since the flrst gambler bet on H and won 2 dollars and then bet on T and won 4 dollars more. The secondgambler bet on H and lost. I fA=H Ha n dB= HTH, then A B = 2 since the flrst gambler bet on H and won but then bet on T and lost and the secondgambler bet on H and won. I fA=B=H T H then AB = B B=8+2=1 0 . Now for each gambler coming in, the casino takes in 1 dollar. Thus the casino takes inT Bdollars. How much does it pay out? The only gamblers who go ofi with any money are those who arrive during the time the pattern B occursand they win the amount BB. But since all the bets made are perfectly fairbets, it seems quite intuitive that the expected amount the casino takes inshould equal the expected amount that it pays out. That is, E(T B) = BB. Since we have seen that fo r B = HTH, BB = 10, the expected time to reach the pattern HTH for the flrst time is 10. If we had been trying to get thepattern B = HHH, then BB = 8+4+2=1 4 since all the last three gamblers are paid ofi in this case. Thus the expected time to get the pattern HHH is 14.To justify this argument, Li used a theorem from the theory of martingales(fair games). We can obtain these expectations by considering a Markov chain whose states are the possible initial segments of the sequence HTH; these states are HTH,HT, H, and;, where;is the empty set. Then, for this example, the transition matrix is 0 BB@HTH HT H ; HTH 1 0 0 0 HT:50 0 :5 H0 :5:50 ; 00:5:51 CCA; and if B = HTH, E(T B) is the expected time to absorption for this chain started in state;. Show, using the associated Markov chain, that the values E(TB)=1 0a n d E(TB) = 14 are correct for the expected time to reach the patterns HTH and HHH, respectively. 29We can use the gambling interpretation given in Exercise 28 to flnd the ex- pected number of tosses required to reach pattern B when we start with pat-tern A. To be a meaningful problem, we assume that pattern A does not havepattern B as a subpattern. Let E A(TB) be the expected time to reach pattern B starting with pattern A. We use our gambling scheme and assume that theflrst k coin tosses produced the pattern A. During this time, the gamblers 430 CHAPTER 11. MARKOV CHAINS made an amount AB. The total amount the gamblers will have made when the pattern B occurs is BB. Thus, the amount that the gamblers made afterthe pattern A has occurred is BB - AB. Again by the fair game argument,E A(TB) = BB-AB. For example, suppose that we start with patter nA=H Ta n da r e trying to get the patter n B = HTH. Then we saw in Exercise 28 that A B = 4 and BB =1 0s oEA(TB) = BB-AB= 6. Verify that this gambling interpretation leads to the correct answer for all starting states in the examples that you worked in Exercise 28. 30Here is an elegant method due to Guibas and Odlyzko10to obtain the expected time to reach a pattern, say HTH, for the flrst time. Let f(n) be the number of sequences of length nwhich do not have the pattern HTH. Let fp(n)b et h e number of sequences that have the pattern for the flrst time after ntosses. To each element of f(n), add the pattern HTH. Then divide the resulting sequences into three subsets: the set where HTH occurs for the flrst time attimen+ 1 (for this, the original sequence must have ended with HT); the set where HTH occurs for the flrst time at time n+ 2 (cannot happen for this pattern); and the set where the sequence HTH occurs for the flrst time at timen+ 3 (the original sequence ended with anything except HT). Doing this, we have f(n)=f p(n+1 )+fp(n+3 ): Thus, f(n) 2n=2fp(n+1 ) 2n+1+23fp(n+3 ) 2n+3: IfTis the time that the pattern occurs for the flrst time, this equality states that P(T>n )=2P(T=n+1 )+8P(T=n+3 ): Show that if you sum this equality over all nyou obtain 1X n=0P(T>n )=2+8=1 0 : Show that for any integer-valued random variable E(T)=1X n=0P(T>n ); and conclude that E(T) = 10. Note that this method of proof makes very clear thatE(T) is, in general, equal to the expected amount the casino pays out and avoids the martingale system theorem used by Li. 10L. J. Guibas and A. M. Odlyzko, \String Overlaps, Pattern Matching, and Non-transitive Games," Journal of Combinatorial Theory, Series A, vol. 30 (1981), pp. 183{208. 11.2. ABSORBING MARKOV CHAINS 431 31In Example 11.11, deflne f(i) to be the proportion of G genes in state i. Show thatfis a harmonic function (see Exercise 27). Why does this show that the probability of being absorbed in state (GG ;GG) is equal to the proportion of G genes in the starting state? (See Exercise 17.) 32Show that the stepping stone model (Example 11.12) is an absorbing Markov chain. Assume that you are playing a game with red and green squares, inwhich your fortune at any time is equal to the proportion of red squares atthat time. Give an argument to show that this is a fair game in the sense thatyour expected winning after each step is just what it was before this step. Hint: Show that for every possible outcome in which your fortune will decrease byone there is another outcome of exactly the same probability where it willincrease by one. Use this fact and the results of Exercise 27 to show that the probability that a particular color wins out is equal to the proportion of squares that are initiallyof this color. 33Consider a random walker who moves on the integers 0, 1, ...,N, moving one step to the right with probability pand one step to the left with probability q=1¡p. If the walker ever reaches 0 or Nhe stays there. (This is the Gambler’s Ruin problem of Exercise 23.) If p=qshow that the function f(i)=i is a harmonic function (see Exercise 27), and if p6=qthen f(i)=µq p¶i is a harmonic function. Use this and the result of Exercise 27 to show that the probability biNof being absorbed in state Nstarting in state iis biN=(i N; ifp=q; (q p)i¡1 (q p)N¡1;ifp6=q: For an alternative derivation of these results see Exercise 24. 34Complete the following alternate proof of Theorem 11.6. Let sibe a tran- sient state and sjbe an absorbing state. If we compute bijin terms of the possibilities on the outcome of the flrst step, then we have the equation bij=pij+X kpikbkj; where the summation is carried out over all transient states sk. Write this in matrix form, and derive from this equation the statement B=NR: 432 CHAPTER 11. MARKOV CHAINS 35In Monte Carlo roulette (see Example 6.6), under option (c), there are six states (S,W,L,E,P1, andP2). The reader is referred to Figure 6.2, which contains a tree for this option. Form a Markov chain for this option, and usethe program AbsorbingChain to flnd the probabilities that you win, lose, or break even fo r a 1 franc bet on red. Using these probabilities, flnd the expected winnings for this bet. For a more general discussion of Markov chains appliedto roulette, see the article of H. Sagan referred to in Example 6.13. 36We consider next a game called Penney-ante by its inventor W. Penney. 11 There are two players; the flrst player picks a pattern A of H’s and T’s, and then the second player, knowing the choice of the flrst player, picks a difierentpattern B. We assume that neither pattern is a subpattern of the other pattern.A coin is tossed a sequence of times, and the player whose pattern comes upflrst is the winner. To analyze the game, we need to flnd the probability p A that pattern A will occur before pattern B and the probability pB=1¡pA that pattern B occurs before pattern A. To determine these probabilities we use the results of Exercises 28 and 29. Here you were asked to show that, theexpected time to reach a pattern B for the flrst time is, E(T B)=BB ; and, starting with pattern A, the expected time to reach pattern B is EA(TB)=BB¡AB : (a) Show that the odds that the flrst player will win are given by John Conway’s formula12: pA 1¡pA=pA pB=BB¡BA AA¡AB: Hint: Explain why E(TB)=E(TAorB)+pAEA(TB) and thus BB=E(TAorB)+pA(BB¡AB): Interchange A and B to flnd a similar equation involving the pB. Finally, note that pA+pB=1: Use these equations to solve for pAandpB. (b) Assume that both players choose a pattern of the same length k. Show that, ifk= 2, this is a fair game, but, if k= 3, the second player has an advantage no matter what choice the flrst player makes. (It has beenshown that, for k‚3, if the flrst player chooses a 1,a2, ...,ak, then the optimal strategy for the second player is of the form b,a1, ...,ak¡1 wherebis the better of the two choices H or T.13) 11W. Penney, \Problem: Penney-Ante," Journal of Recreational Math, vol. 2 (1969), p. 241. 12M. Gardner, \Mathematical Games," Scientiflc American, vol. 10 (1974), pp. 120{125. 13Guibas and Odlyzko, op. cit. 11.3. ERGODIC MARKOV CHAINS 433 11.3 Ergodic Markov Chains A second important kind of Markov chain we shall study in detail is an ergodic Markov chain, deflned as follows. Deflnition 11.4 A Markov chain is called an ergodic chain if it is possible to go from every state to every state (not necessarily in one move). 2 In many books, ergodic Markov chains are called irreducible . Deflnition 11.5 A Markov chain is called a regular chain if some power of the transition matrix has only positive elements. 2 In other words, for some n, it is possible to go from any state to any state in exactlynsteps. It is clear from this deflnition that every regular chain is ergodic. On the other hand, an ergodic chain is not necessarily regular, as the followingexamples show. Example 11.16 Let the transition matrix of a Markov chain be deflned by P=µ12 10 1 21 0¶ : Then is clear that it is possible to move from any state to any state, so the chain is ergodic. However, if nis odd, then it is not possible to move from state 0 to state 0i nnsteps, and if nis even, then it is not possible to move from state 0 to state 1 innsteps, so the chain is not regular. 2 A more interesting example of an ergodic, non-regular Markov chain is provided by the Ehrenfest urn model. Example 11.17 Recall the Ehrenfest urn model (Example 11.8). The transition matrix for this example is P=0 BBBB@01234 0 01000 11=403=40 0 20 1=201=20 30 0 3 =401=4 4 000101 CCCCA: In this example, if we start in state 0 we will, after any even number of steps, be in either state 0, 2 or 4, and after any odd number of steps, be in states 1 or 3. Thusthis chain is ergodic but not regular. 2 434 CHAPTER 11. MARKOV CHAINS Regular Markov Chains Any transition matrix that has no zeros determines a regular Markov chain. How- ever, it is possible for a regular Markov chain to have a transition matrix that haszeros. The transition matrix of the Land of Oz example of Section 11.1 has p NN=0 but the second power P2has no zeros, so this is a regular Markov chain. An example of a nonregular Markov chain is an absorbing chain. For example, let P=µ10 1=21=2¶ be the transition matrix of a Markov chain. Then all powers of Pwill have a 0 in the upper right-hand corner. We shall now discuss two important theorems relating to regular chains. Theorem 11.7 LetPbe the transition matrix for a regular chain. Then, as n! 1, the powers Pnapproach a limiting matrix Wwith all rows the same vector w. The vector wis a strictly positive probability vector (i.e., the components are all positive and they sum to one). 2 In the next section we give two proofs of this fundamental theorem. We give here the basic idea of the flrst proof. We want to show that the powers Pnof a regular transition matrix tend to a matrix with all rows the same. This is the same as showing that Pnconverges to a matrix with constant columns. Now the jth column of PnisPnywhere yis a column vector with 1 in the jth entry and 0 in the other entries. Thus we need only prove that for any column vector y;Pnyapproaches a constant vector as ntend to inflnity. Since each row of Pis a probability vector, Pyreplaces yby averages of its components. Here is an example: 0 @1=21=41=4 1=31=31=3 1=31=201 A0 @1 231 A=0 @1=2¢1+1=4¢2+1=4¢3 1=3¢1+1=3¢2+1=3¢3 1=3¢1+1=2¢2+0¢31 A=0 @7=4 2 3=21 A: The result of the averaging process is to make the components of Pymore similar than those of y. In particular, the maximum component decreases (from 3 to 2) and the minimum component increases (from 1 to 3/2). Our proof will show thatas we do more and more of this averaging to get P ny, the difierence between the maximum and minimum component will tend to 0 as n!1 . This means Pny tends to a constant vector. The ijth entry of Pn,p(n) ij, is the probability that the process will be in state sjafternsteps if it starts in state si. If we denote the common row of Wbyw, then Theorem 11.7 states that the probability of being insjin the long run is approximately wj, thejth entry of w, and is independent of the starting state. 11.3. ERGODIC MARKOV CHAINS 435 Example 11.18 Recall that for the Land of Oz example of Section 11.1, the sixth power of the transition matrix Pis, to three decimal places, P6=0 @RNS R:4:2:4 N:4:2:4 S:4:2:41 A: Thus, to this degree of accuracy, the probability of rain six days after a rainy day is the same as the probability of rain six days after a nice day, or six days aftera snowy day. Theorem 11.7 predicts that, for large n, the rows of Papproach a common vector. It is interesting that this occurs so soon in our example. 2 Theorem 11.8 LetPbe a regular transition matrix, let W= lim n!1Pn; letwbe the common row of W, and let cbe the column vector all of whose components are 1. Then (a) wP =w, and any row vector vsuch that vP=vis a constant multiple of w. (b) Pc =c, and any column vector xsuch that Px=xis a multiple of c. Proof. To prove part (a), we note that from Theorem 11.7, Pn!W: Thus, Pn+1=Pn¢P!WP: ButPn+1!W, and so W=WP, and w=wP. Letvbe any vector with vP=v. Then v=vPn, and passing to the limit, v=vW. Letrbe the sum of the components of v. Then it is easily checked that vW=rw. So,v=rw. To prove part (b), assume that x=Px. Then x=Pnx, and again passing to the limit, x=Wx. Since all rows of Ware the same, the components of Wx are all equal, so xis a multiple of c. 2 Note that an immediate consequence of Theorem 11.8 is the fact that there is only one probability vector vsuch that vP=v. Fixed Vectors Deflnition 11.6 A row vector wwith the property wP=wis called a flxed row vector forP. Similarly, a column vector xsuch that Px=xis called a flxed column vector forP. 2 436 CHAPTER 11. MARKOV CHAINS Thus, the common row of Wis the unique vector wwhich is both a flxed row vector for Pand a probability vector. Theorem 11.8 shows that any flxed row vector forPis a multiple of wand any flxed column vector for Pis a constant vector. One can also state Deflnition 11.6 in terms of eigenvalues and eigenvectors. A flxed row vector is a left eigenvector of the matrix Pcorresponding to the eigenvalue 1. A similar statement can be made about flxed column vectors. We will now give several difierent methods for calculating the flxed row vector wfor a regular Markov chain. Example 11.19 By Theorem 11.7 we can flnd the limiting vector wfor the Land of Oz from the fact that w1+w2+w3=1 and (w1w2w3)0 @1=21=41=4 1=201=2 1=41=41=21 A=(w1w2w3): These relations lead to the following four equations in three unknowns: w1+w2+w3=1; (1=2)w1+( 1=2)w2+( 1=4)w3=w1; (1=4)w1+( 1=4)w3=w2; (1=4)w1+( 1=2)w2+( 1=2)w3=w3: Our theorem guarantees that these equations have a unique solution. If the equations are solved, we obtain the solution w=(:4:2:4); in agreement with that predicted from P6, given in Example 11.2. 2 To calculate the flxed vector, we can assume that the value at a particular state, say state one, is 1, and then use all but one of the linear equations from wP=w. This set of equations will have a unique solution and we can obtain wfrom this solution by dividing each of its entries by their sum to give the probability vector w. We will now illustrate this idea for the above example. Example 11.20 (Example 11.19 continued) We set w1= 1, and then solve the flrst and second linear equations from wP=w. We have (1=2 )+( 1=2)w2+( 1=4)w3=1; (1=4 )+( 1=4)w3=w2: If we solve these, we obtain (w1w2w3)=(1 1=21 ): 11.3. ERGODIC MARKOV CHAINS 437 Now we divide this vector by the sum of the components, to obtain the flnal answer: w=(:4:2:4): This method can be easily programmed to run on a computer. 2 As mentioned above, we can also think of the flxed row vector was a left eigenvector of the transition matrix P. Thus, if we write Ito denote the identity matrix, then wsatisfles the matrix equation wP=wI; or equivalently, w(P¡I)=0: Thus, wis in the left nullspace of the matrix P¡I. Furthermore, Theorem 11.8 states that this left nullspace has dimension 1. Certain computer programminglanguages can flnd nullspaces of matrices. In such languages, one can flnd the flxedrow probability vector for a matrix Pby computing the left nullspace and then normalizing a vector in the nullspace so the sum of its components is 1. The program FixedVector uses one of the above methods (depending upon the language in which it is written) to calculate the flxed row probability vector forregular Markov chains. So far we have always assumed that we started in a speciflc state. The following theorem generalizes Theorem 11.7 to the case where the starting state is itselfdetermined by a probability vector. Theorem 11.9 LetPbe the transition matrix for a regular chain and van arbi- trary probability vector. Then lim n!1vPn=w; where wis the unique flxed probability vector for P. Proof. By Theorem 11.7, lim n!1Pn=W: Hence, lim n!1vPn=vW: But the entries in vsum to 1, and each row of Wequals w. From these statements, it is easy to check that vW=w: 2 If we start a Markov chain with initial probabilities given by v, then the proba- bility vector vPngives the probabilities of being in the various states after nsteps. Theorem 11.9 then establishes the fact that, even in this more general class ofprocesses, the probability of being in s japproaches wj. 438 CHAPTER 11. MARKOV CHAINS Equilibrium We also obtain a new interpretation for w. Suppose that our starting vector picks statesias a starting state with probability wi, for alli. Then the probability of being in the various states after nsteps is given by wPn=w, and is the same on all steps. This method of starting provides us with a process that is called \stationary."The fact that wis the only probability vector for which wP=wshows that we must have a starting probability vector of exactly the kind described to obtain astationary process. Many interesting results concerning regular Markov chains depend only on the fact that the chain has a unique flxed probability vector which is positive. Thisproperty holds for all ergodic Markov chains. Theorem 11.10 For an ergodic Markov chain, there is a unique probability vec- torwsuch that wP=wandwis strictly positive. Any row vector such that vP=vis a multiple of w. Any column vector xsuch that Px=xis a constant vector. Proof. This theorem states that Theorem 11.8 is true for ergodic chains. The result follows easily from the fact that, if Pis an ergodic transition matrix, then „P=( 1=2)I+( 1=2)Pis a regular transition matrix with the same flxed vectors (see Exercises 25{28). 2 For ergodic chains, the flxed probability vector has a slightly difierent inter- pretation. The following two theorems, which we will not prove here, furnish aninterpretation for this flxed vector. Theorem 11.11 LetPbe the transition matrix for an ergodic chain. Let A nbe the matrix deflned by An=I+P+P2+¢¢¢+Pn n+1: Then An!W, where Wis a matrix all of whose rows are equal to the unique flxed probability vector wforP. 2 IfPis the transition matrix of an ergodic chain, then Theorem 11.8 states that there is only one flxed row probability vector for P. Thus, we can use the same techniques that were used for regular chains to solve for this flxed vector. Inparticular, the program FixedVector works for ergodic chains. To interpret Theorem 11.11, let us assume that we have an ergodic chain that starts in state s i. LetX(m)= 1 if themth step is to state sjand 0 otherwise. Then the average number of times in state sjin the flrst nsteps is given by H(n)=X(0)+X(1)+X(2)+¢¢¢+X(n) n+1: ButX(m)takes on the value 1 with probability p(m) ijand 0 otherwise. Thus E(X(m))=p(m) ij, and theijth entry of Angives the expected value of H(n), that 11.3. ERGODIC MARKOV CHAINS 439 is, the expected proportion of times in state sjin the flrstnsteps if the chain starts in statesi. If we call being in state sjsuccess and any other state failure, we could ask if a theorem analogous to the law of large numbers for independent trials holds. Theanswer is yes and is given by the following theorem. Theorem 11.12 (Law of Large Numbers for Ergodic Markov Chains) Let H (n) jbe the proportion of times in nsteps that an ergodic chain is in state sj. Then for any†>0, P‡ jH(n) j¡wjj>†· !0; independent of the starting state si. 2 We have observed that every regular Markov chain is also an ergodic chain. Hence, Theorems 11.11 and 11.12 apply also for regular chains. For example, thisgives us a new interpretation for the flxed vector w=(:4;:2;:4) in the Land of Oz example. Theorem 11.11 predicts that, in the long run, it will rain 40 percent ofthe time in the Land of Oz, be nice 20 percent of the time, and snow 40 percent ofthe time. Simulation We illustrate Theorem 11.12 by writing a program to simulate the behavior of aMarkov chain. SimulateChain is such a program. Example 11.21 In the Land of Oz, there are 525 days in a year. We have simulated the weather for one year in the Land of Oz, using the program SimulateChain . The results are shown in Table 11.2. SSRNRNSSSSSSNRSNSSRNSRNSSSNSRRRNSSSNRRSSSSNRSSNSRRRRRRNSSS SSRRRSNSNRRRRSRSRNSNSRRNRRNRSSNSRNRNSSRRSRNSSSNRSRRSSNRSNRRNSSSSNSSNSRSRRNSSNSSRNSSRRNRRRSRNRRRNSSSNRNSRNSNRNRSSSRSSNRSSSNSSSSSSNSSSNSNSRRNRNRRRRSRRRSSSSNRRSSSSRSRRRNRRRSSSSRRNRRRSRSSRRRRSSRNRRRRRRNSSRNRSSSNRNSNRRRRNRRRNRSNRRNSRRSNRRRRSSSRNRRRNSNSSSSSRRRRSRNRSSRRRRSSSRRRNRNRRRSRSRNSNSSRRRRRNSNRNSNRRNRRRRRRSSSNRSSRSNRSSSNSNRNSNSSSNRRSRRRNRRRRNRNRSSSNSRSNRNRRSNRRNSRSSSRNSRRSSNSRRRNRRSNRRNSSSSSNRNSSSSSSSNRNSRRRNSSRRRNSSSNRRSRNSSRRNRRNRSNRRRRRRRRRNSNRRRRRNSRRSSSSNSNS State Times Fraction R 217 .413 N 109 .208S 199 .379 Table 11.2: Weather in the Land of Oz. 440 CHAPTER 11. MARKOV CHAINS We note that the simulation gives a proportion of times in each of the states not too difierent from the long run predictions of .4, .2, and .4 assured by Theorem 11.7.To get better results we have to simulate our chain for a longer time. We do thisfor 10,000 days without printing out each day’s weather. The results are shown inTable 11.3. We see that the results are now quite close to the theoretical values of.4, .2, and .4. State Times Fraction R 4010 .401 N 1902 .19S 4088 .409 Table 11.3: Comparison of observed and predicted frequencies for the Land of Oz. 2 Examples of Ergodic Chains The computation of the flxed vector wmay be di–cult if the transition matrix is very large. It is sometimes useful to guess the flxed vector on purely intuitivegrounds. Here is a simple example to illustrate this kind of situation. Example 11.22 A white rat is put into the maze of Figure 11.4. There are nine compartments with connections between the compartments as indicated. The ratmoves through the compartments at random. That is, if there are kways to leave a compartment, it chooses each of these with equal probability. We can representthe travels of the rat by a Markov chain process with transition matrix given by P=0 BBBBBBBBBBBBB@123456789 10 1=2 000 1 =2 000 21=301=301=3 0000 30 1=201=2 00000 40 0 1 =301=3 000 1 =3 50 1=401=401=401=40 61=3 000 1 =301=30 0 7 00000 1 =201=20 8 0000 1 =301=301=3 9 000 1 =2 000 1 =201 CCCCCCCCCCCCCA: That this chain is not regular can be seen as follows: From an odd-numbered state the process can go only to an even-numbered state, and from an even-numberedstate it can go only to an odd number. Hence, starting in state ithe process will be alternately in even-numbered and odd-numbered states. Therefore, odd powersofPwill have 0’s for the odd-numbered entries in row 1. On the other hand, a glance at the maze shows that it is possible to go from every state to every otherstate, so that the chain is ergodic. 11.3. ERGODIC MARKOV CHAINS 441 1 2 3 4 5 6 7 89 Figure 11.4: The maze problem. To flnd the flxed probability vector for this matrix, we would have to solve ten equations in nine unknowns. However, it would seem reasonable that the timesspent in each compartment should, in the long run, be proportional to the numberof entries to each compartment. Thus, we try the vector whose jth component is the number of entries to the jth compartment: x= ( 232343232 ) : It is easy to check that this vector is indeed a flxed vector so that the unique probability vector is this vector normalized to have sum 1: w=( 1 121 81 121 81 61 81 121 81 12): 2 Example 11.23 (Example 11.8 continued) We recall the Ehrenfest urn model of Example 11.8. The transition matrix for this chain is as follows: P=0 BBBB@01234 0:000 1:000:000:000:000 1:250:000:750:000:000 2:000:500:000:500:000 3:000:000:750:000:250 4:000:000:000 1:000:0001 CCCCA: If we run the program FixedVector for this chain, we obtain the vector w=¡01234 :0625:2500:3750:2500:0625¢ : By Theorem 11.12, we can interpret these values for w ias the proportion of times the process is in each of the states in the long run. For example, the proportion of 442 CHAPTER 11. MARKOV CHAINS times in state 0 is .0625 and the proportion of times in state 1 is .375. The astute reader will note that these numbers are the binomial distribution 1/16, 4/16, 6/16,4/16, 1/16. We could have guessed this answer as follows: If we consider a particularball, it simply moves randomly back and forth between the two urns. This suggeststhat the equilibrium state should be just as if we randomly distributed the fourballs in the two urns. If we did this, the probability that there would be exactlyjballs in one urn would be given by the binomial distribution b(n;p;j ) withn=4 andp=1=2. 2 Exercises 1Which of the following matrices are transition matrices for regular Markov chains? (a)P=µ:5:5 :5:5¶ . (b)P=µ:5:5 10¶ . (c)P=0 @1=302=3 01001=54=51 A. (d)P=µ01 10¶ . (e)P=0 @1=21=20 01=21=2 1=31=31=31 A. 2Consider the Markov chain with transition matrix P=0 @1=21=31=6 3=401=4 0101 A: (a) Show that this is a regular Markov chain. (b) The process is started in state 1; flnd the probability that it is in state 3 after two steps. (c) Find the limiting probability vector w. 3Consider the Markov chain with general 2 £2 transition matrix P=µ1¡aa b 1¡b¶ : (a) Under what conditions is Pabsorbing? (b) Under what conditions is Pergodic but not regular? (c) Under what conditions is Pregular? 11.3. ERGODIC MARKOV CHAINS 443 4Find the flxed probability vector wfor the matrices in Exercise 3 that are ergodic. 5Find the flxed probability vector wfor each of the following regular matrices. (a)P=µ:75:25 :5:5¶ . (b)P=µ:9:1 :1:9¶ . (c)P=0 @3=41=40 02=31=3 1=41=41=21 A. 6Consider the Markov chain with transition matrix in Exercise 3, with a=b= 1. Show that this chain is ergodic but not regular. Find the flxed probabilityvector and interpret it. Show that P ndoes not tend to a limit, but that An=I+P+P2+¢¢¢+Pn n+1 does. 7Consider the Markov chain with transition matrix of Exercise 3, with a=0 andb=1=2. Compute directly the unique flxed probability vector, and use your result to prove that the chain is not ergodic. 8Show that the matrix P=0 @100 1=41=21=4 0011 A has more than one flxed probability vector. Find the matrix that Pnap- proaches as n!1 , and verify that it is not a matrix all of whose rows are the same. 9Prove that, if a 3-by-3 transition matrix has the property that its column sums are 1, then (1 =3;1=3;1=3) is a flxed probability vector. State a similar result forn-by-ntransition matrices. Interpret these results for ergodic chains. 10Is the Markov chain in Example 11.10 ergodic? 11Is the Markov chain in Example 11.11 ergodic? 12Consider Example 11.13 (Drunkard’s Walk). Assume that if the walker reaches state 0, he turns around and returns to state 1 on the next step and, simi-larly, if he reaches 4 he returns on the next step to state 3. Is this new chainergodic? Is it regular? 13For Example 11.4 when Pis ergodic, what is the proportion of people who are told that the President will run? Interpret the fact that this proportionis independent of the starting state. 444 CHAPTER 11. MARKOV CHAINS 14Consider an independent trials process to be a Markov chain whose states are the possible outcomes of the individual trials. What is its flxed probabilityvector? Is the chain always regular? Illustrate this for Example 11.5. 15Show that Example 11.8 is an ergodic chain, but not a regular chain. Show that its flxed probability vector wis a binomial distribution. 16Show that Example 11.9 is regular and flnd the limiting vector. 17Toss a fair die repeatedly. Let S ndenote the total of the outcomes through thenth toss. Show that there is a limiting value for the proportion of the flrst nvalues ofSnthat are divisible by 7, and compute the value for this limit. Hint: The desired limit is an equilibrium probability vector for an appropriate seven state Markov chain. 18LetPbe the transition matrix of a regular Markov chain. Assume that there arerstates and let N(r) be the smallest integer nsuch that Pis regular if and only if PN(r)has no zero entries. Find a flnite upper bound for N(r). See if you can determine N(3) exactly. *19 Deflnef(r) to be the smallest integer nsuch that for all regular Markov chains withrstates, thenth power of the transition matrix has all entries positive. It has been shown,14thatf(r)=r2¡2r+2 . (a) Deflne the transition matrix of an r-state Markov chain as follows: For statessi, withi=1 ,2 ,..., r¡2,P(i;i+ 1 )=1 , P(r¡1;r)=P(r¡1;1) = 1=2, and P(r;1) = 1. Show that this is a regular Markov chain. (b) Forr= 3, verify that the flfth power is the flrst power that has no zeros. (c) Show that, for general r, the smallest nsuch that Pnhas all entries positive isn=f(r). 20A discrete time queueing system of capacity nconsists of the person being served and those waiting to be served. The queue length xis observed each second. If 0 <x<n , then with probability p, the queue size is increased by one by an arrival and, inependently, with probability r, it is decreased by one because the person being served flnishes service. If x= 0, only an arrival (with probability p) is possible. If x=n, an arrival will depart without waiting for service, and so only the departure (with probability r) of the person being served is possible. Form a Markov chain with states given by the number ofcustomers in the queue. Modify the program FixedVector so that you can inputn,p, andr, and the program will construct the transition matrix and compute the flxed vector. The quantity s=p=ris called the tra–c intensity. Describe the difierences in the flxed vectors according as s< 1,s=1 ,o r s>1. 14E. Seneta, Non-Negative Matrices: An Introduction to Theory and Applications, Wiley, New York, 1973, pp. 52-54. 11.3. ERGODIC MARKOV CHAINS 445 21Write a computer program to simulate the queue in Exercise 20. Have your program keep track of the proportion of the time that the queue length is jfor j=0 , 1 , ..., nand the average queue length. Show that the behavior of the queue length is very difierent depending upon whether the tra–c intensity s has the property s<1,s=1 ,o rs>1. 22In the queueing problem of Exercise 20, let Sbe the total service time required by a customer and Tthe time between arrivals of the customers. (a) Show that P(S=j)=( 1¡r)j¡1randP(T=j)=( 1¡p)j¡1p, for j>0. (b) Show that E(S)=1=randE(T)=1=p. (c) Interpret the conditions s<1,s= 1 ands>1 in terms of these expected values. 23In Exercise 20 the service time Shas a geometric distribution with E(S)= 1=r. Assume that the service time is, instead, a constant time of tseconds. Modify your computer program of Exercise 21 so that it simulates a constanttime service distribution. Compare the average queue length for the twotypes of distributions when they have the same expected service time (i.e.,taket=1=r). Which distribution leads to the longer queues on the average? 24A certain experiment is believed to be described by a two-state Markov chain with the transition matrix P, where P=µ:5:5 p1¡p¶ and the parameter pis not known. When the experiment is performed many times, the chain ends in state one approximately 20 percent of the time and instate two approximately 80 percent of the time. Compute a sensible estimatefor the unknown parameter pand explain how you found it. 25Prove that, in an r-state ergodic chain, it is possible to go from any state to any other state in at most r¡1 steps. 26LetPbe the transition matrix of an r-state ergodic chain. Prove that, if the diagonal entries p iiare positive, then the chain is regular. 27Prove that if Pis the transition matrix of an ergodic chain, then (1 =2)(I+P) is the transition matrix of a regular chain. Hint: Use Exercise 26. 28Prove that Pand (1=2)(I+P) have the same flxed vectors. 29In his book, Wahrscheinlichkeitsrechnung und Statistik,15A. Engle proposes an algorithm for flnding the flxed vector for an ergodic Markov chain whenthe transition probabilities are rational numbers. Here is his algorithm: For 15A. Engle, Wahrscheinlichkeitsrechnung und Statistik, vol. 2 (Stuttgart: Klett Verlag, 1976). 446 CHAPTER 11. MARKOV CHAINS (4 2 4) (5 2 3)(8 2 4)(7 3 4)(8 4 4)(8 3 5)(8 4 8)(10 4 6)(12 4 8)(12 5 7)(12 6 8)(13 5 8)(16 6 8)(15 6 9)(16 6 12)(17 7 10)(20 8 12)(20 8 12) : Table 11.4: Distribution of chips. each statei, leta ibe the least common multiple of the denominators of the non-zero entries in the ith row. Engle describes his algorithm in terms of mov- ing chips around on the states|indeed, for small examples, he recommendsimplementing the algorithm this way. Start by putting a ichips on state ifor alli. Then, at each state, redistribute the aichips, sending aipijto statej. The number of chips at state iafter this redistribution need not be a multiple ofai. For each state i, add just enough chips to bring the number of chips at stateiup to a multiple of ai. Then redistribute the chips in the same manner. This process will eventually reach a point where the number of chips at eachstate, after the redistribution, is the same as before redistribution. At thispoint, we have found a flxed vector. Here is an example: P=0 @123 11=21=41=4 21=201=2 31=21=41=41 A: We start with a=( 4;2;4). The chips after successive redistributions are shown in Table 11.4. We flnd that a= (20;8;12) is a flxed vector. (a) Write a computer program to implement this algorithm. (b) Prove that the algorithm will stop. Hint: Let bbe a vector with integer components that is a flxed vector for Pand such that each coordinate of 11.4. FUNDAMENTAL LIMIT THEOREM 447 the starting vector ais less than or equal to the corresponding component ofb. Show that, in the iteration, the components of the vectors are always increasing, and always less than or equal to the correspondingcomponent of b. 30(Cofiman, Kaduta, and Shepp 16) A computing center keeps information on a tape in positions of unit length. During each time unit there is one request tooccupy a unit of tape. When this arrives the flrst free unit is used. Also, duringeach second, each of the units that are occupied is vacated with probability p. Simulate this process, starting with an empty tape. Estimate the expectednumber of sites occupied for a given value of p.I fpis small, can you choose the tape long enough so that there is a small probability that a new job will haveto be turned away (i.e., that all the sites are occupied)? Form a Markov chain with states the number of sites occupied. Modify the program FixedVector to compute the flxed vector. Use this to check your conjecture by simulation. *31 (Alternate proof of Theorem 11.8) Let Pbe the transition matrix of an ergodic Markov chain. Let xbe any column vector such that Px=x. LetMbe the maximum value of the components of x. Assume that x i=M. Show that if pij>0 thenxj=M. Use this to prove that xmust be a constant vector. 32LetPbe the transition matrix of an ergodic Markov chain. Let wbe a flxed probability vector (i.e., wis a row vector with wP=w). Show that if wi=0 andpji>0 thenwj= 0. Use this to show that the flxed probability vector for an ergodic chain cannot have any 0 entries. 33Find a Markov chain that is neither absorbing or ergodic. 11.4 Fundamental Limit Theorem for Regular Chains The fundamental limit theorem for regular Markov chains states that if Pis a regular transition matrix then lim n!1Pn=W; where Wis a matrix with each row equal to the unique flxed probability row vector wforP. In this section we shall give two very difierent proofs of this theorem. Our flrst proof is carried out by showing that, for any column vector y,Pny tends to a constant vector. As indicated in Section 11.3, this will show that Pn converges to a matrix with constant columns or, equivalently, to a matrix with all rows the same. The following lemma says that if an r-by-rtransition matrix has no zero entries, andyis any column vector with rentries, then the vector Pyhas entries which are \closer together" than the entries are in y. 16E. G. Cofiman, J. T. Kaduta, and L. A. Shepp, \On the Asymptotic Optimality of First- Storage Allocation," IEEE Trans. Software Engineering, vol. II (1985), pp. 235-239. 448 CHAPTER 11. MARKOV CHAINS Lemma 11.1 LetPbe anr-by-rtransition matrix with no zero entries. Let dbe the smallest entry of the matrix. Let ybe a column vector with rcomponents, the largest of which is M0and the smallest m0. LetM1andm1be the largest and smallest component, respectively, of the vector Py. Then M1¡m1•(1¡2d)(M0¡m0): Proof. In the discussion following Theorem11.7, it was noted that each entry in the vector Pyis a weighted average of the entries in y. The largest weighted average that could be obtained in the present case would occur if all but one of the entriesofyhave valueM 0and one entry has value m0, and this one small entry is weighted by the smallest possible weight, namely d. In this case, the weighted average would equal dm0+( 1¡d)M0: Similarly, the smallest possible weighted average equals dM0+( 1¡d)m0: Thus, M1¡m1•‡ dm0+( 1¡d)M0· ¡‡ dM0+( 1¡d)m0· =( 1¡2d)(M0¡m0): This completes the proof of the lemma. 2 We turn now to the proof of the fundamental limit theorem for regular Markov chains. Theorem 11.13 (Fundamental Limit Theorem for Regular Chains) IfPis the transition matrix for a regular Markov chain, then lim n!1Pn=W; where Wis matrix with all rows equal. Furthermore, all entries in Ware strictly positive. Proof. We prove this theorem for the special case that Phas no 0 entries. The extension to the general case is indicated in Exercise 5. Let ybe anyr-component column vector, where ris the number of states of the chain. We assume that r> 1, since otherwise the theorem is trivial. Let Mnandmnbe, respectively, the maximum and minimum components of the vector Pny. The vector Pnyis obtained from the vector Pn¡1yby multiplying on the left by the matrix P. Hence each component of Pnyis an average of the components of Pn¡1y.T h u s M0‚M1‚M2‚¢¢¢ 11.4. FUNDAMENTAL LIMIT THEOREM 449 and m0•m1•m2•¢¢¢: Each sequence is monotone and bounded: m0•mn•Mn•M0: Hence, each of these sequences will have a limit as ntends to inflnity. LetMbe the limit of Mnandmthe limit of mn. We know that m•M.W e shall prove that M¡m= 0. This will be the case if Mn¡mntends to 0. Let d be the smallest element of P. Since all entries of Pare strictly positive, we have d>0. By our lemma Mn¡mn•(1¡2d)(Mn¡1¡mn¡1): From this we see that Mn¡mn•(1¡2d)n(M0¡m0): Sincer‚2, we must have d•1=2, so 0•1¡2d<1, so the difierence Mn¡mn tends to 0 as ntends to inflnity. Since every component of Pnylies between mnandMn, each component must approach the same number u=M=m. This shows that lim n!1Pny=u; where uis a column vector all of whose components equal u. Now let ybe the vector with jth component equal to 1 and all other components equal to 0. Then Pnyis thejth column of Pn. Doing this for each jproves that the columns of Pnapproach constant column vectors. That is, the rows of Pnapproach a common row vector w, or, lim n!1Pn=W: It remains to show that all entries in Ware strictly positive. As before, let y be the vector with jth component equal to 1 and all other components equal to 0. Then Pyis thejth column of P, and this column has all entries strictly positive. The minimum component of the vector Pywas deflned to be m1, hencem1>0. Sincem1•m, we havem> 0. Note flnally that this value of mis just the jth component of w, so all components of ware strictly positive. 2 Doeblin’s Proof We give now a very difierent proof of the main part of the fundamental limit theorem for regular Markov chains. This proof was flrst given by Doeblin,17a brilliant young mathematician who was killed in his twenties in the Second World War. 17W. Doeblin, \Expos¶ ed el aT h ¶ eorie des Chaines Simple Constantes de Markov µ a un Nombre Fini d’Etats," Rev. Mach. de l’Union Interbalkanique, vol. 2 (1937), pp. 77{105. 450 CHAPTER 11. MARKOV CHAINS Theorem 11.14 LetPbe the transition matrix for a regular Markov chain with flxed vector w. Then for any initial probability vector u,uPn!wasn!1: Proof. LetX0;X 1; ::: be a Markov chain with transition matrix Pstarted in statesi. LetY0;Y1; ::: be a Markov chain with transition probability Pstarted with initial probabilities given by w. TheXandYprocesses are run independently of each other. We consider also a third Markov chain P⁄which consists of watching both the XandYprocesses. The states for P⁄are pairs (si;sj). The transition probabilities are given by P⁄[(i;j);(k;l)] =P(i;j)¢P(k;l): Since Pis regular there is an Nsuch that PN(i;j)>0 for alliandj. Thus for the P⁄chain it is also possible to go from any state ( si;sj) to any other state ( sk;sl) in at mostNsteps. That is P⁄is also a regular Markov chain. We know that a regular Markov chain will reach any state in a flnite time. Let T be the flrst time the the chain P⁄is in a state of the form ( sk;sk). In other words, Tis the flrst time that the Xand theYprocesses are in the same state. Then we have shown that P[T>n ]!0a sn!1: If we watch the XandYprocesses after the flrst time they are in the same state we would not predict any difierence in their long range behavior. Since this willhappen no matter how we started these two processes, it seems clear that the longrange behaviour should not depend upon the starting state. We now show that thisis true. We flrst note that if n‚T, then since XandYare both in the same state at timeT, P(X n=jjn‚T)=P(Yn=jjn‚T): If we multiply both sides of this equation by P(n‚T), we obtain P(Xn=j; n‚T)=P(Yn=j; n‚T): (11.1) We know that for all n, P(Yn=j)=wj: But P(Yn=j)=P(Yn=j; n‚T)+P(Yn=j;n<T ); and the second summand on the right-hand side of this equation goes to 0 as ngoes to1, sinceP(n<T )g o e st o0a s ngoes to1. So, P(Yn=j; n‚T)!wj; asngoes to1. From Equation 11.1, we see that P(Xn=j; n‚T)!wj; 11.4. FUNDAMENTAL LIMIT THEOREM 451 asngoes to1. But by similar reasoning to that used above, the difierence between this last expression and P(Xn=j)g o e st o0a s ngoes to1. Therefore, P(Xn=j)!wj; asngoes to1. This completes the proof. 2 In the above proof, we have said nothing about the rate at which the distributions of theXn’s approach the flxed distribution w. In fact, it can be shown that18 rX j=1jP(Xn=j)¡wjj•2P(T>n ): The left-hand side of this inequality can be viewed as the distance between the distribution of the Markov chain after nsteps, starting in state si, and the limiting distribution w. Exercises 1Deflne Pandyby P=µ:5:5 :25:75¶ ; y=µ1 0¶ : Compute Py,P2y, and P4yand show that the results are approaching a constant vector. What is this vector? 2LetPbe a regular r£rtransition matrix and yanyr-component column vector. Show that the value of the limiting constant vector for Pnyiswy. 3Let P=0 @100 :25 0:75 0011 A be a transition matrix of a Markov chain. Find two flxed vectors of Pthat are linearly independent. Does this show that the Markov chain is not regular? 4Describe the set of all flxed column vectors for the chain given in Exercise 3. 5The theorem that Pn!Wwas proved only for the case that Phas no zero entries. Fill in the details of the following extension to the case that Pis regular. Since Pis regular, for some N;PNhas no zeros. Thus, the proof given shows that MnN¡mnNapproaches 0 as ntends to inflnity. However, the difierence Mn¡mncan never increase. (Why?) Hence, if we know that the difierences obtained by looking at every Nth time tend to 0, then the entire sequence must also tend to 0. 6LetPbe a regular transition matrix and let wbe the unique non-zero flxed vector of P. Show that no entry of wis 0. 18T. Lindvall, Lectures on the Coupling Method (New York: Wiley 1992). 452 CHAPTER 11. MARKOV CHAINS 7Here is a trick to try on your friends. Shu†e a deck of cards and deal them out one at a time. Count the face cards each as ten. Ask your friend to lookat one of the flrst ten cards; if this card is a six, she is to look at the card thatturns up six cards later; if this card is a three, she is to look at the card thatturns up three cards later, and so forth. Eventually she will reach a pointwhere she is to look at a card that turns up xcards later but there are not xcards left. You then tell her the last card that she looked at even though you did not know her starting point. You tell her you do this by watchingher, and she cannot disguise the times that she looks at the cards. In fact youjust do the same procedure and, even though you do not start at the samepoint as she does, you will most likely end at the same point. Why? 8Write a program to play the game in Exercise 7. 11.5 Mean First Passage Time for Ergodic Chains In this section we consider two closely related descriptive quantities of interest for ergodic chains: the mean time to return to a state and the mean time to go fromone state to another state. LetPbe the transition matrix of an ergodic chain with states s 1,s2,...,sr. Let w=(w1;w2;:::;wr) be the unique probability vector such that wP=w. Then, by the Law of Large Numbers for Markov chains, in the long run the process willspend a fraction w jof the time in state sj. Thus, if we start in any state, the chain will eventually reach state sj; in fact, it will be in state sjinflnitely often. Another way to see this is the following: Form a new Markov chain by making sjan absorbing state, that is, deflne pjj= 1. If we start at any state other than sj, this new process will behave exactly like the original chain up to the flrst time thatstates jis reached. Since the original chain was an ergodic chain, it was possible to reachsjfrom any other state. Thus the new chain is an absorbing chain with a single absorbing state sjthat will eventually be reached. So if we start the original chain at a state siwithi6=j, we will eventually reach the state sj. LetNbe the fundamental matrix for the new chain. The entries of Ngive the expected number of times in each state before absorption. In terms of the originalchain, these quantities give the expected number of times in each of the states beforereaching state s jfor the flrst time. The ith component of the vector Ncgives the expected number of steps before absorption in the new chain, starting in state si. In terms of the old chain, this is the expected number of steps required to reachstates jfor the flrst time starting at state si. Mean First Passage Time Deflnition 11.7 If an ergodic Markov chain is started in state si, the expected number of steps to reach state sjfor the flrst time is called the mean flrst passage time fromsitosj. It is denoted by mij. By convention mii=0 . 2 11.5. MEAN FIRST PASSAGE TIME 453 1 2 3 4 5 6 7 89 Figure 11.5: The maze problem. Example 11.24 Let us return to the maze example (Example 11.22). We shall make this ergodic chain into an absorbing chain by making state 5 an absorbingstate. For example, we might assume that food is placed in the center of the mazeand once the rat flnds the food, he stays to enjoy it (see Figure 11.5). The new transition matrix in canonical form is P=0 BBBBBBBBBBBBBB@123467895 10 1=20 01 =2 000 0 21=301=3 00000 1=3 30 1=201=2 0000 0 40 0 1 =30 01 =301=31=3 61=3 0000000 1=3 7 0000 1 =201=20 0 8 00000 1 =301=31=3 9 000 1 =20 01 =20 0 5 00000000 11 CCCCCCCCCCCCCCA: If we compute the fundamental matrix N, we obtain N=1 80 BBBBBBBBBB@1 4 9439432 6 1 4 64422249 1 4 93234246 1 4 22466422 1 4 64243239 1 4 94222446 1 4 62349349 1 41 CCCCCCCCCCA: The expected time to absorption for difierent starting states is given by the vec- 454 CHAPTER 11. MARKOV CHAINS torNc, where Nc=0 BBBBBBBBBB@6 56556561 CCCCCCCCCCA: We see that, starting from compartment 1, it will take on the average six steps to reach food. It is clear from symmetry that we should get the same answer forstarting at state 3, 7, or 9. It is also clear that it should take one more step,starting at one of these states, than it would starting at 2, 4, 6, or 8. Some of theresults obtained from Nare not so obvious. For instance, we note that the expected number of times in the starting state is 14/8 regardless of the state in which westart. 2 Mean Recurrence Time A quantity that is closely related to the mean flrst passage time is the mean recur- rence time, deflned as follows. Assume that we start in state si; consider the length of time before we return to sifor the flrst time. It is clear that we must return, since we either stay at sithe flrst step or go to some other state sj, and from any other state sj, we will eventually reach sibecause the chain is ergodic. Deflnition 11.8 If an ergodic Markov chain is started in state si, the expected number of steps to return to sifor the flrst time is the mean recurrence time forsi. It is denoted by ri. 2 We need to develop some basic properties of the mean flrst passage time. Con- sider the mean flrst passage time from sitosj; assume that i6=j. This may be computed as follows: take the expected number of steps required given the outcomeof the flrst step, multiply by the probability that this outcome occurs, and add. Ifthe flrst step is to s j, the expected number of steps required is 1; if it is to some other state sk, the expected number of steps required is mkjplus 1 for the step already taken. Thus, mij=pij+X k6=jpik(mkj+1 ); or, sinceP kpik=1 , mij=1+X k6=jpikmjk: (11.2) Similarly, starting in si, it must take at least one step to return. Considering all possible flrst steps gives us ri=X kpik(mki+ 1) (11.3) 11.5. MEAN FIRST PASSAGE TIME 455 =1 +X kpikmki: (11.4) Mean First Passage Matrix and Mean Recurrence Matrix Let us now deflne two matrices MandD. Theijth entrymijofMis the mean flrst passage time to go from sitosjifi6=j; the diagonal entries are 0. The matrix M is called the mean flrst passage matrix. The matrix Dis the matrix with all entries 0 except the diagonal entries dii=ri. The matrix Dis called the mean recurrence matrix. LetCbe anr£rmatrix with all entries 1. Using Equation 11.2 for the casei6=jand Equation 11.4 for the case i=j, we obtain the matrix equation M=PM +C¡D; (11.5) or (I¡P)M=C¡D: (11.6) Equation 11.6 with mii= 0 implies Equations 11.2 and 11.4. We are now in a position to prove our flrst basic theorem. Theorem 11.15 For an ergodic Markov chain, the mean recurrence time for state siisri=1=wi, wherewiis theith component of the flxed probability vector for the transition matrix. Proof. Multiplying both sides of Equation 11.6 by wand using the fact that w(I¡P)=0 gives wC¡wD=0: HerewCis a row vector with all entries 1 and wDis a row vector with ith entry wiri.T h u s (1;1;:::; 1 )=(w1r1;w2r2;:::;wnrn) and ri=1=wi; as was to be proved. 2 Corollary 11.1 For an ergodic Markov chain, the components of the flxed proba- bility vector ware strictly positive. Proof. We know that the values of riare flnite and so wi=1=ricannot be 0. 2 456 CHAPTER 11. MARKOV CHAINS Example 11.25 In Example 11.22 we found the flxed probability vector for the maze example to be w=(1 121 81 121 81 61 81 121 81 12): Hence, the mean recurrence times are given by the reciprocals of these probabilities. That is, r= ( 1 281 28681 281 2 ) : 2 Returning to the Land of Oz, we found that the weather in the Land of Oz could be represented by a Markov chain with states rain, nice, and snow. In Section 11.3we found that the limiting vector was w=( 2=5;1=5;2=5). From this we see that the mean number of days between rainy days is 5/2, between nice days is 5, andbetween snowy days is 5/2. Fundamental Matrix We shall now develop a fundamental matrix for ergodic chains that will play a rolesimilar to that of the fundamental matrix N=(I¡Q) ¡1for absorbing chains. As was the case with absorbing chains, the fundamental matrix can be used to flnda number of interesting quantities involving ergodic chains. Using this matrix, wewill give a method for calculating the mean flrst passage times for ergodic chainsthat is easier to use than the method given above. In addition, we will state (butnot prove) the Central Limit Theorem for Markov Chains, the statement of whichuses the fundamental matrix. We begin by considering the case that Pis the transition matrix of a regular Markov chain. Since there are no absorbing states, we might be tempted to tryZ=(I¡P) ¡1for a fundamental matrix. But I¡Pdoes not have an inverse. To see this, recall that a matrix Rhas an inverse if and only if Rx=0implies x=0. But since Pc=cwe have ( I¡P)c=0, and so I¡Pdoes not have an inverse. We recall that if we have an absorbing Markov chain, and Qis the restriction of the transition matrix to the set of transient states, then the fundamental matrixNcould be written as N=I+Q+Q 2+¢¢¢: The reason that this power series converges is that Qn!0, so this series acts like a convergent geometric series. This idea might prompt one to try to flnd a similar series for regular chains. Since we know that Pn!W, we might consider the series I+(P¡W)+(P2¡W)+¢¢¢: (11.7) We now use special properties of PandWto rewrite this series. The special properties are: 1) PW =W, and 2) Wk=Wfor all positive integers k. These 11.5. MEAN FIRST PASSAGE TIME 457 facts are easy to verify, and are left as an exercise (see Exercise 22). Using these facts, we see that (P¡W)n=nX i=0(¡1)iµn i¶ Pn¡iWi =Pn+nX i=1(¡1)iµn i¶ Wi =Pn+nX i=1(¡1)iµn i¶ W =Pn+ˆnX i=1(¡1)iµn i¶! W: If we expand the expression (1 ¡1)n, using the Binomial Theorem, we obtain the expression in parenthesis above, except that we have an extra term (which equals1). Since (1¡1) n= 0, we see that the above expression equals -1. So we have (P¡W)n=Pn¡W; for alln‚1. We can now rewrite the series in 11.7 as I+(P¡W)+(P¡W)2+¢¢¢: Since thenth term in this series is equal to Pn¡W, thenth term goes to 0 as n goes to inflnity. This is su–cient to show that this series converges, and sums tothe inverse of the matrix I¡P+W. We call this inverse the fundamental matrix associated with the chain, and we denote it by Z. In the case that the chain is ergodic, but not regular, it is not true that P n!W asn!1 . Nevertheless, the matrix I¡P+Wstill has an inverse, as we will now show. Proposition 11.1 LetPbe the transition matrix of an ergodic chain, and let W be the matrix all of whose rows are the flxed probability row vector for P. Then the matrix I¡P+W has an inverse. Proof. Letxbe a column vector such that (I¡P+W)x=0: To prove the proposition, it is su–cient to show that xmust be the zero vector. Multiplying this equation by wand using the fact that w(I¡P)=0andwW =w, we have w(I¡P+W)x=wx=0: 458 CHAPTER 11. MARKOV CHAINS Therefore, (I¡P)x=0: But this means that x=Pxis a flxed column vector for P. By Theorem 11.10, this can only happen if xis a constant vector. Since wx= 0, and whas strictly positive entries, we see that x=0. This completes the proof. 2 As in the regular case, we will call the inverse of the matrix I¡P+Wthe fundamental matrix for the ergodic chain with transition matrix P, and we will use Zto denote this fundamental matrix. Example 11.26 LetPbe the transition matrix for the weather in the Land of Oz. Then I¡P+W =0 @100 0100011 A¡0 @1=21=41=4 1=201=2 1=41=41=21 A+0 @2=51=52=5 2=51=52=5 2=51=52=51 A =0 @9=10¡1=20 3=20 ¡1=10 6=5¡1=10 3=20¡1=20 9=101 A; so Z=(I¡P+W) ¡1=0 @86=75 1=25¡14=75 2=25 21=25 2=25 ¡14=75 1=25 86=751 A: 2 Using the Fundamental Matrix to Calculate the Mean First Passage Matrix We shall show how one can obtain the mean flrst passage matrix Mfrom the fundamental matrix Zfor an ergodic Markov chain. Before stating the theorem which gives the flrst passage times, we need a few facts about Z. Lemma 11.2 LetZ=(I¡P+W)¡1, and let cbe a column vector of all 1’s. Then Zc=c; wZ=w; and Z(I¡P)=I¡W: Proof. Since Pc=candWc=c, c=(I¡P+W)c: If we multiply both sides of this equation on the left by Z, we obtain Zc=c: 11.5. MEAN FIRST PASSAGE TIME 459 Similarly, since wP=wandwW =w, w=w(I¡P+W): If we multiply both sides of this equation on the right by Z, we obtain wZ=w: Finally, we have (I¡P+W)(I¡W)= I¡W¡P+W+W¡W =I¡P: Multiplying on the left by Z, we obtain I¡W=Z(I¡P): This completes the proof. 2 The following theorem shows how one can obtain the mean flrst passage times from the fundamental matrix. Theorem 11.16 The mean flrst passage matrix Mfor an ergodic chain is deter- mined from the fundamental matrix Zand the flxed row probability vector wby mij=zjj¡zij wj: Proof. We showed in Equation 11.6 that (I¡P)M=C¡D: Thus, Z(I¡P)M=ZC¡ZD; and from Lemma 11.2, Z(I¡P)M=C¡ZD: Again using Lemma 11.2, we have M¡WM =C¡ZD or M=C¡ZD+WM: From this equation, we see that mij=1¡zijrj+(wM)j: (11.8) Butmjj= 0, and so 0=1¡zjjrj+(wM)j; 460 CHAPTER 11. MARKOV CHAINS or (wM)j=zjjrj¡1: (11.9) From Equations 11.8 and 11.9, we have mij=(zjj¡zij)¢rj: Sincerj=1=wj, mij=zjj¡zij wj: 2 Example 11.27 (Example 11.26 continued) In the Land of Oz example, we flnd that Z=(I¡P+W)¡1=0 @86=75 1=25¡14=75 2=25 21=25 2=25 ¡14=75 1=25 86=751 A: We have also seen that w=( 2=5;1=5;2=5). So, for example, m12=z22¡z12 w2 =21=25¡1=25 1=5 =4; by Theorem 11.16. Carrying out the calculations for the other entries of M,w e obtain M=0 @04 1 0=3 8=308=3 10=34 01 A: 2 Computation The program ErgodicChain calculates the fundamental matrix, the flxed vector, the mean recurrence matrix D, and the mean flrst passage matrix M. We have run the program for the Ehrenfest urn model (Example 11.8). We obtain: P=0 BBBB@01234 0:0000 1:0000:0000:0000:0000 1:2500:0000:7500:0000:0000 2:0000:5000:0000:5000:0000 3:0000:0000:7500:0000:2500 4:0000:0000:0000 1:0000:00001 CCCCA; w=¡01234 :0625:2500:3750:2500:0625¢ ; 11.5. MEAN FIRST PASSAGE TIME 461 r=¡0 123 4 16:0000 4:0000 2:6667 4:0000 16:0000¢ ; M=0 BBBB@0 123 4 0:0000 1:0000 2:6667 6:3333 21:3333 11 5:0000:0000 1:6667 5:3333 20:3333 21 8:6667 3:6667:0000 3:6667 18:6667 32 0:3333 5:3333 1:6667:0000 15:0000 42 1:3333 6:3333 2:6667 1:0000:00001 CCCCA: From the mean flrst passage matrix, we see that the mean time to go from 0 balls in urn 1 to 2 balls in urn 1 is 2.6667 steps while the mean time to go from 2 balls inurn 1 to 0 balls in urn 1 is 18.6667. This re°ects the fact that the model exhibits acentral tendency. Of course, the physicist is interested in the case of a large numberof molecules, or balls, and so we should consider this example for nso large that we cannot compute it even with a computer. Ehrenfest Model Example 11.28 (Example 11.23 continued) Let us consider the Ehrenfest model (see Example 11.8) for gas difiusion for the general case of 2 nballs. Every second, one of the 2 nballs is chosen at random and moved from the urn it was in to the other urn. If there are iballs in the flrst urn, then with probability i=2nwe take one of them out and put it in the second urn, and with probability (2 n¡i)=2nwe take a ball from the second urn and put it in the flrst urn. At each second we letthe number iof balls in the flrst urn be the state of the system. Then from state i we can pass only to state i¡1 andi+ 1, and the transition probabilities are given by p ij=8 < :i 2n; ifj=i¡1; 1¡i 2n;ifj=i+1; 0; otherwise. This deflnes the transition matrix of an ergodic, non-regular Markov chain (see Exercise 15). Here the physicist is interested in long-term predictions about thestate occupied. In Example 11.23, we gave an intuitive reason for expecting thatthe flxed vector wis the binomial distribution with parameters 2 nand 1=2. It is easy to check that this is correct. So, w i=¡2n i¢ 22n: Thus the mean recurrence time for state iis ri=22n ¡2n i¢: 462 CHAPTER 11. MARKOV CHAINS 0 200 400 600 800 1000404550556065 0 200 400 600 800 1000404550556065Time forward Time reversed Figure 11.6: Ehrenfest simulation. Consider in particular the central term i=n. We have seen that this term is approximately 1 =p…n. Thus we may approximate rnbyp…n. This model was used to explain the concept of reversibility in physical systems. Assume that we let our system run until it is in equilibrium. At this point, a movieis made, showing the system’s progress. The movie is then shown to you, and youare asked to tell if the movie was shown in the forward or the reverse direction.It would seem that there should always be a tendency to move toward an equalproportion of balls so that the correct order of time should be the one with themost transitions from itoi¡1i fi>n anditoi+1i fi<n . In Figure 11.6 we show the results of simulating the Ehrenfest urn model for the case of n= 50 and 1000 time units, using the program EhrenfestUrn . The top graph shows these results graphed in the order in which they occurred and thebottom graph shows the same results but with time reversed. There is no apparentdifierence. 11.5. MEAN FIRST PASSAGE TIME 463 We note that if we had not started in equilibrium, the two graphs would typically look quite difierent. 2 Reversibility If the Ehrenfest model is started in equilibrium, then the process has no apparent time direction. The reason for this is that this process has a property called re- versibility. DeflneXnto be the number of balls in the left urn at step n. We can calculate, for a general ergodic chain, the reverse transition probability: P(Xn¡1=jjXn=i)=P(Xn¡1=j;Xn=i) P(Xn=i) =P(Xn¡1=j)P(Xn=ijXn¡1=j) P(Xn=i) =P(Xn¡1=j)pji P(Xn=i): In general, this will depend upon n, sinceP(Xn=j) and also P(Xn¡1=j) change with n. However, if we start with the vector wor wait until equilibrium is reached, this will not be the case. Then we can deflne p⁄ ij=wjpji wi as a transition matrix for the process watched with time reversed. Let us calculate a typical transition probability for the reverse chain P⁄=fp⁄ ijg in the Ehrenfest model. For example, p⁄ i;i¡1=wi¡1pi¡1;i wi=¡2n i¡1¢ 22n£2n¡i+1 2n£22n ¡2n i¢ =(2n)! (i¡1)! (2n¡i+ 1)!£(2n¡i+1 )i!( 2n¡i)! 2n(2n)! =i 2n=pi;i¡1: Similar calculations for the other transition probabilities show that P⁄=P. When this occurs the process is called reversible. Clearly, an ergodic chain is re- versible if, and only if, for every pair of states siandsj,wipij=wjpji. In particular, for the Ehrenfest model this means that wipi;i¡1=wi¡1pi¡1;i. Thus, in equilib- rium, the pairs ( i;i¡1) and (i¡1;i) should occur with the same frequency. While many of the Markov chains that occur in applications are reversible, this is a verystrong condition. In Exercise 12 you are asked to flnd an example of a Markov chainwhich is not reversible. The Central Limit Theorem for Markov Chains Suppose that we have an ergodic Markov chain with states s1;s2;:::;sk.I t i s natural to consider the distribution of the random variables S(n) j, which denotes 464 CHAPTER 11. MARKOV CHAINS the number of times that the chain is in state sjin the flrst nsteps. The jth component wjof the flxed probability row vector wis the proportion of times that the chain is in state sjin the long run. Hence, it is reasonable to conjecture that the expected value of the random variable S(n) j,a sn!1 , is asymptotic to nwj, and it is easy to show that this is the case (see Exercise 23). It is also natural to ask whether there is a limiting distribution of the random variablesS(n) j. The answer is yes, and in fact, this limiting distribution is the normal distribution. As in the case of independent trials, one must normalize these randomvariables. Thus, we must subtract from S (n) jits expected value, and then divide by its standard deviation. In both cases, we will use the asymptotic values of thesequantities, rather than the values themselves. Thus, in the flrst case, we will usethe valuenw j. It is not so clear what we should use in the second case. It turns out that the quantity ¾2 j=2wjzjj¡wj¡w2 j (11.10) represents the asymptotic variance. Armed with these ideas, we can state the following theorem. Theorem 11.17 (Central Limit Theorem for Markov Chains) F o ra ne r - godic chain, for any real numbers r<s ,w eh a v e Pˆ r<S(n) j¡nwjq n¾2 j<s! !1p 2…Zs re¡x2=2dx ; asn!1 , for any choice of starting state, where ¾2 jis the quantity deflned in Equation 11.10. 2 Historical Remarks Markov chains were introduced by Andre ‚i Andreevich Markov (1856{1922) and were named in his honor. He was a talented undergraduate who received a goldmedal for his undergraduate thesis at St. Petersburg University. Besides beingan active research mathematician and teacher, he was also active in politics andpatricipated in the liberal movement in Russia at the beginning of the twentiethcentury. In 1913, when the government celebrated the 300th anniversary of theHouse of Romanov family, Markov organized a counter-celebration of the 200thanniversary of Bernoulli’s discovery of the Law of Large Numbers. Markov was led to develop Markov chains as a natural extension of sequences of independent random variables. In his flrst paper, in 1906, he proved that for aMarkov chain with positive transition probabilities and numerical states the averageof the outcomes converges to the expected value of the limiting distribution (theflxed vector). In a later paper he proved the central limit theorem for such chains.Writing about Markov, A. P. Youschkevitch remarks: Markov arrived at his chains starting from the internal needs of prob- ability theory, and he never wrote about their applications to physical 11.5. MEAN FIRST PASSAGE TIME 465 science. For him the only real examples of the chains were literary texts, where the two states denoted the vowels and consonants.19 In a paper written in 1913,20Markov chose a sequence of 20,000 letters from Pushkin’s Eugene Onegin to see if this sequence can be approximately considered a simple chain. He obtained the Markov chain with transition matrix µvowel consonant vowel :128:872 consonant :663:337¶ : The flxed vector for this chain is ( :432;:568), indicating that we should expect about 43.2 percent vowels and 56.8 percent consonants in the novel, which was borne out by the actual count. Claude Shannon considered an interesting extension of this idea in his book The Mathematical Theory of Communication,21in which he developed the information- theoretic concept of entropy. Shannon considers a series of Markov chain approxi-mations to English prose. He does this flrst by chains in which the states are lettersand then by chains in which the states are words. For example, for the case ofwords he presents flrst a simulation where the words are chosen independently butwith appropriate frequencies. REPRESENTING AND SPEEDILY IS AN GOOD APT OR COME CAN DIFFERENT NATURAL HERE HE THE A IN CAME THE TOOF TO EXPERT GRAY COME TO FURNISHES THE LINE MES-SAGE HAD BE THESE. He then notes the increased resemblence to ordinary English text when the words are chosen as a Markov chain, in which case he obtains THE HEAD AND IN FRONTAL ATTACK ON AN ENGLISH WRI- TER THAT THE CHARACTER OF THIS POINT IS THEREFOREANOTHER METHOD FOR THE LETTERS THAT THE TIME OFWHO EVER TOLD THE PROBLEM FOR AN UNEXPECTED. A simulation like the last one is carried out by opening a book and choosing the flrst word, say it is the. Then the book is read until the word theappears again and the word after this is chosen as the second word, which turned out to be head. The book is then read until the word head appears again and the next word, and, is chosen, and so on. Other early examples of the use of Markov chains occurred in Galton’s study of the problem of survival of family names in 1889 and in the Markov chain introduced 19SeeDictionary of Scientiflc Biography, ed. C. C. Gillespie (New York: Scribner’s Sons, 1970), pp. 124{130. 20A. A. Markov, \An Example of Statistical Analysis of the Text of Eugene Onegin Illustrat- ing the Association of Trials into a Chain," Bulletin de l’Acadamie Imperiale des Sciences de St. Petersburg, ser. 6, vol. 7 (1913), pp. 153{162. 21C. E. Shannon and W. Weaver, The Mathematical Theory of Communication (Urbana: Univ. of Illinois Press, 1964). 466 CHAPTER 11. MARKOV CHAINS by P. and T. Ehrenfest in 1907 for difiusion. Poincar¶ e in 1912 dicussed card shu†ing in terms of an ergodic Markov chain deflned on a permutation group. Brownianmotion, a continuous time version of random walk, was introducted in 1900{1901by L. Bachelier in his study of the stock market, and in 1905{1907 in the works ofA. Einstein and M. Smoluchowsky in their study of physical processes. One of the flrst systematic studies of flnite Markov chains was carried out by M. Frechet. 22The treatment of Markov chains in terms of the two fundamental matrices that we have used was developed by Kemeny and Snell23to avoid the use of eigenvalues that one of these authors found too complex. The fundamental matrix N occurred also in the work of J. L. Doob and others in studying the connectionbetween Markov processes and classical potential theory. The fundamental matrix Z for ergodic chains appeared flrst in the work of Frechet, who used it to flnd thelimiting variance for the central limit theorem for Markov chains. Exercises 1Consider the Markov chain with transition matrix P=µ1=21=2 1=43=4¶ : Find the fundamental matrix Zfor this chain. Compute the mean flrst passage matrix using Z. 2A study of the strengths of Ivy League football teams shows that if a school has a strong team one year it is equally likely to have a strong team or averageteam next year; if it has an average team, half the time it is average next year,and if it changes it is just as likely to become strong as weak; if it is weak ithas 2/3 probability of remaining so and 1/3 of becoming average. (a) A school has a strong team. On the average, how long will it be before it has another strong team? (b) A school has a weak team; how long (on the average) must the alumni wait for a strong team? 3Consider Example 11.4 with a=:5 andb=:75. Assume that the President says that he or she will run. Find the expected length of time before the flrsttime the answer is passed on incorrectly. 4Find the mean recurrence time for each state of Example 11.4 for a=:5 and b=:75. Do the same for general aandb. 5A die is rolled repeatedly. Show by the results of this section that the mean time between occurrences of a given number is 6. 22M. Frechet, \Th¶ eorie des ¶ ev¶enements en chaine dans le cas d’un nombre flni d’¶ etats possible," inRecherches th¶ eoriques Modernes sur le calcul des probabilit¶ es,vol. 2 (Paris, 1938). 23J. G. Kemeny and J. L. Snell, Finite Markov Chains. 11.5. MEAN FIRST PASSAGE TIME 467 24 3 6 51 Figure 11.7: Maze for Exercise 7. 6For the Land of Oz example (Example 11.1), make rain into an absorbing state and flnd the fundamental matrix N. Interpret the results obtained from this chain in terms of the original chain. 7A rat runs through the maze shown in Figure 11.7. At each step it leaves the room it is in by choosing at random one of the doors out of the room. (a) Give the transition matrix Pfor this Markov chain. (b) Show that it is an ergodic chain but not a regular chain. (c) Find the flxed vector. (d) Find the expected number of steps before reaching Room 5 for the flrst time, starting in Room 1. 8Modify the program ErgodicChain so that you can compute the basic quan- tities for the queueing example of Exercise 11.3.20. Interpret the mean recur-rence time for state 0. 9Consider a random walk on a circle of circumference n. The walker takes one unit step clockwise with probability pand one unit counterclockwise with probability q=1¡p. Modify the program ErgodicChain to allow you to inputnandpand compute the basic quantities for this chain. (a) For which values of nis this chain regular? ergodic? (b) What is the limiting vector w? (c) Find the mean flrst passage matrix for n= 5 andp=:5. Verify that m ij=d(n¡d), wheredis the clockwise distance from itoj. 10Two players match pennies and have between them a total of 5 pennies. If at any time one player has all of the pennies, to keep the game going, he givesone back to the other player and the game will continue. Show that this gamecan be formulated as an ergodic chain. Study this chain using the programErgodicChain . 468 CHAPTER 11. MARKOV CHAINS 11Calculate the reverse transition matrix for the Land of Oz example (Exam- ple 11.1). Is this chain reversible? 12Give an example of a three-state ergodic Markov chain that is not reversible. 13LetPbe the transition matrix of an ergodic Markov chain and P⁄the reverse transition matrix. Show that they have the same flxed probability vector w. 14IfPis a reversible Markov chain, is it necessarily true that the mean time to go from state ito statejis equal to the mean time to go from state jto statei?Hint: Try the Land of Oz example (Example 11.1). 15Show that any ergodic Markov chain with a symmetric transition matrix (i.e., pij=pji) is reversible. 16(Crowell24) Let Pbe the transition matrix of an ergodic Markov chain. Show that (I+P+¢¢¢+Pn¡1)(I¡P+W)=I¡Pn+nW; and from this show that I+P+¢¢¢+Pn¡1 n!W; asn!1 . 17An ergodic Markov chain is started in equilibrium (i.e., with initial probability vector w). The mean time until the next occurrence of state siis „mi=P kwkmki+wiri. Show that „ mi=zii=wi, by using the facts that wZ=w andmki=(zii¡zki)=wi. 18A perpetual craps game goes on at Charley’s. Jones comes into Charley’s on an evening when there have already been 100 plays. He plans to play until thenext time that snake eyes (a pair of ones) are rolled. Jones wonders how manytimes he will play. On the one hand he realizes that the average time betweensnake eyes is 36 so he should play about 18 times as he is equally likely tohave come in on either side of the halfway point between occurrences of snakeeyes. On the other hand, the dice have no memory, and so it would seemthat he would have to play for 36 more times no matter what the previousoutcomes have been. Which, if either, of Jones’s arguments do you believe?Using the result of Exercise 17, calculate the expected to reach snake eyes, inequilibrium, and see if this resolves the apparent paradox. If you are still indoubt, simulate the experiment to decide which argument is correct. Can yougive an intuitive argument which explains this result? 19Show that, for an ergodic Markov chain (see Theorem 11.16), X jmijwj=X jzjj¡1=K: 24Private communication. 11.5. MEAN FIRST PASSAGE TIME 469 - 5 B20 C - 30 A 15 GO Figure 11.8: Simplifled Monopoly. The second expression above shows that the number Kis independent of i. The number Kis called Kemeny’s constant. A prize was ofiered to the flrst person to give an intuitively plausible reason for the above sum to beindependent of i. (See also Exercise 24.) 20Consider a game played as follows: You are given a regular Markov chain with transition matrix P, flxed probability vector w, and a payofi function f which assigns to each state s ian amountfiwhich may be positive or negative. Assume that wf= 0. You watch this Markov chain as it evolves, and every time you are in state siyou receive an amount fi. Show that your expected winning after nsteps can be represented by a column vector g(n), with g(n)=(I+P+P2+¢¢¢+Pn)f: Show that as n!1 ,g(n)!gwithg=Zf. 21A highly simplifled game of \Monopoly" is played on a board with four squares as shown in Figure 11.8. You start at GO. You roll a die and move clockwisearound the board a number of squares equal to the number that turns up onthe die. You collect or pay an amount indicated on the square on which youland. You then roll the die again and move around the board in the samemanner from your last position. Using the result of Exercise 20, estimatethe amount you should expect to win in the long run playing this version ofMonopoly. 22Show that if Pis the transition matrix of a regular Markov chain, and Wis the matrix each of whose rows is the flxed probability vector correspondingtoP, then PW =W, and W k=Wfor all positive integers k. 23Assume that an ergodic Markov chain has states s1;s2;:::;sk. LetS(n) jdenote the number of times that the chain is in state sjin the flrst nsteps. Let w denote the flxed probability row vector for this chain. Show that, regardlessof the starting state, the expected value of S (n) j, divided by n, tends towjas n!1 .Hint: If the chain starts in state si, then the expected value of S(n) j is given by the expression nX h=0p(h) ij: 470 CHAPTER 11. MARKOV CHAINS 24Peter Doyle25has suggested the following interpretation for Kemeny’s con- stant (see Exercise 19). We are given an ergodic chain and do not know the starting state. However, we would like to start watching it at a time whenit can be considered to be in equilibrium (i.e., as if we had started with theflxed vector wor as if we had waited a long time). However, we don’t know the starting state and we don’t want to wait a long time. Peter says to choosea state according to the flxed vector w. That is, choose state jwith proba- bilityw jusing a spinner, for example. Then wait until the time Tthat this state occurs for the flrst time. We consider Tas our starting time and observe the chain from this time on. Of course the probability that we start in state j iswj, so we are starting in equilibrium. Kemeny’s constant is the expected value ofT, and it is independent of the way in which the chain was started. Should Peter have been given the prize? 25Private communication. Chapter 12 Random Walks 12.1 Random Walks in Euclidean Space In the last several chapters, we have studied sums of random variables with the goal being to describe the distribution and density functions of the sum. In this chapter,we shall look at sums of discrete random variables from a difierent perspective. Weshall be concerned with properties which can be associated with the sequence ofpartial sums, such as the number of sign changes of this sequence, the number ofterms in the sequence which equal 0, and the expected size of the maximum termin the sequence. We begin with the following deflnition. Deflnition 12.1 LetfX kg1 k=1be a sequence of independent, identically distributed discrete random variables. For each positive integer n, we letSndenote the sum X1+X2+¢¢¢+Xn. The sequencefSng1 n=1is called a random walk. If the common range of the Xk’s isRm, then we say that fSngis a random walk in Rm. 2 We view the sequence of Xk’s as being the outcomes of independent experiments. Since theXk’s are independent, the probability of any particular (flnite) sequence of outcomes can be obtained by multiplying the probabilities that each Xktakes on the specifled value in the sequence. Of course, these individual probabilities aregiven by the common distribution of the X k’s. We will typically be interested in flnding probabilities for events involving the related sequence of Sn’s. Such events can be described in terms of the Xk’s, so their probabilities can be calculated using the above idea. There are several ways to visualize a random walk. One can imagine that a particle is placed at the origin in Rmat timen= 0. The sum Snrepresents the position of the particle at the end of nseconds. Thus, in the time interval [ n¡1;n], the particle moves (or jumps) from position Sn¡1toSn. The vector representing this motion is just Sn¡Sn¡1, which equals Xn. This means that in a random walk, the jumps are independent and identically distributed. If m= 1, for example, then one can imagine a particle on the real line that starts at the origin, and at theend of each second, jumps one unit to the right or the left, with probabilities given 471 472 CHAPTER 12. RANDOM WALKS by the distribution of the Xk’s. Ifm= 2, one can visualize the process as taking place in a city in which the streets form square city blocks. A person starts at onecorner (i.e., at an intersection of two streets) and goes in one of the four possibledirections according to the distribution of the X k’s. Ifm= 3, one might imagine being in a jungle gym, where one is free to move in any one of six directions (left,right, forward, backward, up, and down). Once again, the probabilities of thesemovements are given by the distribution of the X k’s. Another model of a random walk (used mostly in the case where the range is R1) is a game, involving two people, which consists of a sequence of independent, identically distributed moves. The sum Snrepresents the score of the flrst person, say, afternmoves, with the assumption that the score of the second person is ¡Sn. For example, two people might be °ipping coins, with a match or non-match representing +1 or ¡1, respectively, for the flrst player. Or, perhaps one coin is being °ipped, with a head or tail representing +1 or ¡1, respectively, for the flrst player. Random Walks on the Real Line We shall flrst consider the simplest non-trivial case of a random walk in R1, namely the case where the common distribution function of the random variables Xnis given by fX(x)=‰1=2;ifx=§1; 0; otherwise. This situation corresponds to a fair coin being °ipped, with Snrepresenting the number of heads minus the number of tails which occur in the flrst n°ips. We note that in this situation, all paths of length nhave the same probability, namely 2¡n. It is sometimes instructive to represent a random walk as a polygonal line, or path, in the plane, where the horizontal axis represents time and the vertical axisrepresents the value of S n. Given a sequence fSngof partial sums, we flrst plot the points (n;Sn), and then for each k<n , we connect ( k;Sk) and (k+1;Sk+1) with a straight line segment. The length of a path is just the difierence in the time values of the beginning and ending points on the path. The reader is referred to Figure12.1. This flgure, and the process it illustrates, are identical with the example,given in Chapter 1, of two people playing heads or tails. Returns and First Returns We say that an equalization has occurred, or there is a return to the origin at time n,i fSn= 0. We note that this can only occur if nis an even integer. To calculate the probability of an equalization at time 2 m, we need only count the number of paths of length 2 mwhich begin and end at the origin. The number of such paths is clearlyµ2m m¶ : Since each path has probability 2¡2m, we have the following theorem. 12.1. RANDOM WALKS IN EUCLIDEAN SPACE 473 510 15 20 25 30 35 40 -10-8-6-4-2246810 Figure 12.1: A random walk of length 40. Theorem 12.1 The probability of a return to the origin at time 2 mis given by u2m=µ2m m¶ 2¡2m: The probability of a return to the origin at an odd time is 0. 2 A random walk is said to have a flrst return to the origin at time 2 mifm> 0, and S2k6= 0 for allk<m . In Figure 12.1, the flrst return occurs at time 2. We deflne f2mto be the probability of this event. (We also deflne f0= 0.) One can think of the expression f2m22mas the number of paths of length 2 mbetween the points (0;0) and (2m;0) that do not touch the horizontal axis except at the endpoints. Using this idea, it is easy to prove the following theorem. Theorem 12.2 Forn‚1, the probabilities fu2kgandff2kgare related by the equation u2n=f0u2n+f2u2n¡2+¢¢¢+f2nu0: Proof. There areu2n22npaths of length 2 nwhich have endpoints (0 ;0) and (2n;0). The collection of such paths can be partitioned into nsets, depending upon the time of the flrst return to the origin. A path in this collection which has a flrst return tothe origin at time 2 kconsists of an initial segment from (0 ;0) to (2k;0), in which no interior points are on the horizontal axis, and a terminal segment from (2 k;0) to (2n;0), with no further restrictions on this segment. Thus, the number of paths in the collection which have a flrst return to the origin at time 2 kis given by f 2k22ku2n¡2k22n¡2k=f2ku2n¡2k22n: If we sum over k, we obtain the equation u2n22n=f0u2n22n+f2u2n¡222n+¢¢¢+f2nu022n: Dividing both sides of this equation by 22ncompletes the proof. 2 474 CHAPTER 12. RANDOM WALKS The expression in the right-hand side of the above theorem should remind the reader of a sum that appeared in Deflnition 7.1 of the convolution of two distributions. Theconvolution of two sequences is deflned in a similar manner. The above theoremsays that the sequence fu 2ngis the convolution of itself and the sequence ff2ng. Thus, if we represent each of these sequences by an ordinary generating function,then we can use the above relationship to determine the value f 2n. Theorem 12.3 Form‚1, the probability of a flrst return to the origin at time 2mis given by f2m=u2m 2m¡1=¡2m m¢ (2m¡1)22m: Proof. We begin by deflning the generating functions U(x)=1X m=0u2mxm and F(x)=1X m=0f2mxm: Theorem 12.2 says that U(x)=1+U(x)F(x): (12.1) (The presence of the 1 on the right-hand side is due to the fact that u0is deflned to be 1, but Theorem 12.2 only holds for m‚1.) We note that both generating functions certainly converge on the interval ( ¡1;1), since all of the coe–cients are at most 1 in absolute value. Thus, we can solve the above equation for F(x), obtaining F(x)=U(x)¡1 U(x): Now, if we can flnd a closed-form expression for the function U(x), we will also have a closed-form expression for F(x). From Theorem 12.1, we have U(x)=1X m=0µ2m m¶ 2¡2mxm: In Wilf,1we flnd that 1p1¡4x=1X m=0µ2m m¶ xm: The reader is asked to prove this statement in Exercise 1. If we replace xbyx=4 in the last equation, we see that U(x)=1p1¡x: 1H. S. Wilf, Generatingfunctionology, (Boston: Academic Press, 1990), p. 50. 12.1. RANDOM WALKS IN EUCLIDEAN SPACE 475 Therefore, we have F(x)=U(x)¡1 U(x) =(1¡x)¡1=2¡1 (1¡x)¡1=2 =1¡(1¡x)1=2: Although it is possible to compute the value of f2musing the Binomial Theorem, it is easier to note that F0(x)=U(x)=2, so that the coe–cients f2mcan be found by integrating the series for U(x). We obtain, for m‚1, f2m=u2m¡2 2m =¡2m¡2 m¡1¢ m22m¡1 =¡2m m¢ (2m¡1)22m =u2m 2m¡1; since µ2m¡2 m¡1¶ =m 2(2m¡1)µ2m m¶ : This completes the proof of the theorem. 2 Probability of Eventual Return In the symmetric random walk process in Rm, what is the probability that the particle eventually returns to the origin? We flrst examine this question in the casethatm= 1, and then we consider the general case. The results in the next two examples are due to P¶ olya. 2 Example 12.1 (Eventual Return in R1) One has to approach the idea of eventual return with some care, since the sample space seems to be the set of all walks ofinflnite length, and this set is non-denumerable. To avoid di–culties, we will deflnew nto be the probability that a flrst return has occurred no later than time n. Thus, wnconcerns the sample space of all walks of length n, which is a flnite set. In terms of thewn’s, it is reasonable to deflne the probability that the particle eventually returns to the origin to be w⁄= lim n!1wn: This limit clearly exists and is at most one, since the sequence fwng1 n=1is an increasing sequence, and all of its terms are at most one. 2G. P¶ olya, \ ˜Uber eine Aufgabe der Wahrscheinlichkeitsrechnung betrefiend die Irrfahrt im Strassennetz," Math. Ann., vol. 84 (1921), pp. 149-160. 476 CHAPTER 12. RANDOM WALKS In terms of the fnprobabilities, we see that w2n=nX i=1f2i: Thus, w⁄=1X i=1f2i: In the proof of Theorem 12.3, the generating function F(x)=1X m=0f2mxm was introduced. There it was noted that this series converges for x2(¡1;1). In fact, it is possible to show that this series also converges for x=§1 by using Exercise 4, together with the fact that f2m=u2m 2m¡1: (This fact was proved in the proof of Theorem 12.3.) Since we also know that F(x)=1¡(1¡x)1=2; we see that w⁄=F( 1 )=1: Thus, with probability one, the particle returns to the origin. An alternative proof of the fact that w⁄= 1 can be obtained by using the results in Exercise 2. 2 Example 12.2 (Eventual Return in Rm) We now turn our attention to the case that the random walk takes place in more than one dimension. We deflne f(m) 2nto be the probability that the flrst return to the origin in Rmoccurs at time 2 n. The quantityu(m) 2nis deflned in a similar manner. Thus, f(1) 2nandu(1) 2nequalf2nandu2n, which were deflned earlier. If, in addition, we deflne u(m) 0= 1 andf(m) 0= 0, then one can mimic the proof of Theorem 12.2, and show that for all m‚1, u(m) 2n=f(m) 0u(m) 2n+f(m) 2u(m) 2n¡2+¢¢¢+f(m) 2nu(m) 0: (12.2) We continue to generalize previous work by deflning U(m)(x)=1X n=0u(m) 2nxn and F(m)(x)=1X n=0f(m) 2nxn: 12.1. RANDOM WALKS IN EUCLIDEAN SPACE 477 Then, by using Equation 12.2, we see that U(m)(x)=1+U(m)(x)F(m)(x); as before. These functions will always converge in the interval ( ¡1;1), since all of their coe–cients are at most one in magnitude. In fact, since w(m) ⁄=1X n=0f(m) 2n•1 for allm, the series for F(m)(x) converges at x= 1 as well, and F(m)(x) is left- continuous at x= 1, i.e., lim x"1F(m)(x)=F(m)(1): Thus, we have w(m) ⁄= lim x"1F(m)(x) = lim x"1U(m)(x)¡1 U(m)(x); (12.3) so to determine w(m) ⁄, it su–ces to determine lim x"1U(m)(x): We letu(m)denote this limit. We claim that u(m)=1X n=0u(m) 2n: (This claim is reasonable; it says that to flnd out what happens to the function U(m)(x)a tx= 1, just let x= 1 in the power series for U(m)(x).) To prove the claim, we note that the coe–cients u(m) 2nare non-negative, so U(m)(x) increases monotonically on the interval [0 ;1). Thus, for each K,w eh a v e KX n=0u(m) 2n•lim x"1U(m)(x)=u(m)•1X n=0u(m) 2n: By lettingK!1 , we see that u(m)=1X 2nu(m) 2n: This establishes the claim. From Equation 12.3, we see that if u(m)<1, then the probability of an eventual return is u(m)¡1 u(m); while ifu(m)=1, then the probability of eventual return is 1. To complete the example, we must estimate the sum 1X n=0u(m) 2n: 478 CHAPTER 12. RANDOM WALKS In Exercise 12, the reader is asked to show that u(2) 2n=1 42nµ2n n¶2 : Using Stirling’s Formula, it is easy to show that (see Exercise 13) µ2n n¶ »22n p…n; so u(2) 2n»1 …n: From this it follows easily that 1X n=0u(2) 2n diverges, so w(2) ⁄= 1, i.e., in R2, the probability of an eventual return is 1. Whenm= 3, Exercise 12 shows that u(3) 2n=1 22nµ2n n¶X j;kµ1 3nn! j!k!(n¡j¡k)!¶2 : LetMdenote the largest value of n! j!k!(n¡j¡k)!; over all non-negative values of jandkwithj+k•n. It is easy, using Stirling’s Formula, to show that M»c n; for some constant c. Thus, we have u(3) 2n•1 22nµ2n n¶X j;kµM 3nn! j!k!(n¡j¡k)!¶ : Using Exercise 14, one can show that the right-hand expression is at most c0 n3=2; wherec0is a constant. Thus, 1X n=0u(3) 2n converges, so w(3) ⁄is strictly less than one. This means that in R3, the probability of an eventual return to the origin is strictly less than one (in fact, it is approximately.65). One may summarize these results by stating that one should not get drunk in more than two dimensions. 2 12.1. RANDOM WALKS IN EUCLIDEAN SPACE 479 Expected Number of Equalizations We now give another example of the use of generating functions to flnd a general formula for terms in a sequence, where the sequence is related by recursion relationsto other sequences. Exercise 9 gives still another example. Example 12.3 (Expected Number of Equalizations) In this example, we will de- rive a formula for the expected number of equalizations in a random walk of length2m. As in the proof of Theorem 12.3, the method has four main parts. First, a recursion is found which relates the mth term in the unknown sequence to earlier terms in the same sequence and to terms in other (known) sequences. An exam-ple of such a recursion is given in Theorem 12.2. Second, the recursion is usedto derive a functional equation involving the generating functions of the unknownsequence and one or more known sequences. Equation 12.1 is an example of sucha functional equation. Third, the functional equation is solved for the unknowngenerating function. Last, using a device such as the Binomial Theorem, integra-tion, or difierentiation, a formula for the mth coe–cient of the unknown generating function is found. We begin by deflning g 2mto be the number of equalizations among all of the random walks of length 2 m. (For each random walk, we disregard the equalization at time 0.) We deflne g0= 0. Since the number of walks of length 2 mequals 22m, the expected number of equalizations among all such random walks is g2m=22m. Next, we deflne the generating function G(x): G(x)=1X k=0g2kxk: Now we need to flnd a recursion which relates the sequence fg2kgto one or both of the known sequences ff2kgandfu2kg. We consider mto be a flxed positive integer, and consider the set of all paths of length 2 mas the disjoint union E2[E4[¢¢¢[E2m[H; whereE2kis the set of all paths of length 2 mwith flrst equalization at time 2 k, andHis the set of all paths of length 2 mwith no equalization. It is easy to show (see Exercise 3) that jE2kj=f2k22m: We claim that the number of equalizations among all paths belonging to the set E2kis equal to jE2kj+22kf2kg2m¡2k: (12.4) Each path in E2khas one equalization at time 2 k, so the total number of such equalizations is just jE2kj. This is the flrst summand in expression Equation 12.4. There are 22kf2kdifierent initial segments of length 2 kamong the paths in E2k. Each of these initial segments can be augmented to a path of length 2 min 22m¡2k ways, by adjoining all possible paths of length 2 m¡2k. The number of equalizations obtained by adjoining all of these paths to any one initial segment is g2m¡2k,b y 480 CHAPTER 12. RANDOM WALKS deflnition. This gives the second summand in Equation 12.4. Since kcan range from 1 tom, we obtain the recursion g2m=mX k=1‡ jE2kj+22kf2kg2m¡2k· : (12.5) The second summand in the typical term above should remind the reader of a convolution. In fact, if we multiply the generating function G(x) by the generating function F(4x)=1X k=022kf2kxk; the coe–cient of xmequals mX k=022kf2kg2m¡2k: Thus, the product G(x)F(4x) is part of the functional equation that we are seeking. The flrst summand in the typical term in Equation 12.5 gives rise to the sum 22mmX k=1f2k: From Exercise 2, we see that this sum is just (1 ¡u2m)22m. Thus, we need to create a generating function whose mth coe–cient is this term; this generating function is 1X m=0(1¡u2m)22mxm; or1X m=022mxm+1X m=0u2mxm: The flrst sum is just (1 ¡4x)¡1, and the second sum is U(4x). So, the functional equation which we have been seeking is G(x)=F(4x)G(x)+1 1¡4x¡U(4x): If we solve this recursion for G(x), and simplify, we obtain G(x)=1 (1¡4x)3=2¡1 (1¡4x): (12.6) We now need to flnd a formula for the coe–cient of xm. The flrst summand in Equation 12.6 is (1 =2)U0(4x), so the coe–cient of xmin this function is u2m+222m+1(m+1 ): The second summand in Equation 12.6 is the sum of a geometric series with common ratio 4x, so the coe–cient of xmis 22m. Thus, we obtain 12.1. RANDOM WALKS IN EUCLIDEAN SPACE 481 g2m=u2m+222m+1(m+1 )¡22m =1 2µ2m+2 m+1¶ (m+1 )¡22m: We recall that the quotient g2m=22mis the expected number of equalizations among all paths of length 2 m. Using Exercise 4, it is easy to show that g2m 22m»r 2 …p 2m: In particular, this means that the average number of equalizations among all paths of length 4mis not twice the average number of equalizations among all paths of length 2m. In order for the average number of equalizations to double, one must quadruple the lengths of the random walks. 2 It is interesting to note that if we deflne Mn= max 0•k•nSk; then we have E(Mn)»r 2 …pn: This means that the expected number of equalizations and the expected maximum value for random walks of length nare asymptotically equal as n!1 . (In fact, it can be shown that the two expected values difier by at most 1 =2 for all positive integersn. See Exercise 9.) Exercises 1Using the Binomial Theorem, show that 1p1¡4x=1X m=0µ2m m¶ xm: What is the interval of convergence of this power series? 2(a) Show that for m‚1, f2m=u2m¡2¡u2m: (b) Using part (a), flnd a closed-form expression for the sum f2+f4+¢¢¢+f2m: (c) Using part (a), show that 1X m=1f2m=1: (One can also obtain this statement from the fact that F(x)=1¡(1¡x)1=2:) 482 CHAPTER 12. RANDOM WALKS (d) Using Exercise 2, show that the probability of no equalization in the flrst 2moutcomes equals the probability of an equalization at time 2 m. 3Using the notation of Example 12.3, show that jE2kj=f2k22m: 4Using Stirling’s Formula, show that u2m»1p…m: 5Alead change in a random walk occurs at time 2 kifS2k¡1andS2k+1are of opposite sign. (a) Give a rigorous argument which proves that among all walks of length 2mthat have an equalization at time 2 k, exactly half have a lead change at time 2k. (b) Deduce that the total number of lead changes among all walks of length 2mequals 1 2(g2m¡u2m): (c) Find an asymptotic expression for the average number of lead changes in a random walk of length 2 m. 6(a) Show that the probability that a random walk of length 2 mhas a last return to the origin at time 2 k, where 0•k•m, equals ¡2k k¢¡2m¡2k m¡k¢ 22m=u2ku2m¡2k: (The casek= 0 consists of all paths that do not return to the origin at any positive time.) Hint: A path whose last return to the origin occurs at time 2kconsists of two paths glued together, one path of which is of length 2kand which begins and ends at the origin, and the other path of which is of length 2 m¡2kand which begins at the origin but never returns to the origin. Both types of paths can be counted using quantitieswhich appear in this section. (b) Using part (a), show that the probability that a walk of length 2 mhas no equalization in the last moutcomes is equal to 1 =2, regardless of the value ofm.Hint: The answer to part a) is symmetric in kandm¡k. 7Show that the probability of no equalization in a walk of length 2 mequals u 2m. *8Show that P(S1‚0;S2‚0; :::; S 2m‚0) =u2m: 12.1. RANDOM WALKS IN EUCLIDEAN SPACE 483 Hint: First explain why P(S1>0;S2>0; :::; S 2m>0) =1 2P(S16=0;S26=0; :::; S 2m6=0 ): Then use Exercise 7, together with the observation that if no equalization occurs in the flrst 2 moutcomes, then the path goes through the point (1 ;1) and remains on or above the horizontal line x=1 . *9In Feller,3one flnds the following theorem: Let Mnbe the random variable which gives the maximum value of Sk, for 1•k•n. Deflne pn;r=µn n+r 2¶ 2¡n: Ifr‚0, then P(Mn=r)=‰pn;r;ifr·n(mod 2); 1; ifp‚q; P(Mn=r)=‰pn;r; ifr·n(mod 2); pn;r+1;ifr6·n(mod 2): (a) Using this theorem, show that E(M2m)=1 22mmX k=1(4k¡1)µ2m m+k¶ ; and ifn=2m+ 1, then E(M2m+1)=1 22m+1mX k=0(4k+1 )µ2m+1 m+k+1¶ : (b) Form‚1, deflne rm=mX k=1kµ2m m+k¶ and sm=mX k=1kµ2m+1 m+k+1¶ : By using the identity µn k¶ =µn¡1 k¡1¶ +µn¡1 k¶ ; show that sm=2rm¡1 2µ 22m¡µ2m m¶¶ 3W. Feller, Introduction to Probability Theory and its Applications, vol. I, 3rd ed. (New York: John Wiley & Sons, 1968). 484 CHAPTER 12. RANDOM WALKS and rm=2sm¡1+1 222m¡1; ifm‚2. (c) Deflne the generating functions R(x)=1X k=1rkxk and S(x)=1X k=1skxk: Show that S(x)=2R(x)¡1 2µ1 1¡4x¶ +1 2µp 1¡4x¶ and R(x)=2xS(x)+xµ1 1¡4x¶ : (d) Show that R(x)=x (1¡4x)3=2; and S(x)=1 2µ1 (1¡4x)3=2¶ ¡1 2µ1 1¡4x¶ : (e) Show that rm=mµ2m¡1 m¡1¶ ; and sm=1 2(m+1 )µ2m+1 m¶ ¡1 2(22m): (f) Show that E(M2m)=m 22m¡1µ2m m¶ +1 22m+1µ2m m¶ ¡1 2; and E(M2m+1)=m+1 22m+1µ2m+2 m+1¶ ¡1 2: The reader should compare these formulas with the expression for g2m=2(2m)in Example 12.3. 12.1. RANDOM WALKS IN EUCLIDEAN SPACE 485 *10 (from K. Levasseur4) A parent and his child play the following game. A deck of 2ncards,nred andnblack, is shu†ed. The cards are turned up one at a time. Before each card is turned up, the parent and the child guess whetherit will be red or black. Whoever makes more correct guesses wins the game.The child is assumed to guess each color with the same probability, so shewill have a score of n, on average. The parent keeps track of how many cards of each color have already been turned up. If more black cards, say, thanred cards remain in the deck, then the parent will guess black, while if anequal number of each color remain, then the parent guesses each color withprobability 1/2. What is the expected number of correct guesses that will bemade by the parent? Hint: Each of the¡ 2n n¢ possible orderings of red and black cards corresponds to a random walk of length 2 nthat returns to the origin at time 2 n. Show that between each pair of successive equalizations, the parent will be right exactly once more than he will be wrong. Explainwhy this means that the average number of correct guesses by the parent isgreater than nby exactly one-half the average number of equalizations. Now deflne the random variable X ito be 1 if there is an equalization at time 2 i, and 0 otherwise. Then, among all relevant paths, we have E(Xi)=P(Xi=1 )=¡2n¡2i n¡i¢¡2i i¢ ¡2n n¢: Thus, the expected number of equalizations equals EµnX i=1Xi¶ =1¡2n n¢nX i=1µ2n¡2i n¡i¶µ2i i¶ : One can now use generating functions to flnd the value of the sum. It should be noted that in a game such as this, a more interesting question than the one asked above is what is the probability that the parent wins thegame? For this game, this question was answered by D. Zagier. 5He showed that the probability of winning is asymptotic (for large n) to the quantity 1 2+1 2p 2: *11 Prove that u(2) 2n=1 42nnX k=0(2n)! k!k!(n¡k)!(n¡k)!; and u(3) 2n=1 62nX j;k(2n)! j!j!k!k!(n¡j¡k)!(n¡j¡k)!; 4K. Levasseur, \How to Beat Your Kids at Their Own Game," Mathematics Magazine vol. 61, no. 5 (December, 1988), pp. 301-305. 5D. Zagier, \How Often Should You Beat Your Kids?" Mathematics Magazine vol. 63, no. 2 (April 1990), pp. 89-92. 486 CHAPTER 12. RANDOM WALKS where the last sum extends over all non-negative jandkwithj+k•n. Also show that this last expression may be rewritten as 1 22nµ2n n¶X j;kµ1 3nn! j!k!(n¡j¡k)!¶2 : *12 Prove that if n‚0, then nX k=0µn k¶2 =µ2n n¶ : Hint: Write the sum asnX k=0µn k¶µn n¡k¶ and explain why this is a coe–cient in the product (1 +x)n(1 +x)n: Use this, together with Exercise 11, to show that u(2) 2n=1 42nµ2n n¶nX k=0µn k¶2 =1 42nµ2n n¶2 : *13 Using Stirling’s Formula, prove that µ2n n¶ »22n p…n: *14 Prove thatX j;kµ1 3nn! j!k!(n¡j¡k)!¶ =1; where the sum extends over all non-negative jandksuch thatj+k•n. Hint: Count how many ways one can place nlabelled balls in 3 labelled urns. *15 Using the result proved for the random walk in R3in Example 12.2, explain why the probability of an eventual return in Rnis strictly less than one, for alln‚3.Hint: Consider a random walk in Rnand disregard all but the flrst three coordinates of the particle’s position. 12.2 Gambler’s Ruin In the last section, the simplest kind of symmetric random walk in R1was studied. In this section, we remove the assumption that the random walk is symmetric.Instead, we assume that pandqare non-negative real numbers with p+q= 1, and that the common distribution function of the jumps of the random walk is f X(x)=‰p;ifx=1; q;ifx=¡1: 12.2. GAMBLER’S RUIN 487 One can imagine the random walk as representing a sequence of tosses of a weighted coin, with a head appearing with probability pand a tail appearing with probability q. An alternative formulation of this situation is that of a gambler playing a sequence of games against an adversary (sometimes thought of as another person, sometimescalled \the house") where, in each game, the gambler has probability pof winning. The Gambler’s Ruin Problem The above formulation of this type of random walk leads to a problem known as the Gambler’s Ruin problem. This problem was introduced in Exercise 23, but we willgive the description of the problem again. A gambler starts with a \stake" of size s. She plays until her capital reaches the value Mor the value 0. In the language of Markov chains, these two values correspond to absorbing states. We are interestedin studying the probability of occurrence of each of these two outcomes. One can also assume that the gambler is playing against an \inflnitely rich" adversary. In this case, we would say that there is only one absorbing state, namelywhen the gambler’s stake is 0. Under this assumption, one can ask for the proba-bility that the gambler is eventually ruined. We begin by deflning q kto be the probability that the gambler’s stake reaches 0, i.e., she is ruined, before it reaches M, given that the initial stake is k. We note that q0= 1 andqM= 0. The fundamental relationship among the qk’s is the following: qk=pqk+1+qqk¡1; where 1•k•M¡1. This holds because if her stake equals k, and she plays one game, then her stake becomes k+ 1 with probability pandk¡1 with probability q. In the flrst case, the probability of eventual ruin is qk+1and in the second case, it isqk¡1. We note that since p+q= 1, we can write the above equation as p(qk+1¡qk)=q(qk¡qk¡1); or qk+1¡qk=q p(qk¡qk¡1): From this equation, it is easy to see that qk+1¡qk=µq p¶k (q1¡q0): (12.7) We now use telescoping sums to obtain an equation in which the only unknown is q1: ¡1=qM¡q0 =M¡1X k=0(qk+1¡qk); 488 CHAPTER 12. RANDOM WALKS so ¡1=M¡1X k=0µq p¶k (q1¡q0) =(q1¡q0)M¡1X k=0µq p¶k : Ifp6=q, then the above expression equals (q1¡q0)(q=p)M¡1 (q=p)¡1; while ifp=q=1=2, then we obtain the equation ¡1=(q1¡q0)M: For the moment we shall assume that p6=q. Then we have q1¡q0=¡(q=p)¡1 (q=p)M¡1: Now, for any zwith 1•z•M,w eh a v e qz¡q0=z¡1X k=0(qk+1¡qk) =(q1¡q0)z¡1X k=0µq p¶k =¡(q1¡q0)(q=p)z¡1 (q=p)¡1 =¡(q=p)z¡1 (q=p)M¡1: Therefore, qz=1¡(q=p)z¡1 (q=p)M¡1 =(q=p)M¡(q=p)z (q=p)M¡1: Finally, ifp=q=1=2, it is easy to show that (see Exercise 10) qz=M¡z M: We note that both of these formulas hold if z=0 . We deflne, for 0 •z•M, the quantity pzto be the probability that the gambler’s stake reaches Mwithout ever having reached 0. Since the game might 12.2. GAMBLER’S RUIN 489 continue indeflnitely, it is not obvious that pz+qz= 1 for allz. However, one can use the same method as above to show that if p6=q, then qz=(q=p)z¡1 (q=p)M¡1; and ifp=q=1=2, then qz=z M: Thus, for all z, it is the case that pz+qz= 1, so the game ends with probability 1. Inflnitely Rich Adversaries We now turn to the problem of flnding the probability of eventual ruin if the gambler is playing against an inflnitely rich adversary. This probability can be obtained bylettingMgo to1in the expression for q zcalculated above. If q<p , then the expression approaches ( q=p)z, and ifq>p , the expression approaches 1. In the casep=q=1=2, we recall that qz=1¡z=M. Thus, ifM!1 , we see that the probability of eventual ruin tends to 1. Historical Remarks In 1711, De Moivre, in his book De Mesura Sortis , gave an ingenious derivation of the probability of ruin. The following description of his argument is taken fromDavid. 6The notation used is as follows: We imagine that there are two players, A and B, and the probabilities that they win a game are pandq, respectively. The players start with aandbcounters, respectively. Imagine that each player starts with his counters before him in a pile, and that nominal values are assigned to the counters in the followingmanner. A’s bottom counter is given the nominal value q=p; the next is given the nominal value ( q=p) 2, and so on until his top counter which has the nominal value ( q=p)a. B’s top counter is valued ( q=p)a+1, and so on downwards until his bottom counter which is valued ( q=p)a+b. After each game the loser’s top counter is transferred to the top of thewinner’s pile, and it is always the top counter which is staked for thenext game. Then in terms of the nominal values B’s stake is always q=ptimes A’s, so that at every game each player’s nominal expectation is nil. This remains true throughout the play; therefore A’s chance ofwinning all B’s counters, multiplied by his nominal gain if he does so,must equal B’s chance multiplied by B’s nominal gain. Thus, P aµ‡q p·a+1 +¢¢¢+‡q p·a+b¶ =Pbµ‡q p· +¢¢¢+‡q p·a¶ :(12.8) 6F. N. David, Games, Gods and Gambling (London: Gri–n, 1962). 490 CHAPTER 12. RANDOM WALKS Using this equation, together with the fact that Pa+Pb=1; it can easily be shown that Pa=(q=p)a¡1 (q=p)a+b¡1; ifp6=q, and Pa=a a+b; ifp=q=1=2. In terms of modern probability theory, de Moivre is changing the values of the counters to make an unfair game into a fair game, which is called a martingale.With the new values, the expected fortune of player A (that is, the sum of thenominal values of his counters) after each play equals his fortune before the play(and similarly for player B). (For a simpler martingale argument, see Exercise 9.) DeMoivre then uses the fact that when the game ends, it is still fair, thus Equation 12.8must be true. This fact requires proof, and is one of the central theorems in thearea of martingale theory. Exercises 1In the gambler’s ruin problem, assume that the gambler initial stake is 1 dollar, and assume that her probability of success on any one game is p. Let Tbe the number of games until 0 is reached (the gambler is ruined). Show that the generating function for Tis h(z)=1¡p 1¡4pqz2 2pz; and that h(1) =‰q=p; ifq•p; 1; ifq‚p; and h0(1) =‰ 1=(q¡p);ifq>p; 1; ifq=p: Interpret your results in terms of the time Tto reach 0. (See also Exam- ple 10.7.) 2Show that the Taylor series expansion forp1¡xis p 1¡x=1X n=0µ1=2 n¶ xn; where the binomial coe–cient¡1=2 n¢ is µ1=2 n¶ =(1=2)(1=2¡1)¢¢¢(1=2¡n+1 ) n!: 12.2. GAMBLER’S RUIN 491 Using this and the result of Exercise 1, show that the probability that the gambler is ruined on the nth step is pT(n)=( (¡1)k¡1 2p¡1=2 k¢ (4pq)k;ifn=2k¡1, 0; ifn=2k. 3For the gambler’s ruin problem, assume that the gambler starts with kdollars. LetTkbe the time to reach 0 for the flrst time. (a) Show that the generating function hk(t) forTkis thekth power of the generating function for the time Tto ruin starting at 1. Hint: Let Tk=U1+U2+¢¢¢+Uk, whereUjis the time for the walk starting at j to reachj¡1 for the flrst time. (b) Findhk(1) andh0 k(1) and interpret your results. 4(The next three problems come from Feller.7) As in the text, assume that M is a flxed positive integer. (a) Show that if a gambler starts with an stake of 0 (and is allowed to have a negative amount of money), then the probability that her stake reachesthe value of Mbefore it returns to 0 equals p(1¡q 1). (b) Show that if the gambler starts with a stake of Mthen the probability that her stake reaches 0 before it returns to MequalsqqM¡1. 5Suppose that a gambler starts with a stake of 0 dollars. (a) Show that the probability that her stake never reaches Mbefore return- ing to 0 equals 1¡p(1¡q1). (b) Show that the probability that her stake reaches the value Mexactly ktimes before returning to 0 equals p(1¡q1)(1¡qqM¡1)k¡1(qqM¡1). Hint: Use Exercise 4. 6In the text, it was shown that if q<p , there is a positive probability that a gambler, starting with a stake of 0 dollars, will never return to the origin.Thus, we will now assume that q‚p. Using Exercise 5, show that if a gambler starts with a stake of 0 dollars, then the expected number of timesher stake equals Mbefore returning to 0 equals ( p=q) M,i fq>p and 1, if q=p. (We quote from Feller: \The truly amazing implications of this result appear best in the language of fair games. A perfect coin is tossed untilthe flrst equalization of the accumulated numbers of heads and tails. Thegambler receives one penny for every time that the accumulated number ofheads exceeds the accumulated number of tails by m.The ‘fair entrance fee’ equals 1 independent of m.") 7W. Feller, op. cit., pg. 367. 492 CHAPTER 12. RANDOM WALKS 7In the game in Exercise 6, let p=q=1=2 andM= 10. What is the probability that the gambler’s stake equals Mat least 20 times before it returns to 0? 8Write a computer program which simulates the game in Exercise 6 for the casep=q=1=2, andM= 10. 9In de Moivre’s description of the game, we can modify the deflnition of player A’s fortune in such a way that the game is still a martingale (and the calcula-tions are simpler). We do this by assigning nominal values to the counters inthe same way as de Moivre, but each player’s current fortune is deflned to bejust the value of the counter which is being wagered on the next game. So, ifplayer A has acounters, then his current fortune is ( q=p) a(we stipulate this to be true even if a= 0). Show that under this deflnition, player A’s expected fortune after one play equals his fortune before the play, if p6=q. Then, as de Moivre does, write an equation which expresses the fact that player A’sexpected flnal fortune equals his initial fortune. Use this equation to flnd theprobability of ruin of player A. 10Assume in the gambler’s ruin problem that p=q=1=2. (a) Using Equation 12.7, together with the facts that q 0= 1 andqM=0 , show that for 0•z•M, qz=M¡z M: (b) In Equation 12.8, let p!1=2 (and since q=1¡p,q!1=2 as well). Show that in the limit, qz=M¡z M: Hint: Replaceqby 1¡p, and use L’Hopital’s rule. 11In American casinos, the roulette wheels have the integers between 1 and 36, together with 0 and 00. Half of the non-zero numbers are red, the other halfare black, and 0 and 00 are green. A common bet in this game is to bet adollar on red. If a red number comes up, the bettor gets her dollar back, andalso gets another dollar. If a black or green number comes up, she loses herdollar. (a) Suppose that someone starts with 40 dollars, and continues to bet on red until either her fortune reaches 50 or 0. Find the probability that herfortune reaches 50 dollars. (b) How much money would she have to start with, in order for her to have a 95% chance of winning 10 dollars before going broke? (c) A casino owner was once heard to remark that \If we took 0 and 00 ofi of the roulette wheel, we would still make lots of money, because peoplewould continue to come in and play until they lost all of their money."Do you think that such a casino would stay in business? 12.3. ARC SINE LAWS 493 12.3 Arc Sine Laws In Exercise 12.1.6, the distribution of the time of the last equalization in the sym- metric random walk was determined. If we let fi2k;2mdenote the probability that a random walk of length 2 mhas its last equalization at time 2 k, then we have fi2k;2m=u2ku2m¡2k: We shall now show how one can approximate the distribution of the fi’s with a simple function. We recall that u2k»1p …k: Therefore, as both kandmgo to1,w eh a v e fi2k;2m»1 …p k(m¡k): This last expression can be written as 1 …mp (k=m)(1¡k=m): Thus, if we deflne f(x)=1 …p x(1¡x); for 0<x< 1, then we have fi2k;2m…1 mfµk m¶ : The reason for the …sign is that we no longer require that kget large. This means that we can replace the discrete fi2k;2mdistribution by the continuous density f(x) on the interval [0 ;1] and obtain a good approximation. In particular, if xis a flxed real number between 0 and 1, then we have X k<xmfi2k;2m…Zx 0f(t)dt : It turns out that f(x) has a nice antiderivative, so we can write X k<xmfi2k;2m…2 …arcsinpx: One can see from the graph of this last function that it has a minimum at x=1=2 and is symmetric about that point. As noted in the exercise, this implies that halfof the walks of length 2 mhave no equalizations after time m, a fact which probably would not be guessed. It turns out that the arc sine density comes up in the answers to many other questions concerning random walks on the line. Recall that in Section 12.1, a 494 CHAPTER 12. RANDOM WALKS random walk could be viewed as a polygonal line connecting (0 ;0) with (m;Sm). Under this interpretation, we deflne b2k;2mto be the probability that a random walk of length 2mhas exactly 2 kof its 2mpolygonal line segments above the t-axis. The probability b2k;2mis frequently interpreted in terms of a two-player game. (The reader will recall the game Heads or Tails, in Example 1.4.) Player A is saidto be in the lead at time nif the random walk is above the t-axis at that time, or if the random walk is on the t-axis at time nbut above the t-axis at time n¡1. (At time 0, neither player is in the lead.) One can ask what is the most probablenumber of times that player A is in the lead, in a game of length 2 m. Most people will say that the answer to this question is m. However, the following theorem says thatmis the least likely number of times that player A is in the lead, and the most likely number of times in the lead is 0 or 2 m. Theorem 12.4 If Peter and Paul play a game of Heads or Tails of length 2 m, the probability that Peter will be in the lead exactly 2 ktimes is equal to fi 2k;2m: Proof. To prove the theorem, we need to show that b2k;2m=fi2k;2m: (12.9) Exercise 12.1.7 shows that b2m;2m=u2mandb0;2m=u2m, so we only need to prove that Equation 12.9 holds for 1 •k•m¡1. We can obtain a recursion involving the b’s and thef’s (deflned in Section 12.1) by counting the number of paths of length 2mthat have exactly 2 kof their segments above the t-axis, where 1•k•m¡1. To count this collection of paths, we assume that the flrst return occurs at time 2 j, where 1•j•m¡1. There are two cases to consider. Either during the flrst 2 j outcomes the path is above the t-axis or below the t-axis. In the flrst case, it must be true that the path has exactly (2 k¡2j) line segments above the t-axis, between t=2jandt=2m. In the second case, it must be true that the path has exactly 2kline segments above the t-axis, between t=2jandt=2m. We now count the number of paths of the various types described above. The number of paths of length 2 jall of whose line segments lie above the t-axis and which return to the origin for the flrst time at time 2 jequals (1=2)22jf2j. This also equals the number of paths of length 2 jall of whose line segments lie below thet-axis and which return to the origin for the flrst time at time 2 j. The number of paths of length (2 m¡2j) which have exactly (2 k¡2j) line segments above the t-axis isb2k¡2j;2m¡2j. Finally, the number of paths of length (2 m¡2j) which have exactly 2kline segments above the t-axis isb2k;2m¡2j. Therefore, we have b2k;2m=1 2kX j=1f2jb2k¡2j;2m¡2j+1 2m¡kX j=1f2jb2k;2m¡2j: We now assume that Equation 12.9 is true for m<n . Then we have 12.3. ARC SINE LAWS 495 0 10 20 30 4000.020.040.060.080.10.12 Figure 12.2: Times in the lead. b2k;2n=1 2kX j=1f2jfi2k¡2j;2m¡2j+1 2m¡kX j=1f2jfi2k;2m¡2j =1 2kX j=1f2ju2k¡2ju2m¡2k+1 2m¡kX j=1f2ju2ku2m¡2j¡2k =1 2u2m¡2kkX j=1f2ju2k¡2j+1 2u2km¡kX j=1f2ju2m¡2j¡2k =1 2u2m¡2ku2k+1 2u2ku2m¡2k; where the last equality follows from Theorem 12.2. Thus, we have b2k;2n=fi2k;2n; which completes the proof. 2 We illustrate the above theorem by simulating 10,000 games of Heads or Tails, with each game consisting of 40 tosses. The distribution of the number of times thatPeter is in the lead is given in Figure 12.2, together with the arc sine density. We end this section by stating two other results in which the arc sine density appears. Proofs of these results may be found in Feller. 8 Theorem 12.5 LetJbe the random variable which, for a given random walk of length 2m, gives the smallest subscript jsuch thatSj=S2m. (Such a subscript j must be even, by parity considerations.) Let °2k;2mbe the probability that J=2k. Then we have °2k;2m=fi2k;2m: 2 8W. Feller, op. cit., pp. 93{94. 496 CHAPTER 12. RANDOM WALKS The next theorem says that the arc sine density is applicable to a wide range of situations. A continuous distribution function F(x) is said to be symmetric ifF(x)=1¡F(¡x). (IfXis a continuous random variable with a symmetric distribution function, then for any real x, we haveP(X•x)=P(X‚¡x).) We imagine that we have a random walk of length nin which each summand has the distribution F(x), whereFis continuous and symmetric. The subscript of the flrst maximum of such a walk is the unique subscript ksuch that Sk>S0; :::; Sk>Sk¡1;Sk‚Sk+1; :::; Sk‚Sn: We deflne the random variable Knto be the subscript of the flrst maximum. We can now state the following theorem concerning the random variable Kn. Theorem 12.6 LetFbe a symmetric continuous distribution function, and let fi be a flxed real number strictly between 0 and 1. Then as n!1 ,w eh a v e P(Kn<nfi )!2 …arcsinpfi: 2 A version of this theorem that holds for a symmetric random walk can also be found in Feller. Exercises 1For a random walk of length 2 m, deflne†kto equal 1 if Sk>0, or ifSk¡1=1 andSk= 0. Deflne †kto equal -1 in all other cases. Thus, †kgives the side of thet-axis that the random walk is on during the time interval [ k¡1;k]. A \law of large numbers" for the sequence f†kgwould say that for any –>0, we would have Pµ ¡–<†1+†2+¢¢¢+†n n<–¶ !1 asn!1 . Even though the †’s are not independent, the above assertion certainly appears reasonable. Using Theorem 12.4, show that if ¡1•x•1, then lim n!1Pµ†1+†2+¢¢¢+†n n<x¶ =2 …arcsinr 1+x 2: 2Given a random walk Wof lengthm, with summands fX1;X2;:::;Xmg; deflne the reversed random walk to be the walk W⁄with summands fXm;Xm¡1;:::;X 1g: (a) Show that the kth partial sum S⁄ ksatisfles the equation S⁄ k=Sm¡Sn¡k; whereSkis thekth partial sum for the random walk W. 12.3. ARC SINE LAWS 497 (b) Explain the geometric relationship between the graphs of a random walk and its reversal. (It is not in general true that one graph is obtainedfrom the other by re°ecting in a vertical line.) (c) Use parts (a) and (b) to prove Theorem 12.5. 499 NA (0,d) = area of shaded region 0d .00 .01 .02 .03 .04 .05 .06 .07 .08 .09 0.0 .0000 .0040 .0080 .0120 .0160 .0199 .0239 .0279 .0319 .0359 0.1 .0398 .0438 .0478 .0517 .0557 .0596 .0636 .0675 .0714 .0753 0.2 .0793 .0832 .0871 .0910 .0948 .0987 .1026 .1064 .1103 .11410.3 .1179 .1217 .1255 .1293 .1331 .1368 .1406 .1443 .1480 .1517 0.4 .1554 .1591 .1628 .1664 .1700 .1736 .1772 .1808 .1844 .1879 0.5 .1915 .1950 .1985 .2019 .2054 .2088 .2123 .2157 .2190 .2224 0.6 .2257 .2291 .2324 .2357 .2389 .2422 .2454 .2486 .2517 .2549 0.7 .2580 .2611 .2642 .2673 .2704 .2734 .2764 .2794 .2823 .2852 0.8 .2881 .2910 .2939 .2967 .2995 .3023 .3051 .3078 .3106 .3133 0.9 .3159 .3186 .3212 .3238 .3264 .3289 .3315 .3340 .3365 .3389 1.0 .3413 .3438 .3461 .3485 .3508 .3531 .3554 .3577 .3599 .3621 1.1 .3643 .3665 .3686 .3708 .3729 .3749 .3770 .3790 .3810 .3830 1.2 .3849 .3869 .3888 .3907 .3925 .3944 .3962 .3980 .3997 .4015 1.3 .4032 .4049 .4066 .4082 .4099 .4115 .4131 .4147 .4162 .4177 1.4 .4192 .4207 .4222 .4236 .4251 .4265 .4279 .4292 .4306 .4319 1.5 .4332 .4345 .4357 .4370 .4382 .4394 .4406 .4418 .4429 .44411.6 .4452 .4463 .4474 .4484 .4495 .4505 .4515 .4525 .4535 .45451.7 .4554 .4564 .4573 .4582 .4591 .4599 .4608 .4616 .4625 .46331.8 .4641 .4649 .4656 .4664 .4671 .4678 .4686 .4693 .4699 .4706 1.9 .4713 .4719 .4726 .4732 .4738 .4744 .4750 .4756 .4761 .47672.0 .4772 .4778 .4783 .4788 .4793 .4798 .4803 .4808 .4812 .48172.1 .4821 .4826 .4830 .4834 .4838 .4842 .4846 .4850 .4854 .4857 2.2 .4861 .4864 .4868 .4871 .4875 .4878 .4881 .4884 .4887 .4890 2.3 .4893 .4896 .4898 .4901 .4904 .4906 .4909 .4911 .4913 .49162.4 .4918 .4920 .4922 .4925 .4927 .4929 .4931 .4932 .4934 .49362.5 .4938 .4940 .4941 .4943 .4945 .4946 .4948 .4949 .4951 .49522.6 .4953 .4955 .4956 .4957 .4959 .4960 .4961 .4962 .4963 .4964 2.7 .4965 .4966 .4967 .4968 .4969 .4970 .4971 .4972 .4973 .4974 2.8 .4974 .4975 .4976 .4977 .4977 .4978 .4979 .4979 .4980 .49812.9 .4981 .4982 .4982 .4983 .4984 .4984 .4985 .4985 .4986 .49863.0 .4987 .4987 .4987 .4988 .4988 .4989 .4989 .4989 .4990 .49903.1 .4990 .4991 .4991 .4991 .4992 .4992 .4992 .4992 .4993 .49933.2 .4993 .4993 .4994 .4994 .4994 .4994 .4994 .4995 .4995 .4995 3.3 .4995 .4995 .4995 .4996 .4996 .4996 .4996 .4996 .4996 .49973.4 .4997 .4997 .4997 .4997 .4997 .4997 .4997 .4997 .4997 .49983.5 .4998 .4998 .4998 .4998 .4998 .4998 .4998 .4998 .4998 .49983.6 .4998 .4998 .4999 .4999 .4999 .4999 .4999 .4999 .4999 .49993.7 .4999 .4999 .4999 .4999 .4999 .4999 .4999 .4999 .4999 .4999 3.8 .4999 .4999 .4999 .4999 .4999 .4999 .4999 .4999 .4999 .49993.9 .5000 .5000 .5000 .5000 .5000 .5000 .5000 .5000 .5000 .5000 Appendix A Normal distribution table Above . . . . . . . . . . . . . . . . . . . . . . . . 1 3 . . 4 5 . . 72.5 . . . . . . . . . . . . . . 1 2 1 2 7 2 4 19 6 72.2 71.5 . . . . . . . . 1 3 4 3 5 10 4 9 2 2 43 11 69.9 70.5 1 . . 1 . . 1 1 3 12 18 14 7 4 3 3 68 22 69.5 69.5 . . . . 1 16 4 17 27 20 33 25 20 11 4 5 183 41 68.9 68.5 1 . . 7 11 16 25 31 34 48 21 18 4 3 . . 219 49 68.2 67.5 . . 3 5 14 15 36 38 28 38 19 11 4 . . . . 211 33 67.6 66.5 . . 3 3 5 2 17 17 14 13 4 . . . . . . . . 78 20 67.2 65.5 1 . . 9 5 7 11 11 7 7 5 2 1 . . . . 66 12 66.7 64.5 1 1 4 4 1 5 5 . . 2 . . . . . . . . . . 23 5 65.8 Below . . 1 . . 2 4 1 2 2 1 1 . . . . . . . . . . 14 1 . . Totals . . 5 7 32 59 48 117 138 120 167 99 64 41 17 14 928 205 . .Medians . . . . . . 66.3 67.8 67.9 67.7 67.9 68.3 68.5 69.0 69.0 70.0 . . . . . . . . . . Heights of the Mid-parentsin inches. Below 62.2 63.2 64.2 65.2 66.2 67.2 68.2 69.2 70.2 71.2 72.2 73.2 Above chil dren. Heights of the adult children. Total number ofAdult Mid-parents. MediansNumber of adult children of various statures born of 205 mid-parents of various statures.(All female heights have been multiplied by 1.08)Note. In calculating the Medians, the entries have been taken as referring to the middle of the squares in which they stand. The reason why the headings run 62.2, 63.2, &c., instead of 62.5, 63.5, &c., is that the observations are unequally distributed between 62 and 63, 63 and 64, &c., there being a strong bias in favour of integral inches. After careful consideration, I concluded that the headin gs, as adopted, best satisfied the conditions. This inequality was not apparent in the case of the Mid-parents. Source: F. Galton, "Regressio n towards Mediocrity in Hereditary Stature", Royal Anthropological Institute of Great Britain and Ireland, vol.15 (1885), p.248.Appendix B Appendix C Life Table Number of survivors at single years of Age, out of 100,000 Born Alive, by Race and Sex: United States, 1990 . 0 100000 100000 100000 43 94707 92840 96626 1 99073 98969 99183 44 94453 92505 96455 2 99008 98894 99128 45 94179 92147 96266 3 98959 98840 99085 46 93882 91764 96057 4 98921 98799 99051 47 93560 91352 95827 5 98890 98765 99023 48 93211 90908 95573 6 98863 98735 99000 49 92832 90429 95294 7 98839 98707 98980 50 92420 89912 94987 8 98817 98680 98962 51 91971 89352 94650 9 98797 98657 98946 52 91483 88745 94281 10 98780 98638 98931 53 90950 88084 93877 11 98765 98623 98917 54 90369 87363 93436 12 98750 98608 98902 55 89735 86576 92955 13 98730 98586 98884 56 89045 85719 92432 14 98699 98547 98862 57 88296 84788 91864 15 98653 98485 98833 58 87482 83777 91246 16 98590 98397 98797 59 86596 82678 90571 17 98512 98285 98753 60 85634 81485 89835 18 98421 98154 98704 61 84590 80194 89033 19 98323 98011 98654 62 83462 78803 88162 20 98223 97863 98604 63 82252 77314 87223 21 98120 97710 98555 64 80961 75729 86216 22 98015 97551 98506 65 79590 74051 85141 23 97907 97388 98456 66 78139 72280 83995 24 97797 97221 98405 67 76603 70414 82772 25 97684 97052 98351 68 74975 68445 81465 26 97569 96881 98294 69 73244 66364 80064 27 97452 96707 98235 70 71404 64164 78562 28 97332 96530 98173 71 69453 61847 76953 29 97207 96348 98107 72 67392 59419 75234 30 97077 96159 98038 73 65221 56885 73400 31 96941 95962 97965 74 62942 54249 71499 32 96800 95785 97887 75 60557 51519 69376 33 96652 95545 97804 76 58069 48704 67178 34 96497 95322 97717 77 55482 45816 64851 35 96334 95089 97624 78 52799 42867 62391 36 96161 94843 97525 79 50026 39872 59796 37 95978 94585 97419 80 47168 36848 57062 38 95787 94316 97306 81 44232 33811 54186 39 95588 94038 97187 82 41227 30782 51167 40 95382 93753 97061 83 38161 27782 48002 41 95168 93460 96926 84 35046 24834 44690 42 94944 93157 96782 85 31892 21962 41230 Age Both sexes Male Female Age Both sexes Male FemaleAll races All races Index n!, 80 …, estimation of, 43{46 absorbing Markov chain, 415 absorbing state, 416AbsorbingChain (program), 421absorption probabilities, 420Ace, Mr., 241Ali, 178alleles, 348AllPermutations (program), 84ANDERSON, C. L., 157annuity, 246 life, 247terminal, 247 arc sine laws, 493area, estimation of, 42Areabargraph (program), 46asymptotically equal, 81 Baba, 178 babies, 14, 250Banach’s Matchbox, 255BAR-HILLEL, M., 176BARNES, B., 175BARNHART, R., 11BAYER, D., 120Bayes (program), 147Bayes probability, 136Bayes’ formula, 146BAYES, T., 149beard, 153bell-shaped, 47Benford distribution, 195BENKOSKI, S., 40Bernoulli trials process, 96BERNOULLI, D., 227BERNOULLI, J., 113, 149, 310{312 Bertrand’s paradox, 47{50BERTRAND, J., 49, 181BertrandsParadox (program), 49beta density, 168BIENAYM ¶E, I., 310, 378 BIGGS, N. L., 85binary expansion, 69binomial coe–cient, 93binomial distribution, 99, 184 approximating a, 329 Binomial Theorem, 103BinomialPlot (program), 99BinomialProbabilities (program), 98Birthday (program), 78birthday problem, 77blackjack, 247, 253blood test, 254Bose-Einstein statistics, 107Box paradox, 181BOX, G. E. P., 213boxcars, 27BRAMS, S., 179, 182Branch (program), 382branching process, 377 customer, 394 BranchingSimulation (program), 388bridge, 181, 182, 199, 203, 287BROWN, B. H., 38BROWN, E., 425Bufion’s needle, 44{46, 51{53BUFFON, G. L., 9, 44, 50{51BufionsNeedle (program), 45bus paradox, 164 calendar, 38 cancer, 147 503 504 INDEX canonical form of an absorbing Markov chain, 417 car, 137CARDANO, G., 30{31, 110, 249cars on a hig hway, 66 CASANOVA, G., 11Cauchy density, 218, 401cells, 347Central Limit Theorem, 325 for Bernoulli Trials, 330for Binomial Distributions, 328for continuous independent trials process, 357 for discrete independent random variables, 345 for discrete independent trials process, 343 for Markov Chains, 464proof of, 398 chain letter, 389characteristic function, 398Chebyshev Inequality, 305, 316CHEBYSHEV, P. L., 313chi-squared density, 216, 296Chicago World’s Fair, 52chord, random, 47, 54chromosomes, 348CHU, S.-C., 110CHUNG, K. L., 153Circle of Gold, 389Clinton, Bill, 196clover-leaf interchange, 39CLTBernoulliGlobal, 332CLTBernoulliLocal (program), 329CLTBernoulliPlot (program), 327CLTGeneral (program), 345CLTIndTrialsLocal (program), 342CLTIndTrialsPlot (program), 341COATES, R. M., 305CoinTosses (program), 3Collins, People v., 153, 202color-blindness, 424conditional density, 162conditional distribution, 134conditional expectation, 239conditional probability, 133 CONDORCET, Le Marquis de, 12confldence interval, 334, 360conjunction fallacy, 38continuum, 41convolution, 286, 291 of binomial distributions, 289of Cauchy densities, 294of exponential densities, 292, 300of geometric distributions, 289of normal densities, 294of standard normal densities, 299of uniform densities, 292, 299 CONWAY, J., 432CRAMER, G., 227craps, 235, 240, 468Craps (program), 235CROSSEN, C., 161CROWELL, R., 468cumulative distribution function, 61 joint, 165 customer branching process, 394cut, 120 Dartmouth, 27 darts, 56, 57, 59, 60, 64, 71, 163, 164Darts (program), 58DAVID, F. N., 86, 337, 489DAVID, F. N., 32de MOIVRE, A., 37, 88, 148, 336, 489de MONTMORT, P. R., 85de M ¶ER¶E, CHEVALIER, 4, 31, 37 degrees of freedom, 217DeMere1 (program), 4DeMere2 (program), 4density function, 56, 59 beta, 168Cauchy, 218, 401chi-squared, 216, 296conditional, 162exponential, 53, 66, 163, 205gamma, 207joint, 165log normal, 224Maxwell, 215 INDEX 505 normal, 212 Rayleigh, 215, 295t-, 360uniform, 60, 205 derangement, 85DIACONIS, P., 120, 251Die (program), 225DieTest (program), 297distribution function, 1, 19 properties of, 22Benford, 195binomial, 184geometric, 184hypergeometric, 193joint, 142marginal, 143negative binomial, 186Poisson, 187uniform, 183 DNA, 348DOEBLIN, W., 449DOYLE, P. G., 87, 470Drunkard’s Walk example, 416, 419{ 421, 423, 427, 443 Dry Gulch, 279 EDWARDS, A. W. F., 108 Egypt, 30Ehrenfest model, 410, 433, 441, 460, 461 EHRENFEST, P., 410EHRENFEST, T, 410EhrenfestUrn (program), 462EISENBERG, B., 160elevator, 89, 116Emile’s restaurant, 75ENGLE, A., 445envelopes, 179, 180EPSTEIN, R., 287equalization, 472equalizations expected number of, 479 ergodic Markov chain, 433ESP, 250, 251EUCLID, 85Euler’s formula, 202 Eulerian number, 127event, 18events attraction of, 160independent, 139, 164repulsion of, 160 existence of God, 245expected value, 226, 268exponential density, 53, 66, 163, 205extinction, problem of, 379 factorial, 80 fair game, 241FALK, R., 161, 176fall, 131fallacy, 38FELLER, W., 11, 106, 107, 191, 201, 218, 254, 344 FERMAT, P., 4, 32{35, 112{113, 156Fermi-Dirac statistics, 107flgurate numbers, 108flnancial records suspicious, 196 flnite additivity property, 23FINN, J., 178First Fundamental Mystery of Proba- bility, 232 flrst maximum of a random walk, 496flrst return to the origin, 473Fisher’s Exact Test, 193FISHER, R. A., 252flxed column vector, 435flxed points, 82flxed row vector, 435FixedPoints (program), 82FixedVector (program), 437°ying bombs, 191, 201Fourier transform, 398FRECHET, M., 466frequency concept of probability, 70frustration solitaire, 86Fundamental Limit Theorem for Reg- ular Markov Chains, 448 fundamental matrix, 419 506 INDEX for a regular Markov chain, 457 for an ergodic Markov chain, 458 GALAMBOS, J., 303 GALILEO, G., 12Gallup Poll, 14, 336Galton board, 99, 351GALTON, F., 281, 345, 350, 377GaltonBoard (program), 99Gambler’s Ruin, 426, 486, 487gambling systems, 241gamma density, 207GARDNER, M., 181gas difiusion Ehrenfest model of, 410, 433, 441, 460, 461 GELLER, S., 176GeneralSimulation (program), 9generating function for continuous density, 394moment, 366, 395ordinary, 370 genes, 348, 411genetics, 345genotypes, 348geometric distribution, 184geometric series, 29GHOSH, B. K., 160goat, 137GONSHOR, H., 425GOSSET, W. S., 360grade point average, 343GRAHAM, R., 251GRANBERG, D., 161GRAUNT, J., 246Greece, 30GRIDGEMAN, N. T., 51, 181GRINSTEAD, C. M., 87GUDDER, S., 160 HACKING, I., 30, 148 HAMMING, R. W., 283, 284HANES data, 345Hangtown, 279Hanover Inn, 65hard drive, Warp 9, 66 Hardy-Weinberg Law, 349harmonic function, 428Harvard, 27hat check problem, 82, 85, 105heights distribution of, 345 helium, 107HEYDE, C., 378HILL, T., 196Holmes, Sherlock, 91HorseRace (program), 6hospital, 14, 250HOWARD, R. A., 406HTSimulation (program), 6HUDDE, J., 148HUIZINGA, F., 389HUYGENS, C., 147, 243{245hypergeometric distribution, 193hypotheses, 145hypothesis testing, 101 Inclusion-Exclusion Principle, 104 independence of events, 139, 164 mutual, 141 independence of random variables mutual, 143, 165 independence of random variables, 143, 165 independent trials process, 144, 168interarrival time, average, 208interleaving, 120irreducible Markov chain, 433Isle Royale, 202 JAYNES, E. T., 49 JOHNSONBOUGH, R., 153joint cumulative distribution function, 165 joint density function, 165joint distribution function, 142joint random variable, 142 KAHNEMAN, D., 38 Kemeny’s constant, 469, 470KEMENY, J. G., 200, 406, 466 INDEX 507 KENDALL, D. G., 378 KEYFITZ, N., 383KILGOUR, D. M., 179, 182KINGSTON, J. G., 157KONOLD, C., 161KOZELKA, R. M., 344 Labouchere betting system, 12, 13 LABOUCHERE, H. du P., 12LAMPERTI, J., 267, 324LAPLACE, P. S., 51, 53, 350last return to the origin, 482Law (program), 310Law of Averages, 70Law of Large Numbers, 307, 316 for Ergodic Markov Chains, 439Strong, 70 LawContinuous (program), 318lead change, 482LEONARD, B., 256LEONTIEF, W. W., 426LEVASSEUR, K., 485library problem, 82life table, 39light bulb, 66, 72, 172Linda problem, 38LINDEBERG, J. W., 344LIPSON, A., 161Little’s law for queues, 276Lockhorn, Mr. and Mrs., 65log normal density, 224lottery Powerball, 204 LUCAS, E., 119 MAISTROV, L., 150, 310 MANN, B., 120margin of error, 335marginal distribution function, 143Markov chain, 405 absorbing, 415ergodic, 433irreducible, 433regular, 433 Markov ChainsCentral Limit Theorem for, 464 Fundamental Limit Theorem for Regular, 448 MARKOV, A. A., 464martingale, 241, 242, 428 origin of word, 11 martingale betting system, 11, 14, 248matrix fundamental, 419 Matrix Powers (program), 407 maximum likelihood estimate, 198, 202 Maximum Likelihood Principle, 91, 117 Maxwell density, 215maze, 440, 453McCRACKEN, D., 10mean, 226mean flrst passage matrix, 455mean flrst passage time, 452mean recurrence matrix, 455mean recurrence time, 454memoryless property, 68, 164, 206milk, 252modular arithmetic, 10moment generating function, 366, 395moment problem, 368, 398moments, 365, 394Monopoly, 469MonteCarlo (program), 42Monty Hall problem, 136, 161moose, 202mortality table, 246mule kicks, 201MULLER, M. E., 213multiple-gene hypothesis, 348mustache, 153mutually independent events, 141mutually independent random variables, 143 negative binomial distribution, 186 NEGRINI, M., 196New York Times, 340New York Yankees, 118, 253 508 INDEX New-Age Solitaire, 130 NEWCOMB, S., 196NFoldConvolution (program), 287normal density, 47, 212NormalArea (program), 322nursery rhyme, 84 odds, 27 ordering, random, 127ordinary generating function, 370ORE, O., 30, 31outcome, 18Oz, Land of, 406, 439 Pascal’s triangle, 94, 103, 108 PASCAL, B., 4, 32{35, 107, 112{113, 156, 242, 245 paternity suit, 222PEARSON, K., 9, 351PENNEY, W., 432People v. Collins, 153, 202PERLMAN, M. D., 45permutation, 79 flxed points of, 82 Philadelphia 76ers, 15photons, 106Pickwick, Mr., 153Pilsdorfi Beer Company, 279PITTEL, B., 256point count, 287Poisson approximation to the binomial distribution, 189 Poisson distribution, 187 variance of, 262 poker, 95polls, 333Polya urn model, 152, 174ponytail, 153posterior probabilities, 145Powerball lottery, 204PowerCurve (program), 102Presidential election, 336PRICE, C., 86prior probabilities, 145probabilityBayes, 136 conditional, 133frequency concept of, 2of an event, 19transition, 406vector, 407 problem of points, 32, 112, 147, 156process, random, 128PROPP, J., 256PROSSER, R., 200protons, 106P¶OLYA, G., 15, 17, 475 quadratic equation, roots of, 73 quantum mechanics, 107QUETELET, A., 350Queue (program), 208queues, 186, 208, 275quincunx, 351 RABELAIS, F., 12 racquetball, 157radioactive isotope, 66, 71RAND Corporation, 10random integer, 39random number generator, 2random ordering, 127random process, 128random variable, 1, 18 continuous, 58discrete, 18functions of a, 210joint, 142 random variables independence of, 143mutual independence of, 143 random walk, 471 inndimensions, 17 RandomNumbers (program), 3RandomPermutation (program), 82rank event, 160raquetball, 13rat, 440, 453Rayleigh density, 215, 295records, 83, 234 INDEX 509 Records (program), 84 regression on the mean, 282regression to the mean, 345, 352regular Markov chain, 433reliability of a system, 154restricted choice, principle of, 182return to the origin, 472 flrst, 473last, 482probability of eventual, 475 reversibility, 463reversion, 352ri†e shu†e, 120RIORDAN, J., 86rising sequence, 120rnd, 42ROBERTS, F., 426Rome, 30ROSS, S., 270, 276roulette, 13, 237, 432run, 229R¶ENYI, A., 167 SAGAN, H., 237 sample, 333sample mean, 265sample space, 18 continuous, 58countably inflnite, 28inflnite, 28 sample standard deviation, 265sample variance, 265SAWYER, S., 412SCHULTZ, H., 255SENETA, E., 378, 444service time, average, 208SHANNON, C. E., 465SHOLANDER, M., 39shu†ing, 120SHULTZ, H., 256SimulateChain (program), 439simulating a random variable, 211snakeeyes, 27SNELL, J. L., 87, 175, 406, 466snowfall in Hanover, 83spike graph, 6 Spikegraph (program), 6spinner, 41, 55, 59, 162spread, 266St. Ives, 84St. Petersburg Paradox, 227standard deviation, 257standard normal random variable, 213 standardized random variable, 264standardized sum, 326state absorbing, 416of a Markov chain, 405transient, 416 statistics applications of the Central Limit Theorem to, 333 stepping stones, 412SteppingStone (program), 412stick of unit length, 73STIFEL, M., 110STIGLER, S., 350Stirling’s formula, 81STIRLING, J., 88StirlingApproximations (program), 81 stock prices, 241StockSystem (program), 241Strong Law of Large Numbers, 70, 314 suit event, 160SUTHERLAND, E., 182 t-density, 360 TARTAGLIA, N., 110tax returns, 196tea, 252telephone books, 256tennis, 157, 424tetrahedral numbers, 108THACKERAY, W. M., 14THOMPSON, G. L., 406THORP, E., 247, 253time to absorption, 419 510 INDEX TIPPETT, L. H. C., 10 traits, independence of, 216transient state, 416transition matrix, 406transition probability, 406tree diagram, 24, 76 inflnite binary, 69 Treize, 85triangle acute, 73 triangular numbers, 108trout, 198true-false exam, 267Tunbridge, 154TVERSKY, A., 14, 38Two aces problem, 181two-armed bandit, 170TwoArm (program), 171type 1 error, 101type 2 error, 101typesetter, 189 ULAM, S., 11 unbiased estimator, 266uniform density, 205uniform density function, 60uniform distribution, 25, 183uniform random variables sum of two continuous, 63 unshu†e, 122USPENSKY, J. B., 299utility function, 227 VANDERBEI, R., 175 Vandermonde determinant, 369variance, 257, 271 calculation of, 258 variation distance, 128VariationList (program), 128volleyball, 158von BORTKIEWICZ, L., 201von MISES, R., 87von NEUMANN, J., 10, 11vos SAVANT, M., 40, 86, 136, 176, 181Wall Street Journal, 161 watches, counterfeit, 91WATSON, H. W., 378WEAVER, W., 465Weierstrass Approximation Theorem, 315 WELDON, W. F. R., 9Wheaties, 118, 253WHITAKER, C., 136WHITEHEAD, J. H. C., 181WICHURA, M. J., 45WILF, H. S., 91, 474WOLF, R., 9WOLFORD, G., 159Woodstock, 154 Yang, 130 Yin, 130 ZAGIER, D., 485 Zorg, planet of, 90