Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Math / Spectral Theory Book / Work for Aug 2013 Update / App G on random var

lecture-5 stat independence

PDF · 4 pages · 78.5 KB
Open PDF file

Four-page lecture handout dated 14 September 2005, labelled Lecture 5, apparently from someone else's probability course and filed under the random-variable appendix of the spectral theory book work. It covers statistical independence versus mutual exclusivity, probability mass functions, expectation and variance, Bernoulli variables and indicators, linearity of expectation for sums, and the binomial distribution with its mean np.

AI-written summary; may contain errors.

Extracted text (machine-read; may contain errors)
Lecture 5: Statistical Independence, Discrete Random Variables 14 September 2005 1 Statistical Independence If Pr (AjB) = Pr ( A) we say that Aisstatistically independent ofB: whether Bhappens makes no di erence to how often Ahappens Since Pr ( A\B) = Pr ( AjB) Pr (B), ifAis independent of B, then Pr (A\B) = Pr ( A) Pr (B) If this holds, though, then Bis also independent of A: Pr (BjA) =Pr (A\B) Pr (A)=Pr (A) Pr (B) Pr (A)= Pr ( B) so we can just say \ AandBare independent". Note that events can be logically or physically independent but still sta- tistically dependent. Let A= \scored above 700 on the math SAT" and B= \attends CMU". These are logically independent (neither one implies the other), but statistically quite dependent, because Pr ( AjB)>Pr (A). Statistical independence means one event conveys no information about the other; statistical dependence means there is some information. Making this precise is the subject of information theory. Information theory is my area of research, so if I start talking about it I won't shut up; so I won't start. Statistically independent is notthe same as mutually exclusive: if AandB are mutually exclusive, then they can't be independent, unless one of them is probability 0 to start with: Pr (() A\B) = 0 = Pr (() A)Pr (() B) \Mutually exclusive" is de nitely informative: if one happens, then the other can't. (They could still both not happen, unless they're jointly exhaustive.) 1 2 Discrete Random Variables Arandom variable is just one which is the result of a random process, like an experimental measurement subject to noise, or a reliable measurement of some uctuating quantity. What this means in practice is that there is a certain probability for the random variable Xto take on any particular value in the sample space. By convention, we'll use capital letters, X; Y; Z; W; : : : for random variables, and the corresponding lower-case letters for points in the sample space | partic- ular outcomes or realizations of the random variable. The di erence between Xandxis the di erence between \the sum of two dice" and \5". If the sample space is discrete, we completely specify the random variable by giving the probability for each elementary outcome or realization, Pr ( X=x). This is often abbreviated p(x), and called the probability distribution or probability distribution function . Because it's a probability, p(x)0 for allx, andP xp(x) = 1. Conversely, any function which satis es those two rules can be a probability distribution. The probability distribution is also sometimes called the probability mass function (p.m.f.), on the analogy of having a unit mass to spread over the sample space. Any function of a random variable is again a random variable. But it may be a trivial one, like sin2X+ cos2X. 2.1 Expectation The expectation of a random variable is its average value, with weights in the average given by the probability distribution E[X]X xPr (X=x)x Continuing the mass analogy from the probability mass function, the expecta- tion is the location of the center of mass. Expectation is like the population mean, so the basic properties of the mean carry over: E[aX+b] =aE[X] +b IfXY, then E[X]E[Y] If we want to know the expectation of a function of X,E[f(X)], we just apply the formula: E[f(X)] =X xPr (X=x)f(x) The variance is the expectation of ( XE[X])2. Var (X)]X xp(x)(xE[X])2 2 Var (X) gives an indication of just how much spread there is in the population around the average (expectation) value | how much slop we should anticipate around the expectation. (The variance is the moment of inertia around the center of mass.) 3 Bernoulli Random Variables ABernoulli random variable is just one which takes on the values 0 or 1. We completely specify the distribution with one number, Pr ( X= 1) = p. Bernoulli variables are really simple, but a lot of more interesting things can be represented using them, so they're an important place to start. The expectation is p, and the variance is p(1p). Let's see that in detail: E[X] =1X x=0xp(x) = 0p(0) + 1p(1) =p Var (X) = E X2 (E[X])2 = 02p(0) + 12p(1) (p)2 =pp2 =p(1p) Notice that if we change p, we get a di erent probability distribution, but always one which has the same mathematical form. Also, the mean, variance, and all the other properties of Xdepend only on p. We say that pis the parameter of the Bernoulli distribution. 3.1 Indicators For any event A, we can de ne a Bernoulli variable which is 1 when Ahappens and 0 otherwise. This is the indicator variable or indicator function for A, 1A. So Pr ( A) =E[1A]. These prove useful, because in some situations it's easier for us to calculation E[1A], than to directly nd Pr ( A). 4 Sums of Random Variables Suppose we have two random variables, which we'll call XandY. We can add them, to get a new variable, X+Y. It will also be random. It will have some expectation, E[X+Y]. What is that? E[X+Y] =X x;y(x+y)Pr (X=x; Y =y) 3 =X x;yxPr (X=x; Y =y) +X x;yyPr (X=x; Y =y) =X xxX yPr (X=x; Y =y) +X yyX xPr (X=x; Y =y) By total probability,P yPr (X=x; Y =y) = Pr ( X=x), likewiseP xPr (X=x; Y =y) = Pr (Y=y). So, E[X+Y] =X xxPr (X=x) +X yyPr (Y=y) =E[X] +E[Y] 5 Binomial Finally, a use for all that combinatorics! Imagine we have n\trials" or units, each of which can succeed (or be 1) with probability p| that is, each trial is a Bernoulli variable. Assume the trials are independent. Our random variable Xis the number of successes, or the sum of the Bernoulli variables. What is the distribution of X? Let's start by thinking of a simple example, where n= 3. If X= 0, it must be the case that none of the three trials succeeded. There is only one sequence of outcomes which will do this: (0 ;0;0). The probability of this event is thus (1p)3. Now consider X= 1. We can get this by the outcomes (0;0;1), (0 ;1;0), and (1 ;0;0). All three are equally likely: they have probability p(1p)2, since there's one success (at probability p) and two failures (at 1 p each). X= 2, similarly, has three possibilities, each of which has probability p2(1p), and X= 3 only one, probability p3. Clearly X < 0 and X > 3 are both impossible. So we have Pr ( X=x) =3 x px(1p)3x, if 0xn, and Pr (X=x) = 0 otherwise. This is going to be generally true, whatever nis: the probability that X=x will be the number of ntrials sequences with xsuccess, times the probability of any one such sequence, or Pr (X=x) =n x px(1p)nk Notice that this has twoparameters, nandp. Because a binomial variable with parameters n; pis the sum of nindependent random variables with parameter p, we can nd the expectation very simply: E[X] =np. The alternative to this is a quite ugly sum. 4