Home / Math and Physics Files / Math / Spectral Theory Book / Work for Aug 2013 Update / App G on random var
lecture-5 stat independence
PDF · 4 pages · 78.5 KB
Open PDF file
Four-page lecture handout dated 14 September 2005, labelled Lecture 5, apparently from someone else's probability course and filed under the random-variable appendix of the spectral theory book work. It covers statistical independence versus mutual exclusivity, probability mass functions, expectation and variance, Bernoulli variables and indicators, linearity of expectation for sums, and the binomial distribution with its mean np.
AI-written summary; may contain errors.
Extracted text (machine-read; may contain errors)
Lecture 5: Statistical Independence, Discrete
Random Variables
14 September 2005
1 Statistical Independence
If
Pr (AjB) = Pr ( A)
we say that Aisstatistically independent ofB: whether Bhappens makes no
dierence to how often Ahappens
Since Pr ( A\B) = Pr ( AjB) Pr (B), ifAis independent of B, then
Pr (A\B) = Pr ( A) Pr (B)
If this holds, though, then Bis also independent of A:
Pr (BjA) =Pr (A\B)
Pr (A)=Pr (A) Pr (B)
Pr (A)= Pr ( B)
so we can just say \ AandBare independent".
Note that events can be logically or physically independent but still sta-
tistically dependent. Let A= \scored above 700 on the math SAT" and B=
\attends CMU". These are logically independent (neither one implies the other),
but statistically quite dependent, because Pr ( AjB)>Pr (A).
Statistical independence means one event conveys no information about the
other; statistical dependence means there is some information. Making this
precise is the subject of information theory. Information theory is my area of
research, so if I start talking about it I won't shut up; so I won't start.
Statistically independent is notthe same as mutually exclusive: if AandB
are mutually exclusive, then they can't be independent, unless one of them is
probability 0 to start with:
Pr (() A\B) = 0 = Pr (() A)Pr (() B)
\Mutually exclusive" is denitely informative: if one happens, then the other
can't. (They could still both not happen, unless they're jointly exhaustive.)
1
2 Discrete Random Variables
Arandom variable is just one which is the result of a random process, like an
experimental measurement subject to noise, or a reliable measurement of some
uctuating quantity. What this means in practice is that there is a certain
probability for the random variable Xto take on any particular value in the
sample space.
By convention, we'll use capital letters, X; Y; Z; W; : : : for random variables,
and the corresponding lower-case letters for points in the sample space | partic-
ular outcomes or realizations of the random variable. The dierence between
Xandxis the dierence between \the sum of two dice" and \5".
If the sample space is discrete, we completely specify the random variable by
giving the probability for each elementary outcome or realization, Pr ( X=x).
This is often abbreviated p(x), and called the probability distribution or
probability distribution function . Because it's a probability, p(x)0 for
allx, andP
xp(x) = 1. Conversely, any function which satises those two rules
can be a probability distribution.
The probability distribution is also sometimes called the probability mass
function (p.m.f.), on the analogy of having a unit mass to spread over the
sample space.
Any function of a random variable is again a random variable. But it may
be a trivial one, like sin2X+ cos2X.
2.1 Expectation
The expectation of a random variable is its average value, with weights in the
average given by the probability distribution
E[X]X
xPr (X=x)x
Continuing the mass analogy from the probability mass function, the expecta-
tion is the location of the center of mass.
Expectation is like the population mean, so the basic properties of the mean
carry over:
E[aX+b] =aE[X] +b
IfXY, then E[X]E[Y]
If we want to know the expectation of a function of X,E[f(X)], we just apply
the formula:
E[f(X)] =X
xPr (X=x)f(x)
The variance is the expectation of ( X E[X])2.
Var (X)]X
xp(x)(x E[X])2
2
Var (X) gives an indication of just how much spread there is in the population
around the average (expectation) value | how much slop we should anticipate
around the expectation.
(The variance is the moment of inertia around the center of mass.)
3 Bernoulli Random Variables
ABernoulli random variable is just one which takes on the values 0 or 1. We
completely specify the distribution with one number, Pr ( X= 1) = p. Bernoulli
variables are really simple, but a lot of more interesting things can be represented
using them, so they're an important place to start.
The expectation is p, and the variance is p(1 p). Let's see that in detail:
E[X] =1X
x=0xp(x)
= 0p(0) + 1p(1)
=p
Var (X) = E
X2
(E[X])2
=
02p(0) + 12p(1)
(p)2
=p p2
=p(1 p)
Notice that if we change p, we get a dierent probability distribution, but
always one which has the same mathematical form. Also, the mean, variance,
and all the other properties of Xdepend only on p. We say that pis the
parameter of the Bernoulli distribution.
3.1 Indicators
For any event A, we can dene a Bernoulli variable which is 1 when Ahappens
and 0 otherwise. This is the indicator variable or indicator function for A, 1A.
So Pr ( A) =E[1A]. These prove useful, because in some situations it's easier
for us to calculation E[1A], than to directly nd Pr ( A).
4 Sums of Random Variables
Suppose we have two random variables, which we'll call XandY. We can add
them, to get a new variable, X+Y. It will also be random. It will have some
expectation, E[X+Y]. What is that?
E[X+Y] =X
x;y(x+y)Pr (X=x; Y =y)
3
=X
x;yxPr (X=x; Y =y) +X
x;yyPr (X=x; Y =y)
=X
xxX
yPr (X=x; Y =y) +X
yyX
xPr (X=x; Y =y)
By total probability,P
yPr (X=x; Y =y) = Pr ( X=x), likewiseP
xPr (X=x; Y =y) =
Pr (Y=y). So,
E[X+Y] =X
xxPr (X=x) +X
yyPr (Y=y)
=E[X] +E[Y]
5 Binomial
Finally, a use for all that combinatorics!
Imagine we have n\trials" or units, each of which can succeed (or be 1) with
probability p| that is, each trial is a Bernoulli variable. Assume the trials are
independent. Our random variable Xis the number of successes, or the sum of
the Bernoulli variables. What is the distribution of X?
Let's start by thinking of a simple example, where n= 3. If X= 0, it
must be the case that none of the three trials succeeded. There is only one
sequence of outcomes which will do this: (0 ;0;0). The probability of this event
is thus (1 p)3. Now consider X= 1. We can get this by the outcomes
(0;0;1), (0 ;1;0), and (1 ;0;0). All three are equally likely: they have probability
p(1 p)2, since there's one success (at probability p) and two failures (at 1 p
each). X= 2, similarly, has three possibilities, each of which has probability
p2(1 p), and X= 3 only one, probability p3. Clearly X < 0 and X > 3 are
both impossible. So we have Pr ( X=x) = 3
x
px(1 p)3 x, if 0xn, and
Pr (X=x) = 0 otherwise.
This is going to be generally true, whatever nis: the probability that X=x
will be the number of ntrials sequences with xsuccess, times the probability of
any one such sequence, or
Pr (X=x) =n
x
px(1 p)n k
Notice that this has twoparameters, nandp.
Because a binomial variable with parameters n; pis the sum of nindependent
random variables with parameter p, we can nd the expectation very simply:
E[X] =np. The alternative to this is a quite ugly sum.
4