Home / Math and Physics Files / Math / Spectral Theory Book / Work for Aug 2013 Update / App G on random var
Appendix G
DOCX · 450.4 KB
Open DOCX file
Draft appendix dated 3.26.05 from Phil's spectral theory book (August 2013 update folder). It builds up random variables via experiments, sample spaces, events, pmf and pdf, using dice, a spinner and coin tosses, and explains the capital-letter notation (X as a function, x as its value). Later sections cover basic probability theory, ensemble experiments, two-dice experiments, experimental determination of distributions, and pulse train amplitude sequences, generalizing Appendix D.
AI-written summary; may contain errors.
Extracted text (machine-read; may contain errors)
This is the Title PhL 3.26.05
Appendix G: Random Variables, Probability Theory and Pulse Train Amplitudes 2
(a) What is a Random Variable ? Part I 2
(b) What is a Random Variable ? Part II: the Capital Letter Notation 4
(c) Basic Probability Theory 6
(d) Ensemble Experiments 15
(e) Experiments rolling two dice at the same time 17
(f) Experimental determination of discrete distribution functions 20
(g) Experiments with Sequences of Pulse Train Amplitudes 22
Appendix G: Random Variables, Probability Theory and Pulse Train Amplitudes
Appendix D discusses random variables X and Y, and then, for a pulse train, the random variables Yn and Ym . The narrow purpose of that Appendix is to establish the relation among parameters α,β,μ,σ since these relations are needed to verify certain results of our pulse train analysis.
In this Appendix we take a more general view of random variables which then provides a context in which Appendix D is a particular special case. This subject always seems "slippery". One can find a discussion like the one presented below in every book on probability theory, but here we wish perhaps to put our own "spin" on the subject with comments not always appearing in texts.
(a) What is a Random Variable ? Part I
Our opening definition is that a random variable is a real parameter for which there exists some probability distribution. The parameter is "random" in the sense that it does not take a fixed value, but can take a range of values. This definition does not require that the probability distribution be a "flat" distribution in which all possible values of the parameter are equally likely, though that is possible, Sometimes one associates the word "random" with this flat distribution sense, such as the face up number of a rolled unloaded die having probability P = 1/6. But the random in "random variable" just means there is variability in the value of the parameter that can be connected with some distribution. The parameter can of course be viewed as a "variable", since it varies, and in particular the parameter will be the variable which is the argument of the probability distribution function. If that distribution happens to be a Gaussian (normal) distribution, then one might refer to the parameter as a "Gaussian random variable".
The above definition is rather abstract and it helps to put it into a context which involves other concepts which also have to be defined. Perhaps the most general framework is to think of the parameter described above as a "random variable" as being a possible outcome of an experiment, so to two new words have appeared. An experiment is "something one does" and an outcome is "something one measures after the experiment is done". A simple example is experiment = roll one die, outcome = number facing up after the die is rolled = random variable n. The set of all possible outcomes is called the sample space, often indicated by Ω (element = ω) or S (element = s). For the die experiment, we have Ω = {1,2,3,4,5,6} as the sample space. Since this is a set, various concepts regarding sets can be applied. A set we know has subsets. Each subset of the sample space has a peculiar name, each subset is called an event. In the die experiment, a possible event would be {2,4,6} which event is that the die rolled an even number. Another event would be {6} which event is that the die rolled a 6. For this experiment, the sample space is discrete.
Next, consider experiment = one spins a "spinner". This is a traditional (in probability texts) physical object that one might think of as a horizontal Lazy Susan having an arrow from center to some point on the rim, or a Roulette Wheel with some point marked on the rim, or just a metal arrow you spin and drop on a table top. The outcome of this experiment is the direction in which the arrow points in range (0, 2π) perhaps clockwise relative to North. In this example, the sample space is all real numbers in the interval (0,2π). For this experiment, the sample space is a continuous set which one could regard as the limit of a discrete set as the number of elements increases.
We shall now assume that in the die experiment the die was a "loaded" die of some sort, and that in the spinner experiment, the arrow was perhaps influenced by magnetic fields or by bad bearings on the Lazy Susan. If we do each experiment many times and each time we observe and write down the outcome, we can then "bin" these outcomes and create a distribution by dividing each bin count by the total number of experiments done ( see section (f) below) . Here are possible distributions for our two experiments.
Fig G.1
For the spinner, the distribution is continuous and is called a probability density function or pdf. For the die, the distribution is discrete as is called a probability mass function or pmf . One could in fact write down a pdf for the left picture assuming n was a continuous variable in this manner,
pdf(n) = Σi=16 pmf(n) δ(i-n) . (G.1)
Both a pmf and pdf must sum to one, since the probability of some outcome is always 1,
Σk=16 pmf(k) = 1 !Syntax Error, Ipdf(x)dx = 1.
Notice that the above equations are consistent in that
1 = !Syntax Error, Ipdf(n)dn = !Syntax Error, I[ Σi=16 pmf(n) δ(i-n)]dn = Σi=16 pmf(n) !Syntax Error, I δ(i-n)]dn
= Σi=16 pmf(n) = 1 .
The motivation for these names pdf and pmf comes from physics. One can have a continuous distribution of mass in some compressible fluid, say, where it would be called a mass density ρ. But in the physics of idealized point particles of mass mi, the mass is congealed at specific points in space. For a set of particles of mass mi at spatial locations ri one could still write a mass density ρ(r) as
ρ(r) = Σi mi δ(r-ri)
and one sees the analogy with (G.1) above.
Here then we have seen two "random variables" in action. One, n, is the discrete outcome of a die roll, the other θ is the continuous outcome of a spinner spin. Each of these is a real parameter and for each there exists an associated probability distribution, as plotted above. Thus, each of these parameters fulfills the requirements of our opening definition of a "random variable".
(b) What is a Random Variable ? Part II: the Capital Letter Notation
We shall continue to examine various experiments to see what new concepts appear.
Consider experiment = coin toss with sample space S = {heads,tails} = possible outcomes. The new feature here is that the outcomes are not real numbers and therefore don't fit our requirement that a random variable be a real parameter. This is easily remedied by assigning real numbers to each outcome, such as heads → 1 and tails → 0. If s is an element of the sample space S = {heads,tails}, then we can define a little function h = H(s) such that 1 = H(heads) and 0 = H(tails). In general this function is a mapping from the sample space S to the real numbers, so we can say H:S→R. In this situation, we can think of h as being a "random variable" as defined above. It is a real parameter and it has an associated probability distribution which in this case is a pmf since S is discrete. For a "loaded" coin we might have
Fig G.2
We are now going to slightly alter our definition of a random variable so it is more precise. As a preliminary, forget probability and just consider some function y = f(x). In this equation, the object f is clearly a function, while the object y is a value that function can take. Thus, y and f are not the same thing. They are equal in the sense that y = f(x) for some x, but y and f belong to completely different classes of objects. Perhaps y is a real variable in R, whereas f is a mapping (function) from R to R. It is quite useful to have different symbols for y and f. If one were to write y = y(x), the reader would understand what was meant, but now y has two separate meanings ( function, and value of function). To maintain precision, if is better that these symbols be different. One choice is to represent the function f(x) as Y(x), so the function name is now Y, while y is a value the function Y can take. Then y = Y(x) and things are clear.
Now we return to h = H(s). H is a function or mapping from S to R, whereas h is a real number. We would not in general say that h and H are the same object. For the coin toss experiment, it is really the function H that is the random variable, whereas h is a value that random variable can take. In our previous experiments with die and spinner, it happened that the function H was the identity function, so we didn't "see it" in the discussion. But still for the spinner, we would write θ = Θ(arrow position) and, as we set things up, if an arrow position s is in the range (0,2π), then we happen to have Θ(φ) = φ and Θ = 1. The function Θ would have been more visible had we perhaps taken arrow position in degrees and put θ in radians.
So one more time with our improved definition: a random variable X is the function which maps the sample space S to the real axis, X: S→R. The values that the random variable X takes are called x. For each element of the sample space, there is some corresponding real number x. In considering x = X(s), we make a distinction between the function X and the value x that it takes for some particular sample space element s. The other ingredient to the definition of a random variable remains unchanged: there is some probability distribution associated with the values x that the random variable X takes if an experiment is repeated many times. That is, x does not take a single fixed value. The distribution need not be a flat distribution, which one might call a random distribution in which every value of x is equally likely. Thus, for this latter flat concept of "random" we arrive at the famous statement that a random variable is neither a variable (it is a function) nor is it random (distribution can be non-random).
What different name might one use? Perhaps a "probabilistic function" which maps the possibly non-numeric outcomes of an experiment to a real parameter, which parameter has some non-trivial probability distribution which can be observed by doing the experiment many times. The problem with this phrase is that it does not focus on that parameter which is the main item of interest, the variable of random variable.
For some general experiment, the outcomes are likely to be non-numerical in nature, and a function like H(s) will be required to map the outcomes onto the real number axis. For a discrete sample space, we might represent this situation as follows:
Fig G.3
In this case, an event which is any subset of the outcomes in the sample space will map into some set of points on the real axis. For a continuous sample space, the picture is a little different,
Fig G.4
The new feature here is a certain continuity requirement: every interval of the random variable x (like the interval shown) must map back into some subset of the sample space. Since subsets are events, this means that every interval of the x axis must map back to some well-defined event. Going the other way, not every event (subset) maps into an interval of the x axis. The continuity idea is that if two points are close together on the real x axis, then they must be close together in the sample space. This in turn means that the sample space has to have some kind of metric to allow a notion of distance between two points. Furthermore, as the subset is expanded, the interval it maps to cannot becomes smaller!
Having now discussed random variables, outcomes, experiments, sample spaces, and events, we are in a good position to review some general probability theory.
(c) Basic Probability Theory
In the above discussion, if we have random variables A,B,C.... , we can write a joint probability distribution function in this manner
pdf(A=a, B=b, C=c.....) (G.2)
where now we don't distinguish whether the various sample spaces are continuous or discrete, we just write anything as a pdf (with the understanding of (G.1) above). The meaning here is that pdf(...) is the probability that random variable A has value a, while at the same time (the same experiment) random variable B has value b, and so on. For the continuous case, pdf(...) da db dc... is the same probability but for the range da of a and db of b, etc.
If all these random variables are statistically independent (such as A = number of dust particles on your pillow and B = temperature at some location on Pluto), this joint probability distribution factors,
pdf(A=a,B=b,C=c.....) = pdf(A=a) pdf(B=b) pdf(C=c) .... (G.3a)
For two random variables A and B, one would have
pdf(X=x,Y=y) = pdf(X=x) pdf(Y=y) . (G.4a)
In this case, A and B are statistically independent and are said to be uncorrelated. If one regards
pdf(X=x,Y=y) as a function fy(x) for various fixed values of y, (G.4a) says that the shape of this function is not influenced by the values of y, only the overall scale of f(x) is affected by y. When we formally define the notion of correlation below, we will see that the official correlation between two random variables X and Y vanishes if (G.4a) is true.
We now adopt a shorthand notation:
p(a,b,c....) ≡ pdf(A=a,B=b,C=c.....) = pdfABC...(a,b,c...) (G.5)
When one sees p(2, -4, 0. 4...) one must remember that the arguments correspond to values of specific random variables, and if things become unclear, one must revert to the fuller notation. One trick is to use a parameter name that reminds the reader of the random variable name, such as a for A: p(a) = P(A=a).
Using this shorthand notation we may express the two equations above as
p(a,b,c.....) = p(a)p(b)p(c)...... statistical independence (G3.b)
p(x,y) = p(x)p(y) uncorrelated (G.4b)
Any pdf is normalized to 1 since the probability of all possible outcomes (mapped from sample spaces to the random variables) is 1. Thus,
∫∫.... p(x,y....) dx dy... = 1 or Σx,y.... p(x,y....) = 1 . (G.6)
An example is that ∫p(x)dx = 1 or Σxp(x) = 1. If x is discrete and y is continuous, Σx∫p(x,y)dy = 1, but we won't bother to show all such "mixed" cases below.
In general, if p(a,b,c...) is some Nth order joint probability, one can find lower order probabilities by summing over some of the variables:
Fact: To obtain a lower-order joint probability from a higher-order one,
p(x,y,z...) = Σ'a,b,c... p(a,b,c.....) (G.7)
Here {x,y,z...} is some subset of {a,b,c....} and the notation Σ' means that we sum over all variables in {a,b,c.... } except those in the subset { x,y,z...} . One could state (G.7) in a set notation this way
p(S') = ΣS-S' p(S) where S' S (G.8)
Proof: The proof is merely the observation that if we are only interested in {x,y,z...}, we go ahead and let all the other variables {a,b,c....} – {x,y,z...} take all possible values and we add up the probability of each of these cases to get its contribution to p(x,y,z...). This is just a case of adding probabilities of outcomes to get a total probability of interest.
Examples:
p(x) = Σy p(x,y)
p(x) = Σy,z p(x,y,z)
p(x,y) = Σz p(x,y,z)
p(x,y) = Σa,b,c...≠ x,y p(a,b,c......) (G.9)
Before we can continue our little presentation, we need to state and prove an important theorem about random variables which are functions of other random variables. To this end, we must first digress on a set of math Lemmas.
Some Math Lemmas
Consider z = f(x,y) = f(r) where r = (x,y). Here f is a "function" which means it is a single-valued function which means under the mapping f: R2→R, every vector r in the domain lands in some unique location in the range. It is possible and in fact likely that f will be a many-to-one function, meaning for a given z in the range, there might be several ri in the domain such that z = f(ri). If the domain is discrete, then so is the range. We might have this situation:
domain Df range Rf Fig G.5
We could then talk about Σx,y being the sum of all points r = (x,y) in the domain such that f(x,y) = z for some z in the range. We can exhaust all points in the range Rf by doing zi = f(ri) and letting ri exhaust all (x,y) in the domain Df. This is how we discover the extent of the range Rf.
We now make this claim, where we have in mind some unseen quantity being acted upon by both sides of the equation (symbol means "such that")
(G.10)
Proof: If on the left we first sum over the points (x,y) corresponding to z according to z = f(x,y), and then we sum over all z in the range of f, our sum includes every point (x,y) in Df exactly once. The double sum is simply a certain ordering of the total sum shown on the right. The same point (x,y) cannot show up
twice in the double sum for two different values of z, because if it did, f(x,y) would map that point into those two z different values, but f is supposedly single-valued. Nor can any point (x,y) be omitted in the double sum on the left because the range Rf for z was created by exhausting all points (x,y) in Df, so any (x,y) in Df corresponds to some z in the range Rf.
Now for a continuous domain Df, we might have
domain Df range Rf Fig G.6
where now an entire continuous curve C (red) in the domain maps to some z in Rf. The equation of the red curve C is z = f(x,y) for some fixed value of z. The heavy curves including this red one are curves of constant z. We can imagine some other function s = g(x,y) whose curves of constant s form an orthogonal curvilinear coordinate system (right angles at any point) with the curves of z = f(x,y) = constant. The parameter s would then vary along each constant-z curve, marking off points along the curve. The analogous statement to the discrete statement made above is this (only a claim, we shall not prove it)
(G.11)
where J is the Jacobian between coordinates (x,y) and (z,s). This is the same Jacobian idea that appears in taking (x,y) to polar coordinates (r,θ) where dxdy = rdrdθ with J = r. We shall not pursue this Jacobian matter further other than to claim it is possible to find a g(x,y) that works and that certain technical issues arise concerning the reasonableness of the function f(x,y).
Comment: In order for the Jacobian to exist and be well-behaved, function f(x,y) must be continuous and differentiable (C1) in both variables, and the mapping between (x,y) and (z,s) must be essentially one- to- one (invertible) so that given (z,s) one can compute (x,y) and vice versa.
Example: Suppose z = f(x,y) = 2x3 + 3y2 and we are interested in the curve C(z=2). This curve C is the intersection of the surfaces z = 2x3 + 3y2 (red) and z = 2 (gray), as illustrated here:
Fig G.7
If we had z = f(x,y,w) so Z is a function of three random variables X,Y and W, then our two results above would become
(G.12)
The first line seems straightforward, but the second is more complicated. Inside the square bracket we now have dS being a differential patch of area, and S(z) is a 2D surface in the (x,y,z) space on which z is a constant according to z = f(x,y,w). For example, in spherical coordinates we write r = x2+y2+z2 and a surface of constant r is a spherical shell. Now we have to imagine two other coordinates s1 = g(x,y,w) and s2 = g(x,y,w) so that (z,s1,s2) form an orthogonal coordinate system and then the new J is the Jacobian between (x,y,w) and (z,s1,s2).
In a more economical notation, and changing to z = f(a,b,c) we can write (G.10) and (G.11) as
Σz [ Σzx,y] = Σx,y (G.13a)
∫dz [ ∫C(z) ds J ] = ∫∫ da db (G.13b)
and then for (G.9),
Σz [ Σza,b,c ] = Σa,b,c (G.14a)
∫dz [ ∫S(z)dS J ] = ∫∫∫da db dc (G.14b)
where Σza,b,c indicates that the sum is restricted to those (a,b,c) values for which f(a,b,c) = z.
Hopefully it is clear how this general idea can be extended to Z = f(A,B,C,D....) where random variable Z is a function of some arbitrary number N of random variables A,B,C,D... .
Fact: One can organize a total sum/integral over the space of N random variables in this manner:
Σz [ Σza,b,c... ] = Σa,b,c... (G.15a)
∫ dz [ ∫S(z) dS J ] = ∫∫....∫ da db dc ...... (G.15b)
In the first line, Σz means the sum is constrained to be only over those a,b,c... such that f(a,b,c...) = z. In the second line, S(z) is an N-1 dimensional surface located within the N dimensional space of (a,b,c....), which surface is defined by f(a,b,c....) = z .
With these Lemmas out of the way, we can now resume our basic probability theory review.
Fact: Any reasonable real function of random variables is a random variable. (G.16)
Proof: First consider Z = f(X,Y), where X and Y are random variables. If x is an allowed value of X, and y of Y, then the allowed values of Z will be z = f(x,y). Since x and y must be real, and since f is a real function, the allowed values of z are real, one of the requirements for Z to be a random variable. At our level of rigor, it only remains to find the probability distribution associated with Z. We claim this is given by,
p(z) = Σzx,y p(x,y) (G.17a)
p(z) = ∫C(z) ds J p(x(s,z),y(s,z)) = ∫C(z) J p(x,y) . (G.17b)
Looking at (G.17a), for a given value of z, in z = f(x,y) only certain x and y values are possible. Thus, only these values of x and y appearing in p(x,y) can contribute to p(z). Other values of x and y will contribute to p(z') for some other z'. We can check normalization as follows using (G.13a)
Σz p(z) = Σz{ [Σzx,y] p(x,y)} = {Σz [Σzx,y]} p(x,y) = Σx,y p(x,y) = 1 .
What is the sample space for Z ? We can select one according to the following plan. Let RA and RB be the total sets of values of parameters a and b. Then RZ = f(RA,RB) treated as an equation of sets tells us the set RZ. We could then just select the sample space for Z to be SZ = RZ which would be a numerical sample space.
More generally, for Z = f(A,B,C....) with N arguments, the distribution for Z is given by
p(z) = Σza,b,c... p(a,b,c,....) (G.18a)
p(z) = ∫C(z) dS J p(a,b,c....) a = a(z, s1, s2..... sN-1), etc. (G.18b)
The function f must be C1 in all its arguments (as noted above) and must be such that the Jacobian J is well behaved (finite and non-zero) over the entire regions of interest in both spaces (a,b,c...) and (z,s1,s2...). This is why (G.16) contains the word "reasonable". QED
Definition : The expected value of a random variable X is given by
E(X) ≡ ∫x p(x) dx or E(X) ≡ Σx x p(x) = the mean, often called μX (G.19)
Comment 1: The expected value is sometimes called the expectation or the expectation value. In quantum mechanics the phrase expectation value predominates, but elsewhere it is the expected value. In quantum mechanics, all physical observables are random variables (position, momentum, energy, etc) except in quantum states which are eigenstates of the observable's quantum operator, in which case the observable takes a fixed value.
Comment 2: In (G.19) above for the mean (and generally below), we could write E(X) ≡ Σi xi p(xi) (as in Appendix D) where the xi are the possible values that random variable X can take. But in this section we use E(X) ≡ Σx x p(x) where then x itself represents the values X can take. This notation puts the summation form on a little more equal footing with the integral form ∫x p(x) dx .
Fact: If Z = f(X,Y), the expected value of Z is given by (G.20)
E(Z) = Σz z p(z) = Σx,y f(x,y) p(x,y) discrete
or
E(Z) = ∫ z p(z) dz = ∫∫ f(x,y) p(x,y) dx dy . continuous
Proof: For the discrete case
E(z) = Σz z p(z) // by the definition of an expected value of a random variable
= Σz z [ Σzx,y p(x,y)] // by (G.17a)
= Σz [ Σzx,y] f(x,y) p(x,y) // move the x,y sum to the left, and replace z = f(x,y)
= Σx,y f(x,y) p(x,y) // by (G.13a)
For the continuous case
E(Z) = ∫ z p(z) dz = ∫dz z [∫C(z)ds J p(x,y) ] // by (G.17b)
= ∫dz [∫C(z)ds J ] f(x,y) p(x,y) // move ∫ds to the left, replace z = f(x,y)
= ∫∫ dxdy f(x,y) p(x,y) // by (G.13b) QED
We now generalize this fact to obtain:
Theorem: If Q = f(X,Y,Z....) where X,Y,Z... are random variables, and if f is a reasonable real valued function, then
(1) Q is a random variable and
(2) E(f(X,Y,Z....)) = Σx,y,z... f(x,y,z....) p(x,y,z....)
E(f(X,Y,Z....)) = ∫∫∫...dxdydz.... f(x,y,z....) p(x,y,z....) (G.21)
Proof: This is a straightforward generalization (G.20) based on (G.18) and (G.15). One just mimics the proof of (G.20).
Special cases:
E(f(X)) = Σx f(x) p(x)
E(f(X)) = ∫dx f(x) p(x) (G.22)
E(f(X,Y)) = Σx,y f(x,y) p(x,y)
E(f(X,Y)) = ∫∫dx dy f(x,y) p(x,y) (G.23)
Next we define the covariance of X and Y, and evaluate it using (G.23),
cov(X,Y) ≡ E((X-μx)(Y-μy)) = Σx,y (x-μx) (y-μy) p(x,y)
or ∫∫ (x-μx) (y-μy) p(x,y) dx dy (G.24a)
Notice that
cov(X,Y) ≡ E(XY) - μxE(Y)- μyE(x) + μxμy = E(XY) - μxμy - μyμx+ μxμy = E(XY) - μxμy
so we have this alternate method of computing covariance
cov(X,Y) = E(XY) - μxμy . (G.24b)
The variance of X is the covariance of X with itself and is the square of the standard deviation σ(X),
var(X) ≡ [σ(X)]2 = cov(X,X) = E((X-μx)(X-μx)) = E((X-μx)2)
= Σx (x-μx)2 p(x) or ∫ (x-μx)2 p(x) dx (G.25)
where we use (G.22) with f(X) = (X-μx)2.
Finally, the correlation of X and Y is the covariance normalized by the two standard deviations,
corr(X,Y) ≡ cov(X,Y) / [σ(X) σ(Y) ] . (G.26)
One can take the limit of a joint distribution of random variables as two of the variables become the same. For example, for continuous and discrete (where we now use the alternate indexed notation),
limX→Y p(x,y) = p(x)δ(x-y) limX→Y p(xi,yi) = p(xi)δi,j (G.27)
The reason is that if X and Y are the same random variable, they cannot take different values. If X,Y and Z are all the same variable, then we would have (dropping the limit notation)
p(x,y,z) = p(x)δ(x-y)δ(x-z) or p(xi,yj,zk) = p(xi) δi,j δi,k (G.28)
Note that δ(x-y)δ(x-z) = δ(x-y)δ(y-z) = δ(x-z)δ(y-z) and δi,j δi,k = δi,j δj,k = δi,k δj,k .
We can apply (G.27) to obtain an alternate evaluation of the variance,
var(X) ≡ [σ(X)]2 = limX→Y cov(X,Y) = limX→Y [ ∫∫ (x-μx) (y-μY) p(x,y) dx dy ] (G.29)
= ∫∫ (x-μx) (y-μx) p(x)δ(x-y) dx dy
= ∫ (x-μx)2 p(x) dx or Σi (xi -μx)2 p(xi) / Σx (x -μx)2 p(x)
which is the same as (G.25).
Notice that the variance could in theory vanish if p(x) = δ(x - x1) so that the entire pdf is concentrated at a single value and then μX = x1 :
var(X) = ∫ (x-μx)2 p(x) dx = ∫ (x-μx)2 δ(x - x1) dx = (x1-μx)2 = (x1- x1)2 = 0 = σ(X) .
But then X is not a random variable since it's value is precisely determined as x1. A pdf of this form is not very interesting and normally one has σ(X) > 0.
Fact: var(X) = σ2(X) = E(X2) - {E(X)}2 (G.30)
Proof: E(X2) - {E(X)}2 = ∫x 2p(x) dx - {∫x p(x) dx} 2 = ∫x 2p(x) dx - μx2
var(X) = ∫ (x-μx)2 p(x) dx = ∫x 2p(x) dx -2μx∫x p(x) dx + μx2∫ p(x) dx
= ∫x 2p(x) dx - 2μx μx + μx2 1 = ∫x 2p(x) dx - μx2 . QED
Now that we have defined correlation in (G.26), we can see why being uncorrelated corresponds to a factoring pdf as was claimed earlier in (G.4a) or (G.4b). That is one item of the following fact:
Fact: The following statements are all equivalent for random variables X and Y : (G.31)
(a) X and Y are statistically independent
(b) X and Y are uncorrelated
(c) p(x,y) = p(x)p(y)
(d) E(XY) = E(X)E(Y)
(e) cov(X,Y) = 0
(f) corr(X,Y) = 0
Proof: We defined (a) in terms of (c). If we can show (f), then we have shown (b). So we need prove only (d),(e) and (f)
(d) E(XY) = ∫∫ x y p(x,y) dx dy = ∫∫ x y p(x)p(y) dx dy = [∫x p(x) dx] [∫y p(y) dy] = E(X)E(Y) .
(e) cov(X,Y) ≡ E((X-μx)(Y-μy)) = ∫∫ (x-μX) (y-μy) p(x,y) dx dy = ∫∫ (x-μX) (y-μy) p(x)p(y) dx dy
= [∫ (x-μX) p(x) dx] [∫ (y-μY) p(y) dy] = [0][0] = 0
where for example
[∫(x-μX) p(x) dx] = ∫x p(x) dx – μX∫ p(x) dx = μX - μX * 1 = 0.
(f) corr(X,Y) ≡ cov(X,Y) / [σ(X) σ(Y) ] = 0 / [σ(X) σ(Y) ] = 0 QED
Fact: If {X,Y} are uncorrelated, and if {Y,Z} are uncorrelated, {X,Z} may be correlated. (G.32)
Proof: Knowing that p(x,y) = p(x)p(y) and p(y,z) = p(y)p(z) tells us nothing about p(x,z). It is easy to think of trivial examples of this fact. Maybe X = Z so corr(X,Z) = var(X)/ [σ(X)]2 ≠ 0.
Fact: If {X,Y} and {Y,Z} and {X,Z} are all uncorrelated pairs, X Y and Z might not be statistically independent. (G.33)
Proof: Knowing about the uncorrelated pairs says nothing about p(x,y,z). That is to say, knowing that p(x,y) = p(x)p(y) and p(y,z) = p(y)p(z) does not imply that p(x,y,z) = p(x)p(y)p(z).
This concludes our brief review of probability theory, and we now continue in our examination of experiments associated with random variables.
(d) Ensemble Experiments
We consider now a new kind of "experiment". We acquire N dice which we shall refer to as an ensemble of dice. We assume they are all "loaded" differently, perhaps with implanted weights. For each die, we perform the single-roll experiment described above for which the outcome lies in the sample space S = {1,2,3,4,5,6}. After doing these N experiments, the N dice are left lying on the green felt of a craps table, each in its final experimental state. We then survey all N dice and take note of each one's face-up number, and from that data we construct a distribution. Since these dice are all slightly different, the distribution will likely differ from that shown earlier in Fig G.1. Perhaps we get this:
Fig G.8
As a computer algorithm, here is this new "ensemble experiment":
Acquire an ensemble of N dice (perhaps each is loaded differently)
For each die in the ensemble, carry out the single die experiment with outcome in S = {1,2,3,4,5,6}.
When the experiments are done, survey the resulting data and construct a distribution.
If the dice were identical, then this ensemble experiment would be the same as just sequentially rolling the same die N times and writing down the outcome of each roll. But the main point is that we assume the dice are not all the same.
Now we carry out an analogous ensemble experiment :
Acquire an ensemble of N babies (they are likely all different)
For each baby in the ensemble, carry out some experiment with outcome in some sample space S.
When the experiments are done, survey the resulting data and construct a distribution.
The experiment we have in mind is simply to let each baby grow up into an adult and then we treat some parameter of that adult as our random variable. Adults generally don't have face-up numbers, but they do have other "parameters" such as mass, height, and number of children. For each adult, the "experiment" was "growing up", analogous to rolling one die in our previous ensemble experiment. It is convenient to let nature do these experiments for us, so in practice we just assemble some group of adults into our ensemble, and then we are left with just the last item above :
When the experiments are done, survey the resulting data and construct a distribution.
Here for example we carry out three ensemble experiments with the same ensemble of people. Each ensemble experiment uses a different sample space: weight, height, number of children. and each deals with a different random variable: W, H and N. For each experiment we obtain a distribution as shown:
Fig G.9
In all three experiments, the outcomes are real numbers, so there is no need for any functions to translate from non-numerical outcomes to real numbers, as we required in for the {heads,tails} sample space. These experiments are all analyzing historical data, nor current "chance events" like rolling dice.
In the above scenario, sometimes the ensemble of people is called "the sample space" which is then a totally different meaning of that phrase, so we won't use it.
(e) Experiments rolling two dice at the same time
In the next set of experiments, we roll two differently weighted dice at the same time and examine particular outcomes (that is to say, we examine particular random variables, each taking real values in its numerical sample space). These two dice are not only weighted differently, but the embedded weights are magnets, which cause the two die to interact with each other during a roll. Here we get to apply various facts from the brief review of probability theory of section (c) above.
In Experiment #1 we just look at the face up number of the X die and that number is x. The sample space as usual is S = {1,2,3,4,5,6} and there will be some pmfX(x). Review: In this notation, the subscript indicates the random variable X, and the argument x is a value that random variable can take. Another notation is pmf(X = x), and the most compact notation is that of (G.5) which is just p(x) where one must remember what the random variable is. It is suggested by the letter used as argument: p(x)= pmfX(x) = pmf(X = x).
In Experiment #2 we do this for the Y die. It has the same sample space, but a different p(y) = pmfY(y).
In Experiment #3 we look at z = x+y with its random variable Z = X + Y. The sum of two random variables is a random variable according to (G.16) or (G.21). The numerical sample space for this third experiment is {2,3..., 11,12}, since there are no other possible outcomes for the sum of two face-up numbers on two dice. According to (G.17a), the probability of Z taking value z in this sample space is given by pmfZ(z) = Σzx,y pXY(x,y), or in compact notation, p(z) = Σzx,y p(x,y), where the sum includes only those x,y values such that x + y = the z inside p(z) on the left. If the two die were differently weighted but contained no magnets, we would have p(z) = Σzx,y pX(x)pY(y) since the two dice are "independent" or "uncorrelated" as in (G.31). If the two dice were identical, then p(z) = Σzx,y pX(x)pX(y) where the two pmf's are the same. If the two dice are "fair dice" (unloaded), then pmfX(x) = pX(x) = 1/6 for any x and then we find that p(z) = Σzx,y (1/6)(1/6) = (1/36) Σzx,y 1. We then have a classic dice problem where we can enumerate the terms in the sum Σzx,y 1 as follows
2 1,1 = 1
3 1,2 + 2,1 = 2
4 1,3 + 3,1 + 2,2 = 3
...
7 1,6 + 6,1 + 5,2 + 2,5 + 3,4 + 4,3 = 6
...
12 6,6 = 1
Below, pmfX(x) is for Experiment #1, and pmfZ(z) for Experiment #3. In each experiment, we roll a pair of identical fair dice many times and obtain these distributions.
Fig G.10
We now do a few quick hand calculations to exercise some of our basic probability facts. Notice ahead of time that
Σx 1 = 6
Σx x = (1 + 2 + 3 + 4 + 5 + 6) = 21
Σx x2 = (12 + 22 + 32 + 42 + 52 + 62) = 91 .
The expected value of Z for Experiment #3 is given by (G.23),
E(Z) = E(X+Y) = Σx,y (x+y) p(x,y) .
In the case of Fig G.10 where p(x,y) = [p(x)]2 = (1/6)2 = (1/36) we have
E(Z) = (1/36) Σx,y(x+y) = (1/36) [ (Σx x)(Σy1) + (Σx 1)(Σyy) ] = (1/18) (Σx x)(Σy1)
= (1/18) (21) (6) = (1/3)(21) = 7 = μZ
in agreement with the symmetric distribution on the right in Fig G.10. Next we compute,
E(Z2) = Σx,y (x+y)2 p(x,y) = (1/36) Σx,y [ x2 + y2 + 2xy ]
= (1/36) [ 2 (Σxx2)(Σy1) + 2 (Σxx)2 ] = (1/18) [91* 6 + 212] = 329/6 = 54.83.
The variance and standard deviation are then given by (G.30).
var(Z) = E(Z2) - {E(Z)}2 = 329/6 – 49 = 35/6 = 5.833
σ(Z) = = 2.415 (G.34)
This last result seems in line with a visual inspection of Fig G.10.
Finally in Experiment #4 we consider the outcome z = x*y = xy. Things are similar to the above, but now the notation Σzx,y means the sum over x,y values such that the product of x and y is z. In the context of (G.23), we now have f = xy whereas in experiment #3 we had f = x+y. We could of course study the situation with any reasonable function f. For f (x,y) = xy we can write:
pmfZ(z) = Σzx,y pmfXY(x,y) dice are different and weighted with magnets
pmfZ(z) = Σzx,y pmfX(x) * pmfY(y) dice are different but no magnets
pmfZ(z) = Σzx,y pmfX(x) * pmfX(y) dice are identical
pmfZ(z) = Σzx,y (1/6) * (1/6) = (1/36) Σzx,y 1 dice are identical and "fair"
For this Experiment # 4 the sample space is all possible products S = {1, 2, 3......36} with certain values missing. For "identical and fair" we make our list using our new meaning of Σzx,y :
1 1,1 = 1
2 1,2 + 2,1 = 2
3 1,3 + 3,1 = 2
4 1,4 + 4,1 + 2,2 = 3
5 1,5 + 5,1 = 2
6 1,6 + 6,1 + 2,3 + 3,2 = 4
7 = 0
8 2,4 + 4,2 = 2
...
36 6,6 = 1
We ask Maple to create pmfZ(z) for this Experiment #4:
Fig G.11
We then repeat the calculations done for Experiment #3:
E(Z) = E(XY) = Σx,y (xy) p(x,y) = (1/36) Σx,y (xy) = (1/36)[ (Σxx)2 = 212/36 = 49/4 = 12.25 = μZ
E(Z2) = Σx,y (xy)2 p(x,y) = (1/36) Σx,y x2y2 = (1/36) [ (Σxx2)2] = (1/36)(91)2 = 8281/36 = 230.03
var(Z) = E(Z2) - [E(Z)]2 = 8281/36 - (49/4)2 = 11515/144 = 79.97
σ(Z) = 8.942 // standard deviation (G.35)
(f) Experimental determination of discrete distribution functions
In practice, one has some list of experimental results (outcomes converted to real numbers if not already real numbers) and one puts the results into "bins" to obtain a distribution. Here we want to be more explicit about what this means.
For a single variable x which is the face-up value of a die, here is how we would compute p(x) after collecting data rolling the die J times :
which we could plot as a bar chart of the type shown in Fig G.1.
Symbolically we might write this as
p(n) = (1/J) Σj=1J ( xj = n ) n = 1,2,3,4,5,6 . (G.36)
The idea is that for each n, we count the number of times that ( xj = n ) is true.
For two variables, things are a bit more complicated. Now we have a set of J pairs {xi,yi} as our collection of data, so that
p(n,m) = (1/J) Σj=1J ( [xi,yi] = [n,m] ) n,m = 1,2,3,4,5,6 (G.37)
where [...] refers to an ordered sequence. Another way to write this would be
p(n,m) = (1/J) Σj=1J [(n = xi) and (m = xj)] n,m = 1,2,3,4,5,6 (G.38)
Here is a sample Maple program to carry this out
The result can be visualized as a 3D bar chart
Fig G.12
A distribution function like p(a,b,c,d,e) is a 6D bar chart, not easy to display.
(g) Experiments with Sequences of Pulse Train Amplitudes
Our pulse trains have a set of amplitudes yn which form a sequence. If the sequence has N elements, then a sequence can be written [y1, y2.....yN]. When many pulse trains are generated in some line code, one finds that the values of yn in position n of the sequence can vary, and that there is some probability distribution associated with the parameter yn. Thus, the position n in the pulse train or sequence is associated with a random variable we must call Yn which takes values yn.
Here is our first experiment. We have an Apparatus which generates sequences of length N according to some set of rules implemented within the Apparatus (perhaps these are the rules for generating AMI linecode sequences, and the AMI encoder inputs random input streams). Every sequence it generates is "legal" according to its rules. Earlier we discussed an ensemble experiment in which M dice were rolled one at a time and were left on the craps table for study and from what we saw on the table, we were able to construct a distribution function. The ensemble was M dice. In our current ensemble experiment, we let the Apparatus crank away and generate an ensemble of M sequences each of length N, and these M sequences are left sitting on the same craps table for us to inspect. This is our "preliminary" ensemble of sequences.
The notation Yn leads to some confusion. When we had a random variable H which took values h, we might have enumerated the set of h values of the sample space as hi. To maintain clarity, if we have random variable Yn which takes values yn, we should enumerate those values as (yn)i. Here, yn is the name of a parameter, just as h was the name of a parameter. Thus, to write down a specific sequence "i" we really should say
[(y1)i, (y2)i.....(yN)i] = sequence "i"
or (G.39)
[y1(i), y2(i)..... y2(i)] = sequence "i"
where (y2)i = y2(i) is some specific symbol value like "5". We used the latter notation when talking about an ensemble of pulse trains in Section 35 and in Appendix D.
The Useful Ensemble. To construct a "useful" ensemble of pulse trains, we first imagine generating some number M of pulse trains of length N from our Apparatus as described above. We then take this ensemble of M pulse trains and considerably increase the size of the ensemble by including in it all cyclic permutations (rotations) of the M pulse trains, so now the enlarged ensemble contains I = M*N pulse trains.
Depending on the rules of the Apparatus generating the pulse trains, these rotated pulse trains might be "illegal" at the boundary where the pulse train wraps around on itself. For example with N = 5 we obtain this sequence plus its rotations,
A B C D E
B C D E|A
C D E|A B
D E|A B C
E|A B C D . (G.40a)
Here the seam is marked by a bar | . If the seam were "illegal", we could probably repair the seam by replacing A with some other symbol Q to get a legal sequence of symbols,
A B C D E
B C D E Q
C D E Q B
D E Q B C
E Q B C D (G.40b)
Perhaps several symbols near the seam would have to be "repaired" in this manner, but we shall assume just one as shown above. Notice that each column has at most one bad symbol Q, and that is 1/N of the symbols in the column. If N = 100, then at most 1% of the symbols in a column are wrong after adjustment. If a 2-symbol repair were needed, then at most 2% in a column would be wrong.
We imagine doing this for all M of our initial pulse trains to get our final set of I = M*N pulse trains. In this "useful ensemble", in each vertical column, 1/N of the symbols might be wrong. We of course have in mind that N is very large, but our examples will always have N very small, such as N = 5.
Now we associate each column of our ensemble of sequences with a random variable Yn. For (G.40),
Y1 Y2 Y3 Y4 Y5 (G.41)
Fact: The expected value <yn> ≡ E(Yn) is independent of n. (G.42)
Proof: Each column has at most 1 bad symbol like Q for each initial ensemble sequence. So if there are now I = M*N pulse trains in the ensemble, then in each column there will be at most M bad symbols. When the expected value is computed by summing the elements of the column n and dividing by I,
E(Yn) = (1/I) !Syntax Error, Iyn(i) = (1/I) !Syntax Error, I (yn)i // two notations for the same thing
one finds that the sums for different n are almost identical with an error on the order of 1/N. In our example the first column sums to A+B+C+D+E whereas the other columns are Q+B+C+D+E . Of course the exact error depends on what symbol Q was used to do the repair and what the palette of symbol values is, and so on, but the general order is 1/N. As N is made large, this error approaches 0 and then all the columns have the same expected value of their Yn. QED
Now recall the probability distribution for our N=5 ensemble, written in two ways as in (G.5),
p(a,b,c,d,e) ≡ p(Y1=a,Y2= b,Y3= c,Y4= d,Y5= e) . (G.43)
This is the height of a specific "bar" in the 6D bar chart that is p(y1,y2,y3,y4,y5), see Fig G.12
Fact: For our useful ensemble, the probability distribution p(a,b,c ... ) has cyclic symmetry. (G.44)
Example: p(a,b,c,d,e) = p(e,a,b,c,d) = p(d,e,a,b,c) = p(c,d,e,a,b) = p(b,c,d,e,a)
Proof: This might seem obvious, but a formal proof is warranted. We show this for the case N = 5; the general case should then be obvious. Using the notation introduced in (G.37) we can imagine evaluating p(a,b,c,d,e) this way, where [....] indicates an ordered sequence of numbers,
p(a,b,c,d,e) = (1/I) Σi=1I ( [ a,b,c,d,e] = [ (y1)i, (y2)i, (y3)i, (y4)i, (y5)i ] ) .
Since our ensemble contains all cyclic rotations of [ (y1)i, (y2)i, (y3)i, (y4)i, (y5)i) ], we don't change the above sum by replacing this [...] with [ (y2)i, (y3)i, (y4)i, (y5)i, (y1)i) ], we merely re-order the sum. In our example above, we can rewrite the 5 sequences on the left in this fancier notation,
(y1)i (y2)i (y3)i (y4)i (y5)i
Y1Y2Y3Y4Y5
α A B C D E (y1)1 (y2)1 (y3)1 (y4)1 (y5)1
β B C D E A (y1)2 (y2)2 (y3)2 (y4)2 (y5)2
δ C D E A B (y1)3 (y2)3 (y3)3 (y4)3 (y5)3
γ D E A B C (y1)4 (y2)4 (y3)4 (y4)4 (y5)4
ε E A B C D (y1)5 (y2)5 (y3)5 (y4)5 (y5)5
We now rotate the above columns one column to the left to get
(y2)i (y3)i (y4)i (y5)i (y1)i
Y1Y2Y3Y4Y5
β B C D E A (y2)1 (y3)1 (y4)1 (y5)1 (y1)1
δ C D E A B (y2)2 (y3)2 (y4)2 (y5)2 (y1)2
γ D E A B C (y2)3 (y3)3 (y4)3 (y5)3 (y1)3
ε E A B C D (y2)4 (y3)4 (y4)4 (y5)4 (y1)4
α A B C D E (y2)5 (y3)5 (y4)5 (y5)5 (y1)5
As it most easily seen on the left, this rotation by one column has just reordered the ensemble sequences from the original order α,β,γ,δ,ε to the new order β,γ,δ,ε,α. The same set of samples is included in the sum above regardless of how they are ordered.
Consider then our starting expression,
p(a,b,c,d,e) = (1/I) Σi=1I ( [ a,b,c,d,e] = [ (y1)i, (y2)i, (y3)i, (y4)i, (y5)i] ) .
Note that c is a specific number, as is (y5)3 . Neither of these is a variable. Now suppose the ensemble has the property just noted above, that we can reorder the sum this way,
p(a,b,c,d,e) = (1/I) Σi=1I ( [ a,b,c,d,e] = [(y2)i, (y3)i, (y4)i, (y5)i, (y1)i] ) .
The equality of the two sequences of numbers shown in the truth test [...] = [...] does not change if we alter each sequence in the exact same way, so rewrite again as
p(a,b,c,d,e) = (1/I) Σi=1I ( [ e,a,b,c,d] = [(y1)i, (y2)i, (y3)i, (y4)i, (y5)i,] ) .
But the expression on the right is the definition of p(e,a,b,c,d). Thus we have shown that
p(a,b,c,d,e) = p(e,a,b,c,d) .
We can then show in the same way that p(e,a,b,c,d) = p(d,e,a,b,c) and so on, so in the end all cyclic permutations (rotations) of p(a,b,c,d,e) are equal for our useful ensemble. QED
Comment: In the special case that all the sequence random variables Yn are statistically independent, we can write
p(y1,y2,y3,y4,y5) = p(y1) p(y2) p(y3) p(y4) p(y5) .
In this special case it is obvious that p(y1,y2,y3,y4,y5) is symmetric under cyclic permutations, as it is under any permutation. A point to note is that we are not assuming statistical independence of the Yn and that therefore there can exist "correlation" between pairs of random variables in our sequences (pulse train amplitudes). Nevertheless, in this general case p(y1,y2,y3,...) our useful ensemble has cyclic symmetry, and this ensemble is a realizable ensemble from our experiments for large N with only some very small error. As N→∞, that error goes to 0.
Fact: The probability distribution p(Yn = x) is independent of n. (G.45)
Proof by Example: When we write p(Yn = x) = p(x), then p(x) is deceptively independent of n, but we must show this in the true full notation where n dependence is not concealed. So consider using our N= 5 example, where we use (G.7) to obtain p(Y3 = x) by summing over all the "other variables" of the full 5th order joint probability distribution p(y1,y2,y3,y4,y5),
p(Y3 = x)
= Σy1,y2,y4,y5 p(Y1=y1,Y2= y2,Y3 = x,Y4= y4,Y5= y5) // definition of p(Y3 = x)
= Σy2,y3,y5,y1 p(Y1=y2,Y2= y3,Y3= x,Y4= y5,Y5= y1) // rename dummy summation variables
= Σy1,y2,y3,y5 p(Y1=y2,Y2= y3,Y3= x,Y4= y5,Y5= y1) // trivially reorder sums on Σ
\ \ \ \
= Σy1,y2,y3,y5 p(Y1=y1,Y2= y2,Y3= y3,Y4= x,Y5= y5) // use cyclic property of p function
= p(Y4 = x) // definition of p(Y4 = x)
Thus we have shown that p(Y3 = x) = p(Y4 = x) and either by iterating or using other cyclic permutations in the middle step above we conclude that all the p(Yn = x) are equal and thus p(Yn = x) is independent of n. As usual, this is for our "useful distribution".
Corollary 1: <yn> ≡ E(Yn) is independent of n, since E(Yn) = (1/I)Σx x p(Yn = x) and p(Yn = x) is independent of n. Thus we have an alternate proof of (G.42) above. (G.46)
Corollary 2: <yn2> ≡ E(Yn2) is independent of n, since E(Yn2) = (1/I)Σx x2 p(Yn = x) and p(Yn = x) is independent of n. (G.47)
Corollary 3: <f(yn)> ≡ E(f(Yn)) is independent of n, since E(f(Yn)) = (1/I)Σx f(x) p(Yn = x), etc.
See (G.22). (G.48)
Fact: The probability distribution p(Yn = x, Ym = z ) depends only on m-n mod N. (G.49)
Proof by Example: Using the same example as above, and again applying (G.7), we find that
p(Y2 = x, Y4 = z) // m-n = 4-2 = 2
= Σy1,y3,y5 p(Y1=y1,Y2= x,Y3 = y3,Y4= z,Y5= y5) // definition of p(Y3 = x)
= Σy3,y5,y1 p(Y1=y3,Y2= x,Y3 = y5,Y4= z,Y5= y1) // rename dummy sum indices
= Σy1,y3,y5 p(Y1=y3,Y2= x,Y3 = y5,Y4= z,Y5= y1) // trivially reorder sums on Σ
\ \ \ \
= Σy1,y3,y5 p(Y1=y1,Y2= y3,Y3 = x,Y4= y5,Y5= z) // use cyclic property of p function
= p(Y3 = x, Y5 = z) // m - n = 5-3 = 2
In general, one has
p(Yn = x, Ym = z ) = p(Yn+k = x, Ym+k = z )
again for our useful ensemble.
Corollary: <ynym> ≡ E(YnYm) for m ≠ n depends only on m-n mod N. (G.50)
E(YnYm) = (1/I) Σx,z (xz) p(Yn = x, Ym = z) = (1/I) Σx,z (xz) p(Yn+k = x, Ym+k = z)
= E(Yn+kYm+k) .
Conclusion: When we experimentally create a huge "statistical ensemble" of pulse trains which have a very long length N, this ensemble is statistically very similar to our "useful ensemble" described above. Every possible legal sequence of amplitudes will appear many times in the ensemble, and all the cyclic versions of each sequence will appear (ignoring the tiny seam corrections) on an equal footing. As N gets very large, the notion of cyclic permutation is more associated with a shift of the pulse train, when we observe some finite window near the center of the pulse train. For N = ∞, the shift notion is all that remains, and there are no seams to worry about. For such pulse trains, we have established these important facts:
<an> = independent of n // general
<an2> = independent of n // general
<anan+k> = independent of n, dependent in general on k // general
<anam> = <an><am> = <an>2 = independent of n // uncorrelated only (G.51)
_____________________________________________________ summary ______________________
Appendix G: Random Variables, Probability Theory and Pulse Train Amplitudes
Subsection (a) provides an opening definition of a "random variable" as being a parameter associated with a probability distribution. The terms experiment, outcome, sample space and event are defined and two simple experiments considered: rolling a die, and spinning a spinner. The corresponding pdf and pmf distributions are discussed.
Subsection (b) refines the definition of "random variable" as a mapping from the sample space to the real numbers, and the coin toss experiment is used as an example. Drawings show examples of discrete and continuous samples spaces and the random variable mappings.
Subsection (c) contains a brief review of basic probability theory with an emphasis on notation. The notions of statistical independence, correlation and normalization are discussed. Some Math Lemmas are presented relating the ordering of terms in a sum Σx,y in the context of some function z = f(x,y) by first summing over a surface of constant z, then over different z values, so Σx,y = Σz Σsurface(z), These lemmas are needed in the following discussion of random variables defined as functions of other random variables. After this, the usual suspects of probability theory are rolled out for inspection: expected values, mean, covariance, variance, standard deviation and correlation.
Subsection (d) treats "ensemble experiments" such as rolling N loaded dice, and a transition is then made to the analysis of random variables associated with an ensemble of people.
Subsection (e) treats the rolling of two weighted, magnetically interacting dice, after which special cases are considered.
Subsection (f) gives examples of the computation of probability mass density functions from experimental data.
Subsection (g) applies the notions of random variables and probability theory to the sequences which are the amplitudes of pulse trains. A special "useful ensemble" is constructed for which the full probability density function is cyclic in its arguments, and it is shown how this ensemble matches a real world ensemble for long sequences. It is shown that <ym> is independent of m and <ymym+k> can only depend on k.
___________________________________ compact overview ----------------
Appendix D: How α and β are related to μ and σ (6 p)
Appendix D reviews some basic probability theory and relates the α and β parameters of the Section 35 uncorrelated spectral power density formula to statistical properties μ and σ of pulse train amplitudes. The probability theory part is significantly expanded in Appendix G.
Subsection (a) discusses random variables X and Y and their expected values like μx = E(X) = <x> and E(XY). Certain statistical measures are defined and related to each other. The concept of X and Y being independent (uncorrelated) random variables is explored.
Subsection (b) translates the results of subsection (a) into the context where X = Ym and Y = Yn where Ym is the random variable associated with position m in a pulse train, and whose values are the pulse amplitudes ym .
Subsection (c) shows that for an uncorrelated pulse train the following facts are true for the coefficients α and β used in Section 35 on spectral power density,
α = μ2 β = μ2 + σ2 (β-α) = σ2 μ ≡ <ym> σ2 = <ym2> - <ym>2
These relations provide a useful interpretation of the parameters α and β in terms of the statistics of the pulse train amplitudes and allow comparison of our results to those of other sources.