probability theory
DOCX · 1.0 MB
Open DOCX file
Informal notes by Phil dated 7.24.13, prompted by confusion over the phrase "random variable" in his Spectral and Scrambler writings. He reviews Chapter 2 of Proakis's Digital Communications and the Wikipedia definition, then works through examples (height of people, temperature on a surface, coin toss, dice) with sample spaces, pdf and PMF. He also tries to map this onto averaging along a sequence versus across an ensemble. Only the first part of the text was seen.
AI-written summary; may contain errors.
Extracted text (machine-read; may contain errors)
Probability Theory PhL 7.24.13
1. What is a Random Variable?
I have an appendix in Spectral on this subject, and another one in Scrambler (at least right now), and I keep mentioning the phrase "random variable" without knowing what it really means. I have two different dimensions of averaging, horizontal across a sequence or pulse train, and vertically for an ensemble of pulse trains, and these get mixed together.
I downloaded Proakis 4th Ed Digital Communications (2000). I have to add a text index to it as an evening task sometime.
There is now a 5th Ed (2007) which I don't have.
Let's now review his Chapter 2 on Probability stuff.
2.1 Probability. Rolling a die, possible outcomes, mutually exclusive events (die 2 or 4, or die 1 3 or 5), union, events
Then joint probability idea, then conditional as I have it.
Page 21. Statistical independence. He says if P(A,B) = P(A)P(B) then these events are statistically independent. (which seems to conflict with what I have said?) He just says you can "extend" the idea of SI to 3 or more factors
A little hazy.
2-1-1 Random Variables Etc.
Sample space = set of values event can take I think.
Random variable is a FUNCTION mapping S to reals. X(s)
For coin toss, X(s) = 1 if heads, X(s) = -1 if tails, say.
Then extends idea to a continuous sample space.
and then later
Let's stop right here and try to translate this into my little scrambler world.
Imagine a horizontal sequence of values. If it is just a set of fixed numbers, it would seem there is no probability theory involved at all.
sequence = ...1,0,3,-1,2,....
In such a sequence there is some mean value like
<an> = Σn an N(an)/N = Σn an p(an)
Yes, you can say that p(an) is the probability within this sequence that a symbol has value an. But so far it is not a "random sequence", it is just a sequence.
You could think of this as an ensemble horizontally of single symbol events. Like each symbol is a coin toss. Then the fixed sequence above is just one possible "experiment". If you write down a list of experiments, then you have my vertical ensemble of sequences.
Go back to the single sequence as a series of separate 1-symbol events. Then you can think of the horizontal sequence as an ensemble as well, an ensemble of single sequence events. Just saying this does not make those events "independent". The events are n = 1,2,3.... and values are an. Then you can talk about E(an) in the horizontal sense. Once you have some kind of ensemble, then you can apply all of my theory to it. But here we are talking ensemble of single-symbol events. Then my line above makes sense
<an> = Σn an N(an)/N = Σn an p(an) = E(an)
Now what is the "random variable" in this situation? Well, the sample space S is perhaps GF(p) or some other discrete set of allowed values, perhaps just {1,0}.
Go back to random variable X(s) for coin toss
Here is my plot of this function whose domain is discrete.
I do NOT understand what this means. It seems to mean nothing. There is no "probability" involved in this random variable definition. For GF(22) you could write this as
where maybe the horizontal axis represents abstract elements of GF(22) while the vertical axis is real numbers in this case integers.
Basically it is a way to get from whatever your "events are " to a set of numbers. Here is another picture
But in MY case, we have s1 = 0, s2 = 1 etc, and the mapping is then from reals to reals and it is just an identity mapping.
X(s) = s or change variable name X(x) = x
So my "random variable" is the function X(x) = x. Very boring. Maybe write this X(an) = an. So this seems strange. How can you define a "random variable" to be a function? It does not make sense .
Go to another source. Wiki has a completely different definition
So at least they define it as a variable! But later they say
So here the set S is a set of "outcomes", still the sample space. In their example, the sample space might be a set of people including Phil. So Phil is a possible "outcome" in this space. The "variable" might be the height of Phil which is a real number, though Phil is not a real number. So here the random variable is the height of a person. The randomness enters because you have select a person from your ensemble of people. Then you might ask what is the probability that for this ensemble height is in range h to h+dh and that would be a pdf.
Maybe I like to thing of the "outcomes" as the "possible events". Then Phil is a possible event in the group of people. The events need not be numbers.
They go on to make a good comment:
For any range of the "random variable" x (height of person), you have to be able to associate any region of the real line (x1,x2) with a well-defined subset in the sample space, that is, with a particular subset of the 10,000 people in the Sample Space.
But then they refer to the subsets as "events". So an event could be a group of people.
The random variable X(s) is like X(Phil) = 6'1".
A PDF for a discrete value is called a Probability Mass Function or PMF of the values are non-negative integers like the number of children of a person. Then X(phil) = 0. That type of random variable is called a PMF. Fine.
Wiki goes on to say:
So values of height h become the new sample space and the space of people is "suppressed".
Let's try now to interpret the coin toss discussion.
Since the random variable maps the two events heads and tails into a set of non-negative integers, the probability function gets that special name PMF. Fine. Now for the first time we have some kind of "probability function" associated with heads and tails. It need not be 1/2 and 1/2 as shown. I would probably just call this thing a pdf anyway, a discrete PDF.
So again we have this indirection of outcomes → real numbers → probability. Y(w) is the random variable here.
The meaning of a Random Variable
Before attempting a definition, we give a few examples.
One example of a random variable is the height h of a person, where the person is one of many people in some ensemble (set) of people, each of whom has some well-defined height. This ensemble of people is an example of a sample space and the individual people are the samples. Any subset of people forms an event, so {Joe} is an event, and {Joe, Alice} is a different event. This h is a "real variable" because it can take various real values. In this sample space the people are not real numbers, they are people, but h is a real number. In this example, the random variable h is continuous whereas the sample space is discrete.
If we measure the height of all the people in the ensemble, we end up with some probability density function (pdf) for height, where the area under the red curve is 1 :
This pdf would be different for a set of NBA basketball players versus a set of horse racing jockeys. The pfd meaning is that the probability of someone in the sample space having a height between h and h+dh is given by pdf(h)dh.
Clearly height in this situation can be regarded as a mapping from the sample space S to the real numbers R. We might call this mapping H, so H:S→R. This mapping is a function whose domain is S and whose range lies in R. We write then h = H(s), with an example 6'1" = H(Joe). For a discrete sample space one could place the samples of the sample space in some sequential order and make a picture of H(s) like so
It is from the function H(s) that one computes pdf(h) by tossing people into narrow height bins, and then plotting the bin counts divided by the total number of people.
Another example of a random variable would be temperature at all points on a 2D surface. The sample space here is the continuous 2D surface, and a sample is a point (x,y) on that surface. Temperature is a real number, but the samples are not real numbers, they are 2D vectors. The mapping is t = T(s) where s = (x,y), so t = T(x.y) and we could plot this as a 2D surface lying over the x,y planar sample space.
Now we attempt a definition.
Definition: Random Variable. A random variable x over S is some real quantity x which has a probability distribution with respect to some particular sample space S. We have seen above where the word "variable" applies, since our "real quantity x" is a variable which takes some set of real values. The word "random" is associated with the fact that the variable x doesn't have some fixed value, but has a value that is described by a pdf. There is some probability distribution pdf(x) associated with variable x which can be computed from x = X(s). That distribution might have a Gaussian shape and if it did, one might call x a "Gaussian random variable". But the raw definition does not specify any particular probability distribution, it just requires that there be one. Nor does the definition specify a particular sample space.
In both the height and temperature examples the random variables h and t were continuous so the pdf(x) was a continuous function.
One can of course have discrete random variables and the classic example involves tossing coins. One can consider a sample space S not of people, but of coins which have just been tossed. Just as a person has a property called height, a tossed coin has a property called heads or tails. A difference is that this property is not a real number, so in order to talk about a random variable, we have to associate perhaps heads with 1.7 and tails with 2.6 and call this real variable h. As before, h = H(s) and we can write down the value of h for each tossed coin in the sample space, and from this we can construct a pdf which might look like this:
Here either variable h or function H(s) would be called a random variable. This h is a real quantity which has a probability distribution relative to some sample space S (a set of tossed coins). This is a discrete random variable. A discrete pdf function is often called a pmd meaning a probability mass distribution, suggested by the idea that the discrete random variable values are like point particles with all their mass concentrated at a point, rather than like distributed matter having a continuous mass density.
If the coins were equally weighted on both sides, and if S contained 106 such tossed coins, we would expect the heights of the two red lines to be the same and to be 1/2. Maybe the coins are asymmetric so as to produce the 2/3, 1/3 distribution shown, or perhaps the tosses were not "honest". One might imagine an honest toss of an equally balanced coin, and one might call this a "random coin toss". But even it tosses are not random in this equal probability sense, h = H(s) is still called a "random variable" because, no matter how the coins were tossed, there does exist some H(s) and some pdf(h). This seems to be a point of confusion regarding random variables. Random just means "there is some associated probability distribution".
****************************
How does a rolling die fit into this framework? The sample space could be a set of N dice that have just been rolled. A die is not a real number. The random variable could be the number facing up. That variable, like height, is already a real number. We have some random variable n = N(s) which enumerates how the dice have all come out. There is some associated pdf(n) we could construct. It might be a set of 6 bars all of height 1/6 or some other set of bar heights for an uneven die.
Now where would the sum of two die values fit in? Ouch, not working very well. My model is not really very good. I think the numbers {1,2,3...6} should be the sample space, not the die. And for coin toss, the sample space should be {heads,tails}. And what about the word outcome? Is the person's height an outcome? The wiki site is very fuzzy on this. I would like to fit many different examples into the a model that works for all.
Let's now look at a different source.
Here "outcomes" are elements of the "sample space". But still very vague. Author defines an "event" as being a subset of S which makes to an interval on the reals. Each interval maps back into some subset we know that, but not every subset maps into an interval. The ones that map into intervals are "events". This seems to me to be a rather strange definition of an "event". Why is that word chosen? We need examples. These authors all went to get on to things like CDF and PDF, they don't want to define their objects. A common problem.
This does not fit well with the idea of sample space = set of people. If we have one die, then the sample spacing being {1.2.3.4.5.6} does fit with the above. It is the set of outcomes of an experiment. There is only one die. Continuing
The "experiment" involves one die. The sample space Ω is as shown. An event is shown. The random variable here might be n = N(ω) = ω, pretty simple. There is some pdf(ω). "A random variable represents the outcome of an experiment. " Strictly N(ω) is the random variable here.
The "experiment" is two coin tosses. Random variable X represents the outcome of this experiment. The same space is not Ω = {H,T} it is sort of {H,T}{H,T} which author writes in several ways. So here we get away from the idea that Ω = {H,T} for any number of coins being tossed.
An "event" here might be getting 2 heads. This is the same as HH in Ω. But the "event" of getting only one head is HT + TH. An event might be a single die rolling an even number.
The "experiment" I guess is an election. The three outcomes are "A wins", "B wins", or "C wins", and that is what is meant by Ω = {A,B,C}.
This all seems reasonable.
So how do I deal with the people and the heights? What is the experiment? It is like an experiment that has already been done and we are just seeing the outcome? Each person is like a die. So the experiment is a room of 2 people, the outcomes are what? This fits very poorly. If people could only have 6 heights, then the same space would be {123456} {123456} like the two coin tosses.
The experiment is what is explained in the first sentence.
What is this book I am looking at? Grinstead is math prof at Swarthmore College. Snell is was at Dartmouth math, he died in 2011.
I inadvertently downloaded this entire book! So I now save it as perhaps my only book on this subject. It has a nice history section at the end of Chapter 1.
Here experiment = roll four dice, outcome Ω = possible sums of all four = {4,5.....36}, and this would be the sample space.
I have done a reasonable scan of the first chapter on "discrete probability distributions", and now will look into Chapter 2 on continuous ones. He opens with the "spinner" but I don't get it! I think the possible outcomes here are an angle 0 to 360 which he regards as (0,1). Fine. The outcome is continuous.
He then does the usual pdf stuff. But he never talks about an experiment with two spinners. I guess the sample space would be a subset of Rn as he says above. So for a room with 10 people, the Ω would be a chunk of R10 with each real axis between 5 and 6 feet?
He gets on to the idea of "a function of a random variable".
Another source now: http://www.cut-the-knot.org/Probability/SampleSpaces.shtml
Very confused! First, the "sample space" is a set of humans and height is a random variable. But then author says "sample space" is line segment [40,272] cm. So which is it? The second is more in line with a set of outcomes of the experiment of measuring height.
This reinforces the idea of samples space = outcome = height.
Notice here the analogy between
experiment = flipping a coin outcome = heads, tails
experiment = picking a person from an ensemble outcome = height
Ensemble is called a "population"
So I think my first impression of people = sample space was just plain wrong.
OK, after this I think I am ready to write things up, manana.
Question: Compare again the roll of a die experiment to the height experiment. Both have outcomes that are already real numbers, so we don't need "functions" to map outcomes to numbers. If the die is a "fair die", then P = 1/6 for each number and that is the pdf or pmf as it is called. But for the height deal, if we make a "fair selection" from our group of 100 people, we don't get a uniform distribution, we get a certain pdf based on the group. In the die case, the random act is "rolling the die". In the height case, the random act is "selecting a member of the group". In this selection case, we always think of the probability of selecting any of the 100 people is the same, so the resulting pdf is what we expect it to be. Also, the people already exist, whereas the die is just now being rolled.
In the die case there is some physical way to compute the outcome of any roll (micro physics) and it might not be a fair roll so you might not get a uniform distribution. But the heights of the group members are pre-established. You might model why each member has the height they have. Maybe each roll of the die generates a new person in the group with that outcome. Maybe each person is an "experiment" like the rolled die. It just that room full of people has known results.
Experiment: N rolls of a die child counts of group of N people
1 roll of N dice
Outcome: N face-up numbers N child counts
Sample Space: N outcome sets {1,2,3,4,5,6} N counts each in {0,1,2.....100}
{1,2,3,4,5,6}N
Random Variable n = face-up number n = child count
Function: N(s) = s, identity N(s) = s, identity
PDF uniform 1/6 if fair rolls child count distribution
prob(θ) 0 2π
********************************************
I finally found someone who verifies my theorem about expectation values: