ricean and all that stuff
DOCX · 32.3 KB
Open DOCX file
Phil's Word-document notes, dated 5.30.03 and updated 8.25.03, summarizing Chapter 2 of Proakis, Digital Communications (3rd ed.). They cover functions of a random variable with his own derivation of the change-of-variables formula, moments, characteristic functions, and the binomial, Gaussian, central and non-central chi-square, Rayleigh and Rice distributions. They end by applying the results to the spread of a standard-deviation probe in a Xilinx design.
AI-written summary; may contain errors.
Extracted text (machine-read; may contain errors)
Standard Distributions PhL 5.30.03
updated 8.25.03
The book from Randy Sylvester is Proakis, Digital Communications , 3rd Ed, 1995. His Chapter 2 is a 60 page survey of fancy statistics stuff. I have two-up copy of this stuff but it is hard to read since print is very small, but the information is there and it is valuable information at that!
This section has no copy to go with it.
2-1 Probability. The sample space S; joint and conditional probabilities; independence.
2-2-1. The notion of P(X x) defines a cdf, but we are more usually interested in the pdf which is the derivative of the cdf. For discrete stuff, he likes to still think of a continuous pdf using a delta function, as in the following:
p(x) = !Syntax Error, IP(x = xi) (x - xi)
Extends to functions of two random variables X1 and X2. Joint, conditional, Bayes.
Here is where the copy starts:
2-1-2. Functions of a Random Variable. This is what we care about.
"Suppose you know all about the distribution of X, and you want to know about some Y = f(X). "
Example is Y = f(X) = aX + b. The idea is to work with the CDF first, sort of transfer over to the new variable, then take a derivative to get the PDF.
Second example is Y = aX3 + b. Same idea.
Third example is Y = aX2 + b, a very important example. This is 2 to 1, not 1 to 1, so a little messier, but the resultant pdf is in 2-1-47 which gets used later.
Now the key result is really (2-1-49) where he uses Y = g(X) instead of my f(X). He does not derive this general result, and the method he shows earlier for the quadratic example does not generalize very well to the general case! However, here is my own derivation of the general case:
Q(y) |dy| = P(x1) | dx1 | + P(x2) | dx2 |+ etc.
Here, do NOT think of the xi as multiple random variables, that lies in the next section. Instead, think of xi(y) as all the solutions of the equation y = g(x). Think of the case y = x2 as the canonical example. There are two strips under the X distribution that map into the same strip in the Y distribution, see handwritten notes with this Word doc titled Example #1. The absolute values are because direction does not matter of the strip, it is the area of the strips that matter, all probabilities are positive definite. You see in the canonical case that x1(y) = + and x2(y) = - so both dx1 and dx2 will be proportional to dy. In the above, P(x) is the distribution for X (perhaps Gaussian, perhaps not), and Q(y) is what we are looking for, the distribution for y = g(x). Divide both sides by dy and just rewrite to get:
Q(y) = P(x1) / |dy/dx|x=x1 + P(x2) / |dy/dx|x=x2 + etc
But y = g(x), so we really have this: [ where x1 = x1(y), etc ]
Q(y) = P(x1) / |g'(x)|x=x1 + P(x2) / |g'(x)|x=x2 + etc
and this is what 2-1-49 says. This is a far-reaching result!!!
In Example #2, I do the thing explicitly for y = x2 in the special case that P(x) = Gaussian. In this case, the two P(x) are equal, and the derivs are equal, so both terms are equal and erase the 2 that comes from the derivative of x2. If you ever get confused, come back to this example! This shows what a squarer does to a PDF if you insert Gaussian signals into it!
Clearly, once you find your new Q(y), you will need to do integrals to compute your new moments. That is not done at this point in the text.
He then goes on to consider the general case Y = Y(X) where each is multiple variables. Be careful, because this section only applies when you are basically converting from n X variables to n Y variables. You find that the PDF's are related by the Jacobian, as one would expect, so nothing too fancy here. But you could use this perhaps in going from x,y coordinates to r, coordinates in some noise situation. If the relation between Y and X is Y = AX where A is an invertible matrix, then J = 1/det(A).
2-1-3. Averages. This is an unbelievably clear well-written section that tells you everything you need to know about moments, central moments, correlation matrix, covariance matrix, joint moments, you name it. Everything is completely clear and concise!
Don't overlook result (2-1-63). This relates again to our issue of y = g(x) and you know about x. To get moments of y, you could on the one hand go compute PDF(y) which is our Q(y) above. But another way is to simply put g(x) inside the integral as shown with P(x).
Characteristic functions . Very very important stuff! The definition is in 2-1-71 where we see first that the (j) is the Fourier Transform of the PDF(x), with a certain choice on the 2 constant location. But also notice that you can interpret (j) as < exp(jx) > = E(exp(jx)) in the normal notation.
Derivatives bring down powers of x from the exponent and then you evaluate at parameter = 0 and you get various moments of interest. Variable the Fourier conjugate of the PDF(x) variable x.
A very powerful result now: Suppose you know the CF for X, and you want the CF for a sum of X's. Consider y = x1 + x2. To get y(j), you need E(exp(jy)), notice the y in there. But now insert y and you get E(exp(jx1) exp(jx2)). If you assume the joint PDF is independent, things factorize and you find that y(j) = x1(j)x2(j)... a very simple result. But we know why this is. It is because in the x domain when you do y = x1 + x2, you are doing convolution (as in my paper on adding random voltages), so in the frequency domain it has to be diagonal like this. The diagonalization applies of course even if each X has a different distribution! I think I could use this idea to combine that binomial distribution stuff with the Gaussian noise in our rate discovery section, but I did not do it. The idea being that these are different distributions and you are adding the variables. I sort of assumed Gaussian for both with large enough R.
He finishes up with multidimensional CF's which I ignore for now.
2-1-4. Some useful distributions.
Binomial. Here he adds some Xi where each is 0 or 1. Does not use CF, just does the traditional derivation and we get statement of the mean, variances.
Uniform. A constant distribution, not too fancy.
Gaussian (Normal). This is our big friend, he has it properly normalized of course. The CDF is the error function and he gets to mention the Q function here as well. Of course the CF of a Gaussian being the FT is another Gaussian. He uses gaussian lower case! Then we get our famous result: if you add up a set of Xi each with a gaussian, as in white noise, you can must multiply the CF's and you get the rules that you just add means and add variances.
Chi-Square. Now getting into new stuff for me. There are two cases. In both cases, you are dealing with the situation y = x2 and x is assumed to have a gaussian characteristic.
(a) central
Suppose X is a gaussian with zero mean. Suppose Y = X2 . What is the pdf for Y? We did this already in the example above and the result is quoted:
p(y) = 1/() exp(-y/22) // See Example #2 mentioned above
In this case we used that example 2-1-47 to get the result. This time we do the pdf first. The cdf is then an integral I have often run into which has 1/with exp(-x/22). This is not even an erf(x) thing. This is the finite integral. We can do the infinite integral to get the CF, however. as in 2-1-107.
So where are we? If you have Y = X2 , the pdf for Y is a Chi-square as above.
Next, add up a bunch of Xi2 terms (all of zero mean and the same ) , as you might get going through a multiplier. Just raise the CF to the nth power and do a reverse FT to get the 2-1-110 result which is a power of y times an expo. This is the Chi-Square with n degrees of freedom. Very good. If you add just two terms, you get n=2 case and the power goes away and you have just the exp(-y/22) part. The cdf is a mess, related to incomplete Gamma.
For the general case of adding n squared terms, author gives these results: Char Function, PDF, the CDF as an integral which can be done in some cases. The first two moments are also stated, and then the variance is given. VERY useful stuff to me to have ready to use like this.
(b) non-central
Now we start all over again where X is a gaussian with a non-zero mean! The pdf is now an expo times a cosh as in 2-1-115. Now add a series of Xi2 where we allow each one to have a different mean, but the same . Compute the CF as a product of the underlying CFs, do the FT, and you end up with the PDF being a Bessel I function with parameter s2 that is the sum of all the means squared. This result is the non-central Chi-squared. You can relate it to certain Qm(a,b) functions of Marcum. Again we get all these results stated: Char Function, PDF, CDF (no closed form), and again we get the two lowest moments and the variance.
Summary. Chi-Square is a large subject, used to see whether data fit a presumed distribution, for example. But here we just think of it as the PDF for the sum of squares of voltages. It the voltages squared have zero mean Gaussians, then you have the simpler central Chi case. Otherwise the non-central.
Rayleigh. (starts with zero mean Gaussian, related to central Chi square) This is what you get when you do the Chi-Square sum of squares, but in the end you take a square root! Proakis gets the pdf just by changing variables in 2-1-128. The CDF is trivial, so are all the moments. As before, he then moves on to the general case of n terms. The Char Function is a mess which turns out to have a 1F1 component with a negative argument. Author first does simple case being the sum of just two terms to get all this stuff. Simple because the Chi-Square has no powers for n=2, just the expo.
Now generalize to sum of n terms, the usual next step. The result is a pdf with powers and expo.
Rice. (starts with non-zero mean Gaussian, related to non-central Chi square). The PDF is now a Bessel I function even in the general case as in 2-1-143. The square root of the sum of Xi2 variable is called R. The CDF is state. Finally, the moments are ALL given as confluent hypergeometrics!
I think when you have a constellation point perturbed by noise in the two dimensions, you end up with a Rice situation for the magnitude of the vector, and that is why fading and multipath get associated with Rice and Rayleigh.
Application. Now I am going to apply the above to a specific situation where I am inserting statistical probes into my Xilinx RD design. I have the probe add up the squares of N numbers and then divide by N, and then take the square root. I know that in the limit that N , this is the square root of the variance from zero. Since I am interested in the no-correlation case, the mean is zero, so my probe is trying to compute the square root of the variance, which is just the SD. But I don't have time to wait for N=, so I want to know what the distribution looks like of finite SD readings about the true SD, ie, the variance of the standard deviation.
See attached Example #3 where this problem is first attacked. I have left out the divide by N part at this point. (N = n). We have a simple Rayleigh situation since means are 0 and since the data stream is assumed Gaussian for large R, even for no noise. I find these results
<r> = n where n = ( [ n+1]/2) / (n/2)
<r2> = n 2
variance = <r2> - <r>2 = n 2 [ 1 - n2/ (n/2) ]
std =
Now, Maple shows us that the quantity n2/ (n/2) 1 from below, and gets quite close at n = 5 or 10, in fact. I show from an AS reference that n2 n/2 for large n.
Now we want to go back and add in our missing factor of 1/ that should be present in our expression at the start of Example #3 for R. On page 2 of that example, we show what an overall constant of 2 does to things. The PDF is changed as shown [ this comes from my multiplying random numbers paper, one of the examples ]. The net result is that all moments get scaled by 2k for kth moment. That means that our STD result above will get scaled in our case by 1/, so we will then have
STD =
Again, this describes the distance of our "probe reading" from the actual standard deviation of the data. It is the standard deviation of the standard deviation. If n , then STD = 0, and we should be right on the money for whatever the theory says. For example, at the R-chunk outputs, we expect for an N case to have this result:
= // in theory units. We have R = 64, so expect = 8.
When we convert this to hardware units with our 8-bit bus, we get = 8*127 = 1016. This is what we would expect to see with n . And now we have an explicit formula for the expected STD away from this amount for finite m. Here is an excel table,
This shows that with N=90 as I am doing, the 3 sigma range is (789,1243) around 1016. In fact we have to go out to N = 5000 to get the STD down to 10 hardware units!
So, I think this gives me a much improved view of my statistical probes. At N = 90, = 76/1016 = 7.5%, so if we look at 3, we are going about 23% to each side. This is much larger than I thought it would be! This is my do5 simulation. Remember that 3 is 99%, 2 is 95%, and 1 is 68%.