Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Math / Scrambler / Not needed anymore

another 2.5 rewrite that got dumped REVD

DOCX · 23.9 KB
Open DOCX file

A discarded rewrite draft dated 7.24.13 by Phil, intended as section 2.5 on probability, MLS spectral power density and autocorrelation. The text covers joint and conditional pdfs, independence, expectation, covariance, variance and correlation, with proofs, then applies these to symbol sequences and <a_n a_{n+k}> averages. The MLS spectral density and autocorrelation parts are not reached in the excerpt.

AI-written summary; may contain errors.

Extracted text (machine-read; may contain errors)
This is the Title PhL 7.24.13 This is a rewrite of temp55 that crashed and burned early on and led to a massive update of FT. 2.5 Basic Probability Theory, MLS Spectral Power Density, and Autocorrelation (a) Basic Probability Theory The probability of events A and B both being true is denoted by P(AB), or P(A,B). This is known as the joint second-order probability density function (pdf). A related concept is P(A|B) meaning the conditional probability that A is true given that B is true. Here are the relations: P(A,B) = P(AB) = P(A|B) P(B) = P(B|A) P(A) . (2.5.1) If events A and B are statistically independent in second order, one gets: P(A,B) = P(AB) = P(A) P(B) P(A|B) = P(A), P(B|A) = P(B) . (2.5.2) The result P(A|B) = P(A) means that the probability of A being true is not influenced by whether or not B is true. Similarly one can define joint pdf's of third order and beyond. For example, statistical independence for a third-order pdf is indicated by, P(A,B,C) = P(ABC) = P(A)P(B)P(C) . (2.5.3) A possible "event" is that a "random variable" X takes some value x. In this case, instead of writing things like P(A) = P(X = x), we abbreviate and say P(X=x) = p(x). Similarly P(X=x, Y=y) = p(x,y), so the position of the argument of p(....) indicates which random variable one is talking about. This concept applies for the values x being either continuous (such as the reals) or discrete (such as GF(p) ). The joint pdfs are now written p(x,y) and p(x,y,z) and so on. If the random variables X and Y are statistically independent, then all the pdf's must completely "factor". For example, p(x,y,z) = p(x)p(y)p(z). Any pdf is normalized to 1 since the probability of all possible events is 1. Thus, ∫∫.... p(x,y....) dx dy... = 1 or Σij.... p(xi,yj....) = 1 . (2.5.4) An example is that ∫p(x)dx = 1 or Σip(xi) = 1. If f(X,Y...) is some arbitrary function of some random variables X.Y..., one can define the expectation value of f as follows: E(f(X,Y,...)) = ∫∫.... f(x,y...) p(x,y....) dx dy.... or Σij.... f(xi,yj...) p(xi,yj....) . (2.5.5) Cases of interest to us will be the following E(X) = ∫ x p(x) dx or Σixi p(xi) = "the mean", often written μx (2.5.6) E((X-μx)(Y-μy)) = ∫∫ (x-μx) (y-μy) p(x,y) dx dy or Σi,j (xi-μx) (yj - μy) p(xi,xj) = "the covariance of X and Y" = cov(X,Y) . (2.5.7) If the joint second-order pdf factors, so that p(x,y) = p(x)p(y), or equivalently p(xi,xj) = p(xi)p(yj), then one says that the random variables X and Y are uncorrelated. Since the higher joint pdf's might not factor, this does not imply statistical independence. Thus we have only one direction here: Fact: X and Y statistically independent X and Y are uncorrelated (2.5.8) We can now identify three different statements: Fact: X and Y are uncorrelated E(XY) = E(X)E(Y) cov(X,Y) = 0 . (2.5.9) Proof: We show this in the continuous notation but everything applies also in the discrete notation. X and Y uncorrelated means p(x,y) = p(x)p(y) which at once means E(XY) = E(X)E(Y) from the definitions shown above: E(XY) = ∫∫ x y p(x,y) dx dy = ∫∫ x y p(x)p(y) dx dy = [∫x p(x) dx] [∫y p(y) dy] = E(X)E(Y) . The covariance is cov(X,Y) = ∫∫ (x-μx) (y-μy) p(x,y) dx dy = ∫∫ (x-μx) (y-μy) p(x)p(y) dx dy = [ ∫ (x-μx) p(x) dx ] [ ∫ (y-μy) p(y) dy ] . But each factor vanishes for the reason shown below, and thus cov(X,Y) = 0 : [ ∫ (x-μX) p(x) dx ] = ∫x p(x) dx - μx ∫ p(x) dx = μx - μx * 1 = 0 . QED Now suppose X and Y are the same random variable. Then we have, for example, p(x,y) = p(x)δ(x-y) or p(xi,yj) = p(xi)δi,j . (2.5.10) The reason is that if X and Y are the same random variable, they cannot take different values. If X,Y and Z are all the same variable, then we would have p(x,y,z) = p(x)δ(x-y)δ(x-z) or p(xi,yj,zk) = p(xi) δi,j δi,k (2.5.11) Note that δ(x-y)δ(x-z) = δ(x-y)δ(y-z) = δ(x-z)δ(y-z) and δi,j δi,k = δi,j δj,k = δi,k δj,k . The main application of this concept relates to the variance and standard deviation of X: cov(X,X) = ∫∫ (x-μx) (y-μx) p(x,y) dx dy = ∫∫ (x-μx) (y-μx) p(x)δ(x-y) dx dy = ∫ (x-μx)2 p(x) dx or Σi (xi -μx)2 p(xi) ≡ var(X) = "the variance of X" = [σ(X)]2 = "the standard deviation squared" . (2.5.12) Notice that this quantity could in theory vanish if p(x) = δ(x - x1) so that the entire pdf is concentrated at a single value. var(X) = ∫ (x-μx)2 p(x) dx = ∫ (x-μx)2 δ(x - x1) dx = (x1-μx)2 = (x1- x1)2 = 0 = σ(X) . A pdf of this form is not very interesting and normally one has σ(X) > 0. Fact: var(X) = σ2(X) = E(X2) - {E(X)}2 (2.5.13) Proof: E(X2) - {E(X)}2 = ∫x 2p(x) dx - {∫x p(x) dx} 2 = ∫x 2p(x) dx - μx2 var(X) = ∫ (x-μx)2 p(x) dx = ∫x 2p(x) dx -2μx∫x p(x) dx + μx2∫ p(x) dx = ∫x 2p(x) dx - 2μx μx + μx2 1 = ∫x 2p(x) dx - μx2 . QED There is one more object of interest which is the correlation function of X and Y corr(X,Y) ≡ cov(X,Y) / [σ(X) σ(Y) ] . (2.5.14) This is just a rescaled version of the covariance. We may now extend our fact above : Fact: X and Y are uncorrelated E(XY) = E(X)E(Y) cov(X,Y) = 0 corr(X,Y) = 0 (2.5.15) To be consistent with notation in other documents, we shall use this alternate <...> notation for expectation values, < f(x,y...)> ≡ E(f(X,Y,...)) . (2.5.16) Normally the notation <...> means an average of something, and expectation values are averages. (b) Application to Sequences We shall now apply these probability concepts to a sequence {an} . A single such sequence, whether finite or infinite, has some expectation value for an. We could count the number of occurrences of each symbol in the sequence N(an) and divide that by the length of the sequence N, to get p(an) = N(an)/N as the probability of occurrence of symbol an. Then <an>1 ≡ Σnan p(an) = (1/N) Σn anN(an) = E1(an) would be the mean value of the symbols in this one sequence. The notation <..>1 is used to indicate this kind of average. We think of each position in the sequence as corresponding to a random variable Ai which can take values in some set of values, and the value that Ai takes is called ai. Then, for example, <an> ≡ E(An) = Σnan p(an) (2.5.17) <an2> ≡ E(An2) = Σn,m an2 p(an) <anam> ≡ E(AnAm) = Σn,m an am p(an,am) <anam> = <an><am> = <an>2 // from (2.5.15), only if An and Am are uncorrelated where all sums are over the entire sequence. And Fact (2.5.13) says var(An) = σ2(An) = E(X2) - {E(X)}2 = <an2> – <an>2 ≡ σn2 (2.5.18) Question: Does <an> or <an2> depend on n ? Recall that p(an) = P(An = an) so in theory we could have P(A3= an) ≠ P(A5= an), say. In our treatment of sequences we assume that there is nothing distinguishing about a particular symbol position other than its value, so we would have P(A3= an) = P(A5= an) for positions 3 and 5 in any sequence. So the answer is no, <an> and <an2> do not depend on n, they are just averages over the sequence. On the other hand, when we write <anam> = <anan+k>, this could depend on the distance k between two symbol positions, as it does in the AMI line code (see FT Section 37), but <anan+k> does not depend on n. When the An and Am variables are uncorrelated, then <anam> = <an><am> = <an>2 and in this case we find that <anam> depends on neither n nor m, and <anan+k> depends on neither n nor k.