Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Physics / Quantum Mechanics / Relativistic Quantum Mechanics

BD V1 chap 1_6

DOCX · 251.4 KB
Open DOCX file

Word-document notes by Phil, dated 3.26.05, working through Bjorken and Drell Vol. 1 with a chapter-by-chapter contents list. They cover the Dirac equation, covariance, free-particle solutions, Foldy-Wouthuysen, hole theory, propagators and Mott scattering. The visible text derives the probability current, shows why Klein-Gordon fails to give positive density, and shows why Dirac matrices need N=4, with personal asides.

AI-written summary; may contain errors.

Extracted text (machine-read; may contain errors)
Relativistic Quantum Mechanics PhL 3.26.05\ Bjorken and Drell Contents: Chapter 1: The Dirac Equation (2) 2 1.1 Formulation of Relativistic Quantum Theory (2) 2 1.2 Early Attempts (4). 2 1.3 The Dirac Equation (6) 4 1.4 Nonrelativistic Correspondence (10) 6 History 12 Chapter 2: Covariance of the Dirac Equation (16) 12 2.1 Covariant form of the Dirac Equation (16). 13 2.2 Proof of Covariance (18). 14 2.3 Space Reflection (24). 24 2.4 Bilinear Covariants (25) 24 Chapter 3: Solutions of the Dirac Equation for a Free Particle (27) 24 3.1 Plane Wave Solutions (28). 24 3.2 Projection Operators for Energy and Spin 33 3.3 Physical Interpretation of Free-particle Solutions and Packets (35) 36 Chapter 4: The Foldy-Wouthuysen Transformation (46) 41 4.1 Introduction (46). 42 4.1 Free-particle FW Transformation (46) 42 4.2 The FW Transformation with E&M Fields (48) 44 4.3 The Hydrogen Atom (52) 48 Chapter 5: Hole Theory (64) 50 5.1 The Problem of Negative-energy Solutions (64). 50 5.2 Charge Conjugation (66) 50 5.3 Vacuum Polarization (70) 52 5.4 Time Reversal and Other Symmetries (71) 53 Chapter 6: Propagator Theory (78) 55 6.1 Introduction. 55 6.2 The Nonrelativistic Propagator (71) 55 6.3 Formal Definitions and Properties of Green's Functions (still non-relativistic situation) 59 6.4 The Propagator in Positron Theory (89). 61 Chapter 7: Applications (100) 65 7.1 Coulomb Scattering of Electrons (100) 66 Preliminary Comments: Compare to Rutherford result, and discuss BD units. 66 Phase Space: 68 Side Exercise with jμ. 69 7.2 Trace Theorems and Final Evaluation of the Mott cross section (103) 71 7.3 Coulomb Scattering of Positrons (106) 76 "Bj" (James Bjorken) was born 1934, fud Stanford, and Prof there at SLAC. He is now 74 years old. Sidney Drell is now emeritus at SLAC and seems to have political interests, being in the Hoover and JASON and such things like Fred Zach did, did nuclear arms control work, is a violinist, Los Alamos board. Perhaps a person similar to Jim's buddy Bill Fraser, who by the way has a very low web profile, CARA is California Association for Research in Astronomy. Oddly, I never met Bjorken or Drell in my 7 years at UCB, though I probably attended talks one or the other gave. Preface: They are taking the "propagator approach" pioneered by Feynman, heavy on the graphs, some think these are the theory. They would like experimentalists and students to understand all this stuff better. Reader must know QM at Schiff level. A long list of things NOT covered in the book: the action/variational approach of Schwinger, axiomatic field theory (the mathematics), pure S-matrix theory (Chew), bound states, fancy dispersion relation stuff, massive vector mesons (too bad), derivative couplings. We are referred to a long list of books on these not-included subjects, on which we find my Mandl book and Chew's book. Since I was in the Chew world, I never even saw these books, but the names are dimly familiar. Chapter 1: The Dirac Equation (2) 1.1 Formulation of Relativistic Quantum Theory (2) . Authors write out 6 basic facts that we all know from regular quantum mechanics, and want to make sure these things are still true in the relativistic extension. (1) there is a state vector ψ in Hilbert space and prob is |ψ|2 , function of qi, si and time t (2) observables correspond to Hermitian operators, they write "hermitian". (3) These operators have eigenstates with real eigenvalues ωn (4) You can describe an arbitrary system state by an expansion on complete basis functions that probably come from the eigenfunctions of an eigenvalue equation. (5) In such an expansion with coefficients an, prob of measuring ωn is |an|2. (6) The SE tells how ψ moves in time 1.2 Early Attempts (4). In non-rel, we have Hψ = i∂ψ/∂t (and Hψ* = - i∂ψ*/∂t) and H = p2/2m for a free particle, and p = -i, ie, we have the SE. The first idea is to replace H with keeping p = -i, hence 1.9. Authors claim that you then bring in all orders of the spatial derivative, and this makes you have a non-local theory which is a mess. The second idea is to think H2 = p2 + m2 and then say (i∂t)2 ψ = (p2 + m2)ψ so you are sort of saying E2ψ = H2ψ. This is really the Klein-Gordon equation (∂μ∂μ + m2)ψ = 0 with =c=1 for the moment. What is wrong with this? To see, We have to first go back to regular quantum mechanics with the regular SE. I am going to do this including a potential V which may be complex, anticipating a need to later derive 20.1 on page 130. Suppose we define j = /(2im) [ ψ*(ψ) - (ψ*) ψ] = /(2im) [ ψ*(ψ) - cc ] = a "current" (7.3) Apply div on both sides to get j = /(2im) [(ψ*) (ψ) + ψ*2ψ - cc] = /(2im) [ ψ*2ψ - cc] since the two (ψ*) (ψ) terms cancel. Now since H = p2/2m + V = -22/2m + V, we replace 2 = 2m(H-V)/(-2) => j = /(i) [ ψ*(H-V)ψ - ([H-V]ψ )*ψ]/(-2) = (i/) [ ψ*Hψ - (Hψ )*ψ – ψ* V ψ + V* ψψ ] = (i/) [ ψ*Hψ - (Hψ )*ψ – (V–V*) ψ* ψ] = (i/) [ ψ*Hψ - (Hψ )*ψ – 2iIm(V)ψ* ψ] = (i/) [ ψ*Hψ - (Hψ )*ψ] + (2/) Im(V) ψ*ψ We still have that minus sign between the two terms, the second is still "cc" of the first. But now we use these facts: Hψ = i∂tψ (Hψ)* = - i∂tψ* and now that minus sign between the terms becomes a plus j = (2/) Im(V) ψ*ψ = (i/) [ ψ*( i∂tψ ) + (i∂tψ* )ψ]/(-2) = (i/)(i) [ ψ*( ∂tψ ) + (∂tψ* )ψ] = – ∂t(ψ*ψ) = – ∂tρ ρ = ψ*ψ = positive! So we conclude that j + ∂tρ = (2/) Im(V) ρ // agrees with 20.1 where VI ≡ – Im(V) which, for a free particle or for real V, we can write as j + ∂tρ = 0 or ∂μjμ = 0 and we have a "conserved current". This is of course a probability current, not an electric current, and ρ is the probability density. The key idea is that ψ*ψ is positive definite so you can interpret it as a probability. And jμ = /(2im) [ ψ*(∂μψ) - (∂μ ψ*) ψ] = conserved 4-current Now if we try using this same current in the E2ψ = H2ψ theory we run into a problem. We get ρ = j0 = /(2im) [ ψ*(∂tψ) - (∂t ψ*) ψ] which is the same as before in terms of the wavefunctions. But NOW we don't have a way to show that ρ is positive! We can no longer say ∂tψ = +(1/i) Hψ and ∂tψ* = - (1/i) Hψ*, so we cannot get rid of that minus sign! When history saw this problem, it put aside the Klein-Gordon equation and moved to the Dirac equation. The KG paper was in 1927, and Dirac was 1928. It would be fascinating to read these papers, but not now! 1.3 The Dirac Equation (6) I never appreciated before what Dirac did, it is such a simple idea. We want to get back to the linear time dependence so we have Hψ = i∂ψ/∂t and we can maybe rescue the above probability problem. So try adding α and β matrix coefficients as shown in 1.13. Scalar coefficients fail because you are not rotationally covariant, not clear what matrices do on that, but we shall see later. So assume some kind of column vector solution and some NxN matrices for the αi and β coefficients. In order to maintain the requirement that E2ψ = H2ψ = (p2c2 + m2c4)ψ = ([-i]2c2 + m2c4)ψ, we require that the matrices have the properties shown in 1.16. This makes the second last term in equation A be zero, and makes the first give the 2 term and the last give the m2c4term. Very excellent, Mr. Dirac. We require also that α and β be Hermitian matrices to keep H Hermitian, and then easy to show that the αi are traceless. Since α2 = β2 = 1, only eigenvalues can be ±1 (remember they are Hermitian so eigenvalues must be real). In order to make the diagonalized αi have trace 0, it must have an equal number of +1 and -1 eigenvalues, so N must be even! Why not try N=2? If we try α k = σk, the first item in 1.16 is happy. But what can you use for β? Need a matrix such that βαi = – αiβ. The entire space of matrices in N=2 is spanned by σ and 1 so β = aσ + b 1 where all coefficients can be complex, so there are 8 real numbers apropo for 2x2. So we need: [ aσ + b ] σk + αk[ aσ + b ] = 0 => a[ σσk + σkσ] + 2bσk = 0 => ak1 + bσk = 0 But tr(σk) = 0 so this says that 2ak = 0 so we end up with ak = 0 and then b = 0 as the only solution! You could mount a completely formal proof, but this is enough for me, there is no N=2 solution. There is an N=4 solution and it is given by 1.17. I have verified on scratch that 1.16 are all true and the matrices are Hermitian, as required. Page 9 then constructs a conserved current with a ρ that is positive! Recall that our non rel current that worked was this: jμ = /(2im) [ ψ*(∂μψ) - (∂μ ψ*) ψ] ∂μjμ = 0 ρ = ψ*ψ The current that works in the Dirac theory does not have gradients, it is just this: j = c ψ†α ψ ρ = ψ† ψ Let's just do this out the way we did in the non-rel case: j = c [(ψ†) α ψ + ψ†α (ψ)] = ? Our Ham equation says i∂tψ = (c/i) α ψ + βmc2ψ = Hψ - i∂tψ = - (c/i) α ψ† + βmc2ψ† = Hψ† which says that (c/i) α ψ = i∂tψ - βmc2ψ (c/i) α ψ† = i∂tψ + βmc2ψ cα ψ = - ∂tψ - i(βmc2)/h ψ cα ψ† = - ∂tψ† + i(βmc2)/h ψ† Therefore j = c [(ψ†) α ψ + ψ†α (ψ)] = - (∂t ψ†)ψ - ψ†(∂t ψ) = - ∂t[ψ† ψ] QED So somehow cα is a "current operator" in this Dirac world. Of course ψ†ψ = |ψ|2 since we are in a metric space where (ψ,ψ) = ψ†ψ in the usual matrix fashion, and we can interpret ρ again as a probability. So what has Dirac just done here? (1) i∂tψ = (c/i) α ψ + βmc2ψ = Hψ is our candidate new Schrodinger equation which just means that H = (c/i) α + βmc2 is our candidate Hamiltonian. So we use the same old SE, but ψ is now a 4-component column vector of functions and H is a 4x4 matrix. (2) This form gives the result H2ψ = E2ψ = ( i∂t)2ψ = (p2c2 + m2c4)ψ. In other words, we get that H2ψ = ([-i]2c2 + m2c4)ψ . In other words, H2 is what it must be to maintain the p = -i relationship and have the right notion that H2 = E2, relativistically correct. (3) We get a conserved current with ρ = ψ† ψ ≥ 0 as required. The one missing thing we have not shown is that somehow the SE is covariant with this H! I guess this is complex enough to show that B&D have an entire chapter devoted to it, Chapter 2. Another item is this: we have not said what kind of "particle" this Dirac Hamiltonian might work for? I wonder how excited Dirac was when he figured this out. It is something I think I could have done in the same circumstance, nothing too messy really. I think these ψ 4-vectors belong to the ½ 1 + 1 ½ representation of the Poincare Group, but that is a long distance away right now. I wonder what was known about such group representations in the Dirac timeframe of 1928 when he did this work? 1.4 Nonrelativistic Correspondence (10). This will be the big test of the Dirac idea: do we get anything resembling the real world in the NR limit of the Dirac theory?? Before I can do notes for this chapter, we have to pre-derive some equations which are extremely important to the discussion, and which are very non-obvious. So let's do that now: _____________________________________________________________________________________ Aside on equations 1.27: There is a lot packed in here, a typical B&D-ism that requires a 2 hour digression on the part of the reader. In the Schrodinger Picture, operators are fixed and the states move under the force of H. |α(t)> = e-iHt/ |α(0) H // Schiff page 169 24.3 In the Heisenberg Picture, operators move according to the commutator relation and states are fixed. dQ/dt = (i/)[H,Q] + ∂Q/∂t |α> // Schiff page 169 24.4 or 24.10 p 171 In the Interaction Picture with H = H0 + H', states move with H', operators with H0. dQ/dt = (i/)[Ho,Q] + ∂Q/∂t |α(t)> = e-iH't/ |α(0) // Schiff 24.14 and 24.13 approx Now, with regard to 1.27, what "picture" is implied? It cannot be the Schrodinger picture because in that picture, you don't have an H commutator relation at all, so it must be one of the other two pictures. In the text, we are told to think of equation 1.26 as saying H = H0 + H'. Defining the H's in this way, we conclude that 1.27 must refer to the Heisenberg Picture since the "full H" appears in the commutators. 1.27 #1: The commutator [H,r] is nonzero only due to the p item in H, so we have [H,ri] = cα[p,ri] = cαj[pj, ri] = -icαi = (/i) cαi since [p,x] = -i Therefore, dri/dt = (i/)[ H,ri] = cαi or dr/dt = cα ≡ vop So we have now proven this first result and defined vop as shown. We are seeing that cα is the correct velocity operator in this theory. Notice that operator r has no "explicit" time dependence so we did not have any term of the form ∂r/∂t. 1.27 #2: We start of here saying: dπ/dt = (i/)[H,π] + ∂π/∂t π = p - (e/c)A(r,t) In Hamiltonian theory, r and p are the arguments of H, and neither has "explicit time dependence", so as before we have ∂p/∂t = 0. This means ∂π/∂t = -(e/c) ∂tA . Thus we have shown that dπ/dt = (i/)[H,π] -(e/c) ∂tA which verifies the second equation 1.27 #3: Let's compute [H, πi] = [H, pi] - (e/c) [H, Ai(r)]. Here we show the r dependence of A as a reminder that this is not going to commute with p-things. So the first term is this [H, pi] = [ cα (p - [e/c]A(r)) + βmc2 + eφ(r), pi] = -e α [A(r), pi] + e [φ(r), pi] // now use rule [f(r),pi] = + i ∂if(r) = -ie {αj ∂iAj(r) - ∂iφ } In vector form, this says [H, p] = -ie {αj Aj(r) - φ } = -ie {(αA(r)) - φ } where the gradient is the vector sense in each term on the RHS. Meanwhile, we now have to compute [H, Ai(r)] = [ cα (p - [e/c]A(r)) + βmc2 + eφ(r), Ai(r)] = cα [p, Ai(r)] = -i cα Ai(r) // used [p, f(r) ] = -if(r) = -ic αj∂jAi Therefore we have [H, πi] = [H, pi] - (e/c) [H, Ai] = -ie {αj ∂iAj(r) - ∂iφ } + ie αj∂jAi = -ie { - ∂iφ + αj ∂iAj - αj∂jAi } = -ie { - ∂iφ + αj (∂iAj - ∂jAi) } But write (∂iAj - ∂jAi) = εijk [ x A]k = εijkBk . For each choice of i ≠ j, there is only one surviving term in the implied k sum, otherwise there are no terms. Then we have αj (∂iAj - ∂jAi) = αj εijkBk = εijk αjBk = [ α x B ]i and then [H, πi] = -ie { - ∂iφ +[ α x B ]i } = (/i)e { - ∂iφ +[ α x B ]i } Now we use 1.27 #2 which said dπi/dt = (i/)[H,πi] -(e/c) ∂tAi = e { - ∂iφ +[ α x B ]i } - (e/c) ∂tAi = e { - ∂iφ - (1/c) ∂tAi + [ α x B ]i } = e { Ei + [ α x B ]i } or in vector form dπ/dt = e { E + α x B } and recall cα ≡ vop hence #3 is proven. Now let's look again at these three results: #1: dr/dt = cα ≡ vop #2: dπ/dt = (i/)[H,π] -(e/c) ∂tA #3: dπ/dt = e { E + α x B } = e { E + (1/c) vop x B } The first says what "velocity" is in this Dirac theory. The second is intermediate. The third says what "force" is in this theory, and it is exactly the Lorentz Force where we must use the appropriate velocity. Notice that all of these are matrix equations, though this is perhaps not obvious. For example, #1 says dr/dt = c where 1 is the 2x2 identity and σ are the Pauli 2x2's. The #3 equation says dπ/dt = eE+ e _____________________________________________________________________________________ Aside on Equation 1.33: (another non-obvious fact) The problem here is to show this : i π x π φ = (-e/c) B φ or i (p - (e/c)A) x (p - (e/c)A) φ = (-e/c) B φ or -i(e/c) { A x p + p x A }φ = (-e/c) B φ or i{ A x p + p x A }φ = B φ One must remember that p = -i, and in the second term in {...} the is acting on everything to its right, it is not just acting on A. So throw this in: i (-i) { A x + x A }φ =?= B φ or { A x + x A }φ =?= B φ So if we can show this last line, we have it. Still the second term requires disambiguation. We can write { x A }k φ = εkij∂iAjφ means εkij∂i(Ajφ) = { x (Aφ) }k or x A φ means x (Aφ) So here then is what we have to show: A x (φ) + x (Aφ) =?= Bφ of course B = x A But I have a vector identity which says this x (Aφ) = -A x (φ) + ( x A) φ so putting this in the previous equation, the A x (φ) terms cancel and we get ( x A) φ =?= Bφ which is correct, QED. _____________________________________________________________________________________ Aside on Equation 1.35: (another non-obvious fact) This claims to just be a rewrite of 1.34. (1) Show that A = ½ B x r is acceptable for a uniform B field, that is, show x A= B. x A = ½ x (B x r) = ½ { B (r) – (B)r } which follows from the usual vector identity where two terms vanish because B = constant in r. Now (r) = ∂iri = 1 + 1 + 1 = 3 (B)rk = Bi∂irk = Biδik = Bk so (B) r = B Thus, ½ { B (r) – (B)r } = ½ { 3B - B} = B QED (2) Compute A2 AA = ½ B x r ½ B x r = ¼ { (BB) (rr) – (Br)2} = ¼ {B2r2 - (Br)2} But we are going to ignore this term because it has 1/c2 compared to the linear term. (3) The cross term in the π2 part of H is this, where we have the same "to the right" issue we had in our previous derivations above, so we write this as -(e/c) { pA + Ap }φ = +(ie/c) { A + A }φ = (ie/c) { (Aφ) + A (φ) } But vector identity says (Aφ) = A (φ) + (A) φ // where second acts only on A But A = ½ (B x r) = ½ { r x B - B x r } but of course x B = 0 and ( x r)k = εkij∂irj = εkijδij = 0 as well. so A= 0. But I knew that because (xA) = 0 for any A, reminding us of Maxell B = 0. OK, fine. So we then have (Aφ) = A (φ) and the cross term in the π2 part of H is this, = (ie/c) { (Aφ) + A (φ) } = (2ie/c) A (φ) = (2ie/c) ½ (B x r) (φ) = - (e/c) (B x r) (pφ) = - (e/c) B r x (pφ) = - (e/c) B L φ We have to divide by 2m from the first term in 1.34 so we get contribution from π2 term in H = - (e/2mc) B L φ We have no electric potential in our uniform B field, so drop last term. And then S ≡ (/2) σ so the second term is then -(e/2mc)(2/)S = - (e/mc)S = - (e/2mc)2S So the contribution of the above terms in H is then this: - (e/2mc) B L φ - (e/2mc)2S B φ = - (e/2mc) ( L + 2S ) B φ and this concludes our derivation of 1.35. ____________________________________________________________________________ Finally, we may now begin our notes on this section! ("Nonrelativistic Correspondence") The first item considered is "the electron at rest". So suddenly our authors are saying that this Dirac theory applies to electrons, maybe. The argument is made that, for such an electron, we can ignore the spatial gradient terms in H shown in 1.13. Yes, we know that λ = h/p and as p→0 λ → ∞ for the "wave" that describes an electron. I guess I buy this argument. Then I agree that the four spinors shown in 1.24 are the correct independent rest solutions, and be sure to notice that the exponents have opposite sign in the last two solutions, these are the "negative energy solutions" as Mandl also likes to call them, due to the + sign in the exponent. Now what comes next? I think the claim is that that prescription 1.25 pμ → πμ for bringing in the EM potentials is relativistically exact. At least we know it is a covariant thing to do, so let's do it. This makes our H as in 1.26 where we have now separated the space and time parts of the potential we just added. The group of equations I call 1.27 are now interesting, but not quite in the main line of the discussion. First, these are Heisenberg Picture operator motion equations and I have derived them all above. There are two main points to learn here: (1) the correct relativistic velocity operator seems to be the matrix cα ≡ vop (2) the "Lorentz force" is really dπ/dt, not dp/dt, and is given by the "usual formula" where we use the correct relativistic velocity cα . I presume when we take the NR limit of this formula, we will recover the traditional formula somehow for Lorentz force, but BD don't do that here. (3) Notice that the conserved space current is this : j = ψ†vopψ from 1.22, j0 = ψ†ψ. The velocity is the velocity of probability "flow". Now we start afresh on page 12. We break the 4-vector into a pair of 2-vectors and , and we write the Dirac equation exactly as in 1.28. The flipping of the second spinor arises from the off-diagonal appearance of σ in the α matrix as shown in 1.17 page 8. At this point, we pretend that all solutions are positive energy solutions, and we remove that fast time dependence from the barred spinors (same for each spinor) thus defining unbarred ones, as in 1.29. Then 1.30 are the exact equations for the unbarred ones. If you look at the second equation and ignore the two terms that don't have power of c in them, you get 1.31 which says that the χ spinor is smaller than the φ spinor by v/c, so the φ is the one we want to look at for our NR limit. Using 1.31 to eliminate χ from the first equation in 1.30 [ where χ appears in one place] , we get 1.32. Notice that, although χ is small, it gets boosted in the first equation by c so it should not be ignored, and it is the cause of the first term on the RHS of 1.32. At this point, we do lots of algebra as outlined above, and we get 1.34 to drop out. Suddenly a "spin term" has appeared out of thin air, it is the σB term and this is saying that there is some kind of intrinsic magnetic moment proportional to σ on this electron in the NR limit. We then assume a uniform B field and get 1.35 which clinches things. This says that, in the NR limit of the Dirac theory, the "interaction Hamiltonian" involves the famous result H' = -(e/2mc) (L + 2S)B where S = (/2)σ and L = r x p If we take the NR Hamiltonian p2/2m and make the spatial π replacement and keep the cross terms, we get the L term shown above, but we don't get the S term unless we add it "manually" by assuming that the electron has a certain magnetic moment. Here we have derived the presence of the second term from the Dirac equation. By finding a covariant Hψ = Eψ formulation (though we have not yet shown covariance), and by looking at its solutions in the NR limit, the "spin" has appeared out of nowhere. What order in history did things happen? I think the Dirac work above was 1928. The Stern-Gerlach experiment was 1922. Below is some spin history from wiki. Pauli made up the idea of a 2-valued but unknown internal quantum number "up/down" and that each spatial state could have one electron of each internal value -- This was the Pauli Exclusion Principle, they are saying 1924 for this. Pauli rejected Kronig's suggestion that this internal number related to angular momentum of the electron, but the Gerlach thing sure must have suggested that notion. Then in 1925 our friends U and G and E did some theory on this spin idea. Then Pauli changed his mind, and formalized the theory in 1927 and named the "Pauli matrices" and "2-spinors". Then in 1928 Dirac did his relativistic H and saw this spin stuff just pop out! That was a major physics victory of the 1920's era. A lot of stuff happened in a small time. Remember that the SE itself was only found 1926, and Heisenberg's matrix stuff was 1925-6 as well. History Spin was first discovered in the context of the emission spectrum of alkali metals. In 1924 Wolfgang Pauli introduced what he called a "two-valued quantum degree of freedom" associated with the electron in the outermost shell. This allowed him to formulate the Pauli exclusion principle, stating that no two electrons can share the same quantum state at the same time. The physical interpretation of Pauli's "degree of freedom" was initially unknown. Ralph Kronig, one of Landé's assistants, suggested in early 1925 that it was produced by the self-rotation of the electron. When Pauli heard about the idea, he criticized it severely, noting that the electron's hypothetical surface would have to be moving faster than the speed of light in order for it to rotate quickly enough to produce the necessary angular momentum. This would violate the theory of relativity. Largely due to Pauli's criticism, Kronig decided not to publish his idea. In the fall of 1925, the same thought came to two Dutch physicists, George Uhlenbeck and Samuel Goudsmit. Under the advice of Paul Ehrenfest, they published their results. It met a favorable response, especially after Llewellyn Thomas managed to resolve a factor of two discrepancy between experimental results and Uhlenbeck and Goudsmit's calculations (and Kronig's unpublished ones). This discrepancy was due to the orientation of the electron's tangent frame, in addition to its position. Mathematically speaking, a fiber bundle description is needed. The tangent bundle effect is additive and relativistic; that is, it vanishes if c goes to infinity. It is one half of the value obtained without regard for the tangent space orientation, but with opposite sign. Thus the combined effect differs from the latter by a factor two (Thomas precession). Despite his initial objections, Pauli formalized the theory of spin in 1927, using the modern theory of quantum mechanics discovered by Schrödinger and Heisenberg. He pioneered the use of Pauli matrices as a representation of the spin operators, and introduced a two-component spinor wave-function. Pauli's theory of spin was non-relativistic. However, in 1928, Paul Dirac published the Dirac equation, which described the relativistic electron. In the Dirac equation, a four-component spinor (known as a "Dirac spinor") was used for the electron wave-function. In 1940, Pauli proved the spin-statistics theorem, which states that fermions have half-integer spin and bosons integer spin. Chapter 2: Covariance of the Dirac Equation (16) 2.1 Covariant form of the Dirac Equation (16). ___________________________________________________________________________________ Notation: BD use gμν = gμν = diag(1,-1,-1,-1). This makes pμpμ = + (mc)2, for example. I remember from somewhere this fact: ∂/∂xμ = ∂μ that is, the thing on the left is really a lower index thing (covariant vector) Now consider this: pμ = + i ∂/∂xμ = +i ∂μ Then pi = +i ∂i = - i ∂i = -i i = -i ∂/∂xi p = -i // which agrees with the traditional formula p0 = +i ∂t = +i ∂t // sign same for time component Now consider γμ∂μ = γ0∂0 + γi∂i = γ0∂0 + γii = γ0∂0 + γ Aside on the LT in contra/covariant notation. It is hard to find a source on this, but Messiah has it pretty well on page 877. The names are xμ = contravariant vector, xμ = covariant vector, and you use gμν to lower an index, and gμν to raise an index. So, suppose we define gμν = gμν = diag(1,-1,-1,-,1). Then: gαβ = gαγgγβ because we are raising the first index γ by applying gαγ = δαβ But we know that g x g = 1 so we write things out as on the right, a simple Kronecker but with appropriate position of the indices. Now, here is what a normal LT on a contravariant vector must look like: x'μ = aμνxν // summed indices must always be on a diagonal line This is the normal "matrix representation" where we might write in naive simple notation that x'μ = Aμνxν See "physics questions" for more notes on this subject! ( and see my long M&M chap 5 curvilinear coordinates digression notes where we always write x'μ = Tμνxν and x'μ = Tμνxν. ) _________________________________________________________________________________ Now start notes on this section 2.1. It starts off defining proper Lorentz Transformations as those with det = +1, while space or time inversion have -1. Next, we rewrite the Dirac equation as in 2.4 where we have defined γi = βαi and γ0 = β. This is the obvious thing to do, we are trying to get Dirac into something resembling covariant form after all. The fact that the γ commutators come out so nicely is encouraging! I have verified everything here including all the facts about these γμ guys. We have ∂/∂xμ = ∂μ = μ γμ∂μ = = and i = as on page 18 Then equation 2.4 really says [ i (γμ ∂μ) - mc ] ψ = 0 or [ i - mc ] ψ = 0 or [ i - mc ] ψ = 0 2.7 or [ - mc ] ψ = 0 2.8 and this have certainly gotten a lot more compact! It certainly "looks" covariant in that seems to be a scalar in some sense. 2.2 Proof of Covariance (18). The statement of the problem is very clear, and I am happy to try to use the same γμ matrices in both the primed and unprimed frames. If we get a solution assuming that, fine. Here is how the spinors are going to transform: ψ'(x') = S(a)ψ(x) The whole thing is solved if we can find a matrix S(a) such that S(a) γμ S-1(a) aνμ = γν // BD page 20 equation A since this makes equation E on page 19 have the form [ i ∂'μ - mc] ψ'(x') = 0, which is to say, this makes the Dirac equation have the same form in the primed frame as it did in the unprimed frame. This is called "form invariance" I think. If we apply S from the right and S-1 from the left on the above, it becomes ( remember aμν is just a number) aνμ γμ = S-1(a) γν S(a) // BD page 20 equation 2.12 which for comparison to something below I swap sides to write as S-1(a) γν S(a) = aνμ γμ Aside: At this point, I would like to remind myself of the rotation world where we have things like this R-1 V R = R V = V' // compare to S thing above where R = Rx(θ) = exp(-iθJx) for example, and Jx is a generator. The thing R is a 3x3 representation of the rotation group SO(3), see sheet on TK binder. On the other hand, the object R is a 2x2 rotation matrix which operates in the | ½ ;± ½ > j=1/2 "spin space" of Levitt's book. The object V is some "vector operator" and we are trying to see what the sandwich does to this operator. So the point is that we have two different representations of the rotation group sort of tied together by my equation, namely, the 2D and 3D representations. This is summarized on my one-page TK binder item. We have a similar situation going on here in BD. The 3x3 role of Rij is played by aνμ which is a 4x4 representation of the Lorentz group. The object S(a) which corresponds to my R is the representation of a LG element in the "Dirac Spinor" 4D representation. Note that it is a 4x4 representation, but is not the same as the 4x4 "vector" representation aνμ. Just as my R is what rotates the little j=1/2 angular momentum states, this thing S(a) is going to be the thing that "rotates" the 4-spinors of the Dirac theory. I quote from a wiki page on LG representations: (0,0) the Lorentz scalar representation. This representation is carried by relativistic scalar field theories. (1/2,0) is the left-handed Weyl spinor and (0,1/2) is the right-handed Weyl spinor representation. (1/2,0) ⊕ (0,1/2) is the Dirac spinor representation. (1/2,1/2) is the vector representation. The electromagnetic vector potential lives in this rep. It is a 1-form field. (1,0) is the self-dual 2-form field representation and (0,1) is the anti-self-dual 2-form field representation. (1,0) ⊕ (0,1) is the representation of a parity invariant 2-form field. The electromagnetic field tensor transforms under this representation. (1,1/2) ⊕ (1/2,1) is the Rarita-Schwinger field representation. (1,1) is the spin-2 representation of the traceless metric tensor. So, the aνμ matrices are the (1/2,1/2) "vector" representation of the LG, whereas the S(a) matrices are the "Dirac Spinor" representation which is denoted (1/2,0) ⊕ (0,1/2). It happens that these are both 4x4 representations, but they are not the same. So that is the "group theory" of what we are doing here. This is elaborated in Messiah Vol II I am happy to say, where the same metric is used for gμν luckily for me. Now back to the issue of R-1V R = R V noted above. Suppose we knew the 3x3 matrices R and suppose we did not know the 2x2 matrices R. We could construct them starting with infinitesimals and generators, and then exponentiate to get the full R group representation elements. That is what B&D are going to do with S-1(a) γν S(a) = aνμ γμ. My TK binder, by the way, has some old notes on the conventions used by various authors for objects related to the LG, such as parameters and generators. More comments. Consider again S-1(a) γν S(a) = aνμ γμ = γ'ν ψ'(x') = S(a)ψ(x) BD 2.12 Think of γν as an operator in the 4D Dirac space (it after all a 4x4 matrix). S is the operator which -- acting in the 4D Dirac space -- makes the vector operator γν become what it is in the transformed (primed) frame γ'ν. The infinitesimal construction method. First, write in infinitesimal aνμ as in 2.13a. This is a little strange because the parameter and the generator are combined into the same object Δωνμ . For example, if we were doing (passive) rotations, we would say R = e+iθJ ≈ 1 + iθJ or Rij ≈ δij + iθJij So we have to think of aνμ = δνμ + Δωνμ where Δωνμ = +iuKνμ for a boost, or +iθJνμ for a rotation, where J and K are the 4D matrices shown in my handwritten notes "summary of Lorentz Group factors" . Thus, in this strange notation, they are saying Δω12 = +iθ(J3)12 = +iθ{ -iε312} = +iθ{-i} = θ = +Δθ = a number Of course they are using Δθ instead of θ to suggest an infinitesimal rotation at this point. Maybe should think of it this way Δω = Σi=1,3 [ +i uiKi +iθiJi ] Δωμν = +i Σi=1,3 { ui (Ki) μν + θi (Ji) μν } where this is an arbitrary LT with our usual 6 parameters. Later as in 2.18 you see B&D break this into a small parameter times a generator matrix. In general, we can think of Δωνμ as being some linear combination of the J's and K's with appropriate small parameters, and thus Δωνμ describes an arbitrary small Lorentz transformation. Now where does 2.13b come from? Consider 2.3 which says δνσ = aμνaμσ = (δμν + Δωμν) (δμσ + Δωμσ) But the two deltas give the LHS so we must then have, to first order in smallness: δμν Δωμσ + Δωμνδμσ = 0 or Δωνσ + Δωσν = 0 or Δωνσ = – Δωσν We can then lower the ν index on both sides by applying gν'ν to both sides, or by just "doing it", and this gives Δωνσ = – Δωσν as claimed, and this would be true in all four index combinations. Next, how do we justify the strange form 2.14 given for infinitesimal S? This took me a while to figure out, see "proof of Taylor series expansion". My conclusion in those notes is that you can think about expanding a matrix F that is a function of a matrix X which is near a matrix A. Although I can write down terms of any smallness order, only the lowest order linear term is simple to write in a compact form, and we get F(X) = F(A) + Σr,s=1,N (∂F/∂Xrs) (Xrs-Ars) = F(A) + Σrs=1,N Krs δXrs where each Krs is itself a matrix just labeled by "rs". In our case here, we have A = 1 and N=4, so we have F(X) = F(1) + Σr,s=1,4 (∂F/∂Xrs) (Xrs-δrs) = F(1) + Σrs=1,N Krs (Xrs-δrs) Now we change F → S and X → a to come closer to the BD form S(a) = S(1) + Σr,s=1,4 (∂S/∂ars) (ars-δrs) = S(1) + Σrs=1,N Krs (ars-δrs) Next, we know that S(1) = 1, and we know (ars-δrs) = Δωrs so we get S(a) = 1 + Σr,s=1,4 (∂S/∂ars) Δωrs = 1 + Σrs=1,N Krs Δωrs Now we rename the 16 matrices Krs = (-i/4) σrs and we get S(a) = 1 + Σr,s=1,4 (∂S/∂ars) Δωrs = 1 - (i/4) Σrs=1,N σrs Δωrs where - (i/4) σrs = Krs = (∂S/∂ars)|a=1 Then finally we change to Greek letters with implied summation S(a) = 1 + (∂S/∂aμν) Δωμν = 1 - (i/4) σμν Δωμν and we can replace Δωμν = Δωμν and we are there! Now, since we know that Δωμν is antisymmetric, if we decompose σμν into S and A parts, only the A part will survive, so we might as well use that part and then rename it to be just σ σμν Δωμν = [σSμν + σAμν] Δωμν = σAμν Δωμν Comment: I just spent 6AM to 2PM or 8 hours absorbing 4 vertical inches of B&D page 20! That is a pretty amazingly slow rate of progress, but I had to digress into other areas to "get happy". All the time I am reviewing 28 year old history back to 1980 or so. Things are a little rusty as you might expect. Derive bottom Equation G page 20. All we need to do here is insert the three through-first-order expressions for S, S-1 and a into 2.12. The leading terms on both sides cancel, and then we set the first order terms to 0 and this gives the result shown on the bottom of the page. Derive page 21 equation 2.15. Write the LHS of page 20 equation G as shown in pencil in the margin, then both sides have Δωαβ and we set the other parts equal. Before we cancel off the Δωαβ, we use the fact that Δωαβ is A (antisymmetric) so write it's mating factor as A + S and keep only the A part as usual, and we get the 2.15 result. Show that 2.16 satisfies 2.15 . At first look this is totally non-obvious. You start off with this: RHS = (i/2) [γμ, [ γα, γβ] ] where I just raised the indices on 2.16 and 2.15. But this is a sum of four terms each completely general products of three gammas, how do you reduce this to the LHS which has only one gamma and a metric tensor. This right now is very non-obvious to me, but I have the thing checked from last time I was here. The Jacobi identity does not seem to help. You somehow have to use the { γα, γβ } = 2gαβ fact to remove pairs of gammas. For compact notation (this is my own deal), just write the index and omit the γ so we have [ν,[α,β]] = [ν, αβ - βα ] = ν (αβ - βα) - (αβ - βα)ν = ναβ - νβα - αβν + βαν = 1 2 3 4 = ( ναβ - αβν ) - ( νβα - βαν ) 1 3 2 4 Each time we move a letter one position, we pick up a minus sign and some g tensor stuff. Suppose we take the third term and slide ν to the left two positions. We end up with - ναβ which cancels the first term, and we pick up two g things. Then suppose we take the last term and move ν two positions to the left. This will cancel the second term, and we are left only with g stuff, so that is the solution. Now go do it in more detail: the rule is αβ + αβ = 2gαβ so αβ = - αβ+ 2gαβ or αβ = - αβ+ 2ab where ab is my shorthand for gαβ. So, start with term 3. We have - αβν = -α(-νβ + 2bv) = +ανβ - 2α(bv) Now move the ν to the left in ανβ to get ανβ = ( -να + 2va)β = -ναβ + 2β(va) Thus, our third term becomes - αβν = ανβ - 2α(bv) = -ναβ + 2β(va) - 2α(bv) And we can then say 1st + 3rd term = ναβ -ναβ + 2β(va) - 2α(bv) = + 2β(va) - 2α(bv) We can handle the 2nd + 4th term by taking the previous result, swapping α and β (and a and b) and adding an overall minus sign, as shown on second isolated line about 6 inches up this page. So 2nd + 4th term = - { 2α(vB) - 2β(av)} Therefore, the sum of all four terms is this: 1+2+3+4 = + 2β(va) - 2α(bv) - 2α(vb) + 2β(av) = 4β va - 4α bv = 4(β va - α bv) Thus we have shown that [ν,[α,β]] = 4(β va - α bv) which we now decompress to get [γν,[γα,γβ]] = 4(γβ gνα -γα gνβ) Therefore we have [γν, σαβ] = (i/2) [γν,[γα,γβ]] = 2i (γβ gνα -γα gνβ) We can now lower the αβ indices if we want to get [γν, σαβ] = (i/2) [γν,[γα,γβ]] = 2i (γβ gνα -γα gνβ) // verifying σ makes 2.15 and this is the way things appear in 2.15, QED and add this to the gamma page please! So this is an important gamma matrix reduction formula reducing triple products to singles! So at this point, we have found an explicit form for the infinitesimal S. Now consider a passive x-axis boost by u. Look at TK "summary of LG facts" where we have Bx(ω) = exp(+iωK1) where K1 = as shown in those notes. Note that iK1 = BD 2.19 for I. So Bx(ω) = exp(+ωiK1) = exp(+ω I) = my matrix with four + signs. So looking at my notes for Bx(u), we get exactly 2.20 of BD. I can skip the details of getting inf to finite because I know what the answer is already. So we are now mid page 22. But let's look at this interesting method anyway. He considers a set of N small transformations of the infinitesimal form and sets Δω = ω/N as shown. Why does this become exponential in the limit? The implicit theorem is this: limN→∞ (1 + [ω/N] x)N = eωx The way to prove this is to expand the LHS using the binomial expansion, take the large N limit, and this becomes the expansion of eωx. For example (1 + [ω/N] x)N = 1 + N (ωx/N) + N(N-1)/2 (ωx/N)2 + N(N-1)(N-3)/3! (ωx/N)3 + .... → 1 + (ωx) + (ωx)2/2! + (ωx)3/3! + ... QED The last step comes from this fact ch(ωI) = 1 + (ωI)2/2! + (ωI)4/4! + .... = 1 + { ch(ω) - 1 } I2 and when you look at matrices I and I2 on page 21 you can fill everything in to get 2.20 which is the correct form of a passive boost. Suppose we took a particle at rest p = (m,0,0,0) and boost it by v as shown in 2.20. We know that we get that p' = (γmc, γmv,0,0). We know then that γ = cosh(ω) and γβ = sinh(ω) and tanh(ω) = β. So this confirms equation E on page 22. Now let's back up and be more general. Looking back at 2.13 we had a = 1 + Δω + .... = exp(ω (In)) where Δωμν = Δω (In)μν Δω = ω/N Similarly, we are going to have S = 1 - (i/4)σμνΔωμν + .... = exp[ - (i/4) ω σμν (In)μν ] as shown in 2.22. Now look again at 2.19 for Iνμ. What is Iνμ ? Iνμ = gμα Iνα But Iνα = -1 ( δν,0δα,1 + δν,1δα,0) So Iνμ = gμα Iνα = -1 gμα ( δν,0δα,1 + δν,1δα,0) = -1 (δν,0 gμ1 + δν,1 gμ0) = -1 ( - δν,0 δμ,1 + δν,1 δμ,0) = ( δν,0 δμ,1 - δν,1 δμ,0) Therefore σνμ (In)νμ = σνμ ( δν,0 δμ,1 - δν,1 δμ,0) = σ01 - σ10 = 2 σ01 and then in this case we have S = exp[ - (i/4) ω σμν (In)μν ] = exp[ - (i/4) ω 2 σ01 ] = exp[ - (i/2) ω σ01] just as claimed in 2.23. So this tells you what a passive x boost does to a 4-spinor! Now let's repeat this last exercise when I = iJ3 instead of iK1. My hand written notes say that J1 = so looking at my rotation TK summary page, I know that J3 = I = iJ3 = = Iνμ = like 2.19. We can see that Iνα = 1 ( δν,1δα,2 - δν,2δα,1) Iνμ = gμα Iνα = 1 gμα ( δν,1δα,2 - δν,2δα,1) = -1 ( δν,1δμ,2 - δν,2δμ,1) = ( δν,2δμ,1 - δν,1δμ,2) σνμ (In)νμ = σνμ ( δν,2δμ,1 - δν,1δμ,2) = σ21 - σ12 = - 2 σ12 S = exp[ - (i/4) ω σμν (In)μν ] = exp[ + (i/4) ω 2 σ12 ] = exp[ + (i/2) ω σ12] and this agrees with 2.24. Note that σ12= σ12 and they have σ12, fine. Now compute σ12 = (i/2) [ γ1, γ2] = i γ1γ2 = Σ3 from my hand notes and this is as shown below 2.24. I agree that the σij are hermitian and that the rotations are unitary so S† = S-1 . And I agree that the σ0j are antihermitian, so we get S† = S instead. They now make the claim that S-1 = γ0S†γ0 is true for both rotations and boosts. Here is a proof: Lemma: claim that {γo, σ0j} = 0. Proof: {γo, σ0j} { γo , [γ0, γj] } = { 0, 0j - j0} = 0(0j - j0) + (0j - j0)0 = 00j - 0j0 + 0j0 - j00 = j - 0j0 + 0j0 - j = 0 QED This tells us that γ0 σ0jγ0 = - σ0j Therefore, when we take a boost SL and sandwich γ0 around it, we change the sign in the exponent, and that is what we need to show that S-1 = e+ = γ0S†γ0 = γ0Sγ0 = γ0e-γ0 = e+ for boosts For rotations, we know σij commutes with γo so γ0 σijγ0 = + σij so the sandwich does nothing at all, and we have S-1 = γ0S†γ0 = S† for rotations. Therefore, as claimed, 2.26 is valid for both boosts and rotations. BD then show easily and clearly that the 4-current jμ transforms as a 4-vector as we expected, and we use our S rule 2.12 to show this. We already know that ∂μjμ = 0 and this is then seen as a scalar condition. Anticipating that ψ†γ0 shows up a lot, we call it = ψ†γ0. One example is that it shows up in the current so we really have jμ = c γμψ Review what we learned in this section! For the normal vector representation of the Lorentz Group (LG), we denote the usual (passive) transformation matrices by a and elements by aμν with "the usual tilt" for a normal matrix. The matrices a are representations of elements of the LG. They act on 4-vectors (not Dirac 4-spinors). The "vector" representation of the LG has the designation (1/2,1/2). Remember that the Lorentz Group has two casimirs the way the rotation group has one casimir J2 and the group designation lists the values of the casimirs. The 4-spinors ψ do not -- although they have four elements -- transform as 4-vectors. Instead, they transform according to a different representation of the LG known as the "Dirac Spinor" representation which has the group theoretic designation (1/2,0) ⊕ (0,1/2). It would appear that this is a "reducible" representation of the LG. As we change reference frame, we know that a Lorentz Group vector object transforms as x'μ = aμνxν . We could write this as x' = ax. Think of x as if it were a ket, and we are saying x'μ = aμνxν in the sense of acting on a vector in Hilbert space which vector is labeled by the index x. If we wanted to deal with a "vector operator" Vμ in this space, we would say A-1V A = a V = V'. Normally we don't do this, however. For example, the current jμ at least in this book is not an operator, it is a 4-vector ket thing and it transforms the same way that xμ does. The generators of the vector transformations are written in this section as Δωνμ = δω Iνμ so really the Iνμ are the generator matrices which I like to call J and K in my hand written notes, and they are also sometimes written in Mμν form but we have avoided that so far. BD don't say much about the generators In, but we can think of n = 1,2,3,4,5,6 if we like. As we change reference frame, we know that a Lorentz Group Dirac Spinor object transforms as ψ'(x') = S(a)ψ(x) or write this as ψ' = S ψ. This is in the "ket sense" mentioned in the previous paragraph. If we had a "Dirac Spinor operator" D, it would transform as S-1D S = a D = D' . In this chapter, we have discovered that the four matrices γμ do in fact form a Dirac Spinor Operator, and these matrices transform in exactly this way as shown in equation 2.12. The generators of the Dirac Spinor transformations are written as Gn = (-i/4) σμν(In)μν again for n = 1,2,3,4,5,6. The (In)μν are numbers -- elements of the Vector generator matrices. The σμν are matrices for each choice of μν. So of course Gn are 4x4 matrices which act as generators in this Dirac Spinor space! BD never write down the 4x4 Gn matrices explicitly, but instead they look at special cases. For example, when we have a passive x boost, GB1 = (-i/4) 2 σ01, and when we have a passive z rotation, GR3 = - (-i/4) 2 σ12. From these examples, I think we conclude that Gboosti = (-i/4)2σ0i Groti = - 2 (-i/4) εijkσjk so in essence that we can regard σμν as being the Dirac Spinor generators. Since antisymmetric, there are only (16-4)/2 = 6 different matrices, and they are the generators as shown (times constant). It happens that σμv = (i/2) [γμ, γν ] so we could say that the Dirac Spinor generators are given as a commutator of some Dirac Spinor operators! The theme of this section is that we can show the Dirac equation is covariant if we can show it has the same form in primed and unprimed frames, AND if we can show how someone in one frame can compute what the other frame's observer is seeing, ψ'(x') = S ψ(x), but of course this last item is just saying that we want to know all about the details of things, not just the fact that something is covariant in form. We start by showing that we can show the form of the Dirac equation is covariant if we can find S such that 2.12 is true. Now I suppose we might have found that no such S exists, and then our covariance proof would fail. But this is not what we found. When we form the following grouping kμ = γμ ψ where = ψ†γ0 , we have learned that although ψ is a ket of the Dirac Spinor Representation, and although γμ transforms as a Dirac Spinor Operator, the combination kμ transforms as a Vector object! This is shown very precisely in 2.27 ! The reason of course is that the operator and ket transformations "cancel each other". Here is a fast way to see it kμ' = <' | γμ | ψ'> = < | S-1γμ S | ψ> = aμν< | γμ | ψ> = aμν kν = < | γμ'| ψ> 2.3 Space Reflection (24). What S satisfies 2.12 for an "a" (ie, aμν) which just inverts the spatial coordinates? S = γ0 ≡ P does just fine, to which you can add a phase if you want to get 2.32. ψ is not an eigenstate of this P operator. However, if you look at the four rest-state solutions on page 10, the positive energy ones have positive parity, and the negative energy states have negative parity if we choose P = + γ0. This is just because P = γ0 = diag(1,-1) where 1 is a 2x2. Any linear combination of the first two states has positive parity, so this would include any non-rel-limit solution of the Dirac equation. Time Reflection? We have to wait until chapter 6 to ponder this, not sure why we could not just do it right here. 2.4 Bilinear Covariants (25). In this section we examine 16 independent 4x4 matrices which span the Dirac spinor space. They are put into 5 groups where the superscript indicates how the things in that group transform when you sandwich them between and ψ. S = scalar, P = pseudoscalar, A = axial vector (that is to say, a pseudovector), V = vector (ie, a polar vector), T = tensor (of rank 2). Again, we are not saying that the 4x4 matrix ΓVμ = γμ itself transforms as a vector, but that γμ ψ does. BD provide a list of properties on all these matrices (which can all be expressed in terms of the four γμ and γ0. Note that γ5 is a certain combination of the others and γ5 = offdiag(1,1) so it swaps the upper and lower component pairs of whatever Dirac spinor thing it acts on. ) Given the list of properties, they prove easily that the 16 matrices Γ really are linearly independent and thus span the space. Any combination (stuff) ψ is referred to as being "bilinear" because the ψ appears twice. The combinations on page 26 can each be identified with some transformation rule under the Lorentz group "a". Since each of these bilinear forms transforms in some well-defined group manner, they are called "covariants". The factor det(a) is used to provide a minus sign for a parity transformation and a plus sign for all proper Lorentz transformations. Chapter 3: Solutions of the Dirac Equation for a Free Particle (27) 3.1 Plane Wave Solutions (28). Comment: We learned in previous chapters how Dirac stumbled upon the "Dirac equation". The problem was to find a Hamiltonian H to put into the SE which says Hψ = Eψ, both operators. The H had to respect the relativistic formula for energy E such that E2ψ = H2ψ = (p2c2 + m2c4)ψ = ([-i]2c2 + m2c4)ψ and H had to result in a conserved probability current jμ with a positive definite ρ. "Early attempts" had various failure modes. By going with a 4 dimensional matrix approach, Dirac found a theory that met all the requirements. We have since come to understand that ψ "belongs" to the (1/2,0) ⊕ (0,1/2) representation of the Lorentz group. As before, we first deal with technical details, then resume notes on this section. Derive 3.5 . First, we have S = exp(-iωσ01/2) for an x boost which we confirmed in 2.23, but I thought that was passive and now this one is active. Well, keep with the above sign for the moment. We find that σ01= iγ0γ1 // from our sheet = -i γ0γ1 = -iα1 => -iσo1 = (-i)( -iα1) = -α1 so we want S = exp(-iωσ01/2) = exp(-α1ω/2). So now we have to go off and look at powers of α1. α1 = γ0γ1 α12 = γ0γ1γ0γ1 = - γ0γ0γ1γ1 = - (+1)(-1) = 1 α13 = α1 α14 = 1 Thus we have exp(-α1ω/2) = Σn=0,∞ (-ω/2)2/n! α1n sum even terms = Σn=0,2,∞ (-ω/2)2/n! * 1 =ch(-ω/2) * 1 sum odd terms = Σn=1,3,∞ (-ω/2)2/n! * α1 = sh(-ω/2) * α1 Therefore exp(-α1ω/2) = ch(ω/2) * 1 - sh(ω/2) * α1 // agrees with page 29 equation A But the matrix α1 is this one: α1 = => exp(-α1ω/2) = and this is exactly equation 3.5 where the ch(ω/2) has been factored out. BD now want to continue with velocity boost in an arbitrary direction, and I accept page 29 C as the right matrix because 2.19 showed us what it was for a pure x where we had -1 where -cosα sits. I am fine for the first equality in 3.7, and now I have to do this: Compute exp[ -(ω/2) α ] as in 3.7: Well, things are pretty simple because α = and then (α)2 = since (σ)2 = 1 sum even terms = Σn=0,2,∞ (-ω/2)2/n! * 1 =ch(-ω/2) * 1 sum odd terms = Σn=1,3,∞ (-ω/2)2/n! * α1 = sh(-ω/2) * α so that exp[ -(ω/2) α ] = ch(ω/2)*1 - sh(ω/2) * α write this now as ch(ω/2) - sh(ω/2) = ch(ω/2) { - th(ω/2) } Now we know we can write σ = σv/ v = (1/v) where v±= vx ± ivy = (1/p) where p±= px ± ipy Now from previous page we have T -th(ω/2) / p = c/(E+mc2) so we put this factor c/(E+mc2) into all the relevant terms in -th(ω/2) . Then the outside factor ch(ω/2) is as shown in 3.6 and we have now exactly produced 3.7 ! We can write this in a compact notation this way: 3.7 = Review: We start with ψ(x) for particle at rest. We boost our observation frame in - direction so from frame F' we see electron with momentum +p. I prefer just doing an "active boost" on the particle to get it up to this speed, but fine. In the rest frame, F, we already solved the Dirac equation on page 10 [ also page 28 where εr = +1 for r = 1,2] and found the various "unit spinors" times their phases. This phase sign is minus for r = 1,2 and these were called "the positive energy solutions". So the idea here is this: (1) we already solved the Dirac equation for p = 0 and we got unit spinors times covariant phaser. (2) We know what S is for an arbitrary-direction boost -- it is S = exp[ - (i/4) ω σμν(In)μν] where we use In as appropriate for an arbitrary-direction boost as shown in page 29 C. (3) We compute this S as shown in 3.7. (4) This matrix S then converts the rest solutions to non-rest solutions. The unit spinors each become a column in this matrix S. (5) Thus, we now have the solution of the Dirac equation for a free particle of arbitrary momentum p. I guess this was easier than trying to solve the equation directly for a particle doing momentum p. Something to note: Our four starting unit spinors now have the appearance as in 3.7 page 30. Recall now that p = γβmc E = γmc2 so cp/E = β At very high energy, the non-unity factors in 4x4 matrix 3.7 have these limits pz * factor → cosθ // that is, pzc/(E+mc2) → pcosθc/E = cosθ β = cos θ p± * factor → sinθ cosφ ± i sinθ sinφ = sinθ e±iφ so that non-unity spinor entries are now "as large as" the unit entries, and strange things perhaps will happen. BD did not take this limit. As we did our boost, our exponential phase started at -iεr mc2t/ = -iεr (mc)(ct)/. In this starting frame F, we know that xμ = xμ = (x0, r) = (ct, r) and pμ = (mc,0), so phase is -iεr xμ pμ/ at least in the rest frame. Then this is the form in the boosted frame, because the phase has to be a world scalar! If it were something else, I think you would have a big mess. So we now break off the spinors from the phase factor and call them wr(p). Derive 3.9a. This equation says that you get 0 when you apply ( - εrmc) to any of our four solution spinors wr(p), you get 0. Here p is just pμ, not an operator. So this equation says that our solutions solve the Dirac equation "in momentum space". This sort of verifies that we found the correct solutions. From earlier, we had this Dirac equation: [ i - mc ] ψ = 0 [ i - mc ] ψ = 0 [ - mc ] ψ = 0 all operators. Now if we have exp( -iεr xμ pμ/), then i exp( -iεr xμ pμ/) = iγμ∂μ exp( -iεr xμ pμ/) = iγμ ∂/∂xμ exp( -iεr xμ pμ/) = i γμ (-iεr pμ/) exp( -iεr xμ pμ/) = εr γμ pμ exp( -iεr xμ pμ/) = εr exp( -iεr xμ pμ/) Thus, the operator is replaced by non-operator εr for a plane wave solution. So the Dirac equation can be written: [εr- mc] wr(p) = 0 which is basically 3.9a Here is how you would transpose this thing (note that pμ is real) [-εr mc] wr(p) = 0 => [γμpμ-εr mc] wr(p) = 0 => wr(p)† [γμpμ-εr mc]† = 0 => wr(p)† [γμ†pμ-εr mc] = 0 Apply γ0 from the right to get => wr(p)† [γμ†pμ-εr mc] γ0 = 0 => wr(p)† γ0 γ0 [γμ†pμ-εr mc] γ0 = 0 => r(p) γ0 [γμ†pμ-εr mc] γ0 = 0 => r(p) [γμ pμ - εr mc] = 0 QED. Derive 3.9b. The claim here is that r(p) wr'(p) = εr δr,r'. That is to say, the column vectors in 3.7 are orthogonal and almost orthonormal except for the εr sign thing. We know the thing above is a scalar so we can evaluate it in the rest frame. We have r(p) wr'(p) = w†r(p) γo wr'(p) = < r(p) | wr'(p) > The γ0 matrix has the form diag(1,-1) and this is then where the εr comes from QED. So I like to write this result as < r(p) | wr'(p) > = εr δr,r' // orthogonality in the 4x4 Dirac space Derive 3.9c. Claiming that Σr=1,4 εr wrα(p) r(p)β = δα,β My inclination is to write this as follows: Σr=1,4 εr <α| wr(p)>< r(p)|β> = δα,β // completeness in the 4x4 Dirac space where you can think of |β> as being a unit column vector such as |β> = (1,0,0,0) -- you put in whatever unit vectors you want for α and β. Then the operator form would be Σr=1,4 εr | wr(p)>< r(p) | = 1 which has the general form of a "completeness relation". To "prove this", I would first apply | wr'(p) > on the right to get Σr=1,4 εr | wr(p)>< r(p) | wr'(p) > = | wr'(p) > Now use the orthogonality 3.9b shown above to get Σr=1,4 εr | wr(p)> εr δr,r' = | wr'(p) > But the LHS is just | wr'(p)> . Note Added: Here is another way to think of the above completeness relation. We can write Σr=1,2 (+1) | wr(p)>< r(p) | = 1+ Σr=3,4 (–1) | wr(p)>< r(p) | = 1– 1 = 1+ + 1– where we break things down into the positive and negative energy solution spaces. If you are dealing only with positive energy spinors, we know that wr(p) with r = 1 and 2 form a complete basis, so only these need be included in the operator completeness statement above. You formally prove this by considering arbitrary linear combinations of positive energy spinors, etc etc. Now notice this fact 1+ | w3(p)> = 0 because < r(p) | w3(p) > = 0 for r = 1 and 2. 1+ | w1(p)> = | w1(p)> because that is what 1+ means! And just do it. We may therefore conclude that 1+ = Λ+ , the positive energy projection operator (to be introduced on page 33). So although these forms don't look very similar, they must be the same. Thus we have added our own little fact here: Σr=1,4 εr wrα(p) r(p)β = δα,β Λ+(p) = (+mc)/2mc = Σr=1,2 | wr(p)>< r(p) | = Σr=1,2 wr(p) r(p) = 1+ = + Σs u(p,s) (p,s) Λ-(p) = (–+mc)/2mc = – Σr=3,4 | wr(p)>< r(p) | = – Σr=3,4 wr(p) r(p) = 1– = – Σs v(p,s) (p,s) Notice the minus sign in the last line, and notice that everything is a 4x4 matrix. These results are used on page 95 to obtain 6.48 from 6.47. Probably this all has a formal presentation in terms of the direct sum and direct product subspace notations, but let's not do that now! Derive 3.11. If we consider r,r' only from the set 1,2. then 3.11 says this w†r(p) wr'(p) = (E/mc2) δr,r' For r = r' = 1, I have shown this to be true, and we use the fact that pz2 + p+p- = p2. In fact, this fact is obviously true for all four columns dotted with themselves, so we know that w†r(p) wr(p) = (E/mc2) r = 1,2,3,4 The following fact is also true w†r(εrp) wr(εrp) = (E/mc2) r = 1,2,3,4 because for r = 3,4 we are negating p in both terms which makes no bilinear difference. So the issue here is with the orthogonality! For this purpose, we can think of the vectors in the following simplified way: w1 = w2 = w3 = w4 = w1† = (1 0 pz p-) w2† = (0 1 p+ -pz) w3† = (pz p- 1 0) w4† = (p+ -pz 0 1) By brute force I find these things to be true: w1†(p) w2(p) = 0 + 0 + pz p- - pz p- = 0 w1†(p) w3(-p) = pz - pz = 0 w1†(p) w4(-p) = p- - p- = 0 w2†(p) w3(-p) = p+ - p+ = 0 w2†(p) w4(-p) = pz - pz = 0 w3†(p) w4(p) = pz p- - pz p- = 0 and of course w3†(-p) w4(-p) = 0 as well. So by brute force I have shown that everything is orthogonal as claimed in 3.11 and the diagonal values are also as claimed. So I have now proven 3.11 to be true by the brute force method. Obviously there must be a simpler way to show this. Here was one attempt: ****************** Lemma: γo wr(p) = εr wr(-p) r = 1,2,3,4) where γ0 = Proof: we know that γ0 negates the signs of the lower two components of our spinor wr(p). For the first two columns in 3.7, we can undo this sign change by going p → -p, so we can say γo wr(p) = wr(-p) r = 1,2 For the last two columns, we can negate the first two components with p → -p, then we can negate all four components with an overall minus sign, and together these two actions will compensate for the action of γ0 on our spinor, so γo wr(p) = -wr(-p) r = 3,4 We can combine these by saying γo wr(p) = εr wr(-p) r = 1,2,3,4 Now consider εr δr,r' = w†r(p) γo wr'(p) = εr w†r(p) wr'(-p) which tells us that w†r(p) wr'(-p) = δr,r' This gives many of the desired results, but not all of them. This gives the desired answers for cross-group combinations, but within the group r ,r' = 1,2 or 3,4 this is not the desired result, though it is true. We have to show this w†r(p) wr'(+p) = δr,r' r,r' in 1,2 or in 3,4 This is all true, but the only way I know to show it is brute force. Enough! ****************** Comments ψr(p): Our full solution is ψr(p) = exp(-iεr pμxμ / )wr(p) What happens when we apply the 4-momentum operator? μ exp(-iεr pμxμ/) = +i ∂μ exp(-iεr pμxμ/) = + εrpμ exp(-iεr pμxμ/ ) Thus we have this slightly strange situation: (as if the neg energy solutions "go backwards") μ ψr(p) = εrpμ ψr(p) If we look at the 3-momentum operator we have (both sides up now) ψr(p) = εrp ψr(p) So the "negative energy" solutions not only have negative energy, they also have momentum -p. Parse sentence on page 31 middle of page: (1) Equation 3.11 when applied to spinors of opposite energy says wr†(p)wr'(-p) δr,r'. Assume that r' is the negative energy spinor. We know from above that wr'(-p) really has momentum +p. So these two spinors here have opposite energy, but the both have momentum +p . Thus "two plane wave solutions of the same spatial momentum p but of opposite energy are orthogonal in the w†w sense." Comments on page 32 equation A: σs w = w First, think of this in a 2x2 sense. Here is what we would do: σ |α> = σ3|α> = +1 |α> where |α> = column (1 0) σ |β> = σ3|β> = –1 |β> where |β> = column (0 1) Now start with the upper line and do this: Rσ R-1R |α> = +1 R |α> => (RσR-1) |θ,φ> = +1 |θ,φ> where R= R(θ,φ) rotates the up state to wherever you want it. R = 2x2 matrix of SU(2). Then use our quoted fact above R-1V R = R V and we have , where R is the usual 3x3 matrix of O(3), ( R -1σ) |θ,φ> = +1 |θ,φ> Now move to the other side to get R σ |θ,φ> = +1 |θ,φ> And call this thing = R and we get our desired result: σ |θ,φ> = +1 |θ,φ> |θ,φ> = R(θ,φ) |α> which indicates a state where the spin is polarized in the direction. So here are the conclusions: σ |θ,φ, up> = +1 |θ,φ, up> |θ,φ> = R(θ,φ) |α> σ |θ,φ, dn> = -1 |θ,φ, dn> |θ,φ> = R(θ,φ) |β> We also know that |β> = Ry(π) |α) and etc from Levitt world. The point is that if you can make a state that is an eigenstate of σ, then that state has spin polarized in the direction (or opposite). In the 4x4 world, we have Σ = Based on this Σ and the above discussion, the spinors wr(p) for r = 1,2,3,4 have spins +,–.+,– . However, we are going to define the v in such a way that when you list the spinors in their normal order, the spin sequence comes out being +,–.–,+. The goal is to change the spins of the negative energy solutions relative the name you use for the spinor. This happens soon below. The spin vector situation. In a rest frame, we have sμ = (0,s) and pμ = (m,0). This suggests that we have in general sμ pμ = 0 as a world scalar, and also sμ sμ = sjsj = - sjsj = -ss = -1 , also a world scalar. Now consider u(p,s) being a positive energy solution with rest spin s. Then we expect: Σ u(p,s) = u(p,s) σ |θ,φ, up> = +1 |θ,φ, up> // comparing to the above How would you write out u(p,s) ? In rest frame we can do it this way u(p,s) = R(θ,φ) u(p,s = +z) = R(θ,φ) x R(θ,φ) But the point is that we are defining u(p,s) to be positive energy and s = s in rest frame, but of course things are 4-vectors. So if = (0,0,1) we have uzμ = (0,0,0,1) . Then my conjecture above says u(p,uz) = u(p,-uz) = = Ry(π) x Ry(π) Here I am seeing a suggestion of the (1/2,0) ⊕ (0,1/2) sense of this Dirac representation. Of course I have written the above in the rest frame. You can then boost out so that u(p,uz) = w1(p) to any p you want. You can go to the rest frame to observe where s is pointing. Of course in a boosted frame, it gets distorted into some sμ. 3.2 Projection Operators for Energy and Spin Energy Projectors. These are easy and obvious. We have Λ± = (± + mc)/(2mc). This works because we know that - mc = 0 for r=1 and 2, meaning = mc when applied to 1 or 2. Thus,( + mc)/(2mc) = ( mc + mc)/(2mc) = 1 as an operator when going on + states. And for negative energy states, we know that = - mc so again we get 1 when acting on them with Λ-. Spin projectors. Comments on gamma matrices. The four basic γμ are the originally defined ones (in terms of αi and β). We are then allowed to define lower index ones by γμ = gμνγν even though we would not say γμ was a 4-vector by itself. But we know that γμ does transform as a 4-vector when sandwiched between and ψ. So this extension allows us to say things like jμ = γμ ψ => jμ = γμ ψ . [ I later learned that γμ transforms as a 4-vector on the action of Dirac rep xforms, just as pμ does on Vector rep xforms, see LG doc.] The various other gamma matrices do not vary up or down and probably you should always write them the same way. Looking on page 25, we see that you might be tempted to have γ5 = -γ5 from the way γ5 is defined, but B&D go out of their way to clearly state that γ5 = γ5 by definition! Later we encounter a γ6 which is the same idea γ6 = γ6. So the moral here is to be careful with γ3 = – γ3 in the spin discussion below. For example, when we have the notation z the valuation in the rest frame would be + γ3s3 = + γ3 since uzμ = (0,0,0,1) Consider the combination γ5γ3 = = so that γ5γ3 = . What does γ5γ3 do to our four basic spinors? w1 = + w1 w2 = – w2 w3 = – w3 w4 = + w4 These four signs agree with 3.16 and are bringing out the "sign of the spin" in the four cases. Consider then (ignore right two columns below for the moment) ½ (1 + γ5γ3) w1 = w1 spin up Σ(uz) w1 = w1 Σ(uz) u(uz) = u(uz) ½ (1 – γ5γ3) w2 = w2 spin down Σ(-uz) w2 = w2 Σ(-uz) u(uz) = 0 ½ (1 – γ5γ3) w3 = w3 spin down Σ(-uz) w3 = w3 Σ(uz) v(uz) = v(uz) ½ (1 + γ5γ3) w4 = w4 spin up Σ(uz) w4 = w4 Σ(-uz) v(uz) = 0 Here we see the sign sequence noted above for the wr spinors being +, –, –, +. Now, when we define the u and v, we define w3(p) = v(p, -uz) and then when you list the states in the following order, you get the following spin signs: u(p,s) + u(p,-s) – v(p,s) + v(p,-s) – These are the rest frame projections we want. So write Σ(uz) = ½ (1 + γ5γ3) = ½ (1 + γ5z) and we generalize to Σ(sμ) = ½ (1 + γ5) for an arbitrary spin 4-vector. We then get, in the rest frame, the third column above. But in terms of the u and v states, we get the far right column, and we see how the minus of the definition of v cancels out the minus of Σ(-uz) w3 = w3. This rightmost column is equations D on page 34. The point of 3.22 is that the spin-projectors are in "covariant form" just the way (-m)ψ = 0 is in covariant form. That is to say, we had earlier that S(a) S-1(a) = S(a) γμ S-1(a) ∂μ = S(a) γμ S-1(a) (aνμ∂'ν ) = { S(a) γμ S-1(a) aνμ } ∂'ν = ' The slashed object is taken into the new reference frame by the transformation S. What happens here when we do this with a non-operator 4-vector? Same thing. That is to say, for example, S(a) S-1(a) = ' where p is just a 4-vector So the same thing will happen with ANY operator when contains things like or . The equation in question will have the same form in primed as in unprimed, where you are acting on some ψ(x) so that you have ψ' = S(a) ψ. So that is what they mean when they say ½ (1 + γ5) is in covariant form. They don't mean it is a world scalar. So this is why we always try to find a covariant form in the rest frame, then we can take it to any frame in this way. In other words, going from page 34 D to 3.22. Comment: I think this all is nothing more than a naming of states. As shown in 3.15, we define v(p,s) to have spin down. Clarifications added 10.11.08 (1) Question: Is a "world scalar" in some sense, or is it not? In pondering this question, I roll out lots of our now known technology, prove covariance of the Dirac equation two ways, and then state a solid conclusion at the end! I now know the following facts: S-1 γα S = Λαν γν γμ transforms as a 4-vector object under Dirac rep action Λ-1 pμΛ = Λμνpν pμ transforms as a 4-vector object under Vector rep action Here U and γα are "operators" in the Dirac 4x4 space, whereas Λ and p are operators in the ∞x∞ QM HS. The following two facts are true S-1 ' S = p'μ S-1 γμ S = p'μ Λμν γν = pν γν = p' = Λp Λ-1 Λ = γμ Λ-1 pμ Λ = γμ Λμνpν = γμ p'μ = ' where we use these facts from our curvilinear doc x'a = Tab xb x'a = Tab xb xa = Tba x'b = x'b Tba xa = Tba x'b = x'b Tba That is to say, we know these two facts S-1 ' S = p' = Λp Λ-1 Λ = ' A. Dirac covariance in coordinate space Now let's look very carefully at the covariance of the Dirac equation: (γμμ - m)ψ(x) = 0 frame F (γμ'μ - m)ψ'(x') = 0 frame F' Start with ψ'(x') = S ψ(x) and start with the second equation above 1 (γμ'μ - m)ψ'(x') = 0 2 (γμ'μ - m) S ψ(x) = 0 Here we have to think of x = Λ-1x' inside ψ(x), so 'μ does not just give zero when acting on S ψ(x). In fact, we know that ∂μ'f(x) = = = Λμν ∂νf(x) => 'μ f(x) = Λμν ν f(x) which we can abbreviate as 'μ = Λμν μ, just as if p were a 4-vector non-operator object. Of course this really means the equality is true when we act of a function of x where x = Λ-1x' . Applying this idea to the linear combination S ψ(x) we arrive at 3 (γμ Λμν ν - m) S ψ(x) = 0 Now apply S-1 from the left and shuffle the S around a bit to get 4 S-1 (γμ Λμν μ - m) S ψ(x) = 0 5 ( [S-1γμ S] Λμν μ – m) ψ(x) = 0 Now use S -1 γμ S = Λμα γα which we stated above to get 6 (Λμα γα Λμν μ – m) ψ(x) = 0 Then use the fact that Λμα Λμν = δαν for any Tab type transformation and we get 7 (γμ μ – m) ψ(x) = 0 Thus we have shown that 1 => 7 and the Dirac equation is "covariant". B. Dirac covariance in coordinate space Now let's repeat all of this using bra-ket notation. Start in the same place as above 1 (γμ'μ - m)ψ'(x') = 0 or (γμ'μ - m)<x'|ψ'> or ( 'μ - m)<x'|ψ'> Replace 'μ <x'| f> = <x'|Pμ| f> in the usual manner to get ( we used P = P† ) a <x'| (γμPμ - m)|ψ'> = 0 or <x'| ( - m)|ψ'> = 0 Now expose Λ† on the left side, and insert a pair prior to the ket to get ( Λ in the QM HS is unitary) b <x| Λ† (γμPμ - m) Λ Λ† |ψ'> = 0 Now use "explanation of...." Section 4.1 which says R†|ψ'r> = S(R)rs | ψs(p)> to get c <x| Λ-1 (γμPμ - m) Λ S |ψ> = 0 d <x| (γμ [Λ-1Pμ Λ] - m) S |ψ> = 0 Next, use our rule Λ-1 pμΛ = Λμνpν given above for transforming a vector operatort in QM HS e <x| (γμ ΛμνPν - m) S |ψ> = 0 Now apply S-1 to the left of both sides f S-1<x| (γμ ΛμνPν - m) S |ψ> = 0 g <x| ([S-1 γμ S ] ΛμνPν - m) |ψ> = 0 Now apply our rule that S -1 γμ S = Λμα γα which says S -1 γμ S = Λμα γα h <x| (Λμα γα ΛμνPν - m) |ψ> = 0 Then use the orthogonality as usual to get i or <x| - m) |ψ> = 0 now replace <x| Pν |f> by ν <x|f> to et 7 (γνν - m) <x| ψ> = 0 or ( - m) ψ(x) = 0 So here we have two separate "proofs" of the covariance of the Dirac equation. The first is done with wavefunctions, the second is done in the QM HS. In the second proof, both the Λ and S transformations are visible. That is, we encounter both [Λ-1Pμ Λ] in the QM HS, and we encounter [S-1 γμ S ] in the Dirac space, and we have to use both these transformation rules. We can now look back at some of the stages of this second proof: a <x'| (γμPμ - m)|ψ'> = 0 => c' <x| S-1Λ-1 (γμPμ - m) Λ S |ψ> = 0 => i' <x| γμPμ - m) |ψ> = 0 so if we look just at the operator center core here, we see this fact S-1Λ-1 (γμPμ) Λ S = γμPμ or U-1(γμPμ) U = γμPμ or U-1U = where U = Λ S So this now brings us to the answer to an implied question " is a world scalar? " Here we see that under the combined Lorentz tranformations Λ (which acts on QM HS operators and so acts on p) and S (which acts in the Dirac spinor space and so acts on γ), the operator object is indeed a scalar. Under either of these transformations alone, we find something that is not a scalar. Now if we have some non-operator 4-vector pμ, then in one Lorentz frame we could compute and we would get a certain 4x4 matrix. If we go to another Lorentz frame and compute ', we will get a different 4x4 matrix if we use the same γ matrices, just treating them as constants. So in that sense, one would never say was a world-scalar. This is the usual meaning of wrt LT's. That is, we transform the p but not the γ. Consider the following which shows that s is in fact a "world scalar matrix" : s' = p'μγ'μ = (Λ-1 pμΛ) (S-1 γμ S) = Λμνpν Λμαγα = pμγμ = s which we could write as ()' = where ()' ≡ p'μγ'μ But when we write ' , we always mean p'μγμ so we then have ' ≠ and in fact ' = p'μγμ = Λμνpνγμ ≠ So without any confusion, we would say that what we always denote by is not a world scalar. (2) How is the spin operator defined and why? This note concerns spin and the spin projection operator. I think one should realize that the association of "spin up" with w4 is an arbitrary decision made by B&D, it is not forced by anything. Exactly why they do this, I comment at the end of this note, but the fact is that they arbitrarily do it. They could have done it differently. They could have let Σ = be the "spin operator" (all indices are down here, Σi and σi) . It is possible to write Σ= γ6γ so that Σ3= γ6γ3. Our spin projector for aligned spin-up states would have been Λ+ = ½ (1 + Σ3) = ½ (1 + γ6γ3) We could then write, in this rest frame, that = sμγμ = s3γ3 = γ3 = -γ3 to get Λ+ = ½ (1 + Σ3) = ½ (1 + γ6γ3) = ½ ( 1 – γ6 ) and then the covariant projector for any spin would have been Λ+ = ½ ( 1 – γ6 ) But this is not what they did. Instead, they use γ5γ3 = – γ5γ3 = or Σ' ≡ – γ5γ = as the "spin operator". The difference between this and the above is the minus sign on the lower σ. In the rest frame, the spin up projector would be Λ+ = ½ (1 + Σ'3) = ½ (1 – γ5γ3) By the same argument as above, this then leads to the following covariant projector Λ+ =½ ( 1 + γ5 ) Again, I think this definition of the "spin operator" and its resultant covariant form is arbitrary, but they don't ever say that in the book. One motivation for doing what they did is this: if we want to associate a physical positron going forward in time with a negative energy electron going backward in time, then going backward in time ought to flip the spin direction if we think of spin as an "angular momentum". If you have a rotating object, time reversal makes it rotate backwards, and the angular mometum vector reverses. So this might be the reason they want to associate v(p,s) with w4 instead of w3. We know that w4 has spin down and we know it is a negative energy electron state, so the physical positron associated with it should have spin up, and we cause this to happen by the way we define the spin operator. That is to say, we think of v(p,s) = w4 as a spin up positron state, meaning it is spin up when we are in the rest frame and set = . In this same manner, we think of u(p,s) as a spin up physical electron. Combination Projectors These are stated on page 35. If you have some linear combination of the wr, you can project out what you want. For example: f = a1w1 + a2w2 + a3w3 + a4w4 P3 f = a3 Recall that state w3 has negative energy and negative spin. Notice he has Σ(uz) where uzμ = aμν u0zν where u0zν = (0,0,0,1) which is the thing. So equations page 35 A are in some general frame, not just the rest frame. Prove that [ Σ(s), Λ± ] = 0 when s.p = 0. This is easy. Install all the pieces, and we have to show that [γ5, ] = 0 or [ γ5 γμ, γν] sμpν = 0 But [ γ5 γμ, γν] = γ5 { γμ, γν} = γ5 gμν so we get [ γ5 γμ, γν] sμpν = 0 <= γ5 gμν sμpν = 0 or s.p = 0, as claimed. 3.3 Physical Interpretation of Free-particle Solutions and Packets (35) The expansion 3.23 includes the r = 1,2 positive energy states only, hence the super (+) on the ψ. The phase space factors do indeed result in 3.24 which then tells you that b(p,s) is the normalized momentum space wavefunction which you square to get probability. Note that b is just a coefficient, possibly complex, it is not a creation or destruction operator! [ but I am sure it will becomes such eventually ] . The big next step is going to be computation of the 4-current for this packet! First, what about the zeroth component? The current we know is jμ = c γμ ψ so j0 = c γ0 ψ = c ψ†ψ Jμ = ∫d3x jμ The total 0th component of the current is then c∫d3x ψ†ψ so we have J0 = c∫d3x ψ†ψ = c * (3.24) so this J0 is what we have already computed (and we are reminded that it is positive definite). But now, let's do the entire 4-vector using the Gordon (1928) Decomposition 3.26 to get something like 3.28. So now we are doing "derive 3.28" but with μ. [ the decomposition is derived at the top of page 37 for us. ] Jμ = c∫d3x γμ ψ Now, let's use 3.23 as is for ψ, and then prime all the integration/sum variables for and we get this, where I will attempt some shorthand: (f = 1/(2π)3 mc2/) , Jμ = c∫d3x ∫d3p'∫d3p f Σs,s'b*' u†' e+ip'.x γ0γμ b u e-ip.x = c∫d3x ∫d3p'∫d3p f Σs,s'b*' b ' ( e+ip'.x γμ e-ip.x) u = ∫d3x ∫d3p'∫d3p f Σs,s'b*' b ' (c 2(x)γμ ψ1(x)) u where now the object ( e+ip'.x γ0γμ e-ip.x) can subject itself to the Gordon decomposition (and here I have temporarily redefined the meaning of ψ just to fit the Gordon formula) ψ1(x) = e-ip.x/ u(p,s) ψ2(x) = e-ip'.x/ u(p',s') ψ2(x)† = u'† e+ip'.x/ 2(x) = ' e+ip'.x/ So now we want to apply 3.26 to: c2(x) γμ ψ1(x) We will need these computations: μ ψ1(x) = i∂μ (e-ip.x/)u = +pμ (e-ip.x/)u = +pμ ψ1(x) μ 2(x) = i∂μ (e+ip'.x/)' = - p'μ (e-ip.x/) ' = - p'μ2(x) ν( 2(x) σμν ψ1(x)) = (ν( 2(x)) σμν ψ1(x) + 2(x) σμν (ν ψ1(x) ) = - p'ν2(x) σμν ψ1(x) + 2(x) σμν pν ψ1(x) =(- p'ν + pν ) 2(x) σμν ψ1(x) So we get c2(x) γμ ψ1(x) = (1/2m) { 2(x) pμ ψ1(x) + p'μ 2(x) ψ1(x) + i (+ p'ν – pν ) 2(x) σμν ψ1(x) } Now replace ψ's on the right. We can combine the two exponentials into ei(p'-p).x/ and we have c2(x) γμ ψ1(x) = (1/2m) ei(p'-p).x/ ' { pμ + p'μ + i (+ p'ν – pν ) σμν } u = (1/2m) ei(p'-p).x/ '{ pμ + p'μ + i (+ p'ν – pν ) σμν } u Inserting this into the result above gives, = ∫d3x ∫d3p'∫d3p f Σs,s'b*' b (c 2(x)γμ ψ1(x)) = ∫d3x ∫d3p'∫d3p f Σs,s'b*' b (1/2m) ei(p'-p).x/ ' { pμ + p'μ + i (+ p'ν – pν ) σμν } u = ∫d3x ∫d3p'∫d3p (f/2m) ei(p'-p).x/ Σs,s'b*' b ' { pμ + p'μ + i (+ p'ν – pν ) σμν } u where f = 1/(2π)3 mc2/ from above. This agrees with the first long line in 3.28. Now we do the d3x integral to pick up (2π)3 δ3(p'-p) and E = E' we are then left with ∫d3p (F/2m) Σs,s'b*' b ' { pμ + p'μ + i (+ p'ν – pν ) σμν } u where now p = p' and F = mc2/E and then (F/2m) = c2/(2E) so result is c2 ∫d3p(1/2E) Σs,s'b*(p,s') b(p,s) (p,s') { pμ + pμ + i (+ pν – pν ) σμν } u(p,s) It appears that the σμν term goes away, and the first two terms add so we have c2 ∫d3p(1/E) Σ±s b(p,s) pμ Σ±s' b*(p,s') (p,s') u(p,s) Now let's recast (3.9b) into the u-spinor language [ this is row vector contracted against column ] r(p) wr'(p) = εr δr,r' r,r' = 1,2,3,4 which of course implies that r(p) wr'(p) = δr,r' r,r' = 1,2 which says that ( where f is some arbitrary function) Σr=1,2 f(r') r(p) wr'(p) = Σr=1,2 f(r') δr,r' = f(r) for either value of r' which we can translate into Σ±s f(s') (p,s') u(p,s) = f(s) for either value of s Now we think of f(s') = b*(p,s') and our result above becomes, c2 ∫d3p(1/E) Σ±s b(p,s) pμ Σ±s' b*(p,s') (p,s') u(p,s) = ∫d3p(pμ c2/E) Σ±s |b(p,s)|2 And here is our full result then, where ψ = Jμ = c∫d3x (+) γμ ψ(+) = ∫d3p(pμ c2/E) Σ±s |b(p,s)|2 and from this we can set p0 = E/c to get J0 = c∫d3x (+) γ0 ψ(+) = c ∫d3p Σ±s |b(p,s)|2 Ji = c∫d3x (+) γi ψ(+) = ∫d3p(pi c2/E) Σ±s |b(p,s)|2 J = c∫d3x (+) γ ψ(+) = ∫d3p(p c2/E) Σ±s |b(p,s)|2  where notice that the vector γ refers to γi. Now we can write the above as J0 = < ψ†(+) | c |ψ(+)> = c ∫d3p Σ±s |b(p,s)|2 = c // normed in 3.24 J = < ψ†(+) | c γ0 γ |ψ(+)> = ∫d3p(p c2/E) Σ±s |b(p,s)|2 = < c γ0 γ> = <c α > and now we have derived the second line in 3.28. In coordinate space, we imagine 1 = ∫d3x Σ±s | x,s><x,s| as our completeness. In momentum space we have to use 1 = ∫d3p Σ±s | p,s><p,s| I guess (no factor). In momentum space, each u(p,s) component has amplitude b(p,s) and probability |b(p,s)|2 . So yes, we have J = < ψ†(+) | c γ0 γ |ψ(+)>x = < p c2/E>p Recall from relativisitics that p0 = E/c and p = γm dr/dt = γmv = γβmc and E = γmc2 so we have p c2/E = γmv c2 /γmc2 = v So our big result is this: J = < v >p = < ψ| v |ψ>/<ψ|ψ> = < ψ| v |ψ> I guess when you have a packet, this thing here is called the group velocity of the packet, and you have to compute it and we see that it is just this: ∫d3p(p c2/E) Σ±s |b(p,s)|2 Whatever it is, that is how fast the packet moves, ie, that is how fast your electron is moving. Now, recall that jμ (x) = c (x) γμ ψ(x) = " the 4-current density" Jμ = c∫d3x (+) γμ ψ(+) = the 4-current" of the positive-energy packet" = ∫d3x jμ (x) = ∫d3p(pμ c2/E) Σ±s |b(p,s)|2 = ∫d3p vμ Σ±s |b(p,s)|2 So I would not refer to Jμ as "the average current for an arbitrary packet...". I would say that "the four current Jμ of the packet" was a weighted average of the velocity of each momentum-space component of the packet", and this weighted average is known as "the group velocity" of the packet. So in the sentence below 3.29 I think they have something jumbled. B&D now point out this distinction between rel and non-rel theories: (1) non-rel theory, [p,H] = 0 and v = p/m is a constant of the motion. (2) rel theory, [α,H] ≠ 0 so the velocity operator vop = cα is not a constant of the motion. I am missing the point made here. How would you make an eigenfunction of cα? Ah, OK, that is the point. You must linearly combine multiple eigenstates of H to get eigenstate of cα because [α,H] ≠ 0. If this were 0, you could diagonalize both at the same time. So "linearly combine" is no doubt going to mean that you have to mix states of both positive and negative energies to get an eigenstate of cα. Page 38 we forge ahead: Now, in 3.30 we add in some v(p,s) negative energy spinor action with coefficient d*(p,s) and with f = 1/(2π)3 mc2/. Jk = c∫d3x ∫d3p'∫d3p f Σs,s' { b*' ' e+ip'.x + d' ' e-ip'.x } γk {b u e-ip.x + d* v e+ip.x } Now we have 4 terms instead of just 1 term. OK, it would take me about 8 pages to derive 3.31, so instead of doing that, let's take note of the main points which are these: 1) in the first cross term b*' ' e+ip'.x γk d* v e+ip.x we have the same sign on the expo factors. The Delta function δ(p+p') will force p' = -p , but this we still have p0'= +p0 after this delta. Thus, the two exponentials now have a residual exp(2ip0x0/) factor which is new -- fast time dependence! 2) the "spin terms" now have a contribution of the form i (+ p'0 + p0 ) σk0 = 2i p0 σk0. You see both of these results in 3.31 where the cross terms are the last two. The other diagonal term is as simple as the first. Fact: With only positive energy terms, our current J in 3.28 was constant in time for our packet. When we add in the negative energy solutions, it suddenly develops ultra-high frequency time dependence from those cross terms! This is the famous zitterbewebung of Schrodinger 1930 (jittering agitation, etc). It is an interference effect between the positive and negative energy solutions. On page 39, we consider a localized 3D gaussian spatial wavefunction at t = 0 and ask "does it involve Dirac solutions of both signs of energy as t moves forward? " We assume a gaussian shape and w1(0) at t=0 as shown in 3.32. We evaluate 3.30 at t=0 to get equation A, something I can verify. We then do a d3x Fourier on both sides of A. The RHS is done explicitly for us. Authors use a vector l as the variable conjugate to r, something we usually call k instead of l. The LHS just pins l = p and -p in the two terms of A, and we get C which I have checked except for the constant. We can then use orthogonality relations to project out the two amplitudes. We find that, although at t=0 we started with a pure u state, we end up t>0 with both u and v states activated, and so we are going to get this jitter effect. The d* amplitude is small in the NR limit since it shows as a v spinor. However, when the gaussian is made small enough that the packet is confined to a Compton wavelength λ = h/mc, then p becomes large enough so the negative energy terms are appreciable. The deBroglie is λ = h/p so we are just saying when p gets up near mc, effects start showing up. Of course p can get much larger than mc, since p = γβmc in general. λ = h/mc = 2.4263102175×10-12 meters ~ 2.4 pm Compton, scale of electron extent for QED As an example of a paradox, B&D consider the famous Klein Paradox (1929). The idea is to send in a beam of electrons from the left against a 1D steep potential wall, say a square well wall as shown on page 41. So put an electron in a 1D box, and try to localize it more and more. Make the walls high, that would seem a good idea. We look only at the right wall. We want to get the usual SE exponential damping going into the barrier. We assume the incoming scattering wave is 3.34 first, which is a u(p,s) thing with a small lower part, overall amplitude "a". The reflected wave is assumed a linear combination of w1 and w2, both having small 3,4 parts as shown. The transmitted waves are still free-electron solutions if we shift by the potential wall height V0. So we assume the transmitted wave as shown, and now we have five constants a,b,b',d,d' and we want to solve the problem "as usual" by matching things at the wall. Continuity requires that a+b just to the left of the wall equal d, where we are doing this comparison for the w1 terms. He concludes that b' = d' = 0 due to no spin flip. When you match the lower components at the boundary, you get an a-b condition so you can solve for all the constants, as you would expect. Notice that our interest is in the wave number k2 inside the barrier region. When you increase the barrier high enough, suddenly you get non-exponential transmission! Notice that jtrans/jincident = 4r /(1+r)2 where r is the quantity I circled in 3.36. When the barrier gets too high, r goes negative, and we have transmission inside the barrier going to the left! And, the reflected current is larger than what came in. Well, I know that what has happened is we get an electron-positron pair created in addition to the original electron, I guess the positron goes off to the right, making the negative current, and the new electron comes back with the original reflected one and you get larger than came in. In this book, this will be resolved in chapter 5 in terms of "hole theory". So the Dirac Equation "knows about" pair creation from the vacuum. I will be interested to see that chapter. Meanwhile, there are still many interesting problems to look at where we stay far from this pair-creation energy level. This is where the F-W transformation is going to be useful. Chapter 4: The Foldy-Wouthuysen Transformation (46) 4.1 Introduction (46). In hydrogen, the Bohr radius is 137 times larger than the Compton wavelength, so we expect good results with predictions we can make which ignore the negative energy solutions. We are not going to be creating many e+e- pairs in a cool H atom! The general idea now is to add a potential V(r) to our Dirac theory and see what we can get for solutions say for a central force. The claim is that this 1950 FW work simplifies the problem in a systematic fashion by somehow decoupling our 4-spinor equation into something simpler. BD cannot resist calling FW a "canonical transformation" with no explanation, which leads to the next comment: Comment on Canonical vs Unitary Transformation: Let U be some transformation that does not involve time. Then we transform the SE in the following manner: Hψ = ψ HU-1Uψ = ψ UHU-1Uψ = Uψ = Uψ H'ψ' = ψ' One example of this kind of thing is moving between "pictures" where U = eiH' where H' is some Hamiltonian. The question now is: in general, must U be unitary? It has to be if we require that the norm of our states be unchanged. If the norms of all states are unchanged, then we will still have total probability = 1 in the primed world if it is so in the unprimed. Note that <ψ'|ψ'> = < Uψ | Uψ > = <ψ|U†U|ψ> = <ψ|ψ> only if U is unitary. If U were not unitary, I suppose we could renormalize states in the primed world and make things be OK. But then we would be converting U to unitary in effect. In Goldstein we read about "canonical transformations" as changing the P and Q of the Hamiltonian perhaps to constant values, and in any event to different values. We did this by adding some ∂F/∂t to the Hamiltonian to form a new Hamiltonian, where F was called a generator. That was classical mechanics. In quantum mechanics, we have our coordinates x and p and our states ψ. I don't know right now how you translate the notion of the canonical transformation from classical to quantum mechanics. However, I do know that U = eiHt acts as a unitary transformation which can take your wavefunctions back to time 0 where they are the "initial values", and in classical mechanics you seek a transformation which takes you to a place where P and Q are constants. Maybe somehow the wavefunctions are the coordinates in QM. The web suggests that somehow the "unitary transformation" of QM corresponds to the "canonical transformation" of classical mechanics. For now, this is a semantic question. I suspect this question will get answered if I keep pursuing the general line of inquiry I am now on, I don't need this answer now. In this chapter, we are going to study various ψ' = eiSψ transformations which Messiah says are in fact unitary. We will keep finding that S = -iβO/2m where β is the usual 4x4 matrix and O are some "odd" terms in a Hamiltonian. Thus we have S = -iβO/2m S† = +i O†β†/2m = +i Oβ/2m = -i βO/2m = S S = S† eiS = unitary We assumed here the O and β anticommute. This is true for example if O = αp or O = γ5. It is therefore probably true for any possible "odd terms" (defined below) that might occur in the Hamiltonian. So the FW business is going to involve some unitary transformations which we can "call" canonical transformations if we like. Also, the transformations are "small" in that somehow S << 1. 4.1 Free-particle FW Transformation (46). Here we are thinking about finding a nice H' in the context of no potentials, just a "free particle". As a small digression of the authors, consider H = σxBx + σzBz which is non-diagonal. How would you make it be diagonal? Well, consider H = σB in general. We know that R σ R-1 = R-1σ R σB R-1 = B (R σ R-1) = B R-1σ = B σz = Bzσz H' = Bzσz where we have selected R such that R-1σ = σz. In the case they specify, σ is in the x-z plane and we just want to bring it back up to the top with a y rotation of a certain angle we could call θ, fine! Comment: think of R as causing a change in coordinate system in which the problem is studied. When we work with H' as shown above, things are simple. Selecting the z axis to align with B is the best way to work on this problem. The same idea will apply to FW. In the above, σz is diagonal of course which is why we like it in H'. In the 4x4 space, we want to get to an H' where all matrices appearing are diagonal at least in the 2x2 subspace sense. Examples include γ0=β, σij, Σ. These are called the "even" matrices. Those which are off-diagonal in this sense include γi, γ5 and αi -- the odd matrices. These "odd" matrices mix the upper and lower components and make our life more difficult. So, consider our 4D problem. We have H = (αp + βm) which of course contains those "odd" α matrices we don't like. We seek H' where H' = R(αp + βm) R-1 = something with only "even" matrices in it. = p Rα R-1 + m Rβ R-1 where now R acts in the 4x4 Dirac space. The suggested purifying rotation is this one: R = exp([θp]βα) = exp([θp]γ) = cos(pθ) + γ sin(pθ) where in the expo expansion we note that (γ)2 = 1 so powers of this thing are either 1 or γ. We then compute RHR-1 and we get page 47 G and see the possibility of killing off the term with α's. The killing angle θ is a function of m and p as shown in H. In getting there, we used {α,β} = 0. Repeat: in order to zero out the α term, we choose angle θ as shown. Then using the triangle on the next page to compute the cos and sin in the second term, we get 4.1 which is a very simple result indeed: H' = β which is a 4x4 matrix. We have β = diag(1,-1) so the upper and lower component pairs are now decoupled. The eigenstates of H' are going to be four such as (1,0,0,0). The first two have eigenenergy , and the last two have eigenenergy - . We can ignore the negative energy solutions when we are thinking about the positive energy solutions, even at high velocity, in this representation. Consider back earlier where with the "regular H" we had the four complicated spinors for a particle not at rest. Here we also have four spinors for the same problem, but they all are like (1,0,0,0). So yes, things are much simpler in this "representation" of the hamiltonian. As an aside note that this rotation thing R is unitary: R = exp([θp]γ) => R-1 = exp(-[θp]γ) R† = exp([θp]γ†) = exp(-[θp]γ) = R-1 4.2 The FW Transformation with E&M Fields (48). We now turn on the EM field and try again to "diagonalize" our Dirac equation using some kind of transform, again trying to get rid of "cross terms" or "odd matrices" as they say. In 4.2 we see that the "odd matrices" both involve α, as before. Now first, jump over to page 49 top where we define a new transformation S such that ψ' = eiSψ. This is not the same S as earlier when we were looking at how spinors transform under Lorentz. Confusion #1. What happens next in line B is somewhat confusing, so notes are needed. I will do it "step by step" as one always needs to do in BD. On this line we have 4 expressions and 3 equal signs. (1) the first equal sign with its two expressions is a statement that Hψ = ψ where means the usual i∂t operator, so this is the time dependent SE written for ψ with Hamiltonian H -- our starting positions. (2) The third expression is a rewrite of the second expression, so we understand it and the second =. (3) the fourth expression is the unexpected one. This comes from the leftmost expression where you apply the time derivative on the two factors. Then it becomes obvious. So now we are happy with B. (4) Now we take the rightmost equality in B and apply e+iS on the left, and this gives the left equality of C. We then interpret the large square bracket as H' since we then have H'ψ' = ψ'. Notice that the ∂t operator inside H' does NOT act on ψ' (for future reference). So in all this shuffling, we have simply found the appropriate H' which gives a SE on ψ', nothing more and nothing less. Our goal now is to select some kind of S which will "kill off" our odd matrix terms. In D we show the usual CH expansion of the sandwich shown in general to all orders. Then in E, only the first four terms of this CH expansion are shown, the fourth being the one with 1/24. Confusion #2. Now we must understand that claim that S = O(1/m) and what this even means. This seems very unobvious to me because we don't yet know what S is. So now begins some more PhL "interpolative comments" to fill in what BD should have told us but did not to "save space" (our time is not important). The first thing we claim is that, in the NR limit, "S is small" so that ψ' is not too different from ψ. The idea here is that in the NR limit, the "odd matrix" terms are already very small, so it does not take a very large "rotation" to cancel them out. In our statement that H = α(p-eA) + βm+ eφ [ we have turned on some fields, by the way ] we see that the βmc2 term is -- in the NR limit -- very much larger than the other two terms, assuming "reasonable" EM fields. So as just claimed, it is the α term here which is the small odd term we want to get rid of and we will do this by a small rotation S. Looking again at H, my first question is this: how does the first term have dimensions of energy? It is really cp that is sitting there in the first term and this is then c(γβmc) = γβmc2 [ where γ and β here are the special relativity things, unrelated to gamma matrix things ]. Then our ratio of interest is roughly just γβ. We set γ = 1, then as usual our NR limit means β << 1. This I understand. I do not understand where the ratio 1/m enters the NR limit discussion. As Messiah points out on page 944, the actual smallness ratio is this (which I arbitrarily call λ) λ= [ α(p-eA)] / m = [ cα(p-eA)] / mc2 = energy/energy = J/m where J = α(p-eA) Then we see that roughly λ = J/m ~ pc/mc2 = γβmc2/mc2 = γβ ~ β Now in BD, the thing Messiah calls J is called O (for odd terms). Conclusion: formally, we have λ = J/m = O/m as our smallness parameter. For example, we shall see in our first iteration below (BD page 49 F) that S = (-i/2)( O/m) β where β here is the matrix, so there is the ( O/m) ratio making S be small. We can then "track" our terms by power of (1/m), with the understanding that the actual expansion is in the dimensionless ratio O/m = v/c << 1. Note in passing: if we have S = (-i/2)( O/m) β where O = odd terms, { in the Hamiltonian formalism we have H(p,q,t) ???} we will have ∂tS = (-i/2) {∂t O}/m β and this will be non-zero because A within O is in general time dependent. We refer below to ∂tS as , so this shows that in general ≠ 0 so we have to worry about it. Now go back to that CH expansion in page 49 D. Tracking with 1/m as discussed above, we see that -- since H is of order m (from βm), and since S is order 1/m, our CH terms are order m, 1, 1/m, 1/m2 and so on. Thus, in equation E, the i/6 term is order 1/m2 and the 1/24 term is order 1/m3 . In this last term, we are allowed to replace H with just βm while maintaining correctness through order 1/m3. In other words, this last approximation H ~ βm is off by order 1 which would cause a discrepancy of [S,[S,[S, [S, order(1)]]]] which is order 1/m4 which we don't care about. Now looking back at C, we have to worry about eiS(-i∂t)e-iS as well as the eiSHe-iS . Note that this former thing cannot be written as - (eiS e-iS) because, for example, ∂tS2 = S + S ≠ 2 S. So we have to apply BCH to the operator as shown, and we get this eiS(-i∂t)e-iS = (-i∂t) + i [S, (-i∂t)] + i2/2 [S,[S, (-i∂t)]] + i3/6 [S,[S,[S, (-i∂t)]]] + ... where I have kept terms through order S3 = order 1/m3 . As noted earlier, there is nothing "further to the right" like ψ' that we have to worry about. Thus, we get (-i∂t) 1 = 0 i [S, (-i∂t)] = iS (-i∂t) - i (-i∂t)S = 0 - = - i2/2 [S,[S, (-i∂t)]] = i/2 [S, { i[S, (-i∂t)] }] = i/2 [S, - ] = -i/2 [S, ] i3/6 [S,[S,[S, (-i∂t)]]] = i/3 [S,{ i2/2 [S,[S, (-i∂t)]] } ] = i/3 [S, -i/2 [S, ] ] = +1/6 [S, [S, ] ] And these are then the three terms we see that the end of equation E Old Comment: The ∂/∂t term has to be there in C because, if our EM fields are time dependent (which they normally will be except for trivial problems), then S that removes the odd matrices will be a different S for different times, and we have S = S(t). We want H' above to have no odd terms, but B&D claim no S(t) exists which does this for all times t. I will return later to understanding just why that is so. Old Comment: Meanwhile, looking at page 47 H without the fields turned on, for the non-rel region we will have p/m << 1 and that means tan(2pθ) ~ 2pθ = p/m. This means the S = βα ( pθ) ~ p/m = p/mc to be dimensionless. In non rel, p = mv << mc, so we are to think of S as being order(1/m). In the extreme non-rel limit, we know that coupling goes away between pos and neg energy solutions of Dirac, so not surprising that S is very small in the NR limit, and that S ~ 1/m which I think really means p/m. Now, in 4.3 we keep terms only through order m0 = 1. So we keep the entire H ( the first term from the eiS H e-iS expansion) plus an order 1 term from the i[S,H] term in the same expansion. Fine. At this order m0 how do we kill off the odd terms? If you select S as shown in F, you get i[S,β]m = - O using S = -iβ O/2m easy to show Notice that O = α(p-eA) from page 48 and E = eφ and both are regarded as order m0 = 1, that is, we might have normal sized A and φ fields. Therefore, indeed S = -iβ O/2m is order 1/m as claimed earlier. So this tells us that H' = βm + E for 4.3, and we are diagonal. Now, using this "first order S", we compute all the other terms in the big mess 49E. The 6 terms of interest are duly computed on page 50, and I am not deriving all these things, nor did I on the last pass (but I could easily do it). We end up finally with this transcription of equation E H' = βm + E' + O' as in 4.4 where we have new odd and even terms. Now the odd terms O' are order 1/m or higher. Just as we did in the first round, we now define S' = -iβ O'/2m and we know this will knock out our the 1/m portion of O' . So after applying S', we then have H'' = βm + E ' + O '' where now O '' is order 1/m2. And we do it one more time and we get H''' = βm + E ' + O ''' = βm + E ' where we just throw away O ''' which is order 1/m3. So the point is that we keep doing transformations to grind down the size of the odd terms in H. So we end up with what is shown in page 51 A. Now we do more messy computations to see what these terms are, and we end up with 4.5 where now we see things like E and B fields along with φ and A stuff. There are now lots of terms! Claim is that the first term is really the square root shown in 51E, where we have the first three terms of this expansion. This is what appeared in 4.1 page 48 with no fields, but we have modified it. But notice that this is not the whole answer! You can't just insert p-eA into 4.1 and think you are done, that is what has been shown here. The second term in 4.5 is -μBβ σB which I think indicates the normal interaction of an electron with a magnetic field. The moment is the Bohr magneton e/2m. This is what we used in Levitt. The next two terms are spin-orbit. For central force, only the second of these terms is non-zero and gives the result 4.6. This is famously ½ of what you would expect from a classical calculation by saying that the electron sees a B' field from the rotating nucleus. The ½ arises because the electron is rotating around the nucleus and the effect is known as "the Thomas Precession". For non-relativistic speeds, this Thomas effect always gives a factor of ½, independent of velocity of rotation. I saved a PDF on this subject. The very last term in 4.5 is (e/8m2) 2V and is called "the Darwin term". B&D give a crude explanation of this in 4.7 which is pretty darn close. The idea is that the electron does not see V, but instead sees an effective V because the electron is "spread out" over its Compton size λ = 1/m. Commentary: I am a little confused about the conclusions of this section. We start with our well-known Dirac H and we go to "some other frame" and observe things there. All the above terms are in the ''' frame, with H'''. So I want to say that you are allowed to study your problem in any "frame" or "representation" you want. So here we can choose to work in the H''' representation. What can I say about equation 4.5 which shows this H''' ? The first "energy term" has β and this is mainly what makes the 34 states have negative major energy compared to the 12 states. We also see the βΣ matrix and the Σ matrix by itself appearing in this equation. Note that βΣ = diag(σ,-σ) so those neg energy solutions have backwards magnetic moment in a B field (antiparticles, holes). The spin-orbit terms are the same for 34 as for 12. This equation has no matrices which couple 12 to 34, and it is "exact" through order 1/m2. So, if we want to study 1,2 regular electrons, we can just look at the 12 equation, and that is what 4.5 becomes if you replace Σ with σ and β with 1. The 34 term effects are completely removed. BUT, in order to achieve this, new terms have appeared in the 12 equation!!! These are part of the theory. We have just exposed these terms by doing the decoupling. Now go back to page 12 where we did the non-rel reduction of the Dirac equation in its normal form. In 1.30 we have our two coupled equations. Had we simply set χ = 0, instead of 1.31, we would have gotten 1.34 without the magnetic moment with B term. So this term is really being caused by the small 34 components of the solution being non-zero. I think this is a 1/m effect. Now about the gyromagnetic ratio thing. On page 13 and with 1.34 we say g = 2. Why are we saying this? When you compute the LB term in H either classically or in quantum theory, you get the result being H' = -μB LB where μB = e/2mc. The 2 gets there classically because current = ev/(2πr) and area πr2 so moment is cM = current x area = evr/2, while L is rmv, so M/L= (evr/2c)/(rmv) = e/2mc. So we could write this as H' = -gLμB LB and say gL = 1. So g = 1 is the "natural thing" to have. But the term in 1.34 showing the SB interaction appears as H' = - 2 μB SB so gS = 2. So that is what we are talking about with g = 2. Now, our reduction on page 12 and earlier takes some account of the small 34 components and out pops this H' = - 2 μB SB term. And fiddling with the pA terms brings out the - 1 μB LB term so our Dirac equation is predicting the electrons have g=1 for orbital , but have g=2 for spin. This stuff on page 13 is all quantum, by the way. We got all this stuff just from our simple Dirac reduction. Now turn to page 51 and H''' as in 4.5. We get the same H' = - 2 μB SB term we got on page 12. We also see the exact same term (p-eA)2/2m appearing. But now NEW terms appear which result from the fact that we have been more accurate in accounting for the 34 contributions. We have decoupled through order 1/m2 which is stronger than what we did on page 12 which I guess was order 1/m. So one of the big "new terms" is the famous "spin-orbit interaction" ! BD quietly mention this, but the point is this: it is predicted by the Dirac theory!!! I think I have always wondered where the hell it came from in courses I had. As BD point out page 51, it only arises that way with a central potential V(r) which is just fine. Now if V(r) = 1/r, then 1/r ∂V/∂r = -1/r3 and we have H' = - 1/r3 e/4m2 SL . If we imagine classically that the electron "sees" a B field B' = -v x E then our H' = - 2 μB SB term would predict H' = - 2 μB SB'. But crudely L = mvr and vxE = vE = vV = v ∂V/∂r, so B' = v ∂V/∂r which then tells us H' = - 2 μB S [v ∂V/∂r ]. But L = mvr so v = L/mr and then H' = - 2 μB/m S [L (1/r)∂V/∂r ]. = - 2 μB/m S L * (1/r)∂V/∂r. BUT, the correct result as shown in 4.6 is - 1 μB/m S L * (1/r)∂V/∂r which is half of our classical prediction for the spin-orbit term. So they say you can think of g = 1 again for this situation, and then both terms involving L have this same "orbital g" = 1 (the other L term being on page 13]. But if we think of the electron as being in a non-inertial frame and redo the physics classically using the extra "ω x ..." Goldstein term, then you get the factor of ½. But relativity still comes in because it is what causes the B' field in the first place. I have not figured this all out, but have a PDF on the subject. I cannot digress an infinite number of levels. The main point is this: The Dirac equation under our FW correctly predicts the spin-orbit term with the correct factor of ½ present. So this just adds to the credibility of the Dirac equation. We needed the FW to bring out these more subtle terms. The Darwin term I have only dimly heard of and I don't recall people worrying much about it. It is a purely radial term for a central potential, so it just shifts energy levels, it does not do radical things like LS coupling does in terms of state labeling and spectroscopy. I presume we will hear more about it from B&D in next section or two. 4.3 The Hydrogen Atom (52). BD have short-changed this important topic and pretty much state equations 4.9 without any derivation. I think they are claiming that these are the angular parts of a separable solution in spherical coordinates and then four radial functions are installed on page 54 and their equations are given in 4.13 and we are told to go see some old papers given on page 52 bottom. These authors include original papers of Darwin(1928), Gordon(1928), then Bethe & Salpeter (a 1957 book) and Rose (a 1961 book). They then simply quote the energy levels in 4.14. Do I have the missing meat in some other book? The angular stuff appears not in Schiff or Messiah, but it does appear in my Sakurai, and I am reminded that the latter is well written as I peruse it. For the moment, let us then accept the eigenenergies as they appear in 4.14. There are two quantum numbers. First is n = 1,2,3.... and second is j which can range only from j+1/2 up to n. For example, if n = 1, I guess you can have j = ½, and if n is 2 you can have j = ½ and 3/2. All very hazy but OK for now. Notice that α = 1/137 appears in 4.14 in two places, actually in the form Zα for hydrogen-like atom. If this is treated as a smallness parameter, you can expand the square root structure in a series and get 4.15 where now the first term with the 1 gives the usual NR energy levels for Z hydrogen. We see there a correction term that is a function of n and j. Next, it is natural to inquire about the ground state and this is n=1 and j= ½ as the only value. The energy is given in the above expansion in equation A, and the actual up and down ground states are given as 4-spinor detailed solutions in B and C. The angular parts are obscure to me and appear to involve l = 1 but I know l=0 in some sense, so we put that on the list of mysteries at this point. The radial behavior lies in factors outside the bracket, including of course the required exponential decay factor. Here they have defined yet another γ as shown top page 56 (a third usage in this chapter really). If Zα > 1, this square root goes imaginary and we get weird Klein-like behavior. But normally we have Zα < 1 and then we may construct a simple table of the lowest states using the historical l,j,n labels. These four states appear in the page 57 picture along with some other states. The hyperfine stuff is the splitting of all states into doublets due to the nuclear spin coupling with the electron spin in a dipole-dipole interaction a la Levitt. "The 21cm line (1420.4 MHz) was first detected in 1951 by Ewen and Purcell at Harvard University." But that is not really the subject of this chapter, but all lines are shown so doubled in the page 57 figure. The Dirac theory we are now studying predicts that the that 2S1/2 and 2P1/2 should be degenerate (see table energies), but they are not in fact and this is due to the Lamb shift which is a field theory effect. On the other hand, Dirac theory does predict a separation between the 2P3/2 and the 2P1/2 and this observed split is called the fine structure (as opposed to the above hyperfine). This amount is 10.9 GHz, the Lamb shift is 1.057 GHz and the hyperfine is 1.42 GHz. I now resume the BD batting order: (p 57) The Lamb shift was only discovered in 1947 and BD make only a passing note of it without giving the size. I know it comes from these diagrams, and most of it from the first. ( p 57) The hyperfine structure is mentioned. (p 58) BD then "estimate" the size of the hyperfine structure in an obscure way. I think they are saying that the B field is generated by the nuclear spin, but the expression they role out could not possibly be more obscure, they are just showing off here. The reader trying to learn something is never going to know where this comes from, some messy Jackson thing. So this whole page is a complete waste for me. The expression must be doing the relativistic EM field transformation so you get a B field from viewing an electric potential φ = q/r from a moving frame, and all that stuff. (p 58-60, ie through end of the chapter). Here they quote a simple explanation of the Lamb shift from some guy Welton. It relies on thinking about the zero point energy of the vacuum. For me, the question is whether Lamb shift comes out of the Dirac theory or not. I think this BD Vol 1 is going to show that the Feynman diagrams are really part of the Dirac theory, for example see page 148, with the hole idea. So I guess my question is this: is there anything you calculate from field theory that contradicts the Dirac theory? Comment: So I am now disappointed that I invested lots of effort and was not rewarded with a clean derivation of the H atom from Dirac theory. I will have to go read Sakurai or elsewhere for this information. As BD said in their preface, they are not worrying about bound-state problems, so they just fly by this stuff quickly. I think it would be best for me now to just continue in BD and save the H atom for another time, otherwise I will lose the BD thread and I will have to start over yet again. Chapter 5: Hole Theory (64) 5.1 The Problem of Negative-energy Solutions (64). An obvious problem with the Dirac spectrum: what keeps positive energy state electrons from "falling" into those negative states? It is easy to compute the transition rate downward and it is infinite. Dirac in 1930 suggested that this would be fixed if ALL those negative energy states were "filled" and that the Pauli Principle applies to those states so you cannot add more, hence no downward transitions. Pause: exactly "what" is "filling" those negative energy states? Do we suddenly have a theory that is a multi-particle theory and that number of particles is infinite, since that is how many are needed to fill those negative energy states? Are all those negative energy electrons "really there"? I guess if you were to do Dirac in a Box, you would say that the Box as a whole has all those negative energy electrons in it, not each point in space. No doubt these questions arise from the fact that this model is not completely correct. Still, it allows lots of predictions. The most obvious is that if you apply a photon with ≥ 2mc2 energy (> 1.022 MeV, not very much), you ought to be able to excite one of those negative energy electrons into a positive energy state, leaving behind a positively charged "hole" which might behave like a particle itself. This would sort of test whether those negative energy electrons are "really there". Experimental result: this is exactly what can happen, and the positron was duly found in 1932. The exact date of the time positrons were "postulated" might be 1930 by Dirac instead of 1928 as it says here. " The existence of positrons was first postulated in 1928 by Paul Dirac as a consequence of the Dirac equation. In 1932, positrons were discovered by Carl D. Anderson, who gave the positron its name.[3] The positron was the first evidence of antimatter and was discovered by passing cosmic rays through a cloud chamber and a lead plate surrounded by a magnet to distinguish the particles by bending differently charged particles in different directions." So here is the "classic" pattern of particle physics. A new particle type is predicted, then it is found. So Dirac discovered anti-matter, pretty amazing. 5.2 Charge Conjugation (66) The "absence" of a negatively charged of energy -E might be the same as the "presence" of a positively charged particle of energy +E. Now things get a little tricky and I want to at least attempt to follow the logical steps here. Our regular Dirac equation is 5.1. If there were a positive energy positron, it's equation would be 5.2 where the sign of e is changed, all else is the same. Here notice authors have used ψc thinking of the positron as a "charge-conjugated electron". So what operator will generate ψc from ψ? The answer is going to be this: ψc = (iγ2)ψ*. Note added: You need to change the relative sign of just the term in 5.1. We know γ2*γ2 = and γ21 γ2 = -1 so if you do a K sandwich then a γ2 sandwich, the i term negates, the m term negates, and the term stays the same, mission accomplished. In the development of this result, authors show that the * makes the grad term go negative, leaving you with 5.3, Then the matrix iγ2 negates γ*μ and causes the relative change in the sign of the mass term, and thus the combination gives you 5.2 from 5.2 as desired. Now BD manage to confuse things a bit with their notation. They know it is convenient to write this result as ψc = (iγ2)ψ* = (iγ2γ0)γ0ψ* = C (γ0ψ*) where they define C = iγ2γ0. The reason for defining C in this way is this: (note that T is a column vector, though at first you think it is a row vector) γ0ψ* = T = (ψ†γ0)T = γ0 (ψ†)T = γ0 (ψ*) which allows us to say that ψc = CT C = iγ2γ0 = i = i = -i α2 CT = -C C* = C C† = -C C2 = -1 C-1 = -C C-1 γμ C = - (γμ)T where the phase "i" is chosen to make the matrix elements of C be real. The matrix iγ2 is written out in 5.7 and you see them there computing ψc = (iγ2)ψ* starting with ψ being the spin-down negative energy state r=4 shown in A. As shown, you end up with the spin-up state r =1. So "the absence of a negative energy spin-down electron" equals "the presence of a positive energy spin-up positron". So this is the first mention of what happens to "the spin" in charge conjugation. I guess it has to be that way. You remove something with Jz = -1/2 and this leaves the sea with a net Jz = +1/2. On scratch, I have shown that if you take the general p spinors shown on page 30, the result is what you would expect. You find this: (iγ2)[ w4(p) exp(-iε4pμsμ /)] * = w1(p) exp(-iε1pμsμ /) (iγ2) = where I have used the fact that ε4= -1 and ε1 = +1. From the phase part, we would say that the r=4 solution had E = p0 and the resulting r=1 solution also has E = p0, in this tricky ε notation. Also, we can see that the (iγ2)ψ* operation "does not change p". That is to say, the w4 spinor starts off with p- and -pz, but the matrix changes to +pz and the * to p+ which are what you find in w1. We defined w4(p) as a function of p in the way it appears, so we cannot argue with the resulting claim that p is the same. Thus, in the sense just described, the entire pμ is the same for both ψ and ψc ! Again, you might think of ψ as being negative energy and having momentum -p, and its absence is ψc which has positive energy and positive p. So in the above analysis, we have to keep careful track of the ε factor in order to reach our conclusion that pμ is the same for ψ and for ψc. This is reflected in the energy projector part of 5.8 where you have to think of ε < 0 for the first projector acting on ψ, and ε > 0 for the projector in the last line which acts on a positive energy state. A similar interpretive thing happens with the sμ vector. In the rest frame, I have shown on scratch that ½ (1 + γ5) = ½ (1 + γ5γμsμ) = ½ (1 – γ5γisi) = ½ (1 + γ0Σs) which agrees with their slightly erroneous mid-page claim on page 69 (they have the final paren in the wrong place, a rare BD typo). Here Σ is the extended σ matrix. To show the above, you just multiply out the matrices in the larger 2x2 sense. In that world, γ5 is a "reverser", etc. So, since γ0 = we would say that our starting spinor w4, although "spin down" with -s, picks up the -1 from the γ0 matrix in the projector, whereas the w4 spinor with +s gets the +1 from the γ0 matrix, so we have a double negative in the first case. This then is why the covariant projector ½ (1 + γ5) is saying that both states have the same sμ. Again, we can interpret this in the "absence/presence" mindset if we want: the w1 state has positive spin because it is the absence of the negative energy state which has negative spin. This then is why we get the results shown on page 69 A. Remember that w4 (not w3) was v(p,s) and w1 was u(p,s) and we see that there are no minus signs. So now we see why these u and v objects were defined as they were, so that they come out to be "charge conjugates" of each other! Perhaps later authors will fiddle with the arbitrary phase factor shown. Technical Note concerning 5.8. Notice the object T = pμT γμT . What is meant by pμT ? Normally that would change a column vector to a row vector, but I think they really define T to mean pμ γμT so only the matrix is transposed and pμ are just some coefficients. One of my γ rules is that γo γμ† γo = γμ so if we add the effect of the *, then we can say that [ γo γo]* = T, so this explains how these T objects are appearing in the second line of 5.8. To get to the last line, you have to insert C-1C twice and recall from page 67 the result C-1γμ C = - γμT and that clinches the fact that the first terms both get sign-changed and we then end up with the last line of 5.8. Authors state clearly that ψc is the positron wavefunction. In mid page 69 authors describe a more general operator C in an operational sense which I would say is a symmetry operation of the Hamiltonian or Lagrangian. You do *, then Cγ0 on spinors, then negate the entire Aμ potential. I am quite familiar with this kind of symmetry, and the point made here is that if |ψ> is a Hilbert space state for an electron, then there must also exists states C |ψ> for the anti-particle, so we are definitely predicting the existence of positrons. 5.3 Vacuum Polarization (70) Now we have to think of an isolated normal electron as being "surrounded" by the Dirac Sea of negative energy and negatively charged electrons (not positrons!). Does the Coulomb force couple between positive and negative energy electrons? If so, we might expect the sea electrons to be "repelled" away from the isolated pe electron causing a region around our electron to have an effective positive charge distribution as illustrated in the second figure, where the first shows the charge distribution in some sense for the electron. The region of repelled sea electrons seems similar to the way a charge in a dielectric would align electric dipoles, causing a region of opposite polarization charge around the charge in question. For this reason, this effect is called "vacuum polarization" meaning Dirac Sea polarization. The figures indicate some scale R for this polarization, but nothing is said about R. Perhaps the "bare charge" is really 2e, but when you test from far away, you see the "dressed charge" which is only e, with the Dirac sea providing -e>0 from its integral. These are very vague notions that get supported later in field theory. There, if you are very close with your "test charge", energies are high and you can make pairs and these are associated with the "vacuum polarization" Feynman diagram and this is in fact the major part of the Lamb shift (which seems to contradict their comment here). But in field theory, "polarization" means a slightly different thing. You actually create a particle pair, causing a charge separation which is then the "polarization". Again, the Dirac Sea model is not quite physical. They "sidestep" the final question of why we don't see all that negative charge of the Sea with the comment that if it made an E field, what direction could it point in? Sure, an infinite uniform charge distribution can't make an E field. 5.4 Time Reversal and Other Symmetries (71) First, we are reminded that P = γ0 times a phase, and how this expression for P causes the Dirac equation to have form invariance, we did this back in Sec 2.3. Note added: What happens to the spin projector with P? γ0½ (1 + γ5) γ0 = ½(1 + [γ0γ5 γ0] [γ0 γ0]) = ½ (1 – γ5sμ [γ0γμ γ0]) = ½ (1 – γ5sμ γμ†) = ½ (1 – γ5sμ γμ) = ½ (1 – γ5sμ γμ) = ½ (1 + γ5') where s'μ = - sμ. which says that s'0 = -s0 but s' = s, so "the spin" in the rest frame does not change. This seems correct because we think that P converts w1(p) to w1(-p) so both should be "spin up". What about the phasor factor" pμsμ →(pμ)(-sμ) = - pμsμ so s.p must be a pseudoscalar. This seems correct because parity should change p but not an angular momentum thing which is a pseudovector. That is: vector pseudovector = pseudoscalar. Our problem now is to make the time-dependent SE have form invariance under t → t' = -t. BD kind of bollix this up a bit, I redid it on scratch. You start with 5.11, replace ψ = T-1ψ', apply τ from the left and you get this result for 5.12: – (T i T -1) ∂t'ψ' = T H T -1ψ' Trial and error shows that you should make T = TK, where K means complex conjugation. Then the above becomes + i ∂t'ψ = T H* T-1ψ' = T [ α* (+i -eA(t)) + βm + eφ(t) ] T-1ψ' where we know that and e and A and m and β and φ are all real. Then we rewrite, time-reversing the potentials and thinking of a current making A, and also priming , to get + ∂t'ψ = T H* T-1ψ' = T [ α* (+i' +eA(t')) + βm + eφ(t') ] T-1ψ' So we will achieve form invariance with the original equation i∂tψ = Hψ if the following properties are true: T α* T-1 = – α and T β T-1 = β On scratch I showed that a solution to this problem is this T = iγ1γ3 = T-1 (it turns out) T = TK where as usual we pick a phase arbitrarily. Recall the rules for "sliding γi around": (1) sign change every time it passes through some other gamma of the set 0,1,2,3 that is different. (2) get minus when one kills itself as in γjγj = -1. So once again, BD take something simple and make it unclear enough that I have to go off and do it myself. In 5.16 we now look at the effect of T on an E>0 electron state which has pμ and sμ . The T sandwich around a real 4-vector slash item like results in ' where the spatial part changes sign. So this sort of reinforces the idea that time reversal negates both p and s as you would expect "if time went backwards". The buzz phrase for this is "Wigner time reversal". Now suddenly we start talking about something called ψPCT(x'μ = –xμ). First, let's just assemble the pieces to compute such an item. ψPCT(x'μ = –xμ) = PCTψ(-x) = P (CK(TK(ψ-t, -x))) = P (CK(T(ψ*(t, -x))) = P(C(T*(ψ(t, x))) = eiφγ0 i γ2 (-iγ1γ3) ψ(t, x) = eiφ γ0γ2 γ1γ3 = – eiφ γ0 γ1 γ2γ3 = i i eiφ γ0 γ1 γ2γ3 = i eiφ γ5 ψ(x) which agrees with what they got. Without question, if the thing on the right were going backwards in space and time, then the thing on the left is going forwards in space and time. So now let's also assume that the thing on the right has E < 0, ie, is negative energy and is a momentum eigenstate hence ε < 1 and has some spin vector sμ. We can then play with the projectors as shown in 5.18. We start with an ε<1 electron with p and s as shown. After fiddling, we find the ψPCT has ε>0 with the same p and reversed s. Recall that from 5.8 that C alone gave you the same p and s, so the new thing here is that s got negated. So here is the bottom line: "positron ψ with E>0 [ namely, ψPCT(x)] = electron ψ with E<0 multiplied by ieiφγ5 and going backwards in spacetime [ namely ieiφγ5ψ(-x) ]. " Now with the fields on, this claim is confirmed as we move from A to B on page 74. To get from one to the other, you need to do the usual sandwich trick but this time with the PCT operator ieiφγ5. As usual, this is non-trivial so let's write it out: H = α (-i -eA) + βm + eφ with Hψ = - E ψ E > 0 equation A If we sandwich PCT around this H, the complex conjugations I think cancel out from C and T, and we are left worrying about [ note that γ5α γ5 = + α since both γ0 and γi get sign changes ] γ5[α (-i -eA) + βm + eφ] γ5 = [ + α (-i -eA) - βm + eφ ] where we have only worried about the γ5 effect. Next, we need to negate spacetime in all the arguments. The potentials don't change so we have = [ + α (+i' -eA') - βm + eφ' ] = – [ α (-i' +eA') + βm – eφ' ] and this is the exact negative of what we started with except for spacetime reversing primes and the charge sign being reversed. . Then we move the sign to the right and end up with [ α (-i' +eA') + βm - eφ' ] ψ'PCT = +E ψ'PCT // this is equation B Were we to change the sign of e, we recover the same form as the original equation. I will now stick with the previous quote and say: "a positron ψ with E>0 and q = |e| [ namely, ψPCT(x)] = electron ψ with charge -|e| and E<0 multiplied by ieiφγ5 and going backwards in spacetime [ namely ieiφγ5ψ(-x) ]. " Since we know what the charges are of positrons and electrons, we can omit that part and rephrase as: "positron ψ with E>0 [ namely, ψPCT(x)] = electron ψ with E<0 multiplied by ieiφγ5 and going backwards in spacetime [ namely ieiφγ5ψ(-x) ]. " I don't think you can leave out the " multiplied by ieiφγ5" part of this statement. This concept of interpreting a positron as related to an electron going backwards in spacetime is called "the Feynman-Stuckelberg form or positron theory." paper 1941. The last paragraph of this section and the chapter is another comment on the notion of the "symmetries" of the Hamiltonian or Lagrangian, hence "of the theory" ( I should say, symmetries of the equations of motion, I suppose). You can consider different "interaction terms" in the Hamiltonian and ponder what symmetries they have. The feeling is that the Dirac theory ought to apply to "other" spin-1/2 particles like protons such that all three of our basic symmetries P,C, and T are preserved. Having said this, authors make two closing comments: (1) Lee and Yang showed that the weak interaction violates both P and C. The present theory of weak interactions and experiment confirm this fact (my comment). (2) Authors claim that PCT combined must always be a symmetry if you assume proper Lorentz invariance and the usual spin-statistics relationship. But they do not prove this claim! I am not sure where I have a proof of this theorem (but I surely do), and not clear how to find this on the web. My impression is that PCT combined is what gives you the true antiparticle. Chapter 6: Propagator Theory (78) 6.1 Introduction. The aim is "to calculate things" in principle "exactly". Field theory is a big load, so they want to put it off. Follow the Feynman method. 6.2 The Nonrelativistic Propagator (78) Quick Math Background on Green's Functions. Stakgold Volume 1 page 261 defines what is meant by a Green's Function. You have a differential equation Lu = f(x) let's say in 1D, where L is a differential operator. Then the Green's Function g(x',x) = G(x'-x) is the solution to the differential equation when the RHS is "driven by" a delta function, so we have Lg(x,x') = δ(x-x'). The usual method is to first solve a problem for the Green's function with a δ driving source, such that the Green's meets the boundary conditions. Then, once you have the Green's, you can write down the solution to Lu = f(x) as follows: u(x) = ∫dx' g(x',x) f(x') The reason being that Lu(x) = ∫dx' Lg(x',x) f(x') = ∫dx' δ(x-x') f(x') = f(x) This is the basic idea. We can work with a differential equation Lu = f(x) and try to solve it for u(x) with some boundary conditions, or we can find the Green's function from Lg(x,x') = δ(x-x') also subject to some boundary conditions, and then our solution is u(x) = ∫dx' g(x',x) f(x'). I recall doing this a lot in Jackson's book and class where the differential equation was perhaps the Poisson equation 2φ = ρ and we found Green's functions that met certain boundary conditions like constant on a sphere, say. I cannot remember whether there were scattering situations in Jackson that used Green's. In our current context, we have a partial differential equation and multiple variables. Fine. In this chapter, the Green's Functions will be "particle propagators", amplitude to go from x to x'. Application of the above. In this section we will use the following Green's functions: (i∂t – H ) G(x,x') = δ4(x-x') (i∂t – Ho) Go(x,x') = δ4(x-x') We can then show that "Huygens's Principle" of 6.1 is just a restatement of the SE. There we write the Huygen's rule as this θ(t'-t) ψ(x') = i ∫d3x G(x',x) ψ(x) // we only scatter forward in time! Causality. If we apply (i∂t – H) to the LHS, we get 0 (from the SE) + the term generated when i∂t hits the θ function which is just iδ(t-t') ψ. Meanwhile, on the RHS (i∂t – H ) G(x,x') = δ4(x-x') and the δ3 kills off the integral, and the δ(t-t') is left over and we get the identity that iψ = iψ. That is why the i has to be on the RHS, because it is part of i∂t. So my main point here is this: We don't have to "accept" Huygen's Principle 6.1 as some kind of non-derivable starting point for this section, thank goodness. In fact, we just derived 6.1 above, and we find that it is entirely equivalent to the SE; it has the same information, it is an alternate form of the SE. The claim will be that this form is more conducive to having scattering type boundary conditions. Now we can start the fancy presentation of this section. In 6.2 we "turn on" a potential for differential short Δt1 at time t1. In 6.3 we integrate 6.2 over time from t1 to t1+Δt1 and Δψ includes only the additional amount of Δψ caused by the presence of V ≠ 0 during this time, so we don't show the base component of Δψ which we get integrating 6.2. We have φ (plane wave) on the right because that's what ψ was before V hit at time t1,. Then in 6.4 we insert this special Δψ amount on the right of a contribution form of 6.1, the i's cancel. After the tiny Δt1 scattering event, we propagate with no potential hence with Go. So the claim is we are showing a particular "contribution" lying within 6.1 which is this momentary scattering contribution of V turned on for just a moment at time t1. Turn V off all the time, this contribution is zero. Then 6.5 tells us that the total amplitude at x' is the original wave φ plus this scattered contribution. This explains the first equality in 6.5. Here is a derivation of the second equality, where we use some compacted notation: f = x' = xfinal a = x = xa. ψf = φf + ∫d3x1 Gof,1V1φ1 dt1 φ1 = i∫d3xa G01,a φa => ψf = i∫d3xa G0f,a φa + ∫d3x1 Gof,1V1[i∫d3xa G01,a φa] dt1 = i∫d3xa { G0f,a + ∫d3x1 dt1 Gof,1V1 G01,a } φa These two terms appear as diagrams (a) and (b) on page 80. We can regard the quantity in {...} as our full G(f,a) for this particular situation, and that is what 6.6 says. Next we turn on V again at t2 > t1 . This causes two new terms, the single scattering of (c) and the double scattering of (d) and these terms are written in 6.7. And so we build up a mountain of scattering terms of different orders. In 6.9 = 6.10 authors show all single, double and triple scattering terms in compressed notation with the time ordering maintained as shown. If we put a time θ function inside Go ( we then have a retarded propagator with > speed of light) we get 6.11. If you break out just the very last scattering, you get the integral equation 6.12 where we now have some assumptions about convergence of the series and V being small, etc. Of course that will be the case in QED which is mostly what this book is about with the 1/137 factor. We can think of our drawings as "Feynman diagrams" for our current example which is the non-relativistic SE since we have no 4x4 matrix stuff going on (although Ho has not yet been stated). Now we come to the notion of a scattering experiment and boundary conditions. We assumed some initial state φ that was "far from" the action of the potential V which we now assume is somehow localized near a scattering center of some sort. So we already have that BC worked into our equations above. Equation 6.13 shows the full G carrying the propagation load from the distant past to some present spacetime point x'. This is an id3x thing like 6.1. Equation 6.14 inserts 6.12 into 6.13 to get the first expression, but then it uses another version of 6.13 going to point x1 to replace the id3x G(1...)φ(x..) with ψ(1) and this gives the second expression in 6.14. So this thing is now a traditional "integral equation" for the wavefunction ψ, and it involves G0 and V as shown. I recall integral equations which "show the last rung" from my multiperipheral model work. Things are clarified in the next "Note Added". ___________________________________________________________________________________ Note added: Let's look at these things in matrix form a bit: Notation: I denote d3x integrals with a ".", and d4x integrals with no "." We have: (I am also keeping track of the factors or i ). When φ appears on the right, it usually means we have taken it into the far past. If only in the finite past, it would be ψ such as in 6.1 ψ' = iG.ψ 6.1 or 6.17, ψ in the finite past , this is Huygens' = SE G = G0 + G0VG 6.12 integral equation for G, driven by V and Go ψ' = iG.φ 6.13 prime means x' and t' in future φ' = iG0.φ how G0 would move φ into the future ψ' = iG.φ = i(G0 + G0VG).φ = (iG0.φ) + G0V(iG.φ) = φ' + G0Vψ , so we have ψ' = φ' + G0Vψ 6.14 integral equation for ψ, driven by V and Go We need some comments here!! In the above, φ always refers to a simple plane wave or state of system where V = 0. If you apply the full complexity of scattering represented by G, you get at a later time that ψ' = G.φ, so symbol ψ here represents some highly complicated wavefunction that is, if you like, the superposition of an infinite number of Feynman diagrams. But if we use only the straight-through diagram which means G0, we get φ' = G0.φ as the evolution. In the above equations, I had to prime things just to show them as being in the future to avoid writing things like φ = G0φ which, nevertheless, is a valid matrix equation! It says φ(x') = IntxGo(x',x)φ(x) or φx' = (G0)x',x φx. So we have two integral equations above with are really equivalent: G = G0 + G0VG 6.12 integral equation for messy G d4x ψ = φ + G0Vψ 6.14 integral equation for messy ψ d4x What about expressions for the S-matrix? In 6.16 we see this: Sfi = φf†.ψi(+) = φf†. iG.φi = φf†. ( φi + G0Vψi(+) ) = φf†. φi + φf†.G0Vψi(+) = δfi + φf†.G0Vψi(+) This is written in this odd way because the second term has all graphs with 1 or more V interactions, so we have exposed the "straight through" term as δfi. We could replace ψi(+) = iG.φi on the right to get this alternative form Sfi = δfi + i φf†.G0VG.φi And here for fun is a longer derivation of these last results: Sfi = φ'f†.ψi(+) = φ'f† . (iG.φi) = iφ'f†.G.φi = φ'f†. ( iG0 + iG0VG) .φi = φ'f . (iG0 .φi) + iφ'f† .G0VG. φi = φ'f .φ'i +i φ'f† .G0VG .φi = δfi + i φf† .G0VG .φi So we can summarize all these ways of writing Sfi Sfi = δfi + φf†.G0Vψi(+) = δfi + i φf† .G0VG .φi = φf.G .φi and 6.16 is the first equality above. Now let's do one more form for Sfi. Suppose we write out Go as the basis function sum shown in 6.26, then the term shown above becomes φf'†.G0Vψi(+) = φf'†.{-i θ ∫d3p φp' φp†} Vψi(+) = –i φf† V ψi(+) for going into future Notice that we have used orthogonality to say φf†. φp = δ3(pf - p) =∫d3x <pf |x't'><x't'|pi> (the dot means a d3x integral). I have put a few primes on things in the above equation line to show that the orthogonality is allowed because the two factors in question are at the same space and time. In the final result above, notice there are no "dots" now. We have therefore a single d4x integral. So we can add this "one more form" to our list above and summarize as follows: The plane waves in the above are really φp = e-ipx e-iEt /(2π)3/2 which then gives us the delta function as claimed. S-Matrix Summary There are very many ways to write Sfi and they are all "derived" right here. Form 5 was derived about 2" up above. Form 1 is really the definition of the S matrix. Supporting material comes first. θψ' = iG.φ 6.13 prime means x' and t' in future (0) θφ' = iG0.φ how G0 would move φ into the future (1) G = G0 + G0VG 6.12 integral equation for messy G (2) ψ = φ + G0Vψ 6.14 integral equation for messy ψ (3) Sfi = φf†.ψi(+) form 1 6.16 defines Sfi Sfi = i φf†.G.φi form 2 6.30 f1 + (0) Sfi = δfi + i φf†.G0VG.φi form 3 f2 + (2) + (1) Sfi = δfi + φf†.G0Vψi(+) form 4 6.16 f1 + (3) Sfi = δfi – i φf† V ψi(+) form 5 6.34 see note above Sfi = δfi + φf† V G.φi form 6 6.33a f5 + (0) Sfi = δfi – i φf† V (φi + G0V(φ + G0Vψ)) form 7 f5 + (3) twice = δfi – i φf† V φi – i φf† V G0V φi – i φf† V G0V G0V φi + ... 6.33b This last form seems to suggest that (recall 1 + x + x2 + ... = 1/(1-x) ) iG = 1 -iV -iV G0V -iVG0VG0V + ... = 1 -iV(1 + GoV + G0V G0V + ...) = 1 - iV(1/[1- G0V]) => iG - 1 = -iV(1/[1- G0V]) => (iG - 1) (1- G0V) = -iV ???? If V = 0, the last result seems to say that iGo = 1 which is true sort of in the sense of (1) above. Well, BD don't really do justice to a formal presentation of "scattering theory" but I know I have other books that do. So we have to put this in the same category with the hydrogen atom solution. At least the equations above with BD equation numbers must be true. One could imagine all these things written in bra-ket operator notation. Then G would be the S matrix operator! It controls the full amplitude between the φi initial state and the φf final state. Comment on the i's. I said earlier that iG and iGo are natural combinations. Another natural combination is -iV, which I failed to notice. This appears for example in (6.3) on page 79, and this gets there due to the "i" in 6.2 which is i∂t. This means that the combinations VG and VG0 don't change sign if you "show the i's", namely, VG = (-iV)(iG). So wherever these two appear together, there is no "i" to think about. So the integral equation G = G0 + G0VG can be thought of as iG = iG0 + iG0(VG) . This idea gives you a way to check the signs in all the above equations. When the i's are properly bound, all signs are +. For example, we will always have Sfi = δfi + something. The first -order perturbation result is a good example. It says Sfi = δfi + φf† (-iV) φi . ___________________________________________________________________________________ On page 83 we then assume an initial φ of a plane wave, our final observation is a similar plane wave at some different kf wavevector and we define the S matrix Sfi as in 6.16 right at the start. The pass through of the original wave gives the usual k-space delta function, and the second term contains the trailing part of 6.14 still with the full ψ in there, now called ψ+ because it is moving forward in time from the staring position as φi. We can then use the diagrams to make a series of terms for our S matrix amplitude. This is a little rough right now, but this is exactly the path we will take. This entire discussion appears as Chapter 9 in Schiff which is probably my best reference. I wonder if I have some notes on that Schiff chapter? My "scattering" sources and notes are very wide spread! Comments: This S matrix business on page 83 applies to the scattering of a single particle by a scattering center, think of Rutherford scattering of alpha particles by a nucleus with no recoil. The single particle in question has wavefunction ψ+(x,t). The scattering center is described by some V(x,t). So at this point, we don't have an S matrix with several particle legs coming in and several coming out, which is more what I am used to. So in this one-particle case, we have the incoming and outgoing asymptotic plane waves φ to consider and that is about it. 6.3 Formal Definitions and Properties of Green's Functions (still non-relativistic situation) Stakgold has much to say about Green's functions, in fact Chapter 1 is so titled. Here we are dealing with an "equation of evolution" with a linear time derivative which is like the heat = diffusion equation. The logic on page 84 is just showing that our Huygen's Principle now written as 6.17 is equivalent to the Green's function definition that LG = δ4 which is then stated in 6.22. Along the way we use the interesting integral representation of the δ and θ functions shown on page 84. The time θ (Heaviside) turns out to be absolutely crucial to everything, you cannot haphazardly ignore it! I will just restate what I said above (i∂t – H ) G(x,x') = δ4(x-x') (i∂t – Ho) Go(x,x') = δ4(x-x') θ(t'-t) ψ(x') = ∫d3x G(x',x) ψ(x) // we only scatter forward in time! Causality. Starting page 85, authors undertake to actually compute Go(x,x'), something I recall doing long ago somewhere. The first step is to get the problem over into Fourier momentum space (p = k) and we quickly find the result 6.25 that says Go(p,ω) = 1/(ω - p2/2m + iε). By choosing the sign of the ε in this way, in 6.26 we are then able to generate the correctly signed θ function in the resulting Go(x,x'). Here, the resultant Go in spacetime (not momentum space) in 6.26 is written as a spectral sum over our plane wave basis functions that looks a lot like the completeness condition. We see that our plane wave situation is just a special case of a more general situation. Assume that you have a differential operator L in space and time, like (i∂t – H), where we are in an "evolution situation". If we assume that G can be written as in 6.28, which is a θ times a spectral sum of the eigenfunctions of L, then if we apply L to both sides look what happens: the LHS of course just becomes δ4(x-x') . The RHS has a 0 term where L hits on one of the ψ's (having the right x argument), but also the term you get when ∂t hits on the theta function. This creates δ(t-t') which forces the time arguments in the sum to be the same, and they then produce δ3(x-x') as per the completeness condition in 6.27. So we then get LHS = RHS = δ4(x-x'), so our assumption that G can be written as in 6.28 must be true. In the footnote, authors state the closed-form result of integrating the RHS of 6.26 and they comment that this is very similar to a diffusion equation result they know of. Now backing up, the ψn are eigenfunctions of Lψn = λn ψn which is H ψn = En ψn in our case. In general we have these being orthonormal and this fact is then time-independent as I now show: < ψn(t)| ψm(t) > = <U ψn(0)| Uψm(0) > = < ψn(0)| U† U ψm(0) > = < ψn(0)| ψm(0) > = δmn where U = exp(-iHt/) is the usual Hamiltonian time driver which is unitary because H is Hermitian. This same idea of time independence of the inner product is used in the next paragraph! We now come to page 87 top. Authors first show that the same G which takes ψ forward in time, propagates ψ* backwards in time. You have to stare at things a little to see where x and x' are located in 6.29 versus the previous equation. The rest of this section seems to rehash the scattering expansion for G, and for the S matrix S. I don't think there is anything new here. Then at the end, authors show that S is unitary because the orthonormality of the initial plane waves is maintained as they go through the scattering region. This is the famous "unitarity of the S matrix" which I know just says the sum of the probability of all outcomes is 1. Geoff Chew always thought this would act as a serious constraint on the hadronic scattering amplitude perhaps enough so to actually determine it along with things like analyticity and crossing. I think he has given up on that program but of course one can never tell. Equation 6.30 is still a single-particle S-matrix amplitude and you have to keep thinking of the scattering center as a potential V perhaps distributed over some space near r = 0 and time near t=0. A plane wave φi arrives from the distant past, and we look at the various outgoing plane waves φf in the distant future. Equation 6.31 breaks out the H' term of the Ham and puts it on the right, a no brainer. The second line is just fiddling with delta functions, nothing new here. But then as I have marked in the book, you can treat the common delta function on the RHS as LG0 acting with the x' variable, and you can extract the operator portion L out of the x" integral. Then you compare the things acted upon by this operator and set them equal. This is loosely referred to as "doing an integral", and we then get G = G0 + G0VG where I am now using a matrix form to account for spacetime integrations, and this is then 6.32 which is just an integral equation for G whose kernel is K = G0V. We could identify this integral equation with page 195 Stakgold equation 3.6 where we set u = G, k = VG, μ = 1, and f(x) = G0. We think either of just one variable of G to get u. This is a Fredholm integral equation. You can see how this integral equation leads to a sum of series situation, G = G0 + G0VG = G0 + G0V(G0 + G0VG) = G0 + G0V(G0 + G0V(G0 + G0VG)) etc = G0 + G0VG0 + G0VG0VG0 + G0VG0VG0VG0 + ... We saw this thing on page 81 in 6.9 earlier and rewritten in 6.11 with built-in time thetas and then 6.12 was our integral equation G = G0 + G0VG. So this is why I say we are doing a "rehash" here of what we already did before. If we insert the above expansion for G into our S matrix formula 6.30 we get a series for Sfi which reflects this series, and a typical term in this series would be shown in the page 90 picture where, for now, we restrict attention to going sequentially forward in time with our scattering interactions with Vat different spacetime locations, still presumably glommed near x=0 and t=0. 6.4 The Propagator in Positron Theory (89). The title of this section is a little odd, I think it should be The Propagator in the Dirac Theory of the Electron and Positron. The previous section lays the groundwork where we deal with full spacetime d4x stuff and we have the "retarded" Green's function propagator with its internal θ function which maintains causality. In this upcoming long 10-page section, we will have to add all the 4x4 matrix machinery, and we will have to deal with the negative energy solutions as positrons and all that stuff, it should be an interesting voyage. The opening 3 pages through bottom page 92 are considering the same scattering scenario as in the non-rel previous section, but with a few differences. We no longer distinguish time, but imagine a scattering center to be V(x)d4x. Then we get to the pictures on page 91 which we knew were coming. We are told on page 90 that "one may say that at each vertex, a particle is destroyed and another created", anticipating Volume II and QFT. Then on page 92 we are told we can interpret the backwards scattered paths as electrons going backwards in spacetime which we otherwise interpret as positrons going forward in spacetime. Then we can have arbitrary zigzagging of the lines always shown forward on page 90. They are "relying on intuition" rather than "rigor". So finally on bottom page 92 we start in on the 4x4 gritty details. Recall in the previous section we defined our Green's in 6.22 by applying L = (i∂t-) and getting LG = δ4. I put a hat on the H to emphasize that it is a spatial differential operator that acts on G(x,x'). So the analogous thing in the Dirac theory is to take L = ( - e -m) which again is a spatial differential operator, and it is also a 4x4 matrix. We then state this matrix equation LG = δ4 and the full G is called S'F. We are of course allowed to ask for the G of any operator L we come up with, and S'F is the Green's for L as just stated, no matter that we have L = γ0(i∂t-H). Authors are putting primes on everything to indicate x' in the operator which is fine, I might have switched x and x' in this presentation. The prime on S'F is how they want to indicate the "full" propagator, reserving the symbol SF for the free propagator. That is, G = SF' and G0 = SF when we compare to the non-rel earlier presentation. So of course 6.40 then defines SF. We then repeat the earlier trick of going to momentum space and we quickly determine that SF(p) is as shown in 6.42 where the famous 1/(-m) is just a shorthand notation for the item shown, and I note in passing that = p2 which is easy to show and which BD have not mentioned ever. Now when we try to return from p-space to x-space, we have our d4p integration to carry out and, as before, we ask about the contour for the dp0 integration. The integrand has poles in p0 at the locations shown in the drawing, with the displacements also as shown. These displacements are positioned for a special reason. Based on the form SF(x' - x) going from x to x', if t'>t we are moving into the future, we need to close the contour down to get expo damping exp(-ip0(-iR)) = exp(-p0R). This picks up the pole on the right side which is the "positive energy pole". We are going the "wrong way" around this pole so that makes one minus sign. The pole residue is then -2πi (1/2E) and we get the form shown in 6.44 where BD have written out the +m numerator in detail showing the positive energy E. So when you go into the future, this scheme says you only have positive energy contributions in the integral which makes SF. Looking at 6.43 where we see an example of a positive energy positron ψc and a positive energy electron ψ, we note that both these things have positive frequency. Charge conjugation you may recall was ψc = (iγ2)ψ*, so although we started with ψ as a negative energy ψ, the * makes it positive energy. So the point of this is that our first form 6.44 includes only the effect of positive energy positrons and positive energy electrons moving into the future. We then obtain 6.45 for "going into the past" and we find all negative energy components, again, with our pole position decision. It does seem anti-causal to be able to propagate anything into the past, but we always have this interpretation issue. Also, a LT can change the meaning of "the past". Look at page 91 first figure. We can think of pair production as a pe positron and a pe electron going into the future. OR, we can think of starting at the upper left and we have a positron going into the past and then it turns into a positron going into the future. I think we are going to need "both directions" to get things to work right. By the way, in getting 6.45 we had the normal pole contour direction, so you might think there should be an opposite sign. But the residue is now 1/(2*-E) and that makes another minus sign, so these two minus signs cancel and 6.45 then has the same minus sign that 6.44 has. In any event, we see what the pole position choice does, and I agree with 6.44 and 6.45. These two equations are combined in 6.47 using time theta functions and projectors for the +m type factors. Notice that 6.46 is somewhat out of presentation order and has d4p. They are just showing how a single iε in the denominator can cause all the correct pole positions. The (m/E) factor arises from the projector definition as shown (the m) and from the pole residue shown in 6.44 and 6.45 (the E). Now for the first time in the book, in equation A they write the properly normalized plane wave including its: normalization factor, the spinor wr, and the spacetime expo (which has εr in it). We can then verify 6.48 by pulling off the expo and normalization stuff, and using the facts claimed top of page and earlier in my notes regarding the alternate forms for the projectors, see discussion above about equation 3.9b and "completeness". Again, this was something BD did not mention anywhere. So I am very happy with 6.48. What about 6.49? My first hour spent on this had me trying to insert 6.48 for SF into the RHS of 6.49. I kept making two bad mistakes and I was getting bogus results. The first mistake is this: the spinors show up in the form ww† in 6.48, which is really a 4x4 matrix (a completeness partial unity). The equation they quote as "with the aid of" is 3.11 on page 31, but this is of the form w†w which is a scalar. So it is totally unclear to me how I could make use of 3.11 in the process just described. Notice that the actual component indices of the spinors are not written anywhere on this page. My mistake was accidentally using the scalar 3.11 result for the ww† matrix object. My second mistake in the above program was thinking I could do the d3p on the spatial expo to get δ(x-x'). This was wrong because p also shows up in the energy part of the expo in the form , so you don't really have a simple integral representation for a δ3 function! I don't really know how they intended to use 3.11 to verify 6.49, but I have my own way. Just apply the operator (' -m ) to both sides using result 6.40. Since ψ+ satisfies the Dirac SE, on the LHS we pick up a term only where we trip up on the theta function. That is to say (' -m ) = γo(i∂t' – ') and so γo(i∂t') θ(t'-t) = +i γo δ(t' - t). So action with (' -m ) on the LHS of 6.49 gives +i γo δ(t' - t)ψ(+)(x'). Now apply (' -m ) to the RHS. We get the δ4 as in 6.40 and RHS = i γ0 ψ(+)(x') δ(t' - t) and we have then shown that both sides match under the action of this operator. This is good enough proof for me. Notice that the only use of the (+) label is that this reminds us we are propagating into the future and that sets the sign of the argument of the θ function on the LHS. And notice that the γ0 must be there on the RHS of 6.49 as they show, because my proof needs it there! The derivation of 6.50 is exactly the same, but the (-) label goes with the negated θ function argument, and of course that causes a minus sign if you redo the line of algebra shown above. I am sure I could find a better way to "prove" these guys 6.49 and 6.50. I recall a certain trick but not clearly enough to use it. It lets you do the d3p integration even though that energy expo is there. You elevate to a d4p integration by adding another delta function, then something happens and you solve the problem. I could not find this trick in my binder TK notes, but I am sure it will arise sooner or later in my readings. So we are now at the bottom of page 95. The Swiss Ernst Stuckelberg figured things out in 1942 but didn't do much with it, Feynman put the thing together in 1948 and gets most of the credit because he actually did calculations and got right answers. The name Dyson is not mentioned by authors. Now top page 96 repeats our previous gimmick of breaking out the H' term from the Ham and putting it on the RHS. Now this term is eSF' rather than VG. We then get our integral equation for the full SF' as we got before, now we have the 4x4 matrix stuff as well, but we still write: (see note added above) G = G0 + G0VG (6.12) → SF' = SF + SF e SF' // 6.51 and we have our "alternative" integral equation mentioned above ψ = φ + G0Vψ (6.14) → Ψ = ψ + SF e Ψ // 6.53 In this last case, we were already using the symbol ψ for the plane waves, so we had to go to Ψ for the full "messy" wavefunction (spinor). Now what are the next two equations 6.54 and 6.55 saying? The messy object Ψ exists at all times. If we evaluate Ψ(x,t) when t→ – ∞, it is so far back in the past that the only possible contributions to it from the integral Ψ – ψ = SF e Ψ must be " backwards" contributions with negative energy propagation and these come from the part of SF which has the negative energy projector Λ- in 6.48 so we get only the 1- factor which is the r=3,4 sum, and we have then 6.55. The reverse statement is true to get 6.54. Now comes the mystery comment in the paragraph after 6.55. The first comment is about 6.54. We are talking forward scattering here of an electron. We have the Λ+ projector sitting there. So we need now more interpretation of 6.54, having already derived it. In the rightmost integral, we have the incoming messy electron Ψ scattering only into positive energy plane wave states rp for r = 1,2. This integral all by itself is some kind of S scattering amplitude. Of course p is summed over d3p. We then take this amplitude integral and multiply it by ψrp now at x which is the final location to get our final Ψ(x). So somehow this is saying that if we decompose what is happening inside the integral into a momentum integral of plane wave intermediate states, we find that only positive energy states are present. Sort of a filter. Hence the comment that the "electron" Ψ cannot "fall into" negative energy states in this integral after scattering by . OK, these are just interesting words and phrases, the equations have the truth. Now suddenly we have 6.56 but this is just a translation of our earlier non-rel formulas. Recall from above that we had this as one of our ways to write Sfi Sfi = δfi – i φf† V ψi(+) form 5 => Sfi = δfi – i f e Ψ i(+) // 6.56 where again we make the changes V → e, ψ → Ψ, φ → ψ, † → overbar. The only complication is that we have a different sign in 6.56 if the final state is a negative energy state, and this arises from the minus sign in the 1- projector noted above somewhere. Now to page 97 of this action-packed drama. We have our S-matrix amplitude in 6.56, and we can insert into this the integral equation 6.53 "iteratively" (as we have done earlier in these notes) and we get 6.57 for the nth order term. There are n powers of e here, and n occurrences of and n integrals. This is a Feynman graph with n vertices and we know that all time orderings are allowed and in fact contribute. That is to say, we would include a path like that on page 90 as well as a path like that in 6.5 b. ___________________________________________________________________________________ Very Long Note inserted: The purpose here is to obtain all the Dirac results which correspond to all our non-rel results. Many of these items don't get any mention in this chapter. Again, a problem with the haphazard presentation method of our authors. But I am happy to do it. This note goes on for about 7 pages and ends with our a summary table. Let's compare this with our earlier non-rel perturbation expansion Sfi = δfi – i φf† V φi – i φf† V G0V φi – i φf† V G0V G0V φi + ... 6.33b Sfi = δfi – i f e ψi – i f e SF e ψi – i f e SF e SF e ψi + .. . 6.57 Where does the bar come from? Let's review a bit the parallel development in previous and present sections. We have (i∂t – Ho) Go(x,x') = δ4(x-x') Ho = 2/2m for free particle (i∂t – H ) G(x,x') = δ4(x-x') H = H0 + V γ0(i∂t-Ho) SF(x,x') = δ4(x-x') Ho = α + βm = γ0(γ + m ) for free particle ( - m) SF(x,x') = δ4(x-x') γ0(i∂t-H) SF'(x,x') = δ4(x-x') H = H0 + V = H0 + eφ – α (eA) // see 4.2 p48 ( - e – m) SF'(x,x') = δ4(x-x') = γ0(γ + m + e γ0φ – γ (eA) ) = γ0(γ + m + e ) So the first thing we notice (moving toward an answer to our question) is that an extra γ0 appears on the LHS in the Dirac case compared to the non-rel case, and we see that this was put there so that SF comes out being a world-scalar, or at least so the defining equation has a covariant look: ( - m) SF = δ4. So now we have a γo "floating around". What is the correponding Huygens' Principle? θ(t'-t)φ' = iG0.φ how G0 would move φ into the future // see 6.1 with Go not G θ(t'-t)ψ' = iSF γo.φ how SF would move ψ into the future // see 6.49 θ(t'-t)ψ' = iG.ψ the "full" versions of the above θ(t'-t) Ψ ' = iSF' γo. ψ Recall from above that (' -m) = γo(i∂t' – ') and so γo(i∂t') θ(t'-t) = +i γo δ(t' - t) which says we get +i γo δ(t' - t)ψ for the LHS of the second line above. The RHS then uses ( - m) SF(x,x') = δ4 and we get RHS = i δ(t' - t) γo.ψ. This then is WHY the γ0 appears in the Huygens for SF. It has to be there due to the way SF is defined. So we see that γ0 continues to "float into" our scattering equations. Now in the Dirac world, we have a Huygens' for going backwards in time. It is the same as forward except for a minus sign, so θ(t-t') ψ' = – iSF γo.ψ You verify this the same was as before, and now (i∂t')picks up a minus sign on the left, which then matches the mins sign on the right. So perhaps we would combine these together and write θ Ψ ' = ε iSF' γo. ψ (0) where ε is the energy signature of ψ, and θ as apropo What about all the other non-rel equations, how do they appear in Dirac world? We have already seen above (derived BD page 96) that the integral equation for SF is exactly parallel to that for G, G = G0 + G0VG (6.12) → SF' = SF + SF e SF' // 6.51 so there are no γo factors in this integral equation for SF'. But let's review anyway how this works: non-rel: (i∂t' – Ho')G = δ4(x'-x) + VG = ∫d4x" {δ4(x'-x)} [δ4(x"-x) + V(x") G(x";x) ] = ∫d4x"{(i∂t' – Ho')Go} [δ4(x"-x) + V(x") G(x";x) ] = = (i∂t' – Ho') ∫d4x" Go [δ4(x"-x) + V(x") G(x";x) ] => G = Go + GoVG relativistic (' - m)SF' = δ4(x'-x) + e SF' = ∫d4x" {δ4(x'-x)} [δ4(x"-x) + e(x") SF'(x";x) ] = ∫d4x"{(' - m)SF} [δ4(x"-x) + e(x") SF'(x";x) ] = = (' - m) ∫d4x" SF [δ4(x"-x) + e(x") SF'(x";x) ] => SF' = SF + SF e SF' So this shows that no γ0 appear here because nothing we do "exposes" a γ0. Next, let's examine item (3) in our table above which says ψ = φ + G0Vψ in the non-rel case. Let's repeat the above comparision to see what happens with this integral equation for the wave function. non-rel: ψ' = iG.φ = i(G0 + G0VG)φ = ( i G0 φ) + i G0V(i Gφ) = φ' + iG0Vψ rel: Ψ ' = i ε SF'γo. ψ = i ε (SF + SFeSF') γo. ψ = (i ε SFγo.ψ ) + SF e (i ε SF'γo ψ) = ψ' + SF e Ψ So here we see the γo making an appearance, but in the end they disappear. Also, ε appears, but then it too goes away. The above equation appears as 6.53, So now we have formed the start of the long table we had for the non-rel case, to wit: θ Ψ ' = ε iSF' γo. ψ (0) forward in time θ(t'-t) and ε = 1, else both opposite θ ψ' = ε iSF γo.ψ (1) SF' = SF + SF e SF' (2) Ψ ' = ψ' + SF e Ψ (3) Now finally we come to the S matrix stuff. The definition would be this: Sfi = φf†.ψi(+) => Sfi = ψf†. Ψi(+) form 1 where the † is there this reason: The non-Dirac scalar product is <ψ|Ψ> = <ψ|x><x|Ψ> = ψ† Ψ where if we want to use x-space matrix notation, ψ† = ψ*T is a row vector of starred items, the star from the Hilbert space basic inner product definition in QM. In the Dirac case, we have this same notation, but the † has the additional meaning of transpose in the 4x4 Dirac matrix space. Now we can install our Huygen's into the above to get Sfi = ψf†. Ψ ' = ψf†. (iSF' γo. ψi) = i ψf†.SF' γo. ψi form 2 // analog i φf†.G.φi Sfi = δfi + i φf†. SF e SF'. ψi form 3 f2 + (2) + (1) Sfi = δfi + ψf†. SF e Ψ form 4 f1 + (3) Now we come to "form 5" which we know has a fancier derivation than the other forms. We are going to start with the term above ψf†. SF e Ψ and insert SF from 6.48 , (lower line is for comparison only) ψf†'. SF e Ψ = ψf†'.{ -iθ∫d3p Σr=1,2ψpr' pr +iθb ∫d3p Σr=3,4ψpr' pr } e Ψ φf'†.G0Vψi(+) = φf'†.{-i θ ∫d3p φp' φp†} Vψi(+) = –i φf† V ψi(+) for going into future As before, the primes indicate x'μ coordinates space and time. How does suddenly appear. It appears because in the Dirac world, completeness and orthogonality have bars, as shown page 30, and this allows for covariant equations. The bars above come from our 1+ positive energy completeness = the positive energy projector. Now in the Dirac world, we are using the normalization shown in page 95 A. The factor was not present in our non-rel work, nor of course was the spinor, though the π factor was present. Now in the non-rel case, we just asumed φf was a plane wave with pf momentum and the spatial integration yielded up a δ3(pf - p). In the Dirac case, in order to continue we need to specify whether ψf is a pe or a ne plane wave with pf. So let's first assume that ψf = ψfs = plane wave with s = 1 or 2. In this case, according to 3.11 we get no contribution from the second term, and get a hit on the r = s term in the first term, so we get ψsf†'. SF e Ψ = ψsf†'.{ -iθ∫d3p Σr=1,2ψpr'pr} e Ψ = -iθ ∫d3p ψsf†.ψps ( pse Ψ ) where we have again used 3.11 to remove the r=1,2 sum (it hits on one term or the other). At this point, the implicit d3x' of the "." integration does this: ∫d3p ∫d3x * ws†(pf) ws(p) e+ipf.x' e-ipi.x' /(2π)3 = ∫d3p * ws†(pf) ws(p) ∫d3x e+ipf.x' e-ipi.x' /(2π)3 = ∫d3p * ws†(pf) ws(p) δ3(pf – pi) expos_time = (m/E) ws†(pf) ws(pf) = 1 where in the very last step we have used 3.11 for the spinor contraction! Thus, we have shown that, in this positive energy final state case, we have ψsf†'. SF e Ψ = ψsf†'.{ -iθ∫d3p Σr=1,2ψpr'pr} e Ψ = -iθ ps e Ψ and so we get our "form 5" S-matrix expression, Sfi = δfi + ψf†. SF e Ψ = δfi – i fs e Ψ form 5 where our final probing state is the spinor plane wave with s = 1 or 2. We could write I suppose fs = ((1/2π)3/2 e+ipf.x s(pf) s = 1 or 2 Now let's do the other situation where our final state was a negative energy state. Everything is the same, except we get a +i instead of a -i due to the signs in our original equation above, which I repeat here. ψf†'. SF e Ψ = ψf†'.{ -iθ∫d3p Σr=1,2ψpr' pr +iθb ∫d3p Σr=3,4ψpr' pr } e Ψ So here are the two results right next to each other: ( could use εf to write as one equation) Sfi = δfi – i fs e Ψ s=1,2 form 5 6.56 Sfi = δfi + i fs e Ψ s=3,4 form 5 6.56 where fs = ((1/2π)3/2 e+ipf.x s(pf) Now we can obtain our "final two forms" for Sfi. The first requires (0) which is : θ Ψ ' = ± iSF' γo. ψ where the ± depends on which way you are going in time. I think we can say θ Ψ ' = εi iSF' γo. ψ if this represents motion of our initial state, then we get Sfi = δfi – εs i fs e Ψ form 6 6.33a f5 + (0) = δfi – εs i fs e { εi iSF' γo. ψ} = δfi + εfεifs eSF' γo. ψi form 6 Now finally let's go for form 7 which starts again with form 5 [Ψ ' = ψ' + SF e Ψ is (3) ] Sfi = δfi – εf i fs e Ψ form 7 f5 + (3) twice = δfi – εf i fs e { ψ' + SF e [ψ' + SF e Ψ]} = δfi – εf i fs e ψi – εf i fs e SF e ψi + ... form 7 Now let's gather it all up: θ Ψ ' = ε iSF' γo. ψ (0) forward in time θ(t'-t) and ε = 1, else both opposite θ ψ' = ε iSF γo.ψ (1) 6.49 and 6.50 SF' = SF + SF e SF' (2) 6.51 Ψ ' = ψ' + SF e Ψ (3) 6.53 Sfi = ψf†. Ψi(+) form 1 defines Sfi Sfi = i ψf†.SF' γo. ψi form 2 f1 + (0) Sfi = δfi + i ψf†. SF e SF'γo. ψi form 3 f2 + (2) + (1) Sfi = δfi + ψf†. SF e Ψ form 4 f1 + (3) Sfi = δfi + εf i fs e Ψ form 5 6.56 special derivation Sfi = δfi + εfεifs eSF' γo. ψi form 6 f5 + (0) Sfi = δfi – εf i fs e ψi – εf i fs e SF e ψi + ... form 7 6.57 εf= 1 f5 + (3) twice Most of the above "forms" don't even appear in Chapter 6, mainly those involving SF', the full propagator. End of Very Long Note. Resuming original note flow after the following bar. ____________________________________________________________________________________ Now comes the clincher: by selecting your i and f states "properly", you find that you can be talking about four different processes! They are all "scattering" off the potential one or more times, the only question is how you interpret this scattering. Consider the following graph: This is order n = 2 because there are two vertices. The "initial" state (upper left) is an electron with negative energy (εi = -1) and we indicate its quantum numbers as –pi, –si let's say, So we put that in for ψi in 6.57. Since this ne state is going backwards in time, we label it ψi(-) as in page 97 equation A. We know that in hole theory, this is really a positron going forward in time with +pi and +si. We saw how the two 4-vectors p and s change sign as you do the CPT on ψ. I would therefore be inclined to indicate the ψi negative energy spinor by v(–p+, –s+) and I would then interpret this branch as a positron going forward in time with + p+ and + s+. But they use opposite signs in equation A, I don't know why. I do agree with the long red paren sentence which matches my CPT comment just made, so you would think they would want v(-p+, -s+), but that is not what they have written. In the above graph for the final state ψf we would put u(p-, s-) and there is less confusion. One possible explanation of our confusion regarding ψi is this: (1) perhaps the sign change in p+ is being compensated by the εr = -1 ; (2) it is true that v is defined with the spin reversed as shown in 3.16 on page 32. It does seem reasonable that they would set up the formalism so that v(p,s) would represent a positron with p and s. Note Added: Here is how I would treat the above diagram. I assume the arrow indicates an electron flow. So ψi is a neg energy state with εi = -1 and ψf is a positive energy state with εf = +1. From form 7 above I would say you compute this diagram as follows: Sfi = + εf i fr' e SF e ψri fr' = ((1/2π)3/2 e+i εf pf.x r(pf) r' = 2 or 2 εf= +1 ψir = ((1/2π)3/2 e-i εi pi.x wr(pi) r = 3 or 4 εi = -1 Note on the basic spinors. Suppose we did this: u(p,s) = B(p) R(θ,φ) w1(0) = exp(-ip K)exp(-iJ) w1(0) = exp(-ipK) exp(-isJ) w1(0) where we are using 4x4 matrices from the Dirac representation. Then of course u(p,s) is a function of the two items as four-vectors. Then we would have the following four results (notation of 3.16 page 32). u(p,s) = exp(-ipK) exp(-isJ) w1(0) u(p,-s) = exp(-ipK) exp(-isJ) w2(0) v(p,s) = exp(-ipK) exp(-isJ) w4(0) v(p,-s) = exp(-ipK) exp(-isJ) w3(0) I need to get this stuff cleaned up for the Dirac 4x4 representation, it goes on the list. Other books I notice just use the basic spinors above. Apart from this Feynman rule detail, we can view the above graph as a scattering graph rotated 90 degrees. In similar fashion, we have "pair annihilation" represented by this graph Where "i" is a normal u(p,s) incoming electron, and "f" is a negative energy electron v(p,s). The two other "rotations" give us "electron scattering" and "positron scattering" . Comments: (1) There are no photons anywhere, just . The EM field is not quantized at this point in our progress through both BD books. The vertex gets a classical Aμ potential, although it has been "installed into the theory" in a relativistically correct manner. HOWEVER: photons and photon propagators are going to suddenly appear in the upcoming chapter 7 so we can compute all the basic lowest order processes. (2) It seems that the propagators can carry information at an arbitrary velocity, even greater than c. I say this because the theta functions are of the form θ(t-t') with no account made for flight time. (3) A propagator Green's function can be represented as a sum over all paths of eiS where S is the action. This whole concept, also invented by Feynman it seems (1948), has not been mentioned by BD up to this point. I think this eiS method is "the modern method" of doing field theory now, but I am reading a 1965 book. Goldstein (1950) mentioned the action S as the time integral of L, but he never got it up into an exponent. At some point I will want to get into this subject, but don't want to go off on yet another tangent right now.