Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Physics / Mechanics / Goldstein Related

Goldstein mechanics

DOCX · 1.3 MB
Open DOCX file

Reading notes written in Phil's own voice on the 1950 first edition of Herbert Goldstein's Classical Mechanics, organized by the book's 11 chapters and sections with book page numbers. They begin with comments on the author, later editions and errata, then summarize and question the material: Lagrange's equations, variational principles, central forces, rigid bodies, relativity, canonical transformations, Hamilton-Jacobi theory, small oscillations and fields.

AI-written summary; may contain errors.

Extracted text (machine-read; may contain errors)
Classical Mechanics Herbert Goldstein 1950 Table of contents: (numbers in parens indicate book page numbers) Comments on the author and book. 3 Chapter 1: Survey of the Basics (1) (mainly, the Lagrangian formulation of mechanics) 4 1.1 Mechanics of a particle (1). 4 1.2 Mechanics of a System of Particles (4) 5 1.3 Constraints (10). 5 1.4 D'Alembert's principle and Lagrange's equations (14) 6 1.5 Velocity-dependent potentials and the dissipative function (18). 17 1.6 Simple applications of the Lagrangian Formulation (22) 18 Chapter 2: Variational Principles and Lagrange's Equations (30) 20 2.1 Hamilton's Principle (30). 20 2.2 Calculus of Variations (31) 20 2.3 Derivation of Lagrange's Equations from Hamilton's Principle (36) 21 2.4 Extension of Hamilton's Principle to more general systems. (38) 22 2.5 Advantages of a variational principle formulation (44). 24 2.6 Conservation theorems and symmetry properties (47) 25 Chapter 3: The 2-body problem with a central force (58) 27 3.1 Reduction to the equivalent 1-body problem (58) 27 3.2 Central Force Problems (59). 27 3.3 The equivalent one-dim problem, and the classification of orbits (63) 28 3.4 The virial theorem (69). 29 3.5 Differential equation for the orbit (71) 30 3.6 The Kepler problem: inverse square law force. (76) 30 3.7 Central Force Scattering (81) 34 3.8 The Lab Frame (85). 35 Chapter 4: Rigid Body Kinematics (93) 35 4.1 The independent coordinates of a rigid body (93) 35 4.2 Real orthogonal transformations (97). 36 4.3 Formal properties of the transformation matrix (101). 36 4.4 The Euler angles (107). 36 4.5 The Cayley-Klein parameters (109). 36 4.6 Euler's Theorem on the motion of a rigid body (118). 36 4.7 Infinitesimal rotations (124). 37 4.8 Rate of change of a vector (132). 38 Clarification of Notational Ambiguity. 38 4.9 The Coriolis Force (135) 40 Chapter 5: Rigid Body Equations of Motion (143) 43 5.1 Angular Momentum and Kinetic Energy of motion about a point (143). 43 5.2 Tensors, Dyads and Dyadics (146) 44 5.3 The inertial tensor and the moment of inertia (149) 46 5.4 The eigenvalues of the inertia tensor and the principle axis transformation (151). 46 5.5. Methods of solving rigid body problems and the Euler equations of motion (156). 47 5.6 Force-free motion of a rigid body (159). 48 5.7 The heavy symmetrical top with one point fixed (164). 52 The Heavy Top Started with a Drop. 54 Heavy top started vertically (p 173) 55 Some Applications (p174). 56 5.8 Precession of charged bodies in a magnetic field (176) 56 The Classical Motion of a Rigid Charged Body in a Magnetic Field 56 Chapter 6. Special Relativity in Classical Mechanics (p 185) 57 6.1 The basic program(185). 57 6.2 The Lorentz Transformation (187). 57 6.3 The covariance of equations (194). 57 6.4 What happens to F = ma ? (199) 58 6.5 Relativistic Lagrangian situation ( 205) 58 6.6 Covariant Lagrangian situation (207) 59 Chapter 7. The Hamilton Equations of Motion ( p 215) 59 7.1 Legendre Transformations and the Hamilton Equations of Motion (215) 59 7.2 Cyclic coordinates and Routh's procedure (218) 60 7.3 Conservation Theorems and Physical Significance of H (220) 61 7.4 Hamilton's Equations from a variational principle (225). 61 7.5 Principle of Least Action (228) 61 Comment on the word "action". 66 Chapter 8. Canonical Transformations ( p 237 ) 67 8.1 The Equations of Canonical Transformation (237). 67 8.2. Examples of canonical transformations (244) 68 8.3 The integral invariants of Poincare (247). 68 8.4 Lagrange and Poisson Brackets as canonical invariants (250). 69 8.5 Equations of motion in Poisson bracket notation (255). 71 8.6 Infinitesimal contact transformations, constants of the motion, symmetry (258) 72 8.7 The angular momentum Poisson bracket relations (263). 72 8.8 Liouville's Theorem (206). 73 Chapter 9. Hamilton-Jacobi Theory ( p 273) 73 9.1 The Hamilton-Jacobi Identity for the Hamilton's Principle Function (273) 73 9.2 A HJE example: the harmonic oscillator (277). 74 9.3 Hamilton's Characteristic Function W (279) 74 Review of the Table: 76 9.4 Separation of Variables in the HJE (284) 77 9.5 Action-angle variables (288) 77 9.6 Further properties of action-angle variables (294) 79 9.7 The Kepler problem in action-angle variables (299) 82 Part I: The basics 82 Part II: Computing the Ji action variables 85 9.8 Hamilton-Jacobi theory, geometric optics, and wave mechanics (307) 88 History Added: 98 Chapter 10. Small Oscillations ( p 318) 99 10.1 Formulation of the Problem (318). 100 10.2 The eigenvalue equation and the principle axis transformation( 321) 100 10.3 Frequencies of free vibration and normal coordinates (329). 105 Comparison with the inertia tensor situation. 107 10.4 Free vibrations of a linear triatomic molecule (333). 107 10.5 Forced vibrations and dissipative forces (338). 110 Chapter 11. Continuous Systems and Fields ( p 347)∂ 114 11.1. Transition from discrete to continuous system (347). 114 11.2 The Lagrangian formulation for continuous systems (350). 114 11.3 Sound vibrations in a gas (355). 115 11.4 The Hamiltonian Formulation (355) 117 11.5 Descriptions of Fields (364). 118 Comments on the author and book. This fellow died in 2005, basically wrote this book post war while he had a temp job as a Harvard instructor. His main gig was at Columbia where he specialized in nuclear radiation shielding. A second edition of this book appeared in 1980, but I am reading my original copy. A third edition appeared in 2001 with coauthors Poole and Safko. This would be used in a current day course I suppose. Here is the contents of that book: This sight maintains errata for the 3rd edition: http://astro.physics.sc.edu/goldstein/ You can see how this compares to my 1950 edition which has 11 chapters. My chapter 10 on oscillations appears earlier as chapter 6 in the new book. Chapters on chaos and perturbation theory have been added. I suspect much of the book is unchanged! All 11 chapters have the exact same names as in the 1950 book. Note added 10.31.16. I found errata on 2nd and 3rd edition but not first. For 3rd it is a website supposedly maintained by John L. Safko at South Carolina. http://astro.physics.sc.edu/goldstein/ . My PDF has no copyright page, so I don't know its printing number. But it seems to have corrections for 1,2,3 printings, and it seems not to have corretions for 4 and 5; for example in my edition the item for page 13 is not corrected. Thus it seems my GPS copy is either printing 4 or 5. My PDF date is Aug 2006. There are many errata for all the editions! I don't know how to find printing dates online. There is now a 2013 "international edition" with a different cover, from Pearson. My integral error page page 93 does not appear in any of these errata lists. It turns out that all three authors are dead, and all in the not-too-distant past: Herbert Goldstein: Died 1.12.05 http://www.columbia.edu/cu/news/05/02/herbertGoldstein.html Charles P. Poole Jr: Died 11.1.15 at age 88, http://obits.dignitymemorial.com/dignity-memorial/obituary.aspx?n=Charles-Poole&lc=9817&pid=176301825&uuid=0d732362-339a-469c-aae6-65be8d0d08e5 John L. Safko: Died 6.15.16 at age 77, did 38 years at South Carolina, a good teacher it says http://www.legacy.com/obituaries/thestate/obituary.aspx?n=john-safko&pid=180380299&fhid=32171 Preface: Various arguments are given why a solid basis in classical mechanics is good for people about to study quantum mechanics and field theory. He throws in special relativity and velocity-dependent potentials. The preface ends with a Hebrew word that I don't know how to translate or even type. I tried some Yiddish dictionaries, no luck. But then I found it on the web with a fancy search: "I had first encountered this insider's code in a secular environment many years ago when reading herbert goldstein's majesterial text on classical mechanics (published 1950, addison wesley) as he closed his preface with the mysterious hebrew lettering." The word is really an abbreviation, not a word, and you do the letters right to left. So if I ignore the thing that looks like a ", I can get these 6 letters: O - B L Sh N T which you can reverse to say T N Sh L B O which roughly agrees with the following phrase: TVSLBO Tom V'nishlom Shevoch L'Eil Boreih Olom (Hebrew acronym to mark the conclusion of a published work) -- means: "finished and complete is (this?) praise to god, creator of the world" In other places I see it written as " Tam V’nishlam Shevach L-eil Boreih Olam.” Web: There is a touching tribute to the individual that can be derived from the traditional custom of writing in conclusion of a Jewish book, the Hebrew phrase “Tam V’nishlam Shevach L’El Borei Olam” - “The book is complete, praise be to God” The 3rd edition has a similar but different ending to the Preface, this time in Greek instead of Hebrew, which is this from Psalm 19,1: ********************************************************************************* Chapter 1: Survey of the Basics (1) (mainly, the Lagrangian formulation of mechanics) ********************************************************************************* My Lagrange doc paper constains a lot of this stuff. 1.1 Mechanics of a particle (1). We assume F = dp/dt and just work from there. p is conserved if no force. L is defined and shown to be conserved if no torque N. Work as Fdx is integrated to show that work causes a change in kinetic energy T. A conservative force is one whose path integral is zero, which in turn means it can be represented as the gradient of a potential V. We soon end up with T+V being conserved. He has not mentioned that many of his statements are only true non-relativistically. Question: If the force in a problem can be written as the gradient of a potential V which is velocity-dependent, is that still a conservative force? Gold seems to avoid asking this question. The work done in moving in a closed loop must be 0. Obviously if V = kx|vx| so that F = k|vx| describes a friction situation, we know that work is done against friction so positive work is done going around the loop, so such a force cannot be conservative (which really means non-heat energy is conserved I think). So here at least is an example of a "velocity dependent potential" which is non-conservative. On the other hand, EM theory gives an effective potential U which I think IS conservative. Probably this is because the Lorentz force v x B is always perpendicular to v so no work gets done. So if a force can be written as the gradient of a position-only-dependent potential, the force is conservative. If the potential is velocity dependent, then the force may or may not be conservative. 1.2 Mechanics of a System of Particles (4). Assumption here is that Fij = - Fji and lies along the line between the particles i and j (either polarity). With this assumption, many conclusions are reached. The center of mass R is defined and it's momentum is conserved if no external force. Total angular momentum is conserved if no external torque. On page 7, L is written as the sum Lcms + Linternal relative to the CMS. The only dependence on the origin location here is through R, the CMS location. On page 9 he has a potential Vij(rij) which gives rise to the internal forces (between particles i and j), and represents the internal potential energy. The total potential is then given top of page 10 where Vi is from external forces, and the 1/2 just from double counting. We conclude that T+V is again conserved for the system as a whole, no great surprise. Beware E&M fields where forces don't meet the assumptions here, and cause violations since they carry momentum and angular momentum and energy. '' Comments: In this fairly long 6 page section, forces on a particle are partitioned into "internal" and "external" forces. Internal means forces of particles acting on each other. This is different from a later partition of the total force on a particle as "constraint force" and "other forces; constraints are not mentioned in this section. It is only if you make the assumption Fij = - Fji for the internal forces that you can show the famous "rigid body" results quoted in my paragraph above such as the existence of a center of mass vector R and the fact that M = ΣFext and the partitioning of angular momentum and energy into CMS and local components, and so on. So Fij = - Fji is a very major assumption used in this section 1.2, 1.3 Constraints (10). To dispel the notion that you just "turn the crank", this little topic is brought up. Sometimes constraints can be written as equations like top of page 11, in which case they are holonomic. [ I think holo- means whole, suggesting there are no differentials appearing in the constraints. ] Otherwise they are non-holonomic which means they are harder to deal with. Example of the latter is any inequality. Holonomic really means you have a constraint equation which allows you to remove a variable from the problem (an "integrable constraint"). Then on page 12 if we have k constraints, we can come up with a number of "generalized coordinates" qi which equals the true number of degrees of freedom in your problem (here 3N-k), then the 3N true coordinates ri which are not independent can be written in terms of the generalized qi ones (note that in general the qi won't come in vectors like the ri do, often the qi are angles). For example, a bead constrained to a fixed 1D string in 3space, would nominally have 3 degrees of freedom x,y and z, but in reality has only one degree of freedom q1 = distance along the string. Rolling without slipping is nonholonomic because you cannot express it as an equation on the ri. It is a velocity condition of the point of contact (thus involving dri). He considers a rolling vertical disk on a flat surface, comes up with nice coordinates. He can write down the constraints as differentials, but this does not count as holonomic. A good problem, I commented long ago. These problems require "custom solutions". A main point here: If you start with 3N coordinates for a system of N particles, and if you have k holonomic constraints, you can then define a set of 3N-k generalized coordinates qi which are independent variables! This independence is crucial for doing Lagrangian mechanics as we shall see. If your constraints are non-holonomic, you have to modify Lagrangian mechanics with those Lagrange multipliers (as we shall also see). Comment: Nowhere in this 4 page section is there an explicit partitioning of forces acting on particle k as a sum of constraint forces plus other forces. No symbol is used for any force really. Comment: Gold uses the symbol N as the number of particles in a system. 1.4 D'Alembert's principle and Lagrange's equations (14). This section now has 11 pages of raw notes below mainly because I just did not understand the approach which is centered on D'Alembert's Principle. Now I think it all makes sense, but it took a long time. [ 1.10.12 while reviewing M&F was led to review Goldstein a bit. ] [ By 9.29.16 I have written Lagrange App D which helps on a lot of this stuff below, and have read the R&S paper 2006 on δr .] Here we have something new to me: the notion of a virtual displacement ri of the coordinates of a system of particles each having a label i. In my first reading, I was totally confused, but then in Jan 2012 I read a little Indian PDF on the subject and will rewrite my notes here based on what I learned there. That PDF discusses both an "allowed displacement dri" and a "virtual displacement δri". In either case, you think of a SET of displacements, one for each particle in the system, so write as { δri } as a reminder of this fact. As an example, consider a 4 particle rigid-body system: The set of four arrows on the left shows an example of a legal "allowed displacement" set of 4 arrows {dri}. The four particles can move differentially as shown because this comprises an allowed rotation of the rigid body about some origin. The only allowed system displacement sets {dri} in this 2D world would be sets which correspond to either a rotation about some origin, or to a translation (or a combination). The picture on the right shows a non-allowed set {dri} where three of the four arrows are zero displacement. This set would require deformation of the "sticks" holding the 4 particles together and this violates the constraints. The Indian PDF clarifies the "time" aspect of things by saying that a "virtual displacement" is just the difference of two "allowed displacements" (which difference is not allowed to be zero). Their idea is this: from the chain rule acting on the constraint set fi(rj , t) = 0 ( k equations for G, s for the Indians) we have this fact, which I write in four different notations Σj=1N Σa=13 ∂fi/∂(rj)a ∂(rj)a /∂t + ∂fi/∂t = 0 // N = number of particles Σj=1N Σa=13 ∂fi/∂(rj)a (vj)a + ∂fi/∂t = 0 // a = index like x,y,z Σj=1N (j)fi vj + ∂fi/∂t = 0 // i = constraint number i Σj=1N ∂fi/∂rj vj + ∂fi/∂t = 0 // this is the Indian notation For an "allowed" displacement set, one has drj = vjdt in the {..} sense described above for the system. Thus the last equation above can be written Σj=1N ∂fi/∂rj vjdt + (∂fi/∂t) dt = 0 Σj=1N ∂fi/∂rj drj + (∂fi/∂t) dt = 0 By defining a δrj virtual displacement as any difference of allowed drj , the Indians conclude that Σj=1N (∂fi/∂rj) δrj = 0 // time cancels in difference i = 1...k and it is in this "sense" that δrj does not involve a dt movement. Here we can interpret (∂fi/∂rj) as some kind of "force" and then the above is some kind of "principle of zero virtual work". We know that we cannot set the 3n scalar members of the set {δri} to arbitrary values due to the constraints -- we can only have sets {δri} which are "allowed" by the constraints. In other words, the 3n variations δri are not "independent". Maybe later I can tie the above in with Gold's discussion. Fact: The above equation says that the δri are linearly dependent! That is the key point. So now let's move back to the Goldstein discussion on page 14. The plan here is to start with a system in static equilibrium, and then later make it be dynamic. We start this long section with (1-38). Why is this true in equilibrium? It says iWi = 0. If the total force Fi on particle i is zero, then sure, Fi = 0 and so we cannot deny that Fi ri = 0 and then sum to get (1-38). This equation seems to say nothing interesting. He assumes equilibrium here (clearly stated). Again, the individual terms in this sum are 0 because the total force on each particle is 0 because we are at rest. In (1-39) we recognize that some forces actually come from the constraints, called fi . Then Fai refers to the "applied force", meaning any non-constraint force. This separation seems reasonable, and then (1-40) is a restatement of (1-38). Question: For the first time now he has partitioned forces into constraint forces and "applied" forces. If there are internal inter-particle forces Fij, into which category do they go? He uses the same Fi(a) symbol that he used in the previous partition of internal + external forces. Gold does not answer this question. R&S refer to the Fi(a) portion as "external forces". How could Gold ignore this question? A simple example would be a set of N planets moving around. Are the gravitational attractions "applied" or are they "constraint" ? Let's look right now at some other sources: Gold 3rd edition (w Poole, Safko). I just found this as PDF for the first time. The text in this area is unchanged, so nothing said about the partition. But good now to have this book for the first time. I can give refs to multiple editions now. Greenwood Classical Dynamics (in Dover): I cannot find this as PDF, but good search views are provided in google books https://books.google.com/books?id=x7rj83I98yMC&q=alembert#v=snippet&q=alembert&f=false Section 1.4 is on virtual work. Refers to some constraints as "workless constraints". Uses same Ri and Fi used by R&S. Same work=0 for static. In the partition Fi is "applied" force. Good text, lots of examples, but no mention of "internal forces". L&L: Have a djvu, but no good hits on "internal forces" Web search: I found someone asking this same question, but not useful. Forces PDF: Found and saved a pdf that classifies forces into external + internal + constraints. So at least someone is talking about this. But too elementary to mention virtual work. wiki on virtual work: States that the internal forces of a rigid body or those of an ideal joint are classified as constraint forces, mainly because there is no virtual work for them! But no d'Alembert so cannot see my answer directly. Searching now on: constraint forces "internal forces" "virtual work" my answer: If you by fiat make the division R + F as being "constraint plus other", then the division depends on what you define to be the constraint forces. For a rigid body of masses and sticks, you want to classify the internal forces as constraint forces because these constraints are associated with zero work. but for a semi-rigid body with springs instead of sticks between the masses (planets eg), I think you would classify the internal force on a particle (force due to all the other particles) as an "other" force and not as a constraint force, because if you classify as constraint, they you have contraints which do work. In the limit that the springs become infinitely stiff, you would reclassify the internal forces as constraint forces. So I guess I am happy now with this question and its answer. I have added my answer to Lagrange Appendix D. Two Examples: The left example is a particle on an inclined plane with a string tension T holding things at rest. Here, the total force F1 acting on particle 1 is zero, as in 1-38. We write this as F1 = F1(a) + f1 . The constraint force f1 is normal to the surface as shown. Then F1(a) is the sum of the gravity down force and the string tension force, so F1(a) is then equal and opposite to f1 . The allowed or virtual displacement would be along the inclined plane, so one would have f1 r1 = 0 due to perpendicularity. Then 1-41 says that F1(a) r1 = 0 and we see this is true also by perpendicularity. Now let's switch to the example shown on the right of a 4-particle rigid object at rest on an inclined plane still held by a string. In this case, f1 is the sum of the two green arrows forces exerted by the sticks, and F1(a) is the gravity down force all by itself. The allowed δri set for our object would be four arrows down the plane (not shown). So notice right off the bat that now f1 r1 ≠ 0 and also F1(a) r1 ≠ 0 for this individual particle! So suddenly all of G's discussion "goes hazy". How do we know that this is true for our 4-particle example: Σi=14 fi ri = 0 ? I have drawn in all the arrows for particle 2 where things are radically different in terms of force directions. But here is the key point. The two green arrows along the top stick are equal and opposite "internal" forces. Thus we will have f2(top) r2 + f1(top) r1 = 0 and we can regard these two terms as 2 of 8 terms in the sum above. Top means due only to the top stick. A similar pair would be for the bottom stick. The side sticks all have things like f1(side) r1 = 0 due to perpendicularity and this accounts for the other four terms. Thus, the above sum will in fact be valid! Gold just does not give the detail that is needed here, so I will now provide it. What is really true for our example is this Σi=14 fi = 0 where fi is the total "internal force" on particle i, formerly written as F(i). Write [ (i) means internal ] fi = Σj≠i F(i)ij where the total internal force on i comes from all the other particles in the rigid object. Then consider Σi=1N fi = Σi=1N Σj≠i F(i)ij = Σpairs, j≠i F(i)ij This is the sum over all pairs ij where i≠j, so there are (N,2) terms in the sum. now write = Σpairs, i<j F(i)ij + Σpairs, i>j F(i)ij = Σpairs, i<j [ Fij + Fji] Now Gold's assumption p 4 is that these internal forces are equal and opposite within each pair, as would be the case if we had "sticks" connecting each pair (or gravity or electrostatic force etc). Then each term in the above sum over pairs i<j is 0, and thus we have proven that Σi=14 fi = 0 for "this kind of system". Then since the ri virtual displacement set of arrows consists of equal arrows pointing down the plane, we then arrive at Σi=1n fi ri = 0 which is the right part of 1-40. In conclusion, I agree that Σi=1n fi ri = 0 as a sum, whereas the individual terms labelled by i are in fact not zero. We saw this in our second example above. So here we are Σi=1N F(a)i ri = 0 (1-41) The truth of this equation is not trivially obvious, since the individual terms are not 0. For each particle, the symbol F(a)i means all forces applied to that particle OTHER than those of the contraints. In our example 2, particle 1 had only gravity, whereas particle 2 had gravity + string tension. Since the above is not trivially true but is only true for systems in which Σi=1n fi ri = 0, and since the terms in the sum have dimensions of work Fdr , and since the set {ri} must be an allowed "virtual" displacement (not an arbitrary displacement set for the system}, this result (1-41) earns itself a formal name which is "the principle of virtual work". One comment about (1-41) is that we seem to have simplified the problem in some sense by removing constraint forces from the equation. This is static equilibrium only !! The history is unclear here but http://books.google.com/books?id=jmKxffjQtasC suggests that maybe the Flemish statics guy Stevin (1586, 1605) came up with this idea. Newton Principia was 1687. Having read my notes above, I have to say that Gold is just totally unclear in this section p 14-15. No meaning for the symbol δri is given, no definition. I agree with the Indians, Goldstein has just dropped the ball here. Going back to the right-side figure above, forget the string, and the thing can slide on the frictionless surface. Any allowed displacement is 4 downhill arrows, so any Indian virtual displacement would be the same, so call each δr! Then if fi are the four constraint forces, you get Σi=14 fi ri = (Σi=14 fi) r = 0 r = 0 // as noted in detail above where the individual terms are NOT zero but the sum is zero. So this is an example where the inclined plane is not moving, but the individual terms are not zero. So App D needs more work. // I have now upgraded App D with an example of this type. Continue now in Goldstein on page 15: Up to this point, Gold has been talking virtual work = 0 in statics situation. Now comes dynamics: So, now we are going to just start over and replace Fi by (Fi- dpi/dt). Then (1-38) becomes (1-41'), true because Fi- dpi/dt = 0 by Newton's Law. If we break out the constraint forces, we get A at top of page 16. And again if we restrict to our special systems as above in which ifi ri = 0, we get (1-42). Notice that the quantity (Fia- dpi/dt) might not be zero, but the dot product sum is zero. Again, this gets the fancy name of "D'Alembert's Principle": Σi=1n ( F(a)i - i) ri = 0 (1-42) This principle is again non-trivial, just as was its static version. Very importantly, each term in the sum is not separately zero. We do not have F(a)i = miai because that is true is Fi = miai and if you exclude the contraint forces the result is in general not true -- F = ma is true if you include ALL the forces. Nevertheless, the sum above is true. Jean le Rond d'Alembert did this in 1743 which you see is 50 years after Principia. On page 16 after stating (1-42) Gold simplifies his notation so that after this Fi means Fi(a). It is true that in (1-42) the constraint forces fi do not appear at all. Independent "generalized" coordinates (p 16 middle). He now imagines that n is the number of "independent" coordinates so that p 16 B is true: ri = ri(qj, t). He differentiates this to get (1-43) which is just fine. ri = ri(qj, t) dri/dt = vi = Σj=1n (∂ri/∂qj) (∂qj/∂t) + ∂ri/∂t (1-43) dri = Σj=1n (∂ri/∂qj) (∂qj/∂t)dt + (∂ri/∂t)dt dri = Σj=1n (∂ri/∂qj) jdt + (∂ri/∂t)dt dr'i = Σj=1n (∂ri/∂qj) 'jdt + (∂ri/∂t)dt Here my idea is that these are two different allowed displacements and they differ because the generalized velocity sets are set differently, one with prime the other not. Things like (∂ri/∂qj) and (∂ri/∂t) are more or less constants, determined by ri = ri(qj, t), these cannot be different for different displacements. Now subtract to get δri = Σj=1n (∂ri/∂qj) [ 'j - j ]dt Then write [ 'j - j ]dt = ∂t [ q'j- qj] dt = q'j- qj = δqj Then we arrive at δri = Σj=1n (∂ri/∂qj) δqj (1-44) I think this is how the Indians would derive (1.44) Note added 9.29. Could write the above as dri = Σj=1n (∂ri/∂qj) dqj + (∂ri/∂t)dt dr'i = Σj=1n (∂ri/∂qj)dq'j + (∂ri/∂t)dt Then δri = dri - dr'i = Σj=1n (∂ri/∂qj) [dqj - dq'j] = Σj=1n (∂ri/∂qj) δqj This is really a trivial point. The notion of δqj = dqj - dq'j is exactly the same as the R&S notion. Both the virtual δri displacements and the virtual δqi displacements must be taken as differences of the corresponding allowed displacements. Gold comments in parens that this then supports time-dependent constraints, and the Indians would agree. So (1-44) is a very important ingredient and Gold sort of fudges it, but I now accept it. Then (1-45) is a no-brainer and he has then defined his generalized force Qj. So Gold now wants to process D'Alembert 1-42 into something involving only these qj coordinates. He comments that if we could cast this into the form jGjqj = 0 where qj are virtual displacements in the generalized coordinates, then since these are independent, we could conclude that this result tells us that Gj = 0. If we skip the algebra for the moment and cut to the chase, we see that we can in fact make such a formula and it is shown in (1-49) and we have Gj ≡ [ ... ] as shown there. This object Gj is seen to involve various derivatives of kinetic energy T, and the generalized force Qj that corresponds to the generalized coordinate qj. Accepting this for the moment, we then have (1-50) on page 18 top. The operator acting on T is the same as that we know appears in "Lagrange's equation", which will appear in a moment. In some sense, (1-50) says that T is driven by the generalized force Qj. At this point, we have not assumed that forces come from a potential V, or that such a V even exists. Now we go an extra step and assume that our regular force Fi (meaning the applied force) is conservative and so can be obtained as Fi = -iV(..ri...). Again skipping algebra steps, we then arrive at the famous Lagrange Equation of motion (1-53) where we have defined a new interesting object known as "The Lagrangian L" and L = T - V, not to be confused with the Hamiltonian which is H = T + V = total energy. The L thing is some weird energy difference thing. The derivatives in (1-53) are wrt to generalized coordinate qj and the generalized velocity , as well as wrt time. d/dt(L/) L/qj = 0 where each term has dimensions of energy over dim(qj). The generalized forces Qj no longer appear! We get one equation for each of our independent qj coordinates. Assumes conservative forces (which means those derivable from a potential energy gradient. This V in turn could still have velocities in it.) Comment: If you know the transform ri = ri(qj, t) and you know the real forces Fi. you can eaily compute the generalized forces Qi using (1-46), so the Qi are totally well-defined. Now let's go back and look at the details of the algebra. Note: I have done this I think in a simpler and clearer way in my Appendix D, but I leave these details here ************************************* start algebra ***************************** Start with p 16 B and apply d/dt: (dri/dt) = (ri/qj) (dqj/dt) + (ri/t) or = (ri/qj) + (ri/t) (1-43) Next, assume (1-44) based on discussion earlier so that δri = Σj=1n (∂ri/∂qj) δqj (1-44) Now the first term in D'Alambert is this, where we can take Fi as either total or applies (noted above) Σi=1N Fi ri N = number of particles in system Now insert 1-44 to get Σi=1N Fi ri = Σi=1N Fi [Σj=1n (∂ri/∂qj) δqj] = Σj=1n [Σi=1N Fi (∂ri/∂qj)] δqj = Σj=1n [Qj] δqj = 0 (1-45) Qj ≡ Σi=1N Fi (∂ri/∂qj) (1-46) The second term in D'Alembert is (ignoring minus sign for the moment) Σi=1N i ri = Σi=1N mi i ri (p 17 A) = Σi=1N mi i [Σj=1n (∂ri/∂qj) δqj] = Σi=1N Σj=1n mi i (∂ri/∂qj) δqj (p 17 B) Now write d/dt[ mi i (∂ri/∂qj)] = mi i (∂ri/∂qj) + mi i d/dt (∂ri/∂qj) or mi i (∂ri/∂qj) = d/dt[ mi i (∂ri/∂qj)] – mi i d/dt (∂ri/∂qj) Then sum this on i to get Σi mi i (∂ri/∂qj) = Σi { d/dt[ mi i (∂ri/∂qj)] – mi i d/dt (∂ri/∂qj) } (1-47) Now take (1.43) and apply ∂/∂qk to both sides to get ∂/∂qk(dri/dt) = ∂/∂qk Σj[(ri/qj) (dqj/dt)] + ∂/∂qk (ri/t) Now one does not think of (dqj/dt) as being a function of qj so pull that from the bracket ∂/∂qk(dri/dt) = Σj ∂/∂qk [(ri/qj)] (dqj/dt) + ∂/∂qk (ri/t) Or just think ∂/∂qk (dqj/dt) = d/dt (dqj /∂qk) = d/dt (δij) = 0 and then claim is justified. Now swap k↔ j to get ∂/∂qj(dri/dt) = Σk ∂/∂qj [(ri/qk)] (dqk/dt) + ∂/∂qj (ri/t) Interchange on left, rewrite on right d/dt(∂ri/∂qj) = Σk (∂2ri/∂qjqk) k + 2ri/∂qjt p 17 C or (∂vi/∂qj) = Σk (∂2ri/∂qjqk) k + 2ri/∂qjt where the LHS's are trivially the same: d/dt(∂ri/∂qj) = (∂vi/∂qj) (*) Now start again with (1-43) and apply ∂/∂k to both sides to get (dri/dt) = Σj(ri/qj) (dqj/dt) + (ri/t) ∂/∂k (dri/dt) = ∂/∂k Σj (ri/qj) (dqj/dt) + ∂/∂k (ri/t) Now ri is not a function of k (only of qk) and so we get ∂/∂k (dri/dt) = Σj (ri/qj) δkj = (ri/qk) ∂vi/∂k = ri/qk and change k to j ∂vi/∂j = ri/qj (1-48) Now I will add my own line. Rewrite Now start with 1-47 again Σi mi i (∂ri/∂qj) = Σi { d/dt[ mi i (∂ri/∂qj)] – mi i d/dt (∂ri/∂qj) } (1-47) 1-48 (*) Replace as shown to get Σi mi i (∂ri/∂qj) = Σi { d/dt[ mivi ∂vi/∂k] – mi vi (∂vi/∂qj) } p 17 E Now look at {...} in p 17F. It can be written as {...}17F = d/dt [ m vi ∂vi/∂k] - mi vi (∂vi/∂qj) = {...}17E Therefore we have from p 17E Σi mi i (∂ri/∂qj) = Σi {...}17F Σi i (∂ri/∂qj) = Σi {...}17F Σi i Σj(∂ri/∂qj)δqj = Σj Σi {...}17F δqj Σi i δri = Σj Σi {...}17F δqj The LHS is the desired piece of (1-42) , and this RHS is Σj Σi {...}17F δqj = Σj [d/dt (∂T/∂j) - (∂T/∂qj)] } δqj which is to say Σi i δri = Σj [d/dt (∂T/∂j) - (∂T/∂qj)] } δqj (**) Now recall from above that the first term in (1-42) was Σi=1N Fi ri = Σj=1n [Qj] δqj ΣiFi ri = Σj Qj δqj Therefore we can write (1-42) this way Σj Qj δqj – Σj [d/dt (∂T/∂j) - (∂T/∂qj)] } δqj = 0 Σj { Qj –[d/dt (∂T/∂j) - (∂T/∂qj)] } δqj = 0 Σj { [d/dt (∂T/∂j) - (∂T/∂qj)] – Qj } δqj = 0 (1-49) Now recall that we could write p 16B because we could eliminate some of our ri type variables using our holonomic constraints, thus givein us the n independent variables qj and this from above we can say [d/dt (∂T/∂j) - (∂T/∂qj)] = Qj T = Σi ( 1/2 mivi2) (1-50) "Lagrange's Equations" ************************** end of long algebra section ***************************** The above equation involves the generalized coordinates qj and generalized velocities j and generalized forces Qj as determined by 1-46. The equation is true in the presence of our constraints! If we had no constraints and qi ~ ri (vector issue of course), then the first term on the left is just mii and Q = F and we just have ma = F. But the whole point is that 1-50 is valid with constraints where the qi are some wierd but independent coordinates and T may somehow depend on the qi as well as the i. Now assume that the real forces Fi = -iV(ri), for example F2 = - r2 V(r1, r2, r3, ....) Then we know from 1-46 that Qj ≡ ΣiFi (∂ri/∂qj) = Σi -iV (∂ri/∂qj) = – Σi (∂V/∂ri) (∂ri/∂qj) But the chain rule then says ( this assumes that V = V(ri; but not t and not vi) ) Qj = – ∂V/∂qj (1-51) Comment: Conservative force means all closed line integral paths give 0. If force comes from potential V(r1, r2, r3, ....) then this is true. If there is velocity dependence in V, then whether or not the force is conservative probably depends on the details. The force is likely NOT to be conservative. This would be the case if there are any frictional forces. But I could imagine that somehow the velocity dependence could be such that the force somehow remains conservative. A quick web scan did not clarify this point. So our Lagrange equations (1-51) are now [d/dt (∂T/∂j) - (∂T/∂qj)] = Qj = – ∂V/∂qj [d/dt (∂T/∂j) - (∂T/∂qj) +∂V/∂qj ] = 0 [d/dt (∂T/∂j) - (∂[T-V]/∂qj)] = 0 [d/dt (∂[T-V]/∂j) - (∂[T-V]/∂qj)] = 0 [d/dt (∂L/∂j) - (∂L/∂qj)] = 0 (1-53) and these are what most people call "Lagrange's equations", written in terms of the Lagrangian L. So we have finally battled our way through section 1-4 ! Assumptions: holonomic constraints, and contraint forces respect Σi fi δri = 0, and in this form we must have all forces from a conservative potential V. You can be sure that our very dense Section 1.4 here is the subject of much of other mechanics books! Lagrangian dynamics. OK 1.8.12 OK 9.29.16 1.5 Velocity-dependent potentials and the dissipative function (18). Two separate topics are treated here: (1) particle in EM field, (2) particle with linear friction. Both involve "velocity dependent forces" and in fact both cases have forces linear in velocity. These situations cannot be handled by "velocity independent potentials" V because, were that so, we would compute F = -V and then F would be velocity independent. We have to treat the first situation because this case will apply to very standard problems involving EM fields acting on charged particles, which comprise the subject of a huge portion of physics! Let's ignore the opening comments for the moment and look at the E&M algebra. Everything is quoted from Jackson, I will say. We define A in (1-57) and define in equation above (1-58). Then (1-58) gives the E field as a function of the potentials A and , complementing (1-57) which gives the B field in terms of just A. The Lorentz Force can be written in terms of potentials as (1-59) which we got from the quoted result (1-56). Goldstein is not deriving all these E&M results, he is just quoting many and then stating equations he is interested in. We see in any event that the Lorentz force has a term which is a linear function of velocity in (1-59). To avoid using a fancy vector identity, Goldstein does lots of details by hand, all of which I verified, and he is then able to write the Lorentz force as in equation on page 21. Then if he defines U as shown in (1-60), he can write the Lorentz force as in equation . Now let's pause on this as just a pure algebra result, and let's go back to the section's opening comments. The comments say: Suppose you could express the generalized force Qj as the Lagrange-like set of derivatives of some "generalized potential" U as shown in (1-54). If this were possible to do, then you could rewrite (1-50) as the Euler-Lagrange equation (1-53) provided you used U in place of the potential V. Now we can think of the real force and real coordinate as examples of the generalized versions of the same, if there are no constraints. In equation on page 21, we have in fact written the real force for a certain problem (charged particle in an EM field) as exactly that combination of derivatives acting on a generalized potential U which in turn is given in (1-60). Therefore, we can apply the normal Euler Lagrange equation to problems like this in E&M provided we use L = T - U. This final L Lagrangian then has the form L = T - q + q/c A v . So this is NOT an L which you write as L = T - V. Note added 9.29.16. Here is the math to prove the claim of the opening section. We have (1-50) () - = Qj and then (1-54) says () - = Qj So if these equations are both true we could subtract them to get () - = 0 provided we take L = T - U. In this section the generalized and real coordinates are the same, so equation γ on page 21 matches (1-54). Comments: This is of course a purely classical result. However, we know for a point charge in motion at position r1 and velocity v that = q (r-r1) and J = qv(r-r1) and we know that we can think of the combination (,J) = J as a conserved 4-current in that J = 0. Suppose we define (,A/c) = A as a 4-potential of some kind, so that (-,A/c) = A . Then we can say, L = T - q + q/c A v = T + AJ (r-r1) Although a little strange, this shows that we have the Lorentz covariant combination of A and J. This last term is of course more or less the interaction Lagrangian of QED field theory, where we would have something like J = e for the electron current and A being the photon field. We are getting a little preview or "omen" of field theory here. This expression L = T - q + q/c A v is not easy to find in my books, so Goldstein has done us all a big favor by writing it down and showing its significance. Jackson has it on page 407 (how could he NOT write this somewhere). He talks about p instead of J and that fixes up my slight mess above, so we then have L = T - q + q/c A v = T + (q/m)Ap where p = m(1,v) where I have fudged some relativistic details for now because I cannot remember them. We now have a complete subject change in mid section p 21. We know that if we have non-conservative forces, we can use the Lagrange form given earlier in (1-50). We want to show how we can incorporate friction into our analysis, provided frictional force is proportional to velocity as in Fi = -kvi. In this case, we first derive the Rayleigh Dissipative Function F as bottom of page 21. The algebra on page 22 top shows that we can think of 2F as the rate at which the frictional process consumes energy. The main result, however, is that we can write this frictional effect as a generalized force Qj = -(F/j) where we don't at this point really have to specify details of our generalized coordinates. So we can use this form for Qj in our result (1-50) and we end up with equation on page 22. We cannot incorporate F into the Lagrangian as in our previous situation, so we have to carry the extra term shown. But now we can use the Lagrangian method to solve problems which include linear friction! And earlier we learned how to incorporate EM fields as well. So this has been a fairly heavy-duty Goldstein section! We will now finally consider some "problems". [ But we don't get a problem illustrating the friction thing until page 45 in the next chapter. ] OK 1/9/12 OK 9/29/16 Comment: Neither of these examples fits into a simple mold where you have force Q = -V(qi, i, t). In the EM case there are potentials φ and A, but Q isn't just a gradient of these things as (1-59) shows. Maybe the friction example does sort of fit this mold because you end up with Qi = - ∂F/∂i where yes, this is a "velocity gradient" and not a normal position gradient, and F = (1/2) k i2 plays the role of V. 1.6 Simple applications of the Lagrangian Formulation (22) Claim is made that since T and V are scalars, calculations are simpler than those dealing with vector forces. Observation is made that you can always write T in the form given in (1-62) and following, for any choice of generalized coordinates. And if there is no explicit time activity in the r(q) equations (scleronomous = time independent constraints) , only the ajk term survives, quite obvious. Thus, in this case, T is a "homogeneous function of degree 2" in the various j , meaning like x2 - 3xy + 2 y2. So don't forget that this quadratic form of T is related to the idea of time-independent constraints! He now looks at some examples: 1(a) particle in space, arbitrary force F (no potential given), so T = as shown, and everything boils down to F = ma. In this situation, the generalized coordinates are selected to be the same as the Cartesian ones. 1(b) same but in planar polar coords. Now take r and as the coordinates, compute T in terms of them. The r(q) equations are for example x = rcos. But we maintain two generalized coords here, no constraints to reduce the number. We compute the generalized forces and use the (1-50) version of Lagrange. That is, we have no potential V here, but we have some generalized forces. One is the radial force, the other is the tangential force which causes a torque which drives the ang mom. 2. Atwood's machine -- two weights and a pulley. Only one coordinate x, write T and V, turn the crank to get top page 26. There was a constraint in this problem -- the positions of the two weights are correlated so only one variable x is needed, not two. We never worry about this constraint in the L formulation, and correspondingly, the tension in the rope is not revealed in the solution. 3. Bead sliding on a uniformly rotating wire. When treated in r,θ coordinates, the one holonomic constraint is that θ = ωt which write as θ - ωt = 0 = f(θ,t), This is then a time-dependent constraint. So we have two variables and one constraint, this leaves r as the only degree of freedom. Our L formalism tells us that = r2 and we see the bead accelerating outward as expected. The details of the constraint forces on the bead never enter our realm of concern! Author then reviews a series of mechanics texts, with some interesting comments. Whittaker's famous book has no drawings! Milne's makes a mess of vectors. These are mostly very famous names. I at least read through all the problems, they are not bad, nine of them. [ I have the Whittaker 1917 2nd Ed now, and it has a whole chapter on the 3-body problem. I read a little and found that this system consists of 9 2nd order PDE's. That would be a great chapter to read some day. ] Summary: In this Chapter 1, we explained the interpretation of d'Alembert's Principle involving applied forces (and ignoring constraint forces) and virtual displacements. Then we used this to derive two versions of the Lagrange equation of motion. The first involves T and Qi and does not require conservative forces and is (1-50) on page 18. The second (1-53) does require conservative forces derived from a potential V, and this contains L = T-V, a quantity known as the Lagrangian, and no Qi appears. We saw how one can use a "generalized" potential U in place of V in the Lagrangian to handle a charged particle in an EM field, and we saw also how one can add a term to Lagrange's equation to handle linear frictional forces. Then a few simple problems were considered after a bibliography with discussion. Comments 10.28.16. My Lagrange doc Appendix D has clarified nearly all of the mysteries present in Goldtstein's Chapter 1. But I really should prepare a better summary maybe with pasted stuff from G. As of 10.29.16 I am pretty happy with Goldstein Chapter 1. Wrote much of it into Lag doc. ********************************************************************************* Chapter 2: Variational Principles and Lagrange's Equations (30) ********************************************************************************* 2.1 Hamilton's Principle (30). This principle says that I = ∫L dt is stationary along the actual path of motion, where L is of course the Lagrangian T - V. In other words, I = 0. In more detail, you think of t as a parameter and L = L(qi(t), i(t), t). The endpoints are some t1 and t2. We know that "physics" will result in some definite solution for the path of motion which would be the various qi(t). It is certainly not obvious by just staring that the physical path qi(t) minimizes or maximizes this integral I. 2.2 Calculus of Variations (31). Here we work in one dimension and replace the parameter t with parameter x, and the overdot means d/dx. We start with the very general form (2-3). We imagine some parameterization of the alternate paths y(t) as in (2-4) where we add parameter . Then by "doing the math", we find that J = 0 is the same as saying [ f/y - d/dx(df/d) ] = 0 which is (2-11) and which has our familiar Lagrange form if x were time t and if f were L. A key step is the parts integration in 2-8 where the "parts" vanish because we assume the endpoints are fixed so the derivative dy/d = 0 at the endpoints. The picture on page 32 shows two possible paths in some sense separated by "" but this separation goes to 0 at the end points. At this point we pause to do three simple sample problems. In the first p34, he shows that y = mx+b is the form of the path that minimizes our integral J which is set up to be the distance along an arbitrary path connecting two points. The solution method is extremely obscure but it works! We just turn the little crank. The path length is ds = dx and this is what we integrate to get integral J. The second problem asks what curve you would use to minimize a surface of revolution's area. The curve which defines the area has two endpoints. You might think that a conical surface would be the answer, since this would at least minimize the distance between the two endpoints, but that is of course not the answer because the more distant points generate more area. We write the area as an integral with the single parameter x, and then it is in the form (2-3) for a J integral, and we turn the crank. The solution curve is an inverse hyperbolic cosine, the solution has the form y = a ch-1(x/a) + b. We have the two conditions that y(x1) = y1 and y(x2) = y2 and these determine a and b. The result looks a little like what he has plotted, but is more like this: ( solid sym axis is vertical graph axis) You can see that the shape pulls the surface in more quickly than a conical surface, to start cutting down on area! But too much pulling would increase the horizontal component of the surface. This problem is interesting because I don't know how else you would solve it! There are an infinite number of ways to parameterize a curve, you would not likely hit upon the right one. You might do a polynomial and then recognize the series for the ch function in your result from regular calculus where you just try to minimize area. Again, if we write the area as integral J, we know how to make J = 0 with our Lagrange-like equation. It seems clear that this must be a minimum from the problem's situation. The third problem is the brachistochrone problem, the word from Gk meaning "shortest time". [ 10.30.16: My description here is wrong, I have now lots of stuff on this problem in a separate folder in Mechanics. Bead starts at origin and slides to some ending point on the leftmost down-hump. Point A is always the origin, and point B is some point on the first down-hump. ] The problem is this: first pick two points A and B in 2-space like the two dots below. Then imagine a frictionless rigid wire which constrains a bead sliding on it and describe this wire by a red curve. The problem is then to find the shape of the curve that minimizes the time for the bead to slide from A to B. The integral is of dt = ds/v and our author leaves the solution to the problems, but we know the answer is that the curve is a cycloid which starts at a cusp and passes through the second point: plot([t-sin(t), -1 + cos(t), t=0..20]); // How you use Maple to do a parametric curve! We start at the upper dot shown at (x,y) = (6.3) on the horizontal axis, and the solution curve heads straight down, then bends off toward the destination point. Depending on where the destination is, we might have a minimum in the curve or we might not. Again, how would you find this solution by some other method? [ sorry about errors here ] Summary of this section: we defined J = ∫dx f(y, ; x) and we showed that a Lagrange-like equation causes J = 0. We did three examples in which the parameter was a geometric "x", never time. In these examples, we found the curve y(x) that caused J = 0, and in all three cases this represented a minimum of J. We minimized a distance, then an area, and then a time. 2.3 Derivation of Lagrange's Equations from Hamilton's Principle (36). We first extend f to multiple variables as in (2-12), and we imagine the role of our single parameter as shown in (2-13). We do the exact same math and we arrive at the generalized multi-variable result (2-16) which is again "Lagrange like". Again, we assume endpoints are fixed in the nD space and the solution is a set of yi(x). In the function f we start with, we assume that the yi are independent, otherwise the result does not follow. The Lagrange-like equations are called the Euler-Lagrange equations. Now we just make the associations shown below (2-2') [p 38] where parameter x becomes t, and we end up with exactly the "Lagrange equations" of mechanics. So this is just a special case of the E-L equations. Conclusion: We have derived the Lagrange equations of motion from Hamilton's Principle of having I = 0 where I is the integral of the Lagrangian over an arbitrary path. You could of course have done this derivation in the other direction. The two are equivalent. But in our examples, we learned some things outside this mechanics bailiwick. [ See Lag doc Appendix D for detail of this section. ] 2.4 Extension of Hamilton's Principle to more general systems. (38) Above we got our result (Hamilton's Principle Lagrange equations) in terms of the Lagrangian L = T-V, but this assumes the existence of a potential V. In the first logic here, we consider using T+W instead of L = T-V inside the integral I = 0. W has the unusual definition 2-18 which has dimensions of work, but is not the usual differential work. When we think about W, we are supposed to think of the forces as constant, and we are varying the path r or q. So W is the differential "virtual" work associated with varying the path at some point on the path. Then W ~ F r and we get to (2-19). Hamilton now says integral of δT + δW is zero. Hold on this for a second, and consider this side alley: If there exists a potential V, we can write Q = dV/dq as earlier, add a zero term, and we are at the bottom of page 39. In this case, we end up with our previous I(L)=0 version. So here we have just shown that if you put T+W inside, or L inside, you get the same thing, for the case of a conservative force. You would put the T+W inside in the case that you had a non-potential-derived force. In both cases we assume holonomic constraints so we can say the qi are independent. But, if there is no V, you can probably still use the T+W form! So again starting from the T+W form, he derives (2-21) which is the same as the non-potential result (1-50) we got back on page 18. So, the T+W gives this result as well. Holonomic means the qi are independent. We are now done thinking about non-potential-derived forces and the T+W method. Next, suppose the n qi are not independent because we have m differential constraints of the form (2-22). On page 32 top he shows that a holonomic constraint implies a differential constraint, but not vice versa. So the case of m differential constraints is a broader set of problems which includes holonomic constraints. Comment Added: In the previous chapter, we had N particles so 3N variables. There were N ri vectors. In that chapter we managed to cut this down to n = 3N-k independent generlized variables qi because we had k holonomic constraints of the form 1-35. In the current chapter section, there is no mention of the ri anymore. We start off here assuming we have n coordinates qi. Perhaps we got there from the same starting system and a set of holonomic constraints. However we got to the n qi, we NOW want to add to the soup a new set of m non-holonoic constraints as in 2-23. This means that only n-m of the qi are independent, and we are going to pick the first n-m of the qi as the independent qi. If we are doing our virtual differentials, we then get m constraints (2-23). The next 4 equations are obvious, and note that we have introduced m "parameters" called l, one for each of our m differential constraints. They are arbitrary so far. But then in (2-27) we write out "the last m expressions" which appear inside (2-26), and we force them to be zero. But then (2-27) is m equations in the m unknowns l, so we know there will be a set of values for the m l which will allow these m expressions to be zero. Notice that the k sum in (2-26) goes all the way up to n, but because of the 0's in (2-27) we can now run this sum up to n-m, which gives us (2-28). This equation involves only n-m of the qi and we can regard these as independent since we had m constraints, so then we set the inner guts of (2-28) to 0 which gives (2-29), for k = 1,2...n-m. But we already had these being 0 for the last m in (2-27), so we can say now that all n of the factors are 0 and this is what (2-30) says. So this is the thing that replaces the simpler result (1-53) on page 18 if we have differential non-holonomic constraints. We have this extra stuff in the equation involving the Lagrange multipliers l and the constants alk from the differential constraints. We can think then of (2-30) as being n equations with n+m unknowns, which are the n qi and the m l. But the constraint equations written as (2-31) gives us m more equations, so we then have n+m equations in n+m unknowns. The next paragraph (last on page 42) gives a way to interpret the term on the RHS of (2-30). Remember what we did. We usually think of our set of {qi} as being independent [ and maybe we had a larger set of N {ri} with a set of 3N-n holonomic constraints which brought is down to the set of n {qi} ]. But now we are saying that in our set {qi}, only n-m of them are independent due to our m differential constraints. This then gives rise to the RHS of (2-30). If all the {qi} were independent, the RHS of (2-30) would be 0 for k = 1..n. So we can think of the RHS as somehow "arising from the non-holo constraints". Now before continuing on this line, we need to go back to the logic leading up to equation (1-50) on page 18. In deriving this equation, we assumed on page 15 (1-40) that the forces of constraint "did no virtual work", and this limited us to a set of rigid body problems with no friction. Suppose we had decided to maintain the constraint forces in our development there (these forces were there called fi). Then on the RHS of (1-50) we would have ended up with Qj + Qj' where the Qj' = i fi ri/qi. Then in going from (1-50) to (1-53), we would have "removed" Qj from the picture, but the Qj' would have remained on the RHS of (1.53). But looking at the RHS of (2-30), we see that we can then identify Qj' = l l alk. So we can interpret this term as the generalized constraint forces! [ A+ ] So in this kind of solution, unlike our earlier situation, here we will find our solutions qi and we will also find the constraint forces of the problem. That is the n+m unknowns! So when does one use this Lagrange multiplier / differential constraint method? (1) Maybe it is "inconvenient" to reduce your set of {qi} to a smaller set of those that are just independent ; (2) you want to know what the forces of constraint are; I would then add: (3) you have a differential constraint. [ Note added 10.30.16: Most of the above exists in Lag doc App D.] Now we come to the example of the hoop rolling down a plane without slipping! There is a lot that we need to say about this seemingly simple example. (1) This is a problem that involves friction! The constraint here does work on the hoop as it rolls, unlike the case where a hoop slides down a ramp with no friction. The friction causes the hoop to rotate and build up angular energy. Thus, we cannot use (1-53) to solve this equation if we think of the set {qi} containing two variables x and . We cannot use (1-53) because this current problem violates the assumptions about the constraint forces that went into the derivation of (1-53). (2) The hoop has moment of inertia Mr2, it is not a disk, hence T is as shown on page 43. (3) We could use rd = dx to say that r = x, and then we could reduce the problem to just one qi and then we could use (1-53) for just one variable. But here we pretend this is "inconvenient to do" and we treat rd = dx as the differential constraint (r) d + (-1) dx = 0 so we get the two a's top of page 44. So here we have n = 2 variables qi = {x, } and we have m=1 differential constraints. We will thus have only one . Our E-L equations (2-30) then become (2-34) and (2-35). This leads to a complete solution of the problem as shown for our n+m = 3 unknowns which are , x and . The result is fascinating. If we look at the x equation, we can regard (-1) as a constraint force acting on this variable x, and we see that the constraint pushes back against the hoop's center of mass, causing it to roll down the plane at only half the frictionless rate we would have gotten. From the point of view, the constraint force is (r) and this force makes a torque which rotates the hoop, whereas in the frictionless problem the hoop would not have an angular acceleration. Note added 9.30.16. I am now confused about the origin of equation (2-26), more specifically, about the use of (2-20') in its derivation. As he says, (2-20') applies to a conservative system with a potential included in L = T+V. But with non-holonomic constraints, are things still conservative? The equation (2-20) differs from (2-20') and I don't follow. For a conservative system, I agree that the (...) in (2-20') is zero since this is just Lagrange's equations. Maybe he has mixed holo and non-holo constraints here? I'm sorry, the logic flow is lacking here, so I would like to see another source on this subject. [ Lag App D] Poole: The presentation is different here, and they quote a 1996 paper on the subject and bring in semi-holonomic constraints and it is all very ugly. But at the start they do make the claim that the Lagrange equations under the integral are valid with any constraints and that the holo constraint assumption is only needed to claim that the δqi are independent. So with nonholo they won't be independent, but the equation (2-20') is still valid since its same derivation works (starting with Hamilton's principle for L). This derivation unfortunately is presented with other variable names on page 37. At the top of that page it is stated that the yi (the qi) are independent, but as of (2-15) this fact has not yet been used. So (2-15) is valid even if the yi are NOT independent, and changing names you then get (2-20') where the δqi now are NOT independent. Then we do the λi "trick": You can add the λi term in (2-26) "for free" since that term is 0. Then you can select the λi to kill off m of the brackets in Σk for m dependent δqi terms, and then the rest of the terms are independent, and THEN you get the desired result (2-27) if you have the right λi. and you have now derived equation (2-30) which is my goal here. Then the claim that you can interpret Σiλiaik = Q'k to be forces which would simulate the non-holo constraints! On page 47 of Poole it is said "for the simplifying assuption that λi= λi(t), so they add least mention that detail. I tried to find the 2nd Ed but could not. I have now written this up in Lagrange Appendix D, I think it is now OK. 2.5 Advantages of a variational principle formulation (44). Claim is that the variational method works best when you do have independent coordinates. It is "elegant" because a one-liner contains all the physics. Also, L is a Lorentz scalar, so method is fully frame independent. The next topic here is not as dramatic to me as it is to the author. I know that the same differential equations govern mechanics and E&M and other things, so I am used to applying the math methods of one field in another. Here he makes this comment for the Lagrangian approach in general. His example is a very good one, however, of a set of 3 (in this case) fully interlinked RLC circuits each with a voltage driver. It is not at all obvious to me where this Lagrangian comes from! The energy stored in an inductor is ½LI2 and ½CV2 for a cap, so that is not where these terms are coming from. There is no real connection between L = T+V here. Rather, we just have a candidate proposed L which gives the right equations so it must be an acceptable L. Let's write Kirchhoff for the top loop (#1): -E1 + I1R1 + Q1/C + 1L1 + M12 2 + M13 3 = 0 Now apply dt to this to get 1R1 + I1/C + 1L1 + M12 2 + M13 3 = 1 Rearrange terms 1L1 + ( M12 2 + M13 3) + 1R1 + I1/C = 1 1L1 + Σk≠1 M1k k + 1R1 + I1/C = 1 and this then is where Gold's 2-39 comes from. The equations for loops 2 and 3 have the same form, and then if we treat these as E-L equations, we can work backwards somehow and get his L. The model for the Lagrange equatinos is p 22α which includes the dissipative F term which we know is 2-38, the power burned in the resistors. I have never seen this example anywhere! It is not in any of my E&M books, and it is hard to find on the web. So it makes his point quite well. The current Ii is the generalized coordinate qi and the velocity i is i. In his Lagrangian, the constants L and M are like "masses", while 1/C is a "spring constant" and the i are like F = qE = I i, so interpret i as your "force" . He then says there is a more subtle point that the variational idea when transferred to some different type of system implies a structural analogy with your original system. Thus it is that you might "carry over" the idea of quantizing a particle to the idea of quantizing a field. QED was quite new when Goldstein wrote this book [1950], at least the Feynman version which I think was 1949. 2.6 Conservation theorems and symmetry properties (47) You can learn a lot about your problem even if you cannot do the full double integration (F = ma) to solve it exactly. In general, if we have n variables qi, we will have n 2nd order differential equations, each having two constants for 2n required constants, and these are provided for example by the n qi and the n i initial conditions. Gold does not mention that these in general are coupled 2nd order ODE's. If we define the n generalized momenta as L/i pi, then the L equation says i = L/qi. For a problem with no qi dependence in V and the usual T, we find that pi= mi = mvi which is the familiar result. But for other Lagrangians like our E&M one, we get an extra term eAi/c as shown in (2-42), and this can be interpreted as the field momentum. Since our generalized force is Qi = - V/dqi, we see that we have a "generalized Newton's law" which is i = Qi which involves the generalized force and the generalized momentum. [ this is in Lag App D ] If a certain coordinate qi does not appear in L, it is called cyclic. In this case, i = L/qi = 0 and the corresponding momentum is seen to be conserved. G comments on Routh's Method, but we are going to do that later in the book. Our author then does two very well presented sections, taken right from The Book. Conservation of linear momentum: In the first, we set up dqi as apropo for a coordinate translation in some direction as shown in the page 50 picture. He shows first that the generalized momentum pi is the component of mvi in the direction (top p 51). Notice that he is considering a system of many particles, and he assumes that dri = dqj for ALL of the particles i, where dqj is a change in only one of the generalized coordinates which we assume somehow causes this uniform translation of all the Cartesian coordinates. He then says that, if this qi is cyclic and so does not appear in L, then we know i = 0. So we then associate in this case conservation of linear momentum in the n direction with this cyclic qi. Because L is cyclic in qi , pi is conserved. This seems pretty clear even without his pictures. Conservation of angular momentum: He then repeats this discussion for angular momentum. Everything is the same but now the picture is shown on page 52 and n is now the axis of rotation. In this situation, our selected qj causes the change in all the ri as shown (they rotate about n). In this case, he first shows that Qj = nN where N is the total torque on the system. So now the generalized force is the component of torque along the n direction. He goes on to show that pj = nL and we then have j = Qj saying that torque drives ang mom in the usual way. And if L is cyclic in qj , then j = 0 and our projection of ang mom is conserved. Everything in this section is done with very simple math and leans on results from page 17 having to do with the relationship between the ri and the qj. Gold notes that conservation of p and L are associated with translational and rotational symmetries. New topic starts on page 53. If we assume that L has no explicit time dependence, then he claims this is the same as time-independent constraints, scleronomous p 23. As a counter example, on page 55 we have a constraint which is moving in time, so perhaps you would have some V(t) and hence L(t). Assuming no explicit time dependence, and assuming no velocities in V (in order to use the Lagrange equation in L), author shows that a certain quantity H iipi - L is conserved, dH/dt = 0. So far we don't know much about what H is. But, if T is homogeneous of degree 2 in the i [ which author identifies with no explicit t dependence, see p 23], then we can use Euler's theorem on T as shown in (2-51), and we quickly end up showing that in this case, H = T+V = total energy. The section ends with a subtle issue of whether or not the work done by forces of constraint is included in this H. The answer is that it is not, but with time-independent constraints no such work is done anyway. The V that appears in the Lagrangian formulation only involves the "applied" forces. A situation where the constraint is time dependent is shown page 55 where the path of constraint is itself moving. So our actions in this last section would not apply to this situation. He has not yet referred to this H thing as a "Hamiltonian". Chapter ends as usual with some annotated references and some problems, all of which I have at least read through, but I did none of the problems. [ Have now down the brach thing.] Note Added: In this section, H is defined as Σp - L and is given no name. On page 54 if we assume conservative systems (meaning V independent of velocity), we have pi = ∂T/∂i [ a key fact! If V is function of velocity, all you know is pi = ∂L/∂i] . We then have Σpii = Σi∂T/∂i = 2T // By a tricky use of "Euler's Theorem" Let's pause to look at this theorem. Suppose f(xi) is such that f(λxi) = λn f(xi). For example, suppose you had a function f(xi) = f(x1, x2) = A11x12 + A12x1x2 + A13x22. Then f(λx1, λx2) = λ2 f(x1, x2) and this example is then "homogeneous of degree 2". Now, on page 23 G shows that if you express T of a system in any generalized coordinates qi, it will be "quadratic" in this n=2 sense as long as the equations of transformations r = r(qi,t) are not functions of t (scleronomous). Therefore, the above "tricky use" equation is true for scleronomous generalized coordinates. Now with the above in place, we have H = Σp - L = (2T) – (T-V) = T + V = total energy!! So, the association of H with total energy only applies when (1) V is conservative; (2) transformation between physical and generalized coordinates is time-independent; (3) there is an implicit third condition to worry about here, namely, that any possible constraints be holonomic. Only then does the set of generalized coordinates qi contain 3N-k independent coordinates qi. This is needed to make Lagrange equations 1-50 be valid without Lagrange multiplier modification. Presumably, if we have some kind of dissipative constraints, they are blocked by condition 1 above where the dissipation is modeled inside V. So, exactly when can we say that H = T+V is the conserved total energy of a system? I think if the three conditions above are met, this will be true. The example on page 55 violates our second condition because r = r(qi,t) is a function of time. A rotating frame of reference would be another example I think. If we are given a Hamiltonian which is obviously in the form T+V with a conservative potential V and expressed in coordinates which are scleronomously related to physical Cartesian coordinates and we have no non-holonomic constraints, I think we can conclude that H = E = conserved total energy of the system. // review to here on 10.2.16, and again 10.30.16 Comment: There is a ton of advanced Lagrange mechanics stuffed into the first 50 pages of this book! And there are many tons lying ahead, book is 372 pages. I read this book part in 2009 then rest in 2011. ********************************************************************************* Chapter 3: The 2-body problem with a central force (58) ********************************************************************************* 3.1 Reduction to the equivalent 1-body problem (58) We start off assuming a quite general Lagrangian of the form 3-1 where the potential might contain higher order derivatives of r, V(r,, ...). R is the rcms. The potential is assumed to be a function only of the relative r and its derivatives. Moving to the CMS rest frame and defining the famous reduced mass , we find that this limited 2-body problem reduces to a simple equivalent one-body problem for particle of mass and we define an origin at r = 0. Note added 10.31.16. (1) Why does spherical symmetry of V(r) in the equivalent single-particle problem (particle goes around origin somehow) imply L conserved? Answer: since F = -V, the force acting on the particle is in the direction. Torque on the particle is r x F which is then 0, hence L conserved in full 3D. (2) Why is motion planar? We know L = r x p. Then L r = r r x p = r x r p = 0 p = 0. Since this is true for any position r, L is normal to the plane of motion. Thus the plane of motion is the plane that is perp to L. Since L = constant (conserved, no torque), this plane never changes. 3.2 Central Force Problems (59). We now further assume that V has the simple form V(r) which means the force is along . Since the problem is now spherically symmetric, any angle you come up with about an axis through the origin will be a cyclic coordinate. We know that L is conserved as a 3-vector. A simple argument shows that any motion must lie in a plane: for each r on the orbit, we know that rL = 0, which means all such r lie in a plane perp to L. You might as well select a z-axis to line up with L, and then your problem has only two variables we can call r and which described the planar motion. We then write the Lagrangian as in 3-6 (T now expressed in polar coordinates r,θ), see pencil notes bottom of page for v2. Since does not appear in L (as predicted, is cyclic), we know that p is conserved. But p = mr2, and of course this "canonical momentum" is really l about the L axis, so we write l = mr2 which we regard as a "first integral" of one of the two Lagrange equations ( that is to say, 1-53 for qi = ). Goldstein then shows that r2/2 = dA/dt, so conservation of l is really the same as the notion of constant areal velocity -- equal area is swept out in equal time. This Kepler Law #2 is true for an arbitrary central potential V(r), not just inverse square. G then writes down the second Lagrange equation (the one in r) in 3-10, defines f(r) in the obvious way. He replaces using the previous "first integral" to get 3-12 which involves , 1/r3 and f(r). Meanwhile, we know the T+V is the conserved energy so we write 3-13 and define this energy as E. Again we replace and get the form 3-15 for E which now involves , 1/r2 and V(r). Web says that an archaic term for integration is quadrature, derived from an old idea of finding a square whose area was equal to the area under a curve. So these first and second integrations are called "quadratures" by our archaic Goldstein. Equation 3-15 can be formally integrated to find t(r) as shown in 3-18, from which you could presumably obtain the inverse r(t). Notice how the "first integral" got rid of which was sitting in Lagrange equation 3-10. In similar fashion, we integrate 3-8 to get 3-20 to give (t). So, embodied in 3-18 and 3-20 we have formal solutions to the arbitrary central force V(r) two-body problem! Given any V(r), select a value for E and l and do a numerical integration to find your solution. Nice comment is made here about the four integration constants. It is good to have E and l be among these constants, since E and l survive the transition to quantum mechanics from classical mechanics. Nothing is yet said about the signs of these two numbers. Other int constants are r0 and θ0. Note added 10.31.16. I think this means the following. If you were to specify E, l, r0 and θ0, then r(t) and θ(t) are fully determined. So pick a point r0,θ0 in the plane, specify E and l, and there is only one orbit that works, only one possible r(t) for that particle, but there must be a ± hiding somewhere. I guess it would be the sign of the square root in (3-16) which determines the sign of . Note added 10.31.16. The EL equations are 2nd order PDEs, so to get a solution you in effect have to do two levels of integration. Doing the first level ("first integrals") is easy for this type of problem, no work really needed. The first level integrals yielded equations which involve first derivatives of and . It is the second level of integration to get r and θ that requires more work, and these are the formal integrals shown on page 63. 3.3 The equivalent one-dim problem, and the classification of orbits (63). What do we know without doing a lot of work? From 3-16 we know vr as function of r, and from 3-21 we know v, so we can then compute v = . Thus, even without solving the problem, we know all about the velocity vector v as a function of r (but we don't know r as a function of t without integrating.) Remember that v is a "first integral" and we have done that level of the problem. Next: Looking at 3-12, we seem to have an F = ma problem in variable r, but there is an extra force term l2/mr3, so we can think of this as some kind of extra fictitious force term in the one-dim problem and incorporate it into an effective total force f '(r) arising from an effective potential V'(r) as shown in 3-22 and 3-22'. This extra term is in fact the familiar centrifugal force mv2/r where v = v. This reminds me of the "radial equation" for the hydrogen atom which is a "one dimensional problem" for a fixed l . First example: consider attractive inverse square force as shown. Plot the two terms which make up V' (dashed lines in Fig 3-3) and plot their sum (done for some l 0). For small r, the fictitious term wins and V' goes positive, then we have a nadir as shown and then V=0 at infinity. For E 0 have unbound orbits (hyperbolas and parabola), for E<0 can have bound orbits (ellipses). Notion of point of closest approach r1 for unbound orbits as shown page 66. Bound orbits (which become bound states in QM) might not be "closed", but move somehow between two circles in general case. So this example shows how you can "classify" the orbits for a known problem, and you get similar results for a similar V(r). Second example: attractive inverse quartic force has plots for V and V' as top page 68. Now the negative 1/r4 term "wins" at small r and you get V' as shown. In this interesting case, for a given energy E you can have both a bound and a separate unbound orbit type. Third example: attractive linear force (HO), and you get V and V' as shown bottom page 68 for this case. In both cases there are only bound orbits. These are ellipses since each direction has same HO frequency. So the main point: even if you are unable or unwilling to solve a central force problem, you can probably make some clear statements about the nature of its orbits by plotting the effective potential V' of the equivalent one-dim problem. Comments added 10.31.16. Goldstein is paying a lot of preparatory attention to quantum mechanics. In QM you have the same notion of potentials and turning points. Classically these turning points are "hard", but in QM they are soft, allowing penetration ("tunneling") of ψ through the walls of the potential, because as long as the walls are not infinitely steep, ψ has to be smooth at the wall and that means it can go through it a little bit. This has implications for semiconductors and many things. The classical r0 and θ0 become lost in QM since orbits are just ψ clouds, but E and l remain. I can see how the case on p 68 with a central "well" implies possible "bound states" for elementary particles, and particle theory has been littered over decades with "potential models" for things, such as the Yukawa potential. Lots of history. "Controversy" regarding the cover of the 3rd edition (11.12.16) Note page 66 Fig 3.7 which shows some generic classical orbit having two turning points. This Fig 3.7 is drawn differently in GP&S, so wit and this is the picture that appears on the cover. Here is a complaint from a book Galactic Dynamics A 2003 letter (I have in PDF) comments on errors that have survived all 3 editions of Goldstein and which talk about this picture. The writer Martin Tiersten claims that you cannot have out-facing curvature in any "attractive" central force orbit. The corrected picture for the above appears on the cover of Marion and Thornton, competitors in this market, showing that you basically have a rotating ellipse situation for such an orbit, His easy derivation is that Fn = mv2/R as the normal force to the center at any point on the orbit, where R is the instantaneous radius or curvature. Thus since Fn never changes sign, R can never change sign and must always be concave toward the origin! So Goldstein's original on p 66 is wrong as are the follow-on two pictures in 3rd Ed. Very good. The authors Safko and Poole respond apologetically. BUT, I don't see why Fn must always point to the center. For a single-power potential like V = k/rn you would have Fn = -nk/rn+1 so in that case Tiersten's comment applies. But if you have two terms with different powers such as V = k/r - k'/r3 then Fn = -kr +3k'/r4 with k,k' > 0. In this case Fn can have either sign. This then is how Safko ended up stating things in the errata, from which I quote: Here is another version of the comment with the above typo corrected, So I think Tiersten's letter only applies to a purely-attractive central force where the sign cannot change. It is true though that Goldstein on page 66 implies that his picture applies to the simple V = -k/r force shown on page 65. He is in a discussion of inverse square force when he draws his picture. So the authors clarify their covers with the above comments (bailing them out). 3.4 The virial theorem (69). This theorem applies to any set of particles having any total forces Fi. It is not necessary even to have a potential. The only requirement is that the motion be bounded so that in effect after enough time the orbit repeats so that G ≡ Σipi ri at t = 0 repeats at some huge t = T. The theorem states that <T> = - (1/2) < Σi Fi ri> for time averages, as shown on p 70. The RHS here is called the "virial". Wiki: The word virial for the right-hand side of the equation derives from vis, the Latin word for "force" or "energy", and was given its technical definition by Rudolf Clausius in 1870.[1] force in singular, strength in plural We start with some basic algebra and we quickly arrive at 3-26 which is the virial theorem for an arbitrary system of particles, where Fi are applied and constraint forces. The theorem gives a way to compute the time average (periodic or long term) of the kinetic energy T as a time average of the sum of dot products shown, the virial of Clausius. G says this is useful in both ideal and non-ideal gas physics where you do in fact have constraints, like walls of a container, but only gives us a reference. Specialize to forces from a potential V, get 3-27. Of course here you cannot have constraint forces unless they are coming out of the potential somehow, like the U business earlier. Specialize to single particle, get 3-28. Specialize to power-law V(r), get 3-29. Specialize to n = -2 for inverse square force, get 3-30, which is a famous result <T> = -1/2 <V>. 3.5 Differential equation for the orbit (71). Back to central force f(r) again. Can always replace time derivative with derivative as in 3-31. Take our earlier result 3-12 (from the r Euler-Lagrange equation, sub'd in), define u = 1/r, and end up with 3-34 which is an ODE for the orbit. It is second order in , t does not appear but u does, so you can solve this thing for (u) for example. You need l, m and f(1/u) as inputs to this equation. People are usually more interested in the orbit than say r(t). If you arrange for a turning point to be at = 0, then orbit is symmetrical in . If we rewrite 3-17 in terms of u, we can write (u) as an integral as in 3-37. This is in effect the solution of the little ODE, but we reuse earlier work to get this result. The book then gives us a list of values of "n" for which the integral can be done in terms of "simple functions", assuming a power law central potential. The simplest are arcsin and arcsinh integrals which he calls "circular functions". These arise when the radical R is a quadratic. But this only gives you cases n = -2 and -3. For n = -1, you are talking a log potential and nothing really works right, we are "out of the power bailiwick". The next more complex class of cases that can be integrated are the elliptical functions which involve R as a quadratic radical as shown page 74. This brings in the cases n = 0 as the only new one. By doing little changes of variable, one can get n = -5 and -7, and then another change gives +5 and +3. So the bottom line is that if you insist on closed form circular or elliptic functions, you have to have n = 1,3,5 or you need -n = 0,2,3,4,5,7. Some fractional n lead also to elliptics. Goldstein harps on this a bit because in 1950 you could not just throw an integral into "Maple" on your home PC, or do a simple numerical integration to find an answer. And as always, we like closed form results where possible. Summary: Central potential context. We already had t = integral of some F(E,l,r) in (3.18) to find t(r) Due to the simple first integral fact ldt = mr2dθ it is easy to replace "t" by "θ" in any ODE or integral. If you replace in this way t by θ, you end up with ODE's or integrals allowing you to compute θ(r) instead of t(r), and θ(r) or r(θ) is the equation of an orbit. If you specialize to power law central potentials, you can ask for which powers n do the integrals generate known functions, and which do not. For some powers including inverse square n = -2 you get "circular functions" like sin or sh or their arcs, while for other powers you get elliptic functions, and yet others you get functions having no name. It is often convenient to use u ≡ 1/r in place of r when doing work in this area. 3.6 The Kepler problem: inverse square law force. (76) If only Kepler had a copy of this book in 1600. However, Leibnitz and Newton did not discover calculus until the 1670's, so Kepler might have been unable to read the book the way I can read it now. In this section are derived the remaining two of Kepler's laws, they just fall out from the simple theory like rain from the sky. We assume f = -k/r2 and solve the little orbit DE to find result 3-43 which says right off the bat: all orbits are conic sections, with some unknown integration constants. But then we use the integral we had earlier for and "do it again", this time obtaining 3-46. The added feature here is that we now have a formula for the orbit eccentricity as a function of E, l, m and k (remember that m is the reduced mass). I think G has made a sign error in this last calculation but the error does not propagate into the answer. It probably involves a different branch of the square root or something like that. I could not find errata for this 1950 edition on the web, but learned that there is a third 2001 edition. Note added 11.12.16. There is a sign error in (3-45) then then also in equation β. But this has no effect on the main result (3.46) because cos(θ-θ') is an even function of its argumenty. I wrote this up in Euler doc. I have corrected both sign errors in red in Goldstein copy. Now that we have a formula for r(), namely 3-46, we could integrate 3-19 (top p78) to get (t), but this is not done in this section of the book. Here we are content with knowledge of the orbits. So, we take 3-46 = 3-47 and discuss the nature of the orbits relative to the size of . This of course agrees with the general orbit discussion earlier for the central force problem. For a circular orbit, time averages equal quantities so we use the virial theorem to get 3-50, not that exciting, but an application of our little virial theorem of a previous section. One point of confusion that caught me. On page 66 we see some orbits bounded by r1 and r2 circles. One could imagine having an inscribed ellipse and one would be tempted to say a = r2 and b = r1. One must remember, however, that the "r" in this picture is the "relative r" which is r = r2-r1. In the heavy sun with planet picture, this would be distance between planet and sun, and sun is at a focus, so here you see that in fact r1+ r2 = 2a. For a binary star where masses might be similar, the picture in the CMS frame of the resulting orbits is like this: Again, we have r = r2-r1. We know the solution for r is "an ellipse", but when this is translated back into the orbits for r1 and r2 using page 58, we obtain a picture like that above for the CMS frame (non-reduced) problem. Each star does an ellipse of the same eccentricity, and the two stars are always on a line through the CMS as shown, and the two reach closest approach to the CMS at the same time. You can arrange for one ellipse to be entirely within the other. Perhaps in the later editions this situation is examined in more detail. As you increase the mass of one of the binaries, it's ellipse gets smaller and closes down around the focus of the ellipse of the lighter mass. A nice simulation is given at http://ltc.cit.cornell.edu/courses/astro101/java/binary/binary.htm which shows the possible locations of the two ellipses and what happens as you increase mass. For equal masses, the two ellipses always overlap and are identical in shape and size. Now back to the reduced mass or one-mass-is-huge view. We get a list of conics versus on page 78. then on page 79 we relate a, C and and we find that "a" is a function of E and not of l. Finally, using that fact that the area sweep rate is constant and area of ellipse is ab, we derive an expression for the period of the orbit in 3-54. We find that 2 a3 which is the main idea of Kepler Law #3. Again, m = and k = force constant. If we apply this to gravitational force and planet/sun, we get 3-56 which would be the actual Kepler result which says 2 = const * a3 with the same constant for all planets. Only the mass of the sun and G and a enter in this formula. Update 11.14.16. A loose-bolt detail here. If you look at the ellipse equation (3.46) you see a leading constant C = mk/l2 which G fails to identify as "C". The usual ellipse equation is for r, not 1/r, and there we know that the leading constant is a(1-ε2). Thus we can identify 1/C = a(1-ε2) which says a = [C(1-ε2)]-1 . G eventually shows this fact in p 79 (3-51). Now since b = and c = εa we know that b = a. But we just showed that = = a-1/2 . This then leads to our desired result that b = a1/2. The ellipse area is then A = πab = πa3/2 which later figures into the 3rd law. In fact the difference between a and a1 is very small except for Jupiter. In any event, Kepler #3 as he saw it is (3-56). Update 11.14.16. G starts with the simple 3-54 version of Kepler #3 for the 1-body problem. For the 2-body problem you replace mass with μ and assume sun dominates so you get (3-56). But in this formula, G does not point out that a still refers to the 1-body problem orbit, which reduces to a1 and a2 for the real orbits of the two objects as I show in binary doc. However, the key point is that Kepler is observing things basically from the frame of reference of the sun, so a is then the correct thing to use, not a1. Again, it is interesting to contemplate the history of mankind during which the motion of the planets was gradually understood, and moved from religious to scientific reality. Wiki notes that Kepler tried to fit his work in with biblical quotations and existing astrology, his time being 1610. In the Bohr atom, we have roughly that = me and k = e2 so in this classical atom, you ought to have elliptical motion for electrons in bound states. We always have = 1 if l = 0 so S states are "circular" and higher l states are "elliptical". In classical mechanics, orbit is always in a plane, but in the quantum atom, the plane must be statistically distributed at all angles, giving the spherical S orbital. Would be interesting to read more on this subject -- connection between classical and quantum atoms. Another comment on the sign error. GR7 has it right! Note added 10.31.16. Sign Confusion page 77. In 1st edition the integral (3.45) has a plus sign, and in 3rd edition we also have a + sign, // wrong What do GR7 have to say? From their page page 94 I quote: In our application, what GR7 call Δ is negative, so the last entry is relevant. Goldstein has q = -Δ. Now start with the above and use this rule arcsin(α) + arccos(α) = π/2 Thus the above integral can be written, I = [ - arccos() + π/2] But you can drop the π/2 since this is an indefinite integral. We then have I = arccos() Now use another rule which says arccos(α) = π - arccos(-α) Again drop the π and you end up with I = - arccos( - ) right If GR7 is correct, then both 1st and 3rd Ed are missing the leading overall minus sign. [correct] GR7 agrees with Schuam p 72. Now let's assume for the moment that Goldstein is correct and that (3-44) is correct, which in 3rd and 1st is this Now write q = (2mk/l2)2 ( 1 + 2El2/mk2) which I have verified on scratch. Then the ∫.... object in (3-44) is equal to this IG = + arccos( - ) = arccos( - ) c = -1 = [ (2mk/l2) - 2u ] / { (2mk/l2) } = [ 1 - 2u(l2/2mk) ] / { } = [ 1 - ul2/mk ] / { } and then - = [ ul2/mk - 1 ] / { } This implies that θ = θ' - [ arccos( [ 1 - ul2/mk ] / { } ) wrong sign after θ' This matches both editions, and I quote it from 3rd ed So everybody is in agreement except GR7 !! All cleared up in Euler doc! 3.7 Central Force Scattering (81). History puts the Kepler application first (1620). The Bohr atom came later in 1913-15 (Livesey p 158), for inverse square central forces. Here we do the classical inverse square force scattering problem, and this is the Rutherford discovery of 1911-13. We shall capture previous work using 3-64 for the "orbit". After all, the trajectory of a scattered particle is in fact going to be a piece of a hyperbolic orbit, and we have taken from 3-46 and restated in 3-65, where we replace angular momentum l by impact parameter s according to 3-59. Here are details. [ this section was rewritten 11.15.16 ] [81] Lots of words about scattering that I am OK with. A ring of solid angle dΩ is defined, and the cross section concept is defined in 3-57. [82] A decent scattering picture and definition of the scattering angle Θ. Simple connection between impact parameter s and angular momentum l in 3.59, based on L = rxp and E as K.E. Then eq 3.60 describes what I wrote in as dN with ds > 0. It happens that in this case, dΘ < 0 just from the way things work (see picture), so we have a strange-looking minus sign. Solve for σ to get 3-62. Notice that the beam intensity I cancels out. Cross section has dimensions of area and we have the same minus sign. [83] Projectile has Z', target has Z, so we get a value for force constant k as shown (3.36). The force is assumed repulsive here!! Look at 3-41 to see that this means k > 0 so V = |k|/r repulsive. He then directly copies 3-46 with this value of k and it has the standard polar variables r and θ as always. He can replace l with s from 3-59 so ε is as shown in 3-65. To keep r positive, we must have 3-66. Break time. Here I went off and wrote notes in hyperbola.doc. I wanted to know exactly what a hyperbola's two branches look like in polar coordinates. It is a rather long story actually. I only worked with horizontal hyperbolas. You have to take a different path depending on whether you want the origin at the left or at the right focus. Here are the results: Origin at left focus: r1 r+ = cos(θ) > - 1/ε makes the left branch r2 r- = cos(θ) > + 1/ε makes the right branch Origin at right focus: r3 r+ = - cos(θ) < - 1/ε makes the left branch r4 r- = - cos(θ) < +1/ε makes the right branch Notice that r3 = - r1 and r4 = - r2. Numerator always positive for hyperbola. Goldstein's equation 3-64 fits the mold of r3 above. Here is a sample of such a left branch. (1) You think of the hyperbola as the scattering path of a particle which is repelled by a fixed particle at the origin (the right side focus). Polar coordinate r is the distance from that origin to a point on the path, and that is the distance over which your repulsive potential operates. Goldstein does not want to do this, but if you somehow had an attractive potential, the path would be better modelled by r4 which has this curve (2) Here the attractive potential source is located at the origin which is still the "right focal point", and we just ignore the left branch here which is off the page to the left. Goldstein Fig 3-15. This is a bit confusing for several reasons: (1) his equation 3-64 directly implies my Fig 1 above where particle comes in from the left. But in this figure, he shows the particle path instead being on the right side. (2) he does not show the location of the attractive potential, I added it. This is my (2) above. (3) for some reason there are ink blobs on both his asymptotes which just adds confusion and Jim's book has this as well. The third edition completely avoids this picture! So I now take Fig (1) and rotate slightly it to get a scattering picture as follows: I note that the slanted black line and the slanted asymptote are really parallel for a point on the hyerbola which is very far away. This picture shows many useful things. (1) the scattering angle Θ is the same as in G's Fig 3-13, it is the standard scattering angle away from straight ahead. . (2) For far-away incoming and outgoing hyberbola points I show θin and θout in official polar coordinates, same as in my earlier pix above. (3) Angle Φ is then Φ ≡ π - Θ = 2α = θin- θout (4) It is true that cos(θout) = cos(θin) = -1/ε since in general cos(θ) < - 1/ε for r3 but these are both extremal points on the hyperbola. I have marked θin and θout on Gold Fig 3-14. (5) Also it is true that (just look at the above drawing), θout = π - α = π - Φ/2 θin = π + α = π - Φ/2 (6) this math then follows at once (these two equations appear as page 83 A ) - 1/ε = cosθout = cos(π-Φ/2) = -cos(Φ/2) cos(Φ/2) = 1/ε sin(Θ/2) = sin([π-Φ]/2) = sin(π/2 - Φ/2) = cos(Φ/2) = 1/ε (7) then it is trivial that (this is equation B on page 83) cot2(Θ/2) = ε2 - 1 // based on sin(Θ/2) = 1/ε and thus csc(Θ/2) = ε (8) But ε2 - 1 is clearly the second term in the radical in 3-65, and so 3-67 follows. This gives the scattering angle Θ in terms of the impact parameter s. Note that E = (1/2) mv02 > 0 when the particle is far away, and energy conservation means that E = T+V maintains this value on the entire path. When it gets close to the repulsive center, T slows down and V increaes so E is maintained. This is elastic scattering, energy in = energy out for scattered particle. Finally we reach the end of page 83 !! [84] Now rewrite 3-67 as page 84 A. Then compute ds/dΘ using 3-62 and this gives B and then finally simplifying 3-68. This is as he says the very famous Rutherford scattering formula with that inverse 4th power of the sine of the scattering angle. This peaks strongly when 0 which corresponds to large s. Note the cross section naming conventions, Goldstein particle physics sources cross section σ(Θ) = dσ/dΩ This G's equation C really says, in particle notation, σtot = ∫dσ = ∫(dσ/dΩ)dΩ G notes that the Rutherford cross section he gets agrees with the QM version non-rel, and I have done the relativstic calculation many times in Bjorken Drell (I think). For this kind of scattering, σtot = ∞ because there is scattering at any s you want. This is also a well-known fact. This is the Rutherford classical result as derived in my Livesey book page 154 where E = 1/2mv2. Rutherford did the theory in 1911, Geiger and Muller did the experiments in 1913, and this showed the existence of a very small nucleus and the planetary model of the atom (Bohr model) was born. Most beam particles go straight through, only those that come very close to a nucleus get deflected a lot, hence the 1/sin4(/2) factor. It showed that the 1/r2 force law was accurate down to something like 10-14 m. The beam was alpha particles created by a radon source. The Livesey derivation assumes a massive target like a gold nucleus, so the CMS kinematics are very close to the lab kinematics. But Goldstein is now going to do this correctly in the next section. Note retained from 10.12.08: On page 83 top G sets up the problem for a positively charged projectile (like an α particle) scattering off a positively charged center (like a nucleus). On that same page, he comments just in passing that if we had one object positive and the other negative -- the attractive case -- the picture on page 84 still applies, the only difference being that the scattering center is on the right side at F instead of on the left side at F'. So the orbit is unchanged. So we find the interesting result that the cross section is the same in either the attractive or repulsive case, something that is not obvious to me at first glance. The third edition adds more stuff to this section discussing other potential shapes. In particular, if you have a combination of attractive and repulsive potentials you can get orbits which circle the target one or more times before leaving. The authors talk about rainbow and glory scattering in passing. I guess they felt this was the right place to add more content about more general scattering situations. There is always the issue of classical versus quantum analysis. I am happy without this extra stuff. I would like at some point to do the gravity assist analysis when an object takes energy off a moving planet. I don't think I have ever done that. Maybe it is trivial. 3.8 The Lab Frame (85). [p 85] New notes 11.16.16. I have to say that this is NOT a very clear Goldstein section. His notation is not very good, he assumes things that are not obvious at least to me. He just gives waypoints instead of a clean derivation of his results. I derived all these equations "at one time" but have no record of these derivations. Some equations don't really need a record, but others do. On first reading, I am confused all three figures. I agree that in the CMS frame where R = 0 all the time we can use results obtained earlier to relate the effective 1-body problem to the physical 2-body problem. In the 1-body problem the orbit of that body is a hyperbola (branch) and we then have these equations (taken from binary star doc). I follow G by priming the CMS vectors. r = r'1- r'2 R = (m1r'1+m2r'2)/(m1+ m2) = 0 r'1 = (μ/m1) r r'2 = - (μ/m2) r In the binary star situation, the ellipse shape was translated into the ellipse shape of the orbit of each particle because in r'1,θ'1 for example θ'1= θ and r'1 = (μ/m1)r. Similarly θ'2= θ+π and r'2 = (μ/m2)r and this caused the orbits to have opposite orientations. The same thing happens here in the CMS frame. The 1-body hyperbolic orbit translates into a pair of "opposite" physical hyperbolic orbits and that is what is shown in Fig 3-17. Goldstein assumes the reader will figure out what I have just stated. [p 86] Proof that CMS scattering angle is the same as the 1-body scattering angle The first puzzler is why the CMS scattering angle should be the same as the 1-body scattering angle. He draws pictures on top of p 86 and simply states that this is so, as if it were totally obvious. How can I prove this claim? In the CMS consider first for the 1-body position (think of ∞ as a very large but finite #) rin ≡ r(t=-∞) rout ≡ r(t=+∞) rin rout = rinroutcosΘ and this Θ would be the 1-body scattering angle. Now define a bunch more CMS objects, r'1in ≡ r'1(t=-∞) r'1out ≡ r'1(t=+∞) r'2in ≡ r'2(t=-∞) r'2out ≡ r'2(t=+∞) r'1in r'1out = r'1inr'1outcosθ'1s r'2in r'2out = r'2inr'2outcosθ'2s In this last line, θ'1s is the deflection or scattering angle for particle 1. I agree that since we are in the CMS frame, the angles θ'1s and θ'2s have a simple relationship. I have draw these angles on G's Fig 3-17 and it is clear that these angles must be the same. Since the CMS is at rest, if one particle goes off at some angle, the other has to go off at the same angle (each angle measured from particle's initial direction), otherwise the CMS would no longer be at rest after collision. So let's call this angle just θ's. Then we have r'1in r'1out = r'1inr'1outcosθ's r'2in r'2out = r'2inr'2outcosθ's (1) But now use r'1in = (μ/m1) rin r'1out = (μ/m1) rout r'2in = (μ/m2) rin r'2out = (μ/m2) rout and then r'1in r'1out = (μ/m1)2 rin rout = (μ/m1)2rinroutcosΘ = [(μ/m1)rin] [(μ/m1)rout] cosΘ = [r'1in] [r'1out] cosΘ (2) Comparing (1) and (2) we conclude that θ's = Θ . The same result obtains using 2 instead of 1 subscript. r'2in r'2out = (μ/m2)2 rin rout = (μ/m2)2rinroutcosΘ = [(μ/m2)rin] [(μ/m2)rout] cosΘ = [r'2in] [r'2out] cosΘ (3) and then comparing (1) and (3) also shows θ's = Θ . So QED on this proof. You have to make use of equations which relate the 1-body problem to the 2-body problem as I have just done. As noted, G assumes this proof is obvious. His Fig 3-16 seems very strange. I understand the label θ there. The line segment seems to refer to r1(t) - r2(t) = r(t) at some instant of time, but we need things at t = +∞. The figure shows the angle What is Fig 3-16 all about? Thus purports to show the 2-body problem activity in the lab frame. Here is one interpretation of this figure The particle tracks are shown in red, I assume each has an asymptote which I draw, and then there would be an angle between the two asymptotes which I call θL12 which is the relative angle of the two particles at t = +∞. Is he claiming that this angle is the same as the 1-body problem scattering angle Θ? No, I think that is the wrong interpretation of the figure. Let's try another interpretation, close view far away view Here I show the two particles in the lab frame at some finite time t, and r(t) has angle Θ(t). Maybe the implication is that as t→∞, Θ(t) → Θ which is the 1-body angle of r after the scattering. But such a picture seems unhelpful because we don't off-hand now r1(t) and r2(t) so hard to draw this thing at various times t. On the right I back up the camera 10 miles and draw the paths. But again, I have no idea what the asymptotic value of Θ is. This picture just seems absolutely worthless!!! It just says that there is a difference vector r and it ends up having the angle Θ at t = + very large time. Conclusion: I have no idea what the purpose of Fig 3-16 is in the mind of Goldstein. [p 87] In the lab frame, the CMS point is located at some moving position R(t). The CMS coordinates of the two particles are defined by r'1 = r1- R // position relative to the CMS point // (3-68) r'2 = r2- R and differentiating v'1 = v1- // position relative to the CMS point // (3-68) v'2 = v2- Since there is no external force on our 2-particle system, we know that (t) = (m1v1(t)+m2v2(t))/(m1+ m2) = (μ/m2) v1(t) + (μ/m1) v2(t) = constant This is true because total momentum is conserved, P = m1v1(t)+m2v2(t) = constant What is Fig 3-18 showing? Now lets define some new objects v1in = 1in = 1(t = -∞) v'1in = '1in = '1(t =- ∞) v1out = 1out = 1(t = +∞) v'1out = '1out = '1(t =+ ∞) In Figure 3-10 he is representing the equation v1out = v'1out + By definition θ is the angle of v1out and I agree with the two locations in Fig 3-18 where this is shown. I also agree that the angle of v'1out is Θ because I showed that earlier, so I agree with the one Θ label in that picture. Now from the figure I can write these two equations: v'1out cosΘ + = v1out cosθ v'1out sinΘ = v1out sinθ Divide them to bget tanθ = // agrees with (3-69) and this then relates the lab particle 1 scattering angle θ to the CMS scattering angle Θ. Now recall from above that r'1 = (μ/m1) r so that '1 = (μ/m1) or v'1 = (μ/m1) v and of course v'1out = (μ/m1) out. // agrees with p 87 A Now I think G is assuming this is an elastic collosion, so the speeds are the same before and after. Then v'1out = v'1in. But v'1in = (μ/m1) in . I think he uses v0 as a name for in in the 1-body problem, so then v'1in = (μ/m1) v0 . Therefore v'1out = (μ/m1) v0 // agrees with (3-70) Now at t = -∞ we can identify the lab frame with the 1-body problem. A particle is incident on a fixed scattering center, so v0 is also v1in , the main lab parameter. At any time in the lab frame we have some center of mass velocity given by = (m1v1+m2v2)/(m1+ m2) = 0 . At t = -∞ this is just in = (m1v1in)/(m1+ m2) At t = +∞ it is the same but given by out = (m1v1out + m2v2out)/(m1+ m2) Setting these equal gives m1v1in = m1v1out + m2v2out = (m1+ m2) The left = sign shows momentum conservation, the second just comes from above. Equating left and right m1v0 ≡ m1v1in = (m1+ m2) // agrees with page 87 B = [ m1/ (m1+ m2)] v0 = [ m2m1/ (m1+ m2)] (1/m2) v0 = (μ/m2) v0 // (3-71) Now go back to (3-69) which we already proved, tanθ = (3-71) (3-70) and replace = (μ/m2) v0 and v'1out = (μ/m1) v0 , so then tanθ = = = = // agrees with (3-72) and now we have a relatively simple relationship between the cms angle Θ and the lab angle θ. The angles are thus the same if m1 << m2. [p 88] In p 88A, why is I the same on both sides of the equation? I is particles/area/sec and it seems that if frames are boosted, you would have I' ≠ I. This issue is not clarified in 3rd edition. I think the remedy is to go back to 3-57 and rewrite it this way σ(Ω)dΩ = Now there is no mention of intensity or time. The beam is turned on at t = 0 and off at some t = T and during the experiment the total incident particles was 1,000,000 and the number scattered into dΩ was perhaps 100. If two observers in frame S and S' watch this experiment, they both agree that the experiment involve 1,000,000 particles total. The guy in frame S sees 100 into dΩ, and the guy in frame S' sees those same 100 but into his dΩ' and dΩ ≠ dΩ' in general. So then σ(Ω)dΩ = = = σ'(Ω)dΩ' = I might write = cross section in the lab frame = σ(Ω) of Goldstein = cross section in the CMS frame = σ'(Ω) of Goldstein and in this notation we have dΩ = dΩ' = So we then get, in G language, σ'(Ω') = σ(Ω) Now let S = lab and S' = CMS frame. Then we have using the rings for dΩ σCMS(Θ) = σLAB(θ) = σLAB(θ) But let's write this the other way σLAB(θ) = σCMS(Θ) Whereas on the previous page G used primes to indicate the CMS frame and no primes for the LAB frame, here he suddenly exacdtly reverses that notation (horrible!) and writes the above as LAB CMS σ'(θ) = σ(Θ) // agrees with (3-73) OK, fine. If you really wanted to know the fraction you have to do a lot of work based on 3-72 which relates the angles Θ and θ , and G does not carry out this task for general m1 and m2. But for the special case m1= m2 which applies to say neutron-proton scattering, the relation is simply θ = Θ/2. Why? tanθ = = = = = tan(Θ/2) θ = Θ/2 // agrees with (3-74) Example: if you backscatter 180o in the CMS, θ = 90o . How can that be? Well in fact the incoming ball has stopped completely so the 90o does not mean anything. But suppose you are just a little off center so in the CMS you have 179o. Then the initial particle goes off very slowly sideways, it is true. It is also true in billiards which is a much different kind of interaction but I guess is still a central force: It happens that in billiards you always get as 90o angle in the lab frame as shown above from the web. Now, for cross section relation you then get, = (sinΘ/sinθ) dΘ/dθ = 2 (sinΘ/sinθ) = 2 (sin2θ/sinθ) = 4sinθcosθ/sinθ = 4 cosθ and then you get σ'(θ) = 4cos(θ) σ(2θ) // agrees with (3-75) LAB CMS [p 89] G goes on to compute how much KE is transferred from the incident particle to the target. He gives a general result (3-76) which I won't derive here (looks easy) and then notes if m1 = m2 the result is very simple v1/v0 = cosθ Example: for a full backscatter, θ = 90o in effect, as noted above, and then v1 lab after colliosion = 0. Goldstein's point is that ALL the energy of the incoming particle is transferred to the target. G notes that in a nuclear pile moderator the hot neutrons collide with H atoms in the moderator and slow down and this increases their fission probability. H does not work well since it absorbs neutrons, but deuterium and carbon are pretty good if not optimal. G concludes stating that you can consider scattering as a black-box operation and our work here involving cross sections applies regardless of what happens in the box, quantum or not. G got his PhD in 1943, at the peak of the nuclear weapon era (Hiro 1945). He was at the MIT Rad Lab doing radar, but later he actually worked on reactor design and moderators and shielding. A prof at Colombia 61-05 where he died emeritus. So you can see why that subject shows up in this book of 1950. My old notes: We define velocities in both frames, do some geometry, use E and p conservation (not very clearly, however) and derive formula 3-72 which relates the lab scattering angle script- to the CMS scattering angle . A simple counting rule relates the lab and CMS cross sections as shown in 3-73. One can write this out in detail, but I guess it is a bit messy so G does not do that. He looks at the m1 = m2 case and finds the simple result 3-74 which says the lab angle is half the CMS angle, and in the lab you can get at most 90 degree scattering, as in billiards. This is all a consequence of conservation of energy and momentum in an elastic collision where no heat is generated and there is no final change in potential energy either. He comments at the end that this means that all this math applies to "whatever happens" to actually make the collision, be it billiard balls, Coulomb force, or something else. At the time of this book 1950, there was no "particle physics" as there was later, so this simple application of E and p conservation was not found in mechanics books the way it is now found in every particle physics books in a kinematics section. Of course we now have relativistic kinematics in these books of which this non-rel stuff is a special case. If you want to slow down some neutrons, you put in a moderator of light nuclei. Heavy nuclei don't recoil so there is no KE transfer and the neutrons don't slow down. C12 works OK, H1 in paraffin works better. This was Goldstein's area of expertise, so he threw in this comment. Slowed neutrons react better to maintain nuclear reactions in a pile, but I am not wandering off on that subject right now. I did read about the Wigner heating effect and how it caused the Windgate Fire in Britain in 1957, a scary event for all involved. Note: Maple says that, if we use = y and script- = x, the conversion factor for the cross section is given by dcosy/dcosx = below, where b = m1/m2. Obviously the whole thing will a big mess in the general case. For b=1, above becomes 2(cos(y)+1)1/2 and everything simplifies, but I don't care much about these details right now. Chapter ends with the usual comments on some books, Whittaker and Sommerfeld and Slater, names from olden days of text books. He fails to put publication dates on books for some reason! I then read through all the problems, many of which are very interesting. For example, mixing 1/r2 with 1/r3 gives a precessing elliptical orbit. And so ends this long "two-body central force" chapter. ********************************************************************************* Chapter 4: Rigid Body Kinematics (93) ********************************************************************************* 4.1 The independent coordinates of a rigid body (93). G gives several arguments why there are 6 independent coordinates for a rigid body, but this is old hat for me. You position a reference point somewhere, then you align a set of axes with three Euler rotations. He talks all about the direction cosines of unit vectors in two systems rotated relative to each other, primed and unprimed. There are 9 such cosines, but of course only three are independent ones. // I reread this section 11.17.16 and it includes the definition of a "rigid" body, meaning rij = cij for all particles which is N(N-1)/2 non-independent conditions on the 3N coordinates. Yes, there are 9 direction cosines which connect two sets of unit vector axes, but only three are independent. Here just for fun, suppose I write x'i = Rijxj R = matrix of 9 direction cosines like Rij = ' = cos(θij) I know that to preserve unit vector lengths, R must be real orthogonal so RRT = 1 or RijRkj = δik which can be written Rl1Rk1 + Rl2Rk2 + Rl3Rk3 = 0 l ≠ k Rl1Rl1 + Rl2Rl2 + Rl3Rl3 = 1 l = 1,2,3 // this is (4-8) A rotation matrix R has only 3 independent parameters. There is no need ever to reread this section, I grok it. 4.2 Real orthogonal transformations (97). So here he does what I just said above. He does not use the word real. Here we put the nine direction cosines into a 3x3 matrix and we show that AAT = 1 so AT = A-1 and all that stuff. He then talks about the active and passive meanings of a rotation, but does not use those words. Done forever with this section. 4.3 Formal properties of the transformation matrix (101). Then get AB BA for sequential rotations, Finally on page 105 he defines the similarity transformation which changes a matrix A to A' but preserves the determinant. He mentions A† = AT* and talks about a complex unitary matrix, but does not use this idea yet. I guess linear algebra was not as familiar then in 1950 as it is now to technical students of many disciplines, so G felt the need to go through all this stuff. He refers us to M&M Chap 10 which is where I went to get the basics when I was doing my matrix review. All fine, do not reread. 4.4 The Euler angles (107). He calls them "Eulerian angles". He does things in terms of passive rotations, which for me means the axes are moved, not the vectors in 3space. Thus, all sin signs are changed relative to my usual active matrix notes. And he uses "Rx" in the middle, and comments on this a lot in a huge footnote, variations in conventions. He ends up with R(,,) = B()C()A() = Rz()Rx()Rz() with passive rotation forms // (4-43,44,45) He does us the favor of multiplying these things out on page 109 and also writing the inverse. No need to read this again. These last results are (4-46,47). // I had Maple verify both his big matrices, all OK. 4.5 The Cayley-Klein parameters (109). I know already how to write the above Euler rotation in the 2D representation of the rotation group, and the results are as in my notes except again for minus signs from passive versus active. Again we can think of R(,,) as above with the exact same three parameters and the result is as shown in 4-67 p 116 which no doubt agrees with something in Levitt modulo minus signs. We know there is a 1:2 homomorphism between SO(3) and SU(2) so that Q and -Q both correspond to the same rotation. If you write your 2x2 matrix as Q = (,,,) then we know there are not really 8 independent real parameters (only three as he carefully shows, since unitary and unimodular). But written in this general form, the parameters ,,, are called the Cayley-Klein parameters and they were found to be useful in fancy gyroscope problems, probably we shall see some in the next chapter. Felix Klein (1849-1925) was the guy who used them. The 2 rotation causing -1 is noted. The 2D space here is called "the u,v space" and G comments that it relates to spin 1/2 things like electrons. I know that in fact the u,v space is the J = 1/2 representation of the rotation group and yes, we call such things spinors. Finally, he shows that if you write P = (z, x,x+,-z) as a 2x2 matrix, then P' = QPQ† gives you the 3D coordinates in the rotated frame. Also, P = P† and trP = 0 obviously. You can write P = r where are the Pauli matrices. In (4-68) G expresses the Cayley-Kleins in terms of the Euler angles for example = (,,). I am happy to say that by now I have all this stuff completely nailed down in various notes, but it would have seemed strange to me back in 1970 when I probably had this class with Sidney Coleman. Just adding a few notes 11.17.16. Q is his SU(2) matrix (4-49). It must be unitary so it has the simpler form (4-56). You can rotate a 3-vector r into r' by doing P' = QPQT where P = though I have never found this very useful. G shows that using Q = you can produce the 3x3 corresponding rotation matrix which is stated in (4-63). He then writes the SU(2) specific rotation matrices and gets the SU(2) version of the Euler angle rotation in (4-67), so he does the mapping both ways SU(2) ↔ SO(3). This is another section I never need to read again, just read these notes! 4.6 Euler's Theorem on the motion of a rigid body (118). Now we get back to some physics. First, the word "displacement" for a rigid body means any change in its "position/orientation". A rotation is a displacement, so is a translation, so is a combination of the two. I had the idea that it only meant translation dr, but that is wrong. Euler's Theorem says this: any displacement of a rigid body in which some point in the body stays fixed can be represented as a rotation about some axis. It seems obvious to me that any such displacement must be a rotation about that fixed point. But an arbitrary rotation I know can be written R(,,) = exp(+iJn) which is then a rotation of γ about the n axis. So to me, with what I know, this is not a very exciting theorem. And then Chasle's theorem says that the most general displacement is a combination of a translation and a rotation. In our group-savvy modern era, we know that the product of the three Euler rotations is a single rotation about some axis which I would write as R(,,) = exp(+iJn) and I am sure I could find the direction n and the angle by consulting my notes for a very short while. [ Note added 11.19.16: this connection was not in my notes, but it is now, see this directory "Euler Angles compared with exp(-iu n.J) parameters" ]. At one time, this must have seemed interesting enough to call it Euler's Theorem. Similarly, Chastle's Theorem is just saying that the position of a rigid body can be identified with an element of the Galilean Group. Goldstein never mentions SU(2) or groups (so far), but I am sure the latest edition does this more systematically. The details of this section are as follows. We want to show that there is some vector R which is not changed by a selected general rotation matrix A, since that would mean that rotation was about the axis of that vector. So that would mean AR = R ( R = n, the rotation axis ) meaning that perhaps it could change in length by matrix A. We know this cannot happen, so we know = 1. G now uses this problem AR = 1R to present a general discussion of eigenvalues problems for a real-orthogonal matrix. He writes the equation as (A-λI)R = 0 in 4-75 and 4-76. I would write this differently as (R-λI)x = 0. To have a solution, one muast have det(A-λI) = 0 which is the secular equation (secular because first used for long term planetary motions), and then we have a cubic in λ which implies 3 roots. He shows that there is a transformation matrix X which brings A to diagonal form, so then X-1AX = Λ where Λ = diag(1,2,3). He then shows these facts (his "lemmas"). I give my own proof of each: 1. The three eigenvalues have unit length. Rx = λx, R*x = λ*x, |Rx|2 = R*x Rx = x R*TRx = xx = |x|2 , but |λx|2 = |λ|2|x|2 QED 2. At least one of the three eigenvalues is real. Cubic 4-83 must have a real root as graphs shows. 3. The product of the eigenvalues must be 1 (because det(A) = 1). X-1AX = Λ det(A) = det(Λ) = λ1λ2λ3 and det(A) = 1 since "special" SO(3). 4. If λ is an eigenvalue, then so is λ*. Coefficients in secular equation are real, so poly(λ) = 0 poly(λ*) = 0 Based on these lemmas, if there are three roots on the unit circle, one must be real and the other must be a complex pair. The real root can only be +1 or -1. Need detΛ = 0 which means λ1λ2λ2* = 1 so λ1 = 1. The conclusion after all this work is that λ = 1 is an eigenvalue! Therefore, you can find R such that AR = R or I would say Rx = x so you have found a vector which the general rotation R does not alter, so R must be a rotation about bout this axis, and thus you have proved Euler's Theorem. So what exactly is the solution vector n such that Rn = n? Assume without LOG that n is a unit vector. That is, what is the eigenvector here? There must be some rotation Q such that Qn = so write Rn = n QRn = QRQ-1 = QRQ-1 = Rz(Φ) But we know the form of Rz(Φ) and we know tr(Rz(Φ)) = 1 + 2cosΦ. But this is also tr(R), so if you are given R, you can deduce Φ in this way 1 + 2cosΦ = tr(R) = Σi Rii In my doc noted above, this means that 1 + 2cosu = R11+ R22+ R33 and that gives the result I found there. The final paragraph says that Chasle's Theorem that the most general motion is a rotation and a translation is pretty obvious, and so the position/orientation of a rigid body can be described by 6 parameters relative to some starting position. 4.7 Infinitesimal rotations (124). We now continue this sort of backward presentation. An obvious example shows that rotations don't commute, but here he is doing active rotations of an object. I now have to change notation, and I would say A = 1 + as the small rotation, replacing his with . Then in our usual generator notation, I would say = exp(iJ) - 1 iJ where J are the usual SO(3) generator matrices. We then connect to his notation by saying d = and I also know that i (Jk)lm = klm and now we avoid having to use for , etc etc etc. Note that matrix Jk is antisymmetric, thus so is η. We can then say dr = [iJ] r => drj = [iJ]jk rk = [diJ]jk rk = [ dm iJm]jk rk = dm (iJm)jk rk = dm mjkrk = jkm rk dm = [ r x d]j => dr = r x d // as in 4-94 Then in the p 132 figure we realize that d must point along the rotation axis and d must be the amount of rotation, so we are back to d = . Yes, the generators Jk are anti-symmetric matrices, so A and are antisymmetric Note added: A is antisymmetric only in its infinitesimal form. More generally A has three components Rn(u) = exp(-iunJ) = + (cosu-1) + sinu where the first term is both sym and antisym, the second term is sym, and the third term is antisym. For infinitesimal you only have the first and third term so Rn(u) ≈ + u u = small angle Notice that the result dr = r x d is telling us the change in a position point's coordinates due to a small rotation. We just derived this from the passive rotation rule for a vector r, namely that r' = exp(+i dJ) r ≈ r + i dJ r dr = i dJ r = r x d Reminder on the last equality dri = [i dJ r]i = i Ωk(Jk)isrs = εkisrsΩk = εisk rsΩk = [ r x d]i Obviously we can rotate any vector G, not just the position vector r, so the above discussion results in G' = exp(+i dJ) G ≈ G + i dJ G dG = i dJ G = G x d The whole thing is summarized in Fig 4-12 which is my "cone picture" and dΩ = Ω in place of my usual dθ = dθ. I normally think of Ω as being related to solid angle, so a little confusing just for me. Also in this section he comments on pseudo or axial vectors, such as L = r x p and such as n here. I call this presentation backwards only because one would usually begin with infinitesimal rotations, then build up the finite rotations from exponentials, and one would mention Lie Algebras and so on. The following section led me to write my Rotating Frames of Reference document finished Aug 2012. I have added some critical comments in red. ok to here on my 11.20.16 review. 4.8 Rate of change of a vector (132). I am sorry to say that the presentation here is very weak [at least for me] . First, we define a rotating frame (body) and a lab frame (space). The rotating "body" frame is some set of axes fixed to a tumbling object, and this will be a non-inertial frame in general. Imagine that we observe some activity from the this tumbling body frame. We might see a vector change by an amount dGbody. If G = r = a point in the rigid body, then dGbody = drbody = 0 because the body is "rigid" relative to the body system. But for some other vector quantity G (perhaps the position of a fired projectile), we can have dGbody 0. The claim is this: dGspace = dGbody G x d // this is his 4-99 Here now is my own derivation and interpretation of this idea: If we apply a small active rotation of amount γ = dΩ about axis to a vector, then using the usual [Ji]jk = -iεijk we can show that for any vector G we have dG = [-iθJ] G = - G x dΩ where dΩ = γ . If, for example, G were the vector representing a point fixed in a rotating body as seen from the lab frame, then we would say [dG]space = dΩ x G Now suppose G is not a fixed point in the object, but is a point that is moving relative to the rotating object. Suppose during a time interval dt we have a movement [dG]body due to this motion within the body system. Then the overall vector change in the space system will be: [dG]space = [dG]body + dΩ x G and then divide by dt to get [dG/dt]space = [dG/dt]body + x G where = d/dt 4-100 Notice that we can express the vector G itself in either space or body coordinates and we might then denote these things by Gspace and Gbody. But we need to use the same representation in all three places that G appears. [ But what the above really involves is [dG]space /dt and [dG]body/dt . One needs to explain how this is related to the notation above. In retrospect, [dG]space is the change in G "as measured in the space frame", and that is what [dG/dt]space somehow means. ] Clarification of Notational Ambiguity. There are two distinct meanings we could attribute to a label like "body" applied to some object X: Xbody = the quantity X "as expressed in body frame coordinates". [ I think this is wrong.] [ The correct statement would be for components, so could have Xbodyi as component of X, but there is no object called Xbody .] [ The idea of some vector Xbody led me into trouble. ] Xbody = the quantity X "as measured in the body frame" [ this is well defined ] For the moment, I will use superscript and subscript to distinguish these two different meanings. As an example, suppose drbody represents a displacement of a particle as measured in the body frame. We can certainly express the components of this vector in either frame, drbodybodyi and drbodyspacei. If one is nonzero, so is the other, because these vectors are related to each other by the Euler angle rotation which describes the relative orientation of the two frames. On the other hand, suppose drbody = 0. This does not imply that drspace= 0! As another example, consider [ dr/dt ]body = vbody and [ dr/dt ]space = vspace . In either of these equations, we could express our coordinates [components] in either body or space frames by putting on the right superscripts. [correct] Goldstein likes to write this as if it were some kind of "operator equation" as in 4-102, but I find that a little confusing. This operator rule would apply only to vectors. [correct] At this point, G derives an expression for ωbody in terms of the Euler angles. The last 2/3 of page 134 took me a long while to figure out, but here it is. We want to start with a rigid body at some finite Euler angles ,, relative to some standard orientation at 0,0,0, and then, starting from that location, we want to look at three infinitesimal Euler rotations such as d about its Euler axis. Think of things perhaps this way, where R = Rz()Rx()Rz() = exp(i dnJ) : R + dR = exp(i [ n + dn]J) = exp(i dnJ) exp(i nJ) = exp(i dnJ) R = R + R [idnJ ] where R, dR, Ji are 3x3 matrices and the first step is similar to Rz(+d) = Rz(d) Rz(). The infinitesimal rotation we are talking about is then represented by dR = R [idnJ ]. Each of our three cases results in a certain dn such as dn = d = d. We look at the drawings on page 107 to see this fact. Then we need to express this in primed body coordinates ( here we have dn= d(0,0,1) being in space coordinates). We then compute dn/dt = in body coordinates for each of these three cases and we give these names like and I go on to denote the frame of reference, for example space = (d/dt) ( 0,0,1). Finally we think of dn = the sum of these three small changes and we add them all up and we get the result 134. Here now are the details. First, look at a small d rotation. In this case from the first picture on page 107 we have dn = d = d so we would then say that space = = ( 0,0,1). To get this in the body system coordinates, we apply to = (0,0,1) the matrix A shown on page 109 and we find that body = ( sin sin, cos sin, cos) and this agrees with the first equation quoted page 134. Second, consider a small rotation. In this case, from the middle picture on page 107 we know that = that line of nodes = cos + sin . This is the axis about which we are doing our small rotation. We then know that dn = d and so space = = [cos + sin ] . We could then transform and into the body frame, but this is a little messy. Instead, look at page 107. We can see this relationship from the last figure: = cos ' sin ' . Thus, we have ,body = = [cos ' sin '] . I did this on a piece of paper to make sure the sign is right. And this again agrees with second line on page 134. Third, consider a small rotation. In this case we know that = ' and so dn = d = d 'and then we have body = '. So can summarize these three body results as follows: body = ( sin sin, cos sin, cos) = body = ( cos, -sin, 0) = body = ( 0, 0, 1) = Finally, we want to argue that we just add these three vectors up to get body which results from simultaneous small rotations in the three Euler angles. We then have all at once: dΩ = d + d + d = dγ = an arbitrary differential "rotation" so that = dΩ /dt = + + = (sin sin + cos, cos sin sin , cos+ ) = body = (x', y', z') components and finally we have derived 4-103. Notice that the three "Euler unit vectors" are not orthogonal. 4.9 The Coriolis Force (135). First, here is my own derivation, then I go back to my original raw notes. Now for an application. Suppose we apply the above rule first to G = r, [dr/dt]space = [dr/dt]body + x r which we then write as vspace = vbody + x r and as noted above, this is true regardless of whether one expresses the coordinates of all these vectors in the body frame or in the space frame. For example, vspacebody = vbodybody + body x r body would be one of the two possibilities. Now apply the rule again to G = vspace [d vspace /dt]space = [d vspace /dt]body + x vspace = [d (vbody + x r) /dt]body + x (vbody + x r) = [d vbody /dt] body + x [dr/dt]body + x vbody + x x r => aspace = abody + 2 x vbody + x x r and, once again, we have our choice of which of the two frames, body or space, to express all the vectors appearing in this equation. Back now to the raw notes: Subscripts here are s = space and r = body. We apply our rule (dG/dt)s = (dG/dt)r + x G to the case G = r, the position of a perhaps moving projectile. This gives vs = vr + x r as in 4-104. Then we apply the same rule again, this time to G = vs and this gives (dvs/dt)s = (dvs/dt)r + x vs The left hand side is as and the RHS becomes: (dvs/dt)r = (d[vr + x r]/dt)r = (dvr/dt)r + x (dr/dt)r = ar + x vr x vs = x [vr + x r] = x vr + x x r and we end up with the big result as = ar + 2 x vr + x x r = F/m where the last comes from F = ma in the inertial s frame. Then we have F = mar + 2 m x vr + m x x r so in the body frame we can write F' = mar F' = F 2 m x vr m x x r = F + Coriolis force + Centrifugal force Thus, in the rotating frame we see two fictitious forces as just labeled above, hence 4-107. The Centrifugal term points away from the earth in the direction (if we had cylindrical coords). Goldstein spends all of page 136 commenting on the Centrifugal force. Notice that we have "discovered" the presence of this force by our derivation, in addition to the presence of the Coriolis force. He shows that ~ 10-6 sec for the earth, a small number. He shows that Centrifugal is at most about 0.3% of gravity, and he points out that a plumb bob therefore does not point to the center of the earth, but points south of it if you are in the NH. This also affects the length of the resultant vector which one thinks of as the local g, making g smallest at the equator, so people are lighter there. Then on page 137 he starts into the Coriolis force which pushes "projectiles" to the right in the NH regardless of whether the projectile is going north or south. If pure east or west, there is no Cor force. In the SH projectiles are deflected to the left if going north or south. At the equator, there is no Cor force no matter which way you shoot. Goldstein does not attempt to interpret this force, but I will. Imagine the equator situation. When the bullet is launched, it has a horizontal space velocity that matches the earth's rotation. The distance from earth axis to projectile is constant on the trajectory so nothing happens. Now imagine you are up at 45 degrees in the NH. You shoot north and the bullet starts with a certain large azimuthal velocity. As it travels north, this azimuthal velocity soon exceeds the local azimuthal velocity because has shrunk! So the bullet drifts to the right because earth is rotating to the right (RH rule), having this extra azimuthal velocity. If bullet goes south, this effect is reversed, so the drift is the other way. But in both cases we get drift "to the right" of the trajectory. Goldstein talks about four examples: 1. The projectile, which we have already discussed. You might get 1/2 degree deflection for a 100 second air-time of a projectile. 2. Wind and isobars. Because wind is deflected to the right (NH) no matter what direction it approaches a low from, in the NH you get a RH rule CCW cyclone around the low. In the SH get CW cyclone. 3. Falling object. This is a vertical projectile, if you will. You have some x v Coriolis force here and you expect an azimuthal displacement. My physical explanation is very similar to that above. As object falls, its decreases, so its original azimuthal velocity exceeds the "local azimuthal velocity" at a lower altitude, so it appears to drift to the east since that is the direction in which the earth turns. Drop something 100 feet and it will drift about .15 inch. Would perhaps apply to bombs. 4. The Foucault Pendulum. This problem has always mystified me. You know that if at the north pole, the plane of the pendulum will appear to precess one circle each 24 hours, just because the earth turns under the fixed pendulum. It is not obvious to me that there is no precession at the equator, because the problem is complicated. The point of attachment is travelling in a circle. If the pendulum at the equator were N-S in starting motion, it is hard to imagine what could cause that plane to rotate, so seems reasonable that the plane stays fixed. Same with an E-W starting position. To really understand this, you have to solve the problem. Goldstein does not do this, but Marion does do it on pages 353-356. Let's look now at this solution! We start with three forces acting on the pendulum: gravity, string tension, and Coriolis as in (1) page 353. Marion then computes the x,y and z component of all the terms in this equation. He sets pointing up the local vertical, and has pointing to the south and pointing to the east. He is using a figure on page 349 which I penciled in a bit, where is the latitude. All the velocity components are variables in this problem, because obviously the pendulum velocity is not constant as it swings. Using the small-angle approximation for swing angle away from vertical, we get the pair of second order ODE's shown in (8) which become (9) when and z are defined. We later see that is and angular frequency of a free pendulum and z= sin, sort of a scaled down based on latitude, so z = 0 at the equator and at the pole. The two differential equations are "coupled" and a clever complex notation is used to get to (12) which has the solution as described. We can use >> z to simplify a bit and we get the solution shown in (19) rewritten as (20). Here, the vector r'(t) describes the motion of a free space pendulum (ie, our pendulum with Coriolis turned off). It could be swinging in any fixed plane. The matrix notation then shows that the pendulum appears to rotate around in the x,y plane at angular rate z = sin. So, for example, if it were located close to the equator, it might take 3 days to go around. Foucault built a big one in Paris in 1851 and it rotated every 32 hours. It was a big hit! At 30 degrees latitude you would expect 2 day period. You might think of the pendulum in this way: as it swings, it acts like a projectile, and in the NH it is deflected "to the right" on each direction of the swing, and this is what causes the circular motion. You might compute the deflection on one swing and then use that to compute the period. A web site comments that the pendulum does not pass over its EQ position. Here I have tried to draw the "orbit" using straight lines which always deflect to the right, and this shows why that fact is true: On each swing, the path is deflected "to the right" and fails to pass over the center point marked with a dot. In reality each of these lines is a curve bending to the right. 5. Atomic physics: rotations affect vibrations in a Coriolis manner. The rotations will create extra Coriolis forces in addition to the normal harmonic oscillator force of the potential well, this no doubt causes line blurring and various details. Our friend Herzberg has an "exhaustive discussion" of this effect, no surprise, this is the guy who wrote 6 whole books on spectra! References: We see M&M reviewed favorably, but Jeffreys is panned for having a lousy section and a faulty picture, despite witty quotations. The usual math methods books like Courant and Hilbert are mentioned (I never had that book). Whitaker is panned for having NO diagrams at all. The last problem tells the student to go figure out the Foucault pendulum. ********************************************************************************* Chapter 5: Rigid Body Equations of Motion (143) ********************************************************************************* [ Resuming here on 7.14.08, after a break of about 35 days during which we had the Lee-Rosa-Daisy visit, my visit to Cape Cod, the Chew visit, the Marilyn-Alan visit, and then a short Thomson stint. ] At this point I started my meta review, and cleaned up various items. This Chapter 5 contains another antiquated presentation presumably because "the math" was not well known in 1950 and so is here integrated into the text. // This Chapter is 41 pages long and thick with equations all of which I derived. 5.1 Angular Momentum and Kinetic Energy of motion about a point (143). We already know about the "division" idea for each [ ie, X = CMS part + rotational part ] . Equation 1-26 (p7) showed how you can divide up the L, and 1-29 (p8) how you can divide up T. It is stated that "often" V can be similarly divided. Examples: For gravity, we have only the CMS component. For mB dipole, we have only the rotational component. G then computes L for a set of particles in a rigid body rotating at some ω rate. The key fact for getting started is that vi = ω x ri since particles of the body are fixed in the body, this being a case of our general theorem from Chapter 4. By brute force, we find that L = Iω where I is a 3x3 symmetric matrix whose elements I would write as follows: Ijk = Σimi [ ri2 δj,k – ri,j ri,k] where ri,j = ri j = ri Matrix I is the inertia tensor, and we might compare to p = mv where m ~ I. The off diagonal elements are of the form Ixy = - Σimixiyi, diagonals are like Ixx = Σimi(ri2– xi2). The diagonal terms are the moments of inertia, the off-diagonals are the products of inertia. [ Recall that the Cartesian quadrupole moment of a charge distribution is Q= ∫d3r (r) (3rr - r2) which is traceless. The inertia tensor is similar, but without the 3, and is not traceless since it's trace will be the sum of the moments of inertia. Consider the spherical tensor moments and Q2,0 = ∫dV (r) ( 3z2 - r2 ) We would express the inertia tensor as Ijk = ∫dV (r) ( r2δjk - rjrk) for continuous mass density ρ. It then seems that we can make this claim Ijk = ∫dV (r) ( r2δjk - rjrk) = ∫dV (r) (3 r2δjk - rjrk) - 2 δjk ∫dV (r) r2 = Qjk - 2 δjk <r2> so we can analyze Ijk as having a cartesian quadrupole component plus another term involving the weighted average value of r2. Probably not too useful! ] 5.2 Tensors, Dyads and Dyadics (146) A dyad [AB] is a particular class of square matrix. One writes [AB]ij = AiBj In quantum mechanics we would write this as [AB] = |A><B| or as ABT in matrix notation. It is a "factorizable matrix" made from the components of two vectors, and of course not all square matrices are factorizable. Consider this example: [riri]jk = (ri)j(ri)k where ri is the position of the ith particle in a rigid body. We know without thinking very hard that the identity matrix is diag(1,1,1) so the "identity dyad" would be [1]jk = δj,k. Thus, although the inertial tensor is not factorizable, it can be represented as a sum of factorizable terms as follows. Ijk = Σimi ( ri2 δj,k – ri,j ri,k) [I] = Σimi ( ri2 [1] - [riri] ) // dyadic representation of inertia tensor To clarify things, I have tried to always show the dyads with square brackets. Any linear combination of dyads is called a dyadic. It is probably easiest to use the "matrix notation" such that [AB] = ABT. In this notation with real numbers of course we know that ATB = AB, the usual dot product. Here are some examples: [AB]C = ABTC = A(BTC) = A (BC) // G bottom page 147 C[AB] = CTA BT = (CA) BT Think of A as a column vector and AT as a row vector, the transpose of the column vector. Now how might you combine two dyads? Consider first "multiplication" : [AB] [CD] = (ABT) (CDT) = A (BT C) DT = (BC) ADT = (BC) [AD] // noted added 2.25.17: (AB) (CD) = (BC) (AD) I guess. Not sure what this means. Another interesting combination of course is this: A[BC]D = AT BCT D = ATB CTD = (AB)(CD) For some strange reason, G writes this also as [BC]:[AD] which I find is pretty useless. So now here is an example of this last combination n[I]n = Σimi ( ri2 n [1] n - n [riri] n ) = Σimi ( ri2 – (rin)2 ) // G 5-17 // noted added 2.25.17: n[I]n = <n | I | n > In normal matrix notation we would write Ijk = Σimi ( ri2 δj,k – ri,j ri,k) I = Σimi ( ri2 1 – ri riT ) nT I n = Σimi ( ri2 nTn –nTri riT n) = Σimi ( ri2 – (rin)2 ) In my opinion, the dyadic notation adds nothing new whatsoever, the matrix notation is much clearer, though it requires T superscripts to keep things straight. The δi,j notation is also quite clear. One more thing. Consider this: e1e2T = (0 1 0) = = what is meant by So then it is pretty clear that [1] = 1 = = + + as in 5-13. G likes to use i,j,k for 1,2,3, this is probably some old tradition of 1950. In general, is a 3x3 matrix with a single 1 in the mth row and the nth column, and of course this explains the notation such as [AB] = A1B1 + A1B2 + .... /each term is a matrix // G 5-12 So the poor student has to put up with this dyad smokescreen right in the middle of trying to learn rotational dynamics! Question #1: Does matrix I vary in time for a rigid body? I think the answer is no. Suppose we select the principle axes to evaluate I so that I = diag(I1, I2, I3). These three numbers are of course constants determined by the geometry of the rigid body. If you stupidly pick some other set of axes to compute I, you get 9 numbers, but they too are all constants. So the matrix I is like the quantity m = mass. It is a property of a rigid body and it does not vary in time. Yes, matrix I is a function of origin and axes chosen to do the evaluation, but we make that choice and we stick with it. 5.3 The inertial tensor and the moment of inertia (149). Just as we computed L earlier, we now compute T again using vi = ω x ri for particle in a rigid body. The result is T = ωIω/2 as simple algebra shows, which we compare to vmv/2 for linear action. We also find that T = ωL/2 which is comparable to T = vp/2 = (1/2)mv2. Now we know that at some time instant, we can write ω = ω . Thus T = ωIω/2 = (ω2/2) nIn = (ω2/2) Σimi ( ri2 – (rin)2 ) // from above The scalar nIn = Σimi ( ri2 – (rin)2 ) is defined as I, called the moment of inertia about the axis of rotation n. I ≡ nIn = Σimi ( ri2 – (rin)2 ) G then shows that this "I" agrees with the conventional definition as the sum over particle masses weighted by their perp distance to the rotation axis squared. The final act of this section is the proof of a simple theorem: Ia = Ib + Icms,a where a and b refer to two parallel rotation axes, but b passes through the CMS. The last term represents the quantity Md2 where d is the distance between the two parallel axes. So as you move your axis a further from the CMS of a rigid body, you get more inertia. 5.4 The eigenvalues of the inertia tensor and the principle axis transformation (151). This is another obfuscated section that is really quite simple. Since matrix I is symmetric and real, it is hermitian, and we know all about the eigenvalues and eigenvectors of a hermitian matrix from our general matrix studies. First, we consider Iψ = λψ, just making it look like QM. We solve det(I-λ1) = 0 to find the three (known to be real) eigenvalues λ which he calls I1,2,3. If they are all different, we know that the three eigenvectors are orthogonal and that they form a matrix which brings I to diagonal form by similarity. XT I X = Idiag X = { ψ1, ψ2, ψ3 } = what he calls { R1, R2, R3 } If two eigenvalues are equal, we do GSO to still end up with three orthogonal eigenvectors. If all three are equal, then Idiag = λ 1 and then I = X Idiag XT = λ 1 so I is diagonal to begin with, as G states. The eigenvalues I1,2,3 are called the principle moments of inertia, and the three axis Ri are called the principle axes. We still have our scalar I = nIn which the moment of inertia relative to n. If we rescale our unit vector to be ρ = n/, then of course we have 1 = ρIρ which we write out as in 5-28 on page 155 and this is an equation of an inertia ellipsoid in ρ-space having some arbitrary orientation. The three symmetry axes of the ellipsoid are the principle axes of the problem as is not hard to show. The coefficients of the ρ terms are of course the Iij matrix elements, hence the name. Finally, one can write I = MRo2 to define an "effective" perp distance Ro of an point object of CMS mass M which would have the same I as the rigid body, and Ro is called the radius of gyration. Comment added 3.12.17. Defining ρ ≡ / = / where I is the thing in T = (1/2)Iω2, you can then talk about "ρ-space" with coordinates ρx, ρy and ρz. The equation I = nIn becomes 1 = ρIρ and for any tensor Iij this equation describes the surface of an ellipsoid in ω-space. If your "r-space" basis is chosen so that I is a diagonal matrix, then the ρ-space ellipse aligns with the axis of ρ-space. Note that r-space and ρ-space are two completely different spaces. 5.5. Methods of solving rigid body problems and the Euler equations of motion (156). Everything is now in place to solve complicated top problems, it took 157 pages to get to this point! We have our Lagrangian as in 5-31, where now x,y,z without primes refer to the body axis which are chosen to align with the rigid body's principle axes -- don't miss this point! Since we have a fixed point by assumption, the CMS cannot be moving, so T is given by T = (1/2)[ I1ωx2 + I2ωy2 + I3ωz2]. We now use the Lagrange equation of the form 1-50 on page 18 where T not L appears. The coordinates qi are Euler angles which tell how to get from space to principle (body) axes. From page 52 we found that the generalized force Qj in this situation is nN where n is the rotation axis, and N is the torque. In our Euler angle picture page 107, we recall that the ψ rotation is about one of our body axes (z' here called z), so we set = and this gives us our Lagrange equation as in 5-33, driven by torque Nz. Ie, we are simply applying our Lagrange equation for the coordinate ψ and doing the various derivatives. Recall that constraints have to be holonomic to be able to use 1-50. Having obtained 5-33, we appeal to symmetry to get the other two equations. We would permute the definitions of x',y',z'. I like saying it this way: we recognize that equation 5-33 is one component of this vector equation: I + ω x (Iω) = N that is Iz + ωx(I2ωy) – ωy(I1ωx) = Nz From rotation group argument, all three components of the equation must be true and we get 5-34 which are called "Euler's equations of motion for a rigid body with one point fixed" (such as a top). We now get a very smart second derivation of these same equations. We known N = dL/dt in the space inertial frame. We then say Nx = (dLx/dt)space = (dLx/dt)body + (ω x L)x = (dLx/dt)body + ωyLz - ωzLy so we then have (dLx/dt)body + ωyLz - ωzLy = Nx or, since L = Iω, I1 (dωx/dt)body + ωyωzI3 - ωzωyI2 = Nx I1(dωx/dt)body – ωyωz(I2 - I3) = Nx Question 2.25.17: In L = Iω , is L the body or space angular momentum? Remember that we can evaluate ω's components in any frame we want in this equation. The other two equations are now obvious from the same derivation. Let's restate the above on more time: (dL/dt)body + ω x L = N and recall that L = I ω If there were no torques on the system, you would say (dL/dt)space = 0 (dL/dt)body = – ω x L This reminds us that the vector L appears to move as observed from the body frame which is non-inertial. So in the body frame of reference, you would have Lbody = Lbody(t) = varies in time. 5.6 Force-free motion of a rigid body (159). We now want to think about solutions for how a rigid body moves in the absence of any torque, so 5-36 applies. Before doing the actual solution, we are going to take a look at Poinsot's Construction (circa 1830). Goldstein normalizes things a little differently from Poinsot, so we have to pay attention. First, go back to page 155 and notice this definition: ρ = n/ where ω = ωn where n = instantaneous rotation axis The equation I = nIn says that I = Ixxnx2 + Ixynxny + etc. The three vectors n, ω and ρ are always along the same line, they just have different lengths. We can talk therefore about an inertia ellipsoid in any of these three spaces. Consider these various equations: L = Iω 2T = ωIω = ω L = Iω2 If we measure things in the inertial space (lab) frame, we could write things like this L = Iω(t) = constant of motion 2T = L ω(t) = constant of motion Because T and L are constants, the right side equation says that ω(t) moves in such a way that it maintains a constant dot product with L. That means ω(t) always lies on a plane known as the invariable plane. The distance of the vector origin below to this plane is ω cosθ = ω L ω(t)/Lω = Lω /L = 2T/L = Iω2/L. If we compute things from the space frame but using the body axes which are the principle axes, we find that Lx = I1ωx, Ly = I2ωy, Lz = I3ωz, or write this as Li = Iiωi. Because L = Iω(t), we are reminded of the fact that ω need NOT point in the same direction of L, something I have confused in the past. The angular velocity does not point in the same direction as the angular momentum in the general case. [ Geoff Chew was talking to me about the linear version of this? ] Now we first had our ellipsoid on page 155 written in n space, to wit I = Ixxnx2 + Ixynxny + ... in n space = (1/ω2) [Ixxωx2 + Ixyωxωy + ... ] in ω space since ω = ωn . We can write this last as 2T = Iω2 = [Ixxωx2 + Ixyωxωy + ... ] ≡ G(ω) which might then be called the "kinetic energy ellipsoid" in ω-space because its scale is set by T. If we were to take our space axes to be instantaneously aligned with the body axes and to have the origin at the center of the ellipsoid [huh?] , we would then have 2T = I1ωx2 + I2ωy2 + I3ωz2 We can put this in standard form as follows 1 = (ωx/a)2 + (ωy/b)2 + (ωz/c)2 where then a,b,c are the semi-axes of the ellipsoid, see picture page 51 of Schaum. Thus: a2 = 2T/I1 b2 = 2T/I2 c2 = 2T/I3 Now I am confused about the geometric relationship of the ellipsoid to the invariable plane. All points on the ω-ellipsoid are "legal" values of ω from the point of view of the T ellipsoid equation. But we also have that requirement that ω must lie on the invariable plane. This gives the following possibilities: We are now going to show that the case on the right is not viable. Assuming we did our I in the principle axes, we know that Lx = I1ωx, for example. Consider the function G(ω) given above which is the equation for the ellipsoidal surface for the ellipsoid with a,b,c as stated above. It is easy to show that G(ω) = 2 L This says that the "solution" value of ω must be such that the tangent at the solution point on the ellipse points in the direction of L, which is up in our pictures. That is why only the left picture applies. This means that the ellipsoid touches the invariant plane at a single point. We know that this point of contact is a point on the ellipsoid that has zero velocity because, after all, the ellipsoid is rotating instantaneously about the ω axis. [ There is another point on the ellipsoid with this property, but it is not on the invariant plane.] This is similar to a car tire at the point of contact with the ground, velox = 0. Like the tire, the ellipse is therefore "rolling without slipping" on the invariant plane. This means the ellipse cannot roll unless it's point of contact on the plane moves (unless it rolls on a symmetry axis). So this is the construction of Louis Poinsot. The ellipsoid rolls without slipping on the invariant plane. The origin of ω space does not move, so the center of the ellipse stays put, and of course the invariant plane also stays put. Alternatively, the distance from the ellipsoid center to the invariable plane is constant, due to the projection idea mentioned above. Exactly how the ellipsoid tumbles requires a solution of the problem from the Euler equations, but we know the nature of the solution by this Poinsot construction. It is similar to doing a general study of the orbits of a central force. Here is a frame from an animation at http://www.aero.iitb.ac.in/~bhat/asymmetric.html which shows a complex example: The ellipsoid has all three semi-axes different, and is started in an unstable place (rotating around the middle inertial axis), so the two curves of interest are quite complex. The yellow one on the invariant plane (trace of the point of contact) is called the herpolhode. The herpolhode of course lies entirely on the plane and is a history of the point of contact. The red curve on the surface of the ellipsoid is the polhode, which is the history of the contact point on its ellipsoidal surface. The "pink" curve is the history of one of the extremal points of the ellipsoid and shows the amount of violent tumble action that is going on here. Going back to our picture above If the ellipsoid had two axes equal (were an ellipsoid of revolution b=c) you can see that the only possible motion of ω would be in a circle on the invariable plane of radius ωsinθ, and in this case ω travels on a cone. Note that the point of contact above is NOT at the ellipse extremum, so as the ellipse rolls, the point moves in the circle, try this with a zucchini. Here from the same animation site, http://www.aero.iitb.ac.in/~bhat/axisymmetric.html and you see that both the polhode and herpolhode are circular. Perhaps this is what a near-spiral football is doing. Remember that the short ellipse axis corresponds to the largest moment of inertia, which for the football is the long direction. So this spiral is wobbling about the spiral axis. This is an example of precession. Rigid body motion is not particularly simple, but it is solvable. The same site shows some cones rolling on cones, but I have not bothered with that detail right now. The situation where I1 = I2 , meaning object is "symmetric" in x and y. This does not mean the object is a solid of revolution, however. A square post is an example. The ellipsoid is a solid of revolution, but not the actual physical rigid body. G does not make this point, but web does. This discussion starts bottom third of page 161 where we see the simplified Euler equations. At once we find that ωz = constant, and we find that ωT precesses at frequency Ω = (1-I3/I1)ωz, reminds me of my Levitt math. This is exactly the situation shown in my screen clip above. Notice that Ω can have either sign. We can write ωT = ωT(sinΩt +cosΩt ) and the precession amplitude ωT is a motion constant set by initial conditions. At top of page 163 we see these facts: T = ½ I1 ωT2 + ½ I3ωz2 L2 = (I1 ωT)2 + (I3 ωz)2 // Lx = I1ωx eg which seem pretty obvious. Think T and L2 as determining ωT and ωz. G uses the earth as an obvious example. Because it is slightly oblate, it will precess while it rotates. If it really were a perfect rigid body, the prediction is that Ω = ωz/300 which means the north pole should do a circle of some size every 300 days or 10 months. This is the torque-free effect not to be confused with the 26,000 year thing we will look at later. In reality, see a 427 day period with lots of noise, probably due to mass shifts and elastic properties of the earth. PL likes the football as a better example, since familiar to most people. The thing wobbles unless launched with perfect spiral. 5.7 The heavy symmetrical top with one point fixed (164). This is the main case study always used in rigid body motion, we have all seen it, gyroscopes, you name it. The position of the top is described exactly by the Euler angles, which is no doubt why they were invented in the first place. The tilt-down is θ, the fast spin axis angle is ψ, and the precession azimuth is φ. You have to look back to page 107 to fully appreciate these angles. This problem has a torque due to gravity acting on the CMS. The direction of the torque is along the line of intersection of the initial xy plane and the tilted down equatorial plane, the "line of nodes". You see this by doing r x F with your hand. This means that components of angmom in the ψ and φ directions will be conserved, but not that on the θ direction. Correspondingly, we see that ψ and φ are cyclic and don't appear in the Lagrangian shown top of page 165 (I derived all checked equations). So we have constants of the motion Lψ and Lφ, but he uses the generalized coordinate notation pψ and pφ . Notice that both of these are functions of and of , see mid page 165. Again, don't confused L with ω. Jumping to the middle picture on page 168, you might at first wonder now pφ is constant when the precession goes backwards, such that shows both signs during the motion (as does ). But from 5-47, you see that must compensate and go more positive when goes negative to yield a constant pφ. You might think of as a constant, but it is not constant!! In fact, 5-51 shows that is a function of the tilt-down angle θ which varies in time. So we have pψ and pφ as two constants of the motion, which G represents by ω dimensioned constants a and b as in 5-46 and 5-47. A third constant is total energy E as in 5-48. We then derive expressions for (θ) and (θ) in 5-50 and 5-51, and then we go on to write an ODE for θ. This ODE has the form 2 = f(u) where u = cosθ and f(u) is a cubic function of a,b,α,β where now α and β are constants in terms of energy E and Mgl and indirectly a. [ a and b replace constants pψ and pφ]. Notice that this is not a linear ODE, but it is still easily integrable as shown in 5-56. G claims that other books go into lots of details getting exactly elliptical function solutions, but he is not going to repeat that work here. Instead, he is going to do the qualitative stuff which is really much more important. Qualitative Analysis: So we have this cubic f(u; a,b,α,β) which we plot page 167. The places where f(u) = 0 are places where = 0 and hence where = 0, and hence are the θ "turning points" for the motion. Only the roots u1 and u2 are physical since in -1,1 range. The idea then is that u(t) moves in a way that it bounces back and forth between u1 and u2 in some manner, and this means that θ(t) has the same behavior and we see this in all three figures on page 168 top: θ is moving between the two circles in all the cases. By the way, the top's symmetry axis is historically called "the figure axis", and those curves are the path of the top of the top on the surface of a sphere whose angles are the θ and φ Euler angles. We then examine three cases based on the size of u' = b/a relative to u1 and u2. If u' lies between the other two, then (b-au) = a(b/a - u) = a(u' - u) can have either sign as motion progresses, u = cosθ. But from 5-50, this means that will have both signs, and this gives the curly middle picture on page 168. If u' is off to the right [ or I suppose off to the left ], then has a constant sign, and this is the left picture. Finally, if u' = one of the turning points, you get the cusp case figure 3. G then shows that this is the figure that always applies if you "release" a gyroscope in the usual manner. It starts at a cusp with these initial conditions. So we have two phenomena of interest here: precession and nutation. For the top, if you could arrange to have u1 = u2, the nutation action would go away, and you would have motion of the top tip just going around in φ, and this is the precession effect, frequency is given by 5-50 with your constant θ inserted. Note that this frequency is complete revolutions per second of the top tip about the z axis. So first, we note that this precession frequency is a function of that angle. As θ → 0 [Note: go back to other arrows in normal.dot! ] it would seen that should increase, so as top tips down as energy leaks out, expect slower precession, and I think I have seen this. In the more general case, you do get bobbing up and down which means there is θ(t) motion, and this part of the motion is called nutation. This also has a characteristic frequency which, in the examples shown, is faster than the precession frequency, about which nothing has yet been said. The Heavy Top Started with a Drop. Up to this point, we have not assumed that the top is "heavy" or "fast", all parameters have been general, but now we specialize to wring out more useful information. Starting with 5-58 on page 169 which says the figure axis rotational KE is much larger than the gravitational PE. But first, without approximations yet, we are going to think of that specific initial condition that we hold the thing while we spin it up, then we "drop" it to get things going. In this case, we know that our θ0 will be the θ2 shown in the page 168 drawings (in particular, the right one for the drop method). At t=0, we have both = 0 and = 0. The fact that =0 from 5-50 forces a relationship between a and b which is that b/a = u0 = cosθ0. Then the fact that = 0, which means =0, forces f(u0) = 0 which then forces α/β = uo. We then just rearrange f(u) as shown in 5-59, which exposes our known zero at u=u0, and our problem is now to find the other relevant root which will correspond to the lower circle at θ1. Using 5-60, we could write an exact expression for this u1 using the quadratic formula. This quadratic is recast in the form 5-61 where still no approximations have been made, just our starting method has been assumed. At this point, we assume that the dimensionless ratio a2/β is large and we are then in the heavy top approximation. Our quadratic then has the two solutions shown, but only the first is the physical one (so we know the other one is the one to the right of u=1 in the sketch on page 167). We then get the answer to our immediate question which answer is this: u0 - u1 = cosθ0 - cosθ1 = x1 = (I1/I3) (2Mgl/I3ωz2)sin2θ0 Since we know that θ1 is a larger angle, we know that cosθ0 - cosθ1 is positive, the larger angle has the smaller cosine. We can solve to get cosθ1 = cosθ0 – (I1/I3) (2Mgl/I3ωz2)sin2θ0 = cosθ0 - x1 This shows clearly that as ωz → ∞, the nutation amplitude decreases to 0. A slower heavy top has more nutation amplitude that a fast one. Next, we want to solve for u(t) = cosθ(t). The solution is given by 5-64 which we arrange as u(t) = cosθ(t) = cosθ0 - (x1/2) ( 1 - cos(at) ) Now we can interpret u in the page 164 figure as the normalized height of the top's tip, and we see that this height is nutating in a trig manner with nutation frequency a = (I3/I1)ωz. The actual angle θ(t) is a messier function, but we get the idea just fine. In order to get the above, we used the heavy top assumption again to approximate sinθ = sinθ0 to get 5-63 and then 5-64 and hence the above. The precession situation is described now by 5-66. This says that ωprecession = always has the same sign, but varies between 0 and a max value of β/α. You can see this in the rightmost picture on page 168. So unlike the nutation situation in which the "angle", at least in u(t), is a trig of time, here it is the precession frequency which is a trig of time. We get = (β/2a)[1 - cos(at)] = (β/a)sin2[(a/2)t] => φ(t) = (β/2a)[ t - (1/a) sin(at) ] and here are some plots with a=1 and β/2a = 1 so β=2: [ (t) on the left, φ(t) on the right ] Clearly the average ωprecession = <> = (β/2a) = Mgl/I3ωz and we then see that ωz→ ∞ takes the precession frequency down to 0. So, a very very fast top or gyroscope will precess very slowly. As it slows down, expect faster precession. We already saw how fast reduces the amplitude of nutation, but according to 5-65 it increases the frequency of nutation . At this point on page 171, G gives a summary of what we have learned about the drop-the-top case. He notes that the fast-top small nutation usually goes away in practice from frictional effects. Some have called this pseudoregular precession, meaning there would be some nutation if not for the friction. How might you start the top off to really have no nutation so you have true "regular precession" ? This requires that u1 = u2 so f(u) has a double zero. He shows that this is the same as having (0) = 0. He shows that, for this to happen, the values of (0) and (0) and θ(0) must satisfy 5-70 or 5-69. For example, we have to rule out (0) = 0. If you pick some of these items, then there will in general be two solutions for (0) that will work. These give two different precession rates with no nutation. Details are not pursued. But in general, to start a top off at some θ0 > 0 with no nutation, you need to push it in just the right way (two choices), or it will nutate on you. Yes, it was an interesting question. Heavy top started vertically (p 173). This means certainly that u1 = cosθ1 = 1 in the p 168 pictures. The question is: what happens when you just let it go? Since Lψ = Lφ in this case, you get a = b. G also shows that α = β. And he shows that with such relations between the various constants, the f(u) curve in fact has a double zero at u=1, and the third root may or may not be in the physical range depending on the values of ωx, Mgl, I1 and I3. He shows that if ωz > ωcritical , then that third zero is out of range and the top just sits there. For ωz less than this critical speed, the third zero u3 moves in range and the top then does regular nutation between u = 1 and u3. That is, it bobs always up to θ=0 and then down to some θ = θ3. This explains what we have all seen. You start a top, it stays vertical but loses speed from friction. When ω falls below critical, it starts to bob (nutate) and precess. It moves from the sleeping state to the wobbly state. As that zero moves more to the left, the nutation amplitude increases, etc etc. Some Applications (p174). (1) Precession of the Equinoxes. Because the earth is tilted relative to the sun and is oblate, the gravitational pull of the sun is more on the near side than the far side (page 175 picture pencil) and this puts a torque on the earth. The torque direction varies during the year I think (he does not say this, and in my model there would be no torque right at the equinoxes). The moon also does this, and the whole thing is a big mess. The effect is weak, hence the 26,000 year period, and the effect is obviously irregular. Called astronomical nutation. This is just one paragraph, so not much detail. (2) Gyroscopes. This is a torque-free mounting of a top in a 3D gimbals system. As you move the external holder, the top keeps pointing in the same direction. The toys we have all seen do not have this full mounting, they are really just tops with a holder which lets you easily spin them up. I should have one but I don't. With the toy, the whole thing acts as a top. (3) Gyrocompass. Not enough detail is given in this small paragraph to merit notes. I think modern ones are just gyroscopes with digital logic, but older ones were mounted in the x-y earth plane and worked that way somehow. I would need to go read about this in some specialized source. 5.8 Precession of charged bodies in a magnetic field (176). Here we get a nice derivation of Larmor precession. For a rotating classical charge distribution of, say, all electrons, you can compute M and you find that M = (e/2mc)L as in 5-76. Precession is uniform, there is no nutation. In the last paragraph, G says that |M| might depend on orientation in some situation and cause us to need a small correction. This is very obscure, but G wrote a little paper about it and threw it in. Here is the abstract of this paper, which says there is a small nutation involved in the classical solution. The Classical Motion of a Rigid Charged Body in a Magnetic Field Herbert Goldstein Massachusetts Institute of Technology, Cambridge, Massachusetts (Received June 1, 1950) A frequently used classical analog to the quantum precession of a magnetic moment in a magnetic field is the motion in a uniform static magnetic field of a symmetrical rigid body charged so that the ratio of charge to mass density is uniform. In general, the figure axis of such a magnetic top both precesses and nutates, and the mechanism producing the nutation is qualitatively different from that responsible for the nutation of the familiar gravitational top. Unless the body is spherical the angular momentum vector also nutates in the course of the precession, a phenomenon caused by the deviation of the torque on the body from the customary, but here incorrect, expression M×B. The uniform precession frequency of the magnetic top is not the Larmor frequency but involves in addition terms depending on higher powers of the magnetic field, a correction which corresponds to the quadratic Zeeman effect. ©1951 American Association of Physics Teachers I read all the book reference notes, lots of stuff. Amused to see that Klein and Sommerfeld during 1897-1910 wrote a huge 4 volume Theorie des Kreisels, Theory of Tops. This was the kind of stuff people worked on in the late Victorian era. There was no QM to divert efforts. The pure history of all this stuff is in itself interesting, you could do this forever. Comment on boomerangs. Arnold Sommerfeld was certainly a heavy lifter in the physics history, and a serious teacher and documenter. Question: Did I ever read this chapter in my class? I certainly don't remember it. I have squiggled blue ink and also orange underlining on page 143-158, but not on the remaining pages. Probably Sydney Coleman made us read part of this but then we moved on? He probably wanted to get us into the Hamiltonian formalism rather than do all this ancient top stuff. ********************************************************************************* Chapter 6. Special Relativity in Classical Mechanics (p 185) ********************************************************************************* 6.1 The basic program(185). Postulates: #1: speed of light c is the same in all inertial frames; #2: Laws of physics are the same in all inertial frames, you cannot tell whether you are moving constant v or not in any absolute sense. Therefore, Galilean velocity addition must be wrong. The program then has two parts: (1) find the correct transformation law and resulting velocity addition law; (2) adjust equations so they are covariant, ie, have the same form in all inertial frames (including boosts). 6.2 The Lorentz Transformation (187). Consider spherical light wave in two frames to get 6-7. This gives one condition of the transformation. To get the required minus sign, define x4= ict and write as 6-8. Ie, don't use a metric tensor, just do it this way (the unpopular way in my era). Then the transform matrix is going to be orthogonal AAT = 1 and not unitary (since it will contain complex elements), as in 6-10. Pick a boost direction z and reduce the 4x4 desired matrix to p 189 bottom, block diagonal. Then consider the origin of one system as seen by the other, and you end up with 6-5 where Goldstein does not use γ = 1/. The main conclusions are in 6-17 where we see how things change along the boost direction, both z and t are affected, not x and y. These results are of course independent of the metric or notation chosen. G then quickly shows length contraction and time dilation, though he uses the older term time dilatation which is 40x less popular on the web. If you measure a rod which is at rest in another frame, it appears shorter by factor 1/γ along the boost direction. A clock in another frame appears to run slower by factor γ, time dilation. The rule for addition is easily found by applying two boost matrices in sequence, and the result is β' = (β1+ β2)/(1 + β1β2). If both starting betas are 1, so is the sum. 6.3 The covariance of equations (194). If aμ is a 4-vector, then aμaμ is a scalar. Lorentz scalars are called world scalars by G. A nice example is to define dτ2 ≡ dxμdxμ and call τ the proper time. In a rest frame of a clock, dxi = 0 and then dτ = dt. In another frame, a pair of clock ticks are spaced by dt = γ dτ so this is consistent with our time dilation notion. If Xμ is some 4-vector, then the sign of XX can be plus, minus, or zero and the vector is then called space-like, time-like or light-like (G does not use that last term). So now we start to "fix things up". First idea is to define uν ≡ dxν/dτ which will be a 4-vector. We will then have ui = (dxi/(dt/γ)) = γ dxi/dt = γ vi. For low speed, they are the same. It turns out that u0 = iγc to transform correctly, and then uνuv = -c2 = a world scalar of course, and time-like. Next, the wave equation 6-26 we are all used to can be written ∂μ∂μ ψ = 0 and he shows that ∂μ is a 4-vector operator and ψ is then a scalar, hurray. 6.4 What happens to F = ma ? (199) This is the big question of course. Newton says d(mvi)/dt = Fi so we are tempted to fix this up to say d(muν)/dτ = Kν where the LHS = a 4-vector, and we hunt for a RHS 4-vector which will reduce to Fi at low speed. At this point G takes two alternate paths. First path: consider the E&M force which you suspect is already relativistically correct. We know the force Fi on a particle from E&M force and we showed earlier in the book, page 21, that you could write it as shown in 6-31 which seems a bit of a mess. But then you imagine a 4-vector Aμ = (Ai, cφ), though justification is not given, then you can write Fi as in 6-33. Notice 6-32. Well, we are led to postulate a Kμ as shown below 6-34, and then we know that Fi = Ki/γ. So at γ=1, it gives the correct RHS. So here it is: d(muν)/dτ = Kν = (q/c) (∂μ[ uνAν] - ∂τAμ ) Both sides are clearly 4-vectors, and both sides have the right d(mvi)/dt = Filimit for low speed. So we have a good candidate at least for the E&M force. But what about for other forces, like gravity say? For the EM force, we have found that Ki = γFi, so maybe this is true for any force. Fiddlings: d(muν)/dτ = Kν dt = γ dτ γd(mui)/dt = Ki = γFi so d(mui)/dt = Fi at low speed. The Kν force he calls the Minkowski force. Second path: just assume that pi = γmvi = mui. Then get dpi/dt = Ki/γ = Fi and then Newton's law looks the same. In general of course pμ = γmvμ = muμ = the momentum 4-vector of a particle and its mass is a scalar. Now at mid page 202. We have found that Ki= γFi and he then shows what K4= (iγ/c) Fv which is the rate of doing work by the force. If we define kinetic energy T as T = γmc2 , then dT/dt = Fv which rescues the idea that you do work to add to T. The non-rel limit is then T = mc2 + ½ mv2 as we expect. Next, using pν = muν we find that p4 = iT/c so now p and T are in the same 4-vector, and now conservation of pμ implies both conservations of momentum and KE. He then makes some comments about inelastic collisions and changes in mc2. In a non-rel inelastic collision, momentum is but energy is not conserved unless you include heat. Comments here are no very clear since heat is really photons. 6.5 Relativistic Lagrangian situation ( 205). G shows that a possible relativistic L is L = -mc2/γ - V if we want to use the traditional Euler-Lagrange equations. This gives pi = ∂L/∂vi = γmvi which we know is the correct conserved relativistic 3-momentum, and we get dpi/dt = Fi = -∂iV as the correct Newton's law. Also, if we use the same definition of H shown bottom page 206, we find that H = T+V = E, where T = γmc2 as we found on page 203. Finally, if we want to add an EM field's V, we use V = q(φ – Av/c) If we write this out, we get L = -mc2/γ + (q/cγ) Aμuμ but these terms are not really Lorentz scalars because of the γ factors which are only rotational scalars. Still, this L gives the right canonical momentum as shown 6-52. 6.6 Covariant Lagrangian situation (207). The previous section is really only half-baked because we continued to use the non-covariant E-L equations, though we found a candidate Lagrangian that made some things come out right. The correct covariant formulation requires 6-53 for the Lagrangian integrated over proper time τ with 4-vector arguments for position and velocity. This Lagrangian is called L' and is different from the L one of the last section, and the covariant E-L equations of 6-54 are also different. What then is the correct L' ? The answer is this: L' = ½ muνuν + (q/c)uνAν which is clearly a Lorentz scalar, and pν = ∂L'/∂uν = muν + (q/c)Aν which we know must be the correct canonical momentum. If we then apply our new covariant E-L equation, we get the results shown top of page 210, which is basically this: d(muν)dτ = Kν which agrees with our earlier result. Kν is the correct EM force on a particle. So here is the covariant form of Newton's Law, at least for a particle in the presence of an EM force. We don't really know what to do in the presence of gravitational forces. The final shot here is to show that the correct form for KE is T2 shown in 6-59. For a free particle this was T = γmc2, but here you see how the presence of A modifies things. Notice that φ does not affect T. Easy to show that if A=0, you get T = γmc2 again. G has some final comments: special relativity does not work for action-at-a-distance forces like gravity, is messy in accelerating systems (probably messier fictitious forces), and everything becomes a mess for rigid body motion which is completely non-covariant. There are a few books on this subject, but not many. It must be pretty messy if G did not attempt it. I think there is a problem with the whole idea of a rigid body in relativity, since distance between two points is boost dependent, etc. I cannot find a quick web source on this subject. I did find an interesting discussion of what happens when you continue to accelerate a "rigid" wheel up to relativistic speeds. SR is of course a huge subject and this chapter is merely touching the surface, commenting on how earlier mechanics stuff gets generalized or adjusted. One needs to solve many problems to get an understanding of things here. ********************************************************************************* Chapter 7. The Hamilton Equations of Motion ( p 215) ********************************************************************************* 7.1 Legendre Transformations and the Hamilton Equations of Motion (215) The Legendre trick is amazingly simple, I don't think I ever understood it before. Start with 7-3 and say you want to replace dx with some other du differential and some new fellow on the LHS. So just define g = f-ux and that does it! When you say g = f - ux, you mean to throw out x and install u. So in this case, you start with df = udx + vdy, and you end up with dg = - xdu +vdy. Yes, there is a minus sign. Each "coefficient" can be interpreted as a partial derivative in the very obvious manner. So in the first equation, we know that u = ∂f/∂x but in the second equation u is one of the variables. So this would be a change in independent variables from (x,y) to (u,y). So this is exactly what we are going to do. In the analogy, f = L(q,,t) and g = H(q,p,t) where we are throwing out in favor of p as independent variables. There is an overall minus sign so g' = -g is really used here. So 7-8 makes complete sense in terms of it being nothing more than a change of independent variables. Here are the starting and ending positions: dL = A dq + B d + C dt A = ∂L/∂q B = ∂L/∂ = p C = ∂L/∂t (1) dH = A' dq + B' dp + C' dt A' = ∂H/∂q B' = ∂H/∂p C' = ∂H/∂t (2) But we have H =p - L so we have more information: dH = ∂H/∂p dp + ∂H/∂ d { – ∂L/∂ d – ∂L/∂q dq – ∂L/∂t dt } (3) = dp + p d { – ∂L/∂ d – ∂L/∂q dq – ∂L/∂t dt } (4) The 2nd and 3rd terms cancel due to the definition of p, and we are left with dH = dp – ∂L/∂q dq – ∂L/∂t dt (5) But now we use the E-L equation to replace ∂L/∂q with and we have dH = dp – dq – ∂L/∂t dt 7-11 (6) Now let's do some comparisons. Comparing (2) and (6) we can say A' = ∂H/∂q = – B' = ∂H/∂p = C' = ∂H/∂t = – ∂L/∂t and these last three equations are now called Hamilton's Equations of Motion. Not clear at this point that we have gained anything by doing all this. 7.2 Cyclic coordinates and Routh's procedure (218) In the Lagrangian formulation, if you have a cyclic coordinate qn, L can still be a non-trivial function of n so you have not removed this nth coordinate from the problem completely. In the Hamiltonian formulation, if qn is cyclic, it does not appear in H (or L), and also pn = α = a constant, so solving for n = ∂H/∂pn = ∂H/∂α is a completely separated problem, so you can regard pn and qn as both being removed from the problem in the Hamiltonian situation. In passing, G mentions the Routh procedure which is simply this: write H' = Σs p - L but only include the s cyclic coordinates in the sum. Then this H' will be H for the s cyclic coordinates, but will be -L for the other n-s coordinates. You then have a Lagrangian problem in only n-s coordinates and you can get the s other things from the k = ∂H/∂pk = ∂H/∂α method. This H' is called R, the Routhian. It is an interesting idea, but I don't think G is going to really do anything with it. If you really like doing Lagrangian dynamics, this method reduces the problem you have to solve to n-s variables + s simple things to solve. You have one foot in each camp! 7.3 Conservation Theorems and Physical Significance of H (220) Equation 7-19, which G easily derives, shows that dH/dt = ∂H/∂t = -∂L/∂t. The whole formalism so far ( started in E-L equations) requires that V not be a function of velocity. In addition, if the rn equations don't have t, then I guess ∂L/dt = 0 and then dH/dt = 0 which means H = constant of the motion we call "energy". Example 1: particle in central potential V(r). We first write H in terms of r and θ and pr and pθ. We examine the Hamilton equations of motion (4 items), and we duplicate results we got in the previous chapter. Example 2: Compute H for a relativistic particle in this central potential, get 7-20 as T + V. Now let's throw in our velocity-dependent EM potential for V, and we get H as in 7-22 which is the first "recipe" formula for H. This is the non-rel version. Do it again in 7-23 for the rel version. In each case, we arrive at the H result by using the recipe of adding -qA/c to p. Covariant Form: Now, as we did in the relativity chapter, we conjecture that H' = pμuμ - L' where L' was the covariant Lagrangian we used in that chapter. The resulting H' is as in 7-26 for particle in an EM potential. We then goof around with the math to show explicitly that dH/dt = ∂H/∂t, so this is true even in the presence of this special case velocity-dependent potential. We are reminded that in 1950, no one had a covariant formulation for non-EM forces! 7.4 Hamilton's Equations from a variational principle (225). First, we start with our usual L variational thing, put in the replacement H = p - L and write dt = dq and we get 7-28' which he names the modified Hamilton's principle, the unmodified one being the usual L thing. then on page 226 we go through a long series of steps, including the usual parts where ∂q/∂α = 0 at the two endpoints. When the dust settles on page 227, instead of an Euler-Lagrange equation, we get the two Hamilton equations! The point is made that in the L world, really depended on q, just a time derivative, whereas in the H world, our variables p and q really are independent. In the L world then you sort of have n variables q to worry about, but you get a second order ODE to solve (EL) for n variables q. In the H world, you have 2n variables to worry about (the p and q) but the Ham equations are first order and there are 2n of them. This is perhaps an easier problem to deal with. 7.5 Principle of Least Action (228) The p sum in H = p - L, when integrated over time, is called the action. So the claim of least action is that, when H is conserved, the solution to your problem occurs when the integral Δ(action) = Δ( ∫Σipiidt )= 0, but this Δ is a new kind of "variation" that differs from the δ we used in all our earlier variational work in this book. This caused me much pain on first reading, I will now try to detangle it. Suddenly G is using new language. He refers now to the δ variations as "virtual displacements", two new words really. It is true that in the δ case, those η functions could be arbitrary apart from being tacked down at the ends, so if you call these η functions "displacements", or perhaps you refer to this combination as a "displacement" : qi(t,α) = qi(t) + αηi(t), then we are OK on terms. " the varied path in the δ-process need not correspond to a possible path of motion for the system", and H may not be conserved during such a variation. Now in the Δ variation, you are not doing a "virtual thing" (virtual means a δx with dt=0 recall). He suddenly starts talking about a "succession of displacements" each including dt. This is where G just plain becomes unclear, and I will need now an outside source. As an aside, I was right about the word "action", as seen here: I have found the Chand source which is very similar to Goldstein, he probably cribbed from G. Let's start by ignoring the various interpretation comments of G on page 228. and instead just look at the math. The idea is that we think of q = q(t,α) as in the δ world, but we also now thing of α = α(t), though I don't know what that means. We have qi(t,α) = qi(t) + α(t) ηi(t) I guess. This then leads to the extra term such that we have Δq = dα(dq/dα) where dq/dα = ∂q(α,t)/∂α + ∂q(α,t)/∂t dt(α)/dα Now I suppose we have to say ∂q(α,t)/∂t = so we have dq/dα = ∂q(α,t)/∂α + (α,t) dt(α)/dα Then Δq = dα(dq/dα) = dα [∂q(α,t)/∂α + (α,t) dt(α)/dα ] = dα ∂q(α,t)/∂α + (α,t) dα dt(α)/dα then he notes that the first term is our older δq, and in the second term Δt ≡ dα dt(α)/dα . So we have in the end Δq = δq + Δt This Δ thing is a certain operator applied onto q(t,α). Now suppose we go up another level and we have some f(q(t,α), t). Then the claim is this: Δf = ∂f/∂q Δq + ∂f/∂t Δt = ∂f/∂q [δq + Δt] + ∂f/∂t Δt = ∂f/∂q δq + { ∂f/∂q + ∂f/∂t} Δt So far, then, we are just playing with this Δ operator, and we just did a Δ-variation of f(q,t). Derivation of the fact that the correct dynamics result in ΔA = 0 Now instead of doing this with a general function, let's do it to "the action" A = ∫pdt = ∫(L+H)dt = ∫Ldt + H(t2-t1) if we assume H = constant independent of time, like a conserved energy. Now apply Δ. The first term gives: (here, I is the integral of L, whatever it might be, so L = for use soon below) Δ ∫Ldt = Δ [ I(t2) - I(t1) ] =ΔI(t2) - ΔI(t1) (230 a) Then we have ΔA = Δ (∫Ldt + H(t2-t1)) = ΔI(q,t2) - ΔI(q,t1) + H(Δt2-Δt1) // no written It must be that Δ is a linear operator so it can do these things. So now mid page 230. Next, consider: ΔI(q,t2) = δI(q,t2) + Δt2 where I(q) is being treated like our f(q,t). Thus, ΔI(q,t2) = δI(q,t2) + L(q,t2) Δt2 and therefore, from above, Now go back and consider: f = ∫Ldt so that = L Δ ∫L(q,t) dt = δ ∫L(q,t) dt + L L(q,t) Δt Δ ∫Ldt =ΔI(t2) - ΔI(t1) = [ δI(q,t2) - δI(q,t1)] + [ L(q,t2) Δt2 - L(q,t1) Δt1] (230 c) Now the first two terms are δ ∫L(q,t)dt for this reason: δ{ ∫L(q,t)dt } = δ { I(q,t2) - I(q,t1) } = δI(q,t2) - δI(q,t1) Then we finally arrive at 7-36 which says Δ ( ∫L(q,t) dt ) = δ ( ∫L(q,t) dt ) + [ L(q,t2) Δt2 - L(q,t1) Δt1] (7-36) So here you see the new contribution on the right. Not only might the endpoints like t1 vary with α, but L can be different at the two times, since L is probably not "conserved" in time. So two contributions to this extra term. Our next task is to work on the δ ( ∫L(q,t) dt ) term above as on bottom of page 230. We first work entirely in the δ world, using the E-L results, but then in the last step we use this reversed deal, δq = Δq – Δt and this takes us to the bottom result d on page 230. Now, why does Δq vanish at the endpoints? I guess this is part of the setup of things, the way we had δq = 0 at the endpoints of the δ variation. We impose this condition on our "variations of the path" between the endpoints. Now since we are allowed to have variations at the two ends which are different, we might have Δt1 = dα dt1/dα be different from what it is at the other end. So only the second terms survives in equation p230 d. Now use the definition of pi in this equation d and you have arrived at 7-37 δ ( ∫L(q,t) dt ) = – [p(t2)(t2) Δt2 – p(t1)(t1) Δt1 ] (7-37) Now of course we want to install this result into 7-36 to get: Δ ( ∫L(q,t) dt ) = – [p(t2)(t2) Δt2 – p(t1)(t1) Δt1 ] + [ L(q,t2) Δt2 - L(q,t1) Δt1] = [L(q,t2) – p(t2)(t2)] Δt2 – [L(q,t1) – p(t1)(t1)] Δt1 = [– H2] Δt2 – [– H1] Δt1 = – H ( Δt2– Δt1) // since H conserved But above we had, ΔA = Δ (∫Ldt + H(t2-t1)) = Δ (∫Ldt) + H ( Δt2– Δt1) and we thus end up with ΔA = Δ ( ∫pdt ) = 0 Since we have used the E-L equations along the way (the dynamics of the system), our conclusion is that the dynamic solution of our problem causes action = 0 ! QED, Now some examples. Consider a non-rel particle T = ½ mv2 = ½ p = ½ A, so action A = 2T. Then our new rule says Δ( ∫Tdt ) = 0. Suppose there are no forces on our particle so T = constant. Then we get that ΔA = T Δ( ∫dt ) = T[ Δt2 - Δt1] = 0. Can write this as Δ(t2-t1) = 0. Now t2 – t1 = "time of transit" for a particular "path". I guess this means a path going from point 1 to point 2 in space. So for the non-rel free particle, the claim is that the solution path is the one which minimizes the time of transit! This will appear in optics somewhere. Now back to Δ( ∫Tdt ) = 0 if there are forces say from a potential V. On page 232 we first express this idea: (ds)2 = (dx1)2 + (dx2)2 +(dx3)2 so that ds is the arclength along the path. We can replace dt with ds as shown top page 232 and we end up easily with 7-40. So our solution path is going to result in this Δ (∫ds) = 0 The solution path minimizes the arclength weighted by this function! Now G wants to generalize this result to more than 1 particle. What do we use for ds in that case? We first assume no t in the r-equations, some word for that I forget. then we know that T is quadratic in the velocities, and the weighting matrix turns out to be the metric tensor he is calling mik here. Then it turns out that the correct generalization to many particles is this: (dρ)2 = dq m dq = dqT m dq where dq = ( dq1, dq2, ....dqn). where for run I use both matrix notation and dyadic notation. The claim is that for normal coordinate systems we use a lot (orthogonal curvilinear), m is diagonal s we then get (dρ)2 = Σi mii (dqi)2 orthogonal curvilinear and in cartesians we would have (dρ)2 = Σi 1 (dqi)2 = Σi ((dxi)2 + (dxi)2 +(dxi)2 Now divide by (dt)2 to get (dρ/dt)2 = (dq/dt) m (dq/dt) = m = 2T // from 232-a This lets you again replace dt not with ds but with the more general dρ and we get Δ (∫dρ) = 0 // for multiple particles (7-44) This is called Jacobi's Form of the Least Action Principle. Now, you can observe that ~ v, speed at any point on the curve. So in general, when we are staying that Δ (∫dρ) = Δ (∫dρ) = 0, we are saying Δ (∫v dρ) = 0. This says that when you integrate the particle's speed along the solution path, you get a minimum. This is basically an average speed along the path and it is going to be as large as possible. Of course one way to do this is to take the shortest path which G refers to as a geodesic of the space. Notice that Jacobi's Form is really about paths in space, time does not appear, so you might get things like orbits from this form. You usually introduce some parametrization variable like θ where (dρ/dθ)2 = (dq/dθ) m (dq/dθ) so dρ = dθ as shown in 7-45. This θ is a spatial thing only, not a time thing. So if we replace dρ with this dθ, our Jacobit result 7-44 becomes 7-45 and if we take the entire integrand to be F, then the claim is the Δ variation is the same as a δ one (I guess since time not involved), and then Δ (∫Fdθ) = 0 = δ (∫Fdθ) and this as usual gives you E-L equations, but now derivatives are relative to θ, not t. Now finally, back to G's comments on page 228. Why is any of the varied paths an allowed physical path? G just fails to make this clear to me and I suspect to other readers. The web points out this alternative way of writing the action ΔA = Δ ( ∫pdt ) = Δ ( ∫pdq ) = 0 = Δ ( ∫pdq ) where now it is little clearer that you are integrating on paths through the multi-particle phase space. The dt form makes this less clear somehow. There is an excellent website http://www.eftaylor.com/leastaction.html#variational which talks about why presentations like Goldstein's are so painful to students. No simple examples are given, and the subject is never really used for anything later in the book. It is just wedged in there for completeness. I looked at some of the papers on the above site and they are most excellent indeed. So once again, I can go and "learn more" about a Goldstein subject if I want, but I would prefer to continue and finish his book. \I do realize that the action is VERY important for many reasons, and it would be a better thing to study later than, say, rigid body motion. Comment on the word "action". As discussed here http://en.wikipedia.org/wiki/Action_(physics), several definitions are in use, and now they have standard names. The action Goldstein uses is now called the "abbreviated action" and is written S0 and we have S0 = A = ∫pdt = ∫pdq " abbreviated action" The action most people now refer to by this word is called S , the "action" or "classical action", S = ∫Ldt " action" or "classical action" // later we learn that S is the same as Hamilton's Principle Function "S" In a situation of constant H, we know from above that A = ∫pdt = ∫(L+H)dt = ∫Ldt + H(t2-t1) which says S0 = S + E(t2-t1) "abbreviated action" // later we learn that S0 is Hamilton's Characteristic Function "W" so they are certainly closely related to each other. Goldstein defines Hamilton's Principle as δS = 0 S = ∫Ldt ********************************************************************************* Chapter 8. Canonical Transformations ( p 237 ) ********************************************************************************* 8.1 The Equations of Canonical Transformation (237). If you could somehow get ALL coordinates qi to be cyclic, then a problem has the trivial linear in time solution shown in 8.2. The purpose of canonical transformation from p,q,H to some P,Q,K is to increase the number of cyclic coordinates! If you just transform from qi to some Qi(qi), that is called a point transformation, but in general Qi(qi,pi) is available. Of course not only to you need some P and Q, you need your new Hamiltonian here called K such that in the new coordinates and with the new Ham, the Ham equations will still be true, in which case we have done a "canonical" transformation. This can be guaranteed by making sure the δ variation as shown in the modified Ham principle is 0 in the new system of Q,P,K. That means that the (p - H) and (P - K) must differ by dF/dt because then the variation integrands will differ by the constant F(t2) - F(t1) which has zero variation under the δ operator. F is called the generating function of the transformation. We then consider four different ways to argument-up the function F, called F1 through F3 as shown page 240. G then considers F1(q,Q,t) first. We get a set of three equations. The first two involve partials of F1, and the third (the same in all four cases) says that K = H + ∂F/∂t. He goes on then to consider the F2 form where we replace Q by P as an independent variable, and the obvious method is to use a Legendre transformation on F1 to get F2 as shown in 8-10, and in this new form we also get three equations. He then continues to the forms F3 and F4 (where he does a double L T on F1), so in each case we get our little set of 3 equations, which is really 2n+1 equations. In the final section, G asks how we might "add" the idea of a transformation on the time, since we know this is going to happen in special relativity. The idea here is to write our standard modified Ham variation thing, but replace time t with some kind of spatial path marker called θ (arbitrarily). Then can interpret the second term as saying there is an n+1st coordinate called qn+1 = t with conjugate momentum pn+1 = -H. Then you can merge the last term into the first as in equation α on page 243. Then when you do a canonical transform to P, Q, you still have your n+1 terms, and Qn+1 = T, Pn+1 = -K. In this case he suggests a generating function G of the F2 type which gives the F2 type equations shown in β p 243. In this manner, then, we are able to treat time t on an equal footing with the other coordinates. Obviously you could consider all four forms for G, but probably he has picked G2 because that one will be used later in the book I get. 8.2. Examples of canonical transformations (244) Example 1: F2 = qiPi results in P=p and Q=q, an "identity canonical transformation" Example 2: F2 = fi(qj) Pi makes Qi = fi(qi), and pi = Pi fi'(qj) and this is the "point transformation" idea where no p are mixed into Q. A linear example is fi = aikqk so fi'(qj) = aij so that in fact pi = aijPi or p = ATP . So here an orthogonal transformation is an example of a point transformation. Example 3: F1 = qkQk results in p = Q and P = -q so basically you get a swap. This emphasizes that there is no real distinction between coordinates and momenta, those now are just "words". Example 4: Start with the harmonic oscillator usual H(p,q) as in 8-28, then use the strange generatort which says F1 = (mω/2) q2 cot Q. Result is that H = ωP which is cyclic in Q. This leads to the usual solution of the HO as 8-30, very strange. In all four of these examples, there was no t dependence in the Fi generator, so K = H. 8.3 The integral invariants of Poincare (247). There are a sequence of such invariants, and the simplest has the form J1 = ∫∫ dq dp where dq dp = [dq1dp1 + dq2dp2 + .... dqndpn ] What we want to prove is that ∫∫ dq dp = ∫∫ dQ dP so J1 is invariant under a canonical transformation. The integrand is rewritten, without proof, as dq dp = Σi dqidpi = Σi Jac(qi,pi: u,v) dudv Goldstein then uses as an example an F2(q,P,t) generator function to show that, in fact Σi Jac(qi,pi: u,v) = Σi Jac(Qi,Pi: u,v) which then proves that J1 is in fact an invariant. This proof uses obvious facts about determinants and the simple facts about the F2 type generator function. It is quite simple, actually, though lots of notation. When this is all done, he says there are in fact a whole series of such invariants. In 8-36 he writes what he calls J2 , but the summation intended unclear. If it is meant to be a sum over i and k, then he is claiming that J2 = ∫∫∫∫ (dq dp)2 but I don't think that is what his Σ means! If that is what it meant, then you would arrive here Jn = ∫∫∫....∫(dq dp)n but it seems to me this is not the same as the Jn shown which is this Jn = ∫∫∫....∫ Πi dqidpi = ∫∫∫....∫ dq1dq2.... dp1dp2...... I scoured the web, but could only find someone who cloned Goldstein's ambiguous results. This is a very obscure topic, so naturally it's not out there, or there may be some other name it goes by. The last item says that your phase space volume is invariant under a canonical transformation. We could probably prove this one using the same methods given earlier, the Jacobian would be a larger one. Hopefully G will only use J1 or Jn in anything further he does. 8.4 Lagrange and Poisson Brackets as canonical invariants (250). This is the Big Section of Goldstein's book which supplies the embarkation point from classical to quantum mechanics. There is lots of trivial derivation stuff here, but let's quote the results and keep our eye on the ball: define the Lagrange (curly) brackets like this: {u,v}q,p ≡ Σi ( ∂qi/∂u ∂pi/∂v – ∂pi/∂u ∂qi/∂v ) Notice that this is in fact the Jacobian ∂(qi, pi)/∂(u,v) which we have already shown to be an invariant under canonical transformations. Therefore, we know that {u,v}q,p = {u,v}Q,P = {u,v} // just write as {u,v} with no labels Now look at special cases for u and v and you find that (the fundamental Lagrange relations) {qi,qj}q,p = 0 => {qi,qj} = 0 {pi,pj}q,p = 0 => {pi,pj} = 0 {qi,pj}q,p = δij => {qi,pj} = δij now define the Poisson (square) brackets like this: [u,v]q,p ≡ Σi (∂u/∂qi ∂v/ ∂pi – ∂u/∂pi ∂v/ ∂qi ) which has the tops and bottoms of the Lagrange brackets all switched. Now, suppose somebody hands you any set of 2n independent functions of the form uj(q, p) . Then Goldstein proves this fact: Σl { ul , ui }q,p [ul , uj] q,p = δij or Σl Lli Plj = δlj or LP = 1 where we define Lagrange and Poisson matrices L and P as shown. Thus, P = L-1 which supports one's feeling that the two brackets are sort of inverses of each other just from how they were defined. We suspect at this point that [ul , uj] q,p = [ul , uj] Q,P but G has not quite shown this yet. Here is how I might show this: let's just add q,p labels to our matrices, then since we know L is invariant, we get Pq,p = (Lq,p)-1 = (LQ,P)-1 = PQ,P So from now on, no need to put any labels on either Lagrange or Poisson brackets. Now let's pick the following special case situation; ui = qi and uj = pj and the overall set of {ul} is just the full set of all the p's and all the q's. Then the LP=1 theorem just quoted becomes [pi, pj] = 0 By doing similar special cases, we get all three results we anticipate: [qi,qj] = 0 [pi,pj] = 0 [qi,pj] = δij and these are exactly the same as the three results we got for the Lagrange curly brackets above. When you put the actual canonical variables p and q inside either brackets, you get the fundamental Lagrange or Poisson bracket relationships, so the above triplet of results is "fundamental". Now imagine F(p,q) and G(p,q). Goldstein first shows the following [F,Qk]q,p = – ∂F/∂Pk [F,Pk]q,p = + ∂F/∂Qk where P and Q are some coordinates which are canonical to q and p (were they not, you could not do some of the intermediates steps). Then he ends up showing this [F,G]q,p = [F,G]Q,P where F and G are any functions you want. But I think I already knew this invariance. I note in passing that the above results are true for any Q, P set, and in particular [F,qk]q,p = – ∂F/∂pk = [F,qk] [F,pk]q,p = + ∂F/∂qk = [F,pk] One more interesting fact: if c is a "constant" not depending on p or q, then [ u,c] = 0 for any u. In a footnote, he says that in quantum mechanics, one will have this "correspondence" i [F,G] Poisson bracket ↔ [F,G] commutator of two operators. for example, [x,px]QM = i [x,px]PB = i Notice that we are probably going to use Poisson brackets more than Lagrange ones because the Poisson ones have our "functions of interest" in the numerator and have the "variables" in the denominator and this is how things normally appear. Probably we will never use Lagrange brackets themselves for anything. 8.5 Equations of motion in Poisson bracket notation (255). Now we continue on this pre-QM avenue, and we combine our "canonical Hamilton's equations" with our knowledge of Poisson brackets. Right off the bat now we have: – ∂F/∂pk = [F,qk] => [qk,F] = ∂F/∂pk so what happens if F = H, the Hamiltonian. Then [qk,H] = ∂H/∂pk = k What happens with the reverse: + ∂F/∂qk = [F,pk] => [pk,F] = – ∂F/∂qk so we get [pk,H] = – ∂H/∂qk = k Thus, I would say the "Poissonator" of either p or q with H "generators" the time derivative of the item in question! G then shows this is true for any function u(p,q; t) and we have this very famous result: du/dt = = [u,H] + ∂u/∂t In this book, the "quantities of interest" u will not have explicit time dependence so we will have du/dt = [u,H] // as you install various quantities "u", you get "equations of motion" Thus, if you can find a u such that [u,H]= 0, then you have found a constant of the motion! This provides a test. The last section here derives the Jacobi Identity which I know well from commutator theory, and I am not at all surprised to find that it is true using Poissonators, but Goldstein dutifully provides a page long proof which is not bad. The application of interest of this Jacobi Identity is very clear: If u and v are both constants of the motion, then [u,v] is also a constant of the motion. This comes from Jacobi just setting w = H. The result seems awfully reasonable. It is called Poisson's Theorem, but one imagines Poisson has lots of theorems in the larger picture. 8.6 Infinitesimal contact transformations, constants of the motion, symmetry (258) Here we just consider a very small transformation from p,q to some P,Q, and we ask what the "generator" G might look like for such a transformation. He uses the word "contact transformation" for some reason, but he just means any canonical transformation. The smallness parameter is called ε. We use the F2 form and have the generator as εG(q,P) [ see p 240] , but to first order in ε we can use εG(q,p). Our usual F2 equations then tell us that δpi = -ε∂G/∂qi and δqi = ε∂G/∂pi which is 8-64, used much later. We could consider ε = dt and G = H, the Ham, and he makes a big deal of saying that we can interpret the transformation two ways. The way of interest now is that in dt, we are going to actually move a little big in phase space, so this is not a virtual transformation, nor are we just relabeling the same point from p,q to P,Q. The Hamiltonian is thus the "generator" of time movement of the system in phase space. I am used to generator being a quantum operator, but here it is the classical generator of a canonical transformation. So here is the big idea: the Hamiltonian can act to transform your system in incremental steps from t back to to where your system had some initial values qi0 and pi0 and in some sense, then, you have moved from q(t), p(t) back to Q = q(0), P=p(0) and now in these new coordinates, all Q and P are "constants". I am very used to this in QM, but here we are doing it in CM with Poisson brackets. Now look at 8-67 and this time we have taken u = H instead of G = H, and G is still just some G. If G commutes (" Poissons", classically speaking) with H, then δH = 0. This is just another way of saying that if a generator of some sort Poissons with H, then H does not change as G is varied. This is what is meant by a symmetry! Two examples are then given. In the first, suppose we have the only transformation change being δq5 = ε. Then from 8-64b we must have G = p5 . First, we see that p5 is therefore the "generator" of a translation in the q5 direction. Second, if p5 Poissons with H, then δH = 0 from 8-67 and p5 is then a symmetry of H. So this example is linear momentum. The second example is going to be angular momentum. If we take G as in 8-71 (G = Lz), we find that we have generated a rotation about the z axis. And if Lz Poissons with H, then it is a symmetry. Again, we would say that G = Lz is the "generator" of rotations about the z axis. So far, we have nothing exponentiated to build up a finite transformation. 8.7 The angular momentum Poisson bracket relations (263). Consider 8-66 δu = ε[u,G] and let's imagine u as some vector function F, and G as our rotation generator n L for rotations about the n axis, so we then have δF = dθ [ F, n L ] . In the case of n= this can be written out as in 8-75 which shows that, for some abstract F, you get both a change in functional form [ hence A' ] and you get a change in value [ vector is rotated] . But certain quantities F will have the property that A' = A, etc, and L is an example of such a quantity. In general, any vector composed from the p and q of the system will have this special property of same functional form. In this case (only) on page 264 G shows that we can write δF = dθ n x F and then the two dθ cancel and we get [ F, n L ] = n x F which is 8-77. This is then the transformation rule for a certain class of "vectors F" under the generator shown. Remember, these are Poissonators, not commutators. In particular, we have [ L, n L ] = n x L and this gives the usual commutatation relations for angular momentum as in 8-80 but again, Poissonators and no i ! It is then no surprise that [ L2, n L ] = 0. Our previous Poisson's theorem says that if Lx and Lz are both constants of the motion, then Lz must also be so. So somehow we have the notions, within classical mechanics, of vector and a scalar objects. Somehow I think this is more easily done within the framework of group theory and representations, but Goldstein has done things pretty much by brute force here in 1950. I wonder what the newer editions do? G mentions that, when he was at Harvard, Schwinger showed him that all these commutator-like results existed within classical mechanics via the Poissson brackets. 8.8 Liouville's Theorem (206). This theorem relates to statistical mechanics and states that, in a statistical ensemble of systems, the density of systems (ie, number of systems in the ensemble per unit volume of phase space) is constant if you look at any specific point in phase space as time mores forward. Because the systems within an ensemble are not perfectly known, they "swarm" around some target point in phase ********************************************************************************* Chapter 9. Hamilton-Jacobi Theory ( p 273) ********************************************************************************* 9.1 The Hamilton-Jacobi Identity for the Hamilton's Principle Function (273). In the last chapter we discussed CT's and part of that was K = H + ∂F/∂t . Suppose you arrange F such that K = 0. Then in the transformed world P,Q,K you would have = = 0 from Hamilton's equations there. This would seem to imply that in this new world, all the P and all the Q are constants of the motion! If we go with an F2 form, then p = ∂F/∂q and our requirement K = 0 becomes H(q,∂F/∂t) + ∂F/∂t = 0 as shown in 9-3. Can we find such an F? This equation is called the Hamilton-Jacobi equation (HJE), and solutions are called S, Hamilton's Principle Function. G notes that since there are n+1 first-order partial derivatives in this HJ equation, there will be n+1 integration constants. One will be an overall adder to S which does not affect the HJ equations, so we then have n constants he calls αi. A solution S will then be S(q, α, t). Remember, no matter what values you gives these constants, you get a solution S. Suppose then we choose αi = Pi which we know are constants in the Q,P,K world. Our F2 equations 8-11 tell us: pi= ∂S/∂qi and Qi = ∂S/∂Pi = ∂S/∂αi ≡ βi So we emphasize things are constants by saying Pi= αi and Qi = βi. Look at this last equation: ∂S(q, α, t )/∂αi ≡ βi You ought to be able to solves this for q(α,β,t) , a process G calls turning the above equation inside out. So he has just outlined a possible path to solve the problem for the qi which is of course what you want. Although it is not useful for the calculation method just outlined, G shows that S = ∫L dt + const So basically S is the true "action" as used in Hamilton's Principle. So this letter S is consistent with the website noted above, which refers to ∫pdq as S0, the abbreviated action, not S. This final fact is not "useful" here for finding S because you cannot integrate L unless you know the q(t), even if you know the form of L. 9.2 A HJE example: the harmonic oscillator (277). We know H(p,q) for the HO, one dimension only. Replace p = ∂S/∂q and write out the HJ 9-3 which is then 9-12. As a solution, try the form shown in 9-13 and you end up with 9-14 which is simpler than 9-12. Integrate to find W and then you know S, at least as an integral. We could do the integral, but we only care about partial of S so why bother. Note: we intend that α = P, as in our general writeup. Ie, we select α to be P. Then β is as shown top page 278, and this time we do the integral to get 9-16. This is of the form ∂S(q, α, t )/∂αi ≡ βi which we want to turn inside out to get the result q(α,β,t) which now appears in 9-17 which we recognize as the HO solution. We then consider the normal starting conditions for a HO at t=0 and find that α = as in 9-18 = total energy E. G now shows trivially why α = E: namely, the HJ equation says H = –∂S/∂t = + α. Of course in our starting situation, we know that β = P = p0 = 0. In the final section, we write out again our integral S having replaced our constants α and β, and we find indeed that S is the integral of L, as expected. So this is a very good example and well worth three pages of the book. Aside for next section: In the L world, if L was cyclic in some qi, then the corresponding pi was a constant. In the H world, as shown page 217 from the Hamilton's equations, this rule is symmetric. Cyclic in pi means qi is constant, and cyclic in qi means pi is constant 9.3 Hamilton's Characteristic Function W (279) We now specialize to the common situation in which H is not an explicit function of time. We define W by S = W(≠t) – α1t so that ∂F/∂t = -α1 and then the HJE says H = α1, so we know that α1 is the conserved total energy E. We then work directly in terms of W and we think of W as generating its own F2 type canonical transformation to a world of P,Q and K. We select αi = Pi and here are the F2 equations for F2 = W: pi = ∂W/∂qi Qi = ∂W/∂Pi = ∂W/∂αi Pi = αi K = α1 = P1 = cyclic in all the Pi and Qi (except P1), so the only non-constant is Q1. On page 281 G writes all the Qi in terms of the βi and all the Pi = αi. ____________________________________________________________________________________ Aside: Equations 9-22 A here are going to be heavily used in upcoming sections so it is good to understand them. They say this: i = ∂K/∂αi [ 8.11b ] = δi,1 Our CT has taken us to a place where everything is cyclic except P1 which says Q1 is the only non-constant coordinate and is in fact linear in t. The solution to the above is Qi(t) = δi,1 t + βi where βi = ∂W/∂αi = the constants shown page 275 9-7 While we are here, we should also explain the next step. Suppose we can form an independent set of constants γi(α). Then we can think of K(Qi,αi(γj)). I think we are supposed to visualize this as some new function I will for the moment call K'(Qi, γi) where now we have P'i = γi as our new momenta conjugate to the same Qi. Our Hamilton's equations are now: i = ∂K'/∂γi ∂K'/∂Qi = – i = 0. The first equation here can be written i = ∂K'/∂γi = ∂K/∂αj ∂αj/∂γi = δj,1 ∂αj/∂γi = ∂α1/∂γi = a constant ≡ νi So this ∂K'/∂γi = νi is the meaning of the page 282 A, G does not show a prime on K (nor later does he put it on Pi). This of course leads to the solution: Qi = νit + βi where we are defining the βi afresh right here. With this all in mind, we can look at the right column on page 283. In my notation, I would have put primes on K and Pi in the case γ ≠ α, but I guess this is just a semantic thing. In the action-angle variable section, we will have an example where γi = Ji. Note added: In this γi world, we have K'(Qi, γi) = α1(γi). so ALL the new momenta P'i = γi appear in our expression α1(γi) for the Hamiltonian. That means in general that NONE of the conjugate coordinates Qi will be constants in time. As shown above, we know that i is a constant νi, so the Qi are then all linear functions of t (in the most general case). ____________________________________________________________________________________ Going back, above we said HJE says H = α1 but this is H(qi, ∂W/∂qi) = α1, so this is what the HJE looks like in terms of W. So, how now do you solve for the original qi(t) and pi(t) ? You have done your CT to the world of almost all P and Q being constants (or perhaps Q linear in time). We can compute the various constants α and β in terms of qi(0) and pi(0) once we find solutions for qi(t) and pi(t). But this takes the usual work, as shown in the HO example. You need to get some expression for W, then you can solve "inside out" 9-22b for the qi and so on. Since t is not involved, we end up with "orbit equations", and we will be seeing examples soon. Finally, G shows that W is none other than the indefinite integral of the abbreviated action S0. At this point, we get a very long table which summarizes how you solve a problem using S or W. The solution methods are not the same, because S and W don't generate the exact same CT. For constant H, you can use either method, but probably the W method is simpler. For general H, you have to use the S method. In the first S method, all the Pi and Qi end up constants γ and β where you could take γ = α as a special case. So, we now have a great formalism, we need now to apply it to something! The HO example was really an example of the W method! Review of the Table: 1. Let's first look at the left side which can be used for H(p,q,t). We find a generating function F2 = S [ S = Hamilton's Principle Function ] which makes K = 0. Then Hamilton's equations tell us that i and i = 0 so that all the coordinates and all the momenta are constants of the motion, so we write Qi = βi and Pi = αi. [ p 283 top left] We then have to compute S by solving the HJ equation H(q, ∂S/∂q,t) + ∂S/∂t = 0. We can always shift from the αi to some γi as noted above. The CT equations pi= ∂S/∂qi tell us nothing new because they are "built in" to the formalism already. But the other CT equations Qi = ∂S/∂γi = βi [ p 284] can be "inverted" to find solutions for the qi(t) and we solve our problem. Again this final equation has the following form: ∂S(qi,γi,t)/∂γi = βi, which we solve to get qi(γi,βi,t). And S = the classical action. The text however provides no time-dependent H examples where this formalism is used. 2. Now we look at the right side of the table for case H(p,q) and no time. { Certainly with our three assumptions outlined in Chapter 1 above [ conservative potential, scleronomous coordinates qi in relation to the rj , and non-holonomic constraints ] we can say that H = E = conserved total energy. Explicit time would enter H say if you had a time-dependent potential, and one might expect this to do work on your system and make the energy be not conserved.} So we write S = W-α1t and work with the CT generated by W [ W = Hamilton's Characteristic Function] . We then have K = H = α1 = P1, so the new Ham is cyclic in all coordinates and all momenta except P1. We have Pi = αi. and Qi = ∂W/∂αi = δi,1t + βi. If we consider a different but independent set of constants γi(αj), then we can restate these results as P'i = γi and Q'i = ∂W/∂γi = νit + βi where now H = K = α1(γi) is non-cyclic in all the momenta P'i and as a result, all the Q'i have this linear time dependence. We have νi = ∂K/∂γi = ∂H/∂γi = ∂α1/∂γi. In this case, we normally omit the primes on the Pi and Qi. We then have to compute W by solving the HJ equation H(q, ∂W/∂q,t) = α1. The CT equations pi= ∂W/∂qi tell us nothing new because they are "built in" to the formalism already. But the other CT equations Qi = ∂W/∂γi = νit + βi can be "inverted" to find solutions for the qi(t) and we solve our problem. Again this final equation has the following form: ∂W(qi,γi,t)/∂γi = νit + βi, which we solve to get qi(γi, βi, νi,t). And W = the abbreviated action. The text does an example of this method using the harmonic oscillator. 9.4 Separation of Variables in the HJE (284) By "separable" G means that you can write W down as a sum of functions Wi each of which depends only on its own coordinate qi, all the constants αj, and no explicit time. So I accept the definition of separability as equation A on page 284, but I do not accept 9-23 because I am unable to derive it nor can I see what the Hi would be. Are we supposed to have ΣiHi = H? I don't think we shall need this equation, and perhaps later I can interpret it. At this point, on page 285, we review again the case of conserved H. We can think of t as a coordinate, and so we "separate" W as shown mid page into W + S2 where S2 only contains t, separating it from the rest. This is just a way to interpret what we have already done. Now suppose we have only one non-cyclic coordinate q1. We "consider" a separable W of the special very simple form shown top page 286 which I interpret to say W = ΣiWi(qi,P1...Pn). Equation 9-25a is our standard result of the previous work. Since H is cyclic in all but q1, the other qi don't appear in H and the pi appear as constants αi, so our equation p180 B is stated as 9-25b. We imagine that we solve this for W1. The other Wi are obtained by trivially integrating 9-25a to get Wi= αiqi. Thus, our full W has the form shown in 9-26. So in fact we end up with W = W1(q1,P1 = ∂W1/∂q1,α2....αn) + Σi≠1Wi(qi,αi) = W1 + Σi≠1qiαi Notice that W1 is a function of q1 and all the Pi (mostly αi), whereas the other Wi are functions only of qi and Pi= αi. The sum above in W gets all the cyclic coordinates! If you think of q1= t and α1 = -H then you can think of the term -α1t ( as the above sum) we used when only t was cyclic as a special case of the sum in 9.26, as in 9-19. Example: Let's do central force V(r) in a plane with coordinates r and φ. Write H as equation C, done, and verified at page bottom. Cyclic in φ, but not r, so W1 will be for the one non-cyclic r. Then 9-27 is application of 9-26 where φ is the only cyclic. The HJE says H(q, ∂W/∂q,t) = α1, but if we install the fact that W = W1(r) + αφφ, the HJE becomes H(r, ∂W1/∂r,t) = α1. It is cyclic in φ. We can solve for W1 as claimed, so Eq D then states the entire W for this problem! Ie, we have computed Hamilton's Characteristic function W(r,φ; α1, αφ). Page 281 already gives us the general solution in the W formalism and so our example we can now compute t+β1 as ∂W/∂α1 which we directly do to get 9-29a. This gives us t = t(r) which we got with lots of work back in Chapter 3, where we used E and l for α1 and αφ. Our second equation β2 as ∂W/∂αφ of 9-22b then gives the orbit equation we got earlier! It is an orbit equation because time does not appear anywhere in it, and φ is sitting by itself so we get φ = φ(r). This example, G claims, shows how fast you can get to a problem's solution in separable cases. Very good! 9.5 Action-angle variables (288) Phase space trajectories (paths) are not something I have dealt with very much. Let's do a simple HO example: q(t) = A sin(ωt +φ) (t) = -Aω cos(ωt +φ) p(t) = -mAω cos(ωt +φ) Then we find that q2 + p2/(mω)2 = A2. The phase space path is thus an ellipse centered at the origin. This would be an example of a closed orbit. We are periodic in both p and q, and with the same frequency, and this is the libration situation. A more general libration path is shown top left of page 289, again closed. Aside: For the above HO I computed J = pdq = p dt = πmA2ω. The meaning of the line integral is the area of the enclosed region, see my separate doc on that subject, and this is true even for arbitrary convex loops. The next example is the simple pendulum as analyzed on page 289. For low energy, it swings back and forth and we get libration. Probably for very small energy, the phase space path is an ellipse since same as HO. As you increase energy, you eventually get just enough to let it be perfectly "up" and swing back and forth between this as a limit on each side. The pendulum in this case is not doing small motion and so is not a HO but something messier, and the path is shown to have cusps at π and -π. It would be a good problem to derive the shape of this curve and all those on the inside. If you increase E still more, it will go around forever and θ increases without limit, and you then get a "rotation" type phase space path. Another example is shown page 289 top right. The word "rotation" is apropo since θ keeps increasing. Obviously the rotation orbit is not closed. In a multi-variable problem, you can think about the orbit or phase-space path in each separate piqi plane. The overall motion is "periodic" if it is periodic (libration or rotation) in all of these separate planes. A good example is the 3D HO with non-rational ratios. In each pq plane we get a simple elliptical closed motion with a closed orbit, while in the actual 3-space the orbit never closes. So be careful in what space you define an "orbit". This last situation is an example of conditional periodicity. Now, we suddenly define, for each coordinate/momentum pair, Ji = pidqi which is integral once around the closed path for closed orbit, or over one "cycle" for the rotation case. For cyclic where p = constant, curve is a straight line as per top 289 and we arbitrarily define a cycle as Δq = 2π, probably since this is often an angle. Notice that Ji has units of angular momentum which motivates the J symbol. Clarification: see separate doc on line integrals of this type and why they give area inside loop. Next, we tie this notion of periodic phase space orbits and the definition of the Ji with our last section on separable HJE work when H not a function of t (the right column of the big table). We assume that we have found some coordinates where the W function is separable, so we can write 9-32, which is really an equation for the phase space orbit pi = f(qi). So throw this into Ji and get 9-35. You see that Ji is a function of the constants αi so is itself a constant, and we are going to use the Ji's as the γi's of our previous Aside. That is, imagine that P'i = Ji . Now, instead of calling the coordinate conjugate to the Ji by the name Q'i , we call them wi . Since the Ji are angular momentum in dimension, the wi must have dimensions of angle. Ji = pidqi dim = ang mom, looks like abbreviated action = "action variables" wi = ∂W/∂Ji like Qi = ∂W/∂γi in right side table page 284 top = conjugate to Ji with dim - angle = "angle variables" = νi(Ji) t + βi ! [ since i = top right table page 284 ] I would say at this point that you could compute the change in wi over one libration period τ to get Δwi ≡ wi(t+τ) - w(t) = νi(Ji) τ so if we knew Δwi and νi, we would know the frequency of libration or rotation, namely frequency = 1/τ = νi(Ji)/ Δwi. We can compute Δwi as shown bottom of page 292 and we find it is 1 if we integrate around one cycle in the i-plane (and not surprisingly we get 0 if we integrate around a cycle in some other plane). I was perhaps expecting Δwi= 2π since it is an angle, but it must be an angle measured in "cycles". We therefore have learned that frequency = 1/τ = νi(Ji) where νi(Ji) = ∂K/∂Ji top right side 283 So we seem to have a procedure here: (1) compute W(α) from the HJE; (2) use that to compute J(α) from 9-35: (3) invert that to get H = α(J); (4) then ν = ∂H/∂J. G now does the HO as an example, starting bottom page 293. The first step is to compute W, but we already did that back on page 277. Notice that α appearing in W is the α of H = α and of S = W – αt . The second step is then to compute J from 9-35 which is a line integral as shown page 293 bottom. When written with the square root, it is implied that you take the + branch on the top of the ellipse and – on the bottom, but this is swept under the rug. The θ parameterization removes this confusion allowing a simple computation of the integral which gives J = 2πα which we then "invert" to get α = (J/2π) = H, and in the last step ν = ∂H/∂J = /2π = ωHO /2π as expected. The main point is this: Even in a messy problem we cannot solve or don't want to solve, we might be able to easily find W, compute J, invert it, and learn the frequency of periodic motion of our solution! We might not know the r(t) or even the orbit equation φ(r), but we will know the period. That is the point of action-angle variables! We take our original problem and do a canonical transformation to a new system which has K= H, and Pi = Ji and Qi = wi = νi t + βi and νi is the resulting frequency. Comments: this is not a simple subject because we are building higher levels of a huge towering construction involving the Hamiltonian, Hamilton's equations, canonical transformations to coordinates where almost everything is cyclic, then we separate out time and define the W function, then we require that it be separable, then we go to these action-angle variables as the last step. 9.6 Further properties of action-angle variables (294) If we stay in a particular qk-pk plane, and if we know that the motion is periodic with frequency νk, then we know that the motion can only be composed of this fundamental frequency and all harmonics. Thus, we can expand as in 9-45a: qk(t) = Σj aj e2πijwk where wk = νkt + βk The inverse is given by 9-47 aj = dwk qk e-2πijwk Here is my proof of this fact, aj = dwk qk e-2πijwk = dwk [Σj' aj' e2πij'wk] e-2πijwk = Σj' aj' { dwk e2πi(j'-j)wk } = Σj' aj' δ(j'-j),0 = aj For a rotation, assume that on each cycle qk increases by some fixed amount q0k. Then q'k = qk – wkq0k increases by 0 on each cycle, since wk increases by 1 per cycle. So q'k is periodic and we can do the expansion in 9-45b just duplicating what we did in 9-45a, so fine. If we combine the motion on all the subplanes, and we then move back from the n generalized coordinates qi to our actual Cartesian coordinates xi, we will get a mix of all the planar frequency effects and we then have the Fourier series shown in 9-48 which, in general, gives a non-periodic motion in Cartesian space, although we were periodic in all of the sub-planes. This situation was called conditional periodicity earlier, and now G also calls it multiply periodic. Before continuing, on page 296 we have a little digression. Suppose you define W' as shown. If any qi= wi does a complete cycle, we know its characteristic function W increases by Ji because this is what 9-35 on page 291 tells us. This is one form of the definition of Wi. A complete cycle for a wi means it increases by 1. Thus, the function W' = W(q,p) - wJ is periodic in each wi and you could if you wanted Fourier expand this thing the way we did the xi on page 295. G comments that W' as so written is a Legendre transform thing, and you get W'(q,w) which is the F1form. This W' is another transformation that maps you from q,p to w,J, but I don't see why this is useful yet. Now on p 297 we start into the whole issue of how you handle things when at least some of the frequencies in some of the sub-planes qk-pk are harmonically related. You might have 3ν1 + 5ν2 = 0, say. If some frequencies are related like this, the problem is "degenerate". If there are m ≤ n-1 relations of this form, you have m-fold degeneracy, and if m = n-1 you have complete degeneracy. G uses the strange word commensurability to refer to the conditions of degeneracy. The problem is treated on page 298 and I change to a better notation here. The idea is that we are going to do a transformation from w,J to some new w',J' where all the m degenerate frequencies ν' will be 0, and this means that none of the corresponding J' will appear in the H' Hamiltonian, which I guess is going to make it a simpler Hamiltonian to deal with. This is all proven right here: First, write Σi=1n jki νi = 0 for k = 1...m // these are the harmonic relations, assume m of them Here we are using some elements of a square nxn matrix which I called j. The above equation appears as this way, and I have arbitrarily added a chunk of 0 and some identity matrix to this thing: There are only m zeros in the right column, only m equations Now consider the proposed F2 transformation which we can write as: F2 = J' j w = (w j ) J' = 9-54 This transform does this, where I am showing the traditional and current names of things qi pi → Qi Pi wi Ji w'i J'i Since the above F2 has the form F2 = f (w) J' = f(q) P , we see that this matches the form of 8-19 on page 244 and thus this is indeed a "point" transformation. Next, our 8-11b equation tells us: Qj = ∂F2/∂Pj => w'i = ∂F2/∂ J'i = ∂ (J' j w) /∂ J'i = (j w)i = 9-55 => w' = j w // identity for the last non-harmonic ones This shows how the q's are shuffled into Q's by this point transformation. Notice how the last terms in this vector j w use the 0 and 1 portion of the j matrix and the bottom part of the vector w' is just a replication of the original w vector. Next, our 8-11a equation page 241 tells us: pj = ∂F2/∂qj => Ji = ∂F2/∂ wi = ∂ (J' j w) /∂ wi = ( J' j )i = ( jT J')i => J = jT J' 9-57 // identity for the last non-harmonic ones Notice that we are not claiming we can invert matrix j. Probably in general we cannot. Now we can look at the new set of frequencies νi' ν' = ' = j = j ν This is the full matrix version of what we showed graphically above. We can see at once that the first m components of the ν' vector are 0 due to our degeneracy relations above, while the last n-m are not. So, the point is that this "point canonical transformation" has mapped the degenerate frequencies all to the value 0, and has left the non-degenerate ones "as they were". Now consider this Hamilton's equation (and K = H as usual since no t dep) i = ∂H/∂Pi => 'i = ∂H/∂J'i = ν'i = 0 for the first m values of i This means that K is cyclic in the first m values of the generalized momenta J'i. 9.7 The Kepler problem in action-angle variables (299) You have to imagine the little 3D picture of a particle at location r,θ,φ. There is a lot of derivation work to do here, and a fair dose of confusion in the names of symbols. Imagine that a Kepler ellipse lies in the grey plane shown at the bottom of page 107, and the position of a planet is described by θ, φ (to establish the plane) and then ψ and r (planet position in the plane). Part I: The basics First we know we can make our little orthogonal system , and and I have page of useful stuff in my math binder. Ambiguity arises in the definition of physical versus canonical things like velocity and momenta. I will use upper case today for physical, and lower case for canonical. Then we can say Vr = Vθ = r Vφ = rsinθ The last two are very obvious. The radius for the θ circle really is r, but for the φ circle is rsinθ. We then have these physical momenta Pr = m Vr = m Pθ = mVθ = m r Pφ = mVφ = m rsinθ and of course P2 = Pr2 + Pθ2 + Pφ2 = m2 [ 2 + r22 + r2sin2θ 2 ] And so H = P2/2m - k/r and we get the desired result H = (m/2) [ 2 + r22 + r2sin2θ 2 ] - k/r // see 9-58 While we are using these physical things, we can do one more thing: L/r = x P = x [ Pr + Pθ + Pφ ] = Pθ – Pφ where we use our reference sheet for the cross product of the unit vectors. Therefore L2/r2 = Pθ2 + Pφ2 = (m r)2 + (m rsinθ )2 = m2r2 [2 + sin2θ 2] Now we can rewrite H as follows, using the above, H = (m/2) [ 2 + r22 + r2sin2θ 2 ] - k/r = (1/2m) [m2 2 + m2r22 +m2 r2sin2θ 2 ] - k/r = (1/2m) [m2 2 + L2/r2 ] - k/r = (1/2m) [ pr2 + L2/r2 ] - k/r where I have used ahead of time the canonical momenta pr found below. Thus, we have derived equation C on page 300 except L2 is presented as p2 which is immensely confusing, but OK, we shall see soon why this is done. Now enough physical stuff, we turn now to canonical quantities. We have vr = vθ = vφ = where I am just defining these as time derivatives, but they are not used anywhere and are dangerous. Now here is the Lagrangian L = T-V = (m/2) [ 2 + r22 + r2sin2θ 2 ] + k/r We can then directly compute the canonical momenta from p = ∂L/∂ to get pr = m pθ = mr2 pφ = m r2sin2θ // 9-59 Pr = m Vr = m Pθ = mVθ = mr Pφ = mVφ = m rsinθ so Pr = m = pr Pθ = mr = pθ/r Pφ = mrsinθ = pφ/(rsinθ) Notice that we do not have p = mv using canonical quantities. No one talks about canonical v. I replicated the physical P line from above for comparison, and you see how the physical and canonical momenta are related. While we are here, let's re-express L2 as follows, continuing from above, L2/r2 = Pθ2 + Pφ2 = (m r)2 + (m rsinθ )2 = m2r2 [2 + sin2θ 2] = (pθ/r)2 + [pφ/(rsinθ)]2 = (1/r2) [ pθ2 + pφ2/ sin2θ] so we can then identify L2 = [ pθ2 + pφ2/ sin2θ] Looking back at 9-60 we have the following, which will be used a few lines below: H = (1/2m) [ pr2 + L2/r2 ] - k/r which is page 300 C, but he uses p2 in place of L2 Now we can solve these for the dotted quantities and put back into H of 9-58 and we get 9-60 which is H in terms of the canonical coordinates and momenta H(qi, pi) = (1/2m) [ pr2 + pθ2/r2 + pφ2/(r2sin2θ) ] - k/r = (1/2m) [ pr2 + L2/r2 ] - k/r where this last form appears on page 300 as equation C. At once we see H is cyclic in φ so pφ = constant = αφ in a minute here. To get the Hamilton-Jacobi equation, you use pi = ∂W/∂qi and set H = α1 = E. This trivially gives 9-61. We then try to separate W into its three parts to see if things will work out OK as in 9-62. This means we replace each W with Wi in its derivative in 9-61. First, the φ part requires that ∂Wφ/∂φ = constant = αφ ( which is pφ). Reading top of page 300 makes us agree with 9-63b where another constant is defined, αθ, which contains αφ as shown. Our purely radial equation is then 9-63c in terms of our three constants, one hidden. Now G wants to associate each of our "constants" with something "conserved". We already know that pφ = conserved = αφ . Canonical momenta to an angle is always an ang mom, so we find that angular momentum about the φ axis is conserved. Of course we know that L overall is conserved, but ignore that fact for the moment. Looking back we see we have αθ2 = [ pθ2 + pφ2/ sin2θ] = L2 ≡ p2 so αθ = L = p = the total momentum magnitude And of course α1 = E is also conserved. So E, Lz and L2 are conserved quantities. Comment on the alphas: Let's back up a moment and remember a few things: pi = ∂Wi/∂qi says pr = ∂Wr/∂qr pθ = ∂Wθ/∂θ pφ = ∂Wφ/∂φ It happens that in our H, φ does not appear, so pφ is a constant called αφ. The other two pi are NOT constants. You have to solve the page 300 equations for them. The solutions are shown top of page 301 as the integrand arguments. The alpha called αθ is not particular pi or Pi, it is L which is called p. When we do our HJ transformation below, we end up with Pi = Ji and Qi = wi. It turns out that αθ is a linear combination of the J's as shown bottom page 301. Part II: Computing the Ji action variables These are stated bottom of page 300 and are correctly written out top of page 301. As usual, you have to about branches of the square root at least in the last equation. Computation of Jφ is trivial and we get Jφ = 2παφ. Computation of Jθ is going to require a trick of Van Vleck. If we look back to our "planar" central force work on page 60, we find that we could write T = (m/2) [2 + r22 ] // as in 3-6 where to avoid confusion, we change that name of that page 60 angle from θ to ψ. But in line 3-8 we shows that m r2 = L, the total angular momentum magnitude in this problem, which we are agreeing for the moment to call "p" or better , pψ. The simple "p" notation reminds us that this canonical momentum pψ is the same as the total angular momentum L, so a subscript is confusing perhaps. Thus from the planar analysis we know that T = (m/2) [2 + r22 ] = (m/2) [2 + (r2)2/r2 ] = (m/2) [2 +p2/r2 ] But I am off the track. Recall from page 54 that, with conservative forces, we know 2T = Σipii So let's write this in our two systems: 2T = pr + pθ + pφ 2T = pr + pψ = pr + p = pr + p so comparing we have pr + pθ + pφ = pr + p which says pθ + pφ = p or pθdθ + pφdφ = pdψ This is the fact we need to have, a pretty fast derivation, might take a while geometrically. We then have pθdθ = pdψ – pφdφ Therefore, Jθ = pθdθ = pdψ – pφdφ = 2πp – 2π pφ = 2π(αθ – αφ) So all this last stuff was a little trick to do the integral 9-65b without using an integral table. Computation of Jr is going to be more serious. It is written out in 9-68 and we have now to do this integral. Another trick shall be invoked, this one from Sommerfeld, our friend. Since E < 0, let's write the square root this way: = (1/r) = (1/r) = (1/r) = (1/r) =(/r) Notice that r1r2 = C/a. When r is in the range r1 < r < r2 the above square root is real. This is good since this is basically pr during an orbit. Here is what the branch situation is in the complex r plane: We arrange to do the outbound path on the + sheet, and the return inbound path on the – sheet. The function is real analytic on the interval. This is not the pr-r space orbit picture, this is the complex r plane by itself. Because the contour is as shown, we can deform the branch cuts so that they meet and then the entire contour will be on the – sheet, like so: Now we can deform the contour and pick up the residue at r=0 and another at r=∞. This is shown somewhat poorly at the top of page 302 where the regions labeled positive and negative square roots don't make sense to me. Those are sheets, not line segments. But the contour senses are correct for the deformation. So, now replace -a with A and define f(r) = – // the – sheet function At the origin, a Laurent series expansion would show the behavior f(r) = – /r so CCW residue there would be just – , just as G claims page 303 #1. Now the great circle part can be handled as is, or we can convert it to a contour around the origin in z = 1/r space. Doing this changes the rotation sense of the contour, but then dz = -dr/r2 makes a cancelling minus sign, so you get dr = – dz/z2 and we end up with formula 3 including the correct sign shown (same sign as 1). This has a double pole at the origin, and the rule for residue there is f'(z=0) where f = – so f' = – (1/f) (B-Cz) → – B/. Now both contours required are backwards, need CW as shown, not CCW, so get extra minus on both residues, and we end up with f(r) dr = 2πi [ + + B/] // as claimed in 5. We then identify A = 2mE B = mk C = (Jθ+ Jφ)2 / (4π2) so integral becomes Jr = f(r) dr = 2πi [ + + B/] = 2πi [ i- i mk/] = – 2π + 2πmk/ = – (Jθ+ Jφ) + πk/ // as in 9-69 where we have always used this idea: = +i . Obviously more could be said about that, but since we got the right answer, fine. Now solve for E to get (Jr + Jθ+ Jφ)2 = π2k2 2m/(-E) H = E = – 2π2mk2/ (Jr + Jθ+ Jφ)2 as in 9.70 This result seems rather amazingly symmetric to me. Well, remember how you get the frequencies: νi = ∂H/∂Ji so the fact that all appear as a simple sum means that the three derivatives will all be the same, and so all three frequencies will be the same (as we expect for Kepler orbits). Next, we are going to do the extra transformation to Q,P = w', J' to zero out two of the frequencies νi. We use exactly the little matrix formalism presented earlier and I have verified every equation on page 304, mostly there in pencil. We know that we get ν1' = ν2' = 0 and ν3' = ν3. As shown on page 300 bottom, the order of things is 1,2,3 = φ, θ, r (probably he did this to get the tough Jr to be treated last). The new J' are now displayed. We know that the first two cannot appear in H because ν'i = ∂H/∂Ji (p298) = 0 for 1,2. This is seen to be correct, and only J3' appears as shown 9-75. At this point, G enters a 3-page commentary section with very important remarks which I will try to put in my own words. First, in the action-angle world, you end up with wi = νit + βi. I failed to realize how important this statement is. No matter how messy your original coordinate functions ri(t) might be for an exact solution to the problem, these "angle" coordinates at most do simple linear motion in t ! It is rather amazing that you can transform a tough problem to a set of "coordinates" where the motion is nothing more than linear in time. When we get to the final world here with wi' = νi't + βi' using F2 of 9-54, we have ν1' = ν2' = 0, which says that w1' = β1' and w2' = β2' and both these angle variables are constants for Kepler! G points out that since J1' = J1 = Jφ, we can think of w1' as some fixed angle like the φ shown on page 107, the position of the line of nodes. Similarly, J2' = Jφ+ Jθ = 2πL = 2πp is related to total angular momentum L, so its angle variable w2' must be a ψ angle in page 107. It is a "fixed angle" β2' which one often takes to mark the position of the perihelion (closest approach) of an orbit. Finally, since J1'/J2' = pφ/p = cosθ so we get a connection to the third polar angle. So his point is that the two wi' and this ratio are constants which are equivalent to the two Euler angles which define the motion plane, and the third ψ which tells you a point in that plane. Action-angles are currently in use for working with planetary orbits. Called Delaunay elements, I think they involve only r and ψ and you work in the orbital plane, I have seen F and G mentioned. When you involve complex orbital perturbations, you can characterize them as causing the normally static angle variables to start moving slowly according to wi' = νi't + βi' where the 1 and 2 νi' are no longer exactly zero due to the perturbation. So this is even currently a significant role for action-angle tools. ( Maybe since 1950 computers have replaced all this analysis? ) In the Bohr atom, it turns out that the path to quantization is simply to say that the actions Ji are quantized in multiples of Plank h. Sommerfeld's "royal road to quantization". Thus, if we take J3' = nh, we get the Bohr formula for energy levels showing n as the principle quantum number. Very cool. Then since J2' is related to total angular momentum l, we can write J2' = lh and then quantized l gets into our hydrogen atom model. In planets, this is called k and affects precession of the perihelion. Then add a fixed B field for Zeeman effect and you get J1' = mh and that is your magnetic quantum number. At each step, you break a degeneracy and you see it all happening in these action-angle variables. In the early days of the old Bohr model, everyone was using action-angle as a "daily tool". But then when the full quantum theory was discovered, and always required perturbation theory, it was always easier to do it in quantum and not classical mechanics, so action-angle went by the wayside. It remains now in astronomy and in text books like this one! 9.8 Hamilton-Jacobi theory, geometric optics, and wave mechanics (307) Preamble: Before starting here, consider a set of ellipses made simply by scaling z, so we have z = f(x,y) = ax2 + by2 For example, if we set z = 1 and a = 1/A2 and b = 1/B2 , we get the "standard ellipse" with semi-major axis A and so on. In 3D, our function z = ax2 + by2 is an "oval cup" made of every larger ellipses stacked on top of each other (called a paraboloid). If we look from above and project all these ellipses onto the x-y plane, we get our family of concentric ellipses of interest. We can then examine the curves which are normal to all the ellipses. Here is a picture: It turns out that you find the normal curves by solving the equation dy/dx = fy/fx . In our case, the solution to the equation is y = K xb/a where constant K determines the location of the curve. Some of these are drawn above for K = 2, 1 and ½ in the case b/a = 2. You can roughly see in this crude drawing that all intersections are right angles. It took me several days (!) to figure out the above little detail, see "surface geometry questions.doc". I am now going to skip to the optics discussion starting page 310 equation 9-87 which is the wave equation we all know and love. The plane wave solutions are written for constant index n. At this point, due to my unfamiliarity with optics, we define "geometric optics" to be the system of equations you get when you assume that n varies, but varies slowly compared to light wavelength λ. Probably this subset of "optics" is not going to know about diffraction fringes, but will know about Snell's law and so on. I have a whole unread optics book by Born and Wolf, by the way. Another gaping hole in my physics knowledge. So, suppose we imaging that if n varies slowly, we will still have the "plane wave form", but things will vary slowly in phase and amplitude. That is, assume A and L are real and the form is φ(r,t) = exp [ A(r) + iko ( L(r) - ct) ] where ko = ω/c as if we had n= 1 (vaccuum) Meanwhile, our wave equation solution for constant n has this form φ(r,t) = exp [ 0 + iko ( n r - ct) ] normally seen as exp [ i (kr – ωt) ] k = nko We would then say that for constant n, the attenuation factor A(r) = 0 and L(r) = n r = nz for example, if k is in the z direction. You can see that L is some kind of phase thing. In the plane wave, we would write phase = (kz-ωt) and plane wave propagation of a particular "wave front" has (kz-ωt) = constant. We can also write this as (nz-ct) = constant since r= z for our special case. We then have some kind of "phase wave front" which is moving along at velocity u = c/n. To see this, consider the left figure: Again, we are looking at a "surface" nz – ct = constant and it moves to the right for our plane wave. What is going to happen when n varies slowly? The amplitude of φ will vary due to A(r), but more importantly for our purposes, the "wave fronts" will be something like that shown on the right. These are now surfaces on which L-ct = constant phase. We know that L(x,y,z) = ct => a changing shape as t changes, so that is why we have drawn the two surfaces with different shapes. From our several days spent on surface geometry, we know the following about the situation on the right, where imagine that t = dt = very small: We could define phase P such that P(r,t) = L(r) - ct . Our left wavefront then has the equation P(r,0) = 5. The right wavefront is then P(r,t) = 5. Of course this is a differently shaped curve, but all the wavefronts have the property that P = 5, surfaces of "constant phase". As we let time go by, the surface P(r,t) = 5 takes on the various shapes shown. Thus, at time t, P(r,t) aligns with the surface L = 5+ct. the normal at some point on the t=0 ( L=5) surface is given by n = 3DL(r)/ |3DL(r)|. This of course points do the next wavefront at time t=dt. if the distance between the wave fronts is ds, we know that the wave front velocity at some point is given by u = ds/dt When dealing with the case g(x,y,z) = 0, we said that g(r+dr) = g(r) + 3Dgdr = 3Dgdr = dg. So if we go off in the normal direction away from g, we have dg = |3Dg(r)| ds since dr = ds. If we think of g(r) = L(r) - 5 in our left wavefront above, then we have dL = |3DL(r)| ds. But we know that dL = cdt because our wavefronts are surfaces where (L-ct) = 5. Combining the last three items, namely that u = ds/dt and dL = |3DL(r)| ds and dL = cdt, we find u = u(r) = ds/dt = [dL/|3DL(r)|] / dt = c/|3DL(r)| u(r) = { c/|3DL(r)|} = { c/|3DL(r)|} [ 3DL(r)/ |3DL(r)| ] = c3DL(r)/ |3DL(r)|2 So this is pretty good, we know the wavefront velocity at any point as a scalar and as a vector! The question now arises in geometric optics: How are we going to solve for the function L(r) ? By starting with the wave equation 9-87 and inserting our form for φ involving A and L, G shows us that we get the equations shown bottom of page 311. But in the geometric optics limit where δn >> λ, G shows that we can neglect the first two terms in (9-93a), and we end up with this being the equation we need to solve for L (3dL(r))2 = n2(r) or |3DL(r)| = n(r) Notice this is not Laplace's equation, it is this: (there are no second derivatives) (∂L/∂x)2 + (∂L/∂y)2 +(∂L/∂z)2 = n2(r) This is called the eikonal equation and I have never worked with it, but I can see that it "controls" geometric optics. The function L here (called S in Born and Wolf) is called "the eikonal". I cannot locate H. Burns on the web, Hamilton is the Hamilton. OK, so I get the picture. In geometric optics, you need to solve this eikonal equation for L(r), and then you can draw your nice curved wavefronts and look at your normal paths I called h(x) in my surface notes. __________________________________________________________________________________ Note Added: Let's go back and review the above a bit, because things are still hazy. I want to write the little expo φ lots of ways now and "gather up" all the little micro facts on this subject: φ = exp[ i (kr - ωt)] // familiar to me But we have k = nk0 where n is some index ≠ 1, and k0 is what k would be if n= 1, ie, in vacuo. So now φ = exp[ i (n k0r - ωt)] = φ = exp[ i (n k0 0r - ωt)] = exp[ i (k0 {n 0r} - ωt)] Now it is the curly bracket thing that is defined to be L, so we do one more step φ = exp[ i (k0 {n 0r} - ωt)] = exp[ i (k0L(r) - ωt)] L(r) = n 0 r Now in optics problems, ω is a constant. Think of Snell refraction at a boundary between medium and vacuum. On both sides, ω is the same, so we can say ω = ω0 and ν = ν0 if we want. On each side we would have [ u = phase velocity symbol, wavefront speed ] u = λ/T = λν u0 = λ0/T0 = λ0ν0 but ν = ν0 so u/λ = u0/λ0 and of course we have k = 2π/λ and k0 = 2π/λ0. and u0 = c ω/k = λ/T = u ω/k0 = λ0/T0 = c ω = 2π/T and ω0 = 2π/T0 Now let's check our eikonal equation for constant index: ( note that (Ar) = A for constant A) L(r) = { n 0 r } = n 0 = > (L(r))2 = n2 = eikonal equation. Here is one more way to write φ where we use ω = ck0 then we get φ = exp[ i (k0L(r) - ωt)] = exp[ i (k0L(r) - k0ct)] = exp[ i k0 (L(r) - ct)] And let's do yet one more variant of the above ωt form φ = exp[ i (k0L(r) - ωt)] = exp{ i 2π( [ k0/2π] L(r) - [ω/2π]t)} = exp{ i 2π(L(r)/λ0 - νt) } __________________________________________________________________________________ So, what is the connection between this optics stuff and our functions S and W of the Hamilton-Jacobi classical mechanics world? Well, our general HJE says H(qi, pi = ∂S/∂qi) + ∂S/∂t = 0. 9-3 on page 274 If H = E, constant, then we know that ∂S/∂t = E. This leads us to define S = W(qi) - Et Notice that ∂S/∂qi = ∂W/∂qi for the normal coordinates. Thus, the HJE in terms of W becomes H(qi, pi = ∂W/∂qi) = E // E is usually called α1 by G Now, suppose you have a single particle in a potential. Then we have H = p2/2m + V and the HJE then says this: (W)2/2m + V = E or (W)2 = 2m(E-V(r)) So compare this to our eikonal equation of geometric optics: (L(r))2 = n(r) (W(r))2 = 2m(E-V(r)) Since the forms are exactly the same, we can carry over our "wavefront" discussion from the optics world to the W world. I can relabel my picture above as follows We are thus making the following replacement: W-world optics world S(r,t) = W(r) - Et P(r,t) = the phase = L(r) - ct W(r) L(r) E c f = 2m(E-V(r)) = 2mT f = n2(r) (W(r))2 = 2m(E-V(r)) = 2mT (L(r))2 = n2(r) Finally we understand Goldstein's opening picture on page 308. The surface S(r,t) = 5 (he calls it a) is going to move to the right as t increases, taking the position of the various W = 5 + Et. So somehow associated with our little mass particle in a potential is a function S(r,t) which has wavefronts like S(r,t)=5 which move in time and take on the shapes of the surfaces W = a + Et. What on earth does this mean? On page 309 more facts are uncovered. the velocity of the wavefront is u = c/|3DL(r)| which becomes u = E/|W| = E/ But we can replace E-V = T = ½ mv2 for our particle, and then u = E/(mv) = E/p. This shows the strange result that the slower your particle goes (speed v), the faster the S=constant wavefronts move (speed u). From page 280 we know that pi = ∂W/∂qi so that p = W. But we know that W = normal to our wavefronts, so the particle momentum must follow a path like my 2D h(x) which is normal to all the wavefronts. Such a possible path is shown by the arrow in the page 308 figure. Say this again: what I call "the normal curves" to the wavefronts are each a potential spatial path for our particle's momentum p. So imagine an animation of the page 308 figure where the particle moves slowly along the trajectory, but as it does, the S = a specific surface moves very quickly to the right through its positions, passing the particle by. Somehow, we suddenly have some kind of "wave theory" of the motion of a classical particle. At the same time, we have some kind of "particle trajectory theory" for light. And this has nothing to do with quantum mechanics. What about least action? In mechanics on page 231 we showed that the "abbreviated action" = So = A was given by the integral of p dt, but for a particle this is mv2 = 2T so we have integral of 2Tdt. But we know that T = ½ m v2 and v = ds/dt along path, so find dt = ds. We can then write our action as an integral of 2Tdt = 2T ds = ds, as shown page 232 equation 7-40. So, in classical mechanics our principle of least action means to extremize ∫ds , probably we want to minimize this integral to solve our problem. From our table above, we see that in optics it is n that corresponds to , so geometrics optics solutions will minimize ∫nds. This is Fermat's Principle of geometric optics. Big Question. This question could have occurred to people of the late 1800's who knew all the stuff described above, but who never heard of quantum mechanics. We have seen that "classical mechanics", as embodied in the Hamilton-Jacobi formalism, can be fully described by a function W which seem to define some kind of wavefronts which move in time, and we have seen that these waves satisfy a differential equation which corresponds to the "geometric limit" of E&M, it its optics incarnation. We have seen this particle-wave duality in classical mechanics, but the particles are the "senior partner" since nobody knows what to "do with" those waves in classical mechanics. The question is this: is there some more general theory of mechanics which corresponds to the full E&M theory? At this point, we need to go back to page 3-9 and derive some simple facts: For our single particle, = p = mv. We already showed that the wavefront velocity is u = E/|W| = E/ = E/p. p = W which we get either from the HJ formalism, or from above. _________________________________________________________________________________ Question Added: For these mysterious mechanics-associated waves, what is the frequency and the wavelength of the wavefronts? The frequency is just something we assume, ν = ν0 as noted earlier in the Snell comment, and then λ = u/ν = E/(pν) where p = mv still. NOW, suppose the correct "wave function" associated with a classical particle, the thing we have been talking about all this time, involved a constant = h/(2π) and we had this form (just suppose), ψ(r,t) = exp[ i S(r,t)/] where ψ(r,t) is the thing that has the wavefronts. Then we could write this as ψ(r,t) = exp[ i 2πS(r,t)/h] = exp{ i 2π [ W(r) - Et] /h} = exp{ i 2π [ W(r)/h - (E/h)t]} Let's now compare this to our last φ form above, namely φ(r,t) = exp{ i 2π(L(r)/λ0 - νt) } These two forms align if we were to make this assumption W(r)/h ↔ L(r)/λ0 E/h ↔ ν Just above we said that the frequency of the classical mechanics waves "could be anything" and then λ depended on what we selected. Suppose we suddenly assume that the frequency of the classical mechanics waves is not arbitrary, but is in fact given by ν = E/h. The thing above with ↔ is just a correspondence between two analogous theories, but now we are making a very much stronger claim, that the frequency of the classical waves really is ν = E/h. What would follow from this assumption? Well, what is our λ ? λ = u/ν = (E/p)/(E/h) = h/p which looks a lot like the "deBroglie wavelength" of a physical particle. Now go back to our supposed form for ψ(r,t) given above which we rewrite here as ψ(r,t) = exp[ i S(r,t)/] = exp{( i/) [ W(r) - Et]} The wave equation for φ contains the index n, but we know n = c/u = c/[E/|W|] = |W| c/ E ––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––– Therefore ( ignore this section, it went the wrong way) (n/c)2 = |W|2 / E = (W)2 /E2 and then the wave equation for φ, when applied to our ψ, becomes 2ψ – { (W)2 /E2 } = 0 Now we can go off and compute 2ψ and we get 2ψ = [ (i/)2 (W)2 + (i/) 2W ] ψ and we also know that = (-iE/)2ψ so our wave equation becomes [ (i/)2 (W)2 + (i/) 2W ] ψ – { (W)2 /E2 } (-iE/)2ψ = 0 or [ (i/)2 (W)2 + (i/) 2W – (W)2 * (i/)2 ] ψ = 0 and the (W)2 cancel and we are left with 2W ψ = 0 I am nonplussed by this result, it is not what I was expecting. One way to satisfy the equation would be to have 2W = 0 which is a Laplace equation for W. _______________________________________________________________________________ So let's back up a few steps and do things differently. We have n = c/u => (n/c) = 1/u = (p/E) and ψ(r,t) = exp{ i 2π [ W(r)/h - (E/h)t]} = exp{ i [ 2πW(r)/h – ωt]} so our time derivative term in the above is – (n/c)2 = – 1/u2 ∂2ψ/∂t2 = – (p/E)2 (- ω2ψ) = (p/E)24π2 ν2 ψ = (pν/E)2 4π2 ψ = (pν/hν)2 4π2 ψ = (p/h)2 4π2 ψ [ = (1/λ2) 4π2 ψ] = p2 4π2/h2 ψ = 2mT /2 ψ = 2m(E-V)/2ψ where you just verify each tiny algebraic step as shown. Our wave equation was 2ψ – (n/c)2 = 0 which now becomes 2ψ + [2m(E-V)/2]ψ = 0 which we can rewrite as 22ψ/2m + (E-V)ψ = 0 and once more: ([ - i]22/2m +V)ψ = E ψ In quantum mechanics, we know that =- i so we have arrived at [2/2m + V] ψ = Eψ which is of course the Schrodinger equation, time independent. What assumptions and actions did we take in the previous section? (1) We assumed that our "wave function" had the form ψ(r,t) = exp[ i S(r,t)/] and satisfied the normal wave equation. This wave function form is completely within classical mechanics, but the constant could be any number. This wave function is just what makes those surfaces of constant S, so nothing fancy here. And waves always satisfy the wave equation. So apart from , this is all classical. (2) We assumed that our classical waves for a particle have a specific frequency that is ν = E/h and that h = a certain number. This does NOT come out of classical mechanics, it is a completely new idea. In classical mechanics, any frequency would be OK. (3) Our motivation ν = E/h was based on a comparison between classical mechanics and the geometric optics limit of EM wave theory , the comparison being this: ψ(r,t) = exp{ i 2π [ W(r)/h – (E/h)t] } φ(r,t) = exp{ i 2π [ L(r)/λ0 – ν t] } Our comparison was encouraged by the fact that both L(r) and W(r) satisfy the same eikonal equation form in the limit of small wavelength which is the "geometric optics" limit for EM waves, and which is the non-quantum limit for mechanics (ie, classical mechanics). (4) The wave function is the "exponentiated action" ψ(r,t) = exp[ i S(r,t)/], an idea that will have lots of implications I think. So, the proposal here is that maybe there is some general mechanics theory which has this wave equation and has some waves ψ and if you make constant h be "very small" then for normal physical objects for which we do classical mechanics, we will be in the geometric optics limit. Since λ = h/mv, we will be in the "geo optics limit" (ie, "classical mechanics"), if h is small and/or m is large. I think if Hamilton were to read what I have just written, he would be stunned to learn that the constant h is actually one of the constants of the universe and really does exist. Hamilton knew about the analogy between geo optics and classical mechanics, but he did not make the above conjectural steps leading to the conjecture of some kind of "wave function" ψ analogous to the EM potential φ, and an analogous equation shown above : the Schrodinger equation. It says Hamilton in 1834 knew about eikonal optics if not the full Maxwell equations, and I can only presume the wave equation was well known by this time. G points out that Hamilton had no motivation to look for a "more general theory" of mechanics because in his time, it was thought to be complete. We now think we know about that more general theory, and we call it "wave mechanics". In this theory there is a constant h which give the above equation. There are "waves" associated with particles which are surfaces of constant S (constant action), as HJ theory showed. The waves ψ(r) are complex numbers, not just real, and are associated with the probability of a particle's location. A particle does have an associated wavelength which is λ = h/p exactly as HJ theory says, and this is the "deBroglie wavelength". Small particles do "interfere" with each other, just as in the wave theory of light. There is a connection E = hν between a particle's energy E and the frequency of the associated waves. This applies also to photons of light, but this fact was as yet unknown. Atomic physics could not be solved until this more general theory was discovered. Atomic physics served as the motivation to find the QM theory. That is because atomic systems have very small-mass particles and are thus not in the geometric optics limit. G refers to "quantum mechanics" sometimes as "wave mechanics". Maybe the WM name is more associated with the Schrodinger equation approach to QM. Quantized spin for example is part of QM but has nothing to do with wave mechanics. History Added: Hamilton (1805-1865) This suggests that the Hamilton's equations of motion and the Hamilton-Jacobi equation are dated around 1834, here is some history from a Dover $9 book I might buy: (Struik, A Concise History of Mathematics). Hamilton, William Rowan. 1833 [1931]. "On a General Method of Expressing the Paths of Light, and of Planets, by the Coefficients of a Characteristic Function." In The Mathematical Papers of Sir William Rowan Hamilton. Vol. 1. Cambridge: Cambridge University Press. ********************************************************************************* Chapter 10. Small Oscillations ( p 318) ********************************************************************************* This chapter is more complicated that it looks at first blush due to the tricky matrix aspects. There are lots of fine details all of which G mentions, but I do prefer to use my full matrix notation where possible to avoid the subscript details. 10.1 Formulation of the Problem (318). We want to examine small oscillations around an equilibrium point which is a point where ∂V/∂qk = 0 for all coordinates, Normally we think of this as a stable minimum situation, but it could be a min for some and a max for other coordinates. Fine. At the EQ point, set V = 0 by fiat. Then a Taylor expansion of V has the form 10-3 where the ηi are deviations of coordinates qi from EQ. That is to say, we have V ≈ ½ ηTVη = ½ ηVη in dyadic which I won't use. The kinetic energy is less obvious. Back on page 23 we showed that if the Cartesian-to-generalized equations have no time in them (scleronomous), then we know T = ½ TM where Mij = ½ Σk mk ∂rk/∂qi ∂rk/∂qj = Mij(q) again, see page 23. One should really call this a kinetic energy tensor. Notice this fact: Mii = ½ Σk mk ∂rk/∂qi ∂rk/∂qi = ½ Σk mk | ∂rk/∂qi|2 > 0 (c1) You cannot have Mii = 0 because that would require that rk was not a function of qi for all your particles, which means that qi is not even a coordinate of the problem -- it is irrelevant. In our small oscillations analysis, we approximate by saying Mij(q) = Mij(q0) ≡ Tij . Then we have T ≈ ½ TT near EQ. By their definitions, V and T are symmetric matrices, and they are real, hence Hermitian. So, assume equalities now and write V = ½ ηTVη and T = ½ TT . We then have L = T-V = ½ TT –½ ηTVη and the Lagrange equations are then d/dt(L) – ηL = 0 => T + Vη = 0 What exactly is meant by the notation η? It means (η1, η2.....ηn) which are offsets on the generalized coordinates (q1, q2.....qn) from their EQ values (q01, q02.....q0n). So η is a vector of coordinates. Notice that all the coordinates are tangled together in the above ODE, so messy to solve. 10.2 The eigenvalue equation and the principle axis transformation( 321) We then try solutions of the form η(t) = C a e-iωt which leads to (V - ω2T) a = 0 or (V - ω2T) η = 0 and we know this can only have non-zero solutions if det(V - ω2T) = 0 [ so inverse does not exist], and this in turn determines the eigenvalues ωk2. Now, for a particular eigenvalue ωk2 [ k is just a label! ] we can write the above as η(t) = fk(ωk, t) = Ck ak exp(-iωkt) (V - ωk2T) ak = 0 or (V - ωk2T) η = 0 At this point, I want to head off a notational confusion. We put a solution label on our fk here to indicate it is for a particular value of k. But we do NOT put a label on η because this is a vector of coordinates like x,y,z. If you have a particular solution to something, you say x = sinωkt you don't say xk = sinωkt . In other words, the notation ηk is non-sensical, so don't be caught using it! ___________________________________________________________________________________ Digression on T: Let's first think about matrix T all on its own. Since real symmetric, we know we can diagonalize it with a real unitary matrix which we will call B. Imagine we have done this so that BTTB = D where BBT = 1 so B-1 = BT. D = a diagonal matrix TB = BD The columns of the matrix B we can call bk and we know they are linearly independent or can be chosen so if there are degenerate eigenvalues. Aside: We have (bk)i = Bik. The fact that B is unitary tells us that for any k: 1 = (BBT)kk = BkiBki = (bk)i (bk)i = bk bk = | bk|2 so it happens that our columns are normalized vectors as well as being orthogonal. In other words, B is really a rotation. These vectors bk are the normalized eigenvalues of this eigenvalue problem: Tbk = Dkkbk Dkk> 0 from Second Digression below Now define a new set of (non-normalized) eigenvectors by rescaling each separately as follows: fk ≡ bk/ => (fk)i = (bk)i/ => Fik = Bik/ and let F be the matrix with these eigenvectors as columns. Then consider (FTTF)km = Fik TijFjm = (1/)(1/) Bik TijBjm = (1/)(1/)(BTTB)km = (1/)(1/) D = a diagonal matrix = δkm The diagonal elements are all 1, so we have shown that FTTF = 1. Since the eigenvectors are not normalized, we cannot say that F is a rotation which we could say about B above. But we have found (that is, we know exists) a matrix F whose columns are linearly independent which brings T to this special diagonal form of unity! We can write F = BS where S is real diagonal scaling matrix Sij = δij. Then we have FTTF = (BS )TT BS = ST BTTB S = ST D S = 1 Thus it is fair to say that F is matrix operation where we first scaled the axes, and then rotate. Second Digression on T : We know that KE = ½ TT and that T is symmetric. We know we can diagonalize T with a unitary matrix B as above to get BTTB = D. We can think of B of taking us to some new generalized coordinates and velocities = B' and we find that KE = ½ 'TD' = ½ Σk Dkk 'k2 . But the kinetic energy associated with a generalized velocity 'k must be positive, it cannot be 0 or negative. If it is zero, then qk' is not really a generalized coordinate, it is some extra baggage the means nothing, so get rid of it. We conclude therefore that we must have Dkk > 0. This fact was used above when we rescaled the eigenvectors. Had we had Dkk= 0, we would have a problem there. Therefore, we know that the diagonal elements of our original physical matrix T are positive, as noted in comment (c1) above, and we know that in coordinates where T is diagonal, the diagonal values are also positive. Third Digression on T: In general, we write T() = TT . Think of as a vector in a vector space of n dimensions. We know that for any possible selection of this vector in this space the kinetic energy T must be > 0. In M&M this thing was called a "quadratic form" xTTx and the matrix T in this case was called "positive definite". If you select the vector to be along the k-axis and to be a unit vector, then we have T() = Tkk , so we know all the diagonal elements of T are > 0. If we do a congruence on T, the result is still positive definite, which is easy to show if we let T' = CTTC xTT'x = xT CTTC x = (Cx)TT (Cx) = yTTy > 0 // for all y _________________________________________________________________________________ Now back to the main flow. We have this equation to solve (V - ω2T) a = 0 or Va = ω2Ta In order to have non-zero solutions a, we must have det(V - ω2T) = 0 and this will determine the eigenvalues we will call ωk2 which we shall show later are real and positive for a potential minimum situation. Write the eigenvalue equation in matrix form using A and Λ. Assume we have found n distinct eigenvalues and corresponding solutions ak . We then have Vak = Tak ωk2 But write the mth component of this equation: Vmi(ak)i = Tmi(ak)i ωk2 and then use (ak)i = Aik to get Vmi Aik = Tmi Aik ωk2 = Tmi Aik' δk'k ωk2 = Tmi Aik' Λk'k where we have defined a diagonal matrix Λ whose diagonal elements are the ωk2. Then we have shown that VA = TAΛ // an all-matrix equation. Notice that the diagonal matrix ends up on the right here, that is why we moved ωk2 over there early. Show that A†TA is real and diagonal. This is highly non-obvious! Here T is in general not diagonal, but it is symmetric and positive definite, and A is constructed from the solutions to our eigenvalue problem (V - ω2T) a = 0, without regard to normalization, (ak)i = Aik . That is, our claim is true no matter how the eigenvectors ak are normalized. Apply A† to both sides of VA = TAΛ to get A†VA = A†TAΛ = (A† TA)Λ Take the Hermitian adjoint of the above equation and use the fact that V†= V same for T, A†VA = Λ* (A† TA) Subtract these two matrix equations to get 0 = (A† TA)Λ – Λ* (A† TA) Lemma 1: Claim A† TA is real. Proof: Let B = A† TA and let A = X + iY where X and Y are real. Then we have B = A† TA = (X + iY)† T (X + iY) = XT TX – YT TY + i { – YTTX + XTTY } Define R = – YTTX + XTTY, the contents of the {..}. Then R = – YTTX + XTTY Rij = Tmn ( -YTim Xnj + XTim Ynj ) = Tmn ( -Ymi Xnj + Xmi Ynj ) = 0 The 0 follows because Tmn is contracted with an antisymmetric tensor. Thus we have B = A† TA = XT TX – YT TY = real So far then we have show than (A† TA)Λ = Λ* (A† TA) where (A† TA) is real. Lemma 2: Claim BΛ = Λ*B with B real implies that Λ must be real. Proof: BimΛmj = Λ*imBmj => Bij Λjj = Λ*iiBij => Bij (Λjj–Λ*ii) = 0 Setting i = j we get Bii (Λii–Λ*ii) = 0 => Λii=Λ*ii => Λ is real if Bii> 0. In our case, we know that Bii > 0 for each i because B=A†TA is a kinetic energy form, see previous comments. Now we know that BΛ = ΛB which says [ B, Λ ] = 0, where B = A†TA. Lemma 3: if a matrix commutes with an ev's diff diagonal matrix, it must be diagonal as well. Proof: BimΛmj = ΛimBmj => BijΛjj = ΛiiBij => Bij(Λjj – Λii) = 0 If the diagonal elements of Λ are all different, then for i≠j this tells us that Bij = 0 so B must be diagonal. I ignore the case where there are identical diagonal elements in Λ because things are messy enough here. but G does talk about how this is handled. It is a GSO issue where we can choose our A to make things work, and so choosing A will result in a diagonal B. So we now conclude that B = A†TA is real and diagonal, where A is constructed from the eigenvectors of the problem (V - ω2T) a = 0 and this is true no matter how the ak are normalized. Show you can select normalization of the ak such that B = A†TA = 1. Since B is already diagonal and positive definite, all we have to do is apply a matrix S which stretches the axes appropriately to make this work. Call this diagonal matrix S and Sij = δij/. Then use A' = AS and we get B' = A'†TA' = (AS)†T (AS) = S† A†TA S = SBS = 1. Again, we knew that B was diagonal with positive diagonal elements, so the last step works. Show that matrix A' normalized in this way diagonalizes V. We had VA' = TA'Λ => A'†VA' = A'†TA' Λ = Λ QED The main point here is that we have shown that we can find a matrix A' which diagonalizes both T and V in this way. T goes to 1, and V goes to Λ. The transformation A' that does this is called the principle axis transformation for our problem. We saw this idea before with the inertia tensor. A simple way to find the eigenvectors a. We had this eigenvalue equation: (V - ω2T) a = 0 We know that we can diagonalize T, it has positive diagonal elements and of course det(T) ≠ 0. This must then be true of the undiagonalized T. Therefore we know that T-1 exists. So apply it to the above equation to get (T-1V - ω2) a = 0 (Q - ω2) a = 0 Q a = ω2 a Q = T-1V Now we have a "normal" eigenvalue problem with matrix Q and we just solve as usual for eigenvalues and vectors. Maple can do this just fine, see section 10.4 below. 10.3 Frequencies of free vibration and normal coordinates (329). Now go back to our original equation which was this: (V - ωk2T) η = 0 or (V - ωk2T)mi ηi = 0 Suppose we define some new coordinates this way: η = A ζ ζ ≡ A-1η Then we have (V - ωk2T) Aζ = 0 Apply AT on the left to get this vector equation: (Λ - ωk2) ζ = 0 (Λ - ωk2)ij ζj = 0 (Λjj – ωk2) ζ(k)j = 0 This is an trivial equation to solve, and we find that, for a specific ωk, the eigenvectors are: ζ(k)i = δki ζk or ζ(k) = (0,0,....ζk,0,0..) This is a "mode" in which only the kth coordinate is non-zero! This is called a "normal mode" and we have one for each eigenvalue ωk2. Because η = A ζ, it is likely that all the coordinates of the η vector are going to be non-zero. So in a normal mode, you might have lots of offset coordinates moving, even though you have only one normal coordinate moving. But what is our one ζk coordinate doing in this normal mode? Well, we know these things so far, η(t) = fk(ωk, t) = Ck ak exp(-iωkt) η = A ζ Thus we can say ζ = (A-1) η = (A-1) Ck ak exp(-iωkt) = Ck exp(-iωkt) [(A-1) ak] If we look at a particular coordinate, we find ζi = Ck exp(-iωkt) [(A-1)ij (ak)j] = Ck exp(-iωkt) [(A-1)ij Ajk ] = Ck exp(-iωkt) δik so doing it this way we again find out that only ζk is moving, and here is what it is doing ζk = Ck exp(-iωkt) ζi = 0 i ≠ k so our normal mode coordinate ζk is doing simple harmonic motion. What are T and V and L in these new coordinates? T = ½ TT = ½ T ATT A = ½ T = ½ Σi i2 V = ½ ηTVη = ½ ζT ATV A ζ = ½ ζTΛζ = ½ Σi ωi2ζi2 L = ½ Σi (i2 – ωi2ζi2) and the Lagrange equations become the first equation set below for i = 1..N. i + ωi2ζi = 0 Tijj + Vij ηj = 0 ie T + V η = 0 which is a simple harmonic oscillator for this coordinate at frequency ωi, confirming what we just found out by different means. Notice how simple the ODE appears in normal coordinates. We have copied from above what things looked like in the original physical coordinates and we see again thatthe coordinates are tangled together in each equation. So, a general solution to the problem then will be a superposition of SHM in each of these normal coordinates. A general solution then has ζi = Ci exp(–iωit) going on in each of the normal coordinates. ζ = a column vector of the above! If you want a general form then for η, you have η = A ζ or ηi = ΣjAij ζj = Σj Aij Cj exp(–iωjt) which can then have a mixture of all the normal mode frequencies! If only one normal mode k is activated, then you would have ηi = Aik ζk = Aik Ck exp(–iωkt) and then all the η components are going at the same frequency. Notice that there are no harmonics of any of the normal modes, just the fundamental of each one. This is in the small oscillation limit. How did we know the ωk2 were positive? Go back to (V - ωk2T) η = 0 and write as V η = ωk2 T η Therefore, for a solution η(t) = fk(ωk, t) = Ck ak exp(-iωkt), we have ηTV η = ωk2 ηTT η But from page 319 the LHS is really δV away from EQ and, if we are at a minimum of V, then LHS is positive. We already know RHS is positive real because T is a positive definite matrix (digression #3). So as long as we are at a stable minimum of the potential defined as δV = ηTV η > 0, then we know that the ωk2 are real and positive, hence justifying the name. If we are NOT at a minimum, then ωk = imaginary and we have exponential instable behavior in all normal coordinates. Comparison with the inertia tensor situation. In that case, we considered I = nIn where I was the moment of inertia about rotation axis n . We diagonalized I by a rotation R so RTIR = Idiag because I was a Hermitian matrix, so Idiag has real positive diagonal elements . We could interpret this rotation R as taking us to a new rotated set of coordinates r' in which n appears as n' according to n = Rn'. Then we find that I = nIn = Rn'I Rn' = n'Idiag n' = Σin'i2 (Idiag)ii = I11n'x2 + I22n'y2 + I33n'z2. In these new coordinates, the ellipsoid of constant I is aligned with the axes. We referred to R as the principle axis transformation. In the current context, we have 2T = T . We find a combined rotation/scaling transformation we called A' = AS above and define η = A' ζ to get 2T = 1 = T . We had to add the scaling so that A' would also diagonalize V as well as T. So this combination rotation/scaling matrix A' is still called a principle axis transformation. In the space, we have some kind of kinetic energy ellipsoid at some weird orientation. The rotation A (assume orthonormal ak) rotates us to some axes in which this ellipsoid is aligned with the axes. Then the scaling reduces this aligned ellipsoid to a sphere: 2T = T = (A' )T(A') = (A'TTA') = (ST(ATTA)S) = rotate, then scale = T 10.4 Free vibrations of a linear triatomic molecule (333). For this problem, we may quickly set up our basic matrices as follows The mass tensor T is pretty obvious since center mass is of size M, others m. The V tensor follows quickly from the page 333 equations. We want now to find the eigenvectors for our equation (V - ω2T) a = 0 (T-1V - ω2) a = 0 (Q - ω21) a = 0 We do the right equation in Maple as follows: Q = T-1V (Q - ω2) = G Ga = 0 The determinant shows the expected eigenvalues including ω = 0 as shown by G in 10-52 We proceed to get the eigenvectors as follows: The eigenvalues (here fore ω2) agree with G 10.52. The eigenvectors are not normalized in such a way that ATTA=1, but they are normalized here in a way that makes things very clear. Then center eigenvector is [1,1,1] which is the overall translation. The left eigenvector is the contrary motion of page 336 (b) of G. And the right eigenvector is page 336 (c). We now want to construct our matrix A from these eigenvectors such that ATTA = 1. This requires some Maple fiddling, but we manage to get our eigenvalues to be the columns of a matrix A, where we have added a scaling factor λi for each column (left side below) We hope to select scaling values such that ATTA =1. We first compute this product directly as on the right above. Using the solve function, we make Maple tell us the values for the λi and we get And this does indeed make R = 1, as shown on the right above, and we can also display the updated A matrix Note that A31 = 1/, while A32 = 1/. These agree with the 9 matrix values that G gives on page 336 in equation (10-56a,b,c), though his columns are shuffled relative to mine. Finally, we can show that our matrix A does in fact diagonalize the V matrix as the formalism claims: and the diagonal elements are the three ωk2 and this is our matrix Λ as discussed above. G goes on to discuss the fact that this three point-mass system really has 9 coordinates in 3D space, so removing translations there are 6 modes. The thing can rotate non-trivially with point masses in two dimensions, so that leaves 4 modes, and we found 2 already, so there are 2 more. They are the same mode in two dimensions, so really only one new item to think about, it is this one: So overall there are 2+2 = 4 vibrational modes. We have not solved this last problem for the eigenfrequency, but it does not seem too hard. 10.5 Forced vibrations and dissipative forces (338). Suppose we just start all over again and throw in some frictional forces to get T + F + Vη = 0 Put in η(t) = C a e-iωt to get (V -iωF - ω2T) η = 0 Treat this as an eigenvalue equation, det=0 will determine the ωk. We then have (V -iωkF - ωk2T) ak = 0 or (V– γkF + γk2T) ak = 0 Goldstein does not to on to actually solve this problem in a general sense. _______________________________________________________________________________ Digression: Let's try to follow the former path and see where it leads. Write the eigenvalue equation in matrix form using A and Λ. Assume we have found n distinct eigenvalues ωk and corresponding solutions ak . We then have Vak – iFakωk – Tak ωk2 = 0 But write the mth component of this equation: Vmi(ak)i – iFmi(ak)i ωk – Tmi(ak)i ωk2 = 0 and then use (ak)i = Aik to get Vmi Aik –iFmi Aik ωk – Tmi Aik ωk2 = 0 Vmi Aik –iFmi Aik'δk'k ωk – Tmi Aik' δk'k ωk2 = 0 VA –iFAΩ - TAΩ2 = 0 Ω2 = Λ where we have defined a diagonal matrix Ω whose diagonal elements are the ωk. So we have shown that VA = TA Ω2 + iFAΩ // an all-matrix equation. Notice that the diagonal matrix ends up on the right here, that is why we moved ωk and ωk2 over there early. Show that A†TA is real and diagonal ??? Apply A† to both sides of VA = TA Ω2 + iFAΩ to get A†VA = A†TA Ω2 + i A†FAΩ Take the Hermitian adjoint of the above equation and use the fact that V†= V same for T and F, A†VA = Ω2*A†TA + i Ω*A†FA Subtract these two matrix equations to get 0 = [(A† TA) Ω2 – Ω2* (A† TA) ] + i [ (A†FA) Ω – Ω*(A†FA)] As we did before, write A = X + iY where X and Y are real. Exactly as in Lemma1 above, we can show that B = A† TA = real [ and it is also B† = B so BT = B so B is real and symmetric ] so the above equation then says 0 = [B Ω2 – Ω2* B ] + i [B Ω – Ω* B] B = A† TA = real Unlike before, because there are two complex terms above, we cannot say Ω = real. Instead, write ωk = pk + i qk Ω = P + iQ P and Q are diagonal matrices ωk2 = pk2 - qk2 + 2ipkqk Ω2 = P2– Q2 + 2iPQ Stuff this into the above to get 0 = [B Ω2 – Ω2* B ] + i [B Ω – Ω* B] = [B (P2– Q2 + 2iPQ) – (P2– Q2 – 2iPQ) B ] + i [B (P + iQ) – (P – iQ) B] = { B (P2– Q2) – (P2– Q2)B – BQ + QB } + i { 2BPQ + 2PQB + BP – PB } = 0 But now B, P and Q are all real, so we can write this as two separate equations: B (P2– Q2) – (P2– Q2)B – BQ – QB = 0 2BPQ + 2PQB + BP – PB = 0 Suppose we find P and Q from our eigenvalue determinant. Rewrite the above as [ B, (P2– Q2] – [B,Q]+ = 0 Ω = P + iQ (1) 2[B,PQ]+ + [B,P] = 0 B = A† TA = real (2) All matrices are symmetric here. For such matrices we know that [Q,R]T = – [Q,R] and [Q,R]+ = + [Q,R]+ So let's transpose both equations above to get – [ B, (P2– Q2] – [B,Q]+ = 0 (3) 2[B,PQ]+ – [B,P] = 0 (4) Then write down (1) - (3) and (2) - (4) to get [ B, (P2– Q2] = 0 [ B,P] = 0 In either case, we see that B is commuting with a diagonal matrix, so B must be diagonal by Lemma 3, assuming as before that all eigenvalues are different. For CC pairs and other degeneracy we will need to do GSO as before, so let's ignore this for now. So, we have indeed shown that B = A† TA is real and diagonal, just as in the former problem. We then want to do the same trick as before and write A' = AS to get B' = A'† TA' = 1. Let's assume we have done this, and let's drop the prime on A from now on. We had as our starting point: VA = TA Ω2 + iFAΩ Apply A† from the left to get A†VA = A†TA Ω2 + i A†FAΩ = Ω2 + i A†FAΩ Ω2 + i A†FA Ω – A†VA = 0 Well, all I can say is this Ω2 – A†{ V – iFAΩA-1}A = 0 so we can argue that transformation A brings the matrix V – iFAΩA-1 to diagonal form. But that is pretty meaningless I'm afraid. Go back to the previous, Ω2 + i A†FA Ω – A†VA = 0 This is 2N2 real nonlinear equations for the 2N2 real unknowns in Aij . We might just as well go back to our starting point (V -iωkF - ωk2T) ak = 0 or (V -iωkF - ωk2T) A = 0 and regard it as 2N2 real linear equations for the 2N2 real parameters of A. Of course only possible solutions when the ωk eigenvalues are used. STOP. No doubt the theory for doing all this is presented somewhere, maybe using the quadratic forms stuff. I don't want to reinvent that whole bag here, and since Goldstein omitted it, it is likely a bit messy. I might guess that in general there are no "normal modes" in the general case. _______________________________________________________________________________ If it happens that F is diagonalized by the same A that we had in our original problem with no friction, then you get "normal modes" as before with equations as in 10-66. This results in the shifted eigenvalues as shown in 10-68 so we get some damping. So this situation is easy. Now G talks about a general system under the effect of a "driving force" and we then have 10-74. This gives 10-75 which is the equation I have been dealing with above, but the RHS is not 0 so you can use Cramer's Rule to solve it! The whole thing runs at the driving frequency ω and we will get resonant behavior from D(ω) shown in the Cramer's solution 10-76. Factoring this thing gives the complex frequencies which result in resonant peaks which are finite due to the damping. This is more the kind of problem you want to solve when light is shined on atoms. G closes with some comments about how this theory applies to EE work. Instead of talking about that, he is going to shift attention to continuous systems in the last chapter. ********************************************************************************* Chapter 11. Continuous Systems and Fields ( p 347)∂ ********************************************************************************* 11.1. Transition from discrete to continuous system (347). Our model here is to imagine an elastic rod as made of a cross section of little masses on springs, like a muscle fiber sheaf, and we can think about just one of these fibers, though two are shown in the figure. We start with constants m,a,k and we think of those longitudinal modes we just studied in the last chapter. Since the rod is of finite length, we start with a finite number of masses, perhaps L/a, We are going to take the limit a→0 and see what happens. The finite Lagrangian is as shown 11-3,4 in terms of those differential-from-rest longitudinal position variables ηi. You can see that m/a → μ, linear mass density of one fiber. And the paren is going to be the derivative dη/dx squared, and so ka → Y, the Young modulus so called. As a shrinks, k must grow. The positions were ηi which we could express as η(xi) where xi = 0 + i*a. This becomes η(x) and we have a longitudinal displacement then at any value of x. Now, 11-5 shows the EL equations, and you get two Δη terms from separate terms in the series, those adjacent to point xi. The a2 denominator is correct, and the two terms on the right will become -Y times ηxx as seen in 11-7 and we end up with a 1D wave equation, no surprise, velocity of Y and μ. The L has the first derivative ηx squared as shown, coming directly from 11-4. Notice Σia becomes ∫dx. So this integral in 11-6 can be thought of as just summing over all the particles in the object. For each value of x, η(x) is a different "generalized coordinate". The thing inside the integral is a Lagrangian density L. So this model would correctly predict longitudinal waves in the rod, but G does not mention that. 11.2 The Lagrangian formulation for continuous systems (350). In our example we saw L (ηt,ηx). Had we used V = kx3 instead of kx2, we would have gotten ηx3 in 11-6. This form is coming somehow from the nearest-neighbor power law interaction. If V = ksin(x), you could still figure it out from trig expansion formulas and I think in all cases result will be a function of ηx. So all these comments are justification for the general form of a L density as presented in 11-10. Notice in 11-8 that L is a spatial integral only of L. If you were to include dt as well, then you are talking "action", and that is coming next. Now we do a variational business on the action in 11-11, but G does not use the word "action" for object I. The summation characteristics are all constant in terms of variation (here integrals) so δL in 11-12 comes from variations in the coordinates η etc, not variations in the summation indices. He might have used the notation ∂L/∂ηx in 11-12 but he would have to write out all three terms including ∂L/∂ηy. The trick is going to be parts integration on the second two terms to make them linear in δη which is an arbitrary variation of the coordinates η(x) at all points in the spacetime volume (but no variation at the endpoints so parts will go away). He shows the parts in equation α. The parts object would be the product ∂L/∂i δη and δη = 0 at the surface boundary so parts yield 0. Then he does the ηx parts in 11-14 and 11-15 swinging around the ∂x operator. Of course the usual minus signs, so the last two terms in 11.16 are negative. As desired, δη is isolated and arbitrary, so the effective EL equations or Hamilton's Equations are as in 11-17. I might write this as ∂t (∂L) – [ ∂ηL – (η) L ] = 0 We can't call these EL equations because L does not appear in them, and we have no derivation path. So right now these are just "Hamilton's Principle equations on motion". Since everything is a function of r, you can think of there being a separate equation here for each η(r) coordinate. If we were to think of elastic rod in 3D, then we would have had ηx and ηy and ηz coordinates at each location r, and this is then the j index in 11-19. So yes, his notation is well justified. I use superscripts. G then introduces the "functional derivative" which picks up the change in L from change in ηi(x) AND from ηi(x). That is to say, if we somehow change ηi(x), then it is likely we will also change its spatial derivatives ∂kηi(x), and we know from 11-10 that L is a function explicitly of both these things. so this fancy derivative δL/δηi picks up all these terms. So 11-19, are both very clear to me. Notice that by definition, L appears inside the δL/δη, but the RHS contains L. As G points out, this δL/δη thing picks up the ugly gradient terms so you don't have to keep writing them, and then everything becomes simple again. We get 11-22 for example, and equations of motion once again look just like EL equations in 11-23, where we are just inserting our definitions 11-19,20. We now see why we put L in these things, because the results come out very simple. We have recovered our EL equation form! In passing, he notes that we can write the 11-17 Hamilton's equations in covariant form as in 11-24 which certainly is impressive, and if L is a scalar, you are cooking with relativistic gas. G wraps up the section computing these new objects for our simple longitudinal rod, and now he points out this is the wave equation and quotes the velocity in 11-25. So far so good! 11.3 Sound vibrations in a gas (355). This will be our second real-world example of using the Lagrangian density formalism to solve a problem. Unfortunately, this involves dynamics of gases, and my knowledge of "fluid dynamics" is very poor, though I do have a Schaum book on the subject, an early chapter of which deals with the basic equations of fluid dynamics. Luckily, Goldstein helps us out and derives most things. One must understand up front the meaning of η(r,t), even this is not trivial to me. You must picture somehow the gas at complete rest and at position r you imagine a tiny volume, perhaps a cube or a sphere, that is perhaps differentially small. In a high speed dynamic situation such as a high-frequency sound wave passing through the gas, at some time t the quantity η(r,t) tells you how much and in what direction that tiny volume is "displaced" from its rest position. The assumption here is that this displacement is temporary and there is some restoring force from the nearby gas that wants to make the little volume move back to its rest position. G does not bring this subject up, but it must be a stable equilibrium situation. In a longitudinal wave travelling in z, I guess the displacements are in the z direction, and the little volumes are pushed around by pressure variations. So fine. We think of each little microvolume as a generalized coordinate -- a sort of particle which a position and a velocity which would be (r,t). Now G likes to use μ as gm/cm3 density instead of ρ. We know that the mass in our little volume is going to be μΔV and the energy of this mass ½( μΔV) 2. If we leave out the ΔV, we are talking kinetic energy "per unit volume" and he uses T (script T) as the "kinetic energy density". So we like equation α on page 355. The μ0 means the at-rest density and we are assuming that we are not hugely varying this with our disturbance. Obtaining the corresponding potential density V takes G a full 3 pages! I will outline the steps here (1) First, I think Vo is meant to be some small but finite volume of gas which near constant V. Then the claim is that Vo V is the potential energy in this volume which arises from work done PdV integrated as the volume say increases to V0 + ΔV, where now ΔV is perhaps distributed around the boundary of Vo. (2) We need a formula for P(V), and we assume a linear fit as shown 11-28: P(V) = Po + (dP/dV)(V-V0) => integral PdV = Po ΔV + ½ (dP/dV)(ΔV)2 (3) Now we need a formula for dP/dV so we need some kind of gas law equation. Because we are doing relatively high frequency, heat does not flow, we are adiabatic, and we have PVγ = constant, which I vaguely remember from somewhere. So the derivative is as shown 11-31 and we can plug that in. (4) Now he introduces a "dimensionless density change variable" σ ≡ Δμ/μ0 which is just fine with me. We know that μ = V/M and μ0 = V0/M, and here we have another implicit assumption. It is that our volume V contains a certain fixed number of gas molecules with total mass M, so the volume sort of moves with the particles in some sense. That is, the mass M of our volume is fixed, so that μ = V/M=> Δμ = ΔV/M. Then ΔV = M Δμ = M (μ0 σ) = (M μ0) σ = Voσ . So we can replace ΔV in both places in 11-28 with Voσ to get 11-34. (5) The next problem is to relate σ to η. The quantity μ0η must be the amount of mass that is "living on the wrong side of the border" at some time t and at locality r, per unit of border area, and dimensions are correctly gm/cm2. I don't think it would be right to call this a mass "flow" because there is no nothing of speed here, just how much is over the border at our instant of interest. If we now consider a different FIXED and FINITE volume and use the same symbol V for this volume, then I agree that the total mass displaced out of this volume is going to be ∫ μ0η dA = mass displaced out of volume if dA points out from the volume. This means the total mass in the volume ∫μ dV is low by this amount, so Δ(∫μ dV) = – ∫ μ0η dA = – μ0∫ η dA = – μ0∫ div η dV by def of divergence But the LHS is the same as (∫Δμ dV) since Δμ is the local mass density change from EQ. Thus we conclude that Δμ = – μ0 div η => Δμ/μ0 = σ = -div η and so we have related σ to η as requested. Somehow this says that the dimensionless outflow of stuff from a tiny volume reduces the dimensionless density -- continuity. But again, it is not a flow, it is the amount caught outside at time t must equal the amount not present on the inside. So this is 11.36. (6) We then replace σ by -η both places it appears in 11-34 to get 11-37, and combining this with our KE result, we get L as shown in 11-38, and sure enough, we see spatial derivatives of η appearing. (7) We now want to use 11-18 to get our equations of motion from L. The necessary derivatives are computed in 11-39 and it is at this point that we see that the linear η won't contribute to the equations of motion. The reason is that the equations say we need to do ∂x on δik which is a constant. So the simplified L is in 11-40 and we arrive at 11-41 which is a strange looking equation indeed in η. But apply div to both sides and it becomes 11-42 which is the wave equation in σ ≡ Δμ/μ0. We thus have a wave in the mass density, a sine wave in same. The pressure is not so simple due to adiabatic gas law 11-30, and out pops our sound velocity as shown, a function of γ, Po and μo. Higher pressure means faster sound, also lower density means faster sound (but in EQ, these two quantities are related by PV = μRT/Mm for ideal gas, say, where Mm is the molar mass). So in conclusion, this serves as another Lagrangian Density problem example. 11.4 The Hamiltonian Formulation (355) No surprise that this comes next. We go back to the discrete spring and mass chain and compute H in the discrete case in 11-46 and then go to the limit in 11-47. We define the canonical momentum density π, and get 11-48 which relates the densities H, L and π exactly we you would expect, this for a single coordinate function η. In the 3D case we have a dot product of π with sitting in there. It is all as expected. We then do the variation dH and shuffle all the parts around as usual and we end up with 11-52. Then we do dH a different way using the functional derivatives we know understand, and we end up with 11-53. But if we use the Lagrange equations, two terms cancel here giving 11-54. If we then compare this to 11-52, we end up with Hamilton's Equations of motion shown 11-55 page 362. As expected, they contain the functional derivatives just the way we had things with L. We can write these out as in 11-56 which we will use for calculations. So much for the H formalism! We now go back to our sound in gas example and do it the H way. We find that H = T+V as expected. One Hamilton's equation repeats our π definition, the other recovers our equation of motion for the gas, same as we got by the L method. Next, we show that if the H has not explicit time dependence, H is conserved, again as expected. He is just turning his big crank. Next, we consider an unknown object G which is the integral of a density which I guess we have to assume is a function of η and π in the same convective sense that H is. We find easily that dG/dt = [G,H] + ∂G/∂t where these are Poisson brackets but we have a huge sum over the indices which here are continuous space. So this shows that if G "commutes" with H, it is a constant of the motion. We are reminded that this is a sort of global symmetry (conservation rule) and there are local ones like continuity conditions as well, the gauge condition ∂μAμ = 0 comes to mind. 11.5 Descriptions of Fields (364). In our sound example in the elastic rod, η(x) was a scalar field, but it worked in a medium of matter. In E&M, we have fields like Ai, φ which don't need any "matter" to "undulate in". This must have been a big shock in the physics world when it was finally realized (but see Einstein's comments on the ether as seen in general relativity, separate doc I saved). Next topic: any L is acceptable if it gives the right equations of motion, even if does not have the traditional form of T - V. For example, 11-61 provides an alternative to the L of 11-40, but it gives the same equations of motion for the sound in gas problem. That is, we get the same wave equation. Next: G presents an L in 11-65 which, when we apply the Lagrange equations, generates Maxell's equations. This is shown in complete detail, I did it all, not too hard. You have to get used to ∂(∂η/∂x) kind of things. Just assuming the A and φ fields (potentials) satisfies the first two Maxwell's, so we only have to do the meaty Maxwell's which show interaction with charge and current. Next: What happens when we install j = ρv into the picture, which relates to physical charged particles? If we start with our vetted E&M L shown in 11-65, and install this j/ρ relation caused by actual charged particles, we get 11-72. But this L includes only the "charge effects" of these particles, not "inertial effects". That is to say, 11-72 has particles of charge qi, and their motion (such as velocity) affects the equations of motion, as in Maxwell's equations. But if you want to include inertial effects of these particles, you have to "throw in" the usual KE term and this gives 11-73 = 11-73'. We can then interpret "the middle term" as the interaction between the particle world and the light world, and this of course is the famous qivμAμ = jμAμ interaction term which we use in QED. He says nothing about "perturbation theory" here in regard to this interaction term. But we see that E2 - B2 is the L for E&M and this must be FμνFμν which is a world scalar, and of course so is jμAμ. So G launches us in this way off into QED and mentions the construction of field Lagrangian densities as the job of the particle theorist of the 1950 era. They were getting into π mesons for example, and were trying the Yukawa potential perhaps, or scalar field theory, I cannot remember right now. The closing comment is that you need (he feels) to have a mechanical description of something before you can "quantize it". The L and H methods of this book are the paths you follow.