Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Math / Lagrange Multipliers / Appendix D

Appendix D v3

DOCX · 83.5 KB
Open DOCX file

Appendix to a larger document on Lagrange multipliers, following the notation of Ray and Shamanna (2006) and referencing Goldstein. It defines holonomic and non-holonomic constraints, allowed and virtual displacements, and the assumption that constraint forces do no virtual work, with ramp and rigid-stick examples. It then uses multipliers to show the constraint forces equal sums of multiplier-weighted constraint coefficients.

AI-written summary; may contain errors.

Extracted text (machine-read; may contain errors)
Appendix D: Lagrange Multipliers for problems with Constraints Our main focus in this document is the use of Lagrange Multipliers to find candidates ri for which a scalar function f(r) is stationary, df = 0, subject to constraints. In this Appendix, however, we consider a different application of Lagrange Multipliers that relates to classical mechanics for a system of N particles in the presence of constraints. A certain amount of background information is required to put this application into context, which we now present. Constraints and virtual displacements Let rk(t) be the position of the kth particle. At any fixed time t, if the particles are unencumbered with constraints, the variables (r1, r2..... rk) can be regarded as a single point in the space E3N and the "configuration space" of the system indicated by (r1, r2..... rk) is all of this space E3N. All points in E3N are "legal" points for the system. A holonomic constraint has the form a(r1, r2..... rN, t) = 0. (D.1) At a given time t, this equation represents a surface of dimension 3N-1 in EN. If there are s such holonomic constraints fi(r1, r2..... rN, t) = 0 i = 1,2...s (D.2) then at time t the system vector (r1, r2..... rN) is constrained to lie on a surface of dimension 3N-s in EN, as discussed in Appendix C, so each extra constraint lowers the operating surface dimension by 1. One can think of the term holonomic as meaning that the constraints are functions of "whole" coordinates like r2 as opposed to differential coordinates like dr2. A non-holonomic constraint is one that is not holonomic. An inequality constraint like a(r1, r2..... rN, t) < 0 which might bound the particles to one side of a surface is an example of a nonholonomic constraint, but our interest lies more in nonholonomic constraints of the form Σk=1N bk drk + btdt = 0 (D.3) in which differentials rather than "whole" coordinates appear. The "coefficients" bk and bt can in general be functions of (r1, r2..... rk, t). If it happens that such a constraint can be integrated by some hook or crook to give an equation of the form a(r1, r2..... rk, t) = 0, then it is regarded as being a holonomic constraint, so a non-holonomic constraint must be non-integrable. If there are m such non-holonomic constraints imposed on a system, one could write them as Σk=1N Aik drk + Aitdt = 0 i = 1,2...m or (D.4) Σk=1N Aik k + Ait = 0 i = 1,2...m . Here bolded Aik represents 3 matrices each of which is m x N (rows x columns). Since these constraints are non-integrable, one cannot really described them in terms of surfaces in E3N. If one differentiates the holonomic constraints (D.2) one gets Σk=1N [(k)fi] drk + [∂tfi] dt = 0 i = 1,2...s or (D.5) Σk=1N [(k)fi] k + [∂tfi] = 0 i = 1,2...s where (k)≡ (∂/∂(rk)1, ∂/∂(rk)2, ∂/∂(rk)3) is a gradient with respect to the components of rk. These equations have the same form as those in (D.4), but these are holonomic because they are integrable to give fi(r1, r2..... rN, t) = 0. For convenience below, we shall now combine (D.4) and (D.5) into a single set of equations: Σk=1N Bik drk + Bitdt = 0 i = 1,2....C // C = s+m or (D.6) Σk=1N Bik k + Bit = 0 i = 1,2....C where Bik = Aik Bit = Ait for i = 1,2...s non-holonomic Bit = (k)fi Bit = ∂tfi for i = s+1, s+2....C holonomic . (D.7) In general, the functions fi, Bik and Bit are all functions of (r1, r2..... rk, t). In the above we have followed the notation of Ray and Shamanna (2006) and we continue our path in the manner they propose. Suppose the set of N differentials {drk} satisfies all of equations (D.6) for some time interval dt. Such a set {drk} is then called a set of allowed displacements. Imagine that {dr'k} is some other distinct allowed displacement set. Elements of the difference set defined as {δrk} = {drk} - {dr'k} k = 1,2...N (D.8) are then called virtual displacements. This is how Ray and Shamanna define the vague term virtual displacement which appears in many texts (about which they have much to say in their Section I A). It follows then that for any virtual displacement set {δrk} equations (D.6) take this form Σk=1N Bik δrk = 0 (D.9) simply because the last term in (D.6) cancels out when one writes δrk = drk - dr'k. Notice that the δrk are not arbitrary differential displacements. They are differences of allowed displacements. Meanwhile, Newton's Law for a system of N particles subject to constraints has the form mk = Fk + Rk k = 1,2...N (D.10) where Rk is the sum of all constraint forces applied to particle k and Fk is the sum of all other forces applied to particle k. Assumption that constraints do no virtual work Next, assume that ( R&S refer to this as being an "ideal constraint" situation), 0 = Σk=1N Rk δrk . (D.11) In Goldstein this equation corresponds to setting the second term in (1-40) to zero. The main justification for making this assumption is that it is valid for particles which form a rigid body and for many other constraint situations. There seems to be no general justification of this equation for all constraint situations, and it is certainly not valid if friction is present. The equation can be interpreted as saying that the total virtual work done by all the constraint forces on all the particles of the system vanishes. As a first simple example, consider two particles 1 and 2 independently sliding down a frictionless static ramp. In this case for particle 1 the constraint force R1 is the normal force of the ramp acting on the particle, and this is clearly perpendicular to dr1 for any allowed displacement set {drk}, and thus R1 is also perpendicular to any δr1 which is part of any virtual displacement set {δrk}. In this case we have R1 δr1 = 0 and R2 δr2 = 0 so the terms in (D.11) are separately zero. A second example is more interesting. Particles 1 and 2 of arbitrary masses slide down a one dimensional ramp but are glued to the two ends of a massless stick so they comprise a rigid body. The upper mass is attached to a string which provides a constant tension T. Here is a picture, (D.12) Particle 2 exerts a constraint force T12 on particle 1 through the stick, and Particle 1 exerts a constraint force T21 on particle 2. Since there is only one force of tension (or compression) in such a stick, T12 = T21. All allowed displacement sets {dr1, dr2} are of the form {dr, dr}, and therefore all virtual displacement sets are of the form {δr1, δr2} = {δr, δr}. In (D.12) we have drawn the two normal forces Ni and the two total constraint forces Ri. It is no longer true that R1 δr1 = 0 and R2 δr2 = 0 as the picture shows. However, it is true that Σk=12 Rk δrk = [ Σk=12 Rk ] δr = 0 because R1 and R2 have equal and opposite components along the ramp, so equation (D.11) is valid for this example. Now suppose the ramp is translating in some direction not along the ramp. The displacements dr1= dr2 are then no longer along the ramp so that Σk=12 Rk drk ≠ 0. However, any virtual displacement set would still have δr1 = δr2 along the ramp and then again Σk=12 Rk δrk = 0. Remember that a virtual displacement set is the difference of any two allowed displacement sets, and in this subtraction the effect of the ramp translation cancels out. So equation (D.11) is still valid. Other examples are given in Ray and Shamanna Section III. Reader Exercise: show in the last example that T12 = T m2 / (m1+m2) and that this makes sense if m1 >> m2 and vice versa. We proceed then assuming our system of interest respects equation (D.11). Comment on internal forces: In (D.10) we have classified the forces acting on particle k into two groups and we write Rk + Fk, where Rk are "constraint" forces and the Fk are any "other" forces. Goldstein refers to the Fk forces as "applied" forces in (1-39), while Ray and Shamanna call them "external" in Sec IV. Consider a system consisting of N masses all connected by a network of massless springs. The force acting on particle k due to all the springs to which it is attached would be best classified in the Fk "other" category. Were we to classify this force into the Rk group, we could not use the no-constraint-work property (D.12) since spring forces do work as the springs compress and expand. An example of such a system is the Solar System where the spring forces are represented by gravity. It would seem a misnomer to classify the internal forces in this case as "applied" or "external". However, if all the springs are slowly adjusted until they become infinitely stiff, in that limit it would be better to classify the internal spring forces into the Rk constraint forces group, since these springs do no work and (D.11) applies. The system is then a "rigid body" and in this case, with the internal forces classified as Rk constraint forces, one benefits from the many simplifications which arise for rigid body mechanics. The internal forces in our spring example fall into the category where Fij = - Fji and Fij is in the direction ri- rk . When this is not the case, the rigid body limit does not apply, but the internal forces can still be classified as "constraint" plus "other". An Application of Lagrange Multipliers Recall the C = m+s constraint equations (D.9), Σk=1N Bik δrk = 0 i = 1,2...C C ≡ m+s = total number of constraints (D.9) Multiply each of these C equations by a Lagrange Multiplier λi and then add the equations to get 0 = Σi=1C λi [Σk=1N Bik δrk] = 0 = Σk=1N [ Σi=1C λi Bik ] δrk . (D.13) Now subtract equation (D.13) from (D.11) to get 0 = Σk=1N [ Rk - Σi=1C λi Bik] δrk . (D.14) We know that the N vectors in the set {δrk} are not linearly independent because they must satisfy the equations (D.9), so we cannot claim from (D.14) that [ Rk - Σi=1C λi Bik] = 0 for each k. But now write out equation (D.14) in full detail, showing all vector components, 0 = Σk=1N Σj=13 [ (Rk)j - Σi=1C λi (Bik)j] (δrk)j . (D.15) Next, relabel the 3N terms in this sum using a single index n defined in this way (k,j) = (1,1), (1,2), (1,3), (2,1), (2,2), (2,3),..........(N,1), (N,2), (N,3) n = 1, 2, 3, 4, 5, 6 , 3N-2, 3N-1, 3N so (D.16) n = 3(k-1)+j k = 1+ Int[(n-1)/3] j = 1 + Rem[(n-1)/3] . Then define (here the n value is implied by the k,j labels as per above) xn ≡ (rk)j Rn ≡ (Rk)j Bin ≡ (Bik)j . (D.17) Equation (D.15) can then be expressed as 0 = Σn=13N [ Rn - Σi=1C λi Bin] δxn (D.18) and (D.9) becomes 0 = Σk=1N Σj=13 (Bik)j (δrk)j or 0 = Σn=13N Bin δxn . i = 1,2...C . (D.19) Once again, we cannot set the square brackets in (D.18) to 0 because the 3N coordinates δxn are not linearly independent due to the C equations (D.19). At any given time t, due to these equations (D.19), there must exist at least one subset of the 3N variables δxn which are independent and the remaining C δxn are dependent. Assume that the dependent variables have indices n = n1, n2,...nC. Consider then the following set of C equations involving the dependent-term square brackets in (D.18), [ Rn - Σi=1C λi Bin] = 0 n = n1, n2,...nC . (D.20) This represents C linear equations in the C parameters λi so, when all is said and done, there must exist a set of values {λi} which make these C equations be true. We assume these values will be λi(t) and this is justified below when the full set of equations is considered. Now the terms in (D.18) indicated by n = n1, n2,...nC all vanish due to (D.20), so (D.18) can be written, 0 = Σn≠n,n,....n [ Rn - Σi=1C λi Bin] δxn . (D.21) But all the δxn appearing in (D.21) are independent, so the square brackets here must also vanish. We then end up with all the square brackets in (D.18) being 0 (assuming the correct solution values of λi), [ Rn - Σi=1C λi Bin] = 0 n = 1,2....3N or Rn = Σi=1C λi Bin n = 1,2....3N or (Rk)j = Σi=1C λi (Bik)j k = 1,2..N j = 1,2,3 or Rk = Σi=1C λi Bik k = 1,2...N . (D.22) Thus the constraint forces are certain linear combinations of the Bik constraint functions shown in (D.7). Meanwhile, Newton's Law (D.10) says that mkak = Fk + Rk where the Fk are the non-constraint forces for the problem. Inserting (D.22) gives mkk = Fk + Σi=1C λiBik Bik = (D.23) The problem is then summarized in the following set of equations from (D.23) and (D.6), mkk(t) = Fk(r1, r2..... rN, t) + Σi=1C λi(t)Bik(r1, r2..... rN, t) k = 1..N 3N equations k(t) = vk(t) k = 1..N 3N equations Σk=1N Bik(r1, r2..... rN, t) k(t) + Bit(r1, r2..... rN, t) = 0 i = 1,2....C . C equations (D.24) This can be regarded as a set of 6N+C scalar first-order ODEs in the single variable t for the 6N+C unknown scalar functions rk(t), vk(t) and λi(t). Although the first 3N equations happen to be linear in the λi(t), the equations are in general non-linear in the rk(t) so this is then a non-linear system of ODE's. Nevertheless, a well known existence theorem (e.g., Ince Section 3.3) states that, provided the Fk, Bik and Bit functions are continuous and differentiable in all their arguments, a unique solution exists for any set of initial values rk(0), vk(0) and λi(0). As (D.22) shows, the λi(0) are related to the constraint forces at time 0. The uniqueness of the solution agrees with our intuition that a physical system will evolve in a unique manner. Finding an analytic form for the solution might not be possible, but a numeric form can always be obtained. Summary of the above Lagrange Multiplier Application We have used a set of Lagrange Multipliers as an aid to solving a mechanics problem involving N particles with C constraints. As already noted, the above presentation is based on Section IV of Ray and Shamanna. The steps were as follows, considered more generically: 1. One has a sum Σn=1NRnδxn = 0. (D.25) 2. One also has a set of sums Σn=1NBinδxn = 0 for i = 1,2...C with C < N (constraints in the above case). This implies that one can find C of the δxn which are linearly dependent on the other δxn. 3. If λi are C arbitrary factors, certainly Σi=1Cλi(Σn=1NBinδxn) = 0 based on 2. These λi are "Undetermined Lagrange Multipliers". 4. One constructs from steps 1 and 3 the new sum Σn=1N[ Rn - Σi=1Cλi Bin] δxn = 0. 5. The C λi are selected so that [..] = 0 in the C dependent δxn terms of this sum. The remaining sum has only independent δxn terms so [..] = 0 for all those terms as well. Then all [...] = 0 in the sum of step 4. 6. The conclusion is that [ Rn - Σiλi Bin] = 0 for all n, assuming the solution λi values of step 5. In contrast with our main topic of finding stationary points df = 0 of a function f(r), in the Lagrange Multiplier application just presented there is no explicit function f(r) we are making stationary. On the other hand, one can regard (D.11) that ΣkRk δrk = 0 as the statement δW = 0 where δW = ΣkRk δrk is the total differential "virtual work" done by the constraint forces during a differential virtual displacement {δrk}. Using Newton's law (D.10) this equation can then be written 0 = δW = Σk[mkak - Fk] δrk (D.26) in which form it is known as D'Alembert's Principle (1743, Goldstein (1-42)). The analog of f(r) is then an action integral W of the differential virtual work δW which integral one renders stationary to get δW = 0. The notion of "the action" is mentioned at the end of this appendix.) We now present another application of the Lagrange Multiplier technique outlined in (D.25). As before, a certain amount of background is needed before the application can be demonstrated. Generalized coordinates, generalized forces, and Lagrange's Equations Usually the Cartesian components of the particle positions rk are not the most convenient variables to use in solving a problem. Angles are often more useful. With C holonomic constraints one can replace the 3N variables {xn} in (D.17) with a set of 3N-C independent variables {qn} where xn = xn(q1, q2, ....q3N-C,t) for n=1 to 3N (D.27) and one can then rework the above presentation in these new generalized coordinates qj which then won't always have the dimensions of length. For example, one then has, for virtual displacements, δxn = Σj=13N-C(∂xn/∂qj) δqj n = 1,2..3N or δrk = Σj=13N-C(∂rk/∂qj) δqj k = 1,2..N // as in Goldstein (1-44). (D.28) The usual presentation of Lagrangian dynamics with holonomic constraints avoids dealing with the forces of constraint and also avoids the need for using a set of Lagrange Multipliers. That presentation appears in Section V of the Ray and Shamanna and Goldstein p 17 (and many other places) and proceeds like this: 1. Start with ΣkRk δrk = 0 in (D.11) which we write in the form of D'Alembert's Principle (D.26) already mentioned above, 0 = Σk=1M [mkak - Fk] δrk (D.26) where Fk is the total non-constraint force acting on particle k. 2. Install (D.28) into D'Alembert's Principle to get 0 = Σk [mkak - Fk] Σj()δqj = Σj { Σkmkak - Σk Fk () } δqj = Σj { Σkmkak - Qj } δqj (D.29) where Qj ≡ Σk=1M Fk () . (D.30) Qj defined above are the generalized forces associated with the "generalized coordinates" qj . 3. Make use of the following non-obvious result (derived below), Σkmkak = () - (D.31) where T is the total kinetic energy of the N particles, T ≡ (1/2)Σkmk (vk vk) . (D.32) Then (D.29) reads, 0 = Σj { () - - Qj } δqj . (D.33) Since the δqj are independent variables, {..} = 0 and we end up with one form of "Lagrange's Equations" () - = Qj j = 1,2....3N-C . // as in Goldstein (1-50) (D.34) If the applied forces can be derived from a potential V(r1, r2, ..... rN) such that Fk = - (k)V(r1, r2, ..... rN) // no t or i dependence in V (D.35) then from (D.30), Qj ≡ Σk Fk () = - Σk ( ) () = – (D.36) where V(qi) ≡ V(rk(qi)) is a function of the generalized coordinates qi. Since V is not a function of the i we know that = 0. Using this fact, and (D.36) that Qj = – , (D.34) can be written () - = 0 . (D.37) The last step is to define the classical Lagrangian L ≡ T - V (D.38) to obtain () - = 0 j = 1,2....3N-C . // as in Goldstein (1-53) (D.39) which is the more conventional form of "Lagrange's Equations" with C holonomic constraints. One can define generalized (canonical, conjugate) momenta pj ≡ so the Lagrange equations (D.39) are just j = or = (q)L . If V = V(qi), then pj = and then if T = Σkmkk2 one has the familiar result pj = mjqj. When T = T(i) one finds Qj ≡ - = = j so = Q, which is the generalized Newton's Law in terms of generalized coordinates and generalized forces. ok to here Footnote: Derivation of the result (D.31) used above Define the 3N-C independent generalized coordinates qi as in (D.27) according to rk = rk(q1, q2....q3N-C,t) ≡ rk(qi,t) . (D.40) Compute the total time derivative of rk , vk = k = = Σj + = Σj j + // = vk(qi, i, t) . (D.41) Then based on this result compute these two partial derivatives of vk , = Σj j + (D.42) = . (D.43) Next, the total kinetic energy of all k particles is given by (D.32), T ≡ (1/2) Σk mk (vk vk) (D.32) so that = Σk mk vk = Σk mk vk . // using (D.43) (D.44) Apply d/dt to the above, with ak = k : () = Σk mk ak + Σk mk vk () . (D.45) Compute the total time derivative appearing in the second term, () = Σj j + . (D.46) Meanwhile, = Σk mk vk = Σk mk vk [ Σj j + ] // (D.42) = Σk mk vk () . // (D.46) (D.47) Subtract (D.47) from (D.45) to get () - = Σk mk ak (D.48) which is the result quoted above in (D.31). Lagrange Equations with non-holonomic constraints This is our final Lagrange Multiplier example, and it exactly follows the generic prescription of (D.25). The general topic is well described in Goldstein Chapter 2. As the starting point we invoke Hamilton's Principle which is this : δS = 0 where S = !Syntax Error, Idt L(qn(t), n(t), t ) (D.49) where there are M generalized coordinates qn. Here L = T - V is the Lagrangian (D.38) with T( i) and V(qi) being the kinetic and potential energies associated with the M generalized coordinates. We shall assume that M = N-D where there were D holonomic constraints on N initial problem coordinates. As shown in Goldstein p 41 (2-20') based on p 37 (2-15), the equation δS = 0 results in the following sum being 0, 1. Σn=1M [ !Syntax Error, Idt { () - ) } ] δqn(t) = 0 . (D.50) The number 1. on the left refers to the first step outlined in (D.25). In this equation the δqn(t) are M independent functions of t ("path variations"), restricted only by δqn(t1) = δqn(t2) = 0. In order for (D.50) to be true for M arbitrary independent functions δqn(t), one must have {} = 0, () - = 0 (D.51) which are the Lagrange Equations (D.39). Thus it is that the Lagrange Equations can be derived from Hamilton's Principle in the case of only holonomic constraints. But suppose that in addition to the D holonomic constraints there are C non-holonomic constraints, so only M-C of the δqn are independent. In this case the above conclusion is incorrect. 2. The C non-holonomic constraints result in C sums Σn Ani δqn = 0 i = 1,2...C analogous to (D.9) above. 3. If λi are C arbitrary factors, certainly Σiλi(ΣnAinδqn) = 0 based on 2. These λi are "Undetermined Lagrange Multipliers". Then !Syntax Error, Idt ( ΣiλiΣnAinδqn ) = 0 as well. 4. Construct from 1 and 3 the new sum, Σn=1M [ !Syntax Error, Idt { () - - ΣiλiAin } ] δqn(t) = 0 // as in Goldstein (2-26) (D.52) 5. The C λi are selected so that {..} = 0 in the C dependent δqn(t) terms of this sum. The remaining sum has only independent δqn(t) terms so {..} = 0 for all those terms as well. Then all {...} = 0 in the sum of 4. 6. The conclusion is that () - = ΣiλiAin . (D.53) This then is the corrected form of the Lagrange Equations in the presence of non-holonomic constraints, as in Goldstein (2-30). Comparing (D.53) with (D.34), one may interpret Q'n ≡ ΣiλiAin as the generalized forces that would have to be present to simulate the effect of the non-holonomic constraints were they not present. The forces for any holonomic constraints are already accounted for by L as in (D.39). The integral S shown in (D.49) is called "the action" and Hamilton's Principle is the classical mechanics incarnation of the Principle of Least Action. In relativistic classical and quantum field theory the action is a spacetime integral of the Lagrangian density L (φn(xμ), ∂μφn(xμ), xμ) where the coordinates qn of L are replaced by fields φn(xμ), and t is replaced by xμ, a point in spacetime. **************** E. L. Ince, Ordinary Differential Equations (Longmans, Green & Co. (Ltd.), London, 1927). This classic text became a Dover book in 1956 and is available online used for $4, shipping included (576p). S. Ray and J. Shamanna, "On Virtual Displacement and Virtual Work in Lagrangian Dynamics", European Journal of Physics, 27 (2006) 311-329. Also https://arxiv.org/abs/physics/0510204v2 . Ray and Shamanna