Appendix D
DOCX · 93.5 KB
Open DOCX file
Appendix D of a larger document on Lagrange multipliers, apparently written by Phil and marked as proofed on 9.26.16. It contrasts holonomic and non-holonomic constraints, then follows Section IV of a 2003 paper by Ray and Shamanna. It derives virtual displacements, assumes constraints do no virtual work, and uses a matrix partition of the constraint gradients to solve for the multipliers. Examples include a triangle of particles and masses on a ramp joined by a stick.
AI-written summary; may contain errors. This description is approximate.
Extracted text (machine-read; may contain errors)
Proofed in full at 3PM 9.26.16
Appendix D: Lagrange Multipliers for problems with Holonomic Constraints
In the Comment 4 at the end of Section 2 we note that sometimes the constraints on a system with "generalized coordinates" qk can be written Σkaikdqk + aitdt = 0. Such constraints written in terms of differentials of the coordinates are examples of non-holonomic constraints. This situation arises typically when constraints act on velocities rather than positions. An example is a vertical disk rolling on a horizontal plane as shown in Goldstein p 13. Such constraints are difficult to handle because the coordinates qk are not generally independent, meaning that the values of all the qk cannot be set independently. There is no constraint equation of the form a(q1,q2.....) = 0 which is a function of the whole coordinates qi; differential constraints are "non-whole". Inequality constraints are also non-holonomic, such as for particles bounded inside a box.
Nevertheless, as shown in our Comment 4, there is a role to be played by "Lagrange Multipliers". There is one such multiplier λi for each constraint and the values of the λi are determined as part of the problem solution.
Our main focus in this document is the use of Lagrange Multipliers to find candidates ri for which a scalar function f(r) is stationary, df = 0, subject to constraints. This is application of Lagrange multipliers is distinct from that outlined in Comment 4 where there is no obvious function f(r) which we are trying make stationary.
However, by extending the notation of stationarity one can first define a certain function of the possible solution trajectories qk(t) for a problem. That function is called "the action" S which is the time integral of the Lagrangian L(qk,k, t). The Lagrange equations of motion (including constraints) can be generated by setting δS = 0, a fact known as Hamilton's Principle (Goldstein Chapter 2 or the web). This notion is also used in the path integral formulations of quantum mechanics and quantum field theory. But these action applications are really not the same as the relatively simple problem addressed in this document concerning finding extrema of some scalar function f(r) .
A constraint of the form a(q1,q2....qN, t) = 0 is referred to as a holonomic constraint, and our present document deals only with such constraints.
We shall now outline another mechanics application of Lagrange multipliers, again unrelated to making some obvious function f(r) be stationary. The presentation below follows Section IV of a 2003 paper by Ray and Shamanna (R&S). Usually one wants to remove constraint forces from an analysis, but here they are instead the focus of the approach. In any event, this is just another example of the use of Lagrange Multipliers.
Note: In this appendix subscripts are just labels, they do not indicate differentiation as we used elsewhere in this document. For example, fi does not mean ∂f/∂xi, it just refers to the ith constraint function. Similarly rk is the 3D coordinate of the kth particle.
An application of Lagrange Multipliers
Consider a system of M particles each having a 3D coordinate rk, so there are N = 3M scalar variables. Assume there is a set of C holonomic constraints fi(r1, r2,,,,rM, t) = 0 for i = 1,2...C. These C constraint surfaces each of dimension 3M-1 in E3M intersect in some surface of dimension 3M-C in E3M which then restricts the positions of the particles, as outlined in Appendix C: only certain {r1, r2...rM} are allowed in any solution of the problem.
For example, suppose there are M=3 particles constrained to lie at the vertices of an equilateral triangle of edge 3, and in addition particle 1 is only allowed to be on a sphere of radius 5. The four constraints then are | r1- r2 | - 3 = 0, | r2- r3 | - 3 = 0, | r3- r1 | - 3 = 0, and (r1)12+ (r1)22+(r1)32- 52 = 0. Each such constraint fits the holonomic template form fi(r1, r2, r3, t) = 0. Each constraint is a surface of dimension 8 in E9 (there are many extrusions as discussed in Appendix C), and the intersection surface has dimension 5 in E9. One could count these 5 degrees of freedom as follows: location of particle 1 on the sphere (2) [perhaps θ.φ] and orientation of the triangle in space (3) [three Euler angles].
In Newtonian mechanics, the equation of motion of a particle may be written Fk + Gk = mkak where Gk is the total force of the constraints acting on particle k and Fk is the total of all other forces acting on this particle. Parameter mk is the particle's mass and ak is its vector acceleration.
A simple example of a holonomic constraint for one particle (particle 1) has that particle sliding down a frictionless curvy water-slide type surface (but no water) in the presence of gravity. In that case, F1 is the force of gravity acting on the particle, while G1 is the force of the slide surface acting on the particle. This force is normal to the instantaneous path of the particle so then G1 dr1 = 0, which can also be written in terms of the particle's velocity G1 v1 = 0 .
Now differentiate the C constraint equations fi(r1, r2,,,,rM, t) = 0 to get
dfi = 0 = Σk=1M (k)fi drk + (∂tfi) dt i = 1,2...C (D.1)
where (k) ≡ ( ∂/∂[(rk)1], ∂/∂[(rk)2], ∂/∂[(rk)3)]) is the 3D gradient with respect to the vector coordinate rk. Let {drk} for k = 1,,M be a set of differential displacements which satisfy equations (D.1), and call this a set of allowed displacements. Let {dr'k} be some other set of allowed displacements so that
dfi = 0 = Σk=1M (k)fi dr'k + (∂tfi) dt i = 1,2...C . (D.2)
Now define δrk ≡ drk - dr'k to obtain a set {δrk} of so-called virtual displacements. Such a set is the difference of any two different sets of allowed displacements. Then subtracting the above two equations gives
0 = Σk=1M (k)fi δrk i = 1,2...C . (D.3)
This equation is true only as a sum, the individual terms in general do not vanish. Now multiply each of the C sums in (D.3) by a separate (and arbitrary) Lagrange Multiplier λi and add the sums together to get
0 = Σi=1C λi [Σk=1M (k)fi δrk]
or
0 = Σk=1M [Σi=1C λi (k)fi] δrk . (D.4)
This method of defining the meaning of a virtual displacement set follows the prescription (5) of R&S.
Assumption that constraints do no virtual work
Next, assume that
0 = Σk=1M Gk δrk . (D.5)
In Goldstein this equation corresponds to setting the second term in (1-40) to zero. The main justification for making this assumption is that it is valid for particles which form a rigid body and for many other constraint situations. There seems to be no general justification of this equation for all constraint situations, and it is certainly not valid if friction is present. The equation can be interpreted as saying that the "total virtual work" done by the constraints on all the particles of the system vanishes.
As a first simple example, consider two particles 1 and 2 independently sliding down a frictionless static ramp. In this case for particle 1 the constraint force G1 is the normal force of the ramp acting on the particle, and this is clearly perpendicular to dr1 for any allowed displacement set {drk}, and thus G1 is also perpendicular to any δr1 which is part of any virtual displacement set {δrk}. In this case we have G1 δr1 = 0 and G2 δr2 = 0 so the terms in (D.5) are separately zero.
A second example is more interesting. Particles 1 and 2 of arbitrary masses slide down a one dimensional ramp but are glued to the two ends of a massless stick so they comprise a rigid body. The upper mass is attached to a string which provides a constant tension T. Here is a picture,
(D.6)
Particle 2 exerts a constraint force T12 on particle 1 through the stick, and Particle 1 exerts a constraint force T21 on particle 2. Since there is only one force of tension (or compression) in such a stick, T12 = T21. All allowed displacement sets {dr1, dr2} are of the form {dr, dr}, and therefore all virtual displacement sets are of the form {δr1, δr2} = {δr, δr}. In (D.6) we have drawn the two normal forces Ni and the two total constraint forces Gi. It is no longer true that G1 δr1 = 0 and G2 δr2 = 0 as the picture shows. However, it is true that Σk=12 Gk δrk = [ Σk=12 Gk ] δr = 0 because G1 and G2 have equal and opposite components along the ramp, so equation (D.5) is valid for this example.
Now suppose the ramp is translating in some direction not along the ramp. The displacements dr1= dr2 are then no longer along the ramp so that Σk=12 Gk drk ≠ 0. However, any virtual displacement set would still have δr1 = δr2 along the ramp and then Σk=12 Gk δrk = 0. Remember that a virtual displacement set is the difference of any two allowed displacement sets, and in this subtraction the effect of the ramp translation cancels out. So equation (D.5) is still valid.
Other examples are given in R&S Section III.
Reader Exercise: show in the last example that T12 = T m2 / (m1+m2) and that this makes sense if m1 >> m2 and vice versa.
We proceed then assuming our system of interest respects equation (D.5).
Resume Lagrange Mulitpliers application
Now subtract equation (D.4) from (D.5) to get,
0 = Σk=1M [ Gk - Σi=1C λi (k)fi] δrk . (D.7)
We know that the M vectors in the set {δrk} are not linearly independent because they must satisfy the equations (D.3), so we cannot claim from (D.7) that [ Gk - Σi=1C λi (k)fi] = 0 for each k. But now write out equation (D.7) in full detail, showing all vector components,
0 = Σk=1M Σj=13 [ (Gk)j - Σi=1C λi ] (δrk)j . (D.8)
Next, relabel the 3M terms in this sum using a single index n defined in this way
(k,j) = (1,1), (1,2), (1,3), (2,1), (2,2), (2,3),..........(M,1), (M,2), (M,3)
n = 1, 2, 3, 4, 5, 6 , 3M-2, 3M-1, 3M (D.9)
so that n = 3(k-1)+j. Then define
Gn ≡ (Gk)j xn ≡ (rk)j . // n = 3(k-1)+j (D.10)
Equation (D.8) can then be expressed as
0 = Σn=13M [ Gn - Σi=1C λi ] δxn . (D.11)
Once again, we cannot set the square brackets to 0 because the 3M coordinates δxn are not linearly independent.
]
Using the same notation we can write (D.3) as
0 = Σk=1M Σj=13 (δrk)j i = 1,2...C .
or
0 = Σn=13M δxn i = 1,2...C (D.12)
It will be convenient to define the following C x (3M) matrix A
Ain = i = 1,2...C n = 1,2...3M (D.13)
We then define two submatrices A and B of matrix A as follows,
(D.14)
so that C x C matrix A is the upper left corner of A and C x 3M matrix B is the rest of the top strip. In terms of matrix A equations ** and ** become
0 = Σn=13M [ Gn - Σi=1C λi Ain ] δxn . (D.15)
0 = Σn=13MAinδxn i = 1,2...C (D.16)
For simplicity, suppose it happens that the last 3M-C of the δxn can be taken to be the independent variables. We can then write (D.15) as
Σn=1CA1nδxn = - Σn=C+13MA1nδxn
Σn=1CA2nδxn = - Σn=C+13MA2nδxn
.....
Σn=1CACnδxn = - Σn=C+13MACnδxn (D.17)
where the independent variables appear on the right side and the dependent ones on the left. Using the above defined matrices A and B we write the above equations in matrix form as
A = = - B (D.18)
Assuming det(A) ≠ 0 we solve, introducing another C x 3M matrix D, to get
= - A-1 B ≡ M D ≡ - A-1 B
(D.19)
so that
δxi = Σn=C+13M Dijδxj i = 1,2...C (D.20)
so that each dependent δxi on the left is some linear combination of the independent δxn on the right, which we abbreviate by writing the above as
δxi = Σn=C+13M Bin δxn
If we regard the independent variables as "known" and the dependent ones as "unknown", we can solve the above equation as a Cramer's Rule problem to get expressions for the dependent variables
δxi =
δxj = - ()-1 Σn≠j δxn j = 1,2...C and i = 1,2...C (D.13)
or
δxj = Σn≠j { - ()-1 } δxn = Σn≠j Aijn δxn
Now we shall determine the Lagrange Multipliers. Consider this set of C equations,
[ Gn - Σi=1C () λi] = 0 n = 1,2...C (D.12a)
which involve the first C square-bracket quantities in (D.11).
Define the CxC matrix A by Ani ≡ so the above may be written as a matrix equation,
= A = A-1 ≡ B . (D.12b)
Assuming detA ≠ 0 we have found solution values for the C Lagrange Multipliers λi which are linear functions of the lowest C of the Gi
λi = Σj=1C BijGj = λi(G1, G2....GC)
Install these λi into (D.12a) to get
[ Gn - Σi=1C () Σj=1C BijGj] = 0 n = 1,2...C (**)
or
[ Gn - Σj=1C {Σi=1C () Bij} Gj] = 0 n = 1,2...C (**)
or
[ Gn - Σj=1C Qnj Gj] = 0 n = 1,2...C Qnj ≡ Σi=1C () Bij (**)
Suppose C = 2. The above equations are then
G1 - [ Q11G1 + Q12G2] = 0
G2 - [ Q21G1 + Q22G2] = 0
or
(1-Q11)G1 - Q12G2 = 0
-Q21G1 - (1-Q22)G2 = 0
In general, the solution of these two equations is G1 = 0 and G2 = 0, so then λi = 0 and all Gi = 0.
I have assumed a special case where the first C of the δxi are dependent and the last 3M-c are independent. Seems to me that could happen. If it does, then there must be no constraint forces!
Since the first C terms in (D.11) now vanish, that equation becomes,
0 = Σn=C+13M [ Gn - Σi=1C λi ] δxn . (D.13)
But the 3M-C variables δxn appearing in (D.13) are independent, so each [...] in (D.13) must vanish. We then end up with all the coefficients of δxn vanishing,
[ Gn - Σi=1C λi ] = 0 n = 1,2....3M
or
Gn = Σi=1C λi(G1, ....GC) n = 1,2....3M . (D.14)
or
Gn = Σi=1C [Σj=1C BijGj] n = 1,2....3M . (D.14)
= Σj=1C { Σi=1C Bij } Gj
= Σj=1C Qnj Gj
Suppose C = 2, and M = a large number. Then we have
G1 = Σj=1C Q1j Gj = Q11G1 + Q12G2
G2 = Σj=1C Q2j Gj = Q21G1 + Q22G2
G3 = Σj=1C Q3j Gj = Q31G1 + Q32G2
G4 = Σj=1C Q4j Gj = Q41G1 + Q42G2
... = etc
The first two equations
The first two equations
or
(D.14) represents 3M easily solvable linear equations in the 3M unknown constraint forces Gn.
Something wrong. Only solution would be all Gi = 0 !!
In this manner we have managed to determine all the constraint forces Gn ≡ (Gk)j in the problem. If the constraints fi have explicit time dependence, then Fni ≡ will be time dependent and so will the Gk. From Newton's Law Fk + Gk = mkak where Fk are any "applied" forces, one then knows all the ak(t) which can then be time-integrated to obtain vk(t) and then rk(t).
To summarize, we have used an application of Lagrange Multipliers as an aid to solving a mechanics problem involving N particles with C holonomic constraints. As already noted, the above presentation is based on Section IV of R&S. That paper comments (at great length) on the often vague term "virtual displacement" as it appears in Goldstein and other standard texts, and we have followed their suggestion for defining this term as a difference of "allowed displacements" which then removes the dt time term in (D.1).
In contrast with our main topic of finding stationary points df = 0 of a function f(r), in the Lagrange Multiplier application just presented there is no explicit function f(r) we are making stationary. On the other hand, one can regard (D.5) that ΣkGk δrk = 0 as the statement δW = 0 where δW = ΣkGk δrk is the differential "virtual work" done by the constraint forces during a differential virtual displacement. This equation can then be written 0 = δW = Σk[mkak - Fk] δrk in which form it is known as D'Alembert's Principle (1743, Goldstein (1-42)). The analog of f(r) is then presumably some action-like integral W of the virtual work δW.
Generalized coordinates, generalized forces, and Lagrange's Equations
Usually the Cartesian components of the particle positions rk are not the most convenient variables to use in solving a problem. For example, in the triangle example noted above one could use 5 angles to describe the state ("configuration") of the constrained system. One can replace the 3M non-independent variables {xn} in (D.10) with a set of 3M-C independent variables {qn} where xn = xn(q1, q2....q3M-C,t) for n=1 to 3M, and one can then rework the above presentation in these new "generalized coordinates" qj which then won't always have the dimensions of length. For example, one then has δxn = Σj=13M-C(∂xn/∂qj) δqj for n = 1 to 3M, or alternatively, δrk = Σj=13M-C(∂rk/∂qj) δqj as in Goldstein (1-44).
The usual presentation of Lagrangian dynamics with holonomic constraints avoids dealing with the forces of constraint and also avoids the need for using a set of Lagrange Multipliers. That presentation appears in Section V of the R&S and Goldstein p 17 (and many other places) and proceeds like this:
1. Start with ΣkGk δrk = 0 in (D.5) which we write in the form of D'Alembert's Principle already mentioned above,
0 = Σk=1M [mkak - Fk] δrk (D.15)
where Fk is the total non-constraint force acting on particle k (sometimes called the "applied" force).
2. Make the replacement just noted above to convert to 3M-C independent generalized coordinates,
δrk = Σj=13M-C () δqj where rk = rk(q1, q2....q3M-C,t) (D.16)
3. Insert this into (D.15) to get
0 = Σk [mkak - Fk] Σj()δqj
= Σj { Σkmkak - Σk Fk () } δqj
= Σj { Σkmkak - Qj } δqj Qj ≡ Σk=1M Fk () (D.17)
where Qj is the "generalized force" associated with the generalized coordinate qj .
4. Make use of the following non-obvious result (derived below),
Σkmkak = () - (D.34)
where
T ≡ (1/2)Σkmk (vk vk) (D.18)
is the total kinetic energy of the M particles. Then (D.17) reads,
0 = Σj { () - - Qj } δqj . (D.19)
Since the δqj are independent variables, {..} = 0 and we end up with one form of "Lagrange's Equations"
() - = Qj j = 1,2....3M-C . // as in Goldstein (1-50) (D.20)
If the applied forces can be derived from a potential V such that
Fk = - (k)V(r1, r2, ..... rM) (D.21)
then from (D.17),
Qj ≡ Σk Fk () = - Σk ( ) () = – (D.22)
where V(qi) = V(rk(qi)) is a function of the generalized coordinates qi. Since V is not a function of the
i we know that = 0. Using this fact and Qj = – in (D.20) gives
() - = 0 . (D.23)
The last step is to define the classical Lagrangian L ≡ T - V to obtain
() - = 0 j = 1,2....3M-C . // as in Goldstein (1-53) (D.24)
which is the more conventional form of "Lagrange's Equations" with C holonomic constraints.
Derivation of the result (D.34) used above
Define the 3M-C independent generalized coordinates qi according to
rk = rk(q1, q2....q3M-C,t) ≡ rk(qi,t) . (D.25)
Compute the total time derivative of rk ,
vk = k = = Σj + = Σj j + = vk(qi, i, t) . (D.26)
Then based on this result compute these two partial derivatives of vk ,
= Σj j + (D.27)
= . (D.28)
Next, define total kinetic energy of all k particles,
T ≡ (1/2) Σk mk (vk vk) (D.29)
so that
= Σk mk vk = Σk mk vk . // (D.28) (D.30)
Applying d/dt to the above then gives,
() = Σk mk ak + Σk mk vk () . (D.31)
Computing the total time derivative in the second term gives,
() = Σj j + . (D.32)
Meanwhile,
= Σk mk vk
= Σk mk vk [ Σj j + ] // (D.27)
= Σk mk vk () . // (D.32) (D.33)
Then subtracting (D.33) from (D.31) we get our desired result,
() - = Σk mk ak . (D.34)