Appendix D v2
DOCX · 174.6 KB
Open DOCX file
Appendix D of Phil's document on Lagrange multipliers, in a second version. It sets up holonomic and non-holonomic constraints, allowed and virtual displacements, and the assumption that constraint forces do no virtual work, with ramp and rigid-stick examples. It then introduces multipliers to show each constraint force equals a sum of multipliers times constraint coefficients. Only the first part of the text was seen.
AI-written summary; may contain errors. This description is approximate.
Extracted text (machine-read; may contain errors)
Appendix D: Lagrange Multipliers for problems with Constraints
Our main focus in this document is the use of Lagrange Multipliers to find candidates ri for which a scalar function f(r) is stationary, df = 0, subject to constraints. In this Appendix, however, we consider a different application of Lagrange Multipliers that relates to classical mechanics for a system of N particles in the presence of constraints. A certain amount of background information is required to put this application into context, which we now present.
Constraints and virtual displacements
Let rk(t) be the position of the kth particle. At any fixed time t, if the particles are unencumbered with constraints, the variables (r1, r2..... rk) can be regarded as a single point in the space E3N and the "configuration space" of the system indicated by (r1, r2..... rk) is all of this space E3N. All points in E3N are "legal" points for the system.
A holonomic constraint has the form
a(r1, r2..... rN, t) = 0. (D.1)
At a given time t, this equation represents a surface of dimension 3N-1 in EN. If there are s such holonomic constraints
fi(r1, r2..... rN, t) = 0 i = 1,2...s (D.2)
then at time t the system vector (r1, r2..... rN) is constrained to lie on a surface of dimension 3N-s in EN, as discussed in Appendix C, so each extra constraint lowers the operating surface dimension by 1. One can think of the term holonomic as meaning that the constraints are functions of "whole" coordinates like r2 as opposed to differential coordinates like dr2.
A non-holonomic constraint is one that is not holonomic. An inequality constraint like a(r1, r2..... rN, t) < 0 which might bound the particles to one side of a surface is an example of a nonholonomic constraint, but our interest lies more in nonholonomic constraints of the form
Σk=1N bk drk + btdt = 0 (D.3)
in which differentials rather than "whole" coordinates appear. The "coefficients" bk and bt can in general be functions of (r1, r2..... rk, t). If it happens that such a constraint can be integrated by some hook or crook to give an equation of the form a(r1, r2..... rk, t) = 0, then it is regarded as being a holonomic constraint, so a non-holonomic constraint must be non-integrable. If there are m such non-holonomic constraints imposed on a system, one could write them as
Σk=1N Aik drk + Aitdt = 0 i = 1,2...m
or (D.4)
Σk=1N Aik k + Ait = 0 i = 1,2...m .
Here bolded Aik represents 3 matrices each of which is m x N (rows x columns). Since these constraints are non-integrable, one cannot really described them in terms of surfaces in E3N.
If one differentiates the holonomic constraints (D.2) one gets
Σk=1N [(k)fi] drk + [∂tfi] dt = 0 i = 1,2...s
or (D.5)
Σk=1N [(k)fi] k + [∂tfi] = 0 i = 1,2...s
where (k)≡ (∂/∂(rk)1, ∂/∂(rk)2, ∂/∂(rk)3) is a gradient with respect to the components of rk. These equations have the same form as those in (D.4), but these are holonomic because they are integrable to give fi(r1, r2..... rN, t) = 0.
For convenience below, we shall now combine (D.4) and (D.5) into a single set of equations:
Σk=1N Bik drk + Bitdt = 0 i = 1,2....C // C = s+m
or (D.6)
Σk=1N Bik k + Bit = 0 i = 1,2....C
where
Bik = Aik Bit = Ait for i = 1,2...s non-holonomic
Bit = (k)fi Bit = ∂tfi for i = s+1, s+2....C holonomic . (D.7)
In general, the functions fi, Bik and Bit are all functions of (r1, r2..... rk, t).
In the above we have followed the notation of Ray and Shamanna (R&S 2006) and we continue our path in the manner they propose.
Suppose the set of N differentials {drk} satisfies all of equations (D.6) for some time interval dt. Such a set {drk} is then called a set of allowed displacements. Imagine that {dr'k} is some other distinct allowed displacement set. Elements of the difference set defined as
{δrk} = {drk} - {dr'k} k = 1,2...N (D.8)
are then called virtual displacements. This is how R&S define the vague term virtual displacement which appears in many texts (about which they have much to say in their Section I A). It follows then that for any virtual displacement set {δrk} equations (D.6) take this form
Σk=1N Bik δrk = 0 (D.9)
simply because the last term in (D.6) cancels out when one writes δrk = drk - dr'k. Notice that the δrk are not arbitrary differential displacements. They are differences of allowed displacements.
Meanwhile, Newton's Law for a system of N particles subject to constraints has the form
mk = Fk + Rk k = 1,2...N (D.10)
where Rk is the sum of all constraint forces applied to particle k and Fk is the sum of all other forces applied to particle k.
Assumption that constraints do no virtual work
Next, assume that ( R&S refer to this as being an "ideal constraint" situation),
0 = Σk=1N Rk δrk . (D.11)
In Goldstein this equation corresponds to setting the second term in (1-40) to zero. The main justification for making this assumption is that it is valid for particles which form a rigid body and for many other constraint situations. There seems to be no general justification of this equation for all constraint situations, and it is certainly not valid if friction is present. The equation can be interpreted as saying that the total virtual work done by all the constraint forces on all the particles of the system vanishes.
As a first simple example, consider two particles 1 and 2 independently sliding down a frictionless static ramp. In this case for particle 1 the constraint force R1 is the normal force of the ramp acting on the particle, and this is clearly perpendicular to dr1 for any allowed displacement set {drk}, and thus R1 is also perpendicular to any δr1 which is part of any virtual displacement set {δrk}. In this case we have R1 δr1 = 0 and R2 δr2 = 0 so the terms in (D.11) are separately zero.
A second example is more interesting. Particles 1 and 2 of arbitrary masses slide down a one dimensional ramp but are glued to the two ends of a massless stick so they comprise a rigid body. The upper mass is attached to a string which provides a constant tension T. Here is a picture,
(D.12)
Particle 2 exerts a constraint force T12 on particle 1 through the stick, and Particle 1 exerts a constraint force T21 on particle 2. Since there is only one force of tension (or compression) in such a stick, T12 = T21. All allowed displacement sets {dr1, dr2} are of the form {dr, dr}, and therefore all virtual displacement sets are of the form {δr1, δr2} = {δr, δr}. In (D.12) we have drawn the two normal forces Ni and the two total constraint forces Ri. It is no longer true that R1 δr1 = 0 and R2 δr2 = 0 as the picture shows. However, it is true that Σk=12 Rk δrk = [ Σk=12 Rk ] δr = 0 because R1 and R2 have equal and opposite components along the ramp, so equation (D.11) is valid for this example.
Now suppose the ramp is translating in some direction not along the ramp. The displacements dr1= dr2 are then no longer along the ramp so that Σk=12 Rk drk ≠ 0. However, any virtual displacement set would still have δr1 = δr2 along the ramp and then again Σk=12 Rk δrk = 0. Remember that a virtual displacement set is the difference of any two allowed displacement sets, and in this subtraction the effect of the ramp translation cancels out. So equation (D.11) is still valid.
Other examples are given in R&S Section III.
Reader Exercise: show in the last example that T12 = T m2 / (m1+m2) and that this makes sense if m1 >> m2 and vice versa.
We proceed then assuming our system of interest respects equation (D.11).
Comment on internal forces: In (D.10) we have classified the forces acting on particle k into two groups and we write Rk + Fk, where Rk are "constraint" forces and the Fk are any "other" forces. Goldstein refers to the Fk forces as "applied" forces in (1-39), while R&S call them "external" in Sec IV. Consider a system consisting of N masses all connected by a network of massless springs. The force acting on particle k due to all the springs to which it is attached would be best classified in the Fk "other" category. Were we to classify this force into the Rk group, we could not use the no-constraint-work property (D.12) since spring forces do work as the springs compress and expand. An example of such a system is the Solar System where the spring forces are represented by gravity. It would seem a misnomer to classify the internal forces in this case as "applied" or "external". However, if all the springs are slowly adjusted until they become infinitely stiff, in that limit it would be better to classify the internal spring forces into the Rk constraint forces group, since these springs do no work and (D.11) applies. The system is then a "rigid body" and in this case, with the internal forces classified as Rk constraint forces, one benefits from the many simplifications which arise for rigid body mechanics. The internal forces in our spring example fall into the category where Fij = - Fji and Fij is in the direction ri- rk . When this is not the case, the rigid body limit does not apply, but the internal forces can still be classified as "constraint" plus "other".
An Application of Lagrange Multipliers
Recall the C = m+s constraint equations (D.9),
Σk=1N Bik δrk = 0 i = 1,2...C C ≡ m+s = total number of constraints (D.9)
Multiply each of these C equations by a Lagrange Multiplier λi and then add the equations to get
0 = Σi=1C λi [Σk=1N Bik δrk] = 0
= Σk=1N [ Σi=1C λi Bik ] δrk . (D.13)
Now subtract equation (D.13) from (D.11) to get
0 = Σk=1N [ Rk - Σi=1C λi Bik] δrk . (D.14)
We know that the N vectors in the set {δrk} are not linearly independent because they must satisfy the equations (D.9), so we cannot claim from (D.14) that [ Rk - Σi=1C λi Bik] = 0 for each k. But now write out equation (D.14) in full detail, showing all vector components,
0 = Σk=1N Σj=13 [ (Rk)j - Σi=1C λi (Bik)j] (δrk)j . (D.15)
Next, relabel the 3N terms in this sum using a single index n defined in this way
(k,j) = (1,1), (1,2), (1,3), (2,1), (2,2), (2,3),..........(N,1), (N,2), (N,3)
n = 1, 2, 3, 4, 5, 6 , 3N-2, 3N-1, 3N
so (D.16)
n = 3(k-1)+j k = 1+ Int[(n-1)/3] j = 1 + Rem[(n-1)/3] .
Then define (here the n value is implied by the k,j labels as per above)
xn ≡ (rk)j Rn ≡ (Rk)j Bin ≡ (Bik)j . (D.17)
Equation (D.15) can then be expressed as
0 = Σn=13N [ Rn - Σi=1C λi Bin] δxn (D.18)
and (D.9) becomes 0 = Σk=1N Σj=13 (Bik)j (δrk)j or
0 = Σn=13N Bin δxn . i = 1,2...C . (D.19)
Once again, we cannot set the square brackets in (D.18) to 0 because the 3N coordinates δxn are not linearly independent due to the C equations (D.19).
At any given time t, due to these equations (D.19), there must exist at least one subset of the 3N variables δxn which are independent and the remaining C δxn are dependent. Assume that the dependent variables have indices n = n1, n2,...nC. Consider then the following set of C equations involving the dependent-term square brackets in (D.18),
[ Rn - Σi=1C λi Bin] = 0 n = n1, n2,...nC . (D.20)
This represents C linear equations in the C parameters λi so, when all is said and done, there must exist a set of values {λi} which make these C equations be true. We assume these values will be λi(t) and this is justified below when the full set of equations is considered.
Now the terms in (D.18) indicated by n = n1, n2,...nC all vanish due to (D.20), so (D.18) can be written,
0 = Σn≠n,n,....n [ Rn - Σi=1C λi Bin] δxn . (D.21)
But all the δxn appearing in (D.21) are independent, so the square brackets here must also vanish. We then end up with all the square brackets in (D.18) being 0 (assuming the correct solution values of λi),
[ Rn - Σi=1C λi Bin] = 0 n = 1,2....3N
or
Rn = Σi=1C λi Bin n = 1,2....3N
or
(Rk)j = Σi=1C λi (Bik)j k = 1,2..N j = 1,2,3
or
Rk = Σi=1C λi Bik k = 1,2...N . (D.22)
Thus the constraint forces are certain linear combinations of the Bik constraint functions shown in (D.7).
Meanwhile, Newton's Law (D.10) says that mkak = Fk + Rk where the Fk are the non-constraint forces for the problem. Inserting (D.22) gives
mkk = Fk + Σi=1C λiBik Bik = (D.23)
The problem is then summarized in the following set of equations from (D.23) and (D.6),
mkk(t) = Fk(r1, r2..... rN, t) + Σi=1C λi(t)Bik(r1, r2..... rN, t) k = 1..N 3N equations
k(t) = vk(t) k = 1..N 3N equations
Σk=1N Bik(r1, r2..... rN, t) k(t) + Bit(r1, r2..... rN, t) = 0 i = 1,2....C . C equations
(D.24)
This can be regarded as a set of 6N+C scalar first-order ODEs in the single variable t for the 6N+C unknown scalar functions rk(t), vk(t) and λi(t). Although the first 3N equations happen to be linear in the λi(t), the equations are in general non-linear in the rk(t) so this is then a non-linear system of ODE's. Nevertheless, a well known existence theorem (e.g., Ince Section 3.3) states that, provided the Fk, Bik and Bit functions are continuous and differentiable in all their arguments, a unique solution exists for any set of initial values rk(0), vk(0) and λi(0). As (D.22) shows, the λi(0) are related to the constraint forces at time 0.
The uniqueness of the solution agrees with our intuition that a physical system will evolve in a unique manner. Finding an analytic form for the solution might not be possible, but a numeric form can always be obtained.
Summary of the above Lagrange Multiplier Application
We have used a set of Lagrange Multipliers as an aid to solving a mechanics problem involving N particles with C constraints. As already noted, the above presentation is based on Section IV of R&S. The steps were as follows, considered more generically:
1. One has a sum Σn=1NRnδxn = 0. (D.25)
2. One also has a set of sums Σn=1NBinδxn = 0 for i = 1,2...C with C < N (constraints in the above case). This implies that one can find C of the δxn which are linearly dependent on the other δxn.
3. If λi are C arbitrary factors, certainly Σi=1Cλi(Σn=1NBinδxn) = 0 based on 2. These λi are "Undetermined Lagrange Multipliers".
4. One constructs from steps 1 and 3 the new sum Σn=1N[ Rn - Σi=1Cλi Bin] δxn = 0.
5. The C λi are selected so that [..] = 0 in the C dependent δxn terms of this sum. The remaining sum has only independent δxn terms so [..] = 0 for all those terms as well. Then all [...] = 0 in the sum of step 4.
6. The conclusion is that [ Rn - Σiλi Bin] = 0 for all n, assuming the solution λi values of step 5.
In contrast with our main topic of finding stationary points df = 0 of a function f(r), in the Lagrange Multiplier application just presented there is no explicit function f(r) we are making stationary. On the other hand, one can regard (D.11) that ΣkRk δrk = 0 as the statement δW = 0 where δW = ΣkRk δrk is the total differential "virtual work" done by the constraint forces during a differential virtual displacement {δrk}. Using Newton's law (D.10) this equation can then be written
0 = δW = Σk[mkak - Fk] δrk (D.26)
in which form it is known as D'Alembert's Principle (1743, Goldstein (1-42)). The analog of f(r) is then an action integral W of the differential virtual work δW which integral one renders stationary to get δW = 0. The notion of "the action" is mentioned at the end of this appendix.)
We now present another application of the Lagrange Multiplier technique outlined in (D.25). As before, a certain amount of background is needed before the application can be demonstrated.
Generalized coordinates, generalized forces, and Lagrange's Equations
Usually the Cartesian components of the particle positions rk are not the most convenient variables to use in solving a problem. Angles are often more useful. With C holonomic constraints one can replace the 3N variables {xn} in (D.17) with a set of 3N-C independent variables {qn} where
xn = xn(q1, q2, ....q3N-C,t) for n=1 to 3N (D.27)
and one can then rework the above presentation in these new "generalized coordinates" qj which then won't always have the dimensions of length. For example, one then has, for virtual displacements,
δxn = Σj=13N-C(∂xn/∂qj) δqj n = 1,2..3N
or
δrk = Σj=13N-C(∂rk/∂qj) δqj k = 1,2..N // as in Goldstein (1-44). (D.28)
The usual presentation of Lagrangian dynamics with holonomic constraints avoids dealing with the forces of constraint and also avoids the need for using a set of Lagrange Multipliers. That presentation appears in Section V of the R&S and Goldstein p 17 (and many other places) and proceeds like this:
1. Start with ΣkRk δrk = 0 in (D.11) which we write in the form of D'Alembert's Principle (D.26) already mentioned above,
0 = Σk=1M [mkak - Fk] δrk (D.26)
where Fk is the total non-constraint force acting on particle k.
2. Install (D.28) into D'Alembert's Principle to get
0 = Σk [mkak - Fk] Σj()δqj
= Σj { Σkmkak - Σk Fk () } δqj
= Σj { Σkmkak - Qj } δqj (D.29)
where
Qj ≡ Σk=1M Fk () . (D.30)
Qj defined above is the "generalized force" associated with the "generalized coordinate" qj .
3. Make use of the following non-obvious result (derived below),
Σkmkak = () - (D.31)
where T is the total kinetic energy of the N particles,
T ≡ (1/2)Σkmk (vk vk) . (D.32)
Then (D.29) reads,
0 = Σj { () - - Qj } δqj . (D.33)
Since the δqj are independent variables, {..} = 0 and we end up with one form of "Lagrange's Equations"
() - = Qj j = 1,2....3N-C . // as in Goldstein (1-50) (D.34)
If the applied forces can be derived from a potential V(r1, r2, ..... rN) such that
Fk = - (k)V(r1, r2, ..... rN) // no t or i dependence in V (D.35)
then from (D.30),
Qj ≡ Σk Fk () = - Σk ( ) () = – (D.36)
where V(qi) ≡ V(rk(qi)) is a function of the generalized coordinates qi. Since V is not a function of the
i we know that = 0. Using this fact, and (D.36) that Qj = – , (D.34) can be written
() - = 0 . (D.37)
The last step is to define the classical Lagrangian
L ≡ T - V (D.38)
to obtain
() - = 0 j = 1,2....3N-C . // as in Goldstein (1-53) (D.39)
which is the more conventional form of "Lagrange's Equations" with C holonomic constraints.
ok to here
Footnote: Derivation of the result (D.31) used above
Define the 3N-C independent generalized coordinates qi as in (D.27) according to
rk = rk(q1, q2....q3N-C,t) ≡ rk(qi,t) . (D.40)
Compute the total time derivative of rk ,
vk = k = = Σj + = Σj j + // = vk(qi, i, t) . (D.41)
Then based on this result compute these two partial derivatives of vk ,
= Σj j + (D.42)
= . (D.43)
Next, the total kinetic energy of all k particles is given by (D.32),
T ≡ (1/2) Σk mk (vk vk) (D.32)
so that
= Σk mk vk = Σk mk vk . // using (D.43) (D.44)
Apply d/dt to the above, with ak = k :
() = Σk mk ak + Σk mk vk () . (D.45)
Compute the total time derivative appearing in the second term,
() = Σj j + . (D.46)
Meanwhile,
= Σk mk vk
= Σk mk vk [ Σj j + ] // (D.42)
= Σk mk vk () . // (D.46) (D.47)
Subtract (D.47) from (D.45) to get
() - = Σk mk ak (D.48)
which is the result quoted above in (D.31).
Lagrange Equations with non-holonomic constraints
This is our final Lagrange Multiplier example, and it exactly follows the generic prescription of (D.25). The general topic is well described in Goldstein Chapter 2.
As the starting point we invoke Hamilton's Principle which is this :
δS = 0 where S = !Syntax Error, Idt L(qn(t), n(t), t ) (D.49)
where there are M generalized coordinates qn. Here L = T - V is the Lagrangian (D.38) with T( i) and V(qi) being the kinetic and potential energies associated with the M generalized coordinates. We shall assume that M = N-D where there were D holonomic constraints on N initial problem coordinates. As shown in Goldstein p 41 (2-20') based on p 37 (2-15), the equation δS = 0 results in the following sum being 0,
1. Σn=1M [ !Syntax Error, Idt { () - ) } ] δqn(t) = 0 . (D.50)
The number 1. on the left refers to the first step outlined in (D.25). In this equation the δqn(t) are M independent functions of t ("path variations"), restricted only by δqn(t1) = δqn(t2) = 0. In order for (D.50) to be true for M arbitrary independent functions δqn(t), one must have {} = 0,
() - = 0 (D.51)
which are the Lagrange Equations (D.39). Thus it is that the Lagrange Equations can be derived from Hamilton's Principle in the case of only holonomic constraints.
But suppose that in addition to the D holonomic constraints there are C non-holonomic constraints, so only M-C of the δqn are independent. In this case the above conclusion is incorrect.
2. The C non-holonomic constraints result in C sums Σn Ani δqn = 0 i = 1,2...C analogous to (D.9) above.
3. If λi are C arbitrary factors, certainly Σiλi(ΣnAinδqn) = 0 based on 2. These λi are "Undetermined Lagrange Multipliers". Then !Syntax Error, Idt ( ΣiλiΣnAinδqn ) = 0 as well.
4. Construct from 1 and 3 the new sum,
Σn=1M [ !Syntax Error, Idt { () - - ΣiλiAin } ] δqn(t) = 0 // as in Goldstein (2-26) (D.52)
5. The C λi are selected so that {..} = 0 in the C dependent δqn(t) terms of this sum. The remaining sum has only independent δqn(t) terms so {..} = 0 for all those terms as well. Then all {...} = 0 in the sum of 4.
6. The conclusion is that
() - = ΣiλiAin . (D.53)
This then is the corrected form of the Lagrange Equations in the presence of non-holonomic constraints, as in Goldstein (2-30). Comparing (D.53) with (D.34), one may interpret Q'n ≡ ΣiλiAin as the generalized forces that would have to be present to simulate the effect of the non-holonomic constraints were they not present. The forces for any holonomic constraints are already accounted for by L as in (D.39).
The integral S shown in (D.49) is called "the action" and Hamilton's Principle is the classical mechanics incarnation of the Principle of Least Action. In relativistic classical and quantum field theory the action is a spacetime integral of the Lagrangian density L (φn(xμ), ∂μφn(xμ), xμ) where the coordinates qn of L are replaced by fields φn(xμ), and t is replaced by xμ, a point in spacetime.
END END END END END END END END END END END END END END
E. L. Ince, Ordinary Differential Equations (Longmans, Green & Co. (Ltd.), London, 1927). This classic text became a Dover book in 1956 and is available online used for $4, shipping included (576p).
S. Ray and J. Shamanna, "On Virtual Displacement and Virtual Work in Lagrangian Dynamics", European Journal of Physics, 27 (2006) 311-329. Also https://arxiv.org/abs/physics/0510204v2 .
********************************
Comments:
This is from Greenwood Google books.
So I have just found a newer (and published) version of the R&S document, and I printed it from the archiv and read it. It is basically the same except for several items:
1. The Intro and Conclusion sections are much longer than in the 2003 paper.
2. They have generalized their paper to handle s holo and m non-holo constraints. But all conclusions and sections are basically exactly the same. It's just that in certain places they have to state both forms of constraint equations. I am glad I found this later version.
*******************************
Questions on the above: We know for sure that [(k)fi] is a function of r1, r2..... rN, t so we might as well assume as well that Aik and Ait are all functions of r1, r2..... rN, t. The equations we then want to solve are these
mkk(t) = Fk(r1, r2..... rN, t) + Σi=1C λiBik(r1, r2..... rN, t) k = 1,2...N
fi(r1, r2, ...rM, t) = 0 i = 1,2...s
Σk=1N Aik(r1, r2..... rN, t) k(t) + Ait(r1, r2..... rN, t) = 0 . i = s+1 ...C
Are we supposed to think of λi as numbers, or are they λi(t), or are they λi(r1, r2..... rk, t) ???
My guess is that we want them to be λi(t). So how would one approach solving the above generic set of equations. I would replace the second group of equations perhaps with (D.5) which says
Σk=1N [(k)fi(r1, r2..... rk, t) ] k(t) + [∂tfi(r1, r2..... rk, t)] = 0 i = 1,2...s (D.5)
We could then define
Hi = Ait i = 1,2...m
Hi = ∂tfi i = m+1,2...C
Then we have this more uniform set of set of equations to deal with,
mkk(t) = Fk(r1, r2..... rN, t) + Σi=1C λi(t) Bik(r1, r2..... rN, t) k = 1....N
Σk=1N Bik(r1, r2..... rN, t) k(t) + Hi(r1, r2..... rN, t) = 0 . i = s+1 ...C (**)
Remember that we know the Fk and the Bij and the Hi functions of all their variables. Above we still have a set of 3N+C equations in 3N+C unknowns rk(t) and λi(t) . Note that
ik(r1, r2..... rN, t) = Σm=1N((m) Bik) m(t) + (∂tBik)
i(r1, r2..... rN, t) = Σm=1N((m) Hi) m(t) + (∂tHi)
Suppose we apply d/dt to the second equation set (**) above. We would get
Σk=1N Bik(r1, r2..... rk, t) k(t) + Σk=1N ik(r1, r2..... rk, t) k(t) + i(r1, r2..... rk, t) = 0
and then making the replacements,
Σk=1N Bik(r1, r2..... rk, t) k(t) + Σk=1N { Σm=1N((m) Bik) m(t) + (∂tBik) } k(t)
+Σm=1N((m) Hi) m(t) + (∂tHi) = 0
But now we can use k(t) = [Fk + Σi=1C λi(t) Bik]/mk to rewrite the above as
Σk=1N Bik [Fk + Σi=1C λi(t) Bik]/mk
+ Σk=1N { Σm=1N((m) Bik) m(t) + (∂tBik) } k(t)
+ Σm=1N((m) Hi) m(t) + (∂tHi) = 0
This equation has the following form
Σkmj Kikmj (vm)j(vk)j + [Σkmj Lik + ΣmjMmi] (vk)j = - Σk=1N Bik [Fk + Σi=1C λi(t) Bik]/mk
Rewrite showing all r's,
Σkmj Kikmj(ri..)( (m)j(k)j + [Σkmj Lik(ri..) + ΣmjMmi(ri..)] (k)j = - Σk=1N Bik(ri..) [Fk(ri..) + Σi=1C λi(t) Bik(ri..)]/mk
This is a very non-linear equation. The ri could appear in any functional form inside the various factors, and you can see that the i have quadratic times functions of ri , I would call this "hopeless".
****************
That paper comments (at great length) on the often vague term "virtual displacement" as it appears in Goldstein and other standard mechanics texts, and we have followed their suggestion for defining this term as a difference of "allowed displacements" which then removes the dt time term in (D.1).
**********************
In this last case, we can assume solutions rk(t) and then we have in effect λi(r1(t), r2(t)..... rN(t), t) = λ'i(t), so we can rule out the third situation. So in my opinion the λi are supposed to be λi(t).
2. For each of these possibilities, how do we know that a solution to the above equations exists?
Σk=1N [(k)fi(r1, r2..... rk, t) ] k(t) + [∂tfi(r1, r2..... rk, t)] = 0 derivative of the above
(D.5)
Σk=1N Aik(r1, r2..... rN, t) k(t) + Ait(r1, r2..... rN, t) = 0 . i = s+1 ...C
In the Comment 4 at the end of Section 2 we note that sometimes the constraints on a system with "generalized coordinates" qk can be written Σkaikdqk + aitdt = 0. Such constraints written in terms of differentials of the coordinates are examples of non-holonomic constraints. This situation arises typically when constraints act on velocities rather than positions. An example is a vertical disk rolling on a horizontal plane as shown in Goldstein p 13. Such constraints are difficult to handle because the coordinates qk are not generally independent, meaning that the values of all the qk cannot be set independently. There is no constraint equation of the form a(q1,q2.....) = 0 which is a function of the whole coordinates qi; differential constraints are "non-whole".
Inequality constraints are also non-holonomic, such as for particles bounded inside a box, but we shall use the non-holonomic term only in the sense just discussed.
As shown in our Comment 4, there is a role to be played by "Lagrange Multipliers". There is one such multiplier λi for each non-holonomic constraint and the values of the λi are determined as part of the problem solution.
Our main focus in this document is the use of Lagrange Multipliers to find candidates ri for which a scalar function f(r) is stationary, df = 0, subject to constraints. This is application of Lagrange multipliers is distinct from that outlined in Comment 4 where there is no obvious function f(r) which we are trying make stationary.
However, by extending the notation of stationarity one can first define a certain function of the possible solution trajectories qk(t) for a problem. That function is called "the action" S which is the time integral of the Lagrangian L(qk,k, t). The Lagrange equations of motion (including constraints) can be generated by setting δS = 0, a fact known as Hamilton's Principle (Goldstein Chapter 2 or the web). This notion is also used in the path integral formulations of quantum mechanics and quantum field theory. But these action applications are really not the same as the relatively simple problem addressed in this document concerning finding extrema of some scalar function f(r) .
A constraint of the form a(q1,q2....qN, t) = 0 is referred to as a holonomic constraint, and the main text of our document deals only with such constraints.
We shall now outline another mechanics application of Lagrange multipliers, again unrelated to making some obvious function f(r) be stationary. The presentation below follows Section IV of a 2006 paper by Ray and Shamanna (R&S) and tries to stick to their notation. Usually one wants to remove constraint forces from an analysis, but here they are instead the focus of the approach. In any event, this is just another example of the use of Lagrange Multipliers.
Note: In this appendix subscripts are just labels, they do not indicate differentiation as we used elsewhere in this document. For example, fi does not mean ∂f/∂xi, it just refers to the ith constraint function. Similarly rk is the 3D coordinate of the kth particle.
An application of Lagrange Multipliers
Consider a system of N particles each having a 3D coordinate rk, so there are 3N scalar variables. Assume there is a set of s holonomic constraints fi(r1, r2,,,,rM, t) = 0 for i = 1,2...s. Assume in addition that there is a set of m non-holonomic constraints Σk=1MAik(t) drk + Ait(t)dt for i = 1,2...m .
These C constraint surfaces each of dimension 3M-1 in E3M intersect in some surface of dimension 3M-C in E3M which then restricts the positions of the particles, as outlined in Appendix C: only certain {r1, r2...rM} are allowed in any solution of the problem.
For example, suppose there are M=3 particles constrained to lie at the vertices of an equilateral triangle of edge 3, and in addition particle 1 is only allowed to be on a sphere of radius 5. The four constraints then are | r1- r2 | - 3 = 0, | r2- r3 | - 3 = 0, | r3- r1 | - 3 = 0, and (r1)12+ (r1)22+(r1)32- 52 = 0. Each such constraint fits the holonomic template form fi(r1, r2, r3, t) = 0. Each constraint is a surface of dimension 8 in E9 (there are many extrusions as discussed in Appendix C), and the intersection surface has dimension 5 in E9. One could count these 5 degrees of freedom as follows: location of particle 1 on the sphere (2) [perhaps θ.φ] and orientation of the triangle in space (3) [three Euler angles].
In Newtonian mechanics, the equation of motion of a particle may be written Fk + Gk = mkak where Gk is the total force of the constraints acting on particle k and Fk is the total of all other forces acting on this particle. Parameter mk is the particle's mass and ak is its vector acceleration.
A simple example of a holonomic constraint for one particle (particle 1) has that particle sliding down a frictionless curvy water-slide type surface (but no water) in the presence of gravity. In that case, F1 is the force of gravity acting on the particle, while G1 is the force of the slide surface acting on the particle. This force is normal to the instantaneous path of the particle so then G1 dr1 = 0, which can also be written in terms of the particle's velocity G1 v1 = 0 .
Now differentiate the C constraint equations fi(r1, r2,,,,rM, t) = 0 to get
dfi = 0 = Σk=1M (k)fi drk + (∂tfi) dt i = 1,2...C (D.1)
where (k) ≡ ( ∂/∂[(rk)1], ∂/∂[(rk)2], ∂/∂[(rk)3)]) is the 3D gradient with respect to the vector coordinate rk. Let {drk} for k = 1,,M be a set of differential displacements which satisfy equations (D.1), and call this a set of allowed displacements. Let {dr'k} be some other set of allowed displacements so that
dfi = 0 = Σk=1M (k)fi dr'k + (∂tfi) dt i = 1,2...C . (D.2)
Now define δrk ≡ drk - dr'k to obtain a set {δrk} of so-called virtual displacements. Such a set is the difference of any two different sets of allowed displacements. Then subtracting the above two equations gives
0 = Σk=1M (k)fi δrk i = 1,2...C . (D.3)
This equation is true only as a sum, the individual terms in general do not vanish. Now multiply each of the C sums in (D.3) by a separate (and arbitrary) Lagrange Multiplier λi and add the sums together to get
0 = Σi=1C λi [Σk=1M (k)fi δrk]
or
0 = Σk=1M [Σi=1C λi (k)fi] δrk . (D.4)
This method of defining the meaning of a virtual displacement set follows the prescription (5) of R&S.
Assumption that constraints do no virtual work
Next, assume that
0 = Σk=1M Gk δrk . (D.5)
In Goldstein this equation corresponds to setting the second term in (1-40) to zero. The main justification for making this assumption is that it is valid for particles which form a rigid body and for many other constraint situations. There seems to be no general justification of this equation for all constraint situations, and it is certainly not valid if friction is present. The equation can be interpreted as saying that the "total virtual work" done by the constraints on all the particles of the system vanishes.
As a first simple example, consider two particles 1 and 2 independently sliding down a frictionless static ramp. In this case for particle 1 the constraint force G1 is the normal force of the ramp acting on the particle, and this is clearly perpendicular to dr1 for any allowed displacement set {drk}, and thus G1 is also perpendicular to any δr1 which is part of any virtual displacement set {δrk}. In this case we have G1 δr1 = 0 and G2 δr2 = 0 so the terms in (D.5) are separately zero.
A second example is more interesting. Particles 1 and 2 of arbitrary masses slide down a one dimensional ramp but are glued to the two ends of a massless stick so they comprise a rigid body. The upper mass is attached to a string which provides a constant tension T. Here is a picture,
(D.6)
Particle 2 exerts a constraint force T12 on particle 1 through the stick, and Particle 1 exerts a constraint force T21 on particle 2. Since there is only one force of tension (or compression) in such a stick, T12 = T21. All allowed displacement sets {dr1, dr2} are of the form {dr, dr}, and therefore all virtual displacement sets are of the form {δr1, δr2} = {δr, δr}. In (D.6) we have drawn the two normal forces Ni and the two total constraint forces Gi. It is no longer true that G1 δr1 = 0 and G2 δr2 = 0 as the picture shows. However, it is true that Σk=12 Gk δrk = [ Σk=12 Gk ] δr = 0 because G1 and G2 have equal and opposite components along the ramp, so equation (D.5) is valid for this example.
Now suppose the ramp is translating in some direction not along the ramp. The displacements dr1= dr2 are then no longer along the ramp so that Σk=12 Gk drk ≠ 0. However, any virtual displacement set would still have δr1 = δr2 along the ramp and then Σk=12 Gk δrk = 0. Remember that a virtual displacement set is the difference of any two allowed displacement sets, and in this subtraction the effect of the ramp translation cancels out. So equation (D.5) is still valid.
Other examples are given in R&S Section III.
Reader Exercise: show in the last example that T12 = T m2 / (m1+m2) and that this makes sense if m1 >> m2 and vice versa.
We proceed then assuming our system of interest respects equation (D.5).
Resume Lagrange Mulitpliers application
Now subtract equation (D.4) from (D.5) to get,
0 = Σk=1M [ Gk - Σi=1C λi (k)fi] δrk . (D.7)
We know that the M vectors in the set {δrk} are not linearly independent because they must satisfy the equations (D.3), so we cannot claim from (D.7) that [ Gk - Σi=1C λi (k)fi] = 0 for each k. But now write out equation (D.7) in full detail, showing all vector components,
0 = Σk=1M Σj=13 [ (Gk)j - Σi=1C λi ] (δrk)j . (D.8)
Next, relabel the 3M terms in this sum using a single index n defined in this way
(k,j) = (1,1), (1,2), (1,3), (2,1), (2,2), (2,3),..........(M,1), (M,2), (M,3)
n = 1, 2, 3, 4, 5, 6 , 3M-2, 3M-1, 3M (D.9)
so that n = 3(k-1)+j. Then define
Gn ≡ (Gk)j xn ≡ (rk)j . // n = 3(k-1)+j (D.10)
Equation (D.8) can then be expressed as
0 = Σn=13M [ Gn - Σi=1C λi ] δxn (D.11)
and (D.3) becomes
0 = Σn=13M δxn i = 1,2...C . (D.12)
Once again, we cannot set the square brackets in (D.11) to 0 because the 3M coordinates δxn are not linearly independent as shown by (D.12).
At any given time t, there must exist at least one subset of the 3M-C variables δxn which are independent and the remaining δxn are dependent. Assume that the dependent variables have indices n = n1, n2,...nC. Consider then the following set of C equations involving the dependent-term square brackets in (D.11),
[ Gn - Σi=1C () λi] = 0 n = n1, n2,...nC . (D.13)
This represents C linear equations in the C parameters λi so when all is said and done, there must exist a set of values {λi} which make these C equations be true. At this point, however, we don't know the Gn and we don't know the xi which appear in so we cannot yet compute values for the λi. When the problem is fully solved, the Gn(t), xn(t) and λi(t) will be known, where here we momentarily exhibit the fact that all these objects can be functions of time.
Now the terms in (D.11) indicated by n = n1, n2,...nC all vanish due to (D.12), so (D.11) can be written,
0 = Σn≠n,n,....n [ Gn - Σi=1C λi ] δxn . (D.14)
But all the δxn appearing in (D.14) are independent, so the square brackets here must also vanish. We then end up
[ Gn - Σi=1C λi ] = 0 n = 1,2....3M
or
Gn = Σi=1C λi n = 1,2....3M
or
(Gk)j = Σi=1C λi
or
Gk = Σi=1C λi ((k)fi) k = 1,2...M . (D.15)
Meanwhile, Newton's Law says that mkak = Fk + Gk where the Fk are the "applied" forces for the problem. Inserting (D.15) gives
mkk = Fk + Σi=1C λi [(k)fi(r1, ] (D.16)
which is 3M differential equations in 3M+C unknowns rk and λi. But we also have the C constraint equations fi(r1, r2, ...rM, t) = 0, so in total there are 3M+C equations in 3M+C unknowns, which then defines a solvable problem. Once the problem is solved, the λi are known and the constraint forces are given by (D.15).
To summarize, we have used an application of Lagrange Multipliers as an aid to solving a mechanics problem involving M particles with C holonomic constraints. As already noted, the above presentation is based on Section IV of R&S. That paper comments (at great length) on the often vague term "virtual displacement" as it appears in Goldstein and other standard mechanics texts, and we have followed their suggestion for defining this term as a difference of "allowed displacements" which then removes the dt time term in (D.1).
In contrast with our main topic of finding stationary points df = 0 of a function f(r), in the Lagrange Multiplier application just presented there is no explicit function f(r) we are making stationary. On the other hand, one can regard (D.5) that ΣkGk δrk = 0 as the statement δW = 0 where δW = ΣkGk δrk is the differential "virtual work" done by the constraint forces during a differential virtual displacement. This equation can then be written 0 = δW = Σk[mkak - Fk] δrk in which form it is known as D'Alembert's Principle (1743, Goldstein (1-42)). The analog of f(r) is then presumably some action-like integral W of the virtual work δW.