Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Physics / Mechanics / Goldstein Related

sub_sham paper notes

DOCX · 52.5 KB
Open DOCX file

Phil's notes, dated 1.7.12 and updated 9.18.16, on a paper that criticizes the vague treatment of virtual displacements in Goldstein and other textbooks. They outline the paper's sections and work through its equations. The notes define allowed displacements, virtual displacements as differences of allowed ones, and ideal constraints with zero virtual work. They count unknowns against equations and cover moving constraint surfaces and pendulum examples.

AI-written summary; may contain errors.

Extracted text (machine-read; may contain errors)
Subhankar and Shamanna paper PhL 1.7.12 updated 9.18.16 This was a real find for me, a 2003 paper that calls Goldstein and others on their arm-waving BS vague discussions of "virtual displacements". I am precisely one of those "students" who does not buy into the discussion. I. Introduction. In their introduction, these two authors quote from several different books on how the virtual displacement is defined and how it is related to constraints. Not only are the quotes imprecise, some even conflict with each other! These guys quote from 11 rather famous sources listed in the references. I have only 2 of the 13 books mentioned. These guys are very much on my wavelength! The actually quote various vague statements from all the books, a bit of in your face. I could imagine writing this paper, it is a topic I might choose. They have this paper stored in the Arxiv, and at Researchgate, and at OpenLib which I was not aware of, this seems to involve peer review and typesetting so probably would not work for me. Paper has this structure (total 8 pages). I. Introduction opening text A. Ambiguity in virtual displacement (where all the quotes are listed) text then 1,2,3,4 then more text II. Virtual Displacements and Forces of Constraint A. Constraints and Virtual Displacement B. Existence of forces of constraint C. Solvability and ideal constraints III. Examples of Virtual Displacements A. Simple pendulum with stationary support B. Simple pendulum with moving support C. Motion along a fixed inclined plane D. Motion along a moving inclined plane IV Lagrange's Method of Undetermined Multipliers V. Generalized Coordinates and Lagrange's Equations of Motion VI. Conclusion, then Ack, then References with no heading In 2003 the two authors were Physics Profs at different Indian universities. II. Virtual Displacement and Constraint Forces A. They first assume the there are s holonomic constraints fi(rk, t) = 0, i = 1.2..s. (1). Chain rule then gives their (2). They write then drk = vkdt as "an allowed infinitesimal displacement" which is fine by me, and then (3) is obvious for such an allowed displacement. They propose to have a "virtual" displacement be any difference of such allowed ones, as long as you pick two different allowed ones. So they write δrk = drk- dr'k . At least for me this is consistent with motion along a plane and not into the plane which is the constraint. Now what happens when you install such a virtual displacement into the chain rule (4)? You get (7) where the dt term cancels because you took a difference! Now notice that when they talk about either allowed or virtual, drk or δrk , they are really talking about a vector of items, such as dr1, dr2.....drN where each element of the vector is a 3-vector. So perhaps I would write { drk }. Examples: Here are some PL examples. Consider first The set here is a null arrow for the left particle and a right arrow for the right particle. This {dr1, dr2} is neither an allowed nor a virtual system displacement because it violates the constraint. On the other hand, these three system displacements would be allowed: Each one corresponds to a rotation of the rigid object about some origin. Here is an example of an allowed system displacement for a 4-particle rigid object Notes added 9.18.16 : This Section A is completely clear, I have derived all equations. The key idea is to understand the notion of an "allowed set" of {drk} or {vk} : the "system of interest" could have these sets without violating any of the constraint equations. Then {δrk} is any set of differences of allowed sets, and this is then their definition of a set of "virtual" displacements: δrk ≡ drk - (drk)' . In exactly the same way, you can define vk ≡ vk - (vk)' where both {vk} and {v'k} are allowed velocity sets. It is quite clear from my examples above that certain {vk} would NOT be allowed (in fact most choices). Their purpose in using this definition is merely to cancel out the ∂tfi terms on the right side of (2) or (4) so you end up with the simple result (7) for the {δrk} set, and the same equation would also be true fore the {vk} set but they don't write this equation. B. They write F = ma for the kth particle. Differentiation ∂t on (2) gives (9). This last is regarded as a constraint on the possible acceleration sets {ak} for the system. Note that Fk refers to an "external force" and not a constraint force. So then (10) includes external forces Fi and constraint forces Ri, fine. They count then 3N equations of motion, but 6N unknowns if you include all the constraint forces. But the s constraints get you down to 6N unknowns and 3N+s equations. [Added ] The 6N comes from the rk(t) and the Rk(t), which is 3N + 3N functions of time in general. If you know these functions, you can then differentiate to get vk(t) and ak(t). But Newton only provides 3N equations as in (10) and you have s constraint equations as in (1), so you have 6N - (3N+s) = 3N - s degrees of freedom in the solution. So to have an exact solution, you need to find 3N - s extra scalar equations. Where will these come from? C. How do you solve the problem? They talk in terms of constraint surface (marble on table) and comment that both drk and δrk are tangential to the allowed motion surface, while the Rk are perp, and this then gives based on the fact of perpendicularity. They first talk about the constraint surface not moving. But what happens if the table is moving. Then certainly allowed motions are no longer tangential to the surface at some particular time t because the constraint surface has moved "up" say. But if you take their difference idea, the surface motion contribution cancels. So for moving constraints in time we get (11). Now suppose you take the right equation above (or in (11) and replace s of the dri by linear combinations of the other dri (since s are dependent on the others since we have our s holonomic constraints). Assume that these are the higher numbered items. Perhaps you then do some shuffling of the 3N-s variables into some other ones, The claim is that when you are done shuffling, you get something of the form (12) where there are now only n = 3N-s terms in the sum and the δj are now all independent. When you do the elimination and shuffle, you will end up with some j coefficients in (12). I don't think the authors want that dot in (12), we just have this set of n = 3N-s new coordinates {j} and they no longer form 3-vectors. How many equations does (12) represent? Well, the equations are really j = 0 for j = 1 to n, so these represent 3N-s constraints. So we now have 6N variables and 3N+s equations of motion and constraint equations, and now we have 3N-s new equations j = 0 so we then have 6N equations in 6N unknowns, and that is what we want. So the right equation above will be their "principle of zero virtual work". Problems which fit into this picture have "ideal constraints". Comments added 9.18.16: For certain constrained problems, like balls rolling on a curved surface, it is obvious that all allowed drk are tangential to the surface, but I would not assume this concept for all holonomic constraint problems! Problems of this type they say have "ideal constraints". In this case it is pretty obvious for a static surface that Σk=1N Rk drk = 0 and therefore also Σk=1N Rk δrk = 0. For such problems, the forces of constraint Rk therefore "do no work", and this is sort of a vague connection to the literature notion that there is "zero virtual work". Since the 3N numbers in {drk} or {δrk} are not linearly independent (due to the constraints), you cannot conclude that (Ri)k = 0, and of course we know that not all constraint forces are NOT zero for a general situation. The equation Σk=1N Rk δrk = 0 has 3N "component" terms on the left side. But in the set { δrk } there are only 3N-s independent components due to the s constraint equations (4). Rewrite the dependent equations perhaps as Σj=13N Rj (δrj) = 0 equation has 3N terms where the (δrj) are the original 3N components and the Rj are the force components, all now written out in a single linear list. Now suppose you were to first replace the LAST s of the (δrj) by linear combinations of the others using the equations (4). You would then have Σj=13N-s R'j (δrj)= 0 equation has 3N-s terms which include just the FIRST (δrj) There is a prime on R'j because some or all of the Rj got extra contributions when the linear combinations were inserted and terms collected. Now make a second "transformation" where instead of using the first (δrj), we use some arbitrary subset of the original 3N (δrj) as the independents. Call these {δξj} just to have a name, and then we get Σj=13N-s Aj (δξj) = 0 where now the Aj are whatever they come out being. Now do yet another transformation and define a new set of 3N-s independent variables (call them dηj) to be some linear combinations of the dξj. Then we get Σj=13N-s Bj (δηj) = 0 or Σj=13N-s j (δj) = 0 j = Bj δj = δηj where now the Bj are "whatever they come out being" when you insert the linear combinations and collect terms. This then is the meaning of equation (12). In passing we might refer to the δj as "generalized coordinates" and the j as "generalized forces", but let's hold on that for a while. Now since all the δj are linearly independent, we have these 3N-s new scalar equations: j = 0 j = 1,2...3N-s and these are the "missing" 3N-s equations that we need to solve the problem, mentioned at the end of B. Now suppose the surface on which the marbles are rolling is itself "moving". Specifically, assume that for some one of the point particles the surface has some non-zero velocity wk . I think that point on the surface could be accelerating, and the surface could be changing shape as well, but there is some wk no matter what is going on with the surface. Due to this wk we no longer have drk = tangent to surface. As a simple example, suppose the surface is just moving downward without acceleration and without changing shape. Then the marble has to move ON the surface, and drk will have a component tangential to the surface and another component wkdt due to the downward surface motion. The key point is this: if we were to start with some OTHER set {dr'k } of allowed displacements at the same time t and at the same locations rk on the surface at time t, the downward component of dr'k would be the same as that of drk . Then when you construct δrk ≡ drk - dr'k, the result is tangential to the surface because the vertical components cancel out in the subtraction. You are then back to the situation of the non-moving surface in terms of the δrk, and this is shown in the right equation of (11). The analysis then goes forward exactly as described above and you have a "solvable problem" for these "ideal constraints". We then maintain the idea of "zero virtual work" in the sense of (12) or of (11) right equation. So these authors start with a clean definition of a "virtual displacement (set)" and end up even for a moving surface with a "zero virtual work" concept. III. Examples (hurray!!) [ N = 1 ] A. Pendulum with fixed support. I agree with the picture and opening equations. I guess x is to the right and y is up in their picture. The virtual work rule is simply Tδr = 0 for this 1-particle system, where T is the tension which is perp to the motion and is an example of a constraint force. // Language error noted. This example is entirely kinematic only: what are the allowed velocities and the allowed displacements? The problem is not solved. B. Now allow the pendulum support to drop in time at speed u. Now we get my expected picture of the various displacements and the virtual ones, being differences, are still perp to the constraint tension. This will be true for vertical or horizontal motion of the support (or any combination). I went through all the equations on 9.18.16 there is nothing too exciting here. Yes, the displacement dr is not along the velocity v. Yes the virtual displacement δr is collinear with the velocity v. And yes we have that T δr = 0 so there is no "virtual work" done by the string on the mass. C and D. Fixed and then moving inclined plane. He does this in polar coordinates so then the plane constraint is a fix on angle as particle goes down the plane. Again they draw the various allowed and virtual displacements. They then let the plane move to the left at u and redraw the arrows. Whether or not the plane moves, the zero virtual work rule looks the same, in this case it is Nδr = 0 where N is the normal force. So this C,D example is essentially the same as the A,B example. In the AB case string tension is perp to virtual displacements whether the support is moving or not (in any direction, with any speed or acceleration etc etc). In the CD case the normal force is perp to the virtual displacements no matter now the plane is moving. In this example they don't talk about vertical motion of the plane. I think things work out as long as the vertical acceleration keeps the ball on the plane. The same comment applies to moving the plane to the right too fast. These issues are swept under the rug. You would have a similar problem with the pendulum if the support were to descend faster than the ball drops to keep the string taut . IV. Those Lagrange Multipliers. Although I have written an entire paper about the Method of Lagrange Multipliers, I am hard pressed to find a connection between that paper and the discussion here. The main issue is that I cannot figure out for what "function" one is trying to find a stationary point. In my paper that function was f but below I will refer to it instead as F since the Indian constraints are called fi. 1. So in my paper I have r = (x1, x2... xM) being a set of M non-independent coordinates. In the Indian paper we have (r1, r2 .... rN , t) being 3N+1 non-independent coordinates. But I think when you are talking about δ type differentials, you can ignore t, so let's now ignore t for the time being. So let's write r = (r1, r2 .... rN) = ( r1x, r1y, r1z, r2x, r2y, r2z , ..... rNx, rNy, rNz ) with M = 3N items = ( r11, r12, r13, r21, r22, r23 , ..... rN1, rN2, rN3 ) = ( x1, x2, x3, x4, x4, x6 ,....... x3N) M = 3N so that the fancy vector r consists of N 3-vectors so the full list has M = 3N components. I will refer to these components as xi and then we have a clear meaning for ∂i = ∂/∂xi for i = 1,2...3N. Now suppose you are given one of these M objects xn . You could say xn = rkj where k = 1 + Int(n/3) j = Rem(n/3) Then we can write ∂n = ∂/∂xn = ∂/∂xkj = "∂kj" again with k and j determined from n as above 2. In my paper we have a(r) = 0, b(r) = 0, ..... q(r) = 0 being S-1 constraints. In the Indian paper we have fi = 0 being the s constraints, so we are pretty close on this idea. 3. In my paper there is a function F we want to extremize so dF = 0, and this leads to ∂mF(r) + λ1∂ma(r) + λ2 ∂mb(r) + ...... + λS-1 ∂mq(r) = 0 m = 1,2...M . (4.1.2) In the Indian paper this would be written ∂nF(r) + λ1∂nf1(r) + λ2 ∂nf2(r) + ...... + λs ∂nfs(r) = 0 n = 1,2...3N . (4.1.2) or ∂kjF(r) + λ1∂kjf1(r) + λ2 ∂kjf2(r) + ...... + λs ∂kjfs(r) = 0 k = 1,2...N, j = 1..3 or ∂kjF(r) + Σi=1sλi∂kjfi(r) = 0 k = 1,2...N, j = 1..3 The second term here matches the second term in the Indian paper equation (17). We could think of this as a vector equation like so kF(r) + Σi=1sλikfi(r) = 0 k = 1,2...N where now k = (∂k1, ∂k2, ∂k3). 4. Now what does my paper have to say about "differentials"? I do of course have this, 0 = dF = F(r) dr = Σm=1M (∂mF) dxk 0 = da = a(r) dr = Σm=1M (∂ma) dxk 0 = db = b(r) dr = Σm=1M (∂mb) dxk ... 0 = dq = q(r) dr = Σm=1M (∂mq) dxk In the Indian paper notation this could be written 0 = dF = F(r) dr = Σn=13N (∂nF) dxn 0 = dfj = fj(r) dr = Σn=13N (∂nfj) dxn j = 1,2...s one equation for each constraint Now rewrite the above where has a new meaning. In the above lines is dot product of two huge vectors, but now we shall have it refer to a regular 3-vector dot product. Then we can write the above as 0 = dF = Σk=1N (kF) drk 0 = dfj = Σk=1N (kfj) drk j = 1,2...s one equation for each constraint 0 = δF = Σk=1N (kF) δrk (*) 0 = δfj = Σk=1N (kfj) δrk j = 1,2...s one equation for each constraint (**) where now we have defined, rk = (rk1, rk2, rk3) drk = (drk1, drk2, drk3) δrk = (δrk1, δrk2, δrk3) Equation (**) above is what the Indians mean by their equation (16) which they write in this more ambiguous notation 0 = δfj = Σk=1N (∂fj/∂rk) δrk 5. So what is F for the Indian paper? I want to identify my (*) above with their equation (13). So compare 0 = δF = Σk=1N Rk δrk // their (13) 0 = δF = Σk=1N (kF) δrk // my (*) To make this connection, we must identify kF = Rk . So we don't really know F, we only know its gradient for the kth coordinate. But maybe I could construct a possible F. Maybe we can then claim that F = Σk=1N∫Rk drk This tells me that the function F we are trying to make stationary is in fact the sum of all the work done by the constraint forces for a set of displacements drk . In a simpler case we would say F = R F(r) - F(r0) =!Syntax Error, IR dr = !Syntax Error, IF dr = line integral For a very short "paths" we could write F({rk0 + drk}) - F({rk0}) = Σk=1N Rk drk F({rk0 + d'rk}) - F({rk0}) = Σk=1N Rk d'rk Then subtracting we get F({rk0 + drk}) - F({rk0 + d'rk}) = Σk=1N Rk δrk or F( rk0; δrk) = Σk=1N Rk δrk = the total virtual work done by the constraint forces Rk = 0 Let's try this again for finite paths. We have F(rk) – F(rk0) = Σk=1N!Syntax Error, IRk drk (***) F(r'k) – F(rk0) = Σk=1N!Syntax Error, IRk drk  So, if I were to define the full function F by (***), that would lead to kF = Rk which in turn would lead to 0 = δF = Σk=1N (kF) δrk and this is the δF = 0 stationary situation of my Lagrange paper. Notice that in the examples with moving platforms, the work done is NOT zero, but probably still δF = 0. So I guess I am happy now to have made this Connection between my paper and the Indian paper, though there are still some loose ends. Idea. I am looking at the equation set (2.6) in my paper. Suppose I were to sum the first M equations (something I had no reason to do). I would get [F1 + F2 + .. + FM] = + λ1 [ a1+ a2 + ...+ aM] + λ2 [ b1+ b2 + ...+ bM] + ... + λs [ q1+ q2 + ...+ qM] or Σi=1M [ Fi – λ1 ai – λ2 bi – ... - λs qi ] = 0 // Σi=1M Hi = 0 The sum here is zero because each [...] bracket is zero because each of my equations is 0. Recall that you have to FIND a set of λj values which make the individual equations be true for some r. Suppose I rename things so that: a = f(1), b = f(2), .....q = f(s). Then I could write the above as Σi [ Fi – λ1 ∂if(1) – λ2 ∂if(2) – ... - λs ∂if(s)] = 0 or Σi=1M [ Fi – Σj=1sλj ∂if(j)] = 0 // Σi=1M Hi = 0 Another option would be to multiply each of my original equations by some arbitrary function Qi and THEN add them, and the result would be Σi [ Fi – λ1 ai – λ2 bi – ... - λs qi ]Qi = 0 or Σi=1M [ Fi – Σj=1sλj ∂if(j)]Qi = 0 (*) // Σi=1M HiQi = 0 The coordinates are enumerated by index i. I could regard the last s coordinates as being the dependent ones, and the first M-s coordinates as being the independent ones. Is it possible to arrange to have [Fi – Σj=1sλj ∂if(j)] = 0 for i = the last s values of i ending in M ? // Hi = 0 ? This would be s equations (linear in the λi) enumerated by index i, and there are of course s values λj. So yes, you could select the λj to be just those which cause all these last brackets to be zero. Call this solution set λ with elements λj. This is NOT something I do in my paper! But suppose we do this here. Then we get: Σi=1M-s [ Fi – Σj=1sλj ∂if(j)] Qi = 0 // Σi=1M-s Hi(λ) Qi = 0 because the last s brackets of (*) are all zero. Suppose now that the set of {Qi} objects is linearly independent, none can be written as a sum of the others. This means that the Qi can have arbitrary values. In this situation, the only way the above sum can vanish is if the brackets all vanish, that is to say, Fi – Σj=1sλj ∂if(j) = 0 i = 1,2....M-s Using this very particular set of λj constants, we end up with this fact: Fi = Σj=1sλj ∂if(j) i = 1,2....M-s This is very different from what I do in my paper. In that paper, F is a given function, and thus Fi are all predetermined. You cannot just set Fi as in the last equation above! You have some given F(r) function. In particular, you see above that the λj are determined entirely by just the last s values of Fi and ∂if(j) for just the last s partials of all the f(j). This λ set knows nothing about F1 and F2 for example. In my paper, F1 and F2 could be anything at all. So consider the equation resulting from above F1 = Σj=1sλj ∂1f(j) . This equation gives some F1, but that is not likely to be the same F1 that comes from your arbitrarily selected function F(r). So : The Indian application of Lagrange multipliers is not compatible with my paper's idea of having all the Fi being determined ahead of time since you have specified F(r). The Indian solution λ is not associated with any stationary point of the function F(r). Nevertheless, if you select the special set λ, the above WILL be true regardless of the F(r) you start with. Now OK, suppose I go back and replace Qi by dxi so I then have Σi=1M [ Fi – Σj=1sλj ∂if(j)]dxi = 0 where { dxi } = dr = some variation of the composite coordinate r which is allowed by the constraints. It is just some differential displacement, but due to the constraints, all M of the dxi are not linearly independent. But now if I select the special set {λj} just outlined above, I then get the last s brackets vanishing and am left with Σi=1M-s [ Fi – Σj=1sλj ∂if(j)] dxi = 0 Now all the dxi ARE independent, and I then conclude that Fi = Σj=1sλj ∂if(j) i = 1,2....M-s FOR THIS particular {λj} set. The set is computed exactly to make this M-s equations be true! I agree with (13) through (16). Notice that (16) has no time derivative. In SS's version of things we have a reason for this -- δrk is a difference of allowed displacements, this is just (7) again with definition of δfi. I agree with (17) as a sum of equations each of which is 0. You choose the λi so that each coefficient in (17) vanishes and this gives (19). Then the entire problem becomes 3N + s equations in 3N+s unknowns, where the λi are s of those unknowns. V. Generalized coordinates. I agree with (22). They then do Goldstein's derivation of the Euler-Lagrange equations. They just do this for the sake of "completeness". So I think this was an interesting paper. Their main interest I think was the time aspect, whereas mine was just the understanding of what is meant by a system displacement as {drk } of either allowed or virtual kinds.