Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Math / Lagrange Multipliers / Appendix C

App C

DOCX · 93.6 KB
Open DOCX file

Appendix C of Phil's Lagrange multipliers paper, supporting the Theorem 1 proof in Section 3. It works through three examples in E3 (two spheres, a sphere and a cylinder with a missing coordinate) and a six-variable case with three constraints, showing the eliminated variables may be multi-valued. A footnote argues that a smooth constraint a=0 defines an (N-1)-dimensional surface, using tangent hyperplanes. The file ends with rough draft fragments.

AI-written summary; may contain errors.

Extracted text (machine-read; may contain errors)
Appendix C: The Intersection of Constraint Surfaces In the proof of Theorem 1 in Section 3 (a) we rely on our ability, at least in theory, to eliminate x1, x2, x3 from the three constraint equations shown in (3.3). a(x1, x2, x3, x4, x5, x6) = 0 x1 = X1(x4, x5, x6) b(x1, x2, x3, x4, x5, x6) = 0 x2 = X2(x4, x5, x6) c(x1, x2, x3, x4, x5, x6) = 0 x3 = X3(x4, x5, x6) . (3.3) How does one know this is possible? The equations might be very complicated, some of the coordinates might not appear in some of the equations, and so on. We look first at three examples in E3 then in Example 4 we return to our question. Example 1: Consider these two constraint equations in E3 : a(x1,x2,x3) = x12 + x22 + x32 - 22 = 0 b(x1,x2,x3) = (x1-2)2 + x22 + x32 - 22 = 0 . (C.1) Each constraint represents a sphere of (radius 2) and so is a 2-dimensional smooth surface in E3. The first sphere is centered at the origin, while the second has its center at (2,0,0). Subtracting the first equation from the second gives -4x1+4 = 0 so x1= 1, then inserting this into both equations gives 1 + x22 + x32 - 22 = 0 1 + x22 + x32 - 22 = 0 x22 + x32 = 3 . (C.2) The intersection of the two constraint surfaces is a circle, a 1-dimensional smooth surface in E3. Given a(x1,x2,x3) = 0 and b(x1,x2,x3) = 0, is it possible to eliminate x1 and x2 from the two constraint equations? The answer is yes, as follows, x1 = X1(x3) = 1 x2 = X2(x3) = ± . (C.3) The following schematic drawing shows the intersection surface in red, (C.4) Notice in this example that there are two possible solutions for x2. One can regard the intersection surface (the circle x22 + x32 = 3) as having two pieces (front half, back half) and the ± sign selects one of these pieces. It is easy to imagine more complicated examples for two constraints in E3 where perhaps one or both of the constraint surfaces contain disjoint pieces (e.g. a hyperboloid), and where the intersection surface also has several disjoint pieces. Example 2: Consider these two constraint equations in E3 : a(x1,x2,x3) = x12 + x22 + x32 - 22 = 0 b(x1,x2,x3) = x22 + x32 - 12 = 0 . (C.5) The second equation is missing the x1 coordinate. The first equation describes the same origin-centered radius 2 sphere of the previous example. The second equation is that of a circle of radius 1, but as a function of three variables it is in fact a cylinder whose axis is the x1 axis. If x2 and x3 lie on the circle, any value of x1 satisfies the second equation. Missing coordinates result in surfaces which are "extruded" in the dimensions of the missing coordinates. Subtracting the second equation from the first gives x12 - 3 = 0 so x1 = ±, then inserting this into both equations gives 3 + x22 + x32 - 22 = 0 x22 + x32 - 12 = 0 x22 + x32 = 1 . (C.6) Given a(x1,x2,x3) = 0 and b(x1,x2,x3) = 0, is it possible to eliminate x1 and x2 from the two constraint equations? The answer is yes, as follows, x1 = X1(x3) = ± x2 = X2(x3) = ± . (C.7) The following schematic drawing shows the intersection surface in red, (C.8) In this example there are two possible solutions for x1 and two for x2 so overall there are four solutions. The intersection surface consists of the two circles each having a front half and a back half. Example 3: Consider these two constraint equations in E3 : a(x1,x2,x3) = x12 + x22 + x32 - 22 = 0 b(x1,x2,x3) = x12 + x22 - 12 = 0 . (C.9) Now x3 is missing from the second equation instead of x1. Subtracting the second equation from the first gives x32 - 3 = 0 so x3 = ±, then inserting this into both equations gives x12 + x22 + 3 - 22 = 0 x12 + x22 - 12 = 0 x12 + x22 = 1 . (C.10) Given a(x1,x2,x3) = 0 and b(x1,x2,x3) = 0, is it possible to eliminate x1 and x2 from the two constraint equations? The answer is yes, as follows, where α is an arbitrary real parameter in [-1,1], x1 = X1(x3 = ±) = α -1 ≤ α ≤ 1 x2 = X2(x3 = ±) = ± . // the two ± signs are independent (C.11) In the previous two examples the argument x3 was a free parameter whose variation mapped out the smooth constraint intersection surface piece(s). In Example 3 x3 is fixed at ± and a new parameter α must be introduced to map out the intersection surfaces. There are again four solution half-circles. The following schematic drawing shows the intersection surface in red, (C.12) Example 4: Now consider these three constraint equations in E6 : a(x1, x2, x3, x4, x5, x6) = 0 b(x1, x2, x3, x4, x5, x6) = 0 c(x1, x2, x3, x4, x5, x6) = 0 . (3.2) (C.13) We assume that each equation describes a smooth surface of dimension 5 in E6 (each is a hypersurface) and each surface may have several pieces. In a problem with constraints, one assumes that there is some non-null surface of intersection of all the constraint surfaces. This intersection surface may have multiple pieces and each piece (barring pathological cases) is itself a smooth surface of dimension 3 in E6. A candidate solution point r must lie on one of these intersection surface pieces. In general, each constraint knocks down the dimension of the intersection surface by one degree of freedom. So let us assume that x4, x5, x6 are the last three components of some point(s) on the overall intersection surface. Each such point of course has some x1, x2, x3 components. If x4, x5, x6 are varied slightly, x1, x2, x3 will also vary slightly, and this is what we mean by writing the functions x1 = X1(x4, x5, x6) x2 = X2(x4, x5, x6) x3 = X3(x4, x5, x6) . (3.3) (C.14) If the intersection surface has multiple pieces which have the same x4, x5, x6 value, then any of the three functions Xn above may be multi-valued, as occurred in our earlier examples. Conversely, if we select a triplet x4, x5, x6 for which there are no points on the intersection surface, the solution x1, x2, x3 does not exist. We don't care about such points in E6. So this then is what we mean in Section 3 (a) when we say above (3.3) that one can use the three constraint equations to eliminate the three variables x1, x2, x3. In the concluding equations (3.11) one can regard the functions X1, X2, X3 as being any of the multi-valued functions just discussed if the intersection surface has multiple pieces. For example, we had in (3.11) that = part of (3.11) (C.15) and here X14 refers to (∂/∂x4)X1(x4, x5, x6) evaluated at coordinates x4, x5, x6 for some candidate solution point r on the intersection surface. The fact that the matrix in (C.15) must have zero determinant is not affected by the possible existence of multiple solution X1, X2, X3 functions. Notice that the derivatives appearing in the matrix are unaffected by the possibility of multi-valued functions for X1, X2, X3. In the same manner as above, we can "eliminate" any triplet of variables xi, xj, xk ( i ≠ j ≠ k) from the three constraint equations (C.13). In the general case where there are S-1 constraints and r has N components x1,x2...xN with N > S, the dimensionality of the intersection surface is N - (S-1) in EN. The conclusions of the Theorem 1 proof outlined in Section 3 (b) are similarly not affected by the possible existence of an intersection surface having multiple pieces and the functions Xn possibly having multiple values. Footnote: Why is a(x1, x2, x3, ...xN) = 0 an (N-1)-dimensional surface in EN ? (C.16) This fact probably seems obvious to the reader and is certainly obvious in the first three examples above. Here we attempt a simple "engineering" explanation. We assume that a constraint equation is a "smooth" equation, meaning it is continuous and differentiable in all its arguments except perhaps at certain isolated locations which can be handled on an ad hoc basis either by fiat or by taking limits. Consider that, since a(r) = constant (namely 0), da = 0 so 0 = da = Σi=1N aidxi , // ai = ai(r) = ai(x1, x2, x3, ...xN) (C.17) where ai as usual means ∂a/∂xi. Let r be some point for which a(r) = 0. At this point we shall assume that there exists at least one as(r) which is non-zero. If all ai(r)= 0, we show below that r is a point on a surface which has a null normal vector, which is not possible, so such an r value could not lie on a(r) = 0. Then (C.17) may be written dxs = (1/as) Σi≠s aidxi . (C.18) Create a coordinate system whose origin is located at this point r for which a(r) = 0, and whose axes are aligned with those of EN. Now imagine an arbitrary tiny displacement of all the dxi other than dxs. For this set of N-1 dxi there is only one possible dxs which causes the point r + dr to satisfy the equation a(r+dr) = a(r) + da = 0 + 0 = 0, where dr ≡ (dx1, dx2.....dxN). That one possible dxs value is that given by (C.18). Imagine repeating this process for a continuum of values for the dxi other than dxs, and in each case we obtain the unique solution dxs from (C.18). We can think of the xs axis as being "vertical" and all the other axes being "horizontal" inasmuch as they are all perpendicular to the xs axis. The vectors dr generated in this manner comprise a tiny patch of a "plane" of dimension N-1 (a hyperplane) in EN which contains the point r. This is so because Σi=1N aidxi = a dr = 0 is the equation of a plane in EN passing through our origin at r with a being that plane's normal vector. More generally, r = d is the equation of a hyperplane in En having a unit normal and whose closest approach to the origin is distance d. Thus at a point r satisfying a(r) = 0 we have constructed a tiny neighborhood of nearby points r which also satisfy a(r) = 0. This neighborhood is a patch of a hyperplane in EN which is certainly a piece of "surface" in EN having dimension N-1. By repeating this process using a mesh of points ri satisfying a(ri) = 0, we then map out a triangulated surface of dimension N-1 in EN. In the limit the mesh size goes to 0, we arrive at a smooth surface of dimension N-1 in EN and that surface is a(r) = 0. The tiny planar patch at any point on the surface is part of the "tangent plane" to the surface at that point. The constraint a(r) = 0 is assumed to be "locally smooth" in the immediate region of any point r on the operational constraint surface, so a small local planar region (an open set) can be constructed around any such point. This is the basic idea of a surface being a manifold M, and the set of vectors in the tangent plane of a point r on M comprise the "tangent space" of M at r, usually denoted by TrM. ******************************************* What happens if x1 is missing from one or more equations? a(x1,x2,x3) = x12 + x22 + x32 - 22 = 0 b(x1,x2,x3) = x22 + x32 - 22 = 0 . First is a sphere, second is a circle? Well you can regard the first as a sphere dim = 2 in E3 and the second as a cylinder dim=2 in E3. Stop. These functions are evaluated at some and this is what we mean by saying we can in theory eliminate x1, x2, x3 , and the problem solution point r will lie on one of these pieces. More generally, if there are S-1 constraint equations, it is later implied that one can eliminate any subset of S-1 variables from the full set of N equations and this is how we demonstrate that the relevant submatrices of the R matrix all have zero determinant. We shall comment here only on the viability of eliminating x1, x2,x3 in the demonstration case of Section 3 (a), and the reader can easily generalize this viability to the general case. Reasonableness of the constraint surfaces. As was noted in ***, each constraint like a(x1, x2, ....xN) = 0 describes an N-1 dimensional surface in EN. This claim itself is perhaps not totally obvious, so we differentiate to get 0 = da = a1dx1 + a2dx2 + ... + aN dxN (C.1) where ai means ∂ia. If all N of the ai vanish, then a(x1, x2, ....xN) is a function of none of its variables and is some equation like a = π. We do not regard this as a meaningful "constraint", so there must be some as = ∂sa which does not vanish. One can then solve the above equation for dxs , dxs = ( a1dx1 + a2dx2 + ... + aN dxN ) / as . // asdxs missing from the sum (C.2) Let r = (x1,x2...xN) be some point in EN which satisfies the constraint equation. Suppose we select a set of small numbers for all the dxi except dxs. Then (C.2) above gives the value for dxs which sets da = 0. We have then created a displacement dr such that r + dr still satisfies the constraint equation a = 0. In this way we can construct a continuum of such dr vectors whose tails lie at r, and for all of which a(r+dr) = 0. These vectors paint out a tiny patch of the constraint surface near point r. Continuing this process, we can paint out the entire constraint surface in EN. The surface has N-1 degrees of freedom since that is how many dxi are on the right side of (C.2). When there are multiple constraints, the "legal operating region" for a solution point r consists of the intersection of all the constraint surfaces. Each new such constraint surface lowers by one the number of degrees of freedom (the dimensionality) of this surface of intersection. If there is no intersection, then the problem is not well-posed! So we have to make some assumptions about the set of constraint equations. First of all, they are all "smooth" equations (continuous, differentiable) for which no derivatives like ai are infinite. At any intersection point of the set of constraint surfaces, each surface can be modeled in the neighborhood of that point as a "plane" through the point having the same dimension as the surface (the tangent plane), allowing us to analyze things in terms of these linearized surfaces. In our region of interest, all the surfaces are formally manifolds being locally smooth at any point. Consider then the case of three constraint surfaces, a(x1, x2, x3, x4, x5, x6) = 0 b(x1, x2, x3, x4, x5, x6) = 0 c(x1, x2, x3, x4, x5, x6) = 0 (3.2) (C.3) each of which is a smooth surface of dimension 5 in E6. These surfaces intersect in some smooth surface of dimension 3 lying in the same space E6. It is possible that this intersection surface might have several disjoint pieces, in which case we assume our "operating region" is on one of these pieces. Suppose we select a point r on the intersection surface which has certain values x4, x5, x6. Differentiating gives 0 = da = a1dx1 + a2dx2 + a3dx3 + a4dx4 + a5dx5 + a6dx6 0 = db = b1dx1 + b2dx2 + b3dx3 + b4dx4 + b5dx5 + b6dx6 0 = dc = c1dx1 + c2dx2 + c3dx3 + c4dx4 + c5dx5 + c6dx6 (C.4) which rewrite as a1dx1 + a2dx2 + a3dx3 = - a4dx4 - a5dx5 - a6dx6 b1dx1 + b2dx2 + b3dx3 = - b4dx4 - b5dx5 - b6dx6 c1dx1 + c2dx2 + c3dx3 = - c4dx4 - c5dx5 - c6dx6 . (C.5) Consider a coordinate system whose origin lies at some point r on the intersection of the constraint surfaces and whose axes are parallel to the axes of the original coordinate system. In this system, we consider some small vectors of the form dr = (dx1, dx2, dx3, dx4, dx5, dx6). If we are given values for the components dx4, dx5, dx6, can we compute values for dx1, dx2, dx3 ? Using Cramer's Rule on (C.5) we the answer is yes as long as det ≠ 0 . If this determinant were to vanish, it would mean that the three row vectors are not linearly independent. Since these three vectors are ***************** Early in our proof of Theorem 1 in Section 3 (a) we rely on our ability, at least in theory, to eliminate x1, x2, x3 from the three constraint equations shown in (3.3). How does one know that this is even possible? Each of the three constraint equations in (3.3) represents a smooth surface of dimension 5 in E6. The intersection of the three surfaces represents a smooth surface of dimension 3 in E6.