Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Math / Lagrange Multipliers / All Support

gradient section save

DOCX · 222.3 KB
Open DOCX file

Working draft saved by Phil on 1.11.15 as a section of a Lagrange multipliers write-up. It works Example 1, maximizing a hemisphere height f on a radius-2 sphere under the constraint y=1, and shows that the gradients of f and the constraint are collinear at the maximum. It checks the maximum with differential moves and second derivatives of the Lagrangian H, then argues generally that gradients are normal to level surfaces. Some text is unfinished and some symbols are lost in extraction.

AI-written summary; may contain errors. This description is approximate.

Extracted text (machine-read; may contain errors)
Save partially edited Section 4 on gradient stuff. PhL 1.11.15 Note that page numbering is turned on in this template and view is 125%, located in phil/roaming/microsoft/templates size about 219K. 4. The Gradient Interpretation: Example 1 Recall that for r such that rank[R(r)] < S we can find Lagrange multipliers λi such that f(r) + λ1a(r) + λ2 b(r) + ...... + λS-1 q(r) = 0 (2.5) where the bolded letters are the rows of the R matrix shown in (1.2). In component notation this equation reads fi(r) + λ1ai(r) + λ2 bi(r) + ...... + λS-1 qi(r) = 0 i = 1,2...N (4.1) where recall fi = ∂f/∂xi. With the usual gradient operator this can be written f(r) + λ1a(r) + λ2 b(r) + ...... + λS-1 q(r) = 0 = (∂1,∂2...∂N) . (4.2) How might one interpret this gradient sum being 0? If we use N = 2 and S = 2, we can at least obtain some visualization, then we can generalize the conclusion in a logical fashion. A Simple Example with One Constraint Let u = f(x,y) represent the surface of a sphere (radius R = 2) in E3, centered at the origin. We consider only the upper half of this surface, and we want to find (x,y) that maximizes f. If there are no constraints, then (4.2) above says f = 0. The only place on our u = f surface having f = 0 is the north pole of the sphere. This is just a regular "critical point" where ∂xf = 0 and ∂yf = 0, so we are happy with this interpretation of (4.2) with no constraints. We now add a constraint a(x,y) = 0 where a(x,y) = y-1. This constraint is the line y=1 in the x-y plane E2. We can extrude this line into a plane y = 1 in E3. The hemispherical surface is a 2D surface in E3, and the extruded constraint is also a 2D surface in E3. These 2D surfaces intersect in a 1D surface which is a curve. There is hopefully some point on this curve that maximizes f. Below is a picture of the sphere. The intersection of the upper spherical surface with the plane y = 1 is shown as a red curve (a half circle). Due to the constraint, we cannot get to the north pole so the maximum value of f is some value less than that u = 2 at the north pole. The extremum of this constrained problem will be at point A, and point B is not an extremum. (4.3) Since N = 2, the gradients in (4.2) are 2D gradients and so are parallel to the x,y plane. To view these gradients, we draw a top view of the sphere, looking toward the sphere center from the positive u axis: (4.4) The three black arrows show f at the points A, B and B'. It seems clear that on the spherical surface f always points toward the north-south axis of the sphere. We can put the gradient arrow tails at the points of interest, but the gradient arrows are all parallel to the x,y plane since for example f = (∂xf, ∂yf). Meanwhile, the three red arrows show the direction of a at points A, B and B'. Since a = y-1, this direction is of course a = (∂xa ∂ya) = (0,1) which points to the right no matter where on the extruded constraint surface the arrow tail is placed. As point B or B' moves toward point A, the black and red arrows become more aligned and finally at point A (the extremum) they are exactly collinear. Equation (4.2) says that at the solution point r = A for the extremum, we have f(r=A) + λ1a(r=A) = 0 (4.5) and indeed, this equation says that at the solution point the two gradients must be collinear. More details of the Example 1. f(x,y) = since the sphere has u2+x2+y2 = R2 = 4 f(x,y) = - (x/f)- (y/f) f(r,θ) = polar coordinates f(r,θ) = [-r/] points toward the sphere's vertical axis (and f = 0 at r = 0, north pole) a(x,y) = y-1 = 0 the constraint a = [1] points toward the right At point A = (x,y) = (0,1) one has r = 1, = and f = . Then (4.5) reads [-1/] + λ1[ 1/] = 0 λ1 = 1/. (4.6) At the solution point A the Lagrange multiplier λ1 can thus be interpreted as the negative of the ratio of the two gradient vectors f and a. Note that at the solution point, the two gradients really are collinear. At point B, we can consider a differential movement dr = |dx|(-) along the constraint surface toward point A (in Fig (4.4) the x axis points down). We find that df(B) = f dr = [ - (x/f)- (y/f)] |dx|(-) = |dx| (x/f) > 0 since x>0 and f>0 . (4.7) Thus, as we move toward point A, f increases since df>0, suggesting that A is indeed a maximum. Consider the mirror point B' located above point A in Fig (4.4). The differential vector dr = |dx| lies along the constraint surface and points from B' toward A. Then df(B') = f dr = [ - (x/f)- (y/f)] |dx| = |dx| (-x/f) > 0 since -x>0 and f>0 . (4.8) Once again, df > 0 showing indeed that point A is a maximum. Finally, we have claimed that at the solution point, the Lagrangian function H(x,y) should have a maximum, since the whole theory is based on H being the function to maximize without constraints. One has, H(x,y) = f(r) + λ1a(r) = + (1/) (y-1). (4.9) Now u = (1/) (y-1) is a plane sloping up to the right in Fig (4.3). This is different from the plane y = 1 which is a vertical plane in that figure. In (4.9) we are adding a spherical surface to a plane sloping up to the right, and the result is an ellipsoidal-like surface (really a quartic surface) which has a maximum at the point A = (0,1). Here is a plot of that surface: (4.10) A direct method of confirming the maximum is provided by examining Hii : Hi = fi + λ1ai Hii = fii + λ1aii aii = ∂2(y-1)/∂xi2 = 0 f = = fi = - xi/f f > 0 fii = - [f * 1 - xifi]/f2 = - [f - xi(-xi/f)]/f2 = - (1/f) - (xi/f)2 Hii = fii + λ1aii = fii = - (1/f) - (xi/f)2 < 0 . This shows that the function H(x,y) is "cupping down" at all locations, as the graph suggests. The point A where Hi = 0 (see (4.6) ) thus also has Hii < 0 and is thus a maximum. The general case Consider first a 2D surface in 3D space defined by F(x,y,z) = C. If one is on this surface and moves around on the surface, one always has F = C. Fact: At any point r, F(r) is locally normal to the surface. The reason is very simple. Suppose an ant starts at point r on the surface and move dr in some arbitrary direction on the surface. Saying the ant stays on the surface says that dF = 0 for this motion. But dF = dr F. Since dF= 0, and since this is true for any on-surface direction dr, F must be normal to the surface at that point r. Now consider F(x1,x2.....xn) = C. This is an n-1 dimensional surface in En. For any motion dr along this surface, one finds the same thing as above, dF = dr F = 0 so F is normal to the surface in En. In the extremum problem with constraints in n dimensions, for any constant C the equation f(x1,x2.....xn) = C defines an n-1 dimensional surface. This surface is called a level surface of f. If one takes a set values for C, one gets a set of level surfaces. They might look like a set of concentric spherical shells in E3, but in general the shape changes slightly as C makes a small change in value. In Example 1 above, the half sphere shown in Fig (4.3) is a level surface for x2+y2+ For a surface of dimension n-1 in En that same thing is true. Given a surface F(x,y,z) = C in 3D space, at any point on that surfaceF is a normal vector, so if an ant crawls some small distance dr away from the point that surface, dr F = 0. The reason this is true is very simple. A motion of dr along the surface results in a change in F of dF = 0 and since dF = dr F, and since dr could be any direction along the surface. it must be that F must be a normal vector (in general it is not a unit normal vector). If we were to upgrade this example to u = f(x,y,z) and have two constraints a=0 and b=0, the final intersection path of the hypersphere surface with the two constraint surfaces is again a 1D curve. If we consider dr along this curve, that dr will be perpendicular to both a and b, which says (λ1a + λ2b) dr = 0 . (4.11) At an extremum point, if (4.2) is valid, we find that f dr = - (λ1a + λ2b) dr = 0 . (4.12) Since this dr movement is perpendicular to f, there can be no change in f by moving along the intersecting constraint surface, and that is why we are at an extremum. So this then provides an interpretation for the general solution case where this equation is true: f + λ1a + ... λS-1 q = 0 (4.2) and thus, f dr = - (λ1a + λ2b + .... + λS-1q) dr (4.13) In general the intersection surface of f with all the constraint surfaces will be of dimension N-S+1 in En. A tiny displacement dr along this surface is perpendicular to all the local constraint surface gradients, so the right side of (4.8) is 0. At a solution point where equation (4.2) is true, we then have f dr = 0 so a small displacement dr in any constraint-legal direction results in df = 0, hence we are at an extremum.