Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Math / Lagrange Multipliers / All Support

the gradient interpretation

DOCX · 58.6 KB
Open DOCX file

A section from Phil's notes on Lagrange multipliers, titled "The Gradient Interpretation". It works an example with N=2, S=2: the upper half of a sphere of radius 2 constrained by y=1, showing the gradients of f and the constraint are collinear at the extremum and that the multiplier equals 1 there. It then generalizes to several constraints, showing df=0 along constraint-legal displacements dr.

AI-written summary; may contain errors.

Extracted text (machine-read; may contain errors)
Lag Mult Gradient Section 4. The Gradient Interpretation Recall that when rank[R(r)] < S we can find Lagrange multipliers λi such that f(r) + λ1a(r) + λ2 b(r) + ...... + λS-1 q(r) = 0 (2.5) where the bolded letters are the rows of the R matrix shown in (1.2). In derivative notation this equation reads fi(r) + λ1ai(r) + λ2 bi(r) + ...... + λS-1 qi(r) = 0 i = 1,2...N (4.1) where recall fi = ∂f/∂xi. With the usual gradient operator this can be written f(r) + λ1a(r) + λ2 b(r) + ...... + λS-1 q(r) = 0 = (∂1,∂2...∂N) (4.2) How might one interpret this gradient sum being 0? If we use N = 2 and S = 2, we can at least obtain some visualization, then we can generalize the conclusion in a logical fashion. A Simple Example: Let u = f(x,y) represent the surface of a sphere (R = 2) in E3, centered at the origin. We consider only the upper half of this surface, and we want to find (x,y) that maximizes f. If there are no constraints, then (4.2) above says f = 0. The only place on our u = f surface having f = 0 is the north pole of the sphere. This is just a normal "critical point" where ∂xf = 0 and ∂yf = 0, so we are happy with this interpretation of (4.2) with no constraints. We now add a constraint a(x,y) = 0 where a(x,y) = y-1. This constraint is the line y=1 in the x-y plane E2. We can extrude this line into a plane y = 1 in E3. The hemispherical surface is a 2D surface in E3, and the extruded constraint is also a 2D surface in E3. These 2D surfaces intersect in a 1D surface which is a curve. There is hopefully some point on this curve that maximizes f. Below is a picture of the sphere. The intersection of the upper spherical surface with the plane y = 1 is shown as a red curve (a half circle). Due to the constraint, we cannot get to the north pole so the maximum value of f is some value less than that u = 2 at the north pole. The extremum of this constrained problem will be at point A, and point B is not an extremum. (4.3) Since N = 2, the gradients in (4.2) are 2D gradients and so are parallel to the x,y plane. So view these gradients, we draw a top view of the sphere, looking toward the sphere center from the positive u axis: (4.4) The two black arrows show f at the points A and B. It seems clear that on the spherical surface f always points toward the north-south axis of the sphere. We can put the gradient arrow tails at the points of interest, but the gradient arrows are all parallel to the x,y plane since for example f = (∂xf, ∂yf). Meanwhile, the two red arrows show the direction of a at points A and B. Since a = y-1, this direction is of course f = (∂xa ∂ya) = (0,1) and this points to the right no matter where on the extruded constraint surface the arrow tail is placed. As point B moves toward point A, the black and red arrows become more collinear and finally at point A (the extremum) the are exactly collinear. Equation (4.2) says that at the solution point r = A for the extremum, we have f(r=A) + λ1a(r=A) = 0 (4.5) and indeed, this equation says that at the solution point the two gradients must be collinear. In this example the equation of the sphere surface is u2+x2+y2 = R2 so u = f(x,y) = . which in polar coordinates is f(r,θ) = . The 2D gradient is then f = [-r/] and thus points toward the sphere's vertical axis as claimed. Also as claimed, at the north pole f = 0. At point A one finds f = [-1/] = [-1]. Since a = [1] at point A, we write ** as [-1] + λ1[1] = 0 λ1 = 1 (4.6) At the solution point the Lagrange multiplier λ1 can thus be interpreted as the negative of the ratio of the two gradient vectors f and a. Now look back at Fig (4.4). At point B, since f has some component along the constraint path, we know that f increases if we move along that path toward point A. But at point A, f is normal to the constraint path and thus has no component along the path, so there is then no advantage to going away from point A in terms of increasing f. So at point A, we don't have the f = 0 (a normal critical point without constraints), but we do have f dr = 0 where dr is a differential motion along the red intersection curve. If we were to upgrade this example to u = f(x,y,z) and have two constraints a=0 and b=0, the final intersection path of the hypersphere surface with the two constraint surfaces is again a 1D curve. If we consider a dr along this curve, that dr will be perpendicular to both a and b, which says (λ1a + λ2b)dr = 0 (4.7) At an extremum point, if (4.2) is valid, we find that f dr = - (λ1a + λ2b)dr = 0 Since this dr movement is perpendicular to f, there can be no change in f by moving along the intersecting constraint surface, and that is why we are at an extremum. So this then provides an interpretation for the general solution case where this equation is true: f + λ1a + ... λS-1 q = 0 (4.2) and thus, f dr = - (λ1a + λ2b + .... + λS-1q) dr (4.8) In general the intersection surface of f and all the constraint surfaces will be a surface of dimension N-S+1 in En. A tiny displacement dr along this surface is perpendicular to all the local constraint surface gradients, so the right side of (4.8) is 0. At a solution point where equation (4.2) is true, we then have f dr = 0 so a small displacement dr in any constraint-legal direction results in df = 0, hence we are at an extremum.