Buck Chapter 6
DOCX · 603.9 KB
Open DOCX file
Typed reading notes by Phil, dated 1.11.15 with later additions, on Buck's Advanced Calculus Chapter 6 (pp 296-365). The visible part covers section 6.1: the volume rule vol L(D) = det(L) vol(D), his proof via the n-piped appendix of his Tensor Doc, and a Maple check of the area factor for E2 to E3. It then covers the change-of-variables theorems with the Jacobian. The outline also lists curves and arc length, surfaces and surface area, and extremal properties.
AI-written summary; may contain errors.
Extracted text (machine-read; may contain errors)
Buck Chapter 6 : Applications to Geometry and Analysis PhL 1.11.15
Chapter 6: Applications to Geometry and Analysis [ pp 296-365 ] 1
6.1 Transformations of Multiple Integrals [296] 1
6.2 Curves and arc length [313] 7
6.3 Surfaces and surface area [330] 14
6.4 Extremal Properties of functions of several variables [349] 23
Chapter 6: Applications to Geometry and Analysis [ pp 296-365 ]
6.1 Transformations of Multiple Integrals [296]
This section deals with volume and area transformations. The first portion of the section assumes that the transformation is linear. Instead of calling it T, they call it L.
Theorem 1: [296] vol [ L(D) ] = det(L) vol(D) k = det(L)
Bucks do NOT prove this theorem. Instead, they show that it is valid for a tetrahedron with one vertex at the origin and they call upon a rule which says vol(tetra) = 1/6 det(r1,r2, r3) = 1/6 det(U). Their proof is then that vol(tetra') = 1/6 det(r'1,r'2, r'3) = 1/6 det(U') and U' = LU since this is true for the columns, and therefore det(U') = det(L) det(U) and QED.
Before going on, how do I know this already? I don't know that I have a simple proof in my Matrix Notes binder. It is interesting to look at Tensor Doc Appendix B on finite n-pipeds and see what conclusions I reached there! This appendix may need some work, right now is an excuse to read it and edit it as needed.
Tensor Doc Appendix B.
I start with the 1-piped and e1 is an arbitrary vector that defines its edge. With the 2-piped and I have two such vectors, length not specified. So I think these vectors really are at this point (in this Appendix) arbitrary, they are just e = edge vectors. I probably should comment on that. But then I start talking about the E vectors in the full context of tensor doc. This requires then that the ei not be just any vectors, but that they be en ≡ Se'n so they are specific to our linear transformation S. I guess I could go backwards. Let the en be arbitrary vectors, then solve en ≡ Se'n for S. I show in fact that S = [ e1, e2, .....en ] so constructing S was an easy task. Then of course en ≡ Se'n . And Volume = | det(S) | so basically:
PL Theorem: If you have a finite n-piped with arbitrary edges e1, e2, .....en, construct S = [ e1, e2, .....en ] and then we know two things: (1) volume = |det(S)| and (2) en ≡ Se'n where the e'n are axis aligned unit vectors. In x'-space we have a unit cube with volume = 1, and in x-space we have our n-piped with volume |det(S)|. Here is my differential picture.
Now, suppose I were to create some scaled vectors in x'-space
v'i = ai e'i i = 1,2,3....n
The volume in x'-space is now V' = a1a2....an. What does this map into under S ?
Sv'i = S (aie'i) = ai Se'i // since S is linear in that S(αv) = αS(v)
= aiei
So our new n-piped edges have each been scaled by ai. Then
V = det [ a1e1, a2e2 ....] = a1a2... sn det [ e1, e2, .....en ] = V' det(S)
Thus I have just shown that for a finite rectangular solid, under the mapping S the volume changes by det(S) and this is true for finite size objects because the transformation is assumed linear. So in this sense my Appendix B provides a proof of Buck Theorem 1.
Be careful 5/16/15: If the transformation is linear, then the volume transformation rule applies to finite size objects like finite N-pipeds. For general transformations, rule only applies to differential volumes and the volume ratio |det(S)| is a function of position. For linear, we know S is NOT a function of position, so the differential volume ratio is the same at all locations and you can then integrate that to find that the ratio applies for any finite volumes under the linear mapping.
Does it apply to transforming other solid objects? To the extend that a solid object is a sum of lots of tiny cubes, I guess it applies to anything. Better to say that locally, any transformation is linear, and it then applies to any small n-piped which is local, and they you add up all the n-pipeds.
Is there some trivial proof of Theorem 1? Well, in my App B showing it for n = 2 and n = 3 is not very hard to do, the hard part is the case of general n. If there were a trivial proof, I think Bucks would have presented it. More precisely, Bucks have proven if for linear case with n = 2 or 3 (but have they?). They have not proved that this then applies to the non-linear case, though it certainly seems plausible.
My matrix notes claim you can write any S as S = QR where R is a rotation matrix and R is upper triangular. But how does that help me?
OK, I will just stick with my App B tensor doc theorem and be happy.
Bucks do show a special case: Start with a tetrahedron with one corner at the origin. Apply transformation and show the volume of this thing scales as det(S). Bucks only care about E3 right here, so I might have done the little cube instead. An arbitrary S can be written as [ e1, e2, e3] . Show that unit cube maps into this 3-piped and note that its volume is det S.
Note added: (1) Bucks note that area does not transform in such a simple way in E3 → E3, but that it DOES in this simple way for E2 → E2 since then area is the volume. (2) They also point out that the rule is accurate for det(L) = 0 where L has less than full rank, since when the f(D) is reduced in dimension, it has volume 0.
Area
Bucks don't attempt an area rule for general case, but they first do E2→E3 where you start with a little axb rectangle in E2. Assume this maps into 2-piped with edges aP1 and bP2 so this 2-piped just hangs in 3D space. They show that its area is ab times a scale factor k given as shown as p 299A or B. Then they just claim without proof that C is the k factor for E2→En . The result in terms of the aij is not very pleasant.
In tensor doc language, I could perhaps embed the above into an En → En situation and then I would use the fact that
dAn = dA'n so that k = I guess in Standard Notation
But g' = RRT so you get something like
k2 = cof (g'33) = cof[(STS)33]
Why 3? Bucks actually do E2→E3 so when I extend this to E3→E3 , the original area is in the z = 0 plane and for that the area label is 3.
Perhaps I could write this out and obtain the identity shown as p 299 D.
OK, here is some Maple work [ Buck Ch6 a.mws] . I will show for n = 3 that the cofactor idea gives the result p 299 B
First, I compute k2 "my way" as the cofactor thing:
and the last line gives the result. But then I directly enter B see if that is the same as det(h):
So amazingly my cofactor formula replicates Bucks' messy expression.
I don't have any fancy explanation of why D = E on page 299. The E is the sum of the squares of the 2x2 submatrices, a very strange object.
Note added 4.17.16. I have a feeling that you can also get the same expression for k in this way
k = where s is the tall s matrix shown on page 298.
Let's have the same Maple program check this conjecture. I don't have to go very far:
But this is just h from the previous Maple program so get same result.
Another Note added 4.17.16. Consider this from wedge doc
7 F*(αx') = Σ'I fI(F(x)) Σ' J det(RIJ) λ^J // F* pulling back a general k-form from Λ'k
= Σ' J gJ(x)λ^J where Σ'I fI(F(x)) det(RIJ) (10.8.2)
Suppose fI = 1. Then you would have
F*(αx') = Σ' J [Σ'I det(RIJ)] λ^J .
= Σ1≤j1≤ij<2 [ Σ1≤i1≤i2<3 det(RIJ) ] λj1 ^ λj2
= [ Σ1≤i1≤i2<3 det(RI12) ] λ1 ^ λ2
= [ det(R1212) + det(R1312) + det(R2312) ] λ1 ^ λ2
= [ det + det + det ] λ1 ^ λ2
But this is not what I am looking for! I want a sum of the squares of these same det's.
Plan A: recall from tensor doc 5.1.6.17 that
cof('nn) = g' g'nn = det(') g'nn // developmental notation
cof('33) == det(') g'33 = det(S)2 (RRT)33 = det(S)2 R3n2
' = STS g' = RRT
Perhaps write det(S) = S13 det13 - S23 det23 + S33 det33 . Then we get
cof('33) = [S13 det13 - S23 det23 + S33 det33]2 [ R312 + R322 + R332]
= [ S132 det132 + S232 det232 + S332 det332 + cross terms ] [ R312 + R322 + R332]
But this does not seem to give their answer which is just det132 + det232 + det332 in any obvious way. Let's look just at the non-cross terms:
= [ Σn Sn32 detn32] [ Σm R3m2] = ΣnmR3m2 Sn32 detn32
We do know that RS = 1 which says ΣnRinSnk = δik.
We also know that R2S2 = 1 which says R3mRmbSbcSck = δ3k.
This Plan A leads nowhere useful, All those cross terms to deal with.
Plan B. cof('33) = '11 '22 - '12 '21
= (STS)11(STS)22 - (STS)12(STS)21
= ΣnSn12 ΣmSm22 - ΣnSn1Sn2 ΣmSm1Sm2
But this is just the B form I already know. I give up again. The form E I verified is correct, but I have no fancy way to explain its simple appearance.
Bucks claim that you replace E with sum of all 2x2 dets in S and then they do a little example.
[300] So much for linear transformations, we are now back to general T. The contact is made that if you work with differentials, the volume factor will be k = det(S) and I agree. Close in, it is linear and that is Theorem 1. But Bucks are going to prove this in detail.
Theorem 2 [300]. Suppose in domain you have a compact set of zero volume, such as a line segment in 2D. The claim is that T(E) also has zero volume in the range. T has to be C1.
This kind of theorem is not of much interest to me, so I skip the proof for now. It is a limit proof.
Corollary [301]. A similar theorem. If T is C1, then if D "has a volume" in the domain, T(D) then "has a volume" in the range. D must be a compact set within some open set Ω and must have J ≠ 0 on Ω.
Footnote says these are "measure theory" theorems and are best pursued in a different framework. I guess "has a volume" means either has a non-zero volume, or has something that can be called volume.
Theorem 3 [301]. Consider T:E3→E3 and (x,y,z) = T(u,v,w), sort of backwards from the way we usually define variables, but correct for curvilinear coordinates! Let D be in the domain, and T(D) in the range and assume J ≠ 0. The volume of the range set T(D) is ∫T(D)dxdydz . The theorem claims that
volume of T(D) = ∫T(D)dxdydz = ∫D J dudvdw
which is of course the classic result. No integrand function yet, just doing volume.
The proof is a long 3 pages and includes a Lemma. It starts off by considering v(T(D)) as a "set function" over the set D -- a topic which was discussed at the end of Chapter 3. A theorem there says that if certain things are true, you can write a set function as an integral over D of a density-type function, and that is then the statement A on page 301. In this case, for v(T(D)) that density function will end up being J which is then the claim of Theorem 3. In order to prove this, they first prove it in the set-differential form:
Lemma [302] This lemma says that the set-derivative of v(T(D))/v(D) = | J |, which is of course the main idea that you want to show. This as usual is in a limit as set D approaches a point p.
[302] At this point they dive into a long and detailed proof of this Lemma. It involves our favorite remainder object R(Δp). Linearized L makes an appearance, and so does L-1 . I think the proof hinges on the idea that this result is true for the linearized L, and then the dust settles that makes it true for the general T. For me, it seems pretty clear that this should all work as long as functions are "reasonable", and that is the detail set that this proof examines. They wrap up saying (as expected) that this Lemma is basically Theorem 3, given the fundamental earlier set-function theorem proved in Ch 3. Fine.
Theorem 4 [304]. Same setup as Theorem 3. As before D is in the domain of (x,y,z) = T(u,v,w) and D* is the corresponding T(D) in the range, which is now assumed to be compact in x,y,z space. As before, J ≠ 0 on Ω from which D is taken. The new ingredient is f = "a continuous function" (does not have to be C1). The theorem of course says
∫T(D) f(x,y,z) dxdydz = ∫D f [T(u,v,w)] J dudvdw
The proof of Theorem 4 is like that of Theorem 3, they first show the differential form on page 305 A where as usual they use a "mean value theorem" along the way to construct what they want. This 305 A is like the Lemma for Theorem 3, and from it the desired result follows and Thm 4 is proved. They refer to the set T(D) as D* and it has to be closed and bounded in the range which is x,y,z space. Of course we are then going to face all those same issues of the integration chapter about how this works if D* is not compact.
[305 bottom half] Here they compare the above general result to our 1D change of variables result, and they claim for 1D, that result is "stronger". There the Jacobian appears just as some φ'(u) variable change thing. One difference is that here we must have J ≠ 0, but there φ'(u) = 0 was OK. Here we need a 1-to-1 map, but there we did not. Here we have | J | but there we just had φ'(u). Their point is that in Chapter 7 they will develop "stronger" versions of Theorem 4 at least for E2 and E3 cases. They keep referring the reader to some monograph by Rado on this subject.
Comment: What kind of worker specializes in these topics? Not someone who is trying to compute something for a practical reason. This worker is a "math theory" person who wants to firm up the basis of things, like the people who "firmed up" the notion of "the theory of distributions". This line of work has a very long, unbroken and respectable history going back centuries, but I am not so interested in this line of work. But I like having a place to go if I need details on some issue.
The section wraps up with a set of Examples.
Example 1: [306] A certain non-linear T is stated, (x,y) = T(u,v). J is computed. The volume integral is computed in two different ways with result 14/3 in each case. D and D* are shown in p 307 figure. First way is ∫D J dudv using the D region on the left side of the p 309 figure. Second way is ∫D* dxdy where D* is shown right side.
Example 2: [307] Thus uses the same T and same D and D* as in Example 1, but an integrand is added. So we have a certain dxdy integral we would like to compute. This is done using the J dudv method. The step I mark α is not very obvious to the casual reader, but a minor detail.
Example 3: [307]. Here we are given a dxdy integral to evaluate with the specific D* shown in Fig 6-4 (and its corresponding D on the right). Staring at the desired integral does in fact "suggest" the simple linear transformation of page 308 A. Since linear, easy to compute T-1 and J. They need T-1 because it is the one that has the form (x,y) = F(u,v). So this integral is evaluated in J dudv sense.
Example 4: [307] Polar coordinates. Now of course the format (x,y) = T(u,v) is staged properly for doing any curvilinear thing. I will come back here when I rethink J definition in tensor doc. J = r as usual. Formally, you have to avoid any domain D* in dxdy space which contains the origin since J = 0 there. So Fig 6-5 p 309 shows some origin-avoiding D* and then the corresponding D avoids r = 0. (I wonder if these drawings are accurate!) No specific integrals are computed here.
If D* does contain the x,y origin, they give a work-around. You break the integral into two pieces, one above and one below the x axis say. Then the integral above the x-axis might look like Fig 6-6 on the left. In Stakgold-like form, the workaround is to put a disk around the origin of radius ρ, compute things, and then take limit ρ→ 0 and hope the result exists. You have to do this for each half of your decomposed integral. In practice I know this is non-issue, but at least we have a way to do it. If some function were singular at r = 0 in some way, more attention would be needed.
Example 5: [310] Spherical coordinates with θ and φ "reversed" from modern say usage. Wiki says that my convention is the physics one, while the other is the math one. Maple uses the math one I think. Bucks compute the usual J and take note of the singular boundary pieces, and in fact have something very much like my tensor doc picture, but mine is a lot more detailed.
Comments about the above Section.
What did they do in this section?
1. Show the detL volume ratio rule for linear transformation L
2. Show complicated area transformation rule for linear transformation L: E2 → En.
3. Claim the J volume rule for a general transformation and prove this in very lengthy Thms 2,3,4.
4. Compare the J rule with change of variables in 1D
5. Restate their pet box notation for integration volume on page 306
6. Examples 1,2,3 showing how to compute certain 2D integrals two ways. One way is the brute force way doing integral as stated. Second way is to use some kind of simplifying transformation and then do the integral in the new variables with J added.
7. Example 4 shows how you might deal with an integration whose area includes a place where J = 0.
The Bucks are much more interested than I am in fine details and rigor in this business. They even claim the subject was one of current research in 1956 when the first Buck was written. Phrase algebraic topology is mentioned, as is "measure theory". I think the Bucks knew all about measure theory and the fancy methods, just as they know about complex variables, but they are trying to explain things in this chapter in their dummed-down world for the student reader. They wanted to say at least something about the Jacobian Integration Rule and their section title is "transformations of multiple integrals".
I recall when I first encountered the J integration volume rule. I was very uncomfortable about it, teachers just used it without explaining why it worked in general. Students understood the idea in polars from simple geometry, but student had no notion of what I have presented in tensor doc on this subject.
6.2 Curves and arc length [313]
Comment: In these Buck headings, only first word is cap, and there is no period.
Comment: The topic here is one which I have spent very little time on in my life, so good to have a nice source and to learn the main items.
A curve is just a mapping γ : E1 → En and think of it as γ(t), a vector in En. For a range of the parameter t, there is some trace in En, but this is not "the curve" although that is normal speech usage. The reason is that different mappings can give the same curve. Bucks curve is entirely what I have always thought of as the parametric representation of things.
Various odd things can happen. The trace can cross itself at multiple points. Thus, at a single point in En it can have different values of t, each with a different tangent vector. If the domain is [t1, t2], then you have γ(t1) as the start point and γ(t2) as the end point. If these two points coincide, the curve is closed. If there are no multiple points, the mapping is 1-to-1 and the curve is simple. A closed simple curve is also known as a Jordan curve or arc.
It is possible to create a curve which "fills all of En", but no examples are given. Probably some kind of limit thing, like a spiral in 2D with some separation parameter ε which you → 0.
Comment: A curve is then just an example of a transformation where the domain happens to lie in E1. Like any general transformation, there is a linearized version L which has a matrix representation, like the R of tensor doc. But the shape of the matrix is this on the left
If you write r(t) = γ(t) as your curve in E3, then you would have x = γ1 and so on, so that the linearized transformation is written as on the far right above. This is NOT a gradient object, it is ∂tr . This linear matrix at some point t0 and corresponding r0 = γ(t0) can only have rank 1 or rank 0, the latter in the case that all three derivatives vanish at the same value of t. If you think of t = time, then the elements of the above matrix are velocities and rank 0 only occurs when the entire velocity vector has v = 0, so your little projectile at such a t value has zero speed.
They claim without proof that such a zero-speed point will be a "corner" on the curve, so a smooth curve must never have a zero-speed point. You can assure there is no zero speed point and thus that rank = 1 if the condition p 315 A is respected. This just says v2 > 0 so cannot have speed v = 0.
Here is my way to understand the issue here. The tangent vector is always v = dr/dt and the direction of the curve is then (t). Normally (t) varies smoothly with t, giving the direction of the curve at t. Things are usually smooth because they are assuming, without saying, that the functions γt(t) are really C1 so have derivatives which are continuous. But as you approach a point t where v(t) = 0, the direction vector becomes undefined right at that point, and you can see how it might completely change direction at that point, though it might not as well.
Example 1: [315] Curve is shown in B so we have r(t) = (t3, |t3|). Components are continuous at t = 0, but curve has a sudden direction change at that point as shown figure top p315. Since v(t) = (3t2, ±3t2) = 3t2(1,±1) and (t) = (1,±1), where ± is for t 0, we see in this example that v(0) = 0, so this is a possible direction-change point and sure enough, you have that change expressed as (±ε) = (1,±1). A simpler example would be r(t) = (t, |t|) and then v(t) = (1,±1). But in this case, we don't have v(0) = 0 but we still have the change in velocity at t = 0. So that is why they didn't use that example! Other things can also give velocity changes besides v(t) = 0.
Example 2: [315] The Straight Line as special case of a curve. Nothing really new here.
Example 3: [316] This is E1→ E2 simple case, with piece of curve traces shown p 316 top. We have
t = t r(t) = (t-t3, t2-t) v(t) = (1-3t2, 2t-1).
Two interesting points:
t = 0 r(0) = (0,0) v(0) = (1,-1). // points SE
t = 1 r(1) = (0,0) v(1) = (-2,1). // points SSW
This is a double point at the origin as the figure shows, with tangent vectors as shown. Additional facts they note are
vx = 1-3t2 = 0 when t = 1/and this means curve is "vertical" at that point
vy = 2t-1 = 0 when t = 1/2 and this means curve is "horizontal" at that point
I marked these two points on the curve. I plot this curve in buck ch6 b.mws.
Example 4: [317] Here E1 → E3 and x,y are trigs with a z added in. The trace is shown in Fig 6-11 p 317.
We have
r(t) = (sin(2πt), cos(2πt),2t-t2) v(t) = (2πcos(2πt), -2πsin(2πt),2-2t) 0 ≤ t ≤2
Then
rstart = r(0) = (0,1,0) vstart = v(0) = (2π,0,2)
rend = r(2) = (0,1,0) vstart = v(2) = (2π,0,-2)
This shows that the curve is closed and the tangent vectors are different at the start and end points.
[318] You can describe a curve as being class Cm in the obvious way. If functions all have convergent power series expansions, then curve is analytic.
Curvature
Imagine some arbitrary trace in 3D space. If you look in a local neighborhood, the curve is "straight" but if you include the next derivative, it looks like a piece of a circle. Even if the curve is a spiral in 3D, a slinky, you can "fit" a circle to any point on the curve. The little piece of curve has its own special "plane", something not totally obvious. And that is the plane of the fitting circle. How can I show this? Let the curve be
r(t) = r0 + At + Bt2 // ignore the rest for small t.
The three points r0 and A and B define a plane so if you go into that plane with a piece of paper, the three points are right on the paper. Then you could write
x(t) = x0 + Axt + Bxt2
y(t) = y0 + Ayt + Byt2
Rewrite this for Maple as
x(t) = a + bt + ct2
y(t) = d + et + ft2
Now a circle looks like this
[ x(t) - g ]2 + [ y(t) - h]2 - R2 = 0
Let's have Maple install x and y and group powers of t:
At t = 0 we find that,
(a-g)2 + (d-h)2 - R2 = 0 // a relationship between unknowns g,h,R
To make the t1 term vanish, we need to have
2(a-g)b + 2(d-h)e = 0 // equation in unknowns g, h
To make the t2 term vanish we need to have
2(d-h)f + e2 + 2(a-g)c + b2 = 0 // equation in unknowns g,h
So these last two equations would determine the circle center (g,h) and then the first tells you R.
R2 = (a-g)2 + (d-h)2
Can Maple do all this?
I can grab these solutions as g1 and h1 and then compute R from them as follows:
So my answer is this
R2 = (1/4) (e2+b2)3 / (bf-ce)2
In the language of Bucks p 318. my result is saying [ I guess k = 1/R ]
k2 = 4 (bf-ce)2/(e2+b2)3
which looks a little like their result. Go back now to
x(t) = a + bt + ct2 x'(t) = b + 2ct x"(t) = 2c
y(t) = d + et + ft2 y'(t) = e + 2ft y"(t) = 2f
Then we have
r(0) = (a,d) |r(0)|2 = a2+d2
r'(0) = (b,e) |r'(0)|2 = b2+e2
r"(0) = (2c,2f) |r"(0)|2 = 4c2+4f2 r'(0) r"(0) = 2bc + 2ef
|r'(0)|2 |r"(0)|2 - [r'(0) r"(0)]2
= ( b2+e2)( 4c2+4f2) - (2bc + 2ef)2
Also |r'(0)|6 = (b2+e2)3
If I plug these results into THEIR formula for k2 the result is:
kBuck2 = 4(bf-ce)2 / (b2+e2)3
But my formula from above was this,
k2 = 4 (bf-ce)2/(e2+b2)3
and they are the same! So I now have my own little derivation of (6-17).
Review of my derivation:
(1) Write the curve near t = 0 as r(t) = r0 + At + Bt2 in 3D space.
(2) Those 3 points define a plane, so without loss of generality, assume that r0 and A and B are vectors which lie in the x-y plane. Rewrite the curve in components
x(t) = a + bt + ct2
y(t) = d + et + ft2
(3) Look for a circular solution for small t of the form
[ x(t) - g ]2 + [ y(t) - h]2 - R2 = 0
where (g,h) is the circle center and R the radius, and these are three unknowns.
(4) Insert (2) into (3) and group by powers of t. Set the coefficients of t0, t1 and t2 = 0 to get these three equations:
(a-g)2 + (d-h)2 - R2 = 0
(a-g)b + (d-h)e = 0
2(d-h)f + e2 + 2(a-g)c + b2 = 0
where unknowns are shown in red.
(5) Solve the last two equations for g,h to get
(6) Insert these solutions into the first equation to get
R2 = (1/4) (e2+b2)3 / (bf-ce)2
(7) Note that
r(0) = (a,d) |r(0)|2 = a2+d2
r'(0) = (b,e) |r'(0)|2 = b2+e2
r"(0) = (2c,2f) |r"(0)|2 = 4c2+4f2 r'(0) r"(0) = 2bc + 2ef
[r'(0) x r"(0)]3 = b2f - e2c = 2(bf-ec) = | r'(0) x r"(0) |
so we can say that our result is
R2 = (e2+b2)3 / [4(bf-ce)2] = |r'(0)|6 / | r'(0) x r"(0) |2
and then my result is
k = | r'(0) x r"(0) | / |r'(0)|3 // just another way to write things.
(8) Now show that
|r'(0)|2 |r"(0)|2 - [r'(0) r"(0)]2 = ( b2+e2)( 4c2+4f2) - (2bc + 2ef)2 = 4 (bf-ce)2
Our solution is then
R2 = |r'(0)|6 / [|r'(0)|2 |r"(0)|2 - [r'(0) r"(0)]2]
and this agrees with (6-17) which shows k = 1/R.
So I am quite happy with this little curvature solution.
[318] As motivation Bucks do show that R = ∞ for a straight line and R = R at all points on a circle.
[319] They then put the curve into the x-y plane with no comment as to why that is allowed. In this plane they have figure 6-12 which is all correct. They then have a ton of algebra and they end up with their proof finished on p 320. Perhaps my proof is easier.
Arc Length [ p 320]
The formula they give is fine by me:
L(r(t)) = !Syntax Error, I | dr/dt| dt = !Syntax Error, Iv(t)dt = !Syntax Error, I ds = !Syntax Error, I|dr|
where you are basically making a differential polygonal fit to the curve and then adding up the lengths of the edges. But they are going to prove all this in detail.
Question added 2.12.16. Think of a car making a trip along the curve with v = dr/dt. Is the car allowed to drive backwards along the path for some portion of the trip? If it did this, there would at least one time where the velocity was v = 0. But then the curve γ is not a "smooth" curve as defined at the bottom of page 314. There it says the differential matrix (a column vector) has to never be rank 0, and rank 0 would mean that all components of the velocity were zero. In all that follows in this chapter the Bucks always require that the curve γ (that is, the mapping γ) be "smooth". Thus, it is always understood that your car cannot go backwards for a portion of the trip. Another term "simple" means that a curve γ has no multiple points (crosses itself). I think if a car stopped and backed up 3 feet then resumed forward, you could say that the path was not simple since it then crosses itself in a continuum of points! But smooth is the main idea here.
Formal definition says use all polygonal fits and the arc length is the LUB (the min) of all those fits.
Some curves have infinite arc length and are not rectifiable.
Theorem 5: [321] Assuming a curve is rectifiable, the arc length is given by (6-20).
Proof is given, nuts and bolts, uses a MVT, uses ε and δ, gets the limiting sum into the Riemann sum for an integral form, then that does it. My personal proof is more down to earth. Let v(t) = |γ'(t)| be the speed of a tiny curve-driving vehicle at time t. As is always the case, distance d = ∫v(t)dt. If the curve were a piece of non-stretchable string, just pull into a straight line and drive it with the same speed profile and it has the same length.
Example 1: [322] just shows the computation of a simple arc length integral, a form of helix in 3D.
Parametric Equivalent Curves = Smoothly Equivalent Curves
Two curves are equivalent if they have the same trace and the only difference is that they have a different parametric speed function. One curve is r(t) while the other is p(t) = r(f(t)) where f(t) changes the speed at each point in some arbitrary way. We have vp = dp/dt = dr/dt f'(t) = v(t) f'(t). Obviously the two curves have the same (t) but not the same v(t) since in general f'(t) ≠ 1.
To avoid trouble, you want f'(t) > 0 over entire curve so you don't run it backwards!
Theorem 6; [323] Two equivalent curves have the same direction at the same point p in space.
This direction is (t) . It seems obvious since curves have the same trace! Note that v(t) is not the same at the same point exactly because we are re-speeding the curve. [ this noted in footnote]. I have already shown this above. They give a short proof.
Theorem 7; [323] Smoothly equivalent curves have the same arc length.
Again, seems obvious since both have the same trace which is what determines arc length.
___________________________________________________________________________________
After writing the following "note", I realized that I had already written an entire doc on this subject called "respeeding a curve.doc" which I failed to reference here in these Buck Ch 6 notes!
Note added 2.12.16. I will prove the above claim. Things are more complicated that I first thought, so the first thing we need is a picture showing the notion of curves γ and γ* being "parametrically equivalent" ,
I have drawn the two segments on the t axis as non-intersecting, but they could intersect and nothing changes. The mapping from interval (α,β) to (a,b) is not an arbitrary mapping. It must map the endpoints into the endpoints, and it must have f'(t) > 0 over the entire range of t of interest. We also have this relationship that γ*(t) = γ(f(t)). Since f'(t) > 0, the mapping f(t) I think is one-to-one and or course a continuous mapping.
This mapping is what I call a "respeeding". It helps me to imagine the two curves as being two car trips along the same path in En . Car C has a trip running from time t = a to t = b and duration b-a. Car C* has a trip running from time t = α to t = β with duration β-α, along the same trace. The trip durations in general are different. Car C has some arbitrary speed profile v(t) during its trip, v(t) = |γ'(t)| where γ(t) can be any vector valued function you want (which does not vanish).
Car C* has speed v*(t) = |γ*'(t)| = v(t) df/dt at time t in its trip, so car C* has the same arbitrary velocity profile as car C but scaled by df/dt. Since the curve γ is "smooth", that arbitrary velocity profile v(t) can never vanish, so neither car can "back up" during a portion of the trip. That is, you can never have v = 0 at some point during the trip. So we assume that car C always goes "forward" during the trip. Since df/dt > 0, this means that car C* also always goes forward during its trip.
As a special case of the above general situation, you could have (a,b) = (α,β). Then the two cars make the trip in the same time period, but they have different speeds v*(t) = v(t) df/dt . Modulation of the speed by a positive function makes no difference in the distance traveled. This is perhaps where the term "respeeding" makes more sense. So here then is the proof:
The arc lengths of the two curves are ( γ is always bolded, it is a vector γ(t) = r(t))
d = !Syntax Error, Iv(t) dt = !Syntax Error, I|γ'(t)| dt f'(t) > 0 for t in [α,β]
d* = !Syntax Error, Iv*(t) dt = !Syntax Error, I|γ*'(t)| dt = !Syntax Error, I|dt(γ*(t))| dt = !Syntax Error, I| dt( γ(f(t)) ) | dt
We want to show that the arc lengths are the same, that is to say, that d* = d.
Consider:
dt( γ(f(t)) ) = dfγ(f) dtf // chain rule
| dt( γ(f(t)) )| = | dfγ(f)| dtf // assuming dtf > 0 for entire range (α,β) of t
This last line says
|γ*'(t)| = |γ'(f)| df/dt
Installing this last line into the d* expression above gives,
d* = !Syntax Error, I|γ*'(t)| dt = !Syntax Error, I |γ'(f)| (df/dt) dt
Now change integration variables from t to f(t). In this change we find
(df/dt) dt = df t = α → f = f(α)
t = β → f = f(β)
Then we find that
d* = !Syntax Error, I |γ'(f)| (df/dt) dt = !Syntax Error, I |γ'(f)| df = !Syntax Error, I|γ'(f)| df
= !Syntax Error, I|γ'(t)| dt = d QED.
Two parametrically equivalent curves therefore have the same arc length. If you were integrating a function in the integrand, the above argument would not work.
Extra Credit. Show that q = q* where
q = !Syntax Error, IF(γ(t)) |γ'(t)| dt f'(t) > 0 for t in [α,β]
q* = !Syntax Error, IF(γ*(t)) |γ*'(t)| dt = !Syntax Error, I F(γ*(t)) |dt(γ*(t))| dt = !Syntax Error, I F(γ(f(t))) dt( γ(f(t)) ) | dt
Here I have added a new factor in the integrand which was "1" in the above derivation. I proceed as above:
q* = !Syntax Error, I F(γ(f(t))) |γ*'(t)| dt = !Syntax Error, I F(γ(f)) |γ'(f)| (df/dt) dt
Change integration variables as before from t to f to get
q* = !Syntax Error, I F(γ(f)) |γ'(f)| df = !Syntax Error, IF(γ(f)) |γ'(f)| df
= !Syntax Error, IF(γ(t)) |γ'(t)| dt = q QED
The proof is essentially the same but you have added this extra factor in the integrand which is a scalar function of the points γ in En along the curve.
__________________________________________________________________
What happens if you select for parameter t the arc length itself? That means changing from t to g(t) as shown You then have vnew = dr/dt g'(t), but g'(t) = 1/ |r'(t)| so vnew = [r'(t)]^ by which I mean a unit vector. The direction varies, but the speed is always exactly 1 ! So such a parameterized curve has unit speed everywhere.
[324] We then have a qualitative discussion of curve theory. People tend to work with classes of curves like the class described above (respeeding, same trace). Different speed functions give different classes.
An algebraic curve is one of the form y = f(x) which is same as x = t, y = f(t) to get it into standard form. Or you might have F(x,y) = 0. You might involve polynomials or rational functions. Algebraic functions I think in general involve roots of polynomials. Hazy phrase. Bucks are just jawing along here a bit on things we won't be doing.
[325] What happens to a curve under a transformation having Jacobian J = det T? A closed curve maps into a closed curve, but the mapped curve might have more loops than the original one (winding number about some point). Example of this page 326 figure. Another issue is the relative "orientation" of the two curves. For E2→ E2 this is set by the sign of det T. For E3→ E3 you need to make a sphere to hold the original curve (or some closed surface), then watch how curve and surface maps to determine orientation. This is shown bottom page 326.
Next comes discussion of angles between two curves which intersect, sent through a mapping T. If the angle is preserved, the mapping is called conformal, as I well know.
Theorem 8. [326] If a mapping E2→ E2 is conformal, the linear matrix they call dT must have the form shown top page 327.
As the footnote notes, this matrix form is just a statement of the Cauchy-Riemann equations. They give a long proof of why the matrix has this form for people who don't know about complex theory. Finally, Example 1 on page 328 shows that a conformal mapping preserves angles ONLY at points where J is not 0. Their example shows such a violation at the origin, top page 329.
They note that the whole conformal concept only applies to E2 → E2 and they claim there is no similar concept for E3→ E3, a topic I have wondered about from time to time. Only conformal transformations in this case are those for rigid motions. Things are too over-constrained in 3D to allow conformal idea.
6.3 Surfaces and surface area [330]
Bucks proceed on analogy with the curve section. A "surface" Σ is a mapping Σ:E2→ En. Several surfaces may result in the same trace, as with curve mappings. New feature is that you can now follow the boundary of a domain D through the mapping, where it becomes the boundary ∂Σ of the trace in En. Notions of multiple points (as shown p 331 top right), and simple = no multiple points. That is, you have multiple points if a surface trace intersects itself.
[332] Like any transformation, T = Σ has a dT matrix which I would call Sij in tensor doc, but tensor doc only talks about En → En . For E2 → E3 the dT is shown page 332. A surface is smooth if this dT has its maximum rank which is rank 2, which is same condition as (6-25).
[333] Example 1 and Example 2. For each we look at the domain (disk or rectangle) and range (cone), and in particular we look at the boundary mappings which are slightly different for the two cases. I am happy with all this stuff. I like the artist depiction of cones, I still have no way to do that in Visio which is a 2D package, no shading. Well, I can do this which is not too bad.
[334] For Σ2 example Bucks compute dΣ2 as 2x3 matrix and we see as expected that it maintains rank 2 except at (u,v) = origin which is the bottom of the cone, so cone is therefore not "smooth" at that point.
Example 3: Spherical coordinates with r = 1. Here the domain is the (θ,φ) plane and a D there maps into a patch on the spherical surface. The boundary is a little strange: As you go around the boundary in θ,φ space, you run a the prime meridian one way and then back the other way, where the poles account for the top and bottom edges of D.
Cases where the trace in E3 is a plane. This is a little different than what I am used to. Bucks start with this
r = r0 + ua + vb Σ: E2(u,v) → E3(x,y,z) r = f(u,v)
This is a mapping, and of course it is a linear mapping example. In the range consider this candidate normal vector
n = a x b so that na = nb = 0 .
Then the equation above says
nr = nr0 = d
and this is MY equation for a plane in E3. So the mapping shown above always gives a planar trace.
Now we know that
n = a x b => ni = εijkakbk
n1 = a2b3- a3b2 =
n2 = a3b1- a1b3 =
n3 = a1b2- a2b1 =
In this way, Bucks express the components of a normal vector in terms of 2x2 determinants of the dT = dΣ transformation matrix. Note that
a2 = ∂y/∂u b2 = ∂y/∂v and so on.
Thus, we can expect a notion of a tangent plane and then the normal at a point r such that n is the normal of that tangent plane, and we can then write
n1 = = = ∂(y,z)/∂(u,v)
This matches their statement at the bottom of page 335 -- an expression for the normal to a surface at some point on the surface in E3. In the page 335 definition they define the above to be the normal, then comes the theorem
Theorem 9: [336] Let Σ = smooth surface trace containing point r. Let Γ be a any smooth curve on Σ passing through r which has a velocity vector V at r. Then claim is that n V = 0 .
Of course if you took two such curves Γ1 and Γ2 which did not coincide, with V1 and V2, you would have both n V1 = 0 and n V2 = 0 and then V1 and V2 would span the tangent plane at point r.
Proof Part 1: Here is the scenario in Buck language
As t varies, we get a γ(t) curve in (u,v) space and this becomes curve Γ(t) in (x,y,z) space. This is a simple concatenation of mappings, and we know from Buck Theorem 18 page 265 that
d(Σγ ) = (dΣ)(dγ)
What are these two matrices? "row for each var on left, col for each var on right" :
(u,v) = γ(t) dγ =
(x,y,z) = Σ(u,v) dΣ =
Then
d(Σγ ) = (dΣ)(dγ) =
= = agrees with p 336 A
So I have verified p 336 without the starting part which says V =
Proof Part 2:
We know that Γ(t) = Σ(γ(t)) where γ(t) is some smooth curve in the domain D. And we know that
V(t) = Γ'(t) so we want to show that V(t) n(t) = 0. As a preliminary, note that
γ(t) = [u(t),v(t)] = our curve in D = [u,v] γ'(t) = [u'(t),v'(t)] = v(t)
Γ(t) = Σ [ γ(t) ] = our curve in range space (x,y,z) r = Σ(u,v)
V(t) = ∂tΓ(t) = ∂t Σ [ γ(t) ]
Vx(t) = ∂tΓx(t) = ∂t Σx [ γ(t) ] = ∂t Σx [u(t),v(t)] = ∂t x [u(t),v(t)] = ∂x/∂u ∂u/∂t + ∂x/∂v ∂v/∂t
So I have now derived the part of p 336 A which says V = ( ***, ***, ***)
The rest of the Buck proof follows from just doing Bruce Force to show that n V = 0 . But I think I can do better than this?
V(t) = ∂tΓ(t) = ∂t Σ [ γ(t) ] = ∂t r [u(t),v(t) ] = (∂ur)(∂tu) +(∂vr)(∂tv)
Go back to the linearization
r = r0 + ua + vb = r0 + ua + vb
a2 = ∂y/∂u b2 = ∂y/∂v a = ∂ur b = ∂vr
n = a x b = (∂ur) x (∂vr)
Then we have
n V = (∂ur) x (∂vr) { (∂ur)(∂tu) +(∂vr)(∂tv) } = 0 + 0 = 0
because a x b is normal to both a and b.
Now the only missing thing is the Buck notation (dΣ)(dγ) in line A. They are saying
∂t Σ [ γ(t) ] = (dΣ)(dγ)
Why is this true? It has been fighting me for a few hours now. Above I showed it is true, but I don't know why it is true in a general sense. Chain rule says
∂t Σ [ γ(t) ] = ∂t (Σγ(t)) as if were a single function Σγ
∂t Σ [ γ(t) ] = Σk=1q ∂Σ({ui})/∂uk * ∂uk/∂t ui are the intermeds like (u.v)
Go back just to Σ. "row for each var on left, col for each var on right"
(x,y,z) = Σ(u,v)
dΣ = ∂(x,y,z)/∂(u,v) = ∂(x1,x2,x3)/∂(u1,u2) = ∂xi/∂uj = ∂Σi/∂uj
= matrix with 3 rows and 2 columns
(u,v) = γ(t)
dγ = ∂(u,v)/∂(t) = ∂ui/∂t = column vector = matrix with 2 rows and 1 column
(dΣ dγ)i1 = Σk (dΣ)ik (dγ)k = Σk ∂Σi/∂uk ∂uk/∂t
But comparing to above,
(dΣ dγ)i1 = ∂t Σi [ γ(t) ]
OK, I guess it just "happens to be true". Uncle.
[337] Looking at the mapping which produces a plane in E3
(x,y,z) = Σ(u,v)
r = Σ(u,v) = r0 + uα + vβ
we already know that (if you don't know that, see p 337 B)
α = ∂Σ/∂u β = ∂Σ/∂v
Note: using these α and β in the r formula above gives you the tangent plane at a point on a surface.
Meanwhile, for a vector function of two variables we can write a 2-variable Taylor expansion as in p 337 C. This was derived earlier for a scalar function and here we just apply that scalar expansion to each component of the vector function Σ.
The higher derivatives relate to curvature in various directions and so on I guess, but Bucks do not pursue that subject here and tell us to go read elsewhere.
Theorem 10: [338] If surface is defined by g(x,y,z) = 0, then g is the normal.
Bucks proof is fine. I have one from earlier doc as follows:
g(r+dr) = g(r) + g dr
If r lies on surface, then g(r) = 0 so we have g(r+dr) = g dr . If dr lies on the surface, then we know that g(r+dr) = 0 as well, in which case g dr = 0. In order for this to be true for any dr on the surface, we must have n = g QED.
Bucks prove this by treating the surface as a mapping from u,v space from which they conclude that
g ∂r/∂u = 0 and g ∂r/∂v = 0. I agree that this can be written g α = 0 and g β = 0 from page 337 A. But α and β span the tangent plane, so yes, g must be normal to that plane.
I think my proof is more to the point if informal.
Area Formula. [338] Says that area = ∫dudv | n(u,v)| where n is the non-unit vector normal given in (6-29).
My Proof: This normal is n = a x b (as I showed directly, and as footnote page 338 verifies) and those are the spanning vectors of a tangent-plane 2-piped which has area |n| = abcosθ. Apply this linear thing to the localized general transformation.
Comment: Think of n = e1 x e2 and E3 = det(R) e1 x e2 = e3 tensor doc A (b).
dAn = J dA'n en
dA3 = J dA'3 e3 = J dA'3 E3 = det(S) dA'3 det(R) e1 x e2 = e1 x e2 dA'3 = n dA'3
This last line shows that |dA3| = |n| dA'3 so |n| is what "scales area" just as Bucks show.
Tensor Doc Exercise. Start with the idea that n = e1 x e2 as a vector in x-space. Assume x-space is Cartesian so g = 1. In developmental notation write ni = εijk(e1)j(e2)k = εijkSj1Sk2. Then for example
n1 = ε1jkSj1Sk2 = S21S32 – S31S22 =
n2 = ε2jkSj1Sk2 = S31S12 – S11S32 =
and so on, so you then get
|n| =
So I leave out the details, but this shows how tensor doc comes up with this same form shown in (6-30).
How about this instead:
n2 = εijk(e1)j(e2)k εij'k'(e1)j'(e2)k'
= εijk εij'k' (e1)j(e2)k(e1)j'(e2)k'
= [ δjj'δkk' - δjk'δkj'] (e1)j(e2)k(e1)j'(e2)k'
= e1 e1 e2 e2 - (e1 e2)2 = '11'22 - '122 = cof(')33 = not very useful pathway
Bucks don't prove this theorem, they just motivate it. On p 339 the first show it works if the area patch lies directly in the u,v plane so that x = u, y = v, z = 0. This causes |n| = 1 so area = ∫dxdy = correct.
Next they kick it up a notch by assuming a linear transformation of the planar type they have been discussing at length. For this transformation they compute the formula 6-30 and arrive at p 339 B. But this mapping produces the rotated 2-piped shown in figure on page 298. For this picture, which shows result of a linear map on u,v rectangle a,b, they find the area as αβsinθ and they show on p 299 in passing that this is expressible as p 299E, which then is more "motivation" for 6-30. They note that this is not a proof, in the ε and δ sense of their other proofs. But they do say that you can break your curved surface into a mesh of 2-pipeds by mapping over from an even mesh in uv space, and they you add up the areas and you have the desired result, picture page 340 top.
Notice that the area scaling factor requires you to evaluate three 2x2 Jacobians which are the three minors of the 3x2 matrix they call dΣ (the R matrix) , the differential for the transformation, as shown in p 339C.
They then have a few examples:
Example 1: [340] Area of a donut of small radius 1 and large radius to center line R. They compute the dΣ matrix and from it the three 2x2 Jacobians and they use the 6-30 then to get the result shown. Picture is weak, I would have had some smaller radius so dimensions come out right.
Example 2: [341] This one is presented in the form z = f(x,y) but that is easily cast into the general form of x = u, y = v, z = f(u,v) and then the general formula is applied to get |n| . In both examples you of course have to do the area integral, and that was trivially doable for the donut, but is less obvious for this example. Good old Maple says,
which is the final answer they give.
Note 1: Bucks comment that "area theory" is still in the process of development, their books 1956 and 1965 editions. Hard to believe such a claim but must be true. Again, the world of a mathematician
Note 2: It is fascinating that you can exactly calculate these weird surface areas! Newton and friends would be amazed, I am not sure whether they knew these formulas, perhaps yes. But of course adding Maple to the mix is very powerful just to do the integrals (either way). I don't know if area calculations like this are in my first year calculus book, I will have to look.
Note 3: Compare to tensor doc when time. [ did a little of this above ]
Global Aspects of Surfaces [343] Bucks give a quick little tour, fascinating stuff. The field is called algebraic topology.
Definition of a 2-manifold M. This thing is not a surface, or trace, it is a set of mappings which go from domain D to E3. Each little mapping Σi maps open D onto a surface Si in our local sense already discussed. You then assemble the patches to get your global surface in E3. It is like taking pictures of Mars, see figure page 344. Each patch mapping must be class C' and must be 1-to-1 and therefore invertible. If the global trace is some visualizable surface, I guess you can surely cover it in this manner. Somehow I think you are maybe covering the manifold with the unions of open sets, and this brings in Heine-Borel, and recall the definition of a "topology" as the set of all open sets for something. So in this manner, Bucks are extending our local understanding of a "surface" to a global object "in the large" as they say.
Now look at figure page 345. You can of course have some Σi and Σj as defined by this picture, and since each is invertible, you could construct Tij ≡ Σi-1 Σj , all functions of local point p. If all Tij thus formed have Jij ≠ 0, then you have a differentiable 2-manifold of class C' . I presume p with J ≠ 0 would represent a non-invertible point even though we have already said things are 1-to-1. Now if it happens that all the Jij have the same sign, say are positive, then you have an orientable 2-manifold. If the J-sign-rule is violated, your manifold is non-orientable.
Fig p 346 gives an example where we look at some of the Σi mappings and their traces Si. At least for this local piece of surface, a selected normal in D maps all in the same surface direction. You might call this an outward normal. An orientable manifold has an "outward normal" that is well defined. The Mobius strip of page 347 is the usual counterexample, and if you track what starts as an outward normal around the surface, it gradually becomes an inward normal, so there are not then two well-defined "sides" of the surface. In 6-33 Bucks give a global mapping for a Mobius strip as r = Σ(u,v).
[ At this point I went off and wrote a side document on this transformation, then came back here.]
Something is wrong with the Buck discussion bottom page 346. They first seem to do this:
D1 = |v| < 1 and 0 < u < 2π.
D2 = |v| < 1 and |u| < π/2
Here is a picture: (left side)
These domains clearly overlap, and their traces will then overlap. I think the trace of D1 alone covers the entire Mobius strip. So far so good. But then they say "the sets D1 and D2 are not connected" which means they are disconnected. This statement just seems wrong. However, if we take the range of the transformation itself to be 0 < u < 2π, then we really have to redraw the blue piece as shown on the right, and then yes, D2 is disconnected and has two pieces. So at least with this interpretation, D2 is disconnected, though D1 is still connected (in contrast to their statement that both are disconnected). But they the two disconnected pieces of D2 they refer to as D1' and D1" adding further confusion. So let's assume that they got the original D1 and D2 definitions reversed, and assume that this is the picture they have in mind:
Then equations A make perfect sense. D1 is disconnected and has the two pieces claimed. Now let's take these two pieces D1' and D1" to be the two disconnected domains shown in the Fig on page 345. These domains are connected by some T21 as shown in that figure which transformation connects the two disconnected regions D1' and D1" . Now suppose (u,v) is a point in D'1 shown above. We could specifically define T21 by saying this point goes to (2π-u,v). Or a different T21 would say that (u,v) goes to the point (2π-u,-v) and this is the T21 they select. Then we have
u' = 2π-u J = ∂(u',v')/∂(u,v) = = +1
v' = -v
So this seems to be another error. Let's assume that they meant the (2π-u,v) transformation. Then
u' = 2π-u J = ∂(u',v')/∂(u,v) = = - 1 = negative
v' = +v
So clearing up these various Buck errors, here is what they are doing: They have constructed specific disconnected domain sets D1' and D1" which are linked by a 1-to-1 transformation T21 , but this transformation has a negative Jacobian, and therefore the surface is disqualified from being orientable according to the definition on p 345.
But still it is not right. The mappings of D'1 and D"1 do not have overlapping traces as required by the Figure p 345. It is true that D1 and D2 have overlapping traces. So it makes more sense to use the original D1 and D2 in this discussion. I just think the whole thing is messed up and slipped through their proofing. Their goal was to find a T12 that is relevant and which has a negative Jacobian.
Maybe I can peek at the google book 3rd ed to see what they have done about this. In that book surfaces start on page 417, but too many pages are missing to find it. Index says Mobius on p 436 and 490 only. I can see 490 and it just has the Mobius picture with no analysis. p 436 is absent and might have something I suppose.
So I guess I can't tell if they cleaned up this little problem. Moving on. They must hire a company to scrub PDF's off the web! It is at Marriott sitting on a shelf if I am really interested (this is why I live in a city!)
Advanced calculus
Buck, R. Creighton (Robert Creighton), 1920- Buck, Ellen F. 1978
Available at Marriott Library LVL 1: General Collection (QA303 .B917 1978 )
Continue p 347.
A closed manifold is one which has no open boundary pieces. A complete sphere has no boundary at all and is a closed manifold. If you punch out a closed disk, there is an open boundary on the punched sphere which is not part of the punched sphere. Point sequences which approach this boundary do not have limit points in the manifold, so the manifold is an open manifold. So seems like a closed set to me.
I guess if you punched out an open disk, the remaining punched sphere would be closed since it would include the boundary of the hole. But this conflicts with Buck statement " a closed manifold has no edge or boundary", so they are not consistent.
Picture page 6-30 shows how you might glue local surface pieces together. Some boundaries cancel, but you might be left with a resulting boundary overall.
So they are just throwing out a few quick "algebraic surface" facts to the reader. I found a PS file on this subject and installed GhostScript and GhostView and can now at least see such files, but it talks about 2-forms and things Buck's have not yet addressed.
ok to here
6.4 Extremal Properties of functions of several variables [349]
I am used to the idea that at a hilltop or valley bottom, I expect the partial derivatives to be 0. Any point where they all vanish is called a critical point. But beware.
It is true that at any such critical point for a surface z = f(x,y) the normal is "up" as shown page 350. But this can happen at a saddle point as shown there, which is not a hilltop.
Bucks show as I derived long ago that n = (-f1,-f2,1) and this is why the normal is up or (0,0,1) at a critical point.
Example 1: [350] Take a simple 1D function f(x) = 4x3-15x2+18x, graph on the left:
There are visibly two critical points (zero slope places) in the interval, but the true max and min on the interval are at the endpoints! If you could not "see" this plot, you would check each of the critical points' f values and the two boundary point values before you concluded where the maximum of f was on this interval. A second point about this example: there are two interior places where f'(x) = 0, one is a local max, the other is a local min.
Example 1A: f(x) ≈ 4x3-15x2+18.7x shown on the right above. This has a single place where f'(x) = 0, but that place is neither a local min nor a local max! It is just a pause. It is also an inflection point where the curvature changes. So this shows that
f'(x) = 0 at interior point that point is a local min or max
I don't think you can call this a 1D saddle point.
Example 2: [351] Now we have a 2D domain function f(x,y) = 4xy - 2x2 - y4 D = 2,2. We find
f1 = 4y-4x So solve these equations: y-x = 0 y=x
f2 = 4x - 4y3 x - y3 = 0 x(1-x2) = 0 x = 0,1,-1
so yes there are three critical points which are (0,0), (1,1),(-1,-1) which lie in our 2-edge square. Now your study of where the true min and max are located is harder. In theory you have examine these three critical points, and all points on the 2D boundary! Now I happen to have a plot handy:
You can see that for a max, you can ignore Q3, and since f(x,y) = f(-x,-y) also Q2. The little humps above are in Q1 on lower right and Q3. So focus only on Q1. It turns out that critical point (0,0) is a saddle, and that (1,1) is then your true max, with same height as (-1,-1). But you have to study the boundary carefully. For min, it is way down somewhere on the boundary.
Here are some earlier notes I made on this Example's plot, noting the date of the Buck book etc:
The Maple plot for Fig 6-32 is this (1st quadrant only)
You see a little hill in the 1Q which has a maximum at (1,1), and there is one also at (-1,-1) not shown. The artist's depiction on page 351 is totally mysterious to me! Maybe the tear drop is the part of the surface that lies above the z = 0 plane. It is a no-goodnick picture IMHO. Here is a more global view showing all four quadrants
One must remember that in 1965 there was no Maple. It came out in 1980, a full 15 years later! And the HP-35 calculator came out in 1972. I casually make the above picture, but how were the Bucks to do it? There were PDP-1's and I guess it would have been possible. About 53 PDP-1's existed in 1969! I guess they had to do this with someone's PDP-1 or just do points by hand.
Theorem 12: [352] (1D Local Min or Max Theorem). Consider f which is C" on [a,b] and f'(c) = 0 for some interior c, so c is then a critical point. Then:
f"(c) < 0 => c is a local maximum point
f"(c) = 0 => c is not a local minimum
In the second case, c might not be if it happens to be a "saddle point" in 1D -- a place where the slope perhaps decreases to 0 but starts increasing again. A simple proof is given based on Taylor's theorem.
Question: What is the case in 2D for f = (x,y) ?
Example 3: [352] f = xy Here is a plot of xy:
It has a saddle point at (x,y) = (0,0). Note that ∂xf(x,y) = y and ∂x2f(x,y) = 0 and similarly for the y derivatives. I am not sure what the point of this example is. It is a saddle point, and f1 = f2 = f11= f22 = 0 at that point (0,0). Based on these four facts you might have thought it could not be a saddle point since the two curvatures are not opposite sign. The catch here is that these two directions happen to have zero curvature, but if you looked at axes located ±45 degrees, you would find + curvature in one direction and - curvature in the other and you would conclude yes a saddle. So this is a false theorem:
saddle => critical point and curvatures are non-zero and have opposite sign // false
Theorem 13: [353] (2D Local Min or Max Theorem) Let Δ = f122 - f11f22 at some critical point. That is, we also know that f1 = 0 and f2 = 0. Then
Δ > 0 => saddle point
Δ < 0 => local min or max (max if f11< 0, min if f11> 0)
// note that f11 = 0 would say Δ > 0, so that would be the first case
Δ = 0 => no conclusion
Apply this to Example 2 with f(x,y) = 4xy - 2x2 - y4 D = 2x2. Only critical points are (0,0), (1,1),(-1,-1) so we only need test them. We have
f1 = 4y-4x f11 = -4 f12 = -4 Δ = 16 - (-4)(-12y2) = 16 - 48y2
f2 = 4x - 4y3 f22 = -12y2
(0,0): Δ = 16, so saddle point
(1,1): Δ = -32 so extremum, and f11 < 0 so its a max
(-1-1): same as above case
Proof of theorem 13: The proof has an embedded Lemma which is complicated to even state. The proof of that lemma is a half page ending at its black square, then the theorem 13 proof using this lemma goes on for another full page. My guess is that this is purely mechanical proof and I am going to skip it. I have full confidence in the theorem and am happy to have a proof stored in Buck of this NEW theorem to me. Maybe I ran into it somewhere in the distant past.
Qualitative Discussion. Peaks and holes appear on topo maps (level curves) as shown p 355 top left, and a saddle point appears as on the right there. You have uphill E and W, but downhill N and S. Notion of a triple saddle shown bottom page 355, so I guess you could have in general an N-saddle where the 2-saddle is the usual thing. The 3-saddle they call a monkey saddle (two for legs, one for tail, Bucks fail to mention that fun fact).
Example 3: [356] f(x,y) = x2 + y3 - 3xy. Analysis:
f1 = 2x-3y f11 = 2 f12 = -3 Δ = 9 - (2)(6y) = 9 - 12y
f2 = 3y2-3x f22 = 6y f21 = -3
Critical points must solve x = (3/2)y
y2 = x y2 = (3/2)y y = 3/2 or y = 0
so critical points are (9/4,3/2) and (0,0)
(0,0) : Δ = 9 - 12y = 9 Δ>0 saddle point
(9/4,3/2) Δ = 9 - 12(3/2) = 9 -18 = -9 < 0 extremum, f11 = 2 > 0 so min.
Here are Maple plots of this surface. The right shows the saddle and the min, the left shows a topo style for the map which I aligned to match the page 356 figure.
Example: [356] (Least Squares Fit with two parameters a and b). Minimize the cost function f(a,b)!
This is a more complicated example. We have some experimental data points (xi, yi) which Ng would say compose an n-element "training set" with "two features". We want to fit the data with the best linear fit which we describe as F(x) = ax + b. We want to minimize f(a,b) = Σj=1n [ F(xi) - yj]2 which Ng would call the "cost function". Can we find values of a and b which create a min for f(a,b)?
I did all the work in pencil in the book page 357. We compute the usual suspects: f1, f2, f11, f22, f12. To simplify notation, they introduce mean values for the xi and yi in the obvious way. To find the critical points we have to solve f1= 0 with f2 = 0 which means we have to solve (6-37). In (a,b) space each of these equations describes a line, so they generally must intersect at a single point which is then the only critical point. We find that (a,b)critical = p 357 A. We compute Δ and find that Δ < 0. We then note that f11 > 0, so we indeed have a min ! Since there is only one such minimum, and since the boundary blows up in all directions, we know this is a global minimum, and this is the correct a and b which minimizes the least squares fit. In my Ng notes this problem would have a,b replaced by θ0, θ1 and θ2 I think and I show an exact linear algebra solution to the problem, but forget that here.
Bucks then comment that you might look for an extremum numerically with a "high speed computer". One method Bucks call "method of steepest ascent" which means you start some where and numerically start marching up the hill, using my well known fact that the 2D gradient points uphill. If steps are small and topology is not too bad relative to your starting point, it will work just fine. Ng always went downhill since trying to minimize a cost function, so he always talked about "gradient descent".
Theorem 14 [358] If 22D f(x,y) = 0 on bounded and open domain D and if f is C", then the min and max of f must occur on the boundary ∂D. That is to say, a harmonic function has its max and min on the boundary ∂D.
I accept the proof since I know this is true. My physical proof is that a rubber membrane cannot have a local min or max out in the middle where both curvatures are up, say, since nothing to hold it that way. The Laplace equations says the sums of the curvatures are 0 in the two directions, so if cups up one way, then cups down the other way, and you can only have saddle points out in the middle of the domain D.
Constraints. [359] We finally get to something I wanted to put in my tool kit in order to deal with a certain Boltzmann problem. That problem occurred somewhere and caused me to decide to read Buck!
Suppose you have w = f(x,y,z) now as a function of 3-variables. Earlier we dealt with z = f(x,y). We could ask in general where f(x,y,z) might have a maximum in some domain of 3D space. I am not able to visualize this situation since I would need to plot w on the vertical axis over the x,y,z "plane". BUT, if it happens that you have a constraint of the form g(x,y,z) = 0 acting in the problem, then you could maybe solve that for z = G(x,y) and then the original problem becomes w = f(x,y, G(x,y)) = F(x,y) and that is a problem we know already how to deal with! So this is a special case of the general constraint problem, but I am happy to do anything like this right now!
Example 1: [359] We are given a specific g(x,y,z) and solve for z as in A. Our function to minimize is a different function f(x,y,z) as stated in A also. Then F(x,y) is stated. We find that it has exactly one critical point over all space (the intersection of two lines in 2D). Using the Δ test we find that this critical point is in fact a minimum. Since this gives x,y and we get z from A, we have the solution (x,y,z) which minimizes the function F(x,y,z) subject to the constraint g(x,y,z) = 0, and this was the goal.
Obviously this is a very special case solution of a special case problem. Bucks plan to give us a better and more general way to do this. I don't know how general their presentation will be, I am hopeful.
Question: Where is the Buck discussion of rank reduction at singular points?
1. Well, first here is from Ch 5:
Theorem 6 [233] (Rank Theorem) restates the above for general f: En → Em, general rank r. Here r = m is full rank and the range of linear f is all of Em which one can think of as a plane of dimension m. For r < m, the range r is a plane of dimension r. These planes all contain the origin since 0→0 no matter what r is.
This says that if you have the maximum possible rank (full rank = r = min(n,m)) then in the range has dimensionality r. For example, for sphericals with n=m=3, you have r = 3 and then a 3D object in the domain maps into a 3D object in the range. The whole space maps that way with my usual picture showing the highrise building as the 3D object on one side and all space on the other side. This also means that a 2D object (surface) maps into a 2D object (surface), curve into curve, point into point.
However, if at some point in the transformation the rank drops down one rung, then r is smaller, and everything gets compressed: then 3D → 2D object, 2D surface → 1D curve, 1D curve → point and so on. This discussion in in Chapter 5, p 288 "dimension-reducing". and "functional dependence" concepts.
Question: Where is the discussion of systems of equations?
Chapter 5 p 292. There we have this example:
u = f(x,y,z) E3 → E2
v = g(x,y,z)
One interesting idea was to embed this into an E3→E3 transformation this way
u = f(x,y,z) E3 → E3
v = g(x,y,z)
w = 0
Then
dT = has rank 2 instead of rank 3
Because the rank is not full, 3D objects map into 2D objects, etc, as just discussed. This was Theorem 30 on page 292.
Question: [vague] In the above example, what happens if we consider f(x,y,z) = 0 and g(x,y,z) = 0 ? We are then looking for the inverse map of the point (u,v) = (0,0). In general since each equation describes a 2D surface, I would expect this inverse map to be the intersection 1D curve. Then in the forward mapping we have 1D curve → 0D point which is a dimension reduction of 1 unit. This corresponds to the fact that the rank of dT shown above is 2 instead of the full 3. So here we have a tidbit on solving a system of equations. Since we can think of these two equations as z = F(x,y) and z = G(x,y), we can visualize these as surfaces in 3D space and we can visualize how they intersect in a curve.
Question: In the above situation, suppose we want to find a max of f(x,y,z) subject to g(x,y,z) = 0 ? I am just gunning a few shots here before reading on. The max of f must be either at an interior critical point of f on the domain D, or on the boundary of D. We have already treated this problem by assuming we can solve to get z = G(x,y) then we have a max problem of F(x,y) = f(x,y,G(x,y)) using our previous methods. But I am hoping to see an approach that somehow uses the rank ideas above.
I now resume notes on the book at page 360 Theorem 15:
Theorem 15: [360] . Imagine that there are some point r such that g(x,y,z) = 0 [r on surface S] and f(x,y,z) is a local max or min at this same r which is not on the boundary of S. The theorem claims that r must be a "critical point" of the combined transformation
u = f(x,y,z) E3 → E2
v = g(x,y,z)
Now we have a pause. Recently (p 349) Bucks have talked about critical points of scalar function f(x,y) as being locations x,y where f1 = 0 and f2 = 0 and these turned out to be candidate extremum locations. But earlier (p 293, Ch 5) Bucks talked about the critical point of a transformation T as a place where the rank of dT was not the maximum. So far these notions of "critical point" existed in different discussions of different animals. I think we are about to find they are one in the same in some sense.
Theorem 15 is making a statement about min/max of f(x,y,z) with constraint g = 0 in terms of critical points of a certain dT, not in terms of critical points of a function.
Now for the above T we have
dT = // as on page 360
This has max rank 2, but if ALL three 2x2 sub dets were 0, then it would be rank 1 and if that happens at some point r, that point r is a critical point of this transformation! I have shown these two facts:
if any two |2x2| vanish, so does the third one.
the implication of the three vanishing or any two is that (f1,f2,f3) = α (g1,g2,g3)
Example 1 revisited: [ now 360] Write out the same f and the same g from p 359, compute the matrix dT and get the matrix shown, then require (f1,f2,f3) = α (g1,g2,g3) and you get the exact same solution point obtained earlier. So this is a verification of Theorem 15 for a particular example: the solution of the problem is associated with a point where rank = 1.
Two proofs of Theorem 15 are presented.
Proof 1: Assume g3 ≠ 0 and create F(x,y) as usual and as shown. It has "critical points" where F1 = 0 and F2= 0. We see here that these are also the "critical points" where dT has rank 1! So here the two ideas are connected finally.
Proof 2: Very interesting and you have to follow the steps very carefully.
1. p* is assumed to be a place where f has a max and lies on g = 0 (surface S)
2. Since v = g(x,y,z) = 0, we know that T(p*) lies somewhere on the v=0 axis in the range of the p 361 figure. Since u = f(x,y,z) and since p* is a max, we must have T(p*) be the right end of the interval. That is to say, the interval must have a right end and T(p*) is the right endpoint.
3. If T had rank 2, a 3D ball around p* in the domain would map into a 2D disk (well, a neighborhood of some shape with space around B on all sides) in the range. This follows from Theorem 30 on page 292 [ which was proven by adding w = 0 as a third line ]. That 2D disk in the range must contain T(p*), called point B in the figure.
4. Points on S in the domain within the ball would map onto the u axis within the disk. All the points on the disk's horizontal diameter segment must be hit since things are continuous etc etc.
5. But this is a contradiction, since it would imply a solution beyond the max. For a point to the right of B within the disk, there must be a point p** in the domain lying on S where f is larger than our assumed max.
6. Therefore we cannot have rank = 2 so we must have rank 1 (or less).
Bucks claim this is just the outline of a rigorous proof, but it seems very good to me.
Enter Lagrange [362]
Construct H:E4 → E as shown. All critical points of H must satisfy the 4 equations shown. But the first three imply (6-38) which tells us that these must also be critical points of the Theorem 15 problem which is to get extremum of f subject to g = 0. The 4th equation is in fact g = 0. The 4th parameter λ is called a Lagrange multiplier! Again: if a point is a critical point of the Lagrange thing, then it is also a critical point of the Thm 15 problem.
Generalize to two constraints g = 0 and h = 0 and want to max f. Construct H:E5 → E as shown. A point which is a critical point of H is also the solution to this problem
g = 0
h = 0
the three first equations shown in B
Let's now look at those three equations
f1(x,y,z)+ λg1(x,y,z) + μh1(x,y,z) = 0
f2(x,y,z)+ λg2(x,y,z) + μh4(x,y,z) = 0
f3(x,y,z)+ λg3(x,y,z) + μh5(x,y,z) = 0
Since we are handed g and h, we know all the functions of (x,y,z) here, all 6 of them. But how do you solve such a set of equations? Do you have to solve for general values of λ and μ? Or will λ and μ be eigenvalues? In the simpler previous problem we never really solved for λ.
Plan A. Write the above three equations as
MT v = 0
where
M = v =
If detM ≠ 0, we find only the trivial solution v = 0. So in order to get a non-trivial solution, we must have
det = 0
That requires that the matrix shown have rank < 3.
Hold that thought, and consider now this transformation T built from the functions f,g,h above:
u = f(x,y,z)
v = g(x,y,z)
w = h(x,y,z)
For this transformation we have
dT =
Places where det(dT) = 0 are transformation-critical-points of T.
I think I have just proved the following theorem:
Theorem PL1: If r is a critical point of the Lagrange H for the problem of max f with g = h = 0, then r is also a critical point of the transformation T shown below.
(1) Consider the problem of maximizing f with constraints g = 0 and h = 0.
(2) Suppose r is a critical point of the Lagrange Multiplier H(x,y,z,μ,λ) transformation shown p 362. We then know that these equations must be valid ( in addition to g = 0 and h = 0 which are the other 2 eq)
f1(x,y,z)+ λg1(x,y,z) + μh1(x,y,z) = 0
f2(x,y,z)+ λg2(x,y,z) + μh2(x,y,z) = 0
f3(x,y,z)+ λg3(x,y,z) + μh3(x,y,z) = 0
Write as f + λg + μh = 0 which says the rows of the R matrix are linearly dependent so detR < 3.
(3) In order for these equations to be have a non-trivial solution for λ and μ we must have
det(R) = det = 0
(4) The above matrix is dT for this T
u = f(x,y,z)
v = g(x,y,z)
w = h(x,y,z)
(5) Therefore dT must have rank < 3, which means solution r is a critical point of T.
(6) Therefore, if you compute dT and find an r such that | dT| = 0, that r is at least a candidate for a solution to the problem described in (1).
Example: [362] We are given
f(x,y,z) = z
g(x,y,z) = 2x+4y-5
h(x,y,z) = x2 - 2y + z2
We seek a solution which gives a max for f subject to the constraints g = 0 and h = 0. So compute
dT = =
det(dT) = -4 - 8x
A candidate solution must therefore have x = -1/2 since this causes det(dT) = 0 so rank dT < 3. Given this value of x, we can compute y and z from the equations g = 0 and h = 0, and we find a candidate solution point r = (-1/2, 3/2, 11/4) which provides a local max for f subject to constraints g = 0 and h = 0.
This brings Chapter 6 to a close!!! What a voyage it was! I am not sure what I will do next. 3PM 2/28/15
Note Added 5/16/15. So what DID I do next after finishing this chapter on that date? Obviously I did not dive into the Boltzmann motivating problem. Well, exactly on the date, the email from John Boyd arrived asking me about my bipolar doc he wanted to cite. I cleaned up Bipolar doc over the next few days and learned from him about ResearchGate. But then on 3/5 through 3/15 I did a Cod/Lee trip. Upon return, I did the MRL taxes through 3/21 including my own taxes. Then did the fstat and annual funds xls update. Then on 3/25 started adding papers to RG. Decided to get my two thesis papers online. Then had RG upload problems with one of those on 3/26. Then I started the 41 day saga of "adding equation numbers" to tensor doc, which turned into a 6 week major update of tensor doc. That job ended on May 6. I got my thesis Legendre thing finally uploaded, then did proofing on the two AFIB papers and uploaded them too. Uploaded other papers so total RG papers is now 21. Then on May 12 I started my complete review of Buck. That took 4 days which brings me to today! So that is how 2.5 months went by after I wondered what I would do next! And that is why all this Buck stuff feels so cold, that is a big time gap for my little CPU memory unit where other things are flying in all the time.
So now what comes next? I will finish my Ch 6 meta notes, then I will ponder the Lagrange Multiplier method more, maybe look at my PDF. I want to get that subject fully nailed down with any number of variables and constraint equations. I then have two more large Buck chapters to do and some appendices, then that long-deferred (most of a lifetime) task will be complete.
Conjectured More General Lagrange Multiplier deal.
(a) Statement of the Problem: ( M variables, S constraints)
u0 = f(x1, x2, .....xM) . f : RM → R
We are interested in finding where this function is a maximum. However, there are also a set of S constraint equations:
u1 ≡ g(1)(x1, x2, .....xM) = 0
u2 ≡ g(2)(x1, x2, .....xM) = 0
....
uS ≡ g(S)(x1, x2, .....xM) = 0 .
We want the maximum of f, but subject to all the constraints.
Now we want to think of the above set of equations as a transformation from x-space to x'-space = u-space. It is in general possible that we have S < M, S = M or S> M.
(b) What does the R matrix look like?
R = f1 f2 f3 ... fM
g(1)1 g(1)2 g(1)3 ... g(1)M
g(2)1 g(2)2 g(2)3 ... g(2)M
g(S)1 g(S)2 g(S)3 ... g(S)M
There are S+1 rows and M columns. I will assume that S+1 ≤ M so this matrix is at most square, but generally rows are longer than columns. If the matrix is not square, imagine filling in the last rows of the R matrix with all 0's to get a square matrix for R. This is like defining extra variables uS+1 etc whose equations are all like uS+1 = 0, and then you have a square transformation with these dummy extra variables added.
In any event, the full rank for this matrix is S+1 since it has S+1 non-zero rows.
(c) Now define the H function as follows, using S extra variables λi,
H = H(x1, x2, .....xM, λ1, λ2,... λs) ≡ f(x1, x2, .....xM) + Σs=1S λs g(s)(x1, x2, .....xM)
Then write out the full set of M+S so-called "H equations" ,
0 = Hi = ∂H/∂xi = fi + Σs=1S λsg(s)i i = 1,2....M
0 = HM+s = ∂H/∂λs = g(s)(x1, x2, .....xM) s = 1,2....S
So there are M+S of these H equations. Think about what these equations say:
(1) The entire set of the H equations is the "critical point" requirement for the function H. If you were interested in finding an extremum of H with all its variables, you would set all these derivatives to 0. But it is not clear why this corresponds with our problem of interest!
(1) The last S equations just say g(s)(x1, x2, .....xM) = 0, so these equations enforce the S constraints which are part of the problem.
(2) the First M equations state that the M first derivatives of H all vanish.
The last S just say that the S constraints are satisfied. The first M equations say that
f + Σs=1S λsg(s) = 0
This says that the rows of the R matrix are linearly dependent. If the number of rows and columns is the same, namely, S+1 = M, then this just says det(R) = 0.
, so it must have det(R) < S+1 and so we have a critical point for the matrix and presumably for the problem we are trying to solve. I have not shown this connection Theorem 15 in this more general case, will do that later.
(d) Our task then is to consider the set of M+S "H equations" which are in M+S unknowns. These may be messy equations, but in theory they are likely to have some solutions, and then those are the locations in space r where we test to see if we have a maximum or not.
(e) What happens if S+1 > M, so there are more rows than columns? Just go through each step above. There are still M + S variables and still M+S equations to solve, so maybe it is OK?
Counterexample: Suppose M =1 and S = 3. Then we want to maximize f(x) subject to constraints gs(x) = 0 for s = 1,2,3. But each constraint then implies a set of solution values for x. The constraint equations have to have one or more solutions in common or you are toast. This example is too extreme.
Counterexample: Suppose M =2 and S = 4. Then we want to maximize f(x,y) subject to constraints gs(x,y) = 0 for s = 1,2,3.4. Each constraint is really a function y = Gs(x) which is a curve in 2D space. So all four constraint curves would have to have a common point, not very likely. With only S = 2 constraints, you would at least have a chance, but not very interesting. The constraints give you a few discrete points and you pick the one that maximizes f(x,y). If there were only S = 1 constraint, then at least the constraint is a full curve and f(x,y) will vary on this curve and have some max value somewhere. This is then the case S+1 = M which is really the max case considered above.
Theorem 15 Boosted.
[ I bit off too much here, need to do some simpler example first, I cannot find the right thread. ]
Not sure how this will go, want to generalize their Proof #1. here is the setup:
u0 = f(x1, x2, .....xM) . f : RM → R
We are interested in finding where this function is a maximum. However, there are also a set of S constraint equations:
u1 ≡ g(1)(x1, x2, .....xM) = 0
u2 ≡ g(2)(x1, x2, .....xM) = 0
....
uS ≡ g(S)(x1, x2, .....xM) = 0 .
Now in our thought experiment, we use the constraints to eliminate S of the variables, make it the last S variables. Then that thought experiment solution gives these S equations for eliminating S variables,
xM = φM(x1.....xM-S) ≡ φ(M)(x1.....xM-S)
xM-1 = φM-1(x1.....xM-S) ≡ φ(M-1)(x1.....xM-S)
xM-2 = φM-2(x1.....xM-S)
...
xM-(S-1) = φM-(S-1)(x1.....xM-S) ≡ φ(M-S+1)(x1.....xM-S)
There are a total of S of these auxiliary functions φi.
Then we have
F(x1.....xM-S) ≡ f(x1.....xM-S, φ(M-S+1)(x1.....xM-S), .... φ(M-1)(x1.....xM-S), φ(M)(x1.....xM-S))
xM-S+1 xM-1 xM
At this point, we have a conventional problem and we are then interested in the critical points:
0 = F1 = f1 + fM-S+1 φ(M-S+1)1 + ...... + fM-1 φ(M-1)1 + fM φ(M)1
0 = F2 = f2 + fM-S+1 φ(M-S+1)2 + ...... + fM-1 φ(M-1)2 + fM φ(M)2
0 = F3 = f3 + fM-S+1 φ(M-S+1)3 + ...... + fM-1 φ(M-1)3 + fM φ(M)3
0 = FM-S = fM-S + fM-S+1 φ(M-S+1)M-S + ...... + fM-1 φ(M-1)M-S + fM φ(M)M-S
So there are M-S of these equations, one for each surviving coordinate.
Now what? Well, first, for each constraint equation we can do the same replacement, so we get,
0 = g(s)1 + g(s)M-S+1 φ(M-S+1)1+ ...... + g(s)M-1 φ(M-1)1 + g(s)M φ(M)1
0 = g(s)2 + g(s)M-S+1 φ(M-S+1)2+ ...... + g(s)M-1 φ(M-1)2 + g(s)M φ(M)2
0 = g(s)M-S + g(s)M-S+1 φ(M-S+1)M-S + ...... + g(s)M-1 φ(M-1)M-S + g(s)M φ(M)M-S
But we get one equation set like the above for s = 1,2,3....S !
Compared to the original Theorem 15 derivation, things are much tougher here. How many φ objects are there? Answer: There are S φ functions to start with, and for each we have M-S derivatives, so there are then S(M-S) φ derivatives floating around in the above equations. The total number of equations is (M-S) in each group and there are S+1 groups, so I see (S+1)(M-S) equations.
For example, in the original Theorem 15 we had S = 1 and M = 3, so there were 2 φ derivatives, and there were then 2*2 = 4 equations.
OK, lets regroup the equations like this: (by the derivative index). First group is then
0 = F1 = f1 + fM-S+1 φ(M-S+1)1 + ...... + fM-1 φ(M-1)1 + fM φ(M)1
0 = g(s)1 + g(s)M-S+1 φ(M-S+1)1+ ...... + g(s)M-1 φ(M-1)1 + g(s)M φ(M)1 s = 1...S
There are a total of S φ objects here, and S+1 equations. So we should then end up with must one equation which contains no φ objects. Write this set of equations in matrix form
-f1 = fM-S+1 φ(M-S+1)1 + ...... + fM-1 φ(M-1)1 + fM φ(M)1
- g(1)1 = g(1)M-S+1 φ(M-S+1)1 + ...... + g(1)M-1 φ(M-1)1 + g(1)M φ(M)1
- g(2)1 = g(2)M-S+1 φ(M-S+1)1 + ...... + g(2)M-1 φ(M-1)1 + g(2)M φ(M)1
- g(S)1 = g(S)M-S+1 φ(M-S+1)1 + ...... + g(S)M-1 φ(M-1)1 + g(S)M φ(M)1
Write this as a matrix equation
-f1 fM-S+1 fM-1 fM φ(M-S+1)1
- g(1)1 g(1)M-S+1 g(1)M-1 g(1)M
- g(2)1 = g(2)M-S+1 g(2)M-1 g(2)M
- g(S)1 φ(M)1
S+1 rows S+1 rows, S cols S rows
Stop on this and instead do another example.
Theorem 15 Example 2
Lets try M = 3 coordinates and S = 2 constraints. Then,
u = f(x,y,z)
v = g(x,y,z) g = 0 constraint #1
w = h(x,y,z) h = 0 constraint #2
What is the R matrix?
f1 f2 f3
g1 g2 g3
h1 h2 h3
Now want to use the two constraint equations to eliminate y and z. We somehow solve these two equations to get
y = Y(x)
z = Z(x)
Then we can write
U(x) = f(x,Y(x),Z(x))
V(x) = g(x,Y(x),Z(x))
W(x) = h(x,Y(x),Z(x))
Now we can write the critical points for each of these:
0 = U1 = f1 + f2Y1 + f3Z1
0 = V1 = g1 + g2Y1 + g3Z1
0 = W1 = h1 + h2Y1 + h3Z1
Write this as a matrix equation
= or 0 = RV
If detR ≠ 1, then we can solve to get V = R-10 = 0, but this is a contradiction because V1 = 1, not 0. Thus, if we want to solve our three equations above, we must have det(R) = 0, and that is really all I want to show in this example.
What happens if we look just at the last pair of equations, those associated with the constraints,
0 = V1 = g1 + g2Y1 + g3Z1
0 = W1 = h1 + h2Y1 + h3Z1
or
- g1 = g2Y1 + g3Z1
-h1 = h2Y1 + h3Z1
or
=
In order for us to eliminate Y1 and Z1 from our equations, we have to be able to invert this matrix. So we have here an added requirement that
det ≠ 0
Theorem 15 Example 3
Always trying to show that if a point solves your extremal problem, that same point drops rank.
Lets try M = 4 coordinates and S = 2 constraints. Then,
u = f(x,y,z,t)
v = g(x,y,z,t) g = 0 constraint #1
w = h(x,y,z,t) h = 0 constraint #2
What is the R matrix? (add a row of 0's to go with a fourth variable q = 0)
f1 f2 f3 f4
g1 g2 g3 g4
h1 h2 h3 h4 full rank = 3. (4,3) = 4!/3! = 4 possible 3x3 subdets
Now want to use the two constraint equations to eliminate z and t. We somehow solve these two equations to get
z = Z(x,y)
t = T(x,y)
Then we can write
U(x,y) = f(x,y, Z(x,y),T(x,y)) = F(x,y)
V(x,y) = g(x,y, Z(x,y),T(x,y)) = G(x,y)
W(x,y) = h(x,y, Z(x,y),T(x,y)) = H(x,y)
Now we can write the critical points for each of these:
0 = U1 = f1 + f3Z1 + f4T1
0 = U2 = f2 + f3Z2 + f4T2 critical point specification
0 = V1 = g1 + g3Z1 + g4T1
0 = V2 = g2 + g3Z2 + g4T2 g ≡ 0
0 = W1 = h1 + h3Z1 + h4T1
0 = W2 = h2 + h3Z2 + h4T2 h ≡ 0
Regroup in this order
0 = U1 = f1 + f3Z1 + f4T1
0 = V1 = g1 + g3Z1 + g4T1
0 = W1 = h1 + h3Z1 + h4T1
0 = U2 = f2 + f3Z2 + f4T2
0 = V2 = g2 + g3Z2 + g4T2
0 = W2 = h2 + h3Z2 + h4T2
and then write these two sets as two matrix equations
= or 0 = MV1
= or 0 = NV2
In order to avoid contradiction, we must have det(M) = 0 and det(N) = 0.
I have then shown that 2 of the 4 possible 3x3 subdeterminants of R vanish, but I know nothing about the other two.
Question: Does this force the other two subdeterminants to vanish? Let's try one of them:
det(M) = f1(g3h4- g4h3) - f3(g1h4- g4h1) + f4(g1h3- g3h1) = 0
det(N) = f2(g3h4- g4h3) - f3(g2h4- g4h2) + f4(g2h3- g3h2) = 0
det(S) = f1(g2h3- g3h2) - f2(g1h3- g3h1) + f3(g1h2- g2h1) = ?
Mult last line by f4 to get
f4det(S) = f1f4(g2h3- g3h2) - f2f4(g1h3- g3h1) + f4f3(g1h2- g2h1) = ?
Now replace from the M and N lines to get
f4det(S)
= f1{- f2(g3h4- g4h3) + f3(g2h4- g4h2) } - f2{ - f1(g3h4- g4h3)+ f3(g1h4- g4h1)} + f4f3(g1h2- g2h1)
= f1{ f3(g2h4- g4h2) } - f2{f3(g1h4- g4h1)} + f4f3(g1h2- g2h1)
= f3 [ f1 (g2h4- g4h2) - f2(g1h4- g4h1) + f4(g1h2- g2h1) ]
≠ 0
So in this case, just having two 0 does not force the third to be 0. It could be 0.
BUT, I could have started on this problem in may other ways. I chose to eliminate z and t and for this reason, both M and N matrices have f3 and f4 etc, and the f1 and f2 vary. If I had instead eliminated say x and y, then I claim I would get these two matrices instead
Q = and S =
and I would conclude that one must have detQ = 0 and detS = 0 in order to avoid a contradiction.
Thus, just looking at these two possible starting positions
1: eliminate z and t
2: eliminate x and y
I have shown that all four 3x3 subdeterminants must vanish, and therefore R must have rank < 3 in order to yield up a possible solution to my problem.
What about other requirements? Consider these equations from above
0 = V1 = g1 + g3Z1 + g4T1
0 = W1 = h1 + h3Z1 + h4T1
or
=
In order to eliminate Z1 and T1, you would have to have det ≠ 0. I am not sure this has much significance, just adding it.
Theorem 15 Example 4
u = f(x1,x2....xM)
v = g(x1,x2....xM) g = 0 constraint #1
w = h(x1,x2....xM) h = 0 constraint #2
What is the R matrix?
f1 f2 f3 f4 ... fM full rank = 3 (M,3) = possible 3x3 subs
g1 g2 g3 g4 ... gM
h1 h2 h3 h4 ... hM
Suppose I eliminate x1 and x2 with functions X(1) and X(2) of the remaining variables. Then the problem defining equations are
0 = f1X(1)3 + f2X(2)3 + f3
0 = f1X(1)4 + f2X(2)4 + f4
...
0 = f1X(1)M + f2X(2)M + fM
Two other sets of equations are obtained from f→ g and g→ h. Group as previously in triplets:
0 = f1X(1)3 + f2X(2)3 + f3
0 = g1X(1)3 + g2X(2)3 + g3
0 = h1X(1)3 + h2X(2)3 + h3 // this will force det(f1 f2 f3) = 0
0 = f1X(1)4 + f2X(2)4 + f4
0 = g1X(1)4 + g2X(2)4 + g4
0 = h1X(1)4 + h2X(2)4 + h4 // this will force det(f1 f2 f4) = 0
0 = f1X(1)M + f2X(2)M + fM
0 = g1X(1)M + g2X(2)M + gM
0 = h1X(1)M + h2X(2)M + hM // this will force det(f1 f2 fM) = 0
So making this choice will cause all dets of the form det(f1 f2 fm ) = 0 for all m ≠ 1,2
If instead I were to eliminate x1 and x3 , then I get all dets of the form det(f1 f3 fm) = 0.
It seems clear to me that by doing pairwise eliminations, you show that ALL possible 3x3 subdeterminants vanish. Thus, you show that you must have rank < 3 for original matrix R.
Theorem 15 Example 5
f(x1,x2....xM)
g(x1,x2....xM) g = 0 constraint #1
h(x1,x2....xM) h = 0 constraint #2
k(x1,x2....xM) k = 0 constraint #3
R matrix has this form
f1 f2 f3 f4 ... fM full rank = 4 (M,4) possible 4x4 subs
g1 g2 g3 g4 ... gM
h1 h2 h3 h4 ... hM
k1 k2 k3 k4 ... kM
If you choose to eliminate x1 and x2 and x3, then you will end up showing that the following 4x4 matrices must have zero determinant: det(f1 f2 f3 fm) . In general, if you eliminate xi and xj and xk, then you will show that det(fi fj fk fm) = 0 for all m. But this certainly exhausts all possible 4x4 subdeterminants and you reach the conclusion that: r solves the problem rank drops off max value of 4.
Theorem 15 Example 6
f(x1,x2....xM)
g(x1,x2....xM) g = 0 constraint #1
h(x1,x2....xM) h = 0 constraint #2
k(x1,x2....xM) k = 0 constraint #3
....
s(x1,x2....xM) s = 0 constraint #S
R matrix has this form
f1 f2 f3 f4 ... fM full rank = S+1 (M,S+1) possible (S+1)x(S+1)
g1 g2 g3 g4 ... gM
h1 h2 h3 h4 ... hM
k1 k2 k3 k4 ... kM
.......
s1 s2 s3 s4 ... sM
If I eliminate x1 through xS, then det(f1 f2....fS fm) = 0 for all m ≠1,2,...S
By making other choices of what variables to eliminate, you show that all possible (S+1)x(S+1) subdeterminants must vanish. The max feasible S would be one that makes S+1 = M, so S = M-1. Thurs, if you are in x,y,z space, you can have at most 2 constraints and Bucks do this problem on page 362.
OK, enough fiddling. I will write this up in separate doc, and that is then a wrap on Chapter 6!