Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Math / Wedge World / Sjamaar Forms / specific chapter notes

Sja Ch 6 notes

DOCX · 1007.9 KB
Open DOCX file

Phil's personal notes (dated 7.17.15, with additions from 8.13.15) working through Sjamaar Chapter 6, sections 6.1 (definition of a manifold) and 6.2 (regular value theorem). The visible portion covers parametrized surfaces x = ψ(t), the Jacobian columns as tangent vectors, the tangent space, and the three-part definition of an embedding. It also treats the graph-of-f example (ψ(t) = (t, f(t))) and the rank requirement on Dψ.

AI-written summary; may contain errors.

Extracted text (machine-read; may contain errors)
Sjamaar Chapter 6 notes PhL 7.17.15 6.1 The definition of a manifold 1 6.2 The regular value theorem 13 6.1 The definition of a manifold Questions. Look at the picture on page 68 middle. Why are the two vectors shown tangent to the surface? What is the normal to the surface at the point on the surface. I am used to such a surface appearing as g(x,y,t) = 0, then I can say things about it. But this is a parametrized surface which has x1 = ψ1(t1, t2) x2 = ψ2(t1, t2) or x = ψ(t) // was y = φ(x) in Ch 3 x3 = ψ3(t1, t2) . // was x = c(t) in Ch 4, and x = c(t) in Ch 5 ... xN = ψN(t1, t2) // is x = ψ(t) here in Ch 6 Well, consider a neighborhood of our point of interest. Then dx1 =(∂ψ1/∂t1) dt1 + ( ∂ψ1/∂t2)dt2 etc dx = (Dψ)(t) dt // matrix times vector dxi = (Dψ)ijdtj = ∂jψi dtj = (∂ψi/∂tj)dtj If you move just a very small dt, since the surface is planar locally, dx lies on the surface. So let dt = dti ei where ei is a unit vector in t-space. // no implied sum Then we get dxi = (Dψ)(t) [dti ei] . These vectors dxi for i = 1,2 are therefore tangent to the surface at t or x. One could define ξi ≡ (Dψ)(t)ei = some un-normalized tangent vectors i = 1,2...n . This then agrees with the arrows shown on the right side of page 68B. Question: Why are the dxi tangent to the surface? If you move a tiny amount in t-space, you move a tiny amount along the surface of the manifold in x-space, since after all the mapping is from t-space to that surface. If you move a tiny amount along the surface, you are moving tangent to the surface. What happens now if we write (Dψ) = (c1, c2.....cn) (Dψ)i1 = (c1)i = column 1 (Dψ)ij = (cj)i Then, [ξi]k = [ (Dψ) ei]k = (Dψ)kj (ei)j = (Dψ)kj δij = (Dψ)ki = (ci)k => ξi = ci Fact: the tangent vectors are the same as the columns of (Dψ). Thus, we really have ci = (Dψ)(t)ei where (Dψ) = (c1, c2.....cn) The tangent vectors form a basis locally at point x on the target space surface. This is the famous tangent space at point x. Since the basis vectors are the columns ci, this is also a column space. I would say this was an affine space unless you translate your basis vectors ξi = ci to the origin. Then you get a subspace. [ the ξi = ci are not generally orthogonal, but still span the tangent space ] *** Note added 8.13.15. (see "the meaning of (Df)...doc"). You want the tangent base vectors to be linearly independent so they form a full and useful (non-degenerate) space at the point x. According to the above, that means you want the columns ci to be linearly independent, since these are the tangent vectors. The (Dψ) matrix in the embedding scenario is in general non-square and taller than it is wide. Saying that the columns are linearly independent is the same as saying the (Dψ) matrix has full rank. As shown in the quoted doc, this means that the matrix (Dψ), although not square, nevertheless has a unique left inverse allowing the equation u = (Dψ)v to be solved for v. That in turn means that the matrix transformation u = (Dψ)v which maps Rn → RN is one-to-one. The "requirement" that (Dψ) be one-to-one appears as item (ii) in the definition 6.1 of an embedding on page 67. It was this line of logic that eluded me on my first reading. I guess you could form the normal in our example as n = c1 x c2 . Question: How do I interpret the expression on the RHS of page 67 B? (Need a decoder to read Sja) [(Dψ)(t)] = a matrix which acts on vectors in t-space which is some U within Rn. So I certainly could write (recall that U lies in Rn ) ψ: U → RN ψ(t) = x our transformation (Dψ): U → RN [(Dψ)(t)] t' = x' a different mapping In the first mapping, as t runs over all of U, the points x run over all of some surface of dimension n inside the space RN. For example, on page 68 U lies in R2 and the toroidal surface is a 2D surface and we happen to display in inside RN = R3 because that is all we know how to draw. In the second mapping, for the n special points t' = ei in U the mapping (Dψ) produces a set of n tangent vectors at the location x (associated with t in U) on the nD surface in RN. Recall that these tangent vectors are the columns ci of the matrix (Dψ). You can write a general t' in U as t' = Σit'iei and this would map into a general point x lying in the tangent space over in RN . (Dψ) t' = (Dψ)( Σit'iei ) = Σi t'i (Dψ)ei = Σi t'i ci = Σi t'i ξi = point in tangent space In this discussion, I have t being the reference point, and t' being the vector mapped by [(Dψ)(t)] and I would write that mapping as [(Dψ)(t)](t') . Alternate Notation: Let's pick some t0 and corresponding x0 as reference. There exists a tangent space for (t0, x0) and that space has tangent vectors ci(0) = (Dψ)(t0)ei . Everything is specific to t0. Now given this sort of fixed reference point, we can consider t = Σitiei and watch it map into (Dψ) t = Σi ti ξi . Thus, any point in U maps into some point in the tangent space associated with (t0, x0). The point t = 0 then maps into the point (Dψ)0 = 0 in the tangent space, so: "the origin of t-space which is U maps into the origin of the tangent space TxM " . So there are then two different values of the variable t to think about here. [ correct ]. In more detail you might write [(Dψ)(t0)] t = Σi ti ξi for any t = Σitiei This is true because we showed above that [(Dψ)(t0)] ei = ξi = ci = basis of Tx0M. [ correct] If we now let t exhaust all points in U within the t-space Rn, we then would exhaust some portion of the infinite tangent space based on t0 and x0. So then [(Dψ)(t0)] U = a piece of Tx0M. // not a piece of M, but a "plane" tangent to M Then we could say: [(Dψ)(t0)] : U → Tx0M where range is part of Tx0M But since U lies in Rn you could regard this as saying [(Dψ)(t0)] : Rn → Tx0M In fact, if you were to extend U to be all of the t-space Rn, then you would exhaust the tangent space. Of course then you might not have a real manifold due to self-intersection, but ignore that. In this situation, the above line could be written [(Dψ)(t0)] (Rn) = Tx0M Both sides of this equation are sets which include all points of the tangent space associated with (t0, x0). Notice that our test variable t does not appear in this equation! Suppose we change the test point in U to be called t' and we let (t, x) be the reference point. Then we would rewrite the above equation as [(Dψ)(t)] (Rn) = TxM // as t' exhausts Rn, we exhaust the tangent space. (*) THIS must be what Sjamaar means by page 67 B!! You could also write this as [(Dψ)(t)] : Rn → TxM where this would be a 1-to-1 mapping which exhaust Rn on the left and exhausts TxM on the right. Now I have jumped ahead a bit. The surface M is not yet a manifold but we shall still call it M. We do know that we generate this surface by doing M = ψ(U). Thus we could rewrite (*) above as [(Dψ)(t)] (Rn) = Txψ(U) // as t' exhausts Rn, we exhaust the tangent space. (*)' This then matches page 67 B. This tangent space and its basis vectors are different as you change t0 and therefore you change x. F Note added 8.13.15. This notation (Txψ) ≡ (Dψ) simply stresses that the matrix (Dψ) when applied to the basis vectors in U (t-space) generates "tangent vectors" in the range space at point x. So the letter T is for Tangent (space or vectors), and the letter x emphasizes that the tangent vectors generated by (Dψ) are only really tangent vectors at the point x on the range space surface. So (Txψ) ≡ (Dψ) is just a notation. RN is called the target space, it holds your nD surface. [ page 68 has n = 2 and N = 3 ] Finally we come to Sjamaar's definition of an embedding on page 67, in terms of the picture page 68: Embedding Definition of Sjamaar The set U is embedded into RN (a larger space) to create an (smooth) embedding of U into RN . This smooth embedding scenario has three requirements. (i) The embedded surface is defined by x = ψ(t).The mapping x = ψ(t) must be C∞ so the mapping is smooth. If you want this surface to be non-self-intersecting, then you must insist that the mapping x = ψ(t) [ which is ψ:Rn → RN] be one-to-one. If you allow self-intersections [ which would be the torus case for a larger U ] then you are creating an immersion, which is a larger concept than an embedding, as I learned from Darling, see "the meaning...doc". Darling also talks about submersion, but I ignore that extra stuff here. (ii) The mapping (Dψ) must be one-to-one. As noted above ***, this means that for any point t in U, at the point x in RN on your embedded surface you will have a full linearly independent set of tangent vectors. The mapping itself would be written (Dψ):Rn → RN and has the same domain U and range space RN as the mapping ψ. (iii) Since x = ψ(t) is C∞, we know that the mapping ψ is continuous in the forward direction, so that a converging sequence of points in U would map into a converging sequence of points on the embedded surface in RN. In item (iii) we require that the inverse mapping ψ-1 must also be continuous (though not C∞). As Sja points out, if you were to consider a sequence of points on the embedded surface which approach some limit point on that surface, the reverse-mapped sequence of ti values in U would approach some limiting t value in U. This condition just says things are smooth. So with these three parts of the definition, an embedded surface has "nice" properties: the mapping ψ is continuous in both directions, the tangent vectors at any x are non-degenerate and span a full subspace Rn within RN, and the surface is not-self intersecting. To make this all work, (Dψ) must have full rank, meaning its column vectors (column space) are linearly independent. This Dψ is my R matrix. Example 6.2, The domain U is a 2D rectangle, and the embedded surface is a piece of a torus. At a given point on the torus there are two tangent base vectors. Then are correctly shown as (Dψ)e1 and (Dψ)e2 which are of course not unit vectors in general. Example 6.3 Consider a smooth (C∞) mapping f: U → Rm where U is the same U in Rn discussed above, and now we have a new dimension space Rm in the picture. In terms of this mapping f, we now define the embedding's function ψ in this manner: x = ψ(t) = (t, f(t) ). Therefore we have N = n+m in this example. In general, something like g(x) = (x,f(x)) is called "the graph of f". So for this particular ψ, we can view the matrix (Dψ) in this manner This picture is the meaning of page 68 E where Sja uses a compressed vector-like notation. The inverse function (assuming it exists) is defined by t = ψ-1(x) = ψ-1(t, f(t)) // since x = ψ(t) = (t, f(t) ) This appears as page 69 A as the equation ψ-1(t, f(t)) = t . Sja is making no claims about how you might write down this inverse function given f(t), he just says it has this property. Now is this example really an embedding? (i) the mapping x = ψ(t) = (t, f(t) ) is smooth given that f(t) is smooth. We can see that t1 = t2 implies that x1 = x2. In the reverse direction, suppose x2 = x1 . that would require t1 = t2 from the first argument. Therefore the mapping ψ is one-to-one as well as C∞ so requirement (i) is met (ii) Regardless of (Df), the picture above shows that the columns of (Dψ) are linearly independent. This is exactly my Mini-Columns Theorem (A.2) in Lagrange doc. Thus, (Dψ) is one to one, so (ii) is met. (iii) Now we have to show that t = ψ-1(x) is continuous. Consider a sequence xn → x . Looking at the forward map x = ψ(t) = (t, f(t) ), we can write xn = (tn, f(tn) ) . Since the forward map is continuous, we know that tn→ t implies xn → x for any tn sequence. If we have a convergence sequence xn → x, then since xn = (tn, f(tn) ), we must also have a convergent sequence t → tn. Thus ψ-1 is also continuous. So property iii is met. Therefore this example really is an embedding. Spivak Digression: His definition of diffeomorphism and two manifold definitions Before continuing to Sjamaar's definition of a manifold, we digress to another source (just because that is what I did on my first reading attempt). Spivak here describes a certain mapping between two regions within space Rn which mapping I would say was "exceedingly well behaved" and "exceedingly smooth" in every possible sense. Spivak. He opens his Chapter 5 on manifolds with this definition: Here is a picture to go with the above definition Here u and v are vectors in Rn and so I bold them. Now we really have v = h(u) so differentiable means in the mult-variable sense, so one is claiming I think that (Dh) exists at any point in U and V, and I guess so does (D2h) and so on, whatever that means (well, ∂j∂ih exists, etc). The fact that an inverse exists says that this little mapping between U and V must be one-to-one. So at least Spivak defines the word diffeomorphism for his poor sucker reader. OK, more words: when you consider v = h(u), you can compute (Dh)(u) and all higher vector derivatives (like ∂i∂jhk) and all these derivatives are well defined and smooth. Moreover, the function h-1 exists AND you have u = h-1(v) and in this direction (Dh-1)(v) exists and is computable and is smooth. The point is that you can differentiate C∞ in either direction on any variable. So I guess you would say this was a very smooth mapping, and since h-1 exists, it is 1-to-1. So in Sja's language we have x = ψ(t) replacing v = h(u) and we have not only ψ being one to one and having an inverse, but (Dψ) exists and (Dψ-1) exists. In the embedding definition, we only required that ψ-1 be continuous, but in this diffeomorphism in fact ψ-1 is C∞, overkill. Note that in Spivak's diffeomorphism, both spaces are on an equal footing. So you could refer to either h or h-1 as a diffeomorphism. Comment added 8.12.15. In regular 1D calculus, suppose you have y = f(x). We know (from Buck p 270) that the inverse is obtained by reflecting in the x=y plane and can result in several inverse branch functions as on page 270 or 271. If f(x) is one-to-one, that would be like the right side only of the parabola on page 270 y = f(x) = x2 . For just that piece y=f(x) and x = f-1(y) ≡ g(y) would exist, f(x) is invertible, the function is 1-to-1 and all is well. The function y = F(x) ≡ f'(x) = 2x is also 1-to-1. We might refer to this as F(x) = (Df)(x). So for a general smooth situation, one would have both f(x) and (Df)(x) being 1-to-1. The question is how you extend this smoothness concept to nD calculus, and that is what Spivak is discussing. You then want y = f(x) to be 1-to-1 so that x = f-1(y) exists. And you also want that ∂yi/∂xj ≡ ∂jfi = (Df)ij(x) be one-to-one. In the 1D case, there is only one derivative, so both f(x) and f'(x) are scalars. In the nD case, f(x) is a vector and (Df)ij is a rank-2 tensor, so things are much more complicated. You can differentiate each vector component in each of many directions. We then want to have "balls" around the points x and y so we can examine derivatives in all directions. This is what Spivak is doing. The diffeomorphism defined above, the vector function h is assumed to have an inverse, so v = h(u) is 1-to-1. And it is required that both v = h(u) and u = h-1(v) be differentiable. But it is not required (Dh)(u) be 1-to-1 . Differentiable means all derivatives C∞. Comment: A homeomorphism is like a diffeomorphism but has a lesser requirement: h and h-1 have to exist and be just continuous in both directions, not necessarily differentiable. So the Sjamaar embedding conditions require something halfway between a diffeo and a homeo morphism. Another spoonful of Spivak caster oil which is his manifold definition: Here is Spivak's corresponding picture where the manifold M is the toroid donut surface. For his picture, the manifold M is a 2D surface, so k = 2. The dimension n is called N by Sjamaar, it is the dimension of the space in which you are doing the embedding, and in the above picture as well as Sjamaar page 68 this dimension is 3. The dimension k is called n by Sjamaar. In Spivak's picture, both the domain and the range of the mapping are in the larger space of dimension n (Sja N). The dark grey square on the right would be Sja's rectangle on the left of p 68 B. So, the idea is that you can talk about a patch on the manifold as being UM where U provides "elbow room in all directions" for any point on M. Then of course h(UM) is the dark rectangle on the right. Of course V = h(U) is the light graph box on the right, and this provides "elbow room in all directions" for any point on the right in the dark square. Now, on the right, the gray rectangle of course has z = 0 (it lies on this plane on the right). For N > 3, all such coordinates would be zero, and the mapping is just that gray rectangle or its generalization to k dimensions. The idea that z = 0 is what Spivak means when he writes Rk x {0). The mapped grey square is the Rk part, and all the z = 0 coordinates are the z = 0 part. Now the grey plane on the right can be large, but we only care about that part of it which corresponds to (is mapped from) the patch of interest on M, and that is then V [Rk x {0)] . So now all of Spivak's notation makes sense, that was the hard part. What he is saying is this: if you have a candidate manifold surface M, to test whether it really is a manifold, you pick a point on M (Sjamaar's x) and you select your little U and V balls (they can be as small as necessary) and you check to see whether the mapping h (which is Sja ψ-1) is a diffeomorphism. So Spivak is requiring that for any tiny patch on M, the mapping ψ must be exceedingly smooth and well behaved, even C∞ in both directions! Now that we have seen Spivak's manifold definition, let's look at Sjamaar's on page 60. There are several differences: (1) Whereas Spivak's ball around the toroid surface point is called U, Sjamaar calls it V, so there is an immediate U↔V change going on. So Spivak had U in Rn around the toroid surface point, while Sjamaar has V in RN around that same point. (2) Spivak puts a ball V around the parameter space rectangle where the higher coordinates are all zero. Sjamaar does not bother with those extra dimensions and he calls this thing U. So Sjamaar Spivak V dim N U dim n ball around the toroid point in RN U dim n V [Rk x {0)] dim k the parameter space VM dim n UM dim k the toroid patch ψ h-1 mapping from param space to toroid (3) Spivak requires that his mapping h-1 be a full blown diffeomorphism. Sjamaar only requires that this mapping be an embedding, which is not as supersmooth as the diffeomorphism in the inverse direction. Spivak has the inverse mapping being C∞, whereas Sjamaar has it being C1 only. Spivak goes on to give an alternate manifold definition which looks closer to the Sja one: Here is a translation table for this alternative definition Sjamaar Spivak V dim N U dim n ball around the toroid point in RN U dim n W dim k the parameter space VM dim n UM dim k the toroid patch ψ f mapping from param space to toroid Here is the Sjamaar manifold definition, Sjamaar's last bullet says that ψ(U) = VM and that ψ is an embedding. In Spivak, since the word embedding is not used, Spivak has to write out the embedding requirements. Spivak (1) is part of the last bullet of Sjamaar. Spivak (2) uses notation f'(y) to represent Sjamaar's (Dψ) matrix. Spivak requires f'(y) to have full rank k, which is like Sjamaar having (Dψ) be full rank n, so it is one-to-one. Spivak (3) is the embedding requirement that the inverse function as well as the forward function be continuous. The notion of a chart and atlas (Sjamaar page 67, so I have gone backwards) The inverse map ψ-1 for a patch of some size is called a chart or coordinate map. This is the map from the fancy torus patch to its parameter space rectangle. This is like the mapping ψ-1 from a piece of the surface of the Earth to a chart (map) of that piece of the earth. PL Example 1 of a manifold, charts and an atlas. The manifold here is a circle, and in this example wiki considers four different patches which in fact overlap (patch = piece of circle), and there are then four charts which make up one atlas of the circle. The dimensions here are: manifold is a 1D surface (n = 1), parameter space is 1D (n = 1) which they call x for two of the charts and y for the other two charts. The circle is embedded in R2 so N = 2. In this example, one has 4 charts to cover the circle, so the atlas has four charts. The wiki says that almost always you need more than one chart to cover a manifold. It also says that you could do the circle with two charts like this: where the line segment at the top is the E1 space to which we are mapping the red part of the circle. The mapping has to be one to one! If you try to cover the circle with just one chart by extending the red piece and shrinking away the blue piece, you end up with a mapping from circle to line segment which is NOT one to one. The two ends of the line segment map to the same (bottom) point on the circle. So this same idea would apply to any closed curve including the one Sja shows page 69D. So it the problem has noting to do with tangent spaces or places where the curve has the same slope. To summarize: a simple closed curve is a manifold, but you cannot cover it with one chart. You need two or more charts. This of course all relates to mapping the globe that is the earth. You try to do it this way But now the left and right edge map to the same great semicircle on the globe, so the mapping is not one to one! The right thing to do is have a whole book of charts, where each chart covers some local section of the earth, and that book of local maps is indeed called an atlas. Review of things prior to page 70 1. Cover the entire M with patches small enough so the above works for each patch. It is OK for the patches to overlap, and probably desirable that they do so. Then M = P1 P2 ... Pn P = Patch (= chart) ψ1: U1→RN is an embedding making use of V1 ψ2: U2→RN is an embedding making use of V2 ..... If you can do this, then M is a manifold. 2. Fact: if the surface M self-intersects even just at one point x, then for any patch containing that point x, you cannot create an embedding because ψ(t) won't be one-to-one since two points on the surface map back to same t. Thus, this one point excludes surface M from being a manifold. 3. The manifold has dimension n within RN, so the codimension of M is N-n. For the torus example, that codimension is 1. 4. For a given one of the patch embeddings ψ(i) that includes x on M, Sja now wants to say that the notation for the tangent space associated with that x is written this way: TxM = [(Dψ(i))(t)] U(i) ≈ the tangent space at x ≈ Rn but Sja still uses that strange notation I don't understand, see p 69 C. [ but now I do understand it ] 5. He says now that the collection of all the embeddings {ψ(i)} which cover M is what you should call an atlas. But I like the idea of thinking of the U(i) as being the charts for each embedding, and the atlas is the set of all those charts. A chart is the inverse map from the M patch to the U domain, and its image is that U domain. So OK, semantics here. He then repeats the Buck idea that the words curve, surface and hypersurface imply that these things are manifolds. To emphasize this fact, you can add the word "smooth" such as "smooth surface". Sja is really hammering hard on this tangent space idea at every opportunity, but we don't yet know why he is doing that. He is setting us up for something! The rest of this Section 6.1 is treated one page at a time: Page 70 Example 6.5. What happens if x = ψ(t) = t ? Here ψ is the identity map. Then ψ: U → U and surely this is a valid embedding, and only one embedding is needed to cover all of U (all of U is "the patch") and then this U within Rn is a manifold. (Dψ) = In so the tangent space is spanned by the same ei that span U, so tangent space = Rn. Conclusion: Any open region U in Rn is in fact a manifold of dimension n and codimension 0. Example 6.6. What happens if, for ψ:Rn → RN we have x = ψ(t) = (t1,t2, ...tn, 0,0,...0)? Is this an embedding? This seems a slight generalization of 6.5 above. The higher dimensions of RN do nothing, and we basically have ψ: U → U happening within the context of this larger RN. So yes, this seems to be an embedding. This situation is "an n-manifold within RN". This reminds me of the original Spivak manifold definition stated above. Example 6.7. Here we are back to the (t,f(t)) graph thing, now written (x,f(x)). The second example would then have ψ = (x,y,f(x,y)) as a column vector. He computes the (Dψ) matrix and I think he has the wrong results displayed. [agreed on 8/13/15] The graphs on the next page just show two views of the same 3D graph, and he is displaying tangent vectors at a few points. I wonder if these are wrong too? How do you display a simple arrow in Maple? I learned how and here is my version of his graph where I just picked one point on the surface to show tangent vectors. I used my formulas, and my arrows are in fact tangent. Page 71 Example 6.8. Here ψ(t) = et(cost,sint) so we have a spiral thing for ψ:R→R2. As t → -∞, it spirals around the origin an infinite number of times which seems reasonable to me. It never gets there. He shows that this ψ(t) does define an embedding (the surface is the spiral curve in E2 ). He claims this curve is a manifold as long as you stay away from the origin by using large but finite negative t endpoint. Page 72 [72] Example 6.9. I already saw that a circle needs two charts or more. Here the unit sphere Sn-1 is considered, where I think n-1 is the dimensionality of the sphere's surface, so a circle is S1 and a sphere is S2. For the general case Sn-1 we have t in U in Rn-1 . He writes a mapping into Rn (where sphere lives), This function somehow has an image which is the sphere surface, details in Exercise B.7 which I have not done. His mapping is the one used for a stereographic projection of the earth. He does a north pole one and then a south pole one, and so you need two charts to cover the sphere, in any number of dimensions. Because you can cover the sphere in these two embeddings, it is a manifold by our piecing-together definition. And so finally I have finished this very painful section 6.1 6.2 The regular value theorem Warning: In the previous section 6.1 we dealt with x = ψ(t) and matrix (Dψ) had columns ci. In Section 3.2 we had y = φ(x) in our initial pullback discussion with pullback operator φ* and in the pullback structure we encountered the matrix object (Dφ) In Section 4.1 we had c* being this pullback operator for 1-forms only. So we have had the symbols φ and c floating around in various ways. Now below we have a new c and a new φ(x) which we take just a-priori, and there is of course a (Dφ) matrix going with this φ. [72] Suppose c = φ(x) is m equations in N unknowns. The points x here lie in U within RN. Warning: In Section 6.1 we had x = ψ(t) and ψ : Rn → RN with U in Rn. Here we have c = φ(x) with φ: RN → Rm with U in RN, so beware that the letters are all shuffled around. We could compute the following differentials: dci = (Dφ)ij(x)dxj = ∂jφi dxj = (∂φi/∂xj)dxj (Dφ)ij has i = m rows and j = N columns. Sjamaar does not say anything yet about the relative sizes of the integers N and n. To be clear, the vectors φ and c have m components, while the vector x has N components. ok to here Suppose that, for a given fixed vector c, the equation c = φ(x) has a solution locus for x such that for all x in this locus, (Dφ) has full row rank m. I know that this means (Dφ) also has column rank m and as a whole this non-square matrix has rank m. This matter all cleared up in Lagrange Appendix A. This would then suggest that N ≥ m in Section 6.2, On page 73 top we see mention of N-m as a space dimension, so again Sja is intending that N ≥ m. This is the scenario of a wide R matrix, as in Lagrange doc. Note added 2.14.16. You can of course compute [(Dφ)(x)] ignoring c and you get some matrix. You can evaluate this matrix at any x in RN that you want. However, c = φ(x) defines some region within RN which is our candidate manifold M. The idea is that this region will be a manifold if [(Dφ)(x)] is full rank for all points x on the surface M. Now Sja's notation for "the solution locus of x of c = φ(x) " is the slightly strange φ-1(x). So if for all x in φ-1(x) the matrix (Dφ)(x) has full rank (which is rank m), then c is called a "regular value of φ". And if there is somewhere in that locus where that rank falls below m, then c is a "singular value". Recall that Buck refer to a point where f(x) has all partials 0 is called a "critical point", so that is a slightly different idea. However, Bucks also use the term "critical point" (p293) for a place where a transformation matrix R drops below full rank. so the Buck's "critical point" seems to be the same as Sjamaar's "singular value" concept. Aside: Go back to the solution set φ-1(x) for a given c. For every c in its space Rm we are generating some whole solution surface for x in its space RN for each value of c. You can think of these spaces as forming a set of level surfaces within RN (such as a set of concentric spheres). Sja and others like to call these level sets. The idea that you have a whole surface for each value of c is related to the idea of a fiber bundle. Think of a hairbrush with handle a cylinder and radial tongs. For each point on the cylinder (for each value of c in one space) you have a whole other space (here a linear "fiber" which is the brush tong). So in this example you have a "bundle of fibers" making up the brush. Comment: This seems similar to the idea of a Tangent Space Txψ in Section 6.1 where for each value of x, you have a "whole space" which is the tangent space at that point. So maybe the tangent space concept is also a fiber bundle situation? But in this case, we don't have a surface for each x, we have a space with axes and so on, so maybe this is not a fiber situation. So here are the claims of the theorem 6.10. [73] Claims of Theorem 6.10. Consider φ: U → Rm as appropriate for the c = φ(x). This is then φ: RN → Rm with U an open set in RN. In Section 6.1, we had U having the smaller dimension, but now it has the larger dimension (so the R matrix is wide as in Lagrange doc, not tall). Think of the candidate manifold as M = φ-1(x) = the solution set {x} for c = φ(x), as defined above. Clearly this M is something embedded in RN since the x are in RN. If c is a regular point, then the claims of this theorem are: (1) M is a manifold in RN of dimension N-m and thus codimension m. (2) Furthermore, the tangent space at point x is the nullspace of (Dφ) . In English, this says: for surfaces implicitly defined by a function, there is a simple test to see if that surface is really a manifold or not. Write the function in the form c = φ(x) and check to see if c is a regular point by seeing if the mxN R matrix (Dφ) has full rank m, which is to say, there are m linearly independent rows. If yes, then the surface is a manifold and the tangent space at some x is the nullspace of the matrix (Dφ). The manifold has dimensionality n and thus the tangent space at any point on the manifold has dimension n. The codimension is N-n = m. Comment: Since we have m equations in N unknowns c = φ(x) defining our manifold, we certainly need to have m ≤ N in order to be defining a solution space in the first place, so again, N ≥ m. We want to show that M = φ-1(x) is a manifold. From page 69, our first task is to put a ball V around a point x on the manifold M. Then we put another ball U in our parameter space. We need to show that the mapping ψ: U → RM is an embedding and that ψ(U) = V M. The complicated proof of this Theorem has something like 24 little steps (each got a red check). I will outline the first few steps: Proof of Claim 1 1. Think of x = (u,v) where dim(x) = N, dim(u) = n, dim(v) = m, and so N = n+m. We are in effect saying to think of RN as being Rnx Rm. 2. Our reference point of interest on the candidate manifold is x0 = (u0,v0) 3. The matrix (Dφ) which is "wide" certainly has a non-zero minor (mxm) , since the m rows are linearly independent. I showed this in Lagrange doc App B. This minor is an invertible m x m matrix. This is a square piece of the R matrix (the Dφ matrix). I proceed down the proof about 1 1/2 inches and we see that Sja is constructing exactly the scenario which is studied in Appendix B.4 and for which I have detailed notes. The idea is that you think of having x = (u,v) where u has n components, and you then have φ(u,v) = c (here c and φ and v have m components, while v has n components). You wonder if there is an "implicit function" v = f(u) which you could write out which solves this equation. If there were an implicit function v = f(u) then you might want to know its R matrix which is (Df)(u). The Implicit Function Theorem says that this matrix is given by: (Df)(u) = – (Dvφ)-1(u,f(u)) (Duφ)(u,f(u)) // which is p 132 B where (Dvφ) and (Duφ) are partial Jacobians of dimension mxm and mxn (both have m rows). Here the object (u,f(u)) is simply an argument, exactly the same as (u,v) = x. Since the (Dvφ) Jacobian is square, you normally could invert it as shown above. The idea is that if this thing is invertible at some (u0,v0), then it is going to be invertible for (u,v) nearby. In our proof the (Dvφ) matrix is called A, and it is invertible because it has full rank and my Lagrange Theorem 8 (Cramer) says that fact makes A invertible. The Implicit Function Theorem B.4 then says that we indeed then have an implicit function v = f(u) and φ(u,f(u)) = c . Note that u is in U and v is in V, so we then have f: U → V within f: Rn → Rm. [ this is a different U] One can regard φ: x in RN → c in Rm as φ: RnxRm → Rm or specifically φ: UxV→ Rm (the space of c). In our application, we in fact have x = (u,v) and x0 = (u0,v0) being points in RN. As u runs over all of U, v runs over all of V according to our implicit function v = f(u). Confusion: In the theorem premise, Sja says U in RN so u should be in RN, but here we are saying u in Rm which is a subspace of RN. Are these different U's ? I think really it is two different U's ! The second u has n coordinates only. So why do we care about the fact that an implicit function v = f(u) exists in the above context? Hopefully that will be answered in the following barrage of notes. Now note and just hold this idea that this column vector we associate with "the graph of f": (u,v) = (u,f(u)) is, according to page 68, the "graph of f". n m Next: what is the meaning of M (U x V) ? The manifold M is all (u,v) pairs of such points x = (u,v) where v = f(u) [ which means φ(u,v) = c]. The space U x V inside RN is bigger than M because UxV contains (u,v) pairs where v ≠ f(u). We know then that M UxV and therefore I think that M (U x V) = M except maybe for the boundary of M. Why does Sja state that " M (U x V) is the graph of f " ? Well as shown above graph of f is the set of points in RN of the form (u,f(u)) and this is basically M, again apart from boundary issues. Waypoint: At this point in the proof, we have shown that M is the graph of f for this implicit function f which we showed must exist. But earlier we showed that the graph of f is always an embedding, so that is the main thrust of this proof. Now, Exercise 6.3 showed that the graph of f [which is ψ(u) = (u,f(u)) ] defines a valid embedding ψ of U into RN and I guess this means that M is a manifold with a single chart since it is the image of this embedding x = ψ(u) = (u,f(u)). What is the dimensionality of this manifold M? It exists in RN, yes we know that. But when we write the x values as x = (u,f(u)), when we run u over U, we are running over Rn so I would say that the surface M has dimensionality n, and thus codimension = m as claimed. Proof of Claim 2 What is the tangent space here? I agree that one can write φ(ψ(u)) = c since ψ(u) = (u,f(u)) = x. I agree that ψ(u) is our embedding of U in RN. The tangent space is spanned by the column vectors of (Dψ), bottom p 67. Comment on relation between φ and ψ: We have these two facts: x = ψ(t) in earlier notes x = ψ(u) in current context c = φ(x) in current context Combining these we can say, c = φ(ψ(u)) This gives the relation between functions φ and ψ. The relation is surely complicated, but doing a derivative with the chain rule is how we learn more about it. Fact: c = φ(ψ(u)) ≡ φ'(u) [ some new functional form] Dc = 0 = (Dφ')(u) . Fact: The chain rule says [ note that φ' does NOT mean a derivative, it is just a new function name ] (Dφ')(u) = (Dφ)(ψ) (Dψ)(u) . Therefore combining these last two facts: (Dφ)(ψ(u)) (Dψ)(u) = 0 . // the product of two matrices. this is page 73 B Replace ψ(u) = x to get (Dφ)(x) (Dψ)(u) = 0 . (**) If we evaluate this at u = u0, I think you should get (Dφ)(x0) (Dψ)(u0) = 0 close to p 73 C As I showed earlier, the tangent vectors are the columns of (Dψ)(u0) at point x0 on M. I had in fact that ξi ≡ (Dψ)(u)ei = column i. So I agree one can write tangent vector v to M at x0 has the form v = (Dψ)a where a is in U so is in Rn Now write (Dφ) v = (Dφ) (Dψ)a . But in (**) I just showed that (Dφ)(x) (Dψ)(u) = 0 for any u in U, so it must be true for u = a . Thus (Dφ) v = (Dφ) (Dψ)a = 0 . Since then (Dφ) v = 0, and since v is any tangent space vector, we know that these vectors lie within the nullspace of the matrix (Dφ) . Basically then TxM = tangent space of M at point x nullspace of (Dφ) which is a matrix of type N x m But we want to show that TxM = nullspace, not just TxM nullspace. If we can show that TxM and nullspace have the same dimension, then we have shown that TxM = nullspace . What is the dimensionality of this nullspace of (Dφ) ? I show here that it is n. Sylvester for non square matrices says: for k x l matrix A, nullity(A) + rank(A) = l for m x N matrix (Dφ), nullity(Dφ) + rank(Dφ) = N The tangent space dimensionality is the number of columns of (Dφ) which has n columns and N rows. Thus dim(tangent space) = n. The rank of (Dφ) is m because c is a regular point. Thus, nullity(Dφ) + m = N Therefore nullity(Dφ) = n and so tangent space of M at point x = nullspace of (Dφ) = space of dimension n This concludes the proof of Claim 2. [74] The rest of Chapter 6 provides examples of the regular value theorem idea. The first several examples relate to surfaces defined by a single equation φ(x) = c (rather than m equations). As noted earlier, an equation like this defines level surfaces in Rn , the space of x. We are used to such level curves when c = φ(x,y) and you imagine φ is height or temperature (topo lines or isotherms). How does the above regular value theorem apply to a single function like this? Well, Dφ is just a row vector whose transpose is grad φ. Only if all elements of the row vector vanish does Dφ drop below its maximum rank of 1, and if all entries vanish (all gradient components vanish) then you have a singular point for c in φ(x) = c. This is then the Buck idea that a "critical point" is a place where all first partials vanish, so we have made now a connection between (1) singular points c in the general theorem c = φ(x) and (2) critical points for c = φ(x) which is to say f(x) = φ(x)-c. What about dimensions of things for a single equation? We have N = N, m=1 (one equation), n = N-1. The surfaces are N-1 dimensional (candidate manifolds). Codimension is 1. What about the tangent space which is the nullspace of (Dφ) = (grad φ)T. Well, the vectors of that nullspace are those for which (φ)T v = 0, which is to say φ v = 0. Here is the idea that φ is normal to the surface. So the null space is (gradφ) , it all makes fine sense. Note added 2.14.16. Suppose ei form a basis for the tangent space TxM. Then we know (Dφ)T ei = 0 There are n of these ei since dimTxM= n. Dφ has n columns and N rows. Write this out: ∂1φ1 ∂2φ1 ...........∂nφ1 (ei)1 = 0 ∂1φ2 ∂2φ2 ...........∂nφ2 (ei)2 ∂1φ3 ∂2φ3 ...........∂nφ3 (ei)3 .... --- Σj ∂jφk (ei)j = 0 but a far cry from what I am looking for: ek ei = δki which would say Σj (ek)j (ei)j = δki but wait: that works if you select (ek)j = ∂jφk = ∂φk/∂xj. But I don't think this is my silver bullet. Now we start into the examples. Example 6.12, φ(x,y) = xy so Dφ = φ = . Here RN = R2 and n = 1 so surfaces are just curves. This φ vanishes only at the origin, so all other points are regular so the surface φ(x,y) = c (that is, xy = c ) is a manifold for any c other than 0. Each level curve is two hyperbola sections in the x-y plane. Each of these two piece curves is a manifold, but Sja does not mention the two pieces issue -- the fact that it seems these manifolds are disconnected. At the singular point = 0, the solution φ-1(0) is the two axes. This space self intersects and that confirms that it is not a manifold. Plotted as c = z = xy, we get a normal saddle point at x = 0. At this point, both partials of φ = xy vanish, so it is a Buck critical point. But also that makes rank Dφ be 0 so that is a Sjamaar singular point. His plots of the level curves and gradients are accurate. I did it myself just for fun but threw it out. Example 6.13, φ(x,y) = x3+y3-3xy. This similar to the example of Buck page 356. φ =Dφ = . This example has m = 1 equations and N = 2 variables. So φ(x,y) = c specifies a curve in 2D space, which curve is a 1D manifold,. I have plotted these curves on the left below for different values of c. Each of these curves is a horizontal slice of the 3D surface shown on the right -- each slice is a different c (z = c). . Looking at Dφ, singular point then is when x2 = y and y2 = x, so y4 = y so y3 = 1 so y = 1 and then x = 1. So one singular point is at (x,y) = (1,1) and this then is the place Buck shows a minimum, that is, a critical point. Another singular point is at (x,y) = (0,0) and that is a saddle point. Here are my Buck pictures. What is the tangent space here? We must solve (Dφ)v = 0 so TS is the nullspace of Dφ. But we know that φ is the normal to the surface, so we are just solving these equations (x2-y)v1 + (y2-x)v2 = 0 and this tells you the ratio of v2/v1. v2 = -[ (x2-y)/(y2-x)]v1 so v = v1(1,[ (x2-y)/(y2-x)] ) = a tangent vector Remember that you draw these in the level curve space on the left above, not on the right. He has them all plotted. Example 6.14. Here for x in RN we take φ = xx and φ = 2x so only singular point is at x = 0. There is m = 1 equation, N = N, so solution manifold will be dimension N-1. The manifold φ-1(c=r2) is a hypersphere of radius r, and for c< 0 the solution set is empty. Easier to think φ = 2r so the gradient is normal to the surface, and the tangent vectors are such that 2rv = 0. That is to say, to find the tangent space we solve (Dφ)v = 0, but this IS the equation 2rv = 0. Of course rv = 0 describes the obvious tangent space for v. So hypersphere is a manifold certainly for any c>0. Comment. In the above example, the hypersphere surface is N-1 dimensional and codimension = 1 as usual. If you select c = 0, sphere is a point. This is not a manifold of dimension N-1 but it happens to be one of dimension 0. So technically you can have a critical point correspond to a manifold, but that manifold will be "the wrong dimension", here 0 instead of N-1. [76] Example 6.15. Now we have two φ functions (m=2) on R4 as shown. Then Dφ has 2 rows and 4 columns as shown. Critical point would be where all 2x2 dets vanish. We see this requires x1= x2 = 0 and then that in turn requires x3 = x4 = 0 and we end up with x = 0 being the only value where Dφ drops below rank 2, and this corresponds to c = 0 so c = 0 is the only singular point. Staying away from that, we have M = φ-1(c) being an n=2 manifold (a 2-manifold) within R4 so codimension = 2. The tangent space requires (Dφ)v = 0 and he shows that v = e2, e4 are tangent space spanning vectors at all points x on M (unusual). Hard to plot this in R4 to see what that looks like. Since this is a 2-manifold in R4, could you plot it in R3? For example, let coords by x,y,z,w. Then for c = (1,0) we would get x2+y2= 1 xz-yw = 0 w = xz/y . In x,y,z space you have a cylinder of radius 1 parallel to the z axis, z can take any value and then w as shown. For any point on the cylinder, you have normal vector in the x-y plane, and tangent vectors in that plane. But these have nothing to do with the normal and tangent vectors in R4. Example 6.16. This does not really fit into the picture because now we have φ(A) = C where A and C are square matrices which are n x n. So we now have n2 equations, I guess that is OK. The particular function φ of interest in this example is φ(A) = ATA so our equation set is ATA = C. Since ATA is symmetric, we must have C = a symmetric matrix. So our "variable x" here is matrix A which you could I suppose write as a column vector if you wanted. Now what does (Dφ) look like? It is (Dφ(A)) and so is some kind of tensor since φ(A) is itself a matrix. Maybe write Cij = [φ(A)]ij = (ATA)ij = (AT)ikAkj = AkiAkj (A)ij = Aij (A + dA)ij = Aij + (dA)ij small variation in A Now define dA = h B where h is very small. Then φ(A+dA) = φ(A+hB) = (A+hB)T(A+hB) = (AT + hBT)(A+hB) = ATA + h(BTA + ATB) + h2BTB φ(A+dA) - φ(A) = h(BTA + ATB) + h2BTB Dφ(A) = BTA + ATB if Dφ(A) ≡ [φ(A+dA) - φ(A)]/h This is the change in φ(A) if we go in the B direction, so to speak. Try components: recall (Dφ)ij = ∂jφi so first index on (Dφ)ij is the function label. [Dφ(A)]kl,ij = ∂[φ(A)]kl /∂Aij = ∂[ATA]kl /∂Aij = Σs ∂[ AskAsl ]/∂Aij = Σs [ Ask (∂Asl /∂Aij) + (∂Ask /∂Aij)Asl ] = Σs [ Ask δs,iδl,j + δs,iδk,jAsl ] = [ Aikδl,j + δk,jAil ] So far then I have [Dφ(A)]kl,ij = ∂[φ(A)]kl /∂Aij = δl,j Aik + δk,jAil = a rank-4 tensor Try doing this Σij [Dφ(A)]kl,ij Bij = Σij [ δl,j Aik + δk,jAil] Bij = [Σij δl,j AikBij + Σij δk,jAilBij] = [Σi AikBil + ΣiAilBik] = (ATB)kl + (BTA)kl = [ ATB + BTA]kl Then as a matrix equation we can write [Dφ(A)] B = ATB + BTA // this is the meaning of p 77 A 4 2 where on the left I am contracting a rank-4 tensor with a rank-2 tensor to get a rank-2 tensor. Using the other method above I got Dφ(A) = BTA + ATB dA = hB So at least both methods give the same thing, We end up with a matrix which is n x n and I guess we want to show that it is always full rank n so we have no singular points. Question: does a symmetric matrix always have full rank? No, obviously. Somehow you would need to lay out the matrix in a linear n2 vector and talk about this situation. Where does "rank" determine something? OK, now 2 PM Monday, I am out of time. This example is not very important, I was just fiddling here. Today is 7/20/15 and I won't be able to resume until Fri 7/31/15 assuming I get back here. Resume 8/13/15 (off by 14 days!) So we have shown that C = ATA with C being a symmetric constant matrix results in [(Dφ)(A)] = BTA + ATB = a matrix which let us call D. The matrix B is that matrix in whose "direction" A is allowed to change while maintaining C = ATA. Sja names this matrix C which is confusing since we already have a matrix named C! Well, he avoids writing C = ATA, but it is confusing. So consider the matrix D ≡ BTA + ATB. We want to know if this has full rank or not. Seems to me full rank would mean rank = n since these are nxn matrices. But I don't think this is the right rank concept for this example. The equation φ(A) = ATA = C is a mapping from the space RnxRn into the space of all symmetric matrices which he calls W. And he refers to RnxRn as V, so then φ: V → W and of course W is a subset of V. The dimension of V seems to be n2. For space W we have Cij = Cji so how many equations is that? A triangular part of C has (n2-n)/2 dimension, so there are then (n2-n)/2 equations, so I would guess that the dimension of W is dim(W) = n2 - (n2-n)/2 = n2 - (n2/2) +n/2 = (n2 + n)/2. Call this quantity s so s = (n2 + n)/2 . Then φ: V → W is then φ: Rn*n → Rs. If I linearize everything into vectors, I suppose I could obtain (Dφ) as a huge matrix and then ask about its rank in the same sense as all our previous example. Maybe nullity = 0 here so rank+nullity = range-dim, so need rank = s and this would be the full rank. I have seen in "the meaning of" how having full rank means you can solve the equation v = Mu and thus declare things to be 1-to-1. Somehow Sja uses this concept to say that D will have full rank if you can solve D ≡ BTA + ATB for B, but I cannot follow that argument. He goes on to show than that the solution is B = (1/2)DA. Is that really true? BTA + ATB = [(1/2)DA]TA + AT(1/2)DA = (1/2)ATDTA + 1/2ATDA = ATDA So in fact B = (1/2)DA does not solve this equation even if ATA = 1. I see now that this example is only interested in C = 1, I did not realize that at first (orthogonal matrices). Let's try B = (1/2)AD instead: BTA + ATB = [ (1/2)AD]TA + AT[ (1/2)AD] = (1/2) DTATA + (1/2)ATAD = D so that then is the correct solution. So somehow the fact that D ≡ BTA + ATB with symmetric D has a solution for B tells us that our effective DA thing has "full rank" and as a result our set of orthogonal matrices forms a manifold. I will have some notes to give him on this I think. This finally concludes Chapter 6, a long voyage. Now we need some meta notes I suppose. And starting meta notes at 7:40 PM is not a good plan, so wait till tomorrow.