App A and B notes
DOCX · 40.9 KB
Open DOCX file
Notes by Phil written after finishing Chapter 4 of Sjamaar's manuscript. Appendix A covers set-theory glossary items (images, preimages, injective/surjective maps, index sums) and the topology of Rn, including a discussion of compactness and a theorem Phil questions. Appendix B covers the fundamental theorem of calculus, Jacobian matrices, the chain rule in matrix and pullback notation, and a derivation of the implicit function theorem.
AI-written summary; may contain errors.
Extracted text (machine-read; may contain errors)
Appendix A and B Notes (Sjamaar) PhL 7.12.15
I read these appendices after finishing Chapter 4.
Appendix A: Set Theory
A.1 Glossary
This is a great concise summary of all the key definitions and facts. Some items are new to me.
X - Y = all elements in set X minus those in set Y (the complement of Y in X). Wiki says this notation is ambiguous since it might mean all differences of the form x-y, so they use X/Y = X-Y (X not in Y)
Cartesian product: X x Y Examples
S1 x [0,1] = circle x (line segment) = cylinder shell S1 = circle
What this means is your coordinates are (r,θ) x z
S1 x S1 = circle x circle = torus shell (r,θ) x (R,φ)
R x R = the plane
When Sja writes f: X→ Y, he means this: the mapping is defined for all x in X, so X is the domain, no confusion about that. The set of all f(x) in Y is called the range = image. Since Y is typically larger than the range = image, it has a different name. Y is the codomain or target space. Sja talks about image, but does not use the word range. If X and Y are reals, f is a function, otherwise it is a mapping, just a convention in math writing.
A is subset of X: f(A) = image of A under map f
B is subset of Y: f-1(B) = preimage of B under map f ( f may not have an inverse!)
f-1(c) = the preimage in X of a point c in Y. It implies f(x) = c and the solution of this equation for the locus of x I usually call a level set or level curve. Sja calls it "the level set of f at c". Remember that the sets may be discrete so curves may not exist, so he is right. Also called a the fiber of f at c.
f o g = composition, I am already happy with this.
injective = same as one-to-one f : X → Y
Note that the image of f(X) might be less than all of Y.
If f(X) = all of Y so that f(X) = Y, then f is surjective which is the same as onto.
If f:X→Y is both injective and surjective, it is called bijective.
Theorem: f = bijective f-1f(x) = x and f(f-1(y) = y for all x in X and all y in Y
f-1 o f = 1 and f o f-1 = 1 (two-sided inverse)
Consider now f: X → R where X is a finite set (such as 1,2...N) and R = reals.
Consider: Σall x in X f(x) = a real number.
X is called an index set. Image items I might write as fn for n = x in X. So above is Σnfn.
Example 1: X = {1,2.....N} sum of interest = Σi=1N fi = Σi=1N f(i)
Example 2: X = { (i,j) where i in X, j in X} sum = Σi,j fi,j
Example 3: X = { (i,j) where i in X, j in X and i ≤ j} sum = Σi<j fi,j
These sums are what we see when we talk about a general differential form!
The last two examples are a multi-index and an ordered multi-index.
Example 3 where N = 3. The dots in the "tableau" show a graph of the index set (i,j) with i ≤ j.
The numbers by the dots indicate a sample function f(x) = f(i,j) = i + j. Exercise (A.2) asks reader to consider to compute the sum S(N) ≡ Σ0≤i≤j≤N (i+j) for arbitrary N using induction.
A.2 Topology of Rn
B(ε,x) = closed ball radius ε around x
Bo(ε,x) = open ball radius ε around x (every point has some space around it)
open set = not defined by Sja, but I think you can put an open ball around any point.
closed set = every sequence in the set converges to a point in the set
Example: the set [0,1) is neither closed nor open
Example: the set R is both closed and open (not many of these)
A is bounded set = ||x|| < R for some fixed R for all x in A.
compact = closed and bounded (for Rn)
Theorem: integral of continuous function over compact set is always clean and finite.
Seems wrong. Consider integral of 1/x over [0,1] which diverges. I think he means
Theorem: integral of continuous bounded function over compact set is always clean and finite.
Maybe a continuous function has to be bounded? Yeah, that's right. The function 1/x is not continuous at x = 0, I am sure that Bucks show this. Yes, you cannot get |f(x) - f(0)| < ε for any δ since f(0) = ∞ for 1/x.
Restate: the function 1/x is not continuous at x= 0. More generally, an unbounded function is not continuous at any point where the function is infinite.
Appendix B: Calculus Review
B.1 The fundamental theorem of calculus
He states this three ways as shown. The function being integrated is called f(t). The first way is the statement of a definite integral from a to b, integral is called F where f(x) = dxF = F'(x)
!Syntax Error, If(t)dt = F(b)-F(a)
The second way says dx(!Syntax Error, If(t)dt) ) = f(x) and F does not appear. Derivation from first:
!Syntax Error, If(t)dt = F(x)-F(a) dx(!Syntax Error, If(t)dt) ) = dxF(x) = f(x)
"the integral of a continuous function is differentiable" and "that derivative is the integrand f".
Easy to show that !Syntax Error, If(t)dt) is continuous I think.
Third way is same as first way, but replace f(t) by F'(t)
!Syntax Error, IF'(t)dt = F(b)-F(a)
or
!Syntax Error, IF'(t)dt = F(x)-F(a)
or
F(x) = F(a) + !Syntax Error, IF'(t)dt F(t) arbitrary function, they call it g(t)
Expresses a function in terms of its derivative.
B.2 Derivatives of vector valued functions of vectors
Here (B.4) gives a definition of ∂jφi making use of the unit vector ej . A simple way to state this definition I think. We have φ : Rn → Rm.
Warning: If you are discussing ∂jφi on an open subset U of Rn , this definition is fine because there is always room around a point to do the derivative. But if U is a closed set, the object ∂jφi might not be well-defined on boundary points of U. [ this warning is stated bottom page 130 ]
The Jacobian matrix is this (generally not square)
J(x) = ∂1φ1 ∂2φ1 ... ∂nφ1
∂1φ2 ∂2φ2 ... ∂nφ2
...
∂1φm ∂2φm ... ∂nφm
In tensor doc I have
(A)ij ≡ ∂jAi ≡ Ai,j // note index reversal (G.1.2)
so fair to say that
J(x) = (φ) = a matrix // a notation Sjamaar does not use.
J(x) = Dφ = a matrix // a notation Sjamaar does use.
Now what is this:
Jv = (φ)v = directional derivative along v.
Example
Jen = (φ)en
(Jen)i = [(φ)en]i =(φ)ij (en)j = ∂jφi (δn,j) = ∂nφi = ∂φi/∂xn
So this makes sense as the derivative along the n axis.
Example 1: φ : R1 → Rm (curve) t → x
Dφ = [ ∂tφ1, ∂tφ2....∂1φm]T = ∂tφ(x(t))
In this case, the image for some [a,b] is a curve in Rm [ the mapping φ is the actual curve for Bucks]. So in this case Dφ = ∂tφ is the velocity. Bucks might write
γ : R1 → Rm (curve) v = ∂tγ
Example 2: φ : Rn → R1 x → t
Dφ = ∂1φ1 ∂2φ1 ... ∂nφ1 = (φ1)T
So in this case the Jacobian matrix is a row vector equal to the gradient of your scalar function φ. The actual result is a row vector which we indicate with the T thing, since f is normally a column vector.
Consider:
Jv = (φ)v = (φ1)T v = φ1 v = a number
Sja comments that φ points in the direction of max increase of φ, direction of steepest ascent. This is my usual argument due to dφ = φ1 dr .
C1 = f is continuous and differentiable on U in Rn = "continuously differentiable"
One must mind the warning given above if U is a closed set! If open, less worry.
C2 means that ∂2ijf exists and is continuous in all of U.
Cr = obvious
C∞= "smooth"
B.3 The Chain Rule
I will have to ponder this section, it views things differently than my usual view.
Suppose we have
y = φ(x) Rn → Rm → Rs ( Sja uses Rk instead of Rs)
z = ψ(y) x φ y ψ z
U V W
Then one writes
z = ψ(φ(x)) = (ψ o φ)(x) = a composition = θ(x)
Write a component of this latter
zi = ψi(φ(x)) = (ψ o φ)i(x)
I am happy to regard zi = zi(x) as shown above. So
[∂zi/∂xj] = Σk=1m [∂ψi(y)/∂yk] |y=φ(x) * [∂φk/∂xj] i = 1...s and j = 1.. n
This is completely clear to me. The question is how to write this in compressed notation.
(Dφ)kj = (φ)kj = ∂jφk = [∂φk(x)/∂xj] = last factor above
(Dψ)ik = (ψ)ik = ∂kψi = [∂ψi(x)/∂yk] = middle vector when taken at y = φ(x)
(Dθ)ij = (θ)ij = ∂jθi = [∂θi(x)/∂xj] = [∂zi(x)/∂xj] = left factor
So the chain rule then says
(Dθ)ij(x) = Σk=1m (Dψ)ik(y) |y=φ(x) (Dφ)kj(x)
or
(Dθ)ij(x) = Σk=1m (Dψ)ik(φ(x)) (Dφ)kj(x)
or
(Dθ)(x) = (Dψ)(φ(x)) (Dφ)(x) matrix equation, generally non-square but conforming.
or
(D(ψ o φ))(x) = (Dψ)(φ(x)) (Dφ)(x) // agrees with p 131 A
So I guess I have never done this so carefully before. It is just matrix multiplication of the relevant R matrices, and nothing need be square, and we are not talking determinants at this point.
Once again, the component equation is this
[∂zi/∂xj] = Σk=1m [∂ψi(y)/∂yk] |y=φ(x) * [∂φk/∂xj] i = 1...s and j = 1.. n
or
[∂(ψ o φ)i/∂xj] = Σk=1m [∂ψi(y)/∂yk] |y=φ(x) * [∂φk/∂xj] i = 1...s and j = 1.. n
Now you could remove the i subscript and write these two equations as
[∂z/∂xj] = Σk=1m [∂ψ(y)/∂yk] |y=φ(x) * [∂φk/∂xj] i = 1...s and j = 1.. n
or
[∂(ψ o φ)/∂xj] = Σk=1m [∂ψ(y)/∂yk] |y=φ(x) * [∂φk/∂xj] i = 1...s and j = 1.. n
If you leave off the |y=φ(x) you get
[∂(ψ o φ)/∂xj] = Σk=1m [∂ψ(y)/∂yk] * [∂φk/∂xj] i = 1...s and j = 1.. n
which Sja refers to as "sloppy abbreviated notation". He does not use bold fonts despite his introductory section, so this adds a little to confusion. User must understand that y = φ(x).
Some people might write the above as
[∂ψ/∂xj] = Σk=1m [∂ψ(y)/∂yk] * [∂φk/∂xj] i = 1...s and j = 1.. n
where now the reader has to understand that [∂ψ/∂xj] means [∂ψ(φ(x))/∂xj]. The reason this notation is bad is that you might think [∂ψ/∂xj] means [∂ψ(x) /∂xj].
Pullback notation: (ψ o φ) = φ*ψ so then we have
[∂( φ*ψ )/∂xj] = Σk=1m [∂ψ(y)/∂yk] |y=φ(x) * [∂φk/∂xj] i = 1...s and j = 1.. n
(D(φ*ψ))(x) = (Dψ)(φ(x)) (Dφ)(x) // agrees with p 131 A
and the first line is the famous (B.6) but my form is more general.
B.4 Implicit Function Theorem
Consider φ(u,v) = 0 where φ lies in Rm, u in Rn and v also in Rm.
Think of (u,v) as being in Rn+m.
The idea is that φ(u,v) = 0 might define a function v = f(u) or u = g(v) and so these functions are "implicit" in the equation φ(u,v) = 0 .
We are just combining (u,v) into a single coordinate r = (u,v) in Rn+m. We can talk about
(Duφ)ij = ∂φi/∂uj = part of the total R matrix
(Dvφ)ij = ∂φi/∂vj // this matrix is m x m, so square, so likely invertible
Now suppose there is a "solution implicit function" which is v = f(u). Then we could talk about
(Df)ij = ∂fi/∂uj
Are these three Jacobian matrices related? It seems to me that
φ(u,v) = 0
so
φi(u,v) = 0 for all u and v in some range
Then
dφi(u,v) = 0
which then says
Σj=1n(∂φi/∂uj) duj + Σk=1m(∂φi/∂vk) dvk = 0
or
Σj=1n(Duφ)ij duj + Σk=1m(Dvφ)ik dvk = 0
or
(Duφ)du + (Dvφ)dv = 0
Back up and write
Σj=1n(Duφ)ij duj + Σk=1m(Dvφ)ik dvk = 0
Apply ∂/∂ua to both sides
Σj=1n(Duφ)ij δj,a + Σk=1m(Dvφ)ik ∂vk/∂ua = 0
or
(Duφ)ia + Σk=1m(Dvφ)ik ∂vk/∂ua = 0
or
(Duφ)ia = - Σk=1m(Dvφ)ik ∂vk/∂ua // the minus sign appears
Now use v = f(u) so that
∂vk/∂ua = ∂fk(u)/∂ua = ∂afk(u) = (Df)ka(u)
We then have from 2 lines above
(Duφ)ia = - Σk=1m(Dvφ)ik (Df)ka
or
(Duφ)(u,v)|v=f(u) = - (Dvφ)(u,v)v=f(u) (Df)(u) //matrix form, and show arguments
Notice that
(Duφ)(u,f(u)) is wrong because ∂uφi(u,f(u)) picks up unintended terms from 2nd arg.
A bit of a semantic thing here
∂u[ φi(u,f(u)) ] would pick up those extra terms
(∂uφi)(u,f(u)) does not pick up the extra terms
this means compute (∂uφi) and THEN evaluate it as shown
So I will use the second notation and restate
(Duφ)(u,f(u)) = - (Dvφ)(u,f(u)) (Df)(u)
Now I noted already that the (Dvφ) matrix is mxm = square, so matrix (Dvφ)-1 exists with suitable conditions, then we have
(Dvφ)-1(u,f(u)) (Duφ)(u,f(u)) = - (Df)(u)
so final form is then
(Df)(u) = – (Dvφ)-1(u,f(u)) (Duφ)(u,f(u)) // which is p 132 B
Now what is all this business about u and v having to be close to u0 and v0? Well this same issue came up in my Buck notes (see Ch 5 meta). The question has to do with the invertibility of (Dvφ). You have to be assured that J = det(Dvφ) ≠ 0 = like det(S) in tensor doc. If you pick a point (u0,v0) where this det ≠ 0, then you can reasonable imagine that there are at least small balls around u0 and v0 where det ≠ 0 and for u and v each in its small ball, we then have det(Dvφ) ≠ 0 . The invertibility of this thing is reasonable at a local level, but globally there is likely to be some combination of u and v where det = 0. In terms of the function v = f(u), that is like my x' = F(x) and we have dx' = Rdx and dx = Sdx' so the R matrix is invertible locally. OK, I like it.
Why is this result above called "the implicit function theorem" (Dini's theorem in Italy)? It is a statement about the R matrix of the implicit function v = f(u) which is inherent in φ(u,v) = 0 . I guess one point is that if the above expression is valid (meaning if you are near a point where Dvφ is invertible), then the implicit function v = f(u) exists AND is differentiable and I guess C1.
Application: Suppose in φ(u,v) = 0 we have φ(u,v) = g(v) - u , a special case. Then our rule above says
(Df)(u) = – (Dvφ)-1(u,f(u)) (Duφ)(u,f(u))
and we have
(Dvφ) = (Dvg) and (Duφ) = - 1
so it seems that we have
(Df)(u) = – [(Dvg)]-1 (-1) = + (Dvg)-1(u,f(u))
But of course φ = 0 says that we have g(v) = u . Since we already were writing v = f(u) we can perhaps assume locally that things are invertible (as just commented on above), so then f(u) = g-1(u) and then
(Dg-1)(u) = (Dvg)-1(u,g-1(u)) // which is p 133 A
This gives you the "differential" or R matrix of an inverse vector function in terms of the R matrix of the function. Item B then states this result for φ(u,v) = 0 in 1D. Again we have u near some u0 and v near v0 .
This is called "the inverse function theorem" and again if (Dvg) avoids det=0, the inverse function is differentiable. If g-1 is differentiable, then you can probably integrate to show g-1 is continuous.
A simple example where u = g(v) = v2 is considered, and we have to make sure we don't have v0 = 0 because at that point there is no inverse. We have g-1(u) = which has a branch point at u = 0, and the quantity [ g-1]'(u) = 1/[1] is undefined at v0 = u0 = 0.
B.5 The substitution formula for integrals
The Jacobian Integration Rule is stated for the case y = p(x) where it is required that p is invertible (bijective) meaning one-to-one, AND both p and p-1 are C1. He writes [ note x = p-1(y)]
∫V f(y)dy = ∫U f(p(x))dx | det(Dp)(x)| // page 133 C
Now thing of g(v) = u as p(x) = y. Our inverse function theorem above then reads
(Dp-1)(y) = (Dxp)-1(y,p-1(y))
We have to make sure that det(Dxp) ≠ 0 in order that p-1 be differentiable. Also, the sets V and U have to be open sets.
As usual, Sja always states the 1D version, and those abs values ARE there!
I often apply this rule to sets that are not open and there are points where the function is not 1-to-1. I guess one would have to show that the problem areas have no contribution to the integral and so on.
I am just aimlessly playing around here. The claim being made in p 132 B is this
(Df) = - (Dvφ)-1 (Duφ)|v=f(u)
That would seem to say that
(Dvφ)(Df) = - (Duφ)|v=f(u)
or
(Dvφ)ij(Df)jk = - (Duφ)|v=f(u)ik
or
Σj ∂φi/∂vj * ∂fj/∂uk = - ∂φi/∂uk|v=f(u)
Well remember that v = f(u) so vj = fj and then this says
Σj ∂φi/∂vj * ∂vj/∂uk = - ∂φi/∂uk|v=f(u)
and this begins to look like a chain rule, but what about that minus sign? I think Bucks dealt with this already, I will take a look.