stakgold chap 2 meta
DOCX · 298.4 KB
Open DOCX file
Personal study notes by Phil, begun 1.17.09 and revisited 3.14.11, that summarize and comment on Stakgold's Chapter 2 section by section (2.1 to 2.11). Topics include metric, normed, Banach and Hilbert spaces, separable Hilbert spaces, functionals and the Riesz theorem, closed and adjoint operators, symmetric operators, the spectrum, completely continuous operators and extremal principles.
AI-written summary; may contain errors.
Extracted text (machine-read; may contain errors)
Stakgold Chapter 2 Meta Notes PhL 1.17.09
These meta notes are excellent, if I do say so myself -- during a review 3.14.11.
CONTENTS : Chapter 2: "Introduction to Linear Spaces"
Meta Meta Review of Chapter 2 2
Meta Review of Chapter 2. 5
2.1 Functions and Transformations (Operators). ( pp 92-96 ) 5
2.2 Linear Spaces (= Vector Spaces). ( pp 96-99 ) 5
2.3 Metric Spaces, Normed Linear Spaces, Inner Product Spaces ( pp 99-116 ) 5
Metric Spaces 5
Normed Linear Spaces and Banach Spaces 6
Inner Product Spaces and Hilbert Spaces 6
2.4 Properties of Separable Hilbert Spaces (HS). ( pp 116-135 ) 7
2.5 Functionals. ( pp 135-139 ) 8
Riesz Representation Theorem (1907). 9
Extension of a functional from subspace to full space. 9
Dual Basis Idea (p 138). 10
2.6 Operators. ( pp 139-146 ) 10
Closed operators -- a somewhat painful topic. 11
2.7 Operators in the Hilbert Space En(c) ( pp 146-165 ) Finite Dimension Hilbert Spaces 13
Adjoint Operator. 13
Nullspace Facts. 13
Connection between Nullspace and The Eigenvalue Problem. 13
Eigenvalues, secular equation, geometric and algebraic multiplicities. 14
Symmetric Operators. 15
Spectral Theorem for Symmetric Operators in En 16
Extremal principles for Symmetric Operators (162) 16
2.8 The inverse of an linear operator ( pp 165-180 ) Infinite Dimension Hilbert Spaces 17
Regular, essentially regular, and singular 17
Closed operator stuff 18
Classifying Closed Operators. 19
Adjoint operator stuff (A*) 20
Alternative Theorem: 20
Complication for unbounded operators regarding the adjoint. 20
Self-Adjoint Operators 21
Symmetric Operators 21
2.9 The spectrum of an operator ( pp 180-184 ) 24
Point, continuous and residual spectra; the resolvent set. 24
2.10 Completely Continuous Operators. ( pp 184-187 ) 26
2.11 Extremal Properties for Bounded Operators ( pp 187 - 190) 28
___________________________________________________________________________________
Meta Review of Chapter 2.
_________________________________________________________________________________
2.1 Functions and Transformations (Operators). ( pp 92-96 )
The notion of 1 to 1 and "inverse exists", notion of domain and range each within some space. The A:X→Y notation is strangely not used, I like to use the word "mapping". In general, X and Y are some kind of spaces or perhaps just sets. Things are interesting in the case that X and Y are both vector (linear) spaces so you can add and scale items (algebraic structure), which you cannot do with just sets. Then the objects in the spaces are vectors and not just boring points. Even if X and Y are linear spaces, the operator A need not be linear. Range and/or domain can be less than their full respective spaces Y and X. The dimensionality of the spaces could be finite or infinite. Another interesting case is A:X→X. So the subject in general is extremely broad. This book will deal with a subset of the broader subject, but even this subset will be broad relative to the simple matrix theory one gets when X = Y = En. This simple case however does provide a set of concepts which are then adapted and tuned for the more general spaces.
__________________________________________________________________________________
2.2 Linear Spaces (= Vector Spaces). ( pp 96-99 )
We learn here about the notion of independent vs dependent vectors, the notion of the dimension of a space, the fact that in a finite-N dimensional space, any set of N independent vectors provides a complete basis on which any vector in the space may be uniquely expanded. Orthogonality is not required for such a basis. The notation ek indicates unit vectors like e1 = (1,0,0. .)
__________________________________________________________________________________
2.3 Metric Spaces, Normed Linear Spaces, Inner Product Spaces ( pp 99-116 )
Metric Spaces
A metric space is a animal distinct from a vector space. It contains points with a rule for how you indicate the distance between any two points. This rule d(x,y) ≥ 0 is called the metric of the space. In a vector space alone, although you can do things like add vectors, you don't in general know anything about the distance between two vectors because "distance" is not defined. So generally people add a metric to a vector space to get a space then which is both a vector space and a metric space.
Once we are talking distance, we can talk about convergence of a sequence of points. We learn the difference between a regular sequence d(xn,x)→0 and a Cauchy sequence d(xn,xm)→0 where the actual limit we want to call x might not exist in the space X. If the limits exist, or if we somehow add them into our space, the resulting metric space is said to be complete. Thus, in a complete metric space, every convergent Cauchy sequence converges in the regular sense as well. Cauchy's main work was 1820-1828. Reals complete the rationals. An interesting space for X and Y is a certain space of functions called L2(a,b) so we have then a mapping f: L2(a,b)→ L2(a,b). Another space is that of continuous functions C(a,b). The metric you use tells you the distance between two functions defined on the interval (a,b).
Later in this section we consider a region S in a metric space that might not contain its boundary, sort of like an open set. In this case, sequences might converge to a point on the boundary not in S. You can close the region by adding all these missing sequence limits, and the closed S is called . The entire metric space might be closed to start with, or you might make it be closed. Stakgold mentions the notion of a set being sequentially compact on page 114, but does not use the idea much (yet). Every sequence contains a convergent subsequence if space is compact in this sense.
A related metric space concept is that of X being dense in Y. The idea is that for any y Y, you can get arbitrarily close to y with some x X in the sense d(x,y) < ε . If you close X to be , then X is dense in . The rationals are dense in the reals. The Weierstrass Approximation Theorem says that polynomials are dense in the space of all continuous functions on some closed finite interval.
Normed Linear Spaces and Banach Spaces
Here we add some kind of "length of a vector" concept to a vector space, written ||x|| for x X and this length is called a norm. With the norm defined, a vector space becomes a normed linear space. It does not necessarily have a metric unless you add and use one. So, if you then define and make use of d(x,y) = ||x-y|| as "the natural metric" in your space, the space is called a Banach Space, where it is assumed that the space is complete as discussed a few paragraphs above. Of course you might reject this natural metric and use some other metric, but then X is not called a Banach Space. Banach's work was around 1932, a full century after friend Cauchy. Again: a normed linear space is a vector space to which you add a norm for the vectors. If you then use the natural metric this norm implies, your space becomes the combination of a vector space, a normed linear space, and a metric space. If space is complete, it is then Banach.
Inner Product Spaces and Hilbert Spaces
The metric is a mapping d: (X,X) → R. An inner product is a similar mapping written <x,y>, but we have (inner product): (X,X) → C (complex numbers). Whereas the norm is required to satisfy the triangle inequality d(x,z) ≤ d(x,y) + d(y,z), the inner product has to satisfy some simple rules (addition and scaling etc) which lead to the Cauchy-Schwarz Inequality |<x,y>|2 ≤ <x,x><y,y> where the thing on the right is a product, not a sum as with the metric. (CSI 1821, then 1888). An inner product can only exist inside a vector space since addition and scaling of vectors is required. A vector space with an inner product is called an inner product space. You could choose to use the natural norm ||x|| = <x,x>1/2 to cause the inner product space to also be a normed linear space. Then you could use the natural metric d(x,y) = ||x-y|| to make this space also be a metric space. This combined X = vector space / inner product space / normed linear space / metric space object is called a Hilbert Space (1900-1910), provided the space is complete. You can always complete X (make it closed) by adding missing points.
There is a very famous prototype Hilbert Space we call Rn which is a generalization of R3 , the 3D space in which we live. You can extend Rn to allow for complex components on the axes, then you have En(c). These Hilbert Spaces can be exactly modeled in terms of matrix algebra where a vector is a column of components, and an operator is a square matrix. This prototype Hilbert Space is usually referred to as Euclidian Space. Again, normally the components of a vector on each axis of Euclidean Space are real numbers as in our physical 3D space which Stakgold calls En(r) = Rn, but you can then extend that to En(c).
In Quantum Mechanics we use the inner product <a|b> = <b,a>. This peculiar reversal of order arises because in QM we want operators to be linear on the kets |b> and antilinear on the bras.
Here is a picture I made that shows how all these types of spaces are related.
__________________________________________________________________________________
2.4 Properties of Separable Hilbert Spaces (HS). ( pp 116-135 )
The main thrust of this book is with ∞ dimensional Hilbert Spaces which go beyond the finite matrix idea. It proves useful to limit one's interest to HS's where the spanning vectors form a countable set, and so can be written f1, f2 .... . Such a Hilbert Space is called a separable Hilbert Space and this is all we will ever care about. The countable set of spanning vectors is called a spanning set. Using a spanning set fi, you can get arbitrarily close to an element in X by doing a fit x = Σi=1N aifi for large enough N (the word fit of course requires a metric, but HS has a metric.). In general, all the ai coefficients will "move" as you bump up N and crank down on your fit to achieve a smaller ε. If you can find an ∞ spanning set where the ai don't move, then that spanning set is a complete basis and you can take N→∞ and all ai stay put. In this case, basis vectors will turn out to be orthogonal, and the ai are generalized Fourier Coefficients. A favorite counterexample is that in L2, the powers tk form a spanning set but not a basis.
The rest of this section seems to be a general grab bag of HS related facts. Any spanning set of independent vectors forms a basis, even if not orthogonal. Any orthonormal spanning set makes a basis. The Projection Theorem is the idea that if your HS H has a subspace M (= linear manifold in Stakgold), then H = M M and you can write any x H as x1 + x2 etc.
Each set of the famous orthogonal polynomials forms an orthogonal basis on some appropriate interval. Often they are not normalized, such as with the Legendres. They usually have a weight function included in the inner product relating to the concept of measure.
The Riesz-Fischer Theorem (1907) makes a connection between convergence of a sequence of vectors in a HS and the convergence of a sequence of real numbers, but you have to do this with respect to a known basis in the space. The real numbers are the squares of the partial sums of the squared Fourier coefficients. The GSO process for making a spanning set be orthogonal is described.
__________________________________________________________________________________
2.5 Functionals. ( pp 135-139 )
Up to this point we have just been talking about "spaces", and in particular Hilbert Spaces. Now for the first time we talk seriously about transformations that map between our spaces as was quickly outlined in Section 2.1. In this Section we talk about T: H → C where H is some Hilbert Space and C = the complex numbers. Such a transformation is called a functional (it need not be linear) and is written T[x]. In Section 2.6 we will talk instead about transformations which map A: H→H so the Hilbert Space is mapped into itself. These transformations A are called operators and we write the transformation as Ax, ( not as A[x] ).
Since T(x) is in space C, || T[x] || = |T[x] |. And of course since the domain space H is a HS, it too has some norm, so ||x|| is well defined. We can consider the positive real number which is the value of the ratio || T[x] || / ||x|| as x moves around in some S H. If this ratio has a maximum value < ∞, then the functional T is said to be bounded on S, and that max value is called the norm of T or || T ||. Conclusion: || T || ≡ max ( || T[x] || / ||x||) for x S. In the discussion above, we talked about the norm of a vector in a vector space, but || T || is the norm of a transformation between two vector spaces, so not the same animal, but it does fulfill the defining properties of a norm. If the norm is infinite, then the functional is unbounded on S. [ In a sense, the norm is a measure of the maximum "amplification" of the functional T, where ||x|| is the input signal, and || T[x] || the output signal. Different signal inputs are amplified by different amounts. ]
Regular convergence means xn → x. If this is in our domain space, we can consider whether or not we have fn = T[xn] → T[x] = f in the range space ( call this the image sequence.) If we have fn→f in the range space for every possible sequence xn→x in some domain region S, then the functional T is said to be continuous on S.
This reminds us a little of the notion of uniform convergence of a function, but that is not the same concept. Suppose you consider a set of square-integrable functions L2(a,b) for the interval (a,b). This thing L2(a,b) is a Hilbert Space and its vectors are functions f(t). You could consider a sequence fn(t) → f(t). If this convergence happens for all t in (a,b), then we say the sequence converges uniformly on (a,b). Notice, however, that this range (a,b) is not the domain space of some transformation like T[x]. In uniform convergence, the parameter t in fn(t) → f(t) which ranges through an interval is part of the definition of HS L2(a,b). In contrast, when we talk in the previous paragraph about convergence T[xn] → T[x], the transformation argument is a sequence of points in the domain space.
We are often interested in functionals which are linear. In this case, we can prove some useful theorems where S is some region in H: ( both theorems are quickly proven in the text)
Theorem 1. If linear T[x] is continuous at x = 0 in S, it is continuous at all x in S.
Theorem 2: For linear functional T[x] , continuous on S bounded on S
Stakgold says he will only be interested in cases where S = a subspace of H, but I don't think that is a requirement for the above theorems to be true. He often refers to S as DT.
Riesz Representation Theorem (1907).
Recall that T: H → C. We want to say that we can represent T[x] as T[x] = <x,f> for some f in H for all x in H. If this were true, it would say that we can completely characterize functional T by some fixed vector f in H, which seems pretty amazing to me. It turns out that this "representation" is true provided that the functional is linear and continuous on all of H, which Theorem 2 above says is the same a T being linear and bounded on H. This theorem therefore claims there is an isomorphism between two spaces: (1) the set of continuous linear functionals on H; (2) the set of vectors in H. Space (1) is called the dual space, but Stakgold does not get into that subject.
In the proof of this theorem, Stakgold defines the notion of the nullspace of T and calls it N. This space N H is a subspace of H (easy to show he says; a subspace is closed! ). We can then write H = N N where this latter is what I call the perp space. If xN, then T[x] = 0, that is what nullspace means. If there is no perp space, meaning N = H, then T[x] = 0 for all x in H, and we can "represent this" as <x,f> with f = 0. Otherwise, we can surely find and normalize some vector in the perp space N, call it fo. He then shows that the required f to make T[x] = <x,f> is explicitly f = [f0] fo . The proof seems completely trivial as you read it. He also shows trivially that this f is unique, so that means whatever f0 you select in N, you will find that [f0] fo is always the same vector.
This site https://www.math.gatech.edu/~harrell/pde/Rieszapp.html shows trivially that the perp space of a linear functional is one dimensional, so there is really only one possible normalized fo, which explains the mystery of the uniqueness of the f = [f0] fo .
The theorem requires that T be linear and continuous on all of H (which is same as being bounded on H), but as you read the simple proof, you see at once why linearity of T is required, but you do not see why continuity is required. I searched a little on the web and could not find an explanation of where continuity is used, and I see that most sites refer to concepts a little beyond Stakgold in dealing with this theorem. Perhaps this is it: if H were not bounded, then perhaps the solution we find which involves T[f0] would be infinite and therefore non-existent. If T is bounded, we know that | T[f0] | / ||f0|| = | T[f0] | ≤ ||T|| and must therefore exist. Thus, the solution f exists which makes T[x] = <x,f> exist. That must be it!
Now, once you can represent T[x] = <x,f>, I think T[x] is patently continuous because an inner product is continuous in either argument.
Extension of a functional from subspace to full space.
Page 137 talks about how you can extend a bounded linear functional from a subspace DT to the entire space H. We know we can write H = DT DT , and we also can write DT = (N N) because DT being a subspace is also a Hilbert Space and our usual projection theorem applies with respect to T.
It is possible that the subspace DT might not be closed. If that is the case, you close it, and then you can say that DT is a Hilbert Space because it is then complete due to this closure. So close DT to T if it is not closed up front. I will not use the overbar in what follows.
Notice that DT is therefore perp to both N and N. We know from Riesz that we can write T[x] = <x,f> for any x in DT and we will have f be in N . We know that this formula applies for x in N (in which case it will happen that T[x] = <x,f> = 0), and it also applies for x in N, in which case T[x] = <x,f> ≠ 0. The idea is to use this same formula T[x] = <x,f> and extend it to x in DT . Since DT is perp to N, we know that we will have <x,f> = 0. Thus, this method of extension just sets T[x] = 0 for x in DT and it does it all with the single formula T[x] = <x,f> now applying for any x in H ! This method of extension is thus making use of the Riesz Representation Theorem and thus you can only extend a functional in this manner if it is linear and bounded (= continuous).
Theorem 3: All linear functionals T[x] on H = En are bounded (hence continuous).
On page 137-138 Stak shows how this theorem can fail for ∞ dim Hilbert Spaces and he concocts a specific example which is a sequence approaching a delta function.
Dual Basis Idea (p 138).
The claim is that if you start with a basis ei that is non-orthonormal, you can find a set of "dual" (reciprocal) basis vectors ei* which are orthonormal to the ei (but not in general to the other ej*). Later on page 150 it is noted that if your ei are themselves orthonormal, then we have ei* = ei so you don't have to think about the dual basis vectors. Thus subject arose in M&M where they wrote ei* as ei , and this was related to covariant and contravariant and metric tensors. Notice that this is unrelated to the dual space notion of certain functionals on H being isomorphic to H.
A chart on page 139 compares the world of functionals on H = En to that on H = L2(a,b), finite dim versus infinite dim. Here are the key points that are worth noting on the L2(a,b) side:
(1) You can only express T[x = Σ∞aiei] as = Σ∞ ai T[ei] if T is linear and bounded, where ei = some basis. Notice that this equality is an extension to infinity of the definition of T being "linear". Here you are interchanging the order of the operator T with the operator of infinite summation.
(2) Notice that T[x] = <x,f> for bounded T means T[x] = dt f*(t)x(t) . Thus, any bounded T in L2(a,b) can be expressed as a simple integral against some function f over (a,b).
______________________________________________________________________________
2.6 Operators. ( pp 139-146 )
We have A:H→H now, and operator action is written as Ax = f. Concepts of nullspace, domain, range, linear, continuity, norm all carry over from the functional work done in the previous section. In particular we now have
norm = ||A|| = max( || Ax || / ||x||) for x in DA
which is the analog of the functional's norm. We get these three completely analogous theorems:
Theorem 1: Continuity of linear operator at x = 0 implies continuity over all of DA.
Theorem 2: For a linear operator, boundedness continuity (p140)
Theorem 3. In En, every linear operator is bounded (and continuous)
We can extend an operator A defined on DA to operator B defined on all of H by something similar to what we did with the functionals. To make it work, as before we have to close DA to A and then define Bx = 0 out in the DA area. In order to be able to close DA, we must know that operator A is continuous, otherwise the limit points of sequences Axn might not be unique and we have closure problems. The discussion on page 141 also uses the boundedness of A. So if linear A is continuous ( = bounded), then we can always extend an operator A from DA to B defined on all of H, and we write B A. Then B is an extension of A, and A is a restriction of B. In the functional world, the Riesz inner product gave us explicit continuity, but we still had to assume T[z] was bounded.
So in general things are under control for bounded = continuous operators A. But what can be said about unbounded linear operators? We are interested because we know that differentiation is unbounded.
Closed operators -- a somewhat painful topic. ( starts below pencil line on page 141)
Stakgold introduces a new concept here, that of closable and closed operators. If an unbounded linear operator is a closed operator, it is not as "nice" as a bounded linear operator, but it is not as "ugly" as an unbounded linear operator which is not closed or closable. So closed unbounded operators are things we will be at least able to work with, but they are not as nice as bounded operators. The prototype example of a closed unbounded operator is d/dx, which will obviously be very important us.
So what does closable mean? For a bounded linear operator, we know that Axn→0 for a null sequence xn→0, due to continuity. For an unbounded but closable operator, if this is not true, then at least we know that "Axn has no limit" will be the worst case for a null sequence xn. So closable means either Axn→0 or Axn→ does not exist (=∞ below). You cannot have the third bad option which is Axn→g ≠ 0 for xn being a null sequence. So closable is thus defined. [ For a closable operator, the image sequence in the range of any null sequence in the domain either converges to 0 or does not converge. ]
Elaboration: Consider two sequences in the domain, zn → z and yn → z. Then xn ≡ zn - yn is a null sequence and if A is closable, then Axn→ 0 or does not exist. If Axn→ 0, then we have Azn → Az = f and Ayn → Az = f. All is well, the notation "Az" is well defined.
Suppose Axn→ g, some finite vector in the range. This is the case we want to avoid! We then somehow have the situation zn → z => Azn → "Az" = f1 and yn → z => Ayn → "Az" = f2 and g = f1 - f2 and f1 ≠ f2. [ Here, (zn-yn) → 0 is our null sequence and A(zn-yn) → g.] In this case, the notation "Az" does not even make sense, it is not unique. Ugly! You might be able to deal with this, but special notation and methods would be required.
Suppose Axn→ ∞3. We then have the situation zn → z => Azn → ∞1 = f1 and yn → z => Ayn → ∞2 =f2 and g = f1 - f2 or ∞3 = ∞1 - ∞2. But we are happy to have ∞ = ∞-∞. For example consider the following: ∞1 = lim(9n), ∞2 = lim(2n) then ∞1 - ∞2 = lim(7n) = ∞3. Since we have f1 and f2 being ∞ in this case, we can thing of "Az" = ∞ for both cases, and we don't get the contradictory situation shown above when Axn→ g finite. We can include in this case the meaning "limit does not exist" when we write the symbol ∞. Again, we get no contradictory notation.
The point is that we cannot have different limit points f in the range for different qn → z in the domain. For a closable operator, either all those limit points like Az are the same (f) for all domain sequences like zn and yn, or all those limit points don't exist. If the limit points do exist, then we say Azn → Az = f.
Now, a closed operator could have some limit point z right inside DA such that the limit Azn does not exist for zn → z, and that is OK. It sounds sort of like a pole in a function is OK for meromorphic.
Closing a closeable operator to obtain a closed operator.
Notice that z (limit of zn) might not be in DA. In closing an operator, we to close DA so it becomes A so that any such z in zn→z gets included in the domain. This then is the idea of closing an operator. If A is closable, then the closed version of it is called . A closed operator is different from a closed set , though the same word "closed" is used. The domain of A is DA, and the domain of is A.
As the text points out, no interesting operator we ever encounter will by unclosed. So we will always deal either with (1) operators that are bounded = continuous, or which are (2) unbounded = not-continuous but are closable. These are only "mildly discontinuous" not "radically discontinuous". Notice that in case (1), if we take any null sequence xn→ 0, we get Axn = 0, the case Axn = ∞ is ruled out. In this case, if zn → z, we have Azn → Az = well defined. So any bounded =continuous operator is therefore certainly "closeable", since we have just removed one possibility of Azn→ ∞. And if the domain is closed, then A is closed, the subject of the next theorem which we now see is quite obvious.
Theorem 10: A continuous (bounded) operator on a closed domain is a closed operator.
Corollary: A continuous (bounded) operator on all of H is a closed operator.
[ This theorem seems pretty obvious. The image sequence of a null sequence for a continuous operator will be a null sequence in the range, and that is one of the two cases for a closable operator. ]
Now here is the big new fact which will no doubt help us a lot:
Famous Theorem 11: A closed operator on a closed domain is continuous. (closed graph theorem)
[ This seems to suggest that in this case, the image sequence of a null sequence cannot diverge. Somehow, having a closed domain "rescues" a closed operator and makes it be continuous. ]
Note added 4.20.09. I have a big problem with Theorem 11. Suppose I start with a general unbounded operator A which is closable. I close the domain by including in it all possible missing limit points of domain sequences. A is then a closed operator and it acts on a closed domain. Why should A suddenly be bounded, when I assumed from the start that it was unbounded? This seems to imply that there are no unbounded closeable operators. Web says Theorem 11 is the "closed graph theorem". I am misunderstanding some aspect of closed operators. The operator A = d/dt is unbounded but closeable. If I close the domain, does d/dt suddenly become bounded? Maybe the catch here is that for d/dt, you can't close the domain.
OK, here I think is a resolution of this issue: ( Note: see "non-differentiable continuous functions.doc" in "calculus" which shows there are many functions such as the Weierstrass function which are continuous everywhere but differentiable nowhere! )
The Stakgold Issue (p 145 Example 7). The operator d/dt is unbounded and closeable on the domain of "differentiable functions" (Stak calls this set of functions "absolutely continuous" functions). Consider a sequence of functions fn(t) in the domain, where fn(t)→ f(t). He claims on page 145 that functions like f(t) are piecewise continuous (which he includes in his absolutely continuous class). He wants us just to include the limit functions like f(t) in the domain. The limits d/dt (f(t)) exist. Though they will be discontinuous functions, they are still fine L2 functions. So here we are "closing the domain" in a certain sense so that that closable operator d/dt becomes a closed operator.
However (!), the domain of d/dt even including this closure is not a "closed set" within L2 ! As noted above, the set of continuous functions which are differentiable nowhere is dense in L2 ! Consider for example the Weierstrass or some other partial sum as fn(t)→ f(t). In this case, f(t) would not be included in our d/dt closure process mentioned above, since f(t) is differentiable nowhere. Thus, in the d/dt closure process we are NOT including the limits of ALL sequences of functions. We are only including the limits of sequences whose limits are differentiable to yield a function in L2. The limit of the Weierstrass sequence does not make a function which is "differentiable to yield a function in L2".
When you talk about a domain being "a closed set", you have to include the limits of ALL sequences.
So, you can take an open domain of d/dt, you can close it by adding the sequence limits discussed above, but even after this closing process, the domain is still open in the sense that not all sequence limits are in the domain. So the semantic confusion here for me was this: after you "close something" it is still "not closed". In our d/dt closing process, we are "partially closing" the domain, and "fully closing" the operator by so doing, by the definition of a closed operator.
Then we can get happy again about the "closed graph theorem" or Theorem 11 of my Stak Chapter 2 meta notes (text page 142) : an operator which is closed on a "closed domain" is in fact bounded and continuous. The operator d/dt is closed, but its domain of absolutely continuous functions (which includes PC) is not a closed set in L2. Therefore we cannot apply our theorem to say d/dt is bounded.
Consider now an unbounded operator A acting on all of H, which is clearly a closed domain. If A were closed, it would have to be continuous by Theorem 11. But if it is continuous, it must be bounded by Theorem 2 (assuming A is linear, which we may have to assume everywhere here). Thus we have this fact:
Theorem 12: There can be no closed unbounded operators defined over a whole Hilbert Space H.
[ for example, d/dt cannot be defined on all of H ]
However, we can find such operators defined over a space which is dense in H! A dense subset of H is not a closed set, and thus Famous Theorem 11 does not apply.
Fact: We have seen how a closed operator might have an unclosed domain and unclosed range. The word "closed" is still used because the graph of A, which is the set of pairs (x,Ax) in graph space, forms a closed set. This is his remark 3 page 143.
Theorem 13: The nullspace of a closed operator is a closed set. (remark 4).
Fact: The differential operator acting on L2 is the prototype of a closed unbounded operator. Perhaps the dense space contained in H which d/dx can act on is the space of "differentiable" functions in L2. We know that d/dx cannot act on all functions in L2 due to Theorem 12. I think this class of differentiable functions is what Stakgold later calls absolutely continuous functions, and this includes piecewise differentiable functions. This operator is discussed as Example 7 on page 145 and is the only example really that was unbounded.
Possible example of interest. Consider the sequence xn(t) = (1/n) sin(n2t). I would argue this is a null sequence since xn(t) → 0. The integral of the square of this function on 0,1 is perhaps 1/(2n2) which → 0. Now what is the image sequence for A = d/dt ? We get fn(t) = n cos(n2t). This provides us then an example where the limit of the image sequence of a null sequence "does not exist", so the image sequence does not converge. Certainly the norm gets larger ~ n and so a limit if it existed would have infinite norm. This is really the idea of being unbounded. The fact is that the image function in this case is an ever larger amplitude faster moving cosine and I think it is fair to say that this does not approach any reasonable kind of limit.
______________________________________________________________________________
2.7 Operators in the Hilbert Space En(c) ( pp 146-165 ) Finite Dimension Hilbert Spaces
The opening sections are pretty easy, examples are given such as rotation and sheer transforms. Active and passive is mentioned. One good comment is that if Aej = akjek on some basis ek that is not orthogonal, you can still invert to get anj = <en, Aej> using the reciprocal basis element ej.
The following items all are the same for determining invertibility of A: A is invertible if: detA ≠ 0, A has full rank, A has no nullspace, Ax=f has exactly one solution, A is regular. If A is not invertible, then A is singular. Only two choices here.
Adjoint Operator.
Given linear operator A, you can define another linear operator A* (the adjoint) by this relation:
<Ax,y> = <x,A*y>
If A = A*, then A is self-adjoint. This is the same as being complex symmetric = Hermitian.
Nullspace Facts.
The Alternative Theorem says that either Ax=f has a solution for all f (meaning the range is all of H), or the equation Ax = 0 has a non-trivial solution (there exists a nullspace, by which I always mean a non-trivial nullspace beyond the 0 vector). The more precise statement is that the following spaces are equal: RA = (NA*) which is the same as NA = (RA*) . It turns out that A and A* have the same size (dimension) range and the same dimension nullspace, so for example we know that dim(NA) = dim(NA*). The dimension of a nullspace of A is called the nullity of A or ν(A), while the dimension of the range is called the rank of A or ρ(A). So we just claimed that ρ(A*) = ρ(A) and ν(A*) = ν(A). Obviously dim(NA) = n - dim(NA) = n - ν(A), but
dim (NA) = dim (NA*) = dim(RA) = ρ(A)
and therefore ρ(A) + ν(A) = n. In the alternative theorem as stated above, if Ax = f always has a solution, then we are in the case ρ(A)= n and ν(A) = 0. The other alternative is that ρ(A) < n in which case we must have ν(A) > 0.
Connection between Nullspace and The Eigenvalue Problem.
We know we can solve the equation Bx = f for any f as long as detB ≠ 0. The reason is simple: in this case B-1 exists, and x = B-1f and there is your explicit solution.
Now suppose B = A - λI = A - λ. From the above, if detB ≠ 0, we can solve the equation Bx = f to get x = B-1f as the specific solution. Now detB = 0 is a polynomial of degree N in λ and it will have N roots which we call λi and these are called the eigenvalues. What is the significance of these eigenvalues λi ?
Eigenvalues, secular equation, geometric and algebraic multiplicities.
Definition: in this context, det[B(λ)] = 0 is called the secular equation.
Definition: An eigenvalue is a root of the secular equation.
(1) if λ ≠ λi, then detB ≠ 0 and Bx=f has the solution x = B-1f and then ρ(B) = dim(RB) = N which means that B has no nullspace, B has full rank, so Bx=0 has no solutions. But Bx=0 is the same as Ax = λx so when λ ≠ λi, our eigenvalue equation Ax = λx has no solutions.
(2) if λ = λi, then Bix = f might still have some solutions, but they are not of the form x = Bi-1f because
Bi-1does not exist. Perhaps we will have dim(RBi) = ρ(Bi) = n < N and this will correspond go ν(Bi)= N-n > 0. In this case, our eigenvalue equation Ax = λix will have some non-trivial solutions, and these are called the eigenfunctions of A (nullspace solutions of Bix = 0). The number of independent solutions of Bix = 0, or equivalently, the number of independent eigenfunctions for λ = λi, is ν(Bi) but has a fancier name: the geometric multiplicity of λi and we use the notation mi for this multiplicity.
(2, refined). The refinement is that we can remove the qualifier "Perhaps" in the above paragraph.
If dim(RB) = ρ(B) = N (B has full rank), we know that Bx = f has a solution for any f in En (meaning En(c)). This means we can write that solution in the form x = Gf where G is some linear operator (you could show linearity easily I think). Therefore, we have f = Bx = BGf and this is true for any f in H = En. Therefore it must be true that G = B-1. This in turn implies that detB ≠ 0. Therefore, if we knew that detB = 0, we would know that B must be less than full rank by at least 1. This means that the nullity of B must be at least 1. Thus, detB=0 [which occurs when λ = λi , an eigenvalue ] => 1 ≤ ν(B) ≤ N. This means Bx=0 has at least one independent nullspace solution at this eigenvalue λ=λi, and it means that Ax = λix has at least one independent eigenfunction.
Theorem: Every operator A in En has at least one eigenvector.
Proof: We certainly know there is at least one eigenvalue because the secular equation has N roots and at worst they could all be the same, and there is your "at least one eigenvalue". It happens that this would only occur when A was a multiple of the identity, but in any case, there is at least one eigenvalue, so we just pick one. For this eigenvalue, we have detB = 0 and we just showed above that Ax = λix then has at least one independent eigenfunction at that eigenvalue. We have really proven this stronger theorem:
Theorem: For every eigenvalue of A, there is at least one eigenvector. In other words, every eigenvalue of an arbitrary A has a geometric multiplicity of at least 1, so mi ≥ 1.
We know that λi is a root of detB = 0 and that root will occur some number of times ki where ki ranges from 1 to N. This ki number is the algebraic multiplicity of λi. Stakgold shows with an easy derivation that ki ≥ mi for any eigenvalue λi. Based on our theorem above, we really have ki ≥ mi ≥ 1.
Example 3: A simple "shear matrix" A = and secular says ki = 2 with both eigenvalues λi = 1, but trivial solution for the eigenvectors finds that the only eigenvector is x = k (1 0) so mi = 1. So right here we have an explicit example of k > m, in case we were wondering if that were possible. Notice that this matrix is not symmetric.
Theorem: Suppose Ax = λx has 1 ≤ ν(B) ≤ N eigenvectors, and suppose each eigenvector has a different eigenvalue, which means each eigenvalue has algebraic multiplicity = 1. In this case, those ν(B) ≤ N eigenvectors will be linearly independent. We can also say mi = 1 for each eigenvalue.
Stak proves this theorem by induction on page 159.
Corollary: If Ax = λx has ν(B) = N eigenvectors with distinct eigenvalues, then the eigenvectors are independent and form a basis for En.
Symmetric Operators.
After the voyage above we return now to this subject. In En we mean by the term "symmetric operator" one that is self-adjoint = Hermitian = complex symmetric. When we get to ∞ dimensional Hilbert Spaces, there will be a distinction between self-adjoint and symmetric, but in En they are the same. [ Cn would be a better name for En, and Rn for the reals. ]
For symmetric operators A, Stak shows on pages 159-160 these familiar facts:
(1) <Ax,x> = real for all x in Cn -- though components of x are complex numbers.
(2) eigenvalues are real
(3) eigenvectors of different eigenvalues are orthogonal
(4) eigenvectors form a basis ( interesting proof uses perp spaces iteratively), so this means there must be a total of N eigenvectors forming this basis.
We know that Σiki = N always. If there are N independent eigenvectors as in (4) above, then it must be true that Σimi = N as well. This means that for each λi, we must have mi = λi and there are no "missing eigenvectors". A non-symmetric operator A very well might have some missing eigenvectors, like that shear matrix example above.
I think it is item (3) above that lets us treat each group of eigenvectors for a given λi as a subspace or eigenmanifold. These spaces are independent of each other. This is what lets us write:
H = M1 M2 M3 ..... Mk
using the direct sum notation (which Stak does not use). For each eigenmanifold Mi, we have mi = ki and then we have Σi=1k ki = N and same for mi. At most there are k = N eigenmanifolds, and at least there is only 1 eigenmanifold.
Comment: If we don't orthonormalize the mi eigenvectors in one or more eigenmanifolds Mi, then the total set of independent eigenvectors { xi} still form a basis, but one which is not fully orthogonal. Here we reinforce our earlier idea that in a finite dimensional Hilbert Space like Cn, a basis does not have to be orthogonal, and this is just because some eigenmanifolds might have mi > 1. Usually you would use GSO to orthonormalize within each Mi , and then {xi} becomes an orthonormal basis.
Spectral Theorem for Symmetric Operators in En
The above partitioning direct sum concept embodies what Stak calls The Spectral Theorem for Symmetric Operators in En. Here we just summarize it as he does:
(a) to each eigenvalue i is associated an eigenmanifold Mi of some dimension mi=ki (geo = alg)
(b) eigenmanifolds are pairwise orthogonal
(c) these eigenmanifolds partition the space, and you can decompose x = x1+ x2 + ... where xi in Mi.
(d) the eigenvalues i are determined by det(A-λI) = 0
(e) due to (c), you can imagine a projection operator Pi for each eigenmanifold Mi.
(f) we can then write A = Σi λi Pi known as the spectral resolution of A (easy to prove, remark 1)
Notice that we then have
Ax = A (Σixi)= Σi (Axi) #1
= (Σj λj Pj) (Σixi) = ΣjΣi λj Pjxi = ΣjΣi λj δjixi = Σi (λixi) #2
which one could derive in reverse starting with Axi = λixi for the component of x which lies in the eigenmanifold Mi.
Remark #2. The eigenvalues of a symmetric operator are real, we know. So it is sometimes useful to list them in decreasing numerical order. It is also useful to have two different methods of listing them off, which I illustrate here:
λ1 λ1 λ1 λ2 λ2 λ3 λ3 λ3 λ3 λ4 λ5 λ5 .. λq
μ1 μ2 μ3 μ4 μ5 μ6 μ7 μ8 μ9 μ10 μ11 μ12 .. μn
Remark #3 Fact: A symmetric operator cannot have all eigenvalues be 0. If so, then Axi= 0 in every eigenmanifold, and you conclude that A = 0.
Extremal principles for Symmetric Operators (162)
Four functionals are defined bottom of page 162, here they are, where we have two kinds of norm type objects, one being the official norm.
||A|| = max I2 over all x I2 = || Ax || / ||x|| = [ <Ax,Ax>/<x,x>]1/2
||A|| = max I1 over all x such that ||x|| = 1 I1 = || Ax ||
|||A||| = max I4 over all x I4 = |<Ax, x>| / ||x||2
|||A||| = max I3 over all x such that ||x|| = 1 I3 = |<Ax, x>|
The lower two objects are not directly related to ||A|| as far as I know
Page 153 very clearly proves that:
Theorem: for a symmetric operator, ||A|| = |||A||| = | the largest eigenvalue| = |μ1|
You can see this pretty easily. In I2 set x = xi and you get I2 = |λi|, so if you max this, you get |μ1|. Similarly, you get I4 = |λi| and maxing it is the same thing. So for a symmetric operator, both these objects ||A|| and |||A||| are "the norm of A".
For a non-symmetric matrix, he claims without proof that this is all you can show
||A|| ≥ |μ1|
|||A||| ≥ |μ1|
which seems an odd result to me, but he gives an example where this is born out. I could not find a proof of this in a quick web search due to having no good search handle.
_________________________________________________________________________________
2.8 The inverse of an linear operator ( pp 165-180 ) Infinite Dimension Hilbert Spaces
Theorem: A is 1-to-1 A-1 exists [ Ax = 0 x=0, ie, there is no nullspace ]
If there were a nullspace, that would imply multiple domain elements mapping to 0, so the mapping is not 1 to 1.
Now the term below "regular" really means an inverse exists for Ax=f for all f and the inverse is "nice". Thus, you can solve the equation Ax=f for x if A is regular.
Regular, essentially regular, and singular
Definition: regular. Recall that a En operator was either regular or singular, but now things are messier. In dimensions, to be regular we need 3 conditions to be met:
(a) Ax=0 x=0, as with matrices; and this implies that A-1 exists;
(b) the range must be the entire Hilbert space, so Ax = f has a solution for any f in H;
(c) A-1 must be bounded (= continuous).
If your RA does not include the boundary of space H you can do the "usual" extension idea to get the boundary included to meet condition (b). Without the boundary, if all other conditions are met, you are essentially regular.
Notice that our old rule that ρ(A) + ν(A) = n now has problems since n = ∞. It seems from (b) that it is possible to have nullity ν(A) = 0, but range is not the whole space. Condition (c) is a niceness condition that we just tack on so that "regular" means something reasonably nice.
Closed operator stuff
Now we consider once again those closed operators. Recall the various items of interest, operators being continuous, being bounded or unbounded, being closed, various sets being closed, etc.
Here is a densepack little review of our theorems from above: all operators are linear from now on, as far as I am concerned. [ I think except for {..} that these are true for ∞ dimensional Hilbert Spaces. ]
Theorem 1: Continuity of linear operator at x = 0 implies continuity over all of DA.
Theorem 2: For a linear operator, boundedness continuity (p140)
{ Theorem 3. In En, every linear operator is bounded (= continuous) }
Theorem 10: A continuous(=bounded) operator on a closed domain is a closed operator.
Corollary: A continuous(=bounded) operator on all of H is a closed operator.
Famous Theorem 11: A closed operator on a closed domain is continuous (=bounded).
Theorem 12: There can be no closed unbounded operators defined over a whole Hilbert Space H.
Theorem 13: The nullspace of a closed operator is a closed set.
Here then are some new facts concerning closed operators and the inverse A-1:
(a) If A is closed and A-1 exists, then A-1 is also closed.
(b) If A is closed and A-1 is bounded, then RA is closed.
Contrapositive: If A is closed and RA is not closed, then A-1 is unbounded.
Apply to A-1: If A-1 is closed and A is bounded, then DA is closed
(c) If A is closed and DA is closed (normal situation I think), then A is bounded.
Apply to A-1: If A-1 is closed and RA is closed, then A-1 is bounded.
A-1 is bounded and RA is closed
I have tried to encapsulate all these "facts" in the following drawing"
The outer solid oval encompasses all linear operators A. Those on the left have DA = closed, those on the right have open domains. The odd-shaped boundary contains the closed operators, and then the inner solid circle shows the location of bounded = continuous operators. You can see that on the left side, these two regions are the same. The dotted ovals have to be interpreted according to their location.. This is perhaps a tough chart to read, but it helps me to try at least to visualize the classification. I have tried to show where differentiation and integration fit in. The Hilbert Space H is meant to be arbitrary here. Now recall from just above:
Regular requires:
(a) Ax=0 x=0, as with matrices; and this implies that A-1 exists;
(b) the range must be the entire Hilbert space, so Ax = f has a solution for any f inH;
(c) A-1 must be bounded (= continuous).
Classifying Closed Operators.
The picture shows that these four cases partition the space of closed operators.
The first case is "regular", the other three are singular:
(1) A is regular (meets the three criteria listed above) [green oval ]
(2) A-1 does not exist, nontrivial Ax=0; [ violates criterion (a) ] [within blue but outside purple]
(3) A-1 exists but A-1 is unbounded [violates criterion (c) and (b)] [within purple but outside yellow]
In this case, we know from our drawing that RA ≠ A = H
(4) A-1 exists but ≠ H. [ violates criterion (b) ] [within yellow, but outside green]
I have tried to show the four cases as well on the chart above. Integration for example is singular in case 3 because its inverse, differentiation, is unbounded. It is really the inverse of an operator that determines its niceness regarding solving our Ax = f equation! We think if integration as nice, but it has the nasty inverse which is unbounded albeit closed. I think differentiation, in contrast, is regular! Unfortunately, this example was omitted on page 168-169 (but appears in spades starting page 174).
Richard Price Question comment added 3/15/11. Suppose A is the Love equation integration operator with its particular kernel. That integration has fixed endpoints and therefore does not seem to fit into the current Stakgold discussion. The question is: which of the four classes above does this operator A fit into? I suspect it is not invertible and that is why you won't be able to take Ax=f and get x = A-1f as some kind of differential equation. But this question is still open, maybe in Stak Chap 3 we can come back to it. I just want to note it as an "open question" here in the midst of the general operator discussion.
Adjoint operator stuff (A*)
Preliminary Factoid: (S ) = . From a "perp" point of view, we have S = 0 , so that a point on the boundary of an open set S is still perp to S . Therefore, S = 0 => (S ) = . This factoid is used on the last line on page 171 with little comment and causes the bar we are about to mention.
First, a reminder as to why we care about the adjoint operator A*. In the ∞ dim HS world, the Alternative Theorem takes this form (p 171): (the new element here is the bar over RA)
Alternative Theorem:
(similar to page 153 for En): (RA ) = NA* so that by factoid, = (NA*)
Corollary: If RA is a closed set, the above says RA = (NA*) which duplicates our matrix version.
So the reason we care about the adjoint is this fact just quoted: = (NA*) . If we can find the nullspace of this adjoint operator A*, we learn something about the range of A, and thus we learn for which vectors y we might expect to have a solution of Ax = y.
Complication for unbounded operators regarding the adjoint.
If A is unbounded, it is not continuous, and Riesz does not apply to functionals like T[x] = <Ax,y> so y may not exist! We want to define the adjoint of A according to <Ax,y> = <x, A*y>, where A*y = g is some vector in our Hilbert Space H. However, for unbounded A we can only state this equation (intended to be true for all x DA) for certain vectors y, and not for all vectors y in H. In other words, for certain vectors y, we won't be able to find g H such that <Ax,y> = <x, g> is true. Such a value of y is said to be inadmissible. If g exists such that <Ax,y> = <x,g>, then y is admissible and we write A*y = g. Therefore you can say that the space spanned by these admissible vectors y is the domain of A*, or DA*. One says that { y,g } form an admissible pair of y is admissible.
There is then an added technical detail. If DA is not dense in H, then g might not be unique. Suppose g1 ≠ g2 were both solutions, then we have to say A*y = g1 and A*y = g2 but we don't want to have an operator A* which maps a vector into multiple vectors in the range, so we cannot allow this. Stak shows that if DA is dense in H, then g will be unique. So we have to limit our interest to domains of A which have the property A = H .
So our conclusion is this about the adjoint for an unbounded operator A:
<Ax,y> = <x, A*y> when x DA where A = H, and y DA*
STak makes a point that DA* includes 0, so that y=0 is admissible.
So in order to even be able to talk about an A* operator for unbounded A, you have to assume that A = H and x must be in DA and y must be in DA* and you just hope that DA* is not an empty space.
Now let's jump ahead for a moment. Suppose we define
A is symmetric ≡ : <Ax,y> = <x, Ay> for all x,y in DA
If we assume that DA DA* , then for all y in DA we can have an A* object and we can write
<Ax,y> = <x, A*y>
So if A is symmetric, and if DA DA*, we can combine these two results and say
<x, Ay> = <x, A*y> for all x and y in DA (*)
If we don't assume that DA DA* , then A* won't exist for certain y in DA and we cannot write the above. We are trying to work toward a concept of A = A* as we had in the En world. Being able to write (*) is a major step in this direction.
Stak comments (page 172, Th2) that you can interpret the above discussion as saying
A is symmetric DA DA*
That is to say, based on our definition of "symmetric", an operator can only possibly be symmetric if
DA DA* , for only then do we know that a candidate y will an admissible value of y.
If you just add the extra condition that DA = DA* , then you get A being self-adjoint. But we still have that extra required fact that A = H, or we cannot even be talking about A*.
Now Stak says it is possible that you could have DA DA* and not have DA = DA* . In this case, A symmetric does not imply A self-adjoint.
Now suppose it happens that DA = H. Then DA DA* says H DA* and of course that implies that
DA* = H, no other choice. So in this case, we have DA = DA* ( = H ), and then symmetric A is also self-adjoint. So any symmetric A whose domain is the whole Hilbert space is self-adjoint. Here then is my picture:
At this point, we get a very heavy dose of unproven adjoint-related claims that I don't want to get sidetracked on unless one of these claims is later needed. Here they are ( A is assumed linear I presume)
(1) If A =H , then A* exists with some DA* containing 0 (just showed this above)
(2) A* is linear
(3) A* is closed ( regardless of whether A is closed), thus if A=A* self adjoint, A is closed.
(4) B A ( operator B is an extension of A) A* B*
(5) if A is bounded on all of H, then so is A*
(6) If (A*)* exists [ requires that A*= H ] , then A (A*)*
(7) If A is closable, then A* = ()* . [ The overbar means closure of operator A, not c.c. ]
(8) If A is closed, then [A =H A*= H ]] , so then A** exists.
Self-Adjoint Operators
An operator A is self-adjoint iff: (page 172)
(a) DA = DA* (which implies A = A* = H)
(b) Ax = A*x x DA = DA* .
Back in our En world, "self adjoint" was a matrix equality A = A* and DA = En = the domain of all operators in En, and two identical matrices of course implies (b).
<Ax,y> = <x,Ay> = <x,A*y> for all x and y in DA
Symmetric Operator Facts
We then get a cluster of theoremlets (page 172 just in the text)
Theorem 1: A is self-adjoint A is symmetric
Theorem 2: A* A ( A* is an extension of A) A is symmetric
Theorem 3: (DA = DA*) + A is symmetric A is self-adjoint => (3): Ax = A*x x DA = DA*
Theorem 4: If A symmetric on all of H, then A is self-adjoint. (this is the case in En)(see para above)
The first and third follow from my figure. I just accept the second. The third suggests that when the domains are the same, you don't really have to add the requirement that Ax = A*x on the domain. Four was already mentioned above.
EXAMPLES: This section is very long and should probably only appear in the raw notes, but I wrote these notes here because things were confusing so I leave them here. These examples are required if one is to have even the dimmest understanding how you find A* from A, and how you find RA, uses of the Alternative Theorem, restrictions and extensions, and so on . These are very tedious because each 5 word sentence requires brain work on the reader's part with concepts that have just been loaded into the reader's data banks. The concepts are very foreign from those a physics person usually encounters.
In each case, we want to know: (1) given the A shown, what is the adjoint operator A* ? (2) what are the domains and nullspaces of A and of A* ? (3) what does the Alternative Theorem tell us about RA given what we know about NA*? It gives the answer if the range is closed.
Example 1: [Ax](t) = tx(t) where H = L2(0,1). We find A* = A, self adjoint. <tx,y> = <x,ty> makes things pretty clear, since t is real.
Example 2: A = k1. If k is real, A is self-adjoint. Otherwise, get A* = 1 ≠ A. In this case, we are again saying <kx,y> = <x,ky> if k is real.
Example 3: [Ax](t) = !Syntax Error, Ids x(s) . We find [A*x](t) = !Syntax Error, Ids (s) so clearly A ≠ A* . If Ax=0, x = 0 because only the function 0 has indefinite integral 0! Thus, A has no nullspace. Similarly, A* also has no nullspace. What is RA ? It must be a function f(t) such that f '(t) = x(t). Thus, the range is not all of L2, but is the subset of L2 which is differentiable functions which Stakgold calls absolutely continuous functions. Range is thus not closed but is dense in L2. [ Since RA is not closed, we cannot use the alternative theorem to say RA = (NA*) = all of L2, which is consistent with what we just found. ] This operator is singular as case (3) since the inverse we know is unbounded.
Example 4: [Bx](t) = !Syntax Error, Ids x(s) – λ x(t) = (A - λI)(t), λ ≠ 0. In this case, B still has no nullspace, but now the RB is all of L2. In fact, this operator meets all the requirements for B to be regular: no nullspace so inverse exists; range is all of H; B-1 is bounded. I think this last item he proves explicitly by constructing the solution of Bx = f explicitly.
Example 5: [Ax](t) = dx/dt. This is the Mother of all Examples (p174-179), filling 5 full pages of the text! This is the one we are going to be VERY interested in for the rest of Stakgold's two volumes! A new idea to me (it shouldn't be!) is that we are going to regard boundary conditions like x(0) = 0 as restrictions on the domain, beyond the usual restriction to absolutely continuous functions. He is going to study 7 different boundary conditions, which each get a label like IV in the text! Each different BC along with A = d/dt really describes a different operator A!
(a) If there are no BC's, we have only condition I ( x is absolute continuous), and he calls the operator A1. He first shows that A1 is closed. He then asks what are the admissible pairs {y, g} such that <A1x,y> = <x,g=A1*y>. He shows that in order for y to be admissible, it must satisfy condition II which is that y(0) = y(1) = 0. In this case, the adjoint operator turns out to be A1*y = g = – dy/dt and of course this minus sign arises from a parts integration and the condition II makes the parts go away. Obviously A1 ≠ A1* . So the domain of A1* is conditions I + II. A1 has a nullspace since x(t) = C is in it. But A1* has no nullspace because y(t) = C is ruled out by the boundary conditions. Range of A1 is all of L2 hence closed. The alternative theorem says in this case that RA1 = (NA1*), and since (NA1*) = 0 only, we conclude that the RA1 = all of L2. If we find a solution of Ax = f for some f, we can always add to it a solution of Ax = 0 which in this case is x = C. Thus we get x(t) = !Syntax Error, Ids f(s) + C
(b) Add a domain restriction x(0) = 0, condition III, call the operator A2 . We find A2 is closed, as was A1. In looking for A2*, we again ask what are the admissible pairs. This time due to III, we only need y to satisfy y(1) = 0 (condition IV) to make the parts vanish, and again A2*y = g = – dy/dt . With conditions III and IV, this time we find that the nullspaces of both A2 and A2* are empty, RA2 is again closed, and we again conclude that RA2 = all of L2 by the same argument as in (a) above. In this case, for given f we get the unique solution x(t) = !Syntax Error, Ids f(s) ; there is no C to add since A2 has no nullspace. [ The only possibility for having a nullspace is if x = C or y=C can survive the boundary conditions. ]
(c) This time start again with A1 but add the domain restriction x(0) = x(1) [ condition V], call the operator A3 which yet again is closed. In order to make those parts vanish this time, we need the condition y(0) = y(1) which is the same as condition V, so DA3 = DA3*. Both A3 and A3* have some nullspace since x=C is legal. Given this nullspace, what is (NA*) ? Well, we need <x(t), C> = 0, and that means that !Syntax Error, Idt x(t) = 0. So functions with zero integral are in (NA*). The range of A is closed, so we get to use the alternative theorem without the bar, RA = (NA*), and we thus learn that for A3, the range is functions whose integral is zero! x(t) = !Syntax Error, Ids f(s) + C is the solution here.
(d) Start again with A1 but add the domain restriction x(0) = 0 and x(1) = 0 [ condition VI] and call the operator A4. This condition is really a special case of condition V, but it alone now makes the parts vanish, so DA* has no conditions on it other than condition I and A4* = - d/dt as in all previous cases. Now the equation A4*y = 0 has y= C as its non-trivial solution while NA4 = 0 due to VI. As in example (c), we find (NA*) consists of functions such that !Syntax Error, Idt y(t) = 0 [ the "consistency condition" ] And since NA4 = 0, we don't add a constant, so solution is x(t) = !Syntax Error, Ids f(s).
(e) This time our condition is x(0) = 0 and x'(0) = 0 [ condition VII ] and call it A5 . When Stak writes this sequence A4 A2 A1 he is saying that we have been restricting A1 more and more with our boundary conditions. A1 had none, A2 had x(0) = 0, and A4 had both x(0) = 0 and x(1) = 0. That is, we are restricting the domain more and more. And the domain of the A*'s gets more extended at the same time. Clearly A5 A2 because we added x'(0) = 0. But in this case, it will turn out that instead of the expected A2* A5* , we will get A2* = A5* (he proves this at length). This fact tells us A2* has no nullspace, and we can see that A2 also has no nullspace. He claims that the RA5 is not all of L2 so it must be that this range is not closed ( else alt theorem would say range = all of L2 since nullspace NA2* empty, as usual in many of the above examples). It turns out that the range condition is f(0) = 0.
2.9 The spectrum of an operator ( pp 180-184 )
The first step is to review the "classification" of closed operators into one regular and three singular cases, we copy from above, replacing A with B:
(1) B is regular, meaning B-1 exists, = H , and B-1 is bounded
(2) B-1 does not exist, nontrivial Bx=0;
(3) B-1 exists and B-1 is unbounded and RB H but = H
(4) B-1 exists and RB does not close to H so there are points in H which lie outside our range.
Next, we focus now specifically on the operator
B ≡ (A - I)
where A is a closed linear operator on a domain DA that is dense in H. We imagine that as we let λ vary over the complex plane, we might "hit on" any of the above cases. The set of λ values that gives each of the above cases has a name. Although we might be thinking about operator B, these names refer to the operator A. Notice that the word "spectrum" applies only to the three singular cases here. So you can say that { set of all values of λ } = { resolvent set | the spectrum }.
Point, continuous and residual spectra; the resolvent set.
(1) the resolvent set of A ( B is regular)
(2) the point spectrum of A (λi are the eigenvalues, B has some nullspace, B-1 does not exist. )
(3) the continuous spectrum of A (B-1 unbounded, RB dense in H)
(4) the residual spectrum of A (H – RB = (RB ) has dimension called the deficiency of λ. )
In En, either an operator B is regular, or it is singular and B-1 does not exist (det B = 0). These are the first two cases above. The latter two cases arise only in ∞ dim HS problems.
We now get four heavy duty theorems:
Theorem 1: If is in the residual spectrum of A, then is an eigenvalue of A* which has geometric multiplicity m ≡ the deficiency of .
The proof is stated in 3 lines, I follow it, but it omits this logic which I now add: B* has a nullspace of dimension m (= number of indep eigenvectors of A*), so nullity ν(B*) = m. The Alt Theorem says
(RB ) = NB* so the dimension of (RB ) is also m. But this is, by definition, the deficiency of this point λ. Note that the number m as first defined here is the geometric multiplicity of A*, and that for our ∞ dim HS stuff, there is no meaning to the term algebraic multiplicity, so we just say "multiplicity".
Theorem 2: If A is symmetric, then the eigenvalues of A are real, and <x,Ax> is real for x in DA. (simple proof, same as in matrix world).
Corollary 12A: From Theorem 1, if λ is in the residual spectrum of A, then is an eigenvalue of A*. For a self-adjoint case A = A*, eigenvalues of each operator are the same and real, so if λ is in the residual spectrum, it must be an eigenvalue of A. But obviously such a point is in the point spectrum of A, so we have a contradiction. Therefore, for a self-adjoint A, the residual spectrum is empty.
At this point, we get two interesting lemmas which are used to prove theorem 3.
Lemma 1 says: If B-1 is unbounded, you can construct a sequence xn in DB of unit norm whose image sequence in RB is a null sequence, and in particular, such that || Bxn|| < 1/n. If B-1 were bounded, we know that such a range null sequence would have to be null in the domain, so in that case this lemma could not apply. It seems very reasonable.
Lemma 2 says ||Bx||2 ≥ Im(λ)2 ||x||2 which was easy to show using the CSI.
Theorem 3: The continuous spectrum of a symmetric operator A is limited to the real axis.
The proof is readable and quite peculiar. If η ≠ 0, then for all norm-1 sequences xn we have ||Bxn|| ≥ |η| according to Lemma 2. This means there can be no range null sequence as postulated in Lemma 1 (unless |η| = 0). So if η ≠ 0, then B-1 must be bounded, but in this case λ is not in the continuous spectrum by definition of same! So if λ is in the continuous spectrum, we must have η = 0 meaning λ = real.
Theorem 4: The entire spectrum of a self-adjoint operator lies on the real axis, and there is no residual spectrum, meaning that case 4 never occurs.
This is just Theorem 3 plus our Corollary 12A above.
EXAMPLES.
Example 1: A = integral operator. In our Example 4 above, we showed that in this case B = regular unless λ = 0, so λ = 0 is the only point in the spectrum of this A. In the raw notes, we show that λ is not an eigenvalue because nullspace of A is empty: Ie, Ax = 0x (λ=0) only has solution x = 0, no eigenvector here. And since A* also has no nullspace, we know (RA ) = null so λ cannot be in the residual spectrum. Therefore, it must be in the continuous spectrum.
We are now going to have a reprise of the various versions A1,2,3,4 of the d/dt operator, as we studied above in Example 5.
Example 2(a). A1 (no other BC's). Since dx/dt = λx has solution eλt for every λ, then every complex or real λ is an eigenvalue! Fascinating. Here the point spectrum is the entire complex λ plane, and other two spectra sets are empty, and so is the resolvent set. So here the point spectrum is actually a continuous set.
Example 2(b). A2 = d/dt with x(0) = 0. For this problem, since x(t) = C et there are no eigenvalues, so no point spectrum. Using the usual Green's method (remember that the Green's function respects the BC's) , equation (A - I)x = f has a solution for any f, so there can be no residual spectrum (full range). And it turns out that (A - I)-1 is bounded which means no continuous spectrum. [Integral operator is bounded.] Conclusion: the entire spectrum is empty! That means that every is regular.
Example 2(c), A3 = d/dt with x(0) = x(1). Point spectrum is = 2in n = 0,1,2,... (or negative). Since different BC, the Green's is now different, but we still get the full range result ( so no residual spectrum) and (A - I)-1 is bounded (so no continuous spectrum) , so A3 has ONLY this point spectrum. Other are regular and so are in the resolvent set.
Example 2(d). A = d/dt with x(0) = 0 = x(1). As in (b), no eigenvalues so no point spectrum. New situation is that now not all f of Ax = f are allowed, so we can have a residual spectrum. The adjoint operator A* allows y = C in it's domain ( see example 5.d), so the range of A is functions of 0 integral. But also the adjoint of B which is B* = A* - I, which has nullspace solutions exp(-t), so that the range of B is then functions which are orthogonal to this expo which eliminates a lot of functions from the range of B! Since B-1 is bounded, there is no continuous spectrum. We don't have B = H I guess, so there is no resolvent set of regular λ values. Turns out all are in residual spectrum.
2.10 Completely Continuous Operators. ( pp 184-187 )
definition: A set S of elements in a HS is bounded if ||x|| < c for all x in S.
Theorem: Set S is finite S is compact (since any sequence must have a convergent subsequence )
Theorem: set is compact set is bounded.
This is not obvious to me from the above compact definition, and Stak does not prove it. Obviously true in a finite-dim space where compact = closed + bounded. A contrapositive proof is no doubt best. I would have to show that unbounded implies there is a sequence without a convergent subsequence. If a set is unbounded, I just make a sequence that marches off forever in one of the directions in which it is unbounded. Think for example of the positive integers. Find a convergent subsequence! There is none, it just keeps on going forever without converging. Therefore not bounded => not compact, so compact must imply bounded.
In some of Stakgold's statements on page 185 top, I think he assumes that sets are closed. For example, the BW theorem says that a subset of Rn is sequentially compact iff it is bounded and closed. This caused me much confusion for a while.
Theorem: set is bounded ≠> set is compact.
There is a very standard example always given here. The φn are regarded as some orthonormal functions in a sequence, such as the normed Legendres on (-1,1). Since orthonormal, || φn || = 1, so the set of φn is clearly bounded. Riesz-Nagy point out that ||φn- φm|| = 2 for n ≠ m since orthonormal, so pretty clear that there can be no convergent subsequence here! So this set of φn is bounded but not compact. This does seem to require an infinite dimensional Hilbert Space. In a finite dimensional space, we know that bounded + closed compact. In a sense you might argue that the set of φn here is not "closed", but that is not really the right language. It is not a problem with the limit of a convergent sequence being missing from our set (being on the boundary), it is that there IS no limit.
Theorem: Operator A is bounded A transforms bounded sets into bounded sets
Theorem: Operator A is continuous A transforms bounded sets into bounded sets
Same theorem, and I proved it both ways in the raw notes.
Definition: Operator A is completely continuous A transforms bounded sets into compact sets.
Example: The identity operator is not completely continuous. It maps our φn set mentioned above into itself. The φn set was bounded but not compact.
Example : We will see that Hilbert Schmidt integral operators are completely continuous, which is why we care.
Theorem 1. If A is completely continuous and n is infinite orthonormal sequence, then An 0.
So here is our old friend the φn sequence again. We know from the definition that the image sequence has to be compact so has to have I think an accumulation point. But this theorem which appeals to Riemann-Lebesgue, says that the image has such an accumulation point at 0 and converges there.
Implication of theorem 1: certain non-convergent sequences (like φn above) get mapped into convergent ones. But that is not good for the inverse operator! It now maps a convergent sequence into a non-convergent one!
Theorem 2: If A is completely continuous and A-1 exists, then A-1 is unbounded if in dim space.
The proof of this one is easy. We just consider our usual sequence φn and we end up with An 0, so with respect to A-1 we have boundedness set by || φn || / || Aφn || = 1/|| Aφn || → ∞, so A-1 is unbounded.
Theorem 3: Consider An A. If the An are completely continuous, so is A (some restriction).
This has a long 1/2 page proof which I will skip for now (1.18.09) . An A means we are approximating in the norm so that || An - A || → 0. The R-N book also discusses this theorem. I think Stakgold (1967) just got all this stuff from their book (1955 English).
Comment: This subject (completely continuous) is mentioned in my Riesz-Nagy book on page 177. Reisz says that he invented the idea in 1917 (reference 9)
Comment: This section on completely continuous operators was very painful for me, largely due to the word "closed" being omitted in several places. Also, we don't have a single example of a c.c. operator yet, and even the identity operator is not one! I am sure he will reference this section later in the integral equations chapter.
2.11 Extremal Properties for Bounded Operators ( pp 187 - 190)
This section is a set of six theorems most of which have simple proofs that I review in the raw notes.
Theorem 1: If A is bounded and DA = H , then || A || = || A* || .
Theorem 2: If A is bounded and DA = H , then the ratio |<Ax,x>|/||x||2 ||A||, for all x 0.
This ratio I called I4 in the En earlier section of these notes. And I said |||A||| = max I4 over all x. So this theorem says, in my language: |||A||| ≤ ||A|| . So for a bounded operator on the full domain, this triple bar norm is in general ≤ the official double bar norm.
Corollary: Obviously, if MA is the largest that |<Ax,x>|/||x||2 can go as you vary x, then MA ||A||. So MA is a different notation for my |||A||| "norm" thing.
Theorem 3: If is in any of the three kinds of spectra for A, then | | || A ||.
This generalizes a result we had in the En section with a similar conclusion.
Theorem 4: If A is symmetric, MA = ||A||. { that is to say, |||A||| = ||A|| }
Theorem 5: If A is symmetric, there exists a domain sequence xk with ||xk|| = 1 which, in the limit, would become an eigenvector of Ax = λ1x where λ1 = ± ||A||. BUT, this theorem allows that the limit itself might not exist, so theorem is stated as a lim Axk → lim λ1xk.
Theorem 6: If A is symmetric and completely continuous, then: ( p 190 )
(1) At least one of these numbers is an eigenvalue: ||A|| or –||A||.
(2) There is no eigenvalue μ for which |μ| > ||A||.
What exactly is this saying? Suppose μ1 is the eigenvalue having the largest absolute value of all the eigenvalues. We don't know the sign of μ1. But we know that | μ1| = ||A|| .
the largest eigenvalue is || A ||.
The proof is simple. By adding the completely continuous condition we are then able to show that the limit in the previous theorem exists, and then λ1 really is an eigenvalue and xk → x converges.