stakgold chap 2
DOCX · 417.7 KB
Open DOCX file
Phil's chapter-by-chapter notes on Stakgold, first written 11.16.04 and updated 1.18.09, with a detailed table of contents and commentary. They cover functions and transformations, linear spaces, metric, normed and inner product spaces, Cauchy sequences and completeness, separable Hilbert spaces, bases and Gram-Schmidt, functionals, and operators in finite and infinite dimensions. Later sections cover spectra, completely continuous operators and extremal properties.
AI-written summary; may contain errors.
Extracted text (machine-read; may contain errors)
Stakgold Chapter 2 Raw Notes PhL 11.16.04
updated 1.18.09
See separate file with meta notes on Chapter 2
Chapter 2: Introduction to Linear Spaces 2
2.1 Functions and Transformations (Operators). ( pp 92-96 ) 2
Real Function of a Real Variable (92) 2
General Transformations (94) 2
2.2 Linear Spaces (= Vector Spaces). ( pp 96-99 ) 3
Dependence and Independence of Vectors (97) 3
Dimension of a Vector Space (97) 3
Examples of Vector Spaces: 3
2.3 Metric Spaces, Normed Linear Spaces, Inner Product Spaces ( pp 99-116 ) 3
Metric Spaces (99) 3
Examples of metric spaces: (102) 5
Normed Linear Spaces (105) -- including Banach Spaces 5
Inner Product Spaces (107) -- including Hilbert Spaces 6
Examples of Inner Product / Hilbert spaces. (110) 8
Exercises 2.2 through 2.9 9
Additional Properties of Metric Spaces (114) 9
Dense Sets (114) 9
2.4 Properties of Separable Hilbert Spaces. ( pp 116-135 ) 10
Dependence and Independence of Vectors: notion of a Basis (118) 10
Orthogonality (120) 12
Linear Manifolds ( ≡ Subspaces) (120) 12
Projections (121) 13
Doing GSO on an independent set (122) 13
Orthonormal Basis (123) 14
Characterization of an Orthonormal Basis (127) 16
Remarks on Orthonormal and Orthogonal Bases in L2. (129) 16
Exercises 2.11 through 2.25. 17
2.5 Functionals. ( pp 135-139 ) 17
2.6 Transformations (= Operators) ( pp 139-146 ) 20
Extensions (141) 20
Closed Operators. 21
2.7 Operators in the Hilbert Space En(c) ( pp 146-165 ) Finite Dimension Hilbert Spaces 22
Introduction + The Matrix of a Linear Transformation on En(c) (146) 23
Examples of Linear Transformations (148) 23
The Two Roles of Matrices (149) 23
Simple Calculations with Matrices (150) 23
The Inverse of a Linear Transformation on En(c) (151) 23
The Adjoint A* of a Linear Transformation A on En(c) ( 152) 24
The Alternative Theorem (153) 24
Eigenvalues and Eigenvectors (154) 26
Operators with n Distinct Eigenvalues (159) 27
Symmetric Operators (159) 28
Extremal principles for Symmetric Operators (162) 30
Exercises 2.31 through 2.39 31
2.8 The inverse of an linear operator ( pp 165-180 ) Infinite Dimension Hilbert Spaces 31
Classification of Singular Operators (167) 32
Examples of operators on L2 (168) 33
The Adjoint Operator (p 170). 34
Examples ( 8 pages worth!) 37
2.9 The spectrum of an operator ( pp 180-184 ) 41
2.10 Completely Continuous Operators. 43
2.11 Extremal Properties for Bounded Operators ( pp 187 - 190) 46
Chapter 2: Introduction to Linear Spaces
This chapter is about 1/4th of the entire book, pages 92-190, about 100 pages of dense
"math theory" chock full of "theorem thickets". It contains stuff I studied in Math 21, Math 105, and then Math 224 at Berkeley. Oddly, Stakgold does not use the nice mapping notation A: X Y. He is a combination math and engineering guy, so will try to be mathematically rigorous and at the same time will try to link things to the real world with examples.
2.1 Functions and Transformations (Operators). ( pp 92-96 )
Real Function of a Real Variable (92)
Uses word function for f and f(t). Notion of domain and range. y = f(x) is f: R→R. Should include the domain like (a,b) as part of your definition of a function. Examples are given.
General Transformations (94)
Instead of mapping f: R→R, we have f: X→Y where X and Y are just some sets. He calls this a general transformation, not necessarily linear. Notions of function, domain, range, one-to-one, onto, into. Generalization in terms of an operator = transformation = mapping. Notice that X and Y do not in general have to be vector spaces or metric spaces or anything like that, just arbitrary sets. BUT, usually (he says) X and Y will have the "algebraic structure" of a vector space (= a linear space) so we can talk (as in the next paragraph) about adding vectors in X or Y.
Examples of transformations are given. f: R→R again, f: Rn→Rn n-tuples, f: {functions} → R (called a functional), f: {functions} → {functions} [ f could be d/dx ]
Notes added:
(1) Consider Ax = y and A:XY with DA and RA. Obviously, each x lands in only one y = Ax. But each y maps back to many x when there is a non-zero nullspace. That is because any x+x' y where x' is in the nullspace. If A-1 exists (A has full rank), then each y maps back to only one x = A-1y. In this case, we are 1-to-1 in both directions (injective). If RA = Y, we are onto (surjective). If RA Y, we are into.
And if both 1-to-1 and onto all Y, then a bijection and inverse then exists. If f and f-1 are both continuous, then f is called a homeomorphism, not to be confused with a homomorphism which is two equivalent worlds. [ continuous suggests to me smooth, analytic, reasonable, no surprises ] .
(2) A linear space has algebraic structure from its definition. Algebraic means doing algebra, adding things, multiplying by constants, etc. It can also have metric (topological) structure if the notion of distance is somehow added (as we shall see below).
2.2 Linear Spaces (= Vector Spaces). ( pp 96-99 )
A linear space is the same thing as a vector space. You have elements of the space and a + operation that commutes and associates, and there is a unique 0 and for every X, there is a -X. On top of this, you have a * operator defined only on an element times a real or complex number. The * operation associates, distributes both ways. At this point, there is no notion of distance between X and Y elements.
Dependence and Independence of Vectors (97)
You can obviously make linear combinations in this vector space. A set of vectors is linearly dependent if at least one can be expressed as a linear combination of the others. Otherwise independent.
Dimension of a Vector Space (97)
The maximal number of independent vectors possible is the dimension of the space. Any set of N independent vectors forms a basis for an N dimensional vector space.
The elements of a vector space can be discrete things like N-tuple vectors, or can be functions of some particular type.
Important point I omitted in first pass. In an N dimensional vector space V, in order for a set of vectors ek to be a basis, they must be linearly independent AND you must be able to represent any element x V as a linear combination of the ek with some well-defined coefficients k . See notes below on spanning set versus basis.
Two theorems here: (1) When you represent a vector projected onto a basis, that representation is unique, ie, there are unique coefficients. (2) In any finite-dim n vector space, any set of n independent vectors forms a basis.
Examples of Vector Spaces:
n-tuples of reals; space of real functions on (a,b); space of polynomials on (a,b); complex values functions on same, as long as norm exists (does not diverge), etc.
2.3 Metric Spaces, Normed Linear Spaces, Inner Product Spaces ( pp 99-116 )
Metric Spaces (99)
A metric space is a set of "points" (not necessarily vectors) with a distance function d(x,y) [ "the metric" ] defined with the usual three basic rules (p 100). Non-negative, only 0 if x = y, triangle rule. You can have a metric space of objects which do not form a vector space! So right now, metric and vector are completely separate concepts to think about. Notice that you do not use the word "linear" with metric space, unless you mean to also imply that the space is a vector space. Linearity relates to the vector aspect, and has nothing to do with the metric aspect.
Reminder of our context here: Earlier we were talking "vector spaces" with algebra stuff. Here we are talking "metric spaces" with a metric d(x,y). It is only in this metric space world that one can even discuss the general concept of "convergence" because that word implies that some distance is approaching 0, and distance requires that you have a metric.
The two tools used here are the normal convergent sequence d(xm–x) → 0 and the so-called Cauchy convergent sequence idea that d(xm–xn) → 0 for indices large enough. The Cauchy thing allows that maybe the limit point x is not in the space. If limit is always there, metric space is complete. You can make an incomplete space be complete by adding all those missing points.
Theorem (p100) A normally convergent sequence is a also Cauchy convergent sequence.
Augustin Louis Cauchy (1789-1857) was a French mathematician. He started the project of formulating and proving the theorems of infinitesimal calculus in a rigorous manner and was thus an early pioneer of analysis. His major results were published in the timer frame 1820-1828. He wrote 789 papers, and His collected works, Œuvres complètes d'Augustin Cauchy, are published in 27 volumes !!! These writings covered notable topics including: the theory of series, where he developed with perspicuous skill the notion of convergence; the theory of numbers and complex quantities; the theory of groups and substitutions; and the theory of functions, differential equations, and determinants. He clarified the principles of the calculus by developing them with the aid of limits and continuity, and was the first to prove Taylor's theorem rigorously, establishing his well-known form of the remainder. He also contributed significant research in mechanics, substituting the notion of the continuity of geometrical displacements for the principle of the continuity of matter. In optics, he developed the wave theory, and his name is associated with the simple dispersion formula. In elasticity, he originated the theory of stress, and his results are nearly as valuable as those of Simeon Poisson. A pretty impressive legacy.
Other significant contributions include being the first to prove the Fermat polygonal number theorem. Cauchy created the residue theorem, used it to derive a whole host of interesting series and integral formulas, and was the first to define complex numbers as pairs of real numbers.
A metric space is complete only if every Cauchy sequence is also a convergent sequence. Cauchy is the double limit idea that d(xm,xn) 0 for large m,n but convergent means d(xm,x) 0. You can make a metric space complete by adding the missing points. For example, the set of rational numbers is not complete, but the set of real numbers is complete.
Example: Page 101-102 gives a fascinating proof that the sequence formed from the series which adds up to e fails to converge to a rational number, so e is not a rational number. The difference between two such sequences is a Cauchy sequence, but the limit is missing. If you assume that the limit is rational, the logic leads to 0 < integer < 3/4!
Examples of metric spaces: (102)
abs value, squared component sum, max diff of two functions, etc. The usual d1 and d2 metrics are defined for functions being objects in the metric space. Examples of complete and non-complete function spaces are given, summary comments here about Lebesgue vs Riemann integrability which I skip for now. [ See Royster notes for other metric examples. ]
Example 5 on page 103 talks about C(a,b), the set of all real continuous functions on (a,b) as an example of a metric space: the L1 absolute value norm max thing is set up as the metric d1. It is claimed that if you assume that a Cauchy sequence converges on (a,b), that is the same notion as "uniform convergence" over (a,b). Then a theorem is quoted but not proven here which says that any uniformly convergent sequence of C functions → a C function, and thus this space C(a,b) is complete.
Example 6 is the same C(a,b) idea but with the L2 metric. In this case, it is easy to create a Cauchy sequence of functions which has no limit in the space so the space is not complete. The catch here is that we look at the sequence shown in the figure on page 20, and the limit is a step function which is NOT continuous, so this limit is missing from the space. Maybe call this space C2(a,b) and the previous C1.
Example 7: comparison: rationals to reals is like continuous to just regular functions. This example introduces the important space of functions called L2(a,b). They can be real or complex with appropriate superscript and appropriate square integrability. One uses the Lebesgue instead of the Riemann sense of integration (more types of functions are allowed). This space of functions includes weird functions which are limits of things, as in our example 6 above. So L2 means the space of functions which are Lebesgue square integrable, and that is what the letter L stands for. Stakgold discusses some facts about L versus R integration, and encourages the such as perhaps changing the order of a limit and an integration as we see in the next paragraph.
We then get something called the Lebesgue Convergence Theorem. If functions xn(t) converge to function x(t), where we are allowed to ignore spikes maybe "of zero measure", and if there exists some f(t) that bounds the sequence in that | xn(t) | < f(t), then lim [ ∫ dt xn(t) ] → ∫dt [lim xn(t)] = ∫dt x(t). This is just the interchange of limits that wiki mentions. This is more rigor that I want right now, and I always assume that we can do such order interchange, but it is good of S to at least mention this subject.
Henri Lebesgue (1875-1941) : The Lebesgue integral plays an important role in the branch of mathematics called real analysis and in many other fields in the mathematical sciences. Lebesgue's integration theory was originally published in his dissertation, Intégrale, longueur, aire ("Integral, length, area"), at the University of Nancy in 1902.
Georg Friedrich Bernhard Riemann (1826-1866) was a German mathematician who made important contributions to analysis and differential geometry, some of them paving the way for the later development of general relativity. Here we are thinking of his integration work where the integral is the limit of those thin vertical strips we all learn about in calculus class, perhaps this was done in 1850 or so, you see he only lived 40 years. His main work was on higher dimensional spaces and their tensors as in gen rel.
Normed Linear Spaces (105) -- including Banach Spaces
Now go back to the linear (=vector) space idea and add to it a norm, to get a normed linear space. This norm has the usual three rules. The norm is the "length" of a vector and is written as usual || x ||. You then have a normed linear space. But, if you then define this metric: d(x,y) = norm(x-y), your space is then also a metric space. This is the natural metric generated by the norm.
Page 106 item 4: he comes up with a funny metric on the reals which, when you try ||x|| = d(x,0), does not give a valid norm. But I think if you start with a valid norm, the metric d(x,y) = ||x-y|| will always be a valid metric, and this is "the natural metric". He gives a rule in C which, if your metric satisfies it, they the norm it implies is valid.
This combination normed linear space + metric space with the natural metric is called a Banach space, provided the metric space aspect is complete (a notion fully discussed above, you add any missing limit points for Cauchy sequences). Other metrics are of course possible for a normed linear space, but then you would not call it a Banach space.
Stefan Banach (1892–1945) was a Polish mathematician who worked in interwar Poland and in Soviet Ukraine. A self-taught mathematics prodigy, Banach was the founder of modern functional analysis and a founder of the Lwów School of Mathematics. Among his most prominent achievements was the 1932 book, Théorie des opérations linéaires (Theory of Linear Operations), the first monograph on the general theory of linear-metric space. This is where he introduced what we call the Banach Space, 1932, quite recent!
Inner Product Spaces (107) -- including Hilbert Spaces
Now go back to the raw vector space again, and this time define an inner product of the vectors that meets the four rules shown on page 107.
In the vector and metric spaces discussed above, we had no concept of sort of "multiplying" two vectors together. You could multiply a vector by a scalar, and you had a metric d: (X,X) → R. The inner product is another functional d: (X,X) → R.
The inner product rules imply the Cauchy-Schwarz Inequality (CSI), as proven right there, a very simple proof. [ The CSI is something in the framework of an inner product space! ]
// think of ≤ 1.
Cauchy-Schwarz Inequality for sums was published by Augustin Cauchy (1821), while the corresponding inequality for integrals was first stated by Viktor Yakovlevich Bunyakovsky (1859) and rediscovered by Hermann Amandus Schwarz (1888) (often misspelled "Schwartz"). The word "sums" applies here if you think of <x,x> = Σi xi2 where xi are components of the vector x. See above for more Cauchy comments, this is just one of the many things he did in the time frame 1820-1928, more below.
Note Added. Think |ab| ≤ ||a|| ||b|| or <a,b>2 ≤ <a,a> <b,b> as the CSI. In L2 this says that
|∫a*(x)b(x)dx |2 ≤ ∫dx |a(x)|2 * ∫dx |b(x)|2 which we can write as || ab ||2 ≤ ||a||2 ||b||2 where the subscript 2 is because we are using the L2 norm. A more general theorem says || ab ||P ≤ ||a||p ||b||r where 1/P = 1/p + 1/r where ||a||p refers to the Lp norm (written Lp usually). This more general theorem is called Holder's Theorem and appears in my Ted Dushane notes and on the web:
So one aspect of this last is that if f and g are L2 , then fg is L2 ! I guess I never realized that simplef fact.
If you then define ||x|| = +sqrt( <x,x>) you can show that ||x|| is in fact a norm, so you then have a normed linear space (with the natural norm), and you can go on then to use the natural metric to get combination linear and metric space. You could call this thing the natural norm of an inner product space.
For an inner product space you can define the notion of two vectors being orthogonal <x,y>=0, you can define the idea of a line y, and you can uniquely decompose any vector into a piece lying along the line + another piece perpendicular to the line, all as in normal geometry, see picture page 109. You can go on to define the "angle" between two vectors in the usual way. Author is keeping general, here, not being specific about this being "geometry".
For a complex vector space, you bail out everything by replacing the inner product commute rule with this rule : <x,y> = [ Note that in this book, * is always used for Hermitian adjoint, † for me ]
Comment: I just noticed for the first time ever that the official definition has
<ax,y> = a<x,y> linear in the first argument!!!
<x,ay> = <x,y> anti-linear
which is exactly the reverse of quantum mechanics and normal matrix operations! I get confirmation with this nice wiki comment:
It seems strange that engineer Stakgold does not even mention that QM and normal Matrix theory use the reverse ordering!
Here is my take on the relationship between all these "spaces":
Examples of Inner Product / Hilbert spaces. (110)
1 (a) En on the reals (Euclidean space) // Note that all of 1 and 2 are finite dimension n.
1 (b) n-tuples, for me really the same as (a)
1 (c) Pn(a,b) polynomials of degree ≤ n
2 (a) En on the complexes (complex Euclidean space)
2 (b) complex n-tuples (complex Euclidean space)
3 The space L2(a,b)r which is real Lebesgue-square-integrable functions on real range (a,b)
4 The space L2(a,b)c which is complex Lebesgue-square-integrable functions on real range (a,b)
4 The space L2(V)c which is complex Lebesgue-square-integrable functions on volume V
In this last example, we have the notion of a Hilbert Space of functions being defined over some other Hilbert Space of function arguments, perhaps E3. Functions here would be called fields over E3 , but the word field does not appear in Stakgold. So don't confuse these two Hilbert Spaces:
Ω: L2(V)c → L2(V)c for example, Ω might be an operator like
f: V → V one point in V maps into an another, r' = f(r)
Stakgold does not point out this detail.
Exercises 2.2 through 2.9
Additional Properties of Metric Spaces (114)
P 114: Some more terms are now defined. If you add "the boundary" to a set of points in a metric space S, what you get is written and is called the closure of S. A point is on the boundary if it is the limit of some sequence in S. The space would be called a closed space. [ This is not the same as a set being "closed under addition", say ] { note that closure and sequence are metric space concepts! }
A set of points S is compact if every sequence contains a convergent subsequence. All finite sets are compact, so the question relates to infinite sets. The positive integers as S is non-compact for the following reason. Consider the sequence 1,2,3,4.... which is in S. Can you write down a convergent subsequence? No! Thus, this S is non-compact. A space that is not closed (not complete) can never be compact. In Rn, compact is that same as closed and bounded.
[ I have a whole chapter of David Royster notes on the subject of compactness, there are several equivalent definitions -- every cover has a finite subcover, for example. A compact set is "the next best thing to being finite". Open interval (0,1) is not compact, but closed interval is. The world of spaces with metrics is called topology, I took a whole course on this once with a great teacher at HU in Sever Hall. But I don't have any book or notes! Royster's notes on this are good. ]
Dense Sets (114)
Suppose you have two sets with S T. If you can get arbitrarily close to any point in T using points only in S, then S is dense in T. Classic example is S = rationals and T = reals. The rational numbers are dense in the reals.
(a) Another example of denseness: let S = set of polynomials of finite (but arbitrarily large) degree and T = set of functions. Assume the d1 metric. You can get arbitrarily close to any function in T with a polynomial of some degree, so S is dense in T. This is called the Weierstrass Approximation Theorem: "Any continuous (ie, non-spiky) function can be approximated by polynomials on a finite closed interval". He uses the L1 metric to show this, da = max |x-y| over interval. The word "uniform" is associated with this metric since we have a max over the entire range of the interval. See below.
(b) Yet another example. Consider T = continuous L2(a,b) with the d2 norm. Same thing, polys are dense in T. T is dense in the full L2.
(c) L2 is formally defined here as f(x) with d2 metric. T is contained in and is dense in L2. [ p 115 ]
The book does not mention this here, but the term uniform convergence refers to the convergence of a sequence that, in addition to the sequence index k, has some other parameter x, and we insist that we have convergence of the sequence for all values of the parameter x in some space of x elements. An example of course would be a sequence of functions fk(x) that converges to some f(x) for all x in (a,b). I think you are allowed to order-interchange a limit fn f with a Riemann integral if the functions fn(x) converge uniformly to some f(x) over (a,b). In the example (a) above, the difference function x(t)-y(t) → 0 uniformly over the interval in use there.
2.4 Properties of Separable Hilbert Spaces. ( pp 116-135 )
In this section, Stakgold repeats in a strange way many of the things he has already talked about above. As you read, you think "I thought he just said that same thing a while ago somewhere". Maybe he is just trying to make this section self-contained in some sense.
David Hilbert (1862-1943, German) was a very big gun, see wiki. Did physics math support among other things, for QM and gen rel, stated the 23 problems, some still unsolved. Godel friend, etc.
definition: "a spanning set" . Let H = some Hilbert space. Suppose we can find a countable set of vectors f1, f2 ... (such as xn) whose linear combinations form a space which is dense in H. If such a set can be found, then those f's are set to form a spanning set (which is not necessarily a basis, see later).
An immediate example of such a countable set would be the powers tn or some polynomials of degree n in a Hilbert space consisting of functions of a real variable going to same. Here n = 1,2,3.... is where the countable comes in. We have already seen how such polynomials are dense in L2.
definition: "separable Hilbert Space" : one which has at least one spanning set.
Obviously, any finite dimensional H is separable (has spanning sets). In fact, all Hilbert spaces which appear in this book are separable. [ I think if there were a simple example of a non-separable Hilbert space, Stakgold would quote it here. Web mentions a certain space of almost periodic functions which is one. So I don't think I will worry about separability. ]
Dependence and Independence of Vectors: notion of a Basis (118)
Detail on the spanning set: a spanning set f1, f2 ... means that you can get arbitrarily close to an element x in H by doing a finite sum x n akfk. As you require a lower , you in general need more terms, that is, you have to increase n. If, when you do this all the way to the dimensionality of the vector space N, the ak don't change for earlier terms in the series, then the fk also form a basis because then you have written x = nmax akfk . For N = finite, a spanning set is always a basis because once you reach N terms in the sum, you have your ak and you are done and the fk must be a basis. For N = infinite, it is easy to find cases where the ak don't stay put as you crunch down on your closeness requirement. I think, for example, that in L2 , the set of powers tk = fk forms a spanning set, but not a basis. However, any orthogonal set of spanning set functions fk (perhaps made by G-S) do always form a basis, and this is formally stated as Theorem 1 a few pages down in these notes.
Notice that, even if fk are not a basis, a spanning set spans a space that is dense in H since you can get arbitrarily close. Call this space S. If you add all the points that you cannot quite reach, that is the closure , and then you have = H. I think the closure of any space S that is dense in H is H.
On page 117 we have the two-column idea with En(C) on the left, and L2 on the right. Convergence in L2 is traditionally called convergence in the mean. [ Maybe convergence in L1 is crudely called uniform convergence over the interval. ]
On page 119 author discusses distinction between a basis and a spanning set. See notes above on this subject. A good example is that you cannot write |x(t)| as a power series in t over (-1,1), so the set of powers 1,t,t2 ... cannot form a basis in L2 . However, it does form a spanning set. It turns out that if you can find an orthonormal spanning set (such as Legendre polys), then that spanning set is also a basis. Once you have a basis with orthog, then you can write a formula for ak and you see at once that earlier coefficients don't change. With powers, there is no simple projection formula, you have to keep solving a set of equations, and all the ak move as you tick n to n+1 in your fit. [ Stakgold is not clear on this detail: you can of course get arbitrarily close to the function f(t) = | t | with a power series expansion, but as you try to get closer by adding more terms, the coefficients all "move" because you are re-solving the normal equations. Maybe the word "expansion" implies that the coefficients don't move. I tried to get Maple to do something here with a power expansion of | t | but found nothing easy. ]
The tn problem. I pondered this tn problem while running. You could of course define ak = ∫dt tk f(t) over some interval and compute all these ak. But it won't be true that f(t) = Σaktk. The actual fit will have some other coefficients f(t) = Σk=0,N bktk . The ak just defined of course "don't move", but the bk "do move". You have to compute bk to minimize ∫dt [ f(t) - Σbktk]2 which is a little calculus problem. As you increase N, the bk all change, but the ε does get crunched down as N increases. In this spanning set tk you can think of the set of {bk} as the "vector" for f(t) -- the components of the vector on the axes. The problem is that this vector does not approach a limit as N→∞, the components bk keep shifting forever. Thus, if you think of {bk} as a vector, you have a sequence of vectors {bk} which does not converge as N → ∞ ! The convergence failure is not due to a point being missing in the vector space; if it were, we could just add that point. It is a question of stability of the convergence as N→∞ . You could think for example of the sequence of numbers {b7} as N→∞ and this sequence of numbers fails to converge. I would think that maybe when N gets in the 1080 range, the amount of motion of this sequence would be getting very small, but that must not be true, the range of motion does not approach 0 or it would converge. This is a somewhat mysterious situation that I think a good computer simulation could clarify. You could write a Maple program to do this.
As an aside, suppose we tried to do our fit with f(t) = Σaktk. We would have
f(t) = Σkaktk = Σktk∫dt' t'k f(t') = ∫dt'f(t') { Σktk t'k }
But since the spanning set functions are not a basis, we cannot say { Σktk t'k } = δ(t-t') and verify our expansion (even if we were to normalize the powers). The usual term for such a delta rule is called closure or completeness, describing the nature of the spanning set functions. Notice that the words complete and closure are related in our discussion above, in terms of filling in missing pinholes in a space. However, these same words in the δ relation seem to have a related but different meaning. If your spanning set functions don't add up to a delta, they are in some sense not a compete set of functions, meaning they cannot be used to expand some functions like |t|. In fact, they are not a basis, and another word for basis we shall see below is a complete orthonormal set, so this is the δ meaning of the word "complete".
Going the other way we can try:
ak = ∫dt tk f(t) = ∫dt tk Σk'ak'tk' = Σk'ak'{ ∫dt tk tk'}
But since the spanning set functions are not orthonormal, we cannot say { ∫dt tk tk'} = δk,k'and verify our projection.
Comment added 1.22.09: I have misled myself slightly in the above discussion concerning expanding things in powers. The example was given that f(t) = | t | cannot be written as an infinite power series of the form f(t) = Σn=0∞antn because the coefficients keep moving around, whereas you can represent it as a power series f(t) = Σn=0∞anPn(t). The distinction is that the Pn(t) form a basis, while the tn only form a spanning set.
But this does not imply that ALL functions cannot be represented as an infinite power series! For example, we know that f(t) = et = Σn=0∞antn is a fine power series, where an = 1/n! If you start with a finite series for et and do a best fit and then increase N, yes, the coefficients all move around a bit to achieve new least squares fit positions. But, as N→∞, all coefficients will converge to stable values so the infinite series in this case is fine.
In general, if f(t) is analytic in a disk around t = 0 of radius R which contains our real interval of interest, then there will always be a fine infinite power series representation of the form f(t) = Σn=0∞antn valid for |t| < R. For et, R = ∞. The reason f(t) = | t | does not have a power series around t = 0 is that it is not analytic there! Same for the square wave I was messing with today in another document.
So the upshot is this: if you want to represent discontinuous functions as an infinite series, you need to do so with orthogonal basis functions, not some spanning set functions like tn .
Orthogonality (120)
On p 120 an orthogonal set is defined in the obvious way, page 120.
Linear Manifolds ( ≡ Subspaces) (120)
A subspace of any linear space (p 120) which is "closed under addition" and which contains the 0 element is called a linear manifold. [ He does not use the word "subspace", but I think it would be defined to be a linear manifold. ] If a linear manifold is closed in the sense of this book, meaning all sequence convergence points are in the space (space contains its boundary), then the closed linear manifold is a Hilbert space if the larger space is a Hilbert space.
Example: consider the set of solutions of Ly = 0 where L is some linear differential operator. The solutions of this equation, if added, are still solutions, and y = 0 is a solution, so the solution space is a linear manifold, and as noted above, that linear manifold is itself a Hilbert Space. I don't yet know how to find the dimensionality of this manifold. [ this particular linear manifold is the nullspace of L. ]
Notice from compact notes that this kind of manifold is pretty different from a topological manifold which is a space which is locally Euclidean. .
Comment: Stakgold is using the term linear manifold to mean a subspace and it includes the origin of the parent space. The web's "linear manifold" is more general in that the subspace can be translated away from the origin. Thus, Stak would say all lines through the origin are linear manifolds, while web would say that any line is a linear manifold. A subspace of a Hilbert Space is a Hilbert Space.
Projections (121)
He talks on page 121 example 4 about M being a subspace of a Hilbert Space, and M (which I call the perp space) is then also a subspace (linear manifold in his sense). [ Think of these as two perp lines through the origin: both contain the 0 point. ] M and M are called orthogonal complements. He then comments that if A = M M as just described, you can decompose any vector x into its projection into M and its projection into M, and these pieces are unique. Simple example is to think of M as a line through the origin and x some vector off the line. In general, it just seems that you are taking a finite set of axes and putting them into M, and all the other axes into M. But perhaps not all finite sets of axes form a subspace.
The Riesz-Fischer Theorem is proven on page 122. Let k be an orthonormal basis in L2 and consider a partial sum of functions XN = k=1,N akk with coefficients ak, and consider this as the number of terms N . If the partial sum sequence of real numbers SN = k=1,N|ak|2 diverges, then the partial sum of vectors XN also diverges. But if the number sequence converges, then so does XN , and in fact, it will be true that ak = <X,φk>, not surprisingly. [ Note that we are relating the convergence of vectors XN → X to the convergence of real numbers SN→ S. Reals converge vectors converge = RFT.]
This theorem was proven independently in 1907 by Frigyes Riesz (Hungarian) and Ernst Sigismund Fischer (Austrian).
Frigyes Riesz (1880-1956), Hungarian / Budapest, did some of the fundamental work in developing functional analysis and his work has had a number of important applications in physics. His work built on ideas which had been introduced by Fréchet, Lebesgue, Hilbert and others. He also made many contributions to other areas including ergodic theory and he gave an elementary proof of the mean ergodic theorem. Riesz is one of the authors of my book on functional analysis which was written in 1952 when Reisz was 72 years old, and he died 4 years later. This book is now a Dover paperback 1990 for $16. My book was translated from the second French 1953 edition and has a copyright date of 1955.
Ernst Sigismund Fischer (1875 - 1954) was born in Vienna, Austria. He worked alongside both Mertens and Minkowski at the Universities of Vienna and Zurich, respectively. He later became professor at the University of Erlangen, where he worked with Emmy Noether. His main area of research was mathematical analysis, specifically orthonormal sequences of functions which laid groundwork for the emergence of the concept of a Hilbert space.
Doing GSO on an independent set (122)
Page 123 describes Gram-Schmidt, and in particular, shows that doing this to powers gives the non-normalized Legendre polynomials. The GSO method is extremely simple and clear.
1. Pick φ1 = f1/||f1|| so φ1 is normalized
2. Set φ2 = f2 - (f2φ1) φ1 . Then: φ1 φ2 = φ1 [f2 - (φ1 f2) φ1] = 0. Then normalize φ2
3. Set φ3 = f3 - (f3φ1) φ1- (f3φ2) φ2 . Then normalize φ3.
Each one is orthogonal to all that went before because you have subtracted off as shown.
The GSO idea arrived in 1883 from Gram (Danish 1850-1916) and then was later used by Schmidt (German, 1876-1959, student of Hilbert). Well, Laplace and Cauchy really did GSO earlier.
Orthonormal Basis (123)
An odd section. I think a basis in an infinite dimensional HS must be orthogonal, not so in finite.
Best Approximation (Least Squares) (124)
On page 125 we have a result that I have already seen recently. Suppose x is an arbitrary element in the full Hilbert Space. Suppose vector y is a general linear combination of only a subset of orthonormal basis elements i . Then what coefficients for that linear combination give you a y that comes "closest" to x? We want closest in the sense that || x - y || is minimized of course (where || || is some norm, as yet unspecified). The answer is that the coefficients are those that define the unique projection of x onto the linear manifold spanned by your finite set of basis elements. These coefficients are just ai = <x,i>, the Fourier coefficients. Naturally the norm of the approximation y (the projection) will be the norm of x, as for example shown in mid page (Bessel's Inequality).
This "best approximation" is also known as "least squares" for this reason. When the space is L2 , the metric is the "sum" (integral) of | x(t) - y(t) |2 over the set of points of t. In this sum or integral, you are adding up a set of squared differences. If we think of the space of t as discrete, then we are thinking of a least squares fit to some data, perhaps, and this is all in Scheid.
OK, this whole thing is just the concept of doing Fourier Analysis with any set of orthonormal basis functions. The usual example is when those functions are eikx but they could be Legendre polynomials or what have you. Just the usual projection and expansion idea. He makes it sound more mysterious than it is I think. He keeps properly repeating that the previously computed coefficients don't move, and this is why you want to start with a basis, not just some spanning set.
Note: A basis does not have to be orthogonal (in a finite dim space), just needs independence and completeness. He avoids using the term "basis" in the above Fourier discussion, but later shows that you will only get convergence if the set is in fact a basis. That is to say, the Bessel's Inequality will become an equality. [ Well, in an infinite dimensional HS a basis does have to be orthogonal. For a finite dimensional space of dimension N, any set of N independent vectors forms a basis, orthogonal or not. ]
Comment: R1 (the reals) has an infinite number of elements. However, as a vector space, R1 has dimension 1, because you can reach all the reals by stretching a single vector.
___________________________________________________________________________________
More detail on pages 122-126, added 3.16.09
Recall that x = x(t) is some function in L2 and we are pondering || x - Σikaiφi||2. We want to find some ai coefficients that minimize this thing so we have a best fit for a given k. What do we get when we write this thing out?
|| x - Σikaiφi||2 = <x,x> - <x, Σikaiφi> - <Σikaiφi, x> + < Σikaiφi, Σjkajφj>
= ||x||2 - Σiki<x,φi> -Σikai< φi, x> + Σikaij <φi, φj>
where I am using the non-QM form of the inner product to match Stak. Since the φi are orthogonal, we can rewrite as
= ||x||2 - Σiki<x,φi> -Σikai< φi, x> + Σik|ai|2
= ||x||2 - Σiki<φi, x>* -Σikai<φi, x> + Σik|ai|2 // which is p 125 A
To shorten the notation, let's define ci ≡ <φi, x> so the above says
= ||x||2 - Σikii - Σika ci + Σik|ai|2 + Σik|ci|2 - Σik|ci|2
where we add and subtract the last item. Then we have
= ||x||2 - Σik|ci|2 + { Σik|ai|2 - Σikii - Σika ci + + Σik|ci|2 }
= ||x||2 - Σik|ci|2 + Σik { |ai|2 - ii - a ci + + |ci|2 }
= ||x||2 - Σik|ci|2 + Σik { |ai – ci|2 } // which is p 125 B
How would we select ai to minimize this quantity. The answer is clear: ai = ci. So we then have this result:
Theorem: Fourier coefficient ai = <φi, x> causes || x - Σikaiφi||2 to be minimized, and the minimum value is this:
|| x - Σikaiφi||2 = ||x||2 - Σik|ai|2
Corollary: Therefore, we see that ||x||2 ≥ Σik|ai|2 and this is Bessel's Inequality. // page 125 C
Now we switch to page 126 column 2. The Fourier Coefficients are now called ci instead of ai. Due to Bessel's Inequality, we know that the series Sk = Σik|ci|2 must converge to ||x||2.
Theorem: If a series Sk = Σik|ci|2 converges, it must be that ci → 0. (Riemann-Lebesgue Lemma)
Proof: Suppose instead ci→ c, some positive number. Than after some point in the series tail, we know that ci is bracketed by (c-ε,c+ε) so we know that ci ≥ c-ε for the rest of the tail. Then we have
tail = Σi=n∞|ci|2 ≥ Σi=n∞ | c-ε |2 = | c-ε|2 Σi=n∞ 1 = | c-ε |2 ∞
We could set ε = c/10 for example, then we have tail = .81 |c|2 ∞. Thus, if c ≠ 0, the tail of our series diverges and we get a contradiction. So I have proven the Riemann-Lebesgue Lemma.
Riesz-Fisher Theorem : (p 122) If φi is an "infinite orthonormal set" for L2 then
(1) If series Σik|ci|2 diverges, then the series Σik ciφi also diverges.
(2) If series Σik|ci|2 converges, then the series Σik ciφi also converges. Moreover, the series converges to a function g in L2 and the ci are the Fourier coefficients ci = <g,φi>
Proof is on page 122, I skip for now, read it at one time.
___________________________________________________________________________________
Characterization of an Orthonormal Basis (127)
definition: If an orthonormal set which is complete is a basis.
Page 127 then shows several equivalent ways to know if you have a complete orthonormal set n . It all makes complete sense to me. I skip the proofs. On page 128,9 we have a proof of two theorems already quoted. [ Note: the word "complete" here means the orthonormal set forms a "basis". It means among other things that any vector in the space can be written as a lincom of the basis functions, so there are no functions that the basis cannot reach. Counterexample: powers not complete on (-1,1) since | t | cannot be reached in the limit sense -- coeffs never stop moving. I was unable to make Maple try to fit a function with a poly! ]
Theorem 1: An orthonormal set that is a spanning set forms a basis.
Recall that a spanning set is one such that you can get arbitrarily close to any function in L2 , albeit with coefficients that possibly move. With an orthogonal spanning set, the coefficients don't move as you move closer to your function by adding terms.
Corollary: An orthonormal set that is a spanning set just for continuous functions is still a basis.
Theorem 2: (The Projection Theorem) If A = M M, you can decompose any vector x into its projection into M and its projection into M, and these pieces are unique.
Remarks on Orthonormal and Orthogonal Bases in L2. (129)
This section is a set of Examples. On page 130 (Example 3) author casually mentions adding a weight function s(t) and modifying the inner product to include it, and all the results follow through. He does not say it has to be positive definite, however. Example 4 notes that the Legendres form an orthogonal but not orthonormal set.
The section ends with a nice list of complete sets of functions, such as sines, cosines and expos. This then really is Fourier.
Exercises 2.11 through 2.25.
2.5 Functionals. ( pp 135-139 )
I think I care more about transformations (operators) A: H H, but it is best to learn first about the simpler objects functionals T: H C. Many of the theorems are almost identical in the two cases.
Earlier, we defined our most general scalar function f(x) mapping the domain of complex numbers to a range of complex numbers. Suppose we keep the same range idea, but we make the domain be vectors in some subspace S of a Hilbert space A. Then we have mapping T: S C, for example. This mapping is called a functional. Let x be in S, then we talk about T(x).
For a bounded functional, as x wanders over S, we need to have |T(x)| constant * ||x||, and the smallest such constant for a bounded functional is defined to be the norm of T. So bounded means you have a finite norm.
Comment(obs, see following comment): So we have T:S→C as a sort of generalized function where the domain, instead of being numbers like reals, is points in a subset S of some underlying Hilbert Space A. This underlying Hilbert Spaces of course has some norm which we would write as ||x|| = <x,x> which we called the natural norm. In some sense "above" this underlying Hilbert Space, we have another space of some sort which is the space of all these functions T. It is certainly a vector space since we can add T1(x) + T2(x). In fact, if we restrict to functions T which are bounded as described above, then this upper space is in fact a normed vector space as we discussed earlier in this chapter. The norm is this:
|| T || = max over S of | T(x) | / ||x|| // max = supremum (sup) or lub
STOP! I was confused when I wrote this last comment. My underlying or lower space is the domain HS, and my upper space is just the range space which for a functional is C. Obviously T(x) lies in C.
On scratch paper, I showed that this really is a norm. For example, you need to show the triangle rule which says that || T1 + T2 || ≤ || T1 || + || T2 || . But based on the idea that if a ≤ b then max(a) ≤ max(b), this is easy to show because | T1(x) + T2(x) | ≤ | T1(x) | + | T2(x) |. Notice that the definition of thus upper space's norm is dependent on how you select the underlying space's norm, which we assume is the natural norm since we said S is in a Hilbert Space. Stakgold does not prove or even claim that this max thing really is a norm, he just says it is "called the norm".
A functional is continuous at x if, as sequence xn x, we have T[xn] T[x]. If this is true uniformly for all x in some space S, then T is continuous on S. [ This is the "sequence definition" of continuity which I think requires a metric. Of course Hilbert Space always has its natural metric. ]
A functional is linear if two obvious facts are true.
We will really only be interested in S which is a linear manifold (subspace) within A, and the notation for this manifold will be DT. There follow two theorems which I now realize are important:
Theorem 1: Suppose T is defined on linear manifold DT (which is a Hilbert space). Then if T is continuous at x=0, it is continuous at all x in DT, so it is continuous on DT. The proof is trivial from the linearity of T. [ It is also an interesting fact, translation independence of the continuity concept. ]
Theorem 2: For linear functional T, boundedness continuity. Just think || T[x] - T[xn] || being less than || T || * || x - xn||. The existence of a norm (ie, a bound) forces continuity and vice versa. So if xn → x, then we must have T[xn]→ T[x] which IS the meaning of continuity.
Prove this thing right here, both ways. These proofs really apply for T or for operator A.
boundedness continuity
We know that || T[x] - T[xn] || = || T[x – xn] || ≤ || T || * || x - xn|| (linear + T bounded). Thus, as || x - xn|| → 0, we know that || T[x] - T[xn] ||→ 0. In other words, in the domain of T we have x → xn implying that in the range we have T[xn] → T[x], and this is the definition of T being continuous (see just above)
continuity boundedness
I follow the page 136 proof, we do a contrapositive proof. Suppose T is unbounded, then we can find a sequence where || T[xn] || > n ||xn|| which means || T[xn] || / ||xn|| > n. This sequence is just a subset of points in H, but it is a useful subset. Since ||T|| = max(|| T[xn] || / ||xn||) , we see that just looking at the particular points xn we have max(|| T[xn] || / ||xn|| ) > n and this max thing certainly has no finite limit, so this sequence thus exposes the unboundedness of T. Notice that we don't actually write down any particular sequence xn , but we know "we could find such a sequence". We could also find a sequence such that || T[xn] || > n2 ||xn|| if we wanted. But the one we select above is fine. Now construct the following sequence: yn = xn / (n ||xn||). We see that ||yn|| = 1/n and thus yn → 0: it is a null sequence in our domain. Of course if T is continuous, we would expect to have T[yn] → 0. So let's evaluate this thing. We find that T[yn] = T[xn / (n ||xn||)] = T[xn] * 1/ (n ||xn||) [ linearity of T]. Then we have
|| T[yn] || = (1/n) * ( || T[xn] || / ||xn||) . But ( || T[xn] || / ||xn||) > n, so || T[yn] || ≥ (1/n) * n = 1. So we have a situation where yn → 0 and || T[yn] || → 1 which means T is NOT continuous. Thus we have shown that if T is not bounded, then T is not continuous. Thus, if T is continuous, T is bounded.
Notice that the proof in this direction is logically more difficult, it is harder, it is less obvious, it requires constructing a specific sequence to violate the boundedness, and then a corresponding specific null sequence. To show that things are not continuous, a good trick is to construct a null sequence whose image cannot possibly go to 0. That is the big trick.
Riesz Representation Theorem. (p 136) (I remember this from ancient times). Any continuous and linear functional which acts on any S in any Hilbert space A can be "represented" as T[x] = <x,f> where f is some unique element of A. In the proof, author first considers all x with T[x] = 0, and such x form a closed linear manifold that I think is the nullspace N. He does not say this, but I think you can partition the Hilbert space this way: H = N + N where N is a Hilbert space but the perp space is perhaps not. Pick some normalized element f0 in N, and he shows that in fact f = [f0] f0, so this is a proof by construction.
The upshot is that we can make a 1-to-1 association between linear functionals T[x] and points f in the full underlying Hilbert Space A, and the association is this: T[x] = <x,f> with f = [f0] f0 where bar means just complex conjugation. Since f0 is in N, we see that f is also in N, but author does not point this out. [ So somehow our upper level normed linear space of T's seems to be homeomorphic with the underlying Hilbert Space A of x's ? If so, you would sort of think that this upper normed linear space was also a Hilbert Space. ] [ It is, and it is called the dual space! ]
On page 137 we consider now T as a bounded linear functional defined on some DT in Hilbert space A. We know that DT is itself a Hilbert space, so we apply Riesz to say that T[x] = <x,f> where f is in DT. But then we just extend this thing so x can be anywhere in A, defining T[x] to be 0 in DT. So we can then just think of the functional as defined on all of A. This is called an extension of T. If T happens to be bounded on DT , then the extended T is bounded on A. You can also talk in the other direction about an operator being a restriction of another.
Theorem 3. In En, every linear functional is bounded (and continuous). This is proven by showing that a linear functional is continuous. [ The underlying Hilbert Space is En, so we get to use ||x|| as the natural norm in En ]
For infinite dimensional spaces, there are linear functionals which are unbounded = discontinuous. On page 138 we have a sequence of functions which becomes a delta function in the limit n . In the limit, we can see that the result is discontinuous. By direct calculation, we find it is also unbounded. In this case, that is because C = | functional | / || x || = n/ sqrt(2n) ~ sqrt(n) and as n, there is no finite norm C that works, so the functional has no norm.
Dual Basis Idea (p 138). The *'s here are NOT complex conjugation, beware. The claim is that if you start with a basis that is perhaps non-orthonormal, you can still find a set of "dual" (reciprocal) basis vectors ei* which are orthonormal to the non-dual regular basis vectors ei (but not in general to the other ej*). Later on page 150 it is noted that if your ei are themselves orthonormal, then we have ei* = ei so you don't have to think about the dual basis vectors.
[Note added 9.18.08: These things appear in M&M Chap 5. The original non-orthogonal basis is there called ek , while the dual or reciprocal basis is called ek , bolded because these are vectors labeled by k. M&M don't use the term dual but they do use the word reciprocal on page 193. And as claimed here, we had in M&M that ekem = δm,k. These two kinds of basis vectors in an En type space as in M&M are related to the use of contravariant and covariant components for a vector. Nothing is said in either M&M or in Stakgold about the possibility that the em might be orthogonal. In fact, ekem = Gkm which is the lower index metric tensor, and ekem = Gkm is the upper index metric tensor, and in general we don't know that either one is diagonal. ]
Stakgold is pretty quiet about why we should care about functionals. In physics, we often have mappings from a Vector Space into real numbers (energy = integral of electric field), and in the calculus of variations, you vary a function f(x) to minimize some scalar quantity like a least time path integral, so this is a mapping from a vector space of functions f(x) to R. I think he just stuck this functional section in for completeness of his book. // Well, you can think of it as a stepping stone between a starting point of O:C→C and an ending point of A:H→H. That is, a functional does T:H → C (where C = complex numbers, H = Hilbert Space of some sort ), and so is a half-way point in our development.
Theorem. (not stated in text and not proven in text, but shown in table on page 139). If in L2 you have a basis ek and if T is a bounded functional, then when you write x = akek, you can change the order of T and summation, so that T[x] = ak T[ek]. You cannot claim this if T is unbounded.
I guess here would be a proof. If T is bounded and linear, then it is continuous and linear, and then Riesz applies, so we can write T[x] = <x,f> = < akek,f> = ak <ek,f> = akT[ek ]. The idea here is an inner product is always continuous (I think).
Comment: When we later deal with integral equations we are going to care a lot about "bounded linear operators" . Stakgold is just warming us up here with "bounded linear functionals".
Bad Example: The QM probability P(x) = |ψ(x)|2 can be written as P: R3 → R and so is an example of a real valued functional. It is bilinear instead of linear, however, so I am not sure our theory above applies to this particular functional. The associating of continuous with bounded only applies to linear functionals. This one seems continuous, but probably there are cases where it is unbounded, such as if |ψ(x)|2 is infinite at some non-zero value of x (but is still integrable).
On my 1.12.08 reading, I did not feel the need to reread all of the book text because my notes above seem pretty good in that they include all the major points. I may find that it is necessary to actually read all of this text section if things later get fuzzy.
2.6 Transformations (= Operators) ( pp 139-146 )
Now we are on more familiar territory. I now see better why he did functionals earlier. They are simpler in the sense that T: H C, whereas now we have our transformation A: H H (I use H instead of script A for Hilbert space). Everything we did for functionals has a parallel here. For example, we can define what it means for a functional A to be linear, and for it to be bounded, and we can then define the norm of A. We can define a domain DA, a range RA and a nullspace NA of x where Ax = 0. Note that we write Ax, not A[x] as in the last section on functionals. We can define the notion of continuity of A.
The norm here is simply this: norm(A) = max( || Ax || / ||x||) for x in DA
Comments: I suppose if the mapping were somehow A:H1 → H2, the norm would use || Ax ||2 / ||x||1.
These three theorems for operators are exact analogs to the functional theorems above:
Theorem 1: Continuity of operator at x = 0 implies continuity over all of DA.
Theorem 2: For a linear operator, boundedness continuity (p140)
Theorem 3. In En, every linear operator is bounded (and continuous)
Extensions (141)
This process is analogous to how it was done for functionals, but now the Riesz thing is not used (it only applies to functionals). We get the same result that the extended operator B or A is defined to match A on the space DA, and then to be zero on DA . In the functional case where we had T[x] = <x,f>, we could appeal to the continuity of the inner product to say x→0 => T[x] → 0. When we attempt to close DA to make it be A so it becomes a Hilbert Space in H, this issue of continuity comes up. If we did not know that x→0 => Ax→ 0, we could not show that sequences like Axn have a unique limit as xn→ 0 and this means you have closure problems. So, we have to assume an operator is continuous in order to be able to properly extend it. Also, Stak does not use it at this point, but if B is an extension of A, you write B A. Since the extended function is 0 on DA, it seems clear that the norm does not change so ||B|| = ||A||.
Null sequences. This is just xn 0 in DA. What can happen to Axn ? Normally goes to 0 (certainly would do so if A were continuous), but sometimes goes to , which would mean A is discontinuous at x = 0. A third case is possible that it goes to some finite f in H, but author says not to worry about that case since we never will see it in our applications to ODEs etc. He says this case would be very ugly to deal with.
Closed Operators. Consider xn 0 and suppose always Axn 0 or . Then operator A is by def closable. [ If it were possible for some xn 0 that Axn f ≠0 , then A would not be closable. ] Suppose now that xn x but for some reason x is not in DA. We close set DA as above to include x and we have Ax = f. We can call the domain and also we have the closed operator . So you start with an operator A that is "closable" as defined above. You then close the domain, and then A acting on that closed domain is called a closed operator.
Say this again: An operator A is "closeable" as long as it is non-pathological with regard to A(xn→0). We might normally expect A(xn→0)→0 since A0 = 0, but maybe A(xn→0)→ ∞ for some reason. If this happens, A is still closeable. It is not closeable if there is a sequence such that A(xn→0)→ f where f is some non-zero finite space element.
If an operator is closeable, you close it by first closing the domain DA to get A which means that the limit x of every sequence xn→x is in this closed domain. Then Ax must equal something, call it f. The concern here is that we want Ax = the same f regardless of what sequence it was that approached x in the domain. This is the purpose of requiring that A be closeable. Then if you have two sequences approaching x, the difference sequence is a null sequence, and A→0 or ∞ for this difference, not some constant thing.
So if A is closable, we can always close it by this method, and the result is .
Next on page 142-3 comes a long discussion on closing an operator A, and a theorem is stated:
Theorem: A closed operator on a closed domain is continuous.
I have no idea yet why I should care about such things. // Well, I am now getting into that later in this chapter, so coming back here to add some notes for page 142 on this subject.
1. A continuous operator on a closed domain is a closed operator.
2. A closed operator on a closed domain is continuous.
So, if the domain is closed, then continuous closed (assumes operator is closeable).
3. A closed operator on a non-closed domain might not be continuous.
4. Remember that continuous bounded for any linear transformation.
Examples of Operators all involving functions on L2.
[ Note: this list of operators is going to be used many more times in this chapter! ]
1. The zero transformation is bounded.
2. A: x(t) t x(t) on (0,1). This operator is bounded and ||A|| = 1.
3. A: x(t) f(t) x(t) on (0,1). Here we find that ||A|| = max of | f(t) | on interval, etc etc. Bounded.
4. A: x(t) x(t+h) on (0,1) [ translation ] This is bounded and ||A|| = 1
5. A: x(t) !Syntax Error, Idt x(t) on (0,1) [integration] . This A is bounded and ||A|| = 2/
6. A: x(t) !Syntax Error, Idt' k(t,t') x(t') on (a,b). [ an integral operator ] k = kernel. Bounded rule given.
7. A: x(t) dx(t)/dt on (0,1) [ differentiation] . This one is unbounded, example given.
More on this last example #7. The domain DA is functions which are differentiable "in some sense". They include continuous functions that are piecewise differentiable. This class is called absolutely continuous functions. [ p 143] The operator A here is linear, range is all of L2. Operator A = d/dt is closed. But this operator is not continuous and is therefore unbounded. [ The class of a.c. functions is the set of all functions which are indefinite integrals of their derivatives. So if f(x) = dx f '(x) , then f is a.c. This class includes piecewise diff functions. ]
Comment: In the last case of differentiation, if you make a sequence of faster winding sines, the derivative size of course increases without limit, so the ratio dx/dt / x is unbounded.
Possible example of interest. Consider the sequence xn(t) = (1/n) sin(n2t). I would argue this is a null sequence since xn(t) → 0. The integral of the square of this function on 0,1 is perhaps 1/(2n2) which → 0. Now what is the image sequence for A = d/dt ? We get fn(t) = n cos(n2t). This provides us then an example where the limit of the image sequence of a null sequence "does not exist". Certainly the norm gets larger ~ n and so a limit if it existed would have infinite norm. This is really the idea of being unbounded. The fact is that the image function in this case is an ever larger amplitude faster moving cosine and I think it is fair to say that this does not approach any reasonable kind of limit.
2.7 Operators in the Hilbert Space En(c) ( pp 146-165 ) Finite Dimension Hilbert Spaces
Comments: We shall have here a treatment of "linear algebra" from a solid Hilbert space viewpoint. The claim is that you had better understand all this before going on to spaces like L2. Think of this section as being Matrix Algebra, or the algebra of finite Hilbert spaces. Many concepts are defined here, such as "singular". When we move later to infinite dimensional Hilbert Spaces (which is where this book really wants to be), we will have to start all over and redefine concepts like "singular". There will however be a reflection in these new definitions of the finite space ideas discussed in this section.
Introduction + The Matrix of a Linear Transformation on En(c) (146)
We start off and show how matrices arise. You can think of a matrix as transforming your coordinates in a fixed basis as you do Ax = y. You can also think of a matrix as changing you from one basis to another basis. So you can rotate a vector, or you can rotate the coordinate system with the vector fixed.
Examples of Linear Transformations (148)
On page 148 we have some examples of operators. The cos rotation in 2D is mentioned, and the shear matrix. The final example on page 149 is the projection operator P. In the matrix world, you have an upper left hand corner identity matrix which keeps the first set of basis functions, then all else is 0. All these things are familiar to me.
The Two Roles of Matrices (149)
This is just the usual active rotation of x into Ax = y with a fixed basis, versus x staying put and being displayed in a new basis B' that is rotated backwards relative to B. All old hat.
Simple Calculations with Matrices (150)
This matrix thing is belabored, and again I think in 1967 (42 years ago), students were less matrix savvy. He is showing for example how you write out C = AB with indices.
On page 150 we relate the aij matrix elements to the inner product and we must take note that there is a complex conjugation bar over the intuitive expression that I am used to not having this bar. The dual-basis vector ei* appears here for the following reason: if your basis ek is not orthonormal, then you cannot (without using the dual basis) "extract" coefficient aki from the sum which shows the action of A on a lincom, namely, Aej = akjek. If your basis is orthonormal, no * is needed on ei. Again, see comment earlier on this reciprocal basis and M&M.
Comment: See earlier QM note that < a | b > = <b,a>. Notice how horrible 2.39 looks in the non QM notation! You need to overbar sitting there in 2.39. Luckily, this book is not going to be doing any QM.
The Inverse of a Linear Transformation on En(c) (151)
Theorem: If Ax = 0 => x = 0, then Ax = f has exactly one solution for x. Proof is pretty clear. The first fact causes Aei to be a new basis. [ this means there is no nullspace ] { in matrix world, no nullspace means full rank and we have detA ≠ 0 and we can make A-1 so x = A-1f and that is our "exactly one solution" }
Definition: If Ax = 0 => x = 0, then A is regular. Otherwise A is singular.
Fact: If detA 0, then we know that A is regular because we know we can construct A-1 in the usual way. So detA = 0 is the same as singular. Also, detA does not change as you change basis.
This is all comfy stuff. Such an A is of full rank, non-singular, invertible, etc etc etc.
The Adjoint A* of a Linear Transformation A on En(c) ( 152)
Definition: For any operator A, there exists an adjoint operator A* such that <Ax,y> = <x,A*y> for any vectors x and y. In matrix terms this means (a*)k = ak-bar. If A = A*, A is self adjoint. In this case you have the fact that ak = ak-bar. Another word for self-adjoint is Hermitian.
Comments: I would state all this in my familiar QM notation:
T[x] = <h | Ax> = <f |x>
We know that <h | Ax> is a linear functional as shown, and Riesz on page 136 told us that we can represent T[x] as <f |x> where f is some element in the space, uniquely determined by A. We then have
<h | Ax> = <f |x> => f is a function of h, so write <f | = < A†h | and then we have
<h | Ax> = < A†h |x>
where A† is the (Hermitian) adjoint operator he will call A*. If A = A† , the operator is Hermitian (which he calls self-adjoint). The Hermitian matrix is called "complex symmetric". All fine.
The Alternative Theorem (153)
Alternative Theorem. Consider Ax = f. The space of f for which Ax = f has a solution is called RA (the range) as we know. And the space of x for which Ax = 0 is called NA (the nullspace). This theorem says: RA = (NA*), where A* is the adjoint operator of A. The range of A is the perp of the nullspace of A*.
We can of course perp again to get NA = (RA*).
Stak proves this theorem in a few lines on page 153. Here is his proof:
(a) Here he shows that, if z in NA* and f in RA , then zf = 0 which says that NA* RA which in turn tells us that NA* RA . The proof that zf = 0 is this: <f,z> = <Ax,z> = <x,A*z> = <x,0> = 0.
(b) Here he shows that, if z in RA, then z in NA* , which means RA NA* . When this is combined with the result of (a), we have shown that NA* = RA which is the desired final result. Here is his proof of (b). Let z in RA and f in RA. Then of course zf = 0. Write f = Ax, so then <z,Ax> = 0 for all x in NA. But we also know that <z,Ax> = 0 for all x in NA (since Ax=0 for such x). Thus <z,Ax> = 0 for all x in H. Thus we have 0 = <z,Ax> = <A*z, x> for all x in H. That implies A*z = 0 so z in NA* .
Therefore we have shown that NA* = RA so dim(NA*) = dim(RA) = n - dim(RA). Therefore,
nullity(A*) = n - rank(A) which is usually written ν(A*) = n - ρ(A) or ν(A*) + ρ(A) = n. Now on page 153-4 Stak proves the fact that NA and NA* have the same dimensions, and so do RA and RA* In other words, he proves that ν(A) = ν(A*) and ρ(A) = ρ(A*). Therefore we know that ν(A) + ρ(A) = n. The web calls this last thing the Rank-Nullity Theorem.
Now back to the (a) (b) theorem proved above. If z in NA* and f in RA, (a) showed that zf = 0. This means that we must have zif = 0 for i = 1,2...N where N = ν(A*). These conditions zif = 0 that f must satisfy are the N consistency conditions mentioned below. In other words, f has to be perp to a certain set of N vectors, otherwise Ax=f has no solution x whatsoever!
If it happens that ν(A) = 0 so no nullspace, then RA = H and for every f, we have a solution Ax=f. Of course if ν(A) = 0 then ρ(A) = n so matrix A is invertible and x = A-1f. Since this is a prescription for the solution x, that solution must be unique. This is the first "alternative" of the alternative theorem. The second "alternative" is that when ν(A) = N > 0, Ax=f has solutions only for f that meet the N consistency conditions zif = 0.
Comment added: This presentation does not explain why the word "alternative" is used! Let's look at a web description of this theorem:
One case is the FREDHOLM ALTERNATIVE in FUNCTIONAL ANALYSIS, named after Swedish mathematician and physicist Erik Ivar Fredholm (1866–1927). For example, if A is a continuous linear operator with closed range, having adjoint A* (for instance, a matrix and its transpose) then:
" either (the system Ax = b is solvable for any b) or (the equation A* y = 0 has a non-zero solution (ie, A* has a non-trivial nullspace)). Analogous results are fundamental in the study of integral equations."
Suppose we had NA* = φ (null set). Then we know that RA = (NA*) = φ = H, the whole space. Thus, in this case Ax = b is solvable for any b. This is the first alternative. Conversely, if A* y = 0 has a non-zero solution, meaning NA* ≠ φ, then we know RA < H, so there will be some b for which Ax = b has no solution x, and this is the second alternative.
Fredholm is the big integral equations guy, and I guess he did this little theorem along the way. Does not appear in wiki under the alternative name anywhere.
Claim: A and A* are either both regular, or they are both singular.
Definition: the rank of A, called ρ(A), is the dimension of RA . That is, rank(A) = dim(RA).
Definition: the nullity of A, called ν(A), is the dimension of NA.
Theorem : The claim is then made that A and A* both have the same rank, and the same nullity.
[ This is not very surprising since the two matrices are just CC transposes ]
Complete Form of the Alternative Theorem. (p 154)
(a) If nullity(A) = 0, then nullity(A*)=0 and Ax=f has one solution (namely, it will be x = A-1f).
(b) If nullity(A) = ν, so A*x=0 has ν indep solutions z1 thru zν, then Ax = f has a solution iff f lies in RA = perp(NA*), which in turn means that <zk,f> must be 0 for k = 1...ν. These are called consistency conditions. If the solutions to Ax=0 are x1...x, then general solution to Ax=f is the particular solution just noted plus any lincom of the xi.
Comment: OK, so here I guess are the same "alternatives" given above. The first case is clear. In the second case, the zi solve A*zi = 0 and span the nullspace of A*. Suppose Ax=f so f lies in RA = (NA*). Then f (NA*) . That means fzi = 0.
Comment: The rank of A is the dimensionality of the range of A. If you have full rank, then every f maps back to some x via f = Ax. Suppose there were two back map x's called x1 and x2. Then we would have Ax1 = f and Ax2 = f and then A(x1-x2) = 0. But if A has full rank, it has no nullspace, so x1 = x2 and the back map is unique. Then in general everything is 1 to 1.
Nullspace Facts from meta notes. The Alternative Theorem says that either Ax=f has a solution for all f (meaning the range is all of H), or the equation Ax = 0 has a non-trivial solution (there exists a nullspace, by which I always mean a non-trivial nullspace beyond the 0 vector). The more precise statement is that the following spaces are equal: RA = (NA*) which is the same as NA = (RA*) . It turns out that A and A* have the same size (dimension) range and the same dimension nullspace, so for example we know that dim(NA) = dim(NA*). The dimension of a nullspace of A is called the nullity of A or ν(A), while the dimension of the range is called the rank of A or ρ(A). So we just claimed that ρ(A*) = ρ(A) and ν(A*) = ν(A). Obviously dim(NA) = N - dim(NA) = N - ν(A), but
dim (NA) = dim (NA*) = dim(RA) = ρ(A)
and therefore ρ(A) + ν(A) = N. In the alternative theorem as stated above, if Ax = f always has a solution, then we are in the case ρ(A)= N and ν(A) = 0. The other alternative is that ρ(A) < N in which case we must have ν(A) > 0.
Eigenvalues and Eigenvectors (154)
One utility of the "eigenvalue problem" is this. If you can solve Axi = ixi [ (A-λi)xi= 0 ] and if you get N independent solutions xi, these {xi} form a basis. You can then solve the problem (A- )y = f for any i by the trivial method shown on page 155, and there is always a solution. This must mean that the operator B = A- has full rank for λ ≠ an eigenvalue of A.
Comment injected: On page 154 S assumes that the xi of the problem Axi = ixi are independent. We know this is not in general the case for any A, but for this discussion he assumes it is true. Then he considers the problem Ay - λy = f. He lets γi be the components of f, he lets ξi be the components of y, and then he shows directly that the solution is: ξi = γi/(λi- λ). This is the "trivial method" and obviously it only works for λ ≠ eigenvalue. This solution of course requires that the {xi} form a basis. At this point, we are at mid page 155. This would seem to suggest that (A-I)-1 exists and your solution must be y = (A-I)-1f. We know that det(A-I) = 0 gives those i so that is why we cannot allow to hit one of those values. (pp 154,155)
Question on bottom half of page 155: How do we know that both A and (A-λ) are diagonal as shown in the xi basis? Well, this is a little tricky. S starts assuming that xi is some arbitrary basis, and he then derives 2.44. We can then see that we could write 2.44 as (A'-λ) y = f where A' is the diagonal matrix shown at page bottom, which then implies that A' must be the diagonal matrix above that one. So the matrix A is somehow equivalent to A' in terms of solving our equation. We of course know that this is equivalence by similarity by matrix X. We know that AX = XA' where A' is diagonal, but A need not be diagonal. However, we can change to a basis xi' = X-1 xi = ei and in that basis, A' = diagonal.
(p 156) First half page contains Iliadic repetition. It just says that the null space of the operator (A-λ) must have some dimension, and this is named the geometric multiplicity. Second half page now sets us up with x = Σ ξiei where ei is some basis. Then the matrix equation (A-λ)x = 0 appears as 2.48. And yes, solution requires det(A-λ) = 0 called secular equation. We are now done with page 156.
Note added: Let B = A-λ1 . Consider Bf = (A-λ1)f. If detB ≠ 0, then det(A-λ1) ≠ 0 and we can take the equation (A-λ1)f = 0 and invert it to find that f = 0. This would say that Af = λf has only solution f = 0, so there are no eigenfunctions f for such a λ. Therefore, you must have detB = 0 in order to have a chance that λ is an eigenvalue. It turns out that detB = 0 can be solved for λ and these are the eigenvalues, and this is the purpose of the secular equation.
(p 157) Suppose a certain eigenvalue 1 occurs as a root of the secular equation k times. There will be some m solutions then to the equation Ax = 1x, or shall we say, (A- 1I)x = 0. This homogeneous equation has a nullspace of some dimension m, and at the moment, we have no idea if m is related to k in any way. The number k is called the algebraic multiplicity of the eigenvalue 1, while m is called the geometric multiplicity, and it is the dimensionality of the nullspace of (A- 1)x = 0.
In passing, Stak says that he has shown that every A has at least one eigenvalue and one eigenvector in the complex EC. In a worst case, the secular equation produces only one eigenvalue, and we know that it has at least one solution. (fine)
Now let's start with the number m as the dimension of our little nullspace for 1. We associate the first m "unit vectors" like 1,0,0,... with this nullspace. In this "basis" (which probably is not the basis you started in when someone handed you the matrix A), the first m column vectors of A are unit vectors times λ1, I agree. This then gives the first matrix appearing on page 157, and then the second one is trivial. He then writes the secular equation for the matrix in this form and gets 0 = ( - 1)m det(C - I). Since it is possible that 1 occurs in this reduced secular equation for the matrix C (which has square dimension N - m), we could have k > m. So at this point, we have shown only that k m (algebraic geometric multiplicity). [ I have a better proof of this fact in my matrix notes. ]
So the upshot of this entire page 157 is that we know that k m (algebraic geometric). There could be more equal roots of the secular equation (k) than there is dimension of the (A-λ) matrix nullspace (m).
Examples (158)
The examples which follow are very very good, as usual in Stakgold. They are worth reading through several times.
Example 1: The case A = 0 and A = 1 are first done. Each has k = n and I think m = n. In both cases "every vector is an eigenvector".
Example 2: This is A = Rz(θ) in 2D for some fixed θ. The eigenvalues are λ = e±iθ , and the eigenvectors are written in terms of some assumed basis vectors called φ1 and φ2 . Both the eigenvector components and the eigenvalues here are complex. So this problem Ax = λx has no solutions if you examine it in the space En(r) where the component of a vector on each axis must be real. Must use En(c)
Example 3: A simple "shear matrix" A = and secular says k = 2 with both eigenvalues 1, but trivial solution for the eigenvectors finds that the only eigenvectors are x = k (1 0) so m = 1. So right here we have an explicit example of k > m, in case we were wondering if that were possible. Notice that this matrix is not symmetric.
Example 4: Here we use A = P, which projects out the M part of M M where M is some subspace of dimension k it is assumed. The eigenvalues must be 1 and 0. Suppose k = 3 and size = n. Here is what P looks like (imagine that the matrix continues out to some dimension n with all zeros )
P =
The secular equation says (1-λ)3(-λ)n-3 = 0. So λ = 1 has k = 3 and λ = 0 has k = n-3. We have matching m = 3 and m = n-3 for these two eigenvalues. The projector is thus just an embedded dimension k unit matrix in the larger empty matrix.
Operators with n Distinct Eigenvalues (159)
Note how S does not capitalize common words like "with" in a heading. I wonder what convention this follows?
( p 159) Now, the section on page 159 is a good induction proof of the following theorem. I don't really like how Stakgold buries the statement of this theorem where it is nearly impossible to find!
Theorem: In the case that all the eigenvalues are different, the eigenvectors are linearly independent.
Proof: This is not instantly obvious, and M&M do not give a proof. Here is a proof given by Stakgold on his page 159. If the vectors are linearly dependent, there must be some i such that i=1N ixi = 0. If you apply A to both sides, you find that this must also then be true: i=1N iixi = 0. If we multiply the first equation by N (the last eigenvalue) and then subtract that from the second equation, we get i=1N-1 (i - N)ixi = 0, and we have cancelled off the Nth term in the sum.
Now comes the induction proof. For N = 1, we know that the one eigenvector is independent (hard not to be!). Now let's assume that the eigenvectors are independent for N-1. Then consider the last equation written in the last paragraph. Since the (i - N) are not 0, and since the N-1 xi are independent, we must conclude that i = 0 for i = 1 to N-1. But then our first equation in the last paragraph says that NxN = 0. Since xN is not null, we conclude that N = 0. Thus, if i=1N ixi = 0, we find that all the i must be 0, and this shows that all N of the xi are linearly independent.
Here is a direct proof for N=2. If dependent, we must have x1 = x2 . Then 1x1 = 2x2. If you solve these two equations for x1 and compare, you get αx2 = (2/λ2)x2 which says λ1 = λ2. Since they are not equal by assumption of our stated theorem, you conclude that = 0. Thus, you could not write x1 as a lincom of the others, so the two eigenvectors must be independent, contrary to assumption. This is the idea of the above proof.
Comment: Imagine that we have found the full set of eigenvalues λi and their alg mults ki from the secular equation. It is possible that the set of eigenvectors associated with a particular λi contains less than ki independent vectors, as in our little shear matrix example 3 above. So then mi < ki for that particular eigenvalue. If all eigenvalues are different, then each one of them has ki = 1 and of course in this case we have mi = 1 since you cannot have m = 0 -- there is always one eigenvector in the nullspace of (A-λ) in the space E(c).
Symmetric Operators (159)
At this point, we specialize to symmetric operators A in En(C) (ie, complex symmetric, or Hermitian). We learn these basic facts (all outlined in my matrix notes now):
(1) <Ax,x> = real for all x in En(C) -- note that x can be complex.
(2) eigenvalues are real
(3) eigenvectors of different eigenvalues are orthogonal
(4) eigenvectors form a basis ( interesting proof uses perp spaces iteratively), so this means there must be a total of N eigenvectors forming this basis.
Comment: So here we have a symmetric A which we can think of as a direct sum of subspaces of dimension ki for each λi (and of course we could have ki = 1 for some or all i). If overall the xi form a basis, we know that Σ mi = N !!! But of course we also know from secular that Σ ki = N. Thus, there can be no subspace in which mi < ki or else the k sum would be less than N. Thus, we know mi = ki in each subspace for a symmetric matrix A.
Comment regarding 4 above: Notice that you could have, in any or every subspace Mi with dimension > 1, a set of mi basis vectors that are not mutually orthogonal. Still, they form a basis for that subspace, and the set of all basis functions for all subspaces of all dimensions (total of N) form a basis. This is an example of a basis that is not an orthogonal basis (but could of course be made by GS to be orthogonal). I think in the infinite dimensional case, at least for functions like tn, each "subspace" has dimension 1 so the all basis functions have to be orthogonal like the Legendres. The important point about symmetric operators is that their eigenfunctions DO form a basis and you could make this an ortho basis. A non-symmetric matrix might not have a full basis!
Spectral Theorem for Symmetric Operators
(a) to each eigenvalue i is associated an eigenmanifold Mi of some dimension mi=ki (alg = geo)
(b) eigenmanifolds are pairwise orthogonal
(c) these eigenmanifolds partition the space, and you can decompose x = x1+ x2 + ... where xi in Mi.
(d) eigenvalues i are determined by A
Stakgold Remark # 1 (all of page 161). At this point, we are reminded of the projection operators Pi associated with each subspace Mi. I am more favorable to these things than I was on earlier readings. Here is the main fact: Let x be some arbitrary vector which we decompose in as in (c) above. The component in Mi called xi must be a linear combination of a set of mi independent spanning vectors in Mi which we might call φi,n [ these span the nullspace of (A-λi)φ=0 ] . But these spanning vectors are all eigenvectors of A with the particular eigenvalue λi. [ they might not be orthogonal ] Thus, xi must be an eigenvector of A. Therefore, for any x in our Hilbert Space (A is symmetric)
Ax = A (Σ xi) = Σ A xi = Σ λi xi = Σ λi Pi xi = Σ λi Pi x = (Σ λi Pi) x
Since this is true for any x, we must have
A = (Σi λi Pi) " the spectral resolution of A"
This is a fascinating result. If A is diagonal, it is pretty obvious looking at the first matrix on p 157 in which case the Pi are the little unit blocks. But this result is true in any basis, not just in the basis in which A and the Pi are diagonal!
Stakgold goes on to define Q(λj) = sum of projectors from 1 to j. Then ΔQ(λj) = Pj. Then you have the above spectral decomposition appearing as
A = Σi λi Pi = Σi λi ΔQ(λi) → ∫dλ λ dQ(λ) // spectral decomposition of an operator
1 = Σi Pi = Σi ΔQ(λi) → ∫dλ dQ(λ) // spectral decomposition of the identity
where the arrows suggest what might happen in an infinite dimensional Hilbert Space where you might have a continuous eigenvalue distribution. That is of course way beyond our current section on En, but we get the idea, and he says that my Nagy book actually does this development. Stakgold says he will do something like this in Chapters 3 and 4. So just a look here at the future.
Summary of page 161:
you can define projection operator Pi such that Pix = xi and A = i iPi. (spectral resolution of A)
Note that PiPj = i,jPi and 1 = iPi so idempotents.
I think of i iPi as the block diagonal form if we are in the diagonal basis.
We are led to two ideas without proof (see my Func Anal book) : 1 = dQ and A = dQ
These are some kind of spectral decompositions.
Stakgold Remark # 2 (p 161-2).
Here he is defining a set of eigenvalue names μi as follows:
λ1 λ1 λ1 λ2 λ2 λ3 λ3 λ3 λ3 λ4 λ5 λ5 .. λq
μ1 μ2 μ3 μ4 μ5 μ6 μ7 μ8 μ9 μ10 μ11 μ12 .. μn
Perhaps we arrange things in decreasing order since eigenvalues are real for symmetric A.
Stakgold Remark # 3 (p 162). Looking at 2.52, suppose all λi were 0. Then you would have A = 0. Thus, a non-trivial symmetric operator A≠ 0 cannot have all eigenvalues being 0. At least one of them must be non zero. A non-zero non-symmetric operator could have all eigenvalues be 0, and an example is given : A =
Extremal principles for Symmetric Operators (162)
New Notes: Four functionals are defined bottom of page 162, and you can really say
||A|| = max I2 over all x I2 = || Ax || / ||x||
||A|| = max I1 over all x such that ||x|| = 1 I1 = || Ax ||
so these things are really the same. The other two objects are not directly related to ||A|| as far as I know, so I will make up some notation (he gives these things no name)
|||A||| = max I4 over all x I4 = |<Ax, x>| / ||x||2
|||A||| = max I3 over all x such that ||x|| = 1 I3 = |<Ax, x>|
and these two "things" are the same. Some kind of extremum things. Page 153 very clearly proves this fact:
Theorem: for a symmetric operator, ||A|| = |||A||| = | the largest eigenvalue| = |μ1|
You can see this pretty easily. In I2 set x = xi and you get I2 = |λi|, so if you max this, you get |μ1|. Similarly, you get I4 = |λi| and maxing it is the same thing. So for a symmetric operator, both these objects ||A|| and |||A||| are "the norm of A".
The proof here relies on the fact that, because A is symmetric, you can have a basis φi ! For a non-symmetric matrix, he claims without proof that this is all you can show
||A|| ≥ |μ1| // see notes below about p 45
|||A||| ≥ |μ1|
This would be a situation of course where mi < ki, which we know can happen for a non-symmetric matrix. An example is given bottom of page 163 where the matrix A has two eigenvalues λ1,2 = 0, but due to the equation shown, you can directly compute that ||A|| = 1 and |||A||| =1/2 . We see that these are both larger than the largest eigenvalue which is 0. Although this example works out right, I don't see why the norms should be larger than the largest eigenvalue, it seems an unusual result to me. I don't know how to find a proof of these claims because I don't have a good search handle.
Old Notes: For symmetric operators we get the nice n-basis shown in (2.54) page 162. For a non-sym operator, we don't get a basis for En so (2.54) is not true, even for some n1 n. So all these proofs fail in that case. The main upshot of this section is that || A || = | μl | where μ1 is the largest in magnitude real eigenvalue. This norm is the I2 thing. So now we know that the norm of a symmetric matrix is!
Exercises 2.31 through 2.39
At this point comes a set of 9 good exercises which I just read through.
2.33 shows that rank(A) is the number of independent rows of A, while rank(A*) is number of independent columns.
2.35 (d) says: Theorem: <Ax,x> = real for all x in En(C) A is symmetric.
2.36 states basic facts about positive and non-negative operators A (which both must be symmetric because we are saying that <Ax,x> = real when we say <Ax,x> ≥ 0 )
2.39 shows that AA* is non-negative (for any A), hence symmetric. I recall these matrices being in the class called normal matrices. Again, * means Hermitian conjugate in Stakgold! Also, the matrices A*A and AA* have the same eigenvalues. In fact (AA*)* = AA* so is Hermitian obviously.
********** And so ends our long specific study of operators on the nice space En(c) . ***********
2.8 The inverse of an linear operator ( pp 165-180 ) Infinite Dimension Hilbert Spaces
We are now no longer thinking just about En(c) as we were in the last section 2.7. Despite the title of this section, we are going to be "redoing everything" for the tougher case of infinite dimension Hilbert spaces, getting ready to later attack ODE's and integral equations armed with these tools. We are no longer allowed to talk "matrix language" now.
Theorem: A is 1-to-1 A-1 exists [ Ax = 0 x=0, ie, there is no nullspace ]
Author now uses the word "operator" for A. We want to solve Ax = f. (1) What f have solutions? (2) We know we can always add homo solutions to a particular solution. (3) If we do f, we would like to see x. This means we want A-1 to be continuous which is the same as being bounded. (4) How exactly do we find the solution? ( entire remainder of this book)
Aside on continuous therefore bounded. Go back to page 136 on functionals. Bounded means there is a finite norm || A ||. Since you can write || Ax - Axn || || A || || x - xn ||, you see that xxn forces Ax Axn if there is a finite norm, so bounded implies continuous. These two ideas are pretty much the same thing.
Def: regular. Recall that a matrix was either regular or singular, but now things are messier. In dimensions, to be regular we have 3 conditions to meet: (a) Ax=0 x=0, as with matrices; and this implies that A-1 exists; (b) the range must be the entire Hilbert space, so Ax = f has a solution for any f in ; (c) A-1 must be bounded (= continuous). If all three are true, we are will to say "regular". If your RA does not include the boundary of space you can do the "usual" extension idea to get the boundary included. Without the boundary you are essentially regular.
Def: singular. This is what A is if it is neither regular nor essentially regular. ∫
Classification of Singular Operators (167)
There are three cases only (page 167)
(1) A is regular, otherwise it is singular in one of the following three ways:
(2) Ax=0 ⇏ x=0 , in which case A-1 does not exist;
(3) Ok on the previous, but A-1 is unbounded;
(4) The closure of RA does not equal the whole Hilbert space.
So for each of our three conditions on page 166, we have a "way" for an operator to be singular: we can violate any one of the conditions.
Closed, closed, closed!! OK, here are some results from notes above near page 142
1. A continuous (bounded) operator on a closed domain is a closed operator.
2. A closed operator on a closed domain is continuous (bounded).
So, if the domain is closed, then continuous (bounded) closed (assumes operator is closeable).
3. A closed operator on a non-closed domain might not be continuous (bounded).
4. Remember that continuous bounded for any linear transformation.
Some time later I drew this picture which seems consistent with the various theorems. On the left side where we have closed domains. I show the location of differentiation which we know cannot act on full L2 because its domain has to be open, not closed, and we know it is unbounded but closed from simple examples. The arrows point to the closed boundary and the b=c boundary.
And now we are going to add more facts about closed operators:
(a) If A is closed and A-1 exists, then A-1 is also closed.
(b) If A is closed and A-1 is bounded, then RA is closed
Contrapositive: If A is closed and RA is not closed, then A-1 is unbounded.
Apply to A-1: If A-1 is closed and A is bounded, then DA is closed
(c) If A is closed and DA is closed (normal situation I think), then A is bounded.
Apply to A-1: If A-1 is closed and RA is closed, then A-1 is bounded. H
Classifying Closed Operators. These are the only possibilities, one regular and three singular:
(1) A is regular;
(2) A-1 does not exist, nontrivial Ax=0;
(3) A-1 exists and A-1 is unbounded and RA H but = H (ie, range is not closed but dense)
(4) A-1 exists and RA does not close to Hso there are points in Hwhich lie outside our range.
This reader is certainly hard-pressed to understand why we are doing all this detail regarding closed operators, but I can only presume that it will be crucial in our applications later. We are in the "theory section" of the course, and will later do "applications.
Examples of operators on L2 (168)
We last encountered some of these same examples on page 143.
Example 1:
(a) The zero transformation means A = 0 operator. I guess Ox = 0 for all x. Seems to me then that ||0|| = 0 since <0x,0x> = 0, so I guess that means bounded (by 0, a pretty small bound). Bounded means continuous means closed if domain is closed, but domain is entire space so closed. Maps many points into 0, is not 1-1, is not invertible, is singular by item (2) above.
(b) The identify A = 1. In this case ||A|| = 1 exactly, so bounded so continuous so closed. Has no nullspace, is invertible, inverse is bounded. Case (1) "magnificently regular" so, he says.
Example 2: Our friend Ax = t x(t). On page 143 we showed this A is bounded. Thus we know here that it is continuous and thus closed. Has no nullspace, A-1 exists but is unbounded. Many points missing from the range. For example, f = 1 is not in the range because t x(t) = 1 means x(t) = 1/t but this is not in L2 which is the domain. You must have f/t be in the domain for f to be in the range.
Example 3: Ax(t) = x(t+h) translation.
Example 4: A = integration. We know A is bounded, and we know A-1 exists (differentiation), and we know that A-1 is therefore unbounded because we studied it before near page 143. So note that the inverse of a bounded operator can be unbounded.
Examples 5 and 6 concern "shift operators" where the operator takes Σaiφi into Σaiφi+1 . This is example 5, and example 6 takes Σaiφi into Σaiφi+1/i2 . I have no idea why these operators are interesting, so for now I let these examples just sit here unread.
These examples are painful to read, extremely tedious with no seeming payoff for the reader.
Old notes: On pages 168,169 we look at examples of operators A and spaces A and see how each example falls into the above classification. Example 4 is integration which is bounded, and the inverse is differentiation which is unbounded. [ Note: the integral of a continuous function is continuous (bounded), but the derivative of a continuous function could be discontinuous(unbounded). ]
Comment: Probably as we move along later, we will find that you don't really have a theory that can do anything unless operators are "reasonable", and this is where things like closed, bounded, continuous and so forth come into play (and similarly characteristics of the inverse operator if it exists).
The Adjoint Operator (p 170).
This section will discuss how adjoint, self-adjoint and symmetric operators are a bit more subtle in the ∞ dim HS situation than they were in the matrix world En. The game is to figure out what else is needed in this fancier world to make things fly right. Is A bounded? Is DA closed?
Question 1: if A is a bounded operator, why is T[x] = <Ax,y> bounded?
Answer 1: The CSI says <x,y>2 ≤ <x,x><y,y> which here means <Ax,y>2 ≤ <Ax,Ax><y,y>. Then
|| T(x) || = max |<Ax,y>| / ||x|| ≤ max | <Ax,Ax>|/||x|| * || <y,y>|| ≤ || A || || y ||
so if || A || exists, and of course for any fixed y || y || exists, we have a bound for || T(x)||.
Question 2: Mid page 170 claim: If <x,y> = 0 for all x in A = A, then y = 0.
Answer 2: Cheating answer: the idea is that the domain is essentially all of A, you can close it to be A if it is dense. If a vector is orthogonal to every vector in HS A, it has to be 0 because there is no room for a perp space. If we were to write A = DA DA then we might find x DA and y DA (y ≠ 0) such that <x,y> = 0 for all x in DA, which is what a perp space means. But if DA is dense in A, there is no room for a perp space! There cannot be some basis axes in DA since dense DA says you can get as close as you like to any such axis, so such an axis must be in DA. Anyway, I think this is at least the "flavor" of the answer.
We did this on page 152 for En, here we "start all over" . Here we consider two cases separately. We are trying to see what is going to happen with our En result that <Ax,y> = <x, A*y> defines the adjoint operator.
Case I: A is bounded and defined on all of A . If we think about <Ax,y> for fixed y and x in A, we know from Riesz (requires bounded) that this can be "represented" as <x,gy> x for some g in A . If we define A*y gy, we can show that A* is in fact a linear operator so A*y "makes sense". This is the same way things worked in Matrix World page 152. Here || A* || = || A || so A* is bounded. We are allowed to use Riesz because A being bounded means our functional here is continuous and so Riesz is justified. Not so in the next case.
Case II: A is unbounded and defined on DA which is dense in A . We try to do the same thing we did above, and we write <Ax,y> = <x,gy> x, but now we have no Riesz theorem, and it turns out that we can now only make this claim for certain choices of A and y and x. If g exists, we call { x, g } an admissible pair, for obvious reasons. For a reason I don't follow, we need DA to be dense in A. So for certain selected values of x and y we can say <Ax,y> = <x,A*y>. The admissible values of y make up DA* , the domain of our new adjoint linear operator A*. They claim that DA* contains 0. [ See question 2 above for business of DA being dense. ]
Things are pretty complicated now, so Stak summarizes 8 facts on page 171. I have at least read and understood what each one means, but have done no proof attempts.
(1) This is what just proved above: if DA dense in A , then A* exists with some DA* containing 0.
(2) A* is linear
(3) A* is closed regardless of whether A is closed
(4) B A ( operator B is an extension of A) A* B*
(5) if A is bounded on all of A, then so is A*
(6) If (A*)* exists [ meaning DA* is dense in A ] , then A (A*)*
(7) If A is closable, then A* = ()* . [ The overbar means closure of operator A, not c.c. ]
(8) If A is closed, then [ DA dense in A DA* dense in A ]] , so then A** exists.
Why is S spending all this time talking about A* ? Suppose we are trying to solve Ax = f. For what f might we expect to find a solution x ? Well, f must be in RA . The following theorem which is just our new version of what we already saw in En tells us that we can learn about RA by learning about the nullspace of the operator A*, so that is why adjoint A* is so important. We reduce the problem of determining whether a solution to Ax=f exists to a standard nullspace problem of a related operator.
Theorem: (similar to page 153 for En): = (NA*) and (RA ) = NA* .
Corollary: If RA is a closed set, the above says RA = (NA*) which duplicates our matrix version.
He then restates a part of the Alternative Theorem: Ax=f has a solution iff <f,z> = 0 for all z which solve A*z = 0, that is , for all z in the nullspace of A*. He comments on lack of uniqueness as before. More words: If <f,z> = 0 for all z in the nullspace NA*, then f must be in the perp space (NA*). But by our theorem, this is RA, and therefore Ax=f has a solution x. So if you have a candidate f, you need to dot it against all the basis vectors of NA* to see if it in fact allows a solution of Ax = f.
We then jump right to the self-adjoint operator A = A*, and require that both have the same domain and the same value when acting on any x in this domain. If A is self-adjoint, then the Alternative says that Ax=f has a solution if <f,z> = 0 for all z with Az=0 (ie, if f is in the perp space of the nullspace). [ In the En or matrix world, just saying A = A* as a matrix was enough for A to be self-adjoint. In the ∞ dim HS world we have to require that DA = DA* and that Ax = A*x for every x in DA. In En, we had DA = the whole space En so this was not an issue. ]
Now the symmetric operator is not quite the same as the self adjoint one. The symmetric A has to satisfy our friend (x,Ay) = (Ax,y) for all x and y in DA. Symmetric tells us that A* A, not that A* = A, so symmetric⇏self-adjoint, although self-adjointsymmetric. We have: ( self-adjoint symmetric, mainly because you don't know that DA = DA* in the symmetric case, but yes in the self-adjoint case. It is harder for an operator to be self-adjoint than it is for it to just be symmetric. )
Theorem: A is self-adjoint (meaning A = A*) A is symmetric [ since (x,Ay) = (A*x,y) = (Ax,y) ]
Theorem: A* A ( A* is an extension of A) A is symmetric
Theorem: (DA = DA*) + A is symmetric A is self-adjoint
Theorem: If A symmetric on all of A, then of above says A is self-adjoint. (this is the case in En)
Examples ( 8 pages worth!)
I now have two sets of notes on these examples. The second set I made while doing the meta doc, and that set is thus both in the meta doc and here.
OLD SET OF NOTES ON THESE EXAMPLES
At this point on page 172 Stakgold presents 8 brutal pages of examples. This is the same set of classical examples used earlier which I said would appear again. In each case, author wants to answer these questions:
is the operator A self-adjoint and/or symmetric? What is the adjoint operator A* ?
What can be said about the solution to Ax = f?
is A regular or one of the singular cases?
(1) As before, we have "multiply a function x(t) by t" in L2, and this one is self-adjoint. This seems pretty obvious when you write <tx,y> = <x,ty>, a fact I have used many times. Claims that the range is not closed, however, so the fact that there is no nullspace does not tell you every f has a solution Ax = f. If we analyze the particular example here, we can get a condition on f, which if true, means Ax = f does have a solution. This condition just says f/t (which is x) has to be in L2. [ Note that A is not regular since A ≠ H, so #4 singular case. ]
(2) The identity operator is self-adjoint. Example of the above where t = 1. [ regular ]
(3) The integral operator [ 0 to t ] is not symmetric. Show this by just doing it. I would not expect it to be, why would you have < int[x(t)], y(t)> = < x(t), int[y(t)] >. You might try some parts integration to relate these things, but I think the constant terms mess things up. Now what about Ax = f? Here the range is not closed, but is dense in L2. Again, we cannot use the alt theorem, so must explore "directly". The answer is that you have a solution to Ax = f if f is absolutely continuous (defined earlier) and f(0) = 0. Now, since the solution will be x(t) = f '(t), we know that A-1 is unbounded (not continuous), because that is a fact about A-1 = d/dt,. This means that A is singular in Case 3 on page 167.
Confusion on item (3). I agree with both pencil-checked equalities. The adjoint is defined by the equality <Ax,y> = <x,A*y> = !Syntax Error, Idt x(t) {A*y}(t) so we conclude that {A*y}(t) must be!Syntax Error, Ids (s). But we know that {Ay}(t) = !Syntax Error, Ids y(s) so we have A ≠ A* on two counts: complex conjugation, and different integration range. Thus we don't even have the outer oval (1) item, so not symmetric, and certainly therefore not self-adjoint.
(4) Now let A = ( integral operator) - I. You would not think this was much different than case 3, but it is. For any 0, this operator has full range. In fact, we get the explicit solution (2.62) on page 174 which shows that you have a solution for any f in L2. I think the Alt Thm says the same thing in this case, we have just verified it. This means the operator A = ( integral operator) - I is regular for 0. And when = 0 we know from (3) that we are Case 3 singular.
Comments on (4): The homogeneous equation Bx = (A-λI)x = (!Syntax Error, Idt - λ I) x = 0 for λ ≠ 0 says that the integral of the function x(t) is a multiple of x(t) and therefore x(t) must be an exponential x(t) = C et/λ.
Write out the above Bx = 0 as as!Syntax Error, Idt' x(t') = λ x(t). If λ ≠ 0, obvious that x(0) = 0 so must have C = 0, so we conclude that operator B has no nullspace! On page 174 Stak finds an explicit solution for x for this equation Bx = f, for any f in L2 , so he concludes that the range of B is all of L2. Stak says nothing about whether B is self-adjoint (B* = B) or symmetric, it does not look so offhand. Since we know the integral operator itself is non-symmetric and the I operator is symmetric, it seems likely that the sum of these operators would be non-symmetric.
(5) Now comes the horrible Example 5 where A = d/dt. We consider 5 distinct situations.
In case (a), there are no boundary conditions imposed, just that x(t) is differentiable which is what (I) says. In this case we find that A* = -d/dt ( so obviously not self-adjoint) and the - just comes from parts integration. But to make the parts go away, we have to have y(0) = y(1) = 0 in the adjoint world, so this is our "admissible pair" requirement on y. Such y are in DA* .
In case (b) we impose x(0) = 0 at the start. In this case we only need y(1) = 0 in the adjoint world.
In case (c) we impose x(0) = x(1), and this requires y(0) = y(1) in adjoint world.
In case (d) we impose x(0) = x(1) = 0, and this requires nothing extra in adjoint world.
In case (e) we impose x(0) = x'(0) = 0 [ initial conditions] and the conclusion is not clear to me here.
The general remark is that as we add more boundary conditions, we restrict the domain of A, and this causes the domain of A* to get larger.
Comment: My motivation to really study these examples is not high because I don't know if there is a payoff down the road. Right now we are just looking at some examples of A and hobnobbing about them. I can always come back later and upgrade things here, so let's just move along. [ agreed, 1.14.09 ] [ but then on 1.17.09 I felt I had to study these examples more closely. ]
NEW SET OF NOTES ON THESE EXAMPLES
EXAMPLES: This section is very long and should probably only appear in the raw notes, but I wrote these notes here because things were confusing so I leave them here. These examples are required if one is to have even the dimmest understanding how you find A* from A, and how you find RA, uses of the Alternative Theorem, restrictions and extensions, and so on . These are very tedious because each 5 word sentence requires brain work on the reader's part with concepts that have just been loaded into the reader's data banks. The concepts are very foreign from those a physics person usually encounters.
Example 1: [Ax](t) = tx(t) where H = L2(0,1). We find A* = A, self adjoint. <tx,y> = <x,ty> makes things pretty clear, since t is real.
Example 2: A = k1. If k is real, A is self-adjoint. Otherwise, get A* = 1 ≠ A. In this case, we are again saying <kx,y> = <x,ky> if k is real.
Example 3: [Ax](t) = !Syntax Error, Ids x(s) . We find [A*x](t) = !Syntax Error, Ids (s) so clearly A ≠ A* . If Ax=0, x = 0 because only the function 0 has indefinite integral 0! Thus, A has no nullspace. Similarly, A* also has no nullspace. What is RA ? It must be a function f(t) such that f '(t) = x(t). Thus, the range is not all of L2, but is the subset of L2 which is differentiable functions which Stakgold calls absolutely continuous functions. Range is thus not closed but is dense in L2. [ Since RA is not closed, we cannot use the alternative theorem to say RA = (NA*) = all of L2, which is consistent with what we just found. ] This operator is singular as case (3) since the inverse we know is unbounded.
Example 4: [Bx](t) = !Syntax Error, Ids x(s) – λ x(t) = (A - λI)(t), λ ≠ 0. In this case, B still has no nullspace, but now the RB is all of L2. In fact, this operator meets all the requirements for B to be regular: no nullspace so inverse exists; range is all of H; B-1 is bounded. I think this last item he proves explicitly by constructing the solution of Bx = f explicitly.
Example 5: [Ax](t) = dx/dt. This is the Mother of all Examples, filling 5 full pages of the text! This is the one we are going to be VERY interested in for the rest of Stakgold's two volumes! A new idea to me is that we are going to regard boundary conditions like x(0) = 0 as restrictions on the domain, beyond the usual restriction to absolutely continuous functions. He is going to study 7 different boundary conditions, which each get a label like IV in the text! Each different BC along with A = d/dt really describes a different operator A!
(a) If there are no BC's, we have only condition I ( x is absolute continuous), and he calls the operator A1. He first shows that A1 is closed. He then asks what are the admissible pairs {y, g} such that <A1x,y> = <x,g=A1*y>. He shows that in order for y to be admissible, it must satisfy condition II which is that y(0) = y(1) = 0. In this case, the adjoint operator turns out to be A1*y = g = – dy/dt and of course this minus sign arises from a parts integration and the condition II makes the parts go away. Obviously A1 ≠ A1* . So the domain of A1* is conditions I + II. A1 has a nullspace since x(t) = C is in it. But A1* has no nullspace because y(t) = C is ruled out by the boundary conditions. Range of A1 is all of L2 hence closed. The alternative theorem says in this case that RA1 = (NA1*), and since (NA1*) = 0 only, we conclude that the RA1 = all of L2. If we find a solution of Ax = f for some f, we can always add to it a solution of Ax = 0 which in this case is x = C. Thus we get x(t) = !Syntax Error, Ids f(s) + C
(b) Add a domain restriction x(0) = 0, condition III, call the operator A2 . We find A2 is closed, as was A1. In looking for A2*, we again ask what are the admissible pairs. This time due to III, we only need y to satisfy y(1) = 0 (condition IV) to make the parts vanish, and again A2*y = g = – dy/dt . With conditions III and IV, this time we find that the nullspaces of both A2 and A2* are empty, RA2 is again closed, and we again conclude that RA2 = all of L2 by the same argument as in (a) above. In this case, for given f we get the unique solution x(t) = !Syntax Error, Ids f(s) ; there is no C to add since A2 has no nullspace. [ The only possibility for having a nullspace is if x = C or y=C can survive the boundary conditions. ]
(c) This time start again with A1 but add the domain restriction x(0) = x(1) [ condition V], call the operator A3 which yet again is closed. In order to make those parts vanish this time, we need the condition y(0) = y(1) which is the same as condition V, so DA3 = DA3*. Both A3 and A3* have some nullspace since x=C is legal. Given this nullspace, what is (NA*) ? Well, we need <x(t), C> = 0, and that means that !Syntax Error, Idt x(t) = 0. So functions with zero integral are in (NA*). The range of A is closed, so we get to use the alternative theorem without the bar, RA = (NA*), and we thus learn that for A3, the range is functions whose integral is zero! x(t) = !Syntax Error, Ids f(s) + C is the solution here.
(d) Start again with A1 but add the domain restriction x(0) = 0 and x(1) = 0 [ condition VI] and call the operator A4. This condition is really a special case of condition V, but it alone now makes the parts vanish, so DA* has no conditions on it other than condition I and A4* = - d/dt as in all previous cases. Now the equation A4*y = 0 has y= C as its non-trivial solution while NA4 = 0 due to VI. As in example (c), we find (NA*) consists of functions such that !Syntax Error, Idt y(t) = 0 [ the "consistency condition" ] And since NA4 = 0, we don't add a constant, so solution is x(t) = !Syntax Error, Ids f(s).
(e) This time our condition is x(0) = 0 and x'(0) = 0 [ condition VII ] and call it A5 . When Stak writes this sequence A4 A2 A1 he is saying that we have been restricting A1 more and more with our boundary conditions. A1 had none, A2 had x(0) = 0, and A4 had both x(0) = 0 and x(1) = 0. That is, we are restricting the domain more and more. And the domain of the A*'s gets more extended at the same time. Clearly A5 A2 because we added x'(0) = 0. But in this case, it will turn out that instead of the expected A2* A5* , we will get A2* = A5* (he proves this at length). This fact tells us A2* has no nullspace, and we can see that A2 also has no nullspace. He claims that the RA5 is not all of L2 so it must be that this range is not closed ( else alt theorem would say range = all of L2 since nullspace NA2* empty, as usual in many of the above examples). It turns out that the range condition is f(0) = 0.
2.9 The spectrum of an operator ( pp 180-184 )
We are going to work with B = A - I where A is a closed linear operator on some DA dense in A. The range is assumed to depend on so write RA(). For certain values of , B will be regular (these regular values of form the resolvent set of A) and for certain other values of B will be singular ( the spectrum of A). Recall there are three ways that closed B can be singular (p 168), and those three ways correspond to the three pieces of the spectrum of A: Here they are, copying from above:
Classifying Closed Operators. These are the only possibilities, one regular and three singular:
(1) B is regular, meaning B-1 exists, = A , and B-1 is bounded
(2) B-1 does not exist, nontrivial Bx=0;
(3) B-1 exists and B-1 is unbounded and RB but =
(4) B-1 exists and RB does not close to so there are points in which lie outside our range.
Case 1: B is regular, meaning Bx=0 has only trivial and = A and B-1 is bounded. The in this situation comprise the resolvent set, which is not part of the singular spectrum.
Case 2: Consider Bx=0 or Axi = ixi. The eigenvectors xi are non-trivial solutions for the eigenvalues of this equation, i are places where B is singular and has no inverse [ (2) above] , and we have usually seen these values to be isolated points, so i make up the point spectrum of A. [ The eigenfunctions are solutions of Bx = 0, so B has a nullspace, so B-1 cannot exist (not 1 to 1).
Case 3: Consider where Bx = 0 has only trivial solution and B-1 is unbounded and RB A but = A, which is case 3. Such belong to the continuous spectrum of A. Not sure why it is continuous.
Case 4: This leaves only case 4 where Bx=0 has only trivial, but < A so () has dim > 0 and this dimension is called the deficiency of . So deficiency = dim () = some integer. It is something like nullity. When fall into this case, you have the residual spectrum of A. So this case can only occur when the range of B does not fill the whole space A. Note that for matrices En, we only had cases 1 and 2, regular or point spectrum. Nothing is every unbounded (A or A-1 when it exists), so no continuous spectrum. And range is always full, so no residual spectrum.
We are now going to bang out four theorems about the spectrum.
Theorem 1: If is in the residual spectrum of A, then is an eigenvalue of A* which has multiplicity = deficiency of . [ is this geomult or algemult? Since no matrix here, there is no determinant, and I don't think algemult has any meaning. It must therefore be the geometric multiplicity, which would be the number of distinct eigenvectors that A* has for this . ]
Theorem 2: If A is symmetric then eigenvalues of A are real, and <x,Ax> is real for x in DA. (simple proof, really same as in matrix world).
Lemma 1 comments: He writes down a null sequence in the image space (the range) Bxn which maps backwards into a non-null sequence in the domain. If B-1 were bounded, that domain backwards image sequence would have to be a null sequence. So yes, I think this is at least a reasonable possibility given that B-1 is unbounded.
Lemma 2: I have now proven this one. Just write out (using A is symmetric so <Ax,x>= real )
|| Ax - λx||2 = ||Ax||2 + (ξ2 + η2) ||x||2 - 2ξ <Ax,x> = { ||Ax||2 + ξ2||x||2 - 2ξ <Ax,x> } + η2 ||x||2
The problem boils down then to showing that { } ≥ 0. Call Ax = y and we want to show that
||y||2 + ξ2||x||2 - 2ξ <y,x> ≥ 0 or ||y||2 + ξ2||x||2 ≥ 2ξ <y,x>
This is certainly true if <y,x> is negative, so we only have to show it when this is positive. So we have to show in this case that
||y||2 + ξ2||x||2 ≥ 2ξ| <y,x>|
But according to the CSI, the largest | <y,x>| can possibly be is ||x|| ||y||. But even in this case the above inequality is true, as we next show, so it must also be true when | <y,x>| does not take its max value. So want to show that
||y||2 + ξ2||x||2 ≥ 2ξ ||x|| ||y|| but this is true since (||y|| - ξ||x||)2 ≥ 0 QED. here
Theorem 3: The continuous spectrum of a symmetric operator A is limited to the real axis.
The proof is readable and quite peculiar. If η ≠ 0, then for all norm-1 sequences xn we have ||Bxn|| ≥ |η| according to Lemma 2. This means there can be no range null sequence as postulated in Lemma 1 (unless |η| = 0). So if η ≠ 0, then B-1 must be bounded, but in this case λ is not in the continuous spectrum by definition of same! So if λ is in the continuous spectrum, we must have η = 0 meaning λ = real.
Theorem 4: The entire spectrum of a self-adjoint operator lies on the real axis, and there is no residual spectrum, meaning that case 4 never occurs.
And now comes our usual bevy of examples:
Example 1: A = integral operator. Look back at Example 4 on page 168. Ax = 0 has only the trivial solution because only the integral of 0 is 0 for all points t. If 0, the range is the full space as we learned from Example 4 on page 173-4 ("remarkably"). And A is bounded. Thus all 3 conditions for 0 being regular according to page 166 are met. Thus, = 0 is the only possible singular point. This point happens to be in the continuous spectrum. It cannot be an eigenvalue because Ax= 0 has only the trivial solution. Not sure why not in the residual spectrum. [ Added: The range of A is I think absolutely continuous functions which are no doubt dense in L2 . Another way to say this is Alt Thm says (RA ) = NA* and I think the A* nullspace is also empty since a similar integral, so reprisereprise hence no residual spectrum. Since λ = 0 is singular but not an eigenvalue and not in the residual spectrum, it must be in the continuous spectrum! ]
Conclusion: integral operator A has the point = 0 in its continuous spectrum, and that is it! Everything else is in the resolvent set.
Example 2(a). A = d/dt with no BC's. Eigenvalue equation has solution x(t) = C et so every is an eigenvalue so spectrum = point spectrum only = all values of in the complex plane. So the "point spectrum" need not always be just isolated points. Every is in the point spectrum.
Example 2(b). A = d/dt with x(0) = 0. For this problem, since x(t) = C et there are no eigenvalues, so no point spectrum. Using the usual Green's method (remember that the Green's function respects the BC's) , equation (A - I)x = f has a solution for any f, so there can be no residual spectrum (full range). And it turns out that (A - I)-1 is bounded which means no continuous spectrum. [Integral operator is bounded.]
Conclusion: the entire spectrum is empty! That means that every is regular.
Example 2(c), A = d/dt with x(0) = x(1). Point spectrum is = 2in n = 0,1,2,... (or negative). Since different BC, the Green's is now different, but we still get the full range result and (A - I)-1 is bounded, so ONLY this point spectrum. Other are regular.
Example 2(d). A = d/dt with x(0) = 0 = x(1). As in (b), no eigenvalues so no point spectrum. New situation is that now not all f of Ax = f are allowed, so we can have a residual spectrum. The adjoint operator A* allows y = C in it's domain ( see example 5.d), so the range of A is functions of 0 integral. But also the adjoint of B which is B* = A* - I, which has nullspace solutions exp(-t), so that the range of B is then functions which are orthogonal to this expo which eliminates a lot of functions from the range of B! Since B-1 is bounded, there is no continuous spectrum. We don't have B = H I guess, so there is no resolvent set of regular λ values. Turns out all are in residual spectrum.
2.10 Completely Continuous Operators.
Remember that "compact" is the generalization of "finite". There is a shower of statements here that have to each be tested for truthfulness.
definition: A set S of elements in a HS is bounded if ||x|| < c for all x in S.
Set S is finite S is compact (since any sequence must have a convergent subsequence )
Set S is compact S is bounded ( again, if every sequence has a CSS, things must be contained)
I think the idea here is that S is surrounded by a finite boundary in some sense.
Set S is bounded ⇏S is compact. (S not closed)
Bounded ≠> compact. Example is given of a sequence of orthonormal functions. Does such a sequence have a convergent subsequence? The sequence is bounded since all elements have ||φn|| = 1. Suppose there were a convergent subsequence which converged to some function ψ, so ψn→ ψ. He claims then that we would have <ψ,h> = 0 for any h in the HS. Why does he claim this? We could regard cn = <ψn, h> as a Fourier Coefficient of h. Riemann-Lebesgue says we must have cn→ 0. The inner product is continuous, so we therefore get 0 = lim cn = lim <ψn, h> = <lim ψn, h> = <ψ,h> for all h in HS. But this last implies ψ = 0, but ||ψ|| = 1. Thus, there is no convergent subsequence. So here we have a bounded sequence which has no convergent subsequence, so the set of orthonormal functions bounded but not compact. [ Riesz-Nagy point out that ||φn- φm|| = 2 for n ≠ m since orthonormal, so pretty clear that there can be no convergent subsequence here! ]
______________________________________________________________________________
Confusion #1: I now have a problem with a parenthetical remark made on page 185 top which says:
" in En(c), the BWT guarantees that every bounded set is also compact".
" bounded sets are compact in all finite-dimensional Hilbert Spaces"
These statements seem wrong. For example, in E1(c) you could talk about the set of the open unit disk around the origin. This set seems bounded but not closed, so why should it be compact? Stakgold gets support from a Google book:
I agree that every finite SET is compact, but R3 is a set with an finite set of basis vectors but an infinite number of possible vectors, so you could not call R3 a "finite set". So maybe these statements are just assuming that the set in question is closed.
Resolution: Let's go down the list of Stak statements carefully. His definition of compactness is that of sequential compactness: every sequence in a set must have a convergent subsequence, then set compact.
(1) A set with finite number of elements is compact: OK. Obviously any sequence would have a convergent subsequence.
(2) Every compact set is bounded. This is not obvious to me from the above compact definition. I probably would have to show that unbounded implies there is a sequence without a convergent subsequence. If a set is unbounded, I just make a sequence that marches off forever in one of the directions in which it is unbounded. Think for example of the positive integers. Find a convergent subsequence! There is none, it just keeps on going forever without converging. OK, I guess this claim is reasonable.
(3) Every bounded set is compact would be the converse of (2). He says this is true only for finite dimensional spaces. What he really means is that every bounded and closed set is compact in a finite dimensional HS.
(4) He says BWT says every bounded set in En is compact. Again, what he means to say is that every bounded and closed set is compact.
His counterexample with the orthogonal functions φn in L2 gives a sequence which is bounded but which has no convergent subsequence and is therefore not compact. This shows that bounded ≠> compact in general, since he has found an example where bounded => non compact. The space here is L2 which is closed, so in fact we have an example of a set where bounded + closed ≠> compact.
(5) The Google book says bounded sets are compact in R3, and it really means to say that closed bounded sets are compact in R3 . It just assumes the set of interest is a closed set.
_____________________________________________________________________________
Completely Continuous Operators
Theorem: Operator A is bounded A transforms bounded sets into bounded sets
Showing the direction is as follows. A bounded means max ||Ax|| ≤ c1||x||. So if ||x|| ≤ c2 because it is in a bounded set, then ||Ax|| ≤ c1c2 so the range is also a bounded set.
Showing the direction. I think you do this similarly to the way we showed that continuity => bounded for a linear functional, see notes above. The proof is via the contrapositive. Suppose some bounded set S is transformed into an unbounded set R. Then we can find a sequence xn in S which maps into Axn in R where || Axn || > n. Consider another sequence yn in S where yn = (1/n) xn and notice that yn → 0, because we know that xn is bounded, being in S. Then we have || Ayn || = || Axn ||/n > 1. Thus we have yn→ 0 having an image sequence which does not go to 0 since its norm is always > 1. Thus, A is not continuous, and therefore (from earlier theorem) it is not bounded. So, we have shown that if A maps some bounded set into an unbounded one, then A is unbounded. Thus, if A is bounded, it must map any bounded set into a bounded set.
It is understood here that we are always talking about linear operators. So we have another way to say our theorem:
Theorem: Operator A is continuous A transforms bounded sets into bounded sets
Definition: Operator A is completely continuous A transforms bounded sets into compact sets.
Since all bounded sets are compact, completely continuous implies continuous.
Example: The identity operator is not completely continuous. It maps our φn set mentioned above into itself. The φn set was bounded but not compact. So any example of a set being bounded but not compact shows that the identity cannot be "completely continuous".
Example: Consider a bounded operator A whose range is finite dimensional. This maps bounded sets into bounded sets from our theorem above. If we assume that the domain bounded sets are closed sets, then the range sets are also closed as well since A is continuous (being bounded). Then the range sets are closed and bounded, so they are compact since finite-dimensional space. Any A:En → En is therefore completely continuous in this sense.
Example : We will see that Hilbert Schmidt integral operators are completely continuous, which is why we care.
Theorem 1. If A is completely continuous and n is infinite orthonormal sequence, then An 0.
The proof is given. If it were not true, we can create a contradiction. We do know that the image sequence is compact and therefore must have a convergent subsequence, but here we find that the image sequence itself converges and in fact converges to 0. As usual, the idea of Riemann-Lebesgue that the Fourier coefficients → 0 is used.
Implication of theorem 1: certain non-convergent sequences (like φn above) get mapped into convergent ones. But that is not good for the inverse operator! It now maps a convergent sequence into a non-convergent one!
Theorem 2: If A is completely continuous and A-1 exists, then A-1 is unbounded if in dim space.
The proof of this one is easy. We just consider our usual norm one domain sequence φn and we end up with An 0, so regarding A-1 we have boundedness set by || φn || / || Aφn || = 1/|| Aφn || → ∞, so A-1 is unbounded.
Theorem 3: Consider An A. If the An are completely continuous, so is A (some restriction).
This has a long 1/2 page proof which I will skip for now (1.18.09).
Comment: This subject (completely continuous) is mentioned in my Riesz-Nagy book on page 177. Reisz says that he invented the idea in 1917 (reference 9)
2.11 Extremal Properties for Bounded Operators ( pp 187 - 190)
Theorem 1: If A is bounded and DA = H , then || A || = || A* || .
We earlier had a claim that if A is defined on all of H, then A bounded => A* bounded. So in our current theorem, we assume A is bounded, so we know both are bounded. Then we just read the simple proof which uses the CSI and it is proven!
Theorem 2: If A is bounded and DA = H , then the ratio |<Ax,x>|/||x||2 ||A||, for all x 0.
Remember that this is not the definition of the norm -- that involves || Ax ||, see page 140. This theorem 2 is simply the Schwarz inequality. This ratio is the I4 object we dealt with back in the En section, see page 44 or so of these notes.
Corollary: Obviously, if MA is the largest that |<Ax,x>|/||x||2 can go as you vary x, then MA ||A||. Back in the En section, I referred to this MA as |||A||| = max I4 over all x.
Theorem 3: If is in any of the three kinds of spectra for A, then | | the norm || A ||.
Note: We saw way back that for a symmetric operator on En we had ||A|| = |||A||| = | the largest eigenvalue| = |μ1|, so we would conclude there that |λ| ≤ ||A||. Here this result seems to be true for any linear operator on any HS, where λ is in any of the spectrum groups.
Proof: Stak shows this is true for each of the three spectral classes: eigenvalue, residual, and continuous.
Old comment on why reasonable at least for the point spectrum. Expand f = ai i eigenvectors, then <f,Af> = aii <f, i> = ai2i So let 1 be the largest , if we pick a1 as our only non-zero coeff, we get <f,Af> = 1, and this will be the value that sets the norm. Then for all other i we will have i norm. A full proof is given on page 188.
Theorem 4: If A is symmetric, then |(Ax,x)|/||x||2 gets all the way up to ||A||, so we say MA = ||A||. So again I might state this as |||A||| = ||A||. Both these norm measurements are the same for symmetric A.
Proof is a full 1/2 page but seems elementary
Theorem 5: If A is symmetric, there exists a domain sequence xk with ||xk|| = 1 which, in the limit, would become an eigenvector of Ax = λ1x where λ1 = ± ||A||. BUT, this theorem allows that the limit itself might not exist, so theorem is stated as a lim Axk → lim λ1xk.
Theorem 6: If A is symmetric and completely continuous, then the largest eigenvalue is || A ||.
The proof is simple. By adding the completely continuous condition we are then able to show that the limit in the previous theorem exists, and then λ1 really is an eigenvalue and xk → x converges.
Interesting method used in the above proof. We know that the bounded sequence xk has to map into a compact sequence, ie, one which has a convergent subsequence. That convergent subsequence in the range must map back into a subsequence of the xk in the domain which he calls uk . He then shows that in fact this uk sequence converges to some u, and that clinches the proof, and u is the eigenvector.
End of chapter 2, "Introduction to Linear Spaces" , about 100 pages of densepack!
Comments: I can see that Stakgold approached this the way I did in my matrix notes. You want to pretty much just state and prove every theorem you can, and get it down on paper once and for all. He has come up with hundreds of very compact little proofs for things. The problem with this approach is that it is hard to learn the subject. However, he does not claim to be teaching us "linear spaces", he is just summarizing the key facts because we will need these facts when we start doing "boundary value problems". I wonder how many man-hours he spent on this chapter, and how many all researchers spent getting the theorems to this point? Several centuries.
After I use the contents of this chapter, I will be more motivated to come back and ponder things more. Right now, I am just trying to learn WHAT this chapter says about things. I might be doing some integral equations for acoustic scattering, and at that time I will get motivated.
Comment: According to PlanetMath, a completely continuous operator is also called a compact operator, perhaps that is a more common term nowadays:
On the other hand, I think other sites make a distinction between compact and cc. I think they say that all compact operators are cc operators.
Comment 1.18.09. I just concluded another long pass through this long chapter 2, and created a separate document of meta notes in addition to updating the raw notes here. My understanding though not 100% is greatly improved over the previous pass made about 5 years ago. My motivation now in 2009, besides wanting to get this chapter nailed once and for all and in writing, was that I was about to read page 712 of Messiah volume 2 concerning a formal proof of how to handle degenerate perturbation theory. But Messiah started talking in the footnote there about norms of linear operators and a certain resolvent function, and I thought this would be a good excuse to make another Stakgold push. Stak did talk a lot about norms, but the resolvent function idea was not in this chapter, at least by that name.
From wiki http://en.wikipedia.org/wiki/Fredholm_theory:
So in our notation, we usually write B = (A - λI), We might then consider an equation of the form
Bx = f Ax - λx = f f = -λx + Ax // integral equation form
Lu - λu = f // differential equation form
Our formal solution then is
x = B-1 f = (A - λI)-1f
and so B-1 is called the resolvent. I suppose we have to regard this inverse operator as a power series of some sort in A, and that is probably the famous series approach one takes to Fredholm approximation. For the differential equation above we use the more familiar L in place of A.
Fact: If we write B(λ) = (A - λI), then we know that B(λ) has zeros at the eigenvalues of A. We know that for any eigenvector ψ, we have Bψ = 0 since ψ is in the nullspace of B. For the point spectra, these things really are zeros, they are isolated points. Perhaps if there are degenerate eigenvalues, we have power of the zero. In any event, if we define the resolvent as G(λ) = B-1(λ), then the G(λ) has poles at the point spectra eigenvalues, and degeneracy probably causes poles of multiple order.